
Let the Model Move: Pose Reference Contact Sheets for AI Music Videos
Verbose motion prompts control every beat but grade your taste. A six-panel posing contact sheet built in Nano Banana Pro hands the choreography to Seedance 2.5, and the model often moves harder than anything you would have scripted. Here is the recipe, plus a decision rule for when to write motion and when to hand it over.
AI-assisted draft. Reviewed and edited by the Phosphene team before publication.
A few months ago the running joke was "Codex is incredible." The newer one, from the same people, is "Codex, do your job properly." The model did not get worse. The bar moved: every agent-suggested prompt now reads like someone else's idea of a shot.
So a Japanese creator who runs AI drama and music-video production decided to go manual for one experiment. No elaborate motion script, no agent-generated camera language. Instead he generated the poses first as reference images, fed them into Seedance 2.5, and let the video model invent the movement. The result is a rough, deliberately unedited 45-second music-video prototype. It looks like a sketch, and that is the point. The method matters more than the polish.
Two ways to control a video, and what each one costs
There are two honest ways to direct an AI video model, and they charge different currencies.
The first is to write the motion into the prompt. The camera tilts up, she looks into the lens, the fabric catches the light. You control every beat, the model executes it, and the final shot is exactly as good as your sense of staging. When the prompt fails, there is nobody else to blame. The author puts it bluntly: a prompt that moves the subject exactly as you intended puts your taste on display, in full.
The second way is what this article is about. Build a posing reference sheet, pass it to the model, and stop narrating. You no longer control the exact choreography, but the model gets room to improvise inside the constraints you set, and its improvisation is often more dynamic than a careful script. The author's conclusion after the first test: if the direction you write is only "okay", handing the scene to the AI is a legitimate option, not a loss of face. Same thing happens in AI drama production, he notes. Dividing labor with the model is a creative decision, not an efficiency one.
The contact sheet as art direction
A storyboard says what happens. A posing contact sheet says what is possible: one person, one outfit, six camera placements, six poses, all deliberately different.
The sheet was generated with Nano Banana Pro (the author explicitly notes it is not Nano Banana 2) and planned in Figma Weave. Seedance 2.5 accepts up to 30 reference images per generation, so a six-panel sheet is a small fraction of the budget. The prompt builds the sheet as a fashion-editorial casting board: six equal vertical panels separated by white gutters, numbered 01 to 06 in the corner, the same person in every panel with identical face, hair, skin tone and proportions.
Each panel is a complete camera brief. Reading them is the fastest way to see what "giving the model possibility space" actually means:
- Panel 01 — camera flat on the floor, optical axis straight up, 16mm rectilinear ultra-wide. The subject sits on an invisible plane above the lens with knees pulled in, so her soles become the two largest shapes in the frame. Chin tucked, the underside of the jaw reads from below; torso foreshortened; clear blue sky filling the background.
- Panel 02 — camera above the head, axis tilted about 50 degrees down, 24mm wide, bird's-eye. She leans forward from the waist, one arm reaching up to the lens, palm flat with fingers spread about 30cm from the front element. The hand owns the top third; the body recedes in steep vertical foreshortening; seamless white cyclorama.
- Panel 03 — camera about 20cm off the ground, axis about 30 degrees up, 8mm fisheye. Barrel distortion bends the horizon into a convex arc. She leans over the lens, one leg forward so the near shoe becomes the biggest element at the bottom edge, face placed near the optical center where proportions stay natural, rooftop and skyline curving away at the frame edges.
- Panel 04 — camera on the floor, axis about 60 degrees up, 20mm rectilinear, worm's-eye. Mid-stride over the lens: the front shoe's tread fills the lower half, legs and torso converge through steep three-point perspective, head small at the top, face tilted down to the camera, sky behind.
- Panel 05 — camera at ankle height, axis level, 20mm. Deep crouch with one leg extended: the near shoe sits about 40cm from the lens and owns the lower third, the back shoe stays small and sharp behind it, forearm resting on the raised knee, head upright, eyes into the lens, saturated two-color studio background.
- Panel 06 — camera near the floor, axis about 45 degrees up, 18mm rectilinear, worm's-eye. Seated and leaning back on one hip, near arm thrust into the lens with fingers spread, so extension distortion renders the hand and forearm at roughly three times the scale of the head. The other hand answers with a smaller paired gesture at the bottom of the frame.
That is half the recipe. The other half is the production contract every panel must obey: a full-frame digital still, ISO 100, f/8, with deep depth of field so the foreground limb and the face stay sharp together. Outdoor panels get hard noon sun from high behind the camera, crisp shadow edges, speculars on fabric and shoes. Studio panels get one large beauty dish above the lens axis plus a separate gradient wash on the background. Grading is a single commercial pass across all six: skin texture kept, blacks crushed, highlights pushed.
Then come the guard rails, which is where pose sheets usually die: the face fully visible and correctly proportioned in every panel, both hands with five clearly separated fingers, exactly one person, six camera positions all different, six poses all different.
Why the sheet beats prose for motion
The Seedance prompt guides treat your input as two layers: a spatial layer, what is in the frame, and a temporal layer, what changes over time. A posing contact sheet is an almost pure spatial input. By design, you leave the temporal layer empty and let the model fill it. That is the whole trick, and it is easy to miss: this is not weaker control, it is a different split of who owns what.
Written motion travels through language first. The model has to reconstruct physical facts from words, keep your camera grammar straight, and preserve identity over the whole clip at the same time. Reference images skip the translation step. The model reads the spatial facts directly from pixels: where the camera was, how the body was folded, what the light was doing. It still has freedom over the actual movement, and freedom is what makes a music video feel alive.
There is a second, less obvious reason the author went this way. When you script motion precisely, every hesitation the model adds reads as a mistake, because you have already imagined the perfect version. When the model owns the movement, you become a casting director instead of a puppeteer. You evaluate takes instead of enforcing a single one, and a surprising take is a feature, not a defect.
When to script, when to hand over
A practical decision rule falls out of the experiment:
- Script the motion when the shot has a required beat: a specific gesture, a message moment, a brand-safe movement. Accept that your taste is on the line.
- Hand over the motion when the job is mood, energy, or a music video where dynamism is the product. Build the sheet, let the model move, generate several takes, and ship the strongest one.
- Split the difference on hard scenes: script the broad scene, use a sheet for the critical moment, and if a segment misses, edit it with timestamp targeting instead of re-rolling the whole clip.
The author's own summary for the failed-take case is worth quoting in spirit: when the direction I specified is just okay, I cut my losses and let the AI try. Then I judge the result on its own terms.
The job is casting, not cinematography
Once you stop writing motion and start writing possibility, the job changes shape. You are choosing what can happen, then letting the generator be the performer. If the result is merely fine, that is information: your pose selection was mid, not the model's obedience. Adjust the sheet, generate again.
That division of labor is the same shape as reference-driven creation tools like Phosphene's template-first workflow, where you set the references and style controls and the generator resolves the details. The skill to build is choosing what to control, and the pose sheet is a clean way to practice it: decide the constraints, then genuinely let go of the rest.