All articles
Why Everyday Actions Are the Hardest Seedance Prompts (And How to Write Them)

Why Everyday Actions Are the Hardest Seedance Prompts (And How to Write Them)

Explosions and dance shots generate fine. Putting on pants falls apart. Everyday actions are the hardest thing to ask an AI video model for, because the viewer already knows the correct answer. Here is how to write Seedance prompts for them, grounded in the Seedance 2.0 and 2.5 technical reports.

AI-assisted draft. Reviewed and edited by the Phosphene team before publication.

The scene took a minute twenty-two of finished video. More references and more prompt than two fight scenes put together. And what was it depicting? A character getting dressed.

That is the counterintuitive reality of video models. Big motion works. A person flying, fighting, or running usually reads fine, because there is no single correct version of it. But a character pulling on boots, sitting down, or picking something up falls apart, because there is one correct version and everyone watching knows it. Miss a step and the shot is obviously fake.

The Seedance technical reports back this up. Seedance 2.0 splits its physical evaluation into natural phenomena, professional physical phenomena, and everyday physical feedback. On that everyday category, its image-to-video motion quality scored 2.87 out of 5 — the paper's Physical Feedback (Daily) slice of the physics evaluation, not an overall model score; its text-to-video Physical Feedback score, for comparison, is 3.46. The paper itself concedes that motion stability during physics simulation is still hard for every model it compared.

So here is the honest read: "Seedance 2.5 handles complex everyday motion" is wrong. Version 2.5 improved, but ByteDance still lists physical plausibility of complex motion and stability of multiple interacting subjects as open problems. Everyday action is still hard. You just have to prompt it in a way the model can follow.

Why putting on pants is hard, technically

Getting dressed looks trivial because a human does it automatically. To a video model it is a pile of simultaneous problems: a hand touching a garment, the garment touching a leg, a foot entering a boot, the fabric deforming, body parts occluding and reappearing, and left absurdly separate from right. The model has to hold all those contact relationships through time and keep each one plausible.

That is why simulating it well is a real capability, not a granted one.

Write a process sheet, not prose

The Seedance prompt guide frames the model as a multimodal director. It treats your input in two layers: a spatial layer, what is in the frame, and a temporal layer, what changes over time. Given that framing, a good prompt is closer to an engineering instruction than to ad copy.

The official base formula is: precise subject, action details, scene or environment, lighting and color tone, camera movement, visual style, image quality, constraints. Everything has a slot, and the slots are sequenced so the model can plan.

Break the action into observable state changes

If you want a character to put on pants, do not write "puts on pants." The model has to invent the whole sequence and will skip or fuse steps.

Write the states instead, in order: sitting on the floor, gripping the waistband, spreading it open, sliding the left foot in, pushing it up the leg, doing the same on the right, pulling to the waist, and finishing seated. The original article spells out roughly this breakdown for a dressing scene.

The useful habit is exposing the causality: a start state, an approach, a contact, a manipulation, a movement of the body or object, and a completion state. Point at the before and after of each step and the model is no longer guessing.

Plan by motion density, not by seconds

To a human, getting dressed is one action. To the model, it is a long run of contact events. That changes how you budget a clip.

A fifteen-second cap tells you how long the clip is, not how hard it is. Judge difficulty by how many contact-bearing state changes you are asking for. Fewer hand-to-object contact transitions inside one shot means the model has less to keep stable, so it is the reasonable choice for things like dressing, eating, or cooking.

The same logic changes how you use Seedance 2.5's longer output. Do not interpret thirty seconds as permission to stuff a full dressing sequence into one take. Use the long form for story structure, and let each challenging interaction stay small.

Allocate time instead of using speed words

"Slowly put on the jacket" asks the model to guess how much time the action deserves. It has no idea. Give it a window instead: zero to five seconds, put on the jacket.

Timestamps are not fixed instructions. The model is still a generative network, not a scheduler, and a time budget guides its temporal planning rather than guaranteeing exact frames. But a window is far more controllable than a speed adjective, because it removes the ambiguity of how long the motion should take.

A practical checklist for everyday-action prompts

Pull from the original article's rules, and the reasoning holds across Seedance 2.0 and 2.5.

  1. Write a state change, not an action name. "Grip the waistband and pull it to the waist" instead of "get dressed."
  2. Decompose into start, contact, manipulation, completion.
  3. Assign seconds instead of leaving it to "slowly" or "quickly."
  4. Do not start the next step before the previous one finishes.
  5. Simplify the camera during contact-heavy motion.
  6. Serialize left then done, right then done, when left and right must stay distinct.
  7. When several objects share one reference, identify each by two or three static features, like color, shape, or type.
  8. More references do not automatically mean more accuracy. Add them for a reason.
  9. Judge scene difficulty by the number of contact state changes, not by the maximum length.
  10. On 2.0, actively split complex daily actions across clips.
  11. On 2.5, do not cram thirty seconds full. Spend the longer output on story.
  12. On 2.5, edit a failed segment with timestamp targeting rather than regenerating the whole clip.
  13. For a hard daily action, prefer a motion or video reference over prose when one is available.

A note for anime work

Anime gives you more room to move. Shorter clips can be stitched together almost like a picture book, so you can cheat around a hard interaction instead of solving it in one take. But if a shot genuinely has to depict a daily action, the rules are the same as live action. The model is not more forgiving of a skipped contact step just because the character is stylized.

Drop the "daily action is easy" assumption

It is the most natural thing in the world to assume that what a human finds easy, a model finds easy. Video models flip that. The mundane is the brute force task, and the spectacular is the easy win.

Next time you sit down to seed a scene of a character putting on a jacket or tying a shoe, treat it like the hardest shot in the project. Break it into states, budget the time, simplify the camera, and keep the interaction small. That is the framing that turns a plausible-looking fake into a shot that survives a viewer who has put on their own boots a thousand times.

Sources