Every video model opens at least two doors: a text prompt, or a starting image plus a prompt. Some add a third: reference images of a character. Choosing the door per scene is a routing decision, not a taste, and a studio that routes well produces better films for less money than one that always uses the same door.

Text-to-video

The model receives a prompt and invents everything: set, light, subject, motion. This is the right door for establishing shots and for scenes with no recurring character: landscapes, crowds, weather, abstract motion, a city at night. It is also the cheapest door in human time, because there is nothing to prepare.

Its weakness is that it remembers nothing. Every call reinvents the world, so two text-to-video scenes of "the fox's cottage" produce two different cottages. Used for character scenes, it is the main cause of inconsistent films. Its second weakness is control: the prompt can ask for a composition but cannot guarantee it, and models routinely ignore spatial instructions such as "on the left".

Image-to-video

The model receives a starting frame and animates it. The first frame is fixed; the model adds motion and, depending on the model, camera movement. This is the right door whenever a specific composition or a specific character must appear, which in a story film is most scenes. It is also how continuity works: a scene that continues the previous one starts from the previous clip's last frame.

Its weakness is that motion is bounded by the frame. Highly dynamic action, a character running toward the camera or a fight, can look constrained because the model is reluctant to leave the composition it was given. Some models also drift: after five or six seconds the character's face starts to change. Shorter clips and a fresh keyframe for the next scene are the fix.

Character reference endpoints

Some providers accept a few reference images and a prompt, and promise to keep the character consistent while inventing the composition. In practice these endpoints favour human faces, struggle with animals and objects, and have been removed or replaced by providers without notice more than once. Treat them as optional: use them when available for photorealistic human characters, never as the only path, and always keep the keyframe route as the fallback. A pipeline that depends on one provider's proprietary endpoint is a pipeline that breaks on that provider's next product decision.

Routing rules

  • The scene has a cast member: compose a keyframe from the reference sheets, then image-to-video.
  • The scene continues the previous one in the same place: last frame of the previous clip, then image-to-video.
  • The person supplied a frame: image-to-video from that frame.
  • The scene is an establishing shot or has no recurring subject: text-to-video.
  • Otherwise: text-to-video, with the style guide prepended to the prompt so the look matches.

A director model can apply these rules itself if the screenplay marks each scene with its cast and its continuity relation to the previous scene. That is why the screenplay schema carries both fields.

Durations

Fast models typically offer 5 or 10 seconds; cinematic ones 4, 6 or 8. The director must choose only from the allowed list for the tier it is writing for, because a scene written for 7 seconds on a model that only makes 5 or 10 will be silently rounded, and the narration written for 7 seconds will not fit. Longer clips cost more and drift more; for drawn styles, 5–8 seconds per scene is the sweet spot, with narration carrying the story across cuts.

Prompts for each door

Text-to-video prompts should describe the shot as a cinematographer would: subject, action, camera, light, style, in that order. Image-to-video prompts should describe only what changes: the motion, the camera move, the expression. Describing the subject again invites the model to redraw it. For both, the style sentence goes first and the negative instructions ("no text, no border") go last.

Frequently asked

Can a film mix the three doors? Yes, and most good films do: text-to-video for the establishing shot, image-to-video for character scenes, continuity frames for sequences. Assembly normalises resolution and frame rate.

Which door is best for anime? Image-to-video from a keyframe in the anime style. Text-to-video on cinematic tiers tends to add photorealistic texture that breaks the look.

What if the video model ignores the keyframe? Lower the motion strength if the model exposes it, shorten the clip, and check that the keyframe has no border or text; models treat framed images as posters and reinterpret them.