Speech is where automated films most often sound automated. A single voice reading everything, lines that do not match the picture, a narrator who repeats what a character just said: each is a structural fault, not a model fault. The fix is structural too: two layers with distinct jobs, voices chosen deliberately, and text budgeted to the picture.

The two layers

The narrator carries time and place: "That night, the fox reached the river." Characters carry intent: "I am not going back." A line belongs to a character on screen; the narrator never repeats it, and never describes what the picture already shows. Enforce this as a contract in the screenplay data, not as a style hint. The director model gets two separate fields per scene, one narration sentence and one optional spoken line with a speaker, and the schema rejects a narration that contains quotation marks.

Not every scene needs both. Establishing shots usually take narration only; a confrontation may take a line only, with silence around it. A film where every scene has both layers sounds like a documentary about itself.

Casting voices

Give each character a profile: gender, age band (child, young, adult, old) and a tone in words ("warm, unhurried", "sharp, impatient"). Score the provider's voice catalogue against the profile using the voices' own tags, prefer voices that no other character in this film uses, and fall back to the narrator voice if nothing fits. Store the chosen voice with the character so that a second film uses the same voice: a character whose voice changes between films is as jarring as one whose face changes.

The narrator deserves its own casting. It should differ from every character in gender or age band, so the ear separates the layers without effort. A common mistake is to give the narrator the most pleasant voice in the catalogue and then discover it is also the best fit for the protagonist.

Budgeting speech to picture

Synthesised speech runs at roughly nine to ten characters per second in most languages, slower in German and faster in Spanish. A scene's narration must fit its duration with headroom for the line: for an eight-second scene with a two-second line, the narration gets about five seconds, so around 45 characters. Cap lines at a sentence so that no single clip carries a speech. When a story is longer than the target length, the director must cut events, not compress sentences; speeding up speech to fit is the fastest way to make a film unintelligible.

Budget in characters, not words, because the speech model bills and paces by characters, and because word length varies by language far more than character count.

Placement on the timeline

Place each narration clip at the start of its own scene on the timeline, not concatenated into one long narration file. A single file drifts against the cuts as soon as one clip runs a fraction long, and by scene twenty the narrator is describing the previous scene. Per-scene placement also makes single-scene regeneration possible: replace the clip and its narration together.

Lines go where the character speaks, which for a lip-synced clip is the start of the clip. Leave a short gap between narration and line; overlapping the two layers is worse than silence.

Music under speech

Generate music once from the style guide, at the full length, and duck it under speech by a fixed amount rather than mixing by ear. Automated ducking at 8–10 dB is enough for narration; heavier ducking makes the music pump. Fade in over the first scene and out over the last; do not cut music at scene boundaries unless the screenplay marks a hard change of mood.

Graceful degradation

Lip sync depends on one provider, and providers fail: quota, outage, a removed endpoint. When it fails, keep the clip and still voice the line in the audio track. The mouth may not move, but the story is intact and the failed job is refunded. Never let an audio failure discard a finished clip. The same rule applies to music: a film without music is a film; a film that fails because the music model was down is a bug.

Writing for voices

Lines in quotation marks with the speaker adjacent; short sentences; one idea per scene. Emotion tags ("softly", "excited") help models that support them and are ignored harmlessly by those that do not. Avoid numbers and abbreviations in narration; speech models read "3D" and "DXF" inconsistently, and a written-out form is safer. Proper names should appear in the written look so the model has seen them before it has to say them.

Frequently asked

Can the narrator be a character? Yes, in first-person stories. Then the narrator voice is that character's voice, and the character's on-screen lines use the same voice at a slightly different pace. The two-layer rule still applies: narration describes, lines act.

What about languages the voice provider does not cover well? Cast the narrator from the best available voice for the language and keep character lines short. Poor voices are most exposed on long narration.

How is speech verified? Run the synthesised audio through speech-to-text and compare with the script. A mismatch above a few percent means the voice mispronounced a name or the text contained something the model could not read.