Most automated video tools have one voice reading everything. It sounds like an audiobook with pictures. A film has a narrator and it has characters, and they must not say the same thing. This post is about how we designed the sound of a story film: the rule that separates the two layers, how voices are cast, how text is budgeted to the picture, and how the mix is assembled.

The rule

The narrator carries time and place; characters carry intent. The director writes dialogue as spoken lines attached to a character on screen, and narration as short bridges between events. If the story has a line, the narrator does not paraphrase it. If a scene has a line, its narration is short or absent. The rule lives in the screenplay data, not in a style hint: each scene has one narration field and one optional line field with a speaker, and the two are checked separately.

Not every scene needs both. Establishing shots usually take narration only; a confrontation may take a line and silence. A film where every scene has both layers sounds like a documentary about itself.

Casting voices

Each cast member carries a voice profile: gender, age and a tone in words. The studio matches the profile against the voice catalogue by its labels and avoids giving two characters the same voice. The narrator is your choice from the list in the wizard, or an automatic default. It helps if the narrator differs from every character in gender or age, so the ear separates the layers without effort.

Budgeting

Narration is written to fit its scene: about nine characters per second. A 10-second scene gets at most ninety characters, which is one clear sentence. Lines are capped at 120 characters; longer speeches are split across scenes. We budget in characters, not words, because the speech model paces by characters and word length varies by language far more than character count. When a story is longer than the target, the director cuts events, never compresses sentences; speeding up speech to fit is the fastest way to make a film unintelligible.

Mixing

Narration clips are placed at the start of their own scenes, not concatenated into one long file. A single file drifts against the cuts as soon as one clip runs a fraction long, and by scene twenty the narrator is describing the previous scene. Per-scene placement also makes single-scene regeneration possible: replace the clip and its narration together.

Music is generated once from the director's brief and sits at a fixed lower level under everything, dipping under narration and dialogue by a fixed amount rather than mixed by ear. Dialogue is lip-synced onto the clip; when lip-sync is unavailable the line is still voiced in the track so no words are lost, and the failed lip-sync job is refunded.

What to write for the best result

  • Write dialogue in quotation marks with the speaker nearby.
  • Keep narration in the past tense and let it carry time and place.
  • Let scenes carry one idea. A paragraph with two lines becomes two scenes.
  • Avoid numbers and abbreviations in narration; speech models read them inconsistently. Write them out.
  • Introduce each character by name once with one visual fact, so the model has seen the name before it has to say it.

Pitfalls we hit

The most pleasant voice in the catalogue tends to be the best fit for both the narrator and the protagonist, so the studio avoids that collision. Long narration is where a weak voice is most exposed; if a language is covered poorly by the catalogue, keep character lines short and let the narrator do less. And a narration that contains a quotation is almost always a paraphrased line; the director is told not to write one, and it is better not to write it in the story either.

Changing a voice later

Narrator: pick any voice in the wizard before creating the film. A character's line: open the scene card, edit the text or pick another voice, regenerate that scene. Only that scene is charged again.

Frequently asked

Can the narrator be a character? Yes, in first-person stories. The narrator voice is then that character's voice, and the two-layer rule still applies: narration describes, lines act.

Why is there only one spoken line per scene? Because a scene is one clip and one lip-sync. A conversation becomes several short scenes, which also gives the film its rhythm; the director splits a paragraph with two lines into two scenes.

What happens to the words if lip-sync fails? Nothing is lost. The scene keeps its animation, the line is voiced with the character's voice in the audio track, and the failed lip-sync job is refunded to the wallet.