Lip sync in Yeşilçam animation means that when a character speaks, the video model animates that character’s mouth, and the line is then voiced in the character’s own fixed voice and aligned word by word to the mouth movement. Because the voice belongs to the character rather than to the scene, the character sounds the same in every shot. This guide explains how the process works, how to pick voices, and how to write and set up scenes that sync well.
Where lip sync applies
Lip-synced dialogue is part of the eight Cinemation styles animated by a video model: Wuxiamation, Celmation, Neonmation, Toymation, Toonmation, Claymation, Postermation and Panelmation. Shadowmation and Pixelmation are drawn by a 2.5D animation engine instead: their dialogue lines are still voiced in each character’s fixed voice, with captions showing the speaker, but there is no mouth animation, no lip sync and no clip review. It applies when the Speech selector is set to Dialogue or Both, because those are the modes in which characters speak. Narration is voiced too, but a narrator has no mouth on screen, so it is simply mixed over the pictures.
How it works, step by step
- The scene is animated with the speaker talking. The video model generates the clip with the speaking character’s mouth moving.
- The line is voiced in that character’s voice. A speech model reads the line using the voice assigned to the character, the same voice every time that character speaks.
- Word timings are measured. The spoken audio is transcribed to find exactly when each word starts and ends.
- The line is aligned to the mouth. The whole line is time-aligned word by word to the mouth movement in the clip. It is stretched to fit as one piece rather than cut per word, so there are no clipped endings or chopped syllables.
- The film is mixed. Voices, narration and music are combined; music sits under the speech so words stay clear.
Why voice consistency comes first
A common shortcut in AI video is to generate each clip with whatever voice the clip happens to produce. The mouth may match, but the character sounds like a different person in every scene, and viewers notice immediately. Yeşilçam takes the other route: each character has one fixed voice for the whole film, and the timing is adjusted to the picture. The result is a character that feels continuous, which matters far more to a viewer than a perfect lip shape on every syllable.
If a video provider is unavailable during production, the studio switches to a backup provider automatically so the film still completes.
Picking voices
Every character in the Characters panel has a Voice field. Leave it on Voice: auto and the studio chooses a suitable voice, or select one yourself from the list of male and female voices. See adding your own characters for the full setup.
- Contrast voices in a two-hander. If both characters are of the same gender, pick two clearly different voices by hand so the audience can tell them apart with their eyes closed.
- Match the look. A tiny Toymation robot and an old Wuxiamation master call for different registers; choose accordingly.
- Keep the narrator distinct. In Both mode, set a narrator voice in the Narrator field that differs from every character voice.
- Reuse voices for recurring characters. If you make a series of short films with the same cast, note which voice each character uses and select it again next time.
Best practices for clean sync
- Short lines sync best. One or two sentences per line fit a single clip comfortably. A very long line has to be stretched more to fit.
- One speaker per line. Do not write two characters’ words into one line; give each its own line and speaker.
- Put the speaker in the scene. The character who speaks should be the one on screen. The director does this automatically when it writes the dialogue; when you write lines yourself, keep your story text consistent with who is present.
- Use clear reference images. A character whose face is well defined in the reference image gives the video model a clear mouth to animate.
- Choose Premium for the final cut. At Premium quality every clip also gets an automatic review for major defects before the film is assembled.
- Keep captions on for social. Even with good sync, many viewers watch muted; captions show the speaker’s name with the line.
What automatic clip review adds
At Premium quality, each generated clip is checked frame by frame by a reviewing model for major defects, such as a vehicle moving backwards, a detached limb, a disappearing eye or the wrong character on screen. Only major problems trigger a regeneration, and regenerations are capped at 30% of clips. In Panelmation the reviewer also finds where each speaker’s head is, so speech balloons point to the right person. Keyframes that come back with frames, borders or letterboxing are detected and redrawn.
Writing lines with sync in mind
Whether the director writes the dialogue, you write it with I will write the dialogue, or you use Write dialogue with Sam, the same principles help: natural spoken rhythm, short sentences, no stage directions inside the line. Lines you write are used as written, so read them aloud before producing. More in writing dialogue with AI and speech modes and captions, or see the styles on the Animation page.
Frequently asked questions
Does each character keep the same voice across scenes? Yes. Each character has one fixed voice, and every line that character speaks is voiced with it.
Will the ends of my lines be cut off? No. The whole line is stretched to fit the clip as one piece rather than cut per word, which avoids clipped endings.
Which styles have lip-synced dialogue? The eight Cinemation styles animated by a video model, when Speech is set to Dialogue or Both. Shadowmation and Pixelmation voice the lines in fixed voices with captions, but without lip sync.
Can I choose the voice myself? Yes. Pick a male or female voice in the character’s Voice field, or leave it on auto.
