AI lip sync in Yeşilçam Animation works backwards from the usual approach: the video model animates the speaking character’s mouth first, then the line is voiced in that character’s fixed voice and aligned word by word to the movement. We chose this so that a character sounds the same in every scene. This post explains the trade-off, how the alignment works and what it means for your films.

The problem with generic lip sync

There are, broadly, two ways to make an animated character talk. You can generate the mouth from the audio, which gives tight sync but tends to fight whatever else the picture is doing. Or you can let a video model generate the scene with speech in it, which often looks natural but hands the voice to the model. In our early tests the second approach produced a problem nobody forgives: the same character sounded like a different person from one scene to the next. A child became a young adult between two shots. A grandmother lost her accent.

Audiences tolerate a lot in animation. They accept talking foxes and physics that bends. What they do not accept is a familiar character suddenly speaking with a stranger’s voice. It breaks the film instantly, and no amount of mouth accuracy fixes it.

What we chose instead

We decided that the voice belongs to the character, not to the scene. When you add a character and set a Voice, either one you pick from the male and female voices or Voice: auto, that voice is fixed for the whole project. Every line that character speaks, in every scene, is produced in that voice by a speech model. The picture adapts to the voice, not the other way round.

That meant building an alignment step between the two, which is where most of the engineering went.

How the alignment works

  1. The director assigns each line to a speaker, from your script, from Write dialogue with Sam, or from its own draft.
  2. The video model animates the scene with the speaking character’s mouth moving.
  3. The line is voiced in that character’s fixed voice.
  4. The voiced speech is transcribed to get the timing of every word.
  5. The whole line is stretched to fit the mouth movement, and the words are aligned to it.

The important detail is in step five. We stretch the line as a whole; we do not cut it per word. An earlier version tried to snap individual words to individual mouth shapes, and the result was clipped endings: the last syllable of a sentence would disappear because the mouth had closed a fraction too early. Stretching the full line keeps every word intact and keeps the natural rhythm of the voice, while still landing it where the mouth moves.

Where it applies

This alignment runs in the eight Cinemation styles animated by a video model, namely Wuxiamation, Celmation, Neonmation, Toymation, Toonmation, Claymation, Postermation and Panelmation, when a character has dialogue, that is, with the Speech mode set to Dialogue or Both. Narration is a different track: the narrator is not on screen, so there is no mouth to align to, and the narration is simply mixed under the picture. In Silent there is no voice at all and the music comes up. Shadowmation and Pixelmation are drawn by a 2.5D animation engine rather than a video model: their dialogue is still voiced in each character’s fixed voice, with captions showing the speaker, but there is no mouth animation to align to. The procedural styles do not lip-sync; Typemation and Popbook have no dialogue at all, because their text is the content.

Picking voices that work

  • Contrast the cast. Two characters with similar voices are hard to tell apart when the camera is on one of them. Give a pair a clear difference in pitch or gender.
  • Match the look. A voice that fits the description of the character makes the sync feel better than it is, because nothing surprises the ear.
  • Start on auto. Voice: auto chooses sensibly from the character; switch only when a voice is wrong for the part.
  • Keep lines short. One or two sentences per line align more cleanly than long monologues and are easier to follow with captions on.

Our guide on lip sync and character voices goes further into choosing voices and writing lines that speak well.

What it does not do

We want to be precise about limits. The alignment adapts timing; it does not re-draw a mouth that the video model left closed. In most scenes that is not an issue, but a character seen from behind, or very small in a wide shot, will not show much mouth movement to align to. In Premium quality, every clip is also checked by an automatic reviewer for major defects, such as a wrong character or a disappearing eye, and the worst clips are regenerated; you can read about that in every clip gets a second look. The reviewer also finds where each speaker’s head is, which Panelmation uses to point speech balloons at the right person.

Why this is worth it

Voice consistency is what turns a sequence of clips into a film with characters in it. Once the voices stay put, you can write longer stories, bring characters back in a second episode, and let the audience get to know them. That is the kind of animation we want people to make here. The feature is part of the eight video-model Cinemation styles on the Animation page, and the broader cast and dialogue release is described in characters that talk.

Frequently asked questions

Does each character keep the same voice across scenes? Yes. A character’s voice is fixed for the project, and every line that character speaks is voiced in it.

How is the voice matched to the mouth? The voiced line is transcribed for word timings and the whole line is stretched to fit the mouth movement the video model created.

Why are sentence endings not cut off? Because the line is stretched as a whole rather than cut per word, so every syllable is kept.

Which styles have lip sync? The eight Cinemation styles animated by a video model, when characters have dialogue. Shadowmation and Pixelmation voice the lines in fixed voices without mouth animation, and narration and procedural styles are not lip-synced.