Lip sync in animation is the process of making a character’s mouth movements match spoken audio so that the character appears to be saying the line. In AI animation it is done in two broad ways: generating the mouth movement from the audio, or generating the performance first and aligning the voice to the mouth afterwards, much as a dubbing studio does. Both rely on accurate timing of speech sounds.

Good lip sync is almost invisible; bad lip sync is noticed within a fraction of a second, because people are very sensitive to mismatches between what they see and what they hear.

The building blocks: phonemes and visemes

Speech is made of phonemes, the smallest sound units of a language. Many phonemes look the same on the lips: the sounds p, b and m all close the lips; f and v put the upper teeth on the lower lip. Animators group phonemes into visemes, the visible mouth shapes. Traditional hand-drawn animation used charts of roughly a dozen mouth shapes, and many digital rigs still use a set of that size: closed lips, open vowel, rounded vowel, wide vowel, teeth on lip, tongue behind teeth and a rest position.

A lip-sync system therefore needs two pieces of information: which viseme should appear, and when. The “when” is usually the harder part.

Getting the timing: transcription and alignment

Timing comes from analysing the audio. A speech-recognition model transcribes the line and returns timestamps for each word, and sometimes for each phoneme. When the text is already known, a technique called forced alignment matches that text to the audio and returns where each word begins and ends. These timings drive the mouth, the captions and any on-screen text that should appear with the voice.

Technique 1: audio-driven mouth generation

In this approach, the audio comes first and a model generates or modifies the face so that the mouth follows it. Classic rigged animation does this by mapping phoneme timings to visemes on a character rig. Neural approaches go further and repaint the lower face of a video frame by frame to match new audio.

  • Strengths: the voice is never changed, and any line can be applied to any existing footage.
  • Weaknesses: repainting can blur teeth, smear stylised faces or fail on characters that do not look human, such as puppets, animals or strongly stylised cartoons. Each extra pass also adds cost.

Technique 2: performance first, voice aligned

In the second approach, a video model animates the character speaking as part of the scene, so the mouth, the head and the body move together naturally. The line is then produced by a speech model in the character’s own voice and time-aligned to the mouth movement. This is closer to how dubbing works: the picture is fixed, and the audio is fitted to it.

  • Strengths: the whole performance is coherent, stylised characters keep their look, and the voice can be chosen and kept constant for each character.
  • Weaknesses: the audio must be stretched or compressed slightly to fit the speaking window, which has limits before it sounds unnatural.

Yeşilçam Studios uses this second approach in the eight Cinemation styles of its Animation studio that are animated by a video model (Wuxiamation, Celmation, Neonmation, Toymation, Toonmation, Claymation, Postermation and Panelmation) when characters have dialogue. In Shadowmation and Pixelmation, which are drawn by a 2.5D animation engine, lines are still voiced in each character’s fixed voice with captions showing the speaker, but there is no mouth animation. The video model animates the speaking character’s mouth; the line is voiced in that character’s fixed voice and aligned word by word to the mouth movement, so a character sounds the same in every scene. Speech is transcribed to get word timings, and the whole line is stretched to fit rather than cut per word, so endings are not clipped. The reasoning behind that choice is described in why we chose voice consistency over generic lip-sync, and the practical side is in the guide Lip sync and character voices.

Why voice consistency matters

A film with several scenes exposes a problem that single clips hide: if each clip generates its own voice, a character can sound different from scene to scene. Viewers read a changing voice as a different person. Assigning a fixed voice to each character and producing every line with it keeps identity stable, just as a consistent face does. Voice casting is part of character design, not an afterthought.

Common pitfalls

  1. Clipped endings. Cutting audio at word boundaries to force a fit can chop final consonants. Stretching the whole line evenly avoids it.
  2. Over-stretching. A short speaking window and a long line produce unnaturally fast speech. Keep lines short enough for the scene length.
  3. Wrong speaker moving. In a two-shot, the mouth of the listener may move instead of the speaker. Clear scene descriptions and review of the clip help.
  4. Mouth flapping. Opening and closing the mouth on every syllable looks mechanical. Real speech blends shapes and skips minor sounds.
  5. Leading and lagging. Audio slightly ahead of the picture is more noticeable than audio slightly behind it; alignment should err on the side of the picture leading.
  6. Stylised faces. Puppets, paper characters and pixel art need mouth shapes that fit their style, not realistic lips.

Lip sync and captions

The same word timings that drive the mouth can drive captions, so captions appear as each word is heard. In dialogue scenes, captions that include the speaker’s name help viewers who watch without sound, and they make it obvious who is talking even when two characters are on screen. See multi-character dialogue generation for how speakers are attributed.

Frequently asked questions

What is a viseme? A viseme is a visible mouth shape that corresponds to one or more speech sounds. Several phonemes share the same viseme because they look alike on the lips.

Is it better to generate the mouth from the audio or align the audio to the mouth? It depends on the footage. Repainting the mouth suits existing live footage; aligning a fixed character voice to a generated performance keeps stylised characters intact and their voices consistent across scenes.

Why does my character’s last word sometimes sound cut off? This usually happens when audio is trimmed to fit a shorter window. Stretching the whole line slightly, or shortening the line, prevents it.

Do silent or narrated films need lip sync? No. Lip sync matters only when a character on screen speaks. Narration is an off-screen voice and needs only timing with the pictures and captions.