Short answer: For Turkish AI voiceover you first prepare the text for the voice: spell out numbers and dates in words, write abbreviations the way they are read, keep circumflex letters and Turkish characters, and break up long sentences. Then you choose a narrator voice that fits the job, listen to the result without the picture, and fix any word that was read wrong. Most of the quality comes from the text.
Most people have the same experience on their first try: the voice sounds natural, but the phrase "3. kat" (third floor) is read as "üç kat" (three floors), the model stumbles over the digits wherever a price appears, and in "hikaye", written without its circumflex, the k comes out hard. The problem is usually not the voice itself but how the text is handed to it. Turkish text is built for the eye, while a voice model reads exactly what is written. This article gives the rules that close that gap and a checklist you can use before publishing.
How is an AI voiceover made?
An AI voiceover is made in five steps: you rewrite the text for the listener, spell out numbers and abbreviations the way they are read, choose a narrator or character voice, generate the audio and listen to it without the picture, and then fix the misread passage in the text and regenerate only that part. Generating the audio takes seconds; the real effort is in preparing the text.
- Write the text for the ear. Short sentences, one idea per sentence, no nested clauses.
- Spell out the reading. Write digits, dates, times, currencies, abbreviations and symbols in words.
- Choose the voice. Narrator or character; warm, calm or energetic.
- Listen. Turn off the picture and listen once from start to finish; errors are easier to hear without images.
- Fix it piece by piece. Change the faulty word in the text and regenerate only that sentence or scene.
If the voiceover is part of a video, the order changes slightly: the scene plan comes first, the text is split into scenes, and the voice is written to each scene's length. You will find the whole flow in the guide to making videos with AI, and how to plan scenes on paper in the what is a storyboard article.
Where do voice models make the most mistakes with Turkish text?
Voice models stumble most in six places in Turkish: numbers and dates, abbreviations and symbols, circumflex letters, words typed without their Turkish characters, foreign proper names, and question intonation. What they have in common is this: the model cannot always work out the reading a person infers from context, and it often picks the form it has seen most.
Losing Turkish characters is the most insidious. In text copied from another system or typed in a hurry, the letters ı, ğ, ş, ç, ö and ü can drop out. A word typed as "Isik" is read with front vowels rather than like "ışık" (light), and in the spelling "dag" the lengthened sound that ğ gives "dağ" (mountain) is never heard. Running the text through a Turkish spell checker before handing it to the voice catches most of these errors before generation.
How should numbers, dates and abbreviations be written?
Write numbers, dates and abbreviations in words, the way you want the listener to hear them. A voice model may read a dotted date digit by digit, skip the dot of an ordinal number, or spell out an abbreviation such as "A.Ş." (the Turkish equivalent of "Inc.") letter by letter. If digits will appear on screen, keep the digits in the subtitles and spell them out in the voice script. Keeping the two texts separate is the safest approach.
| What the text says | What can happen | Write in the voice script |
|---|---|---|
| 15.04.2026 | The digits are read dot by dot | on beş nisan iki bin yirmi altı (fifteenth of April, two thousand twenty-six) |
| 3. katta | The dot is skipped and it becomes "üç katta" (on three floors) | üçüncü katta (on the third floor) |
| %40 indirim | The symbol is skipped or the order gets mixed up | yüzde kırk indirim (forty percent off) |
| ₺1.499 | The thousands separator is taken for a decimal point | bin dört yüz doksan dokuz lira (one thousand four hundred ninety-nine lira) |
| 2,5 m² | The unit is not read | iki buçuk metrekare (two and a half square meters) |
| saat 14.30 | "on dört nokta otuz" (fourteen point thirty) | saat on dört otuz (fourteen thirty) |
| KDV dahil | The letters are read in a jumble | katma değer vergisi dahil (value added tax included) |
| Demir Yapı A.Ş. | Spelled out as "a ş" | Demir Yapı Anonim Şirketi (Demir Yapı Incorporated) |
| Dr. Ayşe Kaya, vb. | The abbreviation is read letter by letter | Doktor Ayşe Kaya, ve benzeri (Doctor Ayşe Kaya, and so on) |
| TBMM | The model tries to read it as a word | te be me me (the letter names, one by one) |
Some abbreviations are read as words (NATO, TÜBİTAK), others letter by letter (TBMM, the Turkish parliament). Instead of guessing which one the model will choose, write the reading yourself. The same principle applies to currencies: "lira" instead of "TL", "dolar" instead of "USD". Group long strings of digits, such as phone numbers or order codes, in twos and threes with commas between them; the model then finds places to breathe.
How do circumflex letters and foreign names change pronunciation?
The circumflex, the little hat over a vowel, changes pronunciation in Turkish; when it is removed, the model reads a different word. "Hâlâ bekliyoruz" and "hala bekliyoruz" are different sentences: the first means "we are still waiting", while "hala" without the hat means a paternal aunt. "Kâr" is profit, "kar" is the snow that falls. According to the spelling rules of the Turkish Language Association, the mark has three functions, and all three directly affect the sound.
- It distinguishes meaning: hâlâ and hala, kâr and kar, âdet and adet. Here the vowel is read long.
- It softens the consonant before it: kâğıt, dükkân, hikâye, rüzgâr, Elâzığ. Written without the hat, k, g and l can be read hard.
- It marks the relational suffix: "resmî tatil" (an official holiday) and "resmi duvarda" (his picture is on the wall), "millî" (national) and "milli" (having an axle) are not the same word.
Foreign proper names
A foreign person's or place name can be mangled depending on which language's rules the model reads it with. Try it as written first; if it comes out wrong, use a Turkish spelling close to the pronunciation, in the voice script only: "Jak" instead of "Jacques", "Mişel" instead of "Michelle", "Siyatıl" instead of "Seattle". The original spelling stays in the subtitles and on-screen text. For company and product names, follow the owner's own pronunciation.
How do the question particle, commas and sentence length affect intonation?
A voice model reads intonation from punctuation and from how suffixes are written. When the question particle "mi" is written separately, the sentence is built as a question and the stress falls on the word before the particle; in a run-together spelling such as "geliyormusun" (are you coming), the question intonation can be lost. A comma means a short pause, a full stop a long one. Long, nested sentences make the voice flow breathless, flat and monotonous.
- The particle carries the stress: In "Sen mi geldin?" (Was it you who came?) the stressed word is "sen" (you); in "Sen geldin mi?" (Did you come?) it is "geldin" (came). Place the particle according to the word you are asking about.
- A comma is a place to breathe: In "Evet, hazırız" (Yes, we are ready), "evet" gets its own beat. Without the comma, the two words are read in one go.
- The focus sits next to the verb: In Turkish the stressed element usually comes right before the verb. In "Toplantıya Ayşe geldi" (It was Ayşe who came to the meeting), Ayşe is the one that stands out.
- Break up long sentences: Split any sentence you cannot comfortably say in one breath. Chains of "ki" and "olan" clauses that look fine on the page lose the listener when read aloud.
The text also sets the length. As of October 2026, Yeşilçam's director budgets narration at about 9 characters per second; a 10-second scene gets at most 90 characters. Shorten a sentence that does not fit the scene instead of speeding it up: sped-up Turkish turns into a mumble that swallows the suffixes.
Which voice should you choose, and should you use narration or dialogue?
Choose the voice by the type of job: a warm or calm narrator for promos and training, a deep voice for trailers, and in a story film a separate, consistent voice for each character. The narrator should differ from the characters in gender or age so that the ear does not confuse the two layers. Once chosen, a character's voice should not change for the whole film.
The Website Promo flow in Yeşilçam Studio has six narrator voices: warm female, warm male, energetic female, energetic male, calm narrator and deep trailer voice. In the story-to-film pipeline, each character's voice is stored in the character bible together with their face. The rule of two layers, where the narrator carries time and place and the characters carry intent, is explained in detail in the narration or dialogue article, and building a voice cast in the AI narration and voice casting article.
When do you need lip sync?
You need lip sync only if the speaking mouth is visible on screen. In a wide shot, with a character whose back is turned, or in a narrated scene, playing the voice on the audio track is enough and you save the cost of the operation. In a close-up the viewer watches the mouth; there, differences in sync quality are heard and seen immediately. You can compare the tiers in the voiceover and lip sync without a studio article.
Is AI voiceover free, and what does it cost on Yeşilçam?
On Yeşilçam, voiceover does not require a separate subscription; it is a line item in the film or animation estimate, and the amount is shown before generation. The Free plan gives 250,000 tokens a month, but operations are charged with a ×1.3 multiplier. The price of voice-related operations such as lip sync is listed separately in the price table, and the plan multiplier is applied on top of it.
| Operation | List price (tokens) | About, at the Basic rate | Pro (×0.8) | Free (×1.3) |
|---|---|---|---|---|
| Opening the voice library | 1,000 | $0.01 | 800 | 1,300 |
| Lip sync | 29,167 | $0.18 | 23,334 | 37,917 |
| Lip sync (premium voice) | 43,751 | $0.26 | 35,001 | 56,876 |
| Lip sync (pro) | 175,001 | $1.05 | 140,001 | 227,501 |
| Music track | 29,167 | $0.18 | 23,334 | 37,917 |
As of October 2026, the Yeşilçam price list puts standard lip sync at 29,167 tokens and pro lip sync at 175,001 tokens; at the Basic rate these come to about $0.18 and $1.05. The dollar figure is calculated from the Basic plan's monthly allowance: 5 million tokens for $30, so 1 million tokens is about $6.
A small calculation makes the decision easier. In a one-minute film with six dialogue scenes, giving every scene standard sync costs 6 × 29,167 = 175,002 tokens; that is almost exactly the price of a single pro sync. So it makes sense to save the pro tier for the one close-up where the viewer really watches the mouth. The Free plan's 250,000 tokens, with the ×1.3 multiplier, cover six standard syncs (6 × 37,917 = 227,502). Failed generations are refunded automatically; current prices are on the pricing page and the flow itself on the Yeşilçam Studio page.
Is it allowed to clone someone else's voice or imitate a famous voice?
As a general rule, no; you need permission. A person's voice is a personal attribute that identifies them; copying it without permission, making them say words they never said, or using it in commercial work can create legal problems. If you are going to work with any voice other than your own, get permission with a written scope.
- Personality rights: Article 24 of the Turkish Civil Code gives a person whose personality rights are unlawfully infringed the right to seek protection from the courts. An infringement is unlawful unless there is the person's consent, an overriding public or private interest, or authority granted by law.
- Personal data: Article 6 of Türkiye's Personal Data Protection Law No. 6698 lists biometric data among special categories of personal data. Processing a voice to identify or imitate a person may be assessed within this framework; for details, see the publications of the Turkish Personal Data Protection Authority.
- Contract with a voice actor: State clearly in which work, for how long and in which languages the recording will be used, and whether a voice model may be created from the voice.
- Sensitive cases: The voice of a deceased person, a child, a politician or a celebrity needs extra care. Even in parody, do not mislead the audience, and state that the voice was generated with AI.
This section is general information, not legal advice. For a specific job, publication or contract, consult a lawyer.
What does AI voiceover not solve?
AI voiceover carries plain narration, training videos, promo copy and draft dialogue well; but it cannot always capture an actor's interpretation, subtle irony, a tearful voice or the rhythm of a poem. Where the voice catalog is thin, long narration sounds tiring and flat. For work like this, a voice actor is the better route.
- Brand voice and broadcast: For a long-running ad campaign or a television or radio broadcast, a studio recording and a human performer are safer.
- Regional dialects: Voice models speak close to standard Turkish; asking them to imitate a dialect easily turns into caricature.
- Long audiobooks: In narration that runs for hours, monotony builds up and the listener tires.
- Legal and medical content: A pronunciation error can become an error of meaning; a person should listen to every sentence.
There is also a middle way: voice the draft with AI to set up the timing and the edit, then have a performer read the final version. You can also upload your own recording and use it with pro lip sync; the performance is then yours, not synthesized. You will find how the sync tiers apply in animation in the lip sync and character voices guide.
Frequently asked questions
Can AI voiceover be done for free? For trying it out, yes. Yeşilçam's Free plan gives 250,000 tokens a month and charges operations with a ×1.3 multiplier; that is enough to try a short narration and a few standard lip syncs. Files downloaded on the free plan carry a small Yeşilçam Studios logo; subscribers' files are clean.
What is the most common mistake in Turkish voiceover? Handing the text to the voice exactly as written. Digits, dates, abbreviations and symbols are left to the model's guess, and words stripped of their circumflex or Turkish characters are read as different words. Spell these out in the voice script and listen to the result once without the picture.
How many characters should a voiceover script be? Calculate it from the length. Yeşilçam's director budgets narration at about nine characters per second; for 30 seconds of narration, roughly 270 characters of text is enough. Shortening a longer text, rather than speeding it up to fit, keeps it intelligible.
Can I use my own voice? Yes. When you record cleanly in a quiet room and upload it as an asset, pro lip sync uses that recording directly and the performance is yours. Keep the recording to the line, leave a short silence at both ends and avoid room echo.
Do I have to regenerate the whole film for one misread word? No. You open the scene card, fix the word in the text or choose another voice, and regenerate only that scene; you are charged only for that scene. Because narration clips are placed per scene, the fix does not shift the other scenes.

