AI character voices
Updated September 20, 2026

Most AI video workflows treat audio as a second job: generate silent footage, synthesize lines, then fight the lip sync. The video models here generate audio and picture in one pass, which removes the fight but changes what a “voice setting” is. It is not a text-to-speech engine you feed a script into — it is a reference the generation pass is conditioned on.
| Written description | Age, timbre, weight, delivery — generated from the script along with everything else about the character, and performed by the video model. |
|---|---|
| Matched preset | A voice picked from a library of system voices, each with a short audio sample you can listen to. One model call proposes a match per character by gender, language and age; you can re-match or pick another by hand. |
| Your own sample | Upload audio for a character and it becomes that character's reference for every segment they speak in. |
Whichever is attached is used consistently: the same character speaks with the same voice across every segment and every episode, for the same reason their face stays the same — the reference travels with them.
There is no tool here for cloning a voice from a recording of someone else. It is the deliberate kind of gap, not a roadmap item: publishing a synthetic performance of a real person's voice is a likeness problem in most markets we serve, and a product that makes it one click is a product that makes it routine. Our own compliance checklist treats likeness — face and voice alike — as something you must have the right to use before you publish.
Two things that follow from it, worth saying plainly:
The dialogue language is chosen in Step 1 and applies to the whole project: the lines are written in that language at the analysis stage, copied verbatim into the storyboard, and performed in it. This is not a dubbing pass over an English original — there is no English original. The cast, names and settings can move to the matching market at the same time through the target-region setting. Fifteen languages are supported, and the case against dubbing goes into why generating the line beats re-timing it.
Voice matching is a free model call — it does not consume credits and does not count against your concurrent-task limit. Uploading your own sample is free. The audio itself is part of video generation, which starts at 10 credits per second of output (≈ $0.06/s at the lowest credit price); there is no separate per-word or per-character charge for speech. See the pricing page, or the generator overview for where voice sits in the run.
The dialogue language is set once in Step 1 and the whole run is performed in it, with the target-region setting moving names, faces and streets to that market at the same time. Each sample below was produced that way, one per language.
Languages
Every series below was produced by the same pipeline with the cast speaking that language natively: the interface, the working documents and the spoken lines are all in one language, and nothing is dubbed. Open one to hear it.
A Carta de LisboaPortuguêsNative dialogue
윈터 프라미스한국어Native dialogue
El Secreto de CostaEspañolNative dialogue
浪人の誓い日本語Native dialogue
Ikrar JakartaBahasa IndonesiaNative dialogue
Зимняя коронаРусскийNative dialogue
L’Héritier du DéfiléFrançaisNative dialogue
Il Tavolo dell'OlivaItalianoNative dialogue
De Vuurtoren van de FjordNederlandsNative dialogue
العهد الصحراويالعربيةNative dialogue
Çantasındaki SözleşmeTürkçeNative dialogue
Die Istanbul-TäuschungDeutschNative dialogue
Останній сигналУкраїнськаNative dialogue
問劍青雲繁體中文Native dialogue
The Last EnvelopeEnglishIn your languageNo. Voices are matched from a library of sampled preset voices, or supplied by you as an upload of your own audio. There is no path for reproducing a voice from a recording of someone else, and that is intentional rather than unfinished.
No. The video models generate picture and audio in the same pass, so the performance and the mouth movement come from one generation. That is also why changing a line means regenerating that segment rather than re-cutting an audio track.
Yes. Upload a sample for a character and it becomes that character's voice reference for every segment they speak in.
The dialogue language is a project-level setting, so yes within one project — but the language is chosen independently of the language you wrote the script in, and independently of your interface language. One script can be produced for several markets by running it with different settings.
Yes, for the same reason the face does: whatever voice is attached to a character is attached at the project level and travels into every segment that character speaks in.
Picture and performance from one pass, in the language your audience speaks.
Try SceneMixer freeCredit packages from $1.49 · Pro $7.99/mo · No credit card to start
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer