
Open a character in step 2 and the Voice description box is at the bottom of the dialog, under the appearance description. That single box is where an accent lives.
Nothing about the voice is recorded, dubbed or stitched together afterwards. When a segment is generated, the video model renders the picture and says the lines in the same pass, and the description of each speaking character is handed to it as the brief for that voice. That is why the mouth shapes match the words: the voice is part of the generation, not a layer added to it. Dialogue and sound in one generation.
The accent words are not filtered, shortened or dropped: “soft and breathy, a light Taiwanese Mandarin, unhurried” travels with every line that character speaks, in every segment — and nothing asks whether you meant it. If you want Hong Kong-accented Mandarin, you say Hong Kong-accented Mandarin.
The one job the voice library cannot do is accents. It is organized by language — Chinese, English, Cantonese — and a match is chosen by gender, age band and manner. There is no Taiwanese-accented entry, no Scottish entry, no Brooklyn entry, and the accent words in your description play no part in choosing one. Accents come from the description; the library is for holding a voice steady. How character voices work.
Two places show which of the two modes a character is in. On the character card in step 2, a small speaker icon sits next to the name when an audio voice is attached — no icon means the description is in charge. In the storyboard editor, the @ menu has a Voice tab listing the cast: each character has a speaking icon, and one who has no audio also gets a small Text only tag next to it. Hovering the icon says which of the two it is — Voice assigned, or Voice description only (no audio). The icon without a tag is the green light; the speaker means an audio sample is attached and the description, accent included, is not what gets performed.
If you find yourself in the wrong state, open the character and use Delete Voice — the confirmation says the remote file is kept, so nothing is destroyed and you can pick a voice again later. It is free, and it takes effect on the next segment you generate. Then your accent words are back in play.
Say which accent, and say how much of it. “An accent” is not a description; “a light Taiwanese Mandarin” is. These are the same one-line format the description already uses, with the accent appended:
| What you write | What you are asking for |
|---|---|
| Low and warm, unhurried, a light Taiwanese Mandarin | Mandarin with a Taiwanese accent, said gently and without hurry |
| Bright and quick, Hong Kong-accented Mandarin, clipped endings | The Hong Kong flavour on Mandarin, with the clipped rhythm that goes with it |
| Relaxed and a little husky, a Sichuan accent, conversational | A Sichuan-flavoured Mandarin, casual rather than performed |
| Neutral standard Mandarin with only a trace of a Northeastern flavour | Standard Mandarin with the faintest regional colouring |
| Measured and gravelly, a soft Scottish lilt, unhurried | English with a gentle Scottish lilt |
| Warm and slow, a soft Southern US drawl | English with the stretched vowels of the American South |
| Flat and weary, a London accent, clipped | English that sounds like London, tight and tired |
| Light and quick, English with a slight French accent | English spoken by someone whose first language is French |
Three habits make it work.
Three different things, three different places to put them. Mixing them up is the usual reason an accent “does not work”.
| What you want | Where it goes |
|---|---|
| The cast speaks the project’s dialogue language with a regional flavour — Taiwanese, Hong Kong, Sichuan, Northeastern, Shanghai, Beijing-style erhua, a rural English drawl | The character’s Voice description. The lines stay in the project’s dialogue language and wear that region on top |
| The lines themselves are written in that region’s idiom — actual regional expressions rather than standard ones | The dialogue itself, in the script or in the storyboard line. Wording plus accent together is what makes a dialect land |
| The cast speaks a different language altogether — Cantonese, English, Japanese | The project’s dialogue language setting, not the description. Cantonese is available there as a dialogue-only option. How to set the dialogue language |
| This one line is angry, whispered, hurried, or stressed on a word | The delivery note in brackets at the start of that line. Do not repeat the accent there — it is already on the character. Dialogue in the storyboard |
The distinction that matters: an accent changes how the language is spoken, a dialect setting changes what is spoken. Asking for both in one place fights itself — writing “speaks Cantonese” in a description on a project whose dialogue language is Mandarin leaves the model with two instructions about the same sentence that pull against each other. If you want Cantonese, set Cantonese as the dialogue language and put any regional colouring in the description on top.
| What you see | What it comes from | What to do |
|---|---|---|
| Nothing at all changes after you save | The character has audio attached, so the description is not what is performed | Check the badge; on the card use Delete Voice, then generate the segment again |
| The accent is there but too faint | The strength word was vague, or the region was named only in general terms | Name it more precisely and grade it — “a clear, unmistakable Sichuan accent” — then regenerate that segment |
| Old episodes still sound the way they always did | Finished segments keep the voice they were shot with; that is what makes a delivered episode stable | Regenerate the segments where you want the accent. It costs the same as generating them the first time. What each clip fix costs |
| The accent is missing from a clip you fixed after the fact | The finished-clip voice panel synthesizes speech with preset voices from the library, so an accent written in the description does not come out of it | Regenerate the segment instead of redubbing it. Changing a voice in a clip you already have |
| The voice is right but drifts between segments | The description mode trades a little consistency for freedom — each generation performs the words afresh | Accept it, or move that character onto a preset voice from the library and give up the written accent for that character |
| The character’s voice is fine but a stranger’s line is odd | Lines spoken by someone outside the cast do not carry a character’s description | Give that speaker a character card of their own and write their accent there too, or leave it as a one-off line |
One thing to know before you reach for an uploaded voice to solve this: an attached recording and a written accent are alternatives, not partners. Attach audio and the description stops being performed — accent words included. Members who want a specific regional voice use the upload deliberately, in place of the description. How character voices work.
Yes. The voice description travels with every line that character speaks and is performed as written: the video model says the line and renders the shot in one pass, so there is no separate dubbing step that could drop it. What you do not get is a strength setting — if an accent comes out too light, say so more explicitly and generate that segment again.
Almost always because the character has audio attached: a preset you matched or a recording you uploaded. With audio in place the description is not what the model performs. On the character card use Delete Voice to go back to the description, then generate the segment again.
You can, and everything you generate from then on uses it. Segments you already shot keep the voice they were shot with; regenerate the ones where you want to hear the accent, at the usual cost of generating a segment.
No. That panel synthesizes speech with preset voices from the library, so an accent written in the description does not come out of it. An accent only arrives in a segment that is generated again.
No. The library is organized by language — Chinese, English, Cantonese — and a match is chosen by gender, age band and manner, not by the accent words in your description. The description is the only place an accent comes from.
No. An accent or dialect keeps the project’s dialogue language and puts a regional flavour on it, so it goes in the voice description. To change the language itself, set the project’s dialogue language, where Cantonese is available as a dialogue-only option; writing speaks Cantonese in the description while the dialogue language stays Mandarin asks for two things at once, and they pull against each other.
No. It is a property of the character, written once in the voice description, and it covers every line they speak. What changes line by line — mood, pace, a word under stress — goes in the delivery note in brackets at the start of that line.
Both. Each character carries their own voice description, so accents never interfere. Voice-only characters — narrators, the person on the other end of the phone — have a voice description too, and the accent works the same way for the lines they speak.
Open a character, check the badge, and add one sentence of accent to the voice description.
Updated October 10, 2026
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·Affiliate program·© 2026 SceneMixer