Consistent characters
Updated September 20, 2026

A text-to-video model renders each generation independently. Describe a woman in her thirties with a scar through one eyebrow in shot 1 and again in shot 30, and you get two women who match the description and not each other — adjectives underdetermine a face. Every fix below is the same idea in a different place: stop describing the character, and hand the model a picture of them.
| Reference sheet | Each character is generated once as a full-body sheet in 9:16 on a plain background under directional key light — a casting photo, not a mood piece, because whatever light is baked into the sheet gets inherited by every shot that uses it. |
|---|---|
| Attached crops | Face and body crops from that sheet are sent with every segment the character appears in. The model matches a picture instead of interpreting a description. |
| Links by id | Storyboards point at assets by identifier, not by name. Two characters with the same name, or a rename halfway through a project, no longer break the chain silently. |
| Off-screen rule | A character who is only being looked at, spoken to or reached for — and is not in frame — is written as a direction, never named. Naming an absent character is what makes a model paste their face into a single-person shot. |
A fifth, smaller one: each person in a shot carries one or two distinguishing facial notes pulled from their own description, which takes the edge off the classic failure where two characters in the same uniform trade faces mid-shot.
Consistency cannot mean "never changes" — a character who is set on fire in episode 6 has to look burned in episode 7. The dividing line is drawn by how a video model treats a reference image: it copies the face, the clothing, the form and the structure almost literally, and it applies light, weather and time of day as effects layered on top.
Getting this backwards is expensive in both directions: a variant for "holding the cup" is an image you paid for and will never see, and a prompt for "after the fire" is a request the model will politely ignore in favour of the clean room in its reference.
The layered explanation, with what each layer can and cannot catch, is in the character-consistency guide.
Reference sheets are their own line item, separate from script analysis (4 credits per 3,333 characters of English text): each account gets 15 reference images free, and extra or regenerated images after that are 7 credits each on the default image model. State variants carry their own small free allowance on a first project and are billed as images after that. Video itself starts at 10 credits per second of output (≈ $0.06/s at the lowest credit price) whether or not references are attached — consistency is not a surcharge. Details on the pricing page.
Consistency is the one claim in this category you should never take on trust, and you do not have to. A new account carries 50 credits, 15 free reference images across characters, scenes and props, a free episode-1 storyboard on your first project, and one free preview of up to 5 seconds that renders the opening shot of your own episode against your own sheets. That is enough for a real test:
For calibration: the median project on this platform has five named characters, and nine in ten have fewer than ten. If your story needs thirty, the same mechanism still applies — the cost is the sheets, not the consistency.
The dialogue language is set once in Step 1 and the whole run is performed in it, with the target-region setting moving names, faces and streets to that market at the same time. Each sample below was produced that way, one per language.
Languages
Every series below was produced by the same pipeline with the cast speaking that language natively: the interface, the working documents and the spoken lines are all in one language, and nothing is dubbed. Open one to hear it.
A Carta de LisboaPortuguêsNative dialogue
윈터 프라미스한국어Native dialogue
El Secreto de CostaEspañolNative dialogue
浪人の誓い日本語Native dialogue
Ikrar JakartaBahasa IndonesiaNative dialogue
Зимняя коронаРусскийNative dialogue
L’Héritier du DéfiléFrançaisNative dialogue
Il Tavolo dell'OlivaItalianoNative dialogue
De Vuurtoren van de FjordNederlandsNative dialogue
العهد الصحراويالعربيةNative dialogue
Çantasındaki SözleşmeTürkçeNative dialogue
Die Istanbul-TäuschungDeutschNative dialogue
Останній сигналУкраїнськаNative dialogue
問劍青雲繁體中文Native dialogue
The Last EnvelopeEnglishIn your languageThe same face across an entire series is the design target and what the reference-sheet architecture delivers for the cast in your asset library. It is not a guarantee on every frame: identically dressed characters sharing a shot, and characters whose sheet is mostly coat, remain the hard cases. Everything else — episode count, elapsed time, how many shots a character appears in — does not degrade it, because every shot is matched against the same picture.
Yes. Any reference sheet can be replaced with an image you upload, and the rest of the pipeline treats it exactly like a generated one. That is the route for a character you have already designed elsewhere.
Technically the upload path does not care what the image is, but publishing AI video of a real person's likeness is a legal problem in most markets and our compliance guide treats it as one. We do not provide a face-swap tool, and we do not recommend it.
Nothing breaks. Storyboards reference assets by identifier, and the rename is propagated through the stored text at the same time. Before September 2026 references were matched by name, and that is exactly the failure this change removed.
No. Attaching reference art to a segment does not change the per-second rate — you pay once for each reference sheet, then nothing further no matter how many shots use it. Consistency is an architecture here, not an upsell.
No, and you should not. The shot text names the character; the reference art carries what they look like. Repeating a physical description in the shot text pulls the model back towards the words and away from the picture, which is precisely the drift this design removes.
Yes — that is what state variants are for. Each look is its own reference image, and the storyboard decides per segment which one applies, down to the individual segment rather than the whole episode.
Generate the cast once; every shot after that matches a picture, not a description.
Try SceneMixer freeCredit packages from $1.49 · Pro $7.99/mo · No credit card to start
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer