Costume consistency
Updated September 20, 2026

The instinct is to describe the clothes in every shot. It is the wrong move, and expensively so: a paragraph of wardrobe in a shot prompt competes with the face for the model's attention, and in a tall frame it pulls the camera back far enough to shrink the head until identity is gone. A pale overcoat filling seventy percent of the frame with a tiny face on top is a real failure mode, not a hypothetical.
So clothing lives on the character's reference sheet, and shots refer to the person rather than re-describing the outfit. What the sheet establishes, every shot inherits.
| The reference sheet | One outfit, described in cause terms — cut, fabric, construction, closure — and rendered as a casting photo on white. |
|---|---|
| The costume axis | A project-level register (period, tailoring tradition, fabric and craft) injected once during analysis, so the whole cast is dressed in one world. |
| A second outfit | Becomes its own reference image — a state of the character, not a sentence in a shot. |
| Small changes | Rain, dust, a loosened collar, something held: written as text in the shot. No second image needed. |
| Shot prompts | Refer to the character. They do not re-list the clothing, and they do not pull back to display it. |
This is the distinction that decides whether you need a second image, and it comes from how reference-conditioned video models behave: they copy what is in the picture — face, garments, accessories, structure — almost verbatim, and they accept as text anything that sits on top of the picture as a global effect — weather, time of day, light, dirt.
Getting this line wrong in the direction of too many images is the common error: early on, most auto-detected “states” were things like holding a cup or a lamp being lit. Those are text. The pipeline now caps how many states it will propose per character precisely because over-generating them costs money and buys nothing.
Scripts routinely write a character's wardrobe as an arc — early episodes in one thing, later in another, plus a wedding look. Handed all three as a single description, the render draws a figure wearing half of each. So the description is split structurally before any image is generated: the first outfit stays on the base sheet, and the rest become separate state images. Part labels (footwear, accessories, fabric) are recognised as parts of one outfit rather than as a second one.
Per shot, in priority: a state chosen for that specific segment, then a per-episode override image, then a state assigned for that episode, then the base sheet. Only states that have actually been rendered count — if a state exists on paper but has no image, the shot falls back to the base sheet and carries the difference as text instead, so the change is still described even when there is no picture of it.
Assignment is automatic: after a shot list is generated, a pass decides which segments use which state and writes it down. It is done per segment rather than per episode on purpose — a flashback outfit applied to a whole episode would put the character in the wrong clothes for every shot in it.
Each state image is a generated image: 7 credits. Your first project includes 15 free reference images across characters, locations and props, and state images have their own separate free allowance in that first project — so trying one costs nothing. Text-only changes cost nothing at all, which is the main reason the state-versus-sentence line is worth learning. See character sheets for how the base image is built, and the pricing page for rates.
Languages
Every series below was produced by the same pipeline with the cast speaking that language natively: the interface, the working documents and the spoken lines are all in one language, and nothing is dubbed. Open one to hear it.
A Carta de LisboaPortuguêsNative dialogue
윈터 프라미스한국어Native dialogue
El Secreto de CostaEspañolNative dialogue
浪人の誓い日本語Native dialogue
Ikrar JakartaBahasa IndonesiaNative dialogue
Зимняя коронаРусскийNative dialogue
L’Héritier du DéfiléFrançaisNative dialogue
Il Tavolo dell'OlivaItalianoNative dialogue
De Vuurtoren van de FjordNederlandsNative dialogue
العهد الصحراويالعربيةNative dialogue
Çantasındaki SözleşmeTürkçeNative dialogue
Die Istanbul-TäuschungDeutschNative dialogue
Останній сигналУкраїнськаNative dialogue
問劍青雲繁體中文Native dialogue
The Last EnvelopeEnglishIn your languageBecause re-describing it competes with the face for attention and pushes the camera back to display the clothes, shrinking the head until the character is unrecognisable. Wardrobe is established once on the reference sheet and inherited by every shot that refers to that character.
When the garment itself changes — a full change of outfit, armour, a uniform, something burned or torn beyond recognition. If the same garment is just wet, dusty, bloodied or unbuttoned, write it in the shot text; reference-conditioned models accept those as effects layered over the picture.
The description is split before anything is drawn: the first outfit stays on the base sheet and the later ones become their own state images. Handed all of them at once, a single render draws a figure wearing parts of each.
A state chosen for that segment wins, then a per-episode override, then a state assigned to the episode, then the base sheet. States that have no rendered image do not count — the shot uses the base sheet and describes the difference in words.
No. Wardrobe is fixed when the shot is generated, so a change means generating that segment again at the usual per-second price.
Fifteen free reference images is a whole small cast.
Try SceneMixer freeCredit packages from $1.49 · Pro $7.99/mo · No credit card to start
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer