Costume consistency

AI Costume & Outfit Consistency

Updated September 20, 2026

Quick answer: clothing is established once on a character's reference sheet and inherited by every shot that refers to them — shots never re-list the outfit, because doing so competes with the face and pulls the camera back until the head is a dot. A real change of garment becomes its own reference image; weather, dirt and a loosened collar are written as text, because reference-conditioned models copy what is in the picture and accept effects layered over it.
Still from “Neon Debt”, an AI short drama generated end-to-end by SceneMixer (7 characters, 1 episode)
Sample: “Neon Debt” — Game Engine Realism, 1 episode, produced end-to-end by SceneMixer from the script · watch the episodes

Wardrobe is established once, then referred to

The instinct is to describe the clothes in every shot. It is the wrong move, and expensively so: a paragraph of wardrobe in a shot prompt competes with the face for the model's attention, and in a tall frame it pulls the camera back far enough to shrink the head until identity is gone. A pale overcoat filling seventy percent of the frame with a tiny face on top is a real failure mode, not a hypothetical.

So clothing lives on the character's reference sheet, and shots refer to the person rather than re-describing the outfit. What the sheet establishes, every shot inherits.

At a glance

Where wardrobe is decided
The reference sheetOne outfit, described in cause terms — cut, fabric, construction, closure — and rendered as a casting photo on white.
The costume axisA project-level register (period, tailoring tradition, fabric and craft) injected once during analysis, so the whole cast is dressed in one world.
A second outfitBecomes its own reference image — a state of the character, not a sentence in a shot.
Small changesRain, dust, a loosened collar, something held: written as text in the shot. No second image needed.
Shot promptsRefer to the character. They do not re-list the clothing, and they do not pull back to display it.

The line between a state and a sentence

This is the distinction that decides whether you need a second image, and it comes from how reference-conditioned video models behave: they copy what is in the picture — face, garments, accessories, structure — almost verbatim, and they accept as text anything that sits on top of the picture as a global effect — weather, time of day, light, dirt.

Getting this line wrong in the direction of too many images is the common error: early on, most auto-detected “states” were things like holding a cup or a lamp being lit. Those are text. The pipeline now caps how many states it will propose per character precisely because over-generating them costs money and buys nothing.

Two outfits written into one description

Scripts routinely write a character's wardrobe as an arc — early episodes in one thing, later in another, plus a wedding look. Handed all three as a single description, the render draws a figure wearing half of each. So the description is split structurally before any image is generated: the first outfit stays on the base sheet, and the rest become separate state images. Part labels (footwear, accessories, fabric) are recognised as parts of one outfit rather than as a second one.

Which image a given shot uses

Per shot, in priority: a state chosen for that specific segment, then a per-episode override image, then a state assigned for that episode, then the base sheet. Only states that have actually been rendered count — if a state exists on paper but has no image, the shot falls back to the base sheet and carries the difference as text instead, so the change is still described even when there is no picture of it.

Assignment is automatic: after a shot list is generated, a pass decides which segments use which state and writes it down. It is done per segment rather than per episode on purpose — a flashback outfit applied to a whole episode would put the character in the wrong clothes for every shot in it.

What it will not do

What it costs

Each state image is a generated image: 7 credits. Your first project includes 15 free reference images across characters, locations and props, and state images have their own separate free allowance in that first project — so trying one costs nothing. Text-only changes cost nothing at all, which is the main reason the state-versus-sentence line is worth learning. See character sheets for how the base image is built, and the pricing page for rates.

Languages

Native dialogue in 15 languages

Every series below was produced by the same pipeline with the cast speaking that language natively: the interface, the working documents and the spoken lines are all in one language, and nothing is dubbed. Open one to hear it.

See all 15 languages

Frequently asked questions

Why is the outfit not described in every shot?

Because re-describing it competes with the face for attention and pushes the camera back to display the clothes, shrinking the head until the character is unrecognisable. Wardrobe is established once on the reference sheet and inherited by every shot that refers to that character.

When do I need a second image for a costume?

When the garment itself changes — a full change of outfit, armour, a uniform, something burned or torn beyond recognition. If the same garment is just wet, dusty, bloodied or unbuttoned, write it in the shot text; reference-conditioned models accept those as effects layered over the picture.

My script describes a character's wardrobe changing over the season. What happens?

The description is split before anything is drawn: the first outfit stays on the base sheet and the later ones become their own state images. Handed all of them at once, a single render draws a figure wearing parts of each.

Which image gets used in a given shot?

A state chosen for that segment wins, then a per-episode override, then a state assigned to the episode, then the base sheet. States that have no rendered image do not count — the shot uses the base sheet and describes the difference in words.

Can I change the outfit in a shot I already rendered?

No. Wardrobe is fixed when the shot is generated, so a change means generating that segment again at the usual per-second price.

Dress the cast once

Fifteen free reference images is a whole small cast.

Try SceneMixer free

Credit packages from $1.49 · Pro $7.99/mo · No credit card to start

Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer