AI multi-character scenes
Updated September 20, 2026

Generated video has a specific failure with more than one person in a shot: it will put in a face that should not be there. Someone mentioned in the text — looked at, shouted to, thought about — arrives on screen, wearing a face invented on the spot, because nothing told the model the difference between “present” and “referred to”.
So the shot list states it per shot. Each shot carries an explicit list of which characters are in frame, and that list — not the prose, not the previous shot's ending state — is the only thing that decides whose reference art gets attached.
| Who is in frame | A per-shot list. The union across a segment's shots decides which reference sheets are sent with it. |
|---|---|
| Who is off screen | Written as a direction — “towards frame left”, “eyeline to camera” — not as a name. No face is attached. |
| Single-person shots | Marked explicitly, so the renderer is told the frame holds one person and not to invent company. |
| Who is speaking | Named per line; everyone else in frame is marked lips closed, so only the speaker's mouth moves. |
| Background crowd | A declared state of the scene, written once in the segment header, or absent — never implied by ambient sound. |
| Practical ceiling | Four named people in one frame is the point past which quality falls off. The shot list works around it rather than through it. |
This is measurable, not theoretical. Shots that named an off-screen antagonist produced the wrong face or an extra person in five out of eight renders. The same shots rewritten to give a direction instead of a name produced zero such failures in twelve renders. So the pipeline rewrites them: in a single-person shot, a character who exists only as an eyeline or a direction of movement is converted to a spatial phrase and removed from the frame list.
The same logic covers the dead: a character referred to only as a body in this segment is written as a plain noun, not as a character reference, so nobody attaches their reference sheet and stands them up.
There are exactly two ways to shoot it and the list is required to pick one: a profile two-shot, or an over-the-shoulder. What it may not write is “A facing B, with B behind in the background, out of focus” — for B to be in A’s depth of field, A has to have their back to B, and that camera position does not exist in a normal medium shot. Asked for it anyway, the renderer puts one person face-on to camera and the other standing behind them. The axis is set in the first shot where both are actually in frame together.
A room of three hundred guests is a visible state of the scene, so it is written as one — a line in the segment header saying who the background people are, where, and what they are doing. Consistent across every segment in that scene. Without that line the shot may not refer to a crowd at all, because “murmur of guests” in the ambient audio is not a visual instruction, and a renderer told only that will give you guests in one shot and an empty hall in the next.
When a crowd is declared, the frame instruction changes with it: the crowd stays in the depth of field, and the in-frame list still counts only named characters.
Nothing extra — these are properties of the shot list, which is billed by script length at 1 credit per hundred characters of the episode's share. Reference sheets are 7 credits each beyond the 15 free ones in your first project, and rendering starts at 10 credits per second of output. See character consistency for how faces stay the same across shots, and the pricing page for rates.
Languages
Every series below was produced by the same pipeline with the cast speaking that language natively: the interface, the working documents and the spoken lines are all in one language, and nothing is dubbed. Open one to hear it.
A Carta de LisboaPortuguêsNative dialogue
윈터 프라미스한국어Native dialogue
El Secreto de CostaEspañolNative dialogue
浪人の誓い日本語Native dialogue
Ikrar JakartaBahasa IndonesiaNative dialogue
Зимняя коронаРусскийNative dialogue
L’Héritier du DéfiléFrançaisNative dialogue
Il Tavolo dell'OlivaItalianoNative dialogue
De Vuurtoren van de FjordNederlandsNative dialogue
العهد الصحراويالعربيةNative dialogue
Çantasındaki SözleşmeTürkçeNative dialogue
Die Istanbul-TäuschungDeutschNative dialogue
Останній сигналУкраїнськаNative dialogue
問劍青雲繁體中文Native dialogue
The Last EnvelopeEnglishIn your languageFour named people is the practical ceiling before identity starts to degrade, and the shot list is built to stay under it — ensemble scenes are covered in singles and over-the-shoulder shots rather than wide group frames.
That is exactly the failure the per-shot frame list prevents. A character who is only looked at or spoken to is written as a direction, not a name, and is not in the frame list — so no reference art is attached and no face is invented.
Yes, as a profile two-shot or an over-the-shoulder. What does not work is one facing the other with the other blurred in the background: that combination of framing and depth is physically inconsistent, and asking for it produces a person standing behind someone talking to camera.
They are declared once in the segment header — who they are, where, what they are doing — and stay consistent across that scene. If nothing is declared, the shot does not reference a crowd. Ambient crowd noise alone is not treated as a visual instruction.
Only the named speaker of each line. Everyone else in frame is explicitly marked with their lips closed, which is what stops a listener from silently mouthing along with the dialogue.
Generate one crowded scene and count the faces.
Try SceneMixer freeCredit packages from $1.49 · Pro $7.99/mo · No credit card to start
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer