
After a read, open the episode in Step 3 and read its outline: each scene block lists its lines in speaker-colon form.
Speech recognition runs on each shot’s own audio. What it hears is handed to the shot description as ‘these words were spoken here’; the describer then names the speaker from the faces and writes the line into the shot. Music under speech is handled; shouting over a crowd, heavy accents and very quiet lines are where it fails, and those are left out rather than invented. Lines then appear in the outline like a script’s: Write Dialogue So It Is Picked Up.

Left to itself, a vision model asked ‘what is said here’ will sometimes hear and sometimes not, on the same shot, and will invent plausible lines when unsure. Recognised words pinned into the description stop both; the same words are used every time the shot is described.
The read asks once whether the clip has music and what kind. If it does, each segment you generate carries a one-line music brief matching it, so the generated segments have music of the same character; the source track itself is never used. Ambient sound is noted but not carried. Where the Music Comes From.
Chosen on the upload screen, defaulting to your interface language (the Chinese site is always Chinese), and fixed once the read starts. Lines heard in another language are translated into it. It cannot be changed afterwards for a remake, which is why Step 1 does not offer the dropdown. After the Read.
The outline itself cannot be edited, and the storyboard copies its lines word for word. Generate the storyboard, then open the segment where the line belongs and add or fix it there, with its speaker; a line given to the wrong person is fixed the same way. Editing is free. Generate the Storyboard.
| In the clip | In the read | In your episode |
|---|---|---|
| A clear spoken line | Quoted, speaker named | A line in the outline |
| Nobody speaks | Shot marked silent | No line |
| Unclear speech | Left out | Add it in the segment by hand |
| Music | Character noted | Same-character music in every segment |
| Ambient sound | Noted | Not carried; the storyboard writes its own |
The recogniser could not make it out. Compare with your clip and add it in the segment once the storyboard exists.
No. Voices come from the cast's voice settings, as in any project.
Yes; pick the dialogue language on the upload screen and the lines are translated at the read.
No. Its character is described and the segments generate their own.
Group noise is treated as background sound, not dialogue.
Updated September 27, 2026
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer