MiniMax H3

Quick answer: MiniMax H3 is the 2K tier on SceneMixer. Like the other video tiers it generates sound with the picture and can cut between several shots inside one segment; what sets it apart is the 2K picture, segments of up to 15 seconds, and its own reference limits of 9 images and 3 voice clips. It costs 25 credits per second (about $0.16 at the lowest credit price). Pick it when you want a 15-second segment at 2K, and keep the number of references in mind.
Still from “The Last Envelope”, an AI short drama generated end-to-end by SceneMixer (6 characters, 1 episode)
Sample: “The Last Envelope” — Realistic HD, 1 episode, produced end-to-end by SceneMixer from the script · watch the episodes

H3 Fast is the 768P sibling with the same rules at a lower rate.

Where you pick it

Two places, one setting. In the episode editor: the button in the top bar that shows the current tier’s name, next to Composite. In Settings: Model Preferences → Video Model. Both change the same thing — the tier your next segments render on, in this project and in new ones. Segments already made keep the tier they were made with, and each version records it. The price on a segment’s generate button always shows the tier you are about to use.

What a segment on this tier can be

Length: up to 15 seconds in one generation; the 30-second long-take mode is not available here. Resolution: 2K, in the project’s aspect ratio. Sound: always on — lines in each character’s voice and ambient sound come out of every render. Cuts: a segment can hold several shots with genuine cuts between them inside a single generation; the storyboard’s shot list is the edit.

References: the rule that trips people

Two limits to watch: up to 9 images and up to 3 voice clips (about 14 seconds of voice in total). When a segment would go over, the site leaves out the least important references first. A segment that names many characters and props, plus their voices, can reach the limits; when it does, prefer fewer, better references. A segment keeps at most three voice clips (fewer if they add up to more than about 14 seconds); the other speakers use their written voice descriptions. Voice clips need an image alongside: if a segment would carry voices but no image, the site drops the voice clips and falls back to each character’s written voice description.

What it costs

25 credits per second. H3 Fast is 16 per second for the same rules at 768P; for comparison the Wan 1080P tier is 38 per second. Credits are taken on delivery; if a generation fails, the credits come back or are held until support checks it.

What it is good at, and what to watch

Good at: a 15-second segment rendered sharp enough for a large screen — like the Wan tiers, it can hold an exchange of looks, a reaction and an insert with cuts between them. Watch: the reference cap, the 15-second ceiling, and the fact that this family has its own reference limits, different from the Wan tiers’; a segment that worked on a Wan tier may need its references thinned to run here.

Content checks

Text and references are checked before rendering. A rejection is not charged: the credits come back and the credit history says so. The count is kept per account: after 12 rejections within 24 hours, the next one is held for review instead of refunded, and video generation pauses for 1 hour. So edit the flagged material rather than retrying the same segment. Content rules and checks explains what gets flagged.

Common errors

The messages you may meet on this tier:

What you seeUsual causeWhat to do
“Invalid input — check the reference images and duration, then retry.”Usually one of the reference files for this segment could not be readGenerate again once; if it keeps failing, check that the character, location and prop images open, then contact support
“Blocked by content safety rules — edit the script or description and retry.”The script or a reference tripped the content checkEdit the flagged part; see the content-check section for when a rejection is refunded
“The generation service hit a temporary problem. You weren’t charged — feel free to try again.”The service was saturated; the site already tried again once for youWait a few minutes and generate again; nothing was deducted
Video generation pausedMore than 12 rejected generations within 24 hours; the next one is held for reviewWait out the pause (1 hour); fix the material before retrying

When to switch tier

To H3 Fast for the same rules at a lower rate while drafting. To a Wan or Seedance 2.5 tier when a scene needs more than 9 image references, or an unbroken 30-second take — the latter only if the episode’s storyboard was generated with 30-second segments. Each segment remembers its tier.

Reference typeLimit on this tierNote
ImagesUp to 9At least one image is required when any reference is attached
Voice clipsUp to 3, about 14 s in totalVoices need an image alongside; with no image the site drops the voice clips and uses the written voice descriptions
Over a limit—The site leaves out the least important references first; a character whose voice is left out speaks from the written voice description

FAQ

Is 2K higher than the 1080P tiers?

Yes in pixel count. Whether it shows depends on where the clip is watched; on phones the difference is small.

Why was a reference dropped from my segment?

The segment had more than 9 images, more than 3 voice clips or too many seconds of voice, or a voice clip had no image with it. Name fewer characters or props in that segment and regenerate.

Can H3 render a 30-second take?

No. Segments on this tier are up to 15 seconds.

Does H3 always add sound?

Yes. Every render includes dialogue and ambient sound; there is no silent mode.

Do H3 segments mix with Wan segments in one episode?

Yes. Each segment remembers its tier and the composite stitches the active versions.

Render a 15-second segment at 2K

Pick H3 for a segment you want at 2K, keep the references lean, and see how it holds on a large screen.

Start a project

Updated September 27, 2026

Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer