
The ad pipeline runs on the same engine as the drama pipeline, so the storyboard it produces is the same working document you can see in this sample project.
Up to four photos of the same subject: different angles, a detail shot, the thing in use. The first one is treated as the main subject — it is the one the end card is built around, and, for a product, a dish or a garment, the one that gets redrawn as a clean shot on white. A shop, a car or an app screen is not redrawn (see below). The rest become locations.
They are read before anything is charged. If the read comes back unusable — too blurry, nothing recognisable, a screenshot at a strange aspect ratio — you get told on the spot and nothing is spent.
There is a one-line box next to the upload. Use it when the picture alone is misleading: “this is our shop, not a stock photo” changes the whole treatment, because a shop is a place the ad happens inside, while a product is a prop the ad happens around.
Where the sentence and the photo disagree, the sentence wins — this is the one place where what you type outranks what the model sees.
Length is 15, 30 or 60 seconds, which is one, two or four segments of fifteen seconds. Tone picks the narrative recipe — spectacle, funny, warm, minimal, twist, appetite, demo or lifestyle — and each one carries its own art direction as well as its own structure.
Resolution is 480p, 720p or 1080p. The price on each option already includes every step, so the number you see is the number you pay.
Handing it to the AI runs the whole chain unattended and quotes one price up front: clean the photo, write the script, break it down, generate the cast and sets, shoot each segment, composite with the end card.
The step-by-step mode charges less up front and hands you the normal editor after the script is written, so you can rewrite a line, change a character or reshoot one segment before paying for the rest.
Products, food and clothing are redrawn as a clean studio shot on white so they can be lit consistently in every segment. Shops, restaurants, cars and houses are not redrawn — cutting them out of their surroundings destroys the signage, the depth and the sense of scale that make them recognisable. App and web screenshots are not redrawn either; redrawing turns readable interface into mush.
If there are people in a photo of a place or a vehicle, they are removed before anything else happens. A model given a reference frame copies the faces in it, and putting a real bystander into your ad is not something we will do.
Superlatives and absolute claims are stripped from the script before it is produced — “the lowest price anywhere”, “number one in sales”, “cures everything”. Brand and product names go through the same pass, but only against that shorter list of outright claims — an ordinary trademark is left alone, while a name that itself makes a claim (“#1”, “100%”, “guaranteed”) is softened.
Every segment carries at least one spoken line, and the tagline is said out loud before the final freeze. Nothing legible is written into the picture — video models cannot render readable text, so the words that have to land are delivered by voice and by the end card.
| Setting | What it does |
|---|---|
| Photos | 1–4 of the same subject; the first is the main one |
| Length | 15s / 30s / 60s → 1 / 2 / 4 segments |
| Tone | Eight narrative recipes, each with its own art direction |
| Resolution | 480p / 720p / 1080p, price shown per option |
| End card | Brand, product, tagline, call to action |
| Music | Generated with the picture, not added afterwards |
No. The script is written from the photos and the one-line description. In step-by-step mode you can rewrite it before anything is shot.
Yes, and it is handled differently on purpose: a place becomes the location the ad happens inside, and the first scene is rebuilt from your own photo rather than imagined from text.
Only on the end card. Video models cannot render readable text reliably, so anything that must be read exactly is put on the card at the end, and everything else is spoken.
They are removed before the shoot for shops, restaurants and vehicles. Screenshots containing people are declined outright rather than edited.
You can recomposite with different end-card text, and in step-by-step mode you can reshoot individual segments. The finished file can also be downloaded and shared with a link.
The read is free, and the price for the whole ad is shown before you start.
Updated September 21, 2026
Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer