How to Make an Ad from One Photo

Quick answer: Upload one to four photos of whatever you are selling — a product, a dish, a car, a shop front, an app screen — and SceneMixer reads them, works out what the thing is, writes a short script around it, generates the cast and locations, shoots each segment and finishes with an end card carrying your brand, product name, tagline and call to action. You pick the length, the shape and the resolution; everything else is one button.
Still from “Rocket Pup”, an AI short drama generated end-to-end by SceneMixer (7 characters, 1 episode)
Sample: “Rocket Pup” — Feature 3D Animation, 1 episode, produced end-to-end by SceneMixer from the script · watch the episodes

The ad pipeline runs on the same engine as the drama pipeline, so the storyboard it produces is the same working document you can see in this sample project.

Step 1 — Add the photos

Up to four photos of the same subject: different angles, a detail shot, the thing in use. The first one is treated as the main subject — it is the one the end card is built around, and, for a product, a dish or a garment, the one that gets redrawn as a clean shot on white. A shop, a car or an app screen is not redrawn (see below). The rest become locations.

They are read before anything is charged. If the read comes back unusable — too blurry, nothing recognisable, a screenshot at a strange aspect ratio — you get told on the spot and nothing is spent.

Step 2 — Say what it is, if the photo is ambiguous

There is a one-line box next to the upload. Use it when the picture alone is misleading: “this is our shop, not a stock photo” changes the whole treatment, because a shop is a place the ad happens inside, while a product is a prop the ad happens around.

Where the sentence and the photo disagree, the sentence wins — this is the one place where what you type outranks what the model sees.

Step 3 — Pick length, tone, shape and resolution

Length is 15, 30 or 60 seconds, which is one, two or four segments of fifteen seconds. Tone picks the narrative recipe — spectacle, funny, warm, minimal, twist, appetite, demo or lifestyle — and each one carries its own art direction as well as its own structure.

Resolution is 480p, 720p or 1080p. The price on each option already includes every step, so the number you see is the number you pay.

Step 4 — Let it run, or take the wheel

Handing it to the AI runs the whole chain unattended and quotes one price up front: clean the photo, write the script, break it down, generate the cast and sets, shoot each segment, composite with the end card.

The step-by-step mode charges less up front and hands you the normal editor after the script is written, so you can rewrite a line, change a character or reshoot one segment before paying for the rest.

What it does to your photo, and why

Products, food and clothing are redrawn as a clean studio shot on white so they can be lit consistently in every segment. Shops, restaurants, cars and houses are not redrawn — cutting them out of their surroundings destroys the signage, the depth and the sense of scale that make them recognisable. App and web screenshots are not redrawn either; redrawing turns readable interface into mush.

If there are people in a photo of a place or a vehicle, they are removed before anything else happens. A model given a reference frame copies the faces in it, and putting a real bystander into your ad is not something we will do.

The copy rules

Superlatives and absolute claims are stripped from the script before it is produced — “the lowest price anywhere”, “number one in sales”, “cures everything”. Brand and product names go through the same pass, but only against that shorter list of outright claims — an ordinary trademark is left alone, while a name that itself makes a claim (“#1”, “100%”, “guaranteed”) is softened.

Every segment carries at least one spoken line, and the tagline is said out loud before the final freeze. Nothing legible is written into the picture — video models cannot render readable text, so the words that have to land are delivered by voice and by the end card.

SettingWhat it does
Photos1–4 of the same subject; the first is the main one
Length15s / 30s / 60s → 1 / 2 / 4 segments
ToneEight narrative recipes, each with its own art direction
Resolution480p / 720p / 1080p, price shown per option
End cardBrand, product, tagline, call to action
MusicGenerated with the picture, not added afterwards

FAQ

Do I need a script?

No. The script is written from the photos and the one-line description. In step-by-step mode you can rewrite it before anything is shot.

Can I use a photo of my shop rather than a product?

Yes, and it is handled differently on purpose: a place becomes the location the ad happens inside, and the first scene is rebuilt from your own photo rather than imagined from text.

Will the ad show my logo or price on screen?

Only on the end card. Video models cannot render readable text reliably, so anything that must be read exactly is put on the card at the end, and everything else is spoken.

What happens if my photo has customers in it?

They are removed before the shoot for shops, restaurants and vehicles. Screenshots containing people are declined outright rather than edited.

Can I change the ad after it is finished?

You can recomposite with different end-card text, and in step-by-step mode you can reshoot individual segments. The finished file can also be downloaded and shared with a link.

Try it with one photo

The read is free, and the price for the whole ad is shown before you start.

Make an ad

Updated September 21, 2026

Pricing·Novel to Video·Guides·Samples·Support·Privacy·Terms·Legal·Contact·© 2026 SceneMixer