Seven shots from one Wan 3.0 prompt: the format, and where it slips
There is no director mode. Multi-shot is a prompt format, the model cuts on its own unless told not to, and the shot count you ask for is not the one you get.

Start with the thing that saves you an afternoon: there is no director mode.
No toggle, no shots parameter, nothing in the Wan 3.0 request that takes a
shot count. Several well-ranked guides describe a six-shot "AI Director" feature
with per-shot controls; open Alibaba's parameter table and look for
the field
and it is not there, because multi-shot in Wan 3.0 is a prompt format rather
than a setting.
That is better news than it sounds. A format costs nothing to try, works identically on every provider, and you can change it mid-sentence. What it also means is that nothing enforces your shot list, which is the part worth planning around.
There is no shot-count parameter in the Wan 3.0 API reference or the create-task schema. Both are worth opening before believing a page that says otherwise. Read 2026-08-25.
The model cuts whether you ask it to or not
This is the single most useful fact about long Wan 3.0 generations and almost nobody leads with it.
Leave the coverage unspecified and Wan 3.0 picks its own. On a loose prompt it will happily cut between two or three setups inside five seconds, and across thirty it will build something that looks edited. Competent, usually. Entirely its film rather than yours.
So there are two jobs, and they need opposite instructions:
| You want | Write |
|---|---|
| A deliberate sequence of shots | The numbered format below |
| One unbroken take, no cuts at all | The literal phrase "one continuous take" |
That second row is not a stylistic suggestion. Alibaba's own guidance names the phrase, and without it a thirty-second request is a coin flip on whether you get a long take or a small edited film. If you came here from the article on choosing a length, this is the sentence that makes the thirty-second single take actually a single take.
The numbered format
The official Wan 3.0 shape is a one-line summary, then numbered shots each carrying a time range, and a closing line for tone and audio.
A short piece about a baker opening up before dawn, told in seven shots.
Shot 1 [0-4s]: Wide — the shutter rolls up, street still dark, breath visible.
Shot 2 [4-9s]: Medium — she ties an apron, flour dust catching the work light.
Shot 3 [9-13s]: Close-up — dough turned out onto steel, hands pressing once.
Shot 4 [13-18s]: Medium — trays slide into the oven, door thumps shut.
Shot 5 [18-22s]: Insert — the clock hits six, light changing at the window.
Shot 6 [22-26s]: Medium — she sets a loaf in the window, straightens, exhales.
Shot 7 [26-30s]: Wide — the first customer pushes the door, bell over it.
Overall tone: warm, unhurried. Room tone, oven hum, the bell at the end.
No music.Four rules that come out of the official examples:
- Time ranges must tile the whole duration. Gaps and overlaps are where pacing goes wrong.
- One primary event per shot, with a visible end state. "She ties an apron" is a shot. "She gets ready" is a mood.
- Do not write the transitions. No "cut to", no "then". The model infers a cut from the shot boundary, and transition language tends to confuse the structure rather than tighten it.
- State the audio once, at the end. Wan 3.0 adds music and spoken lines on
its own initiative, so
No musicis a real instruction that you will want more often than you expect.
Seven shots across thirty seconds averages 4.3 seconds each, which is about the floor for a shot that reads. Below three seconds you are describing a montage, and the model will treat it as one. If you would rather start from a structure than a blank box, the prompt generator lays a brief out in numbered beats.
Where it slips: the shot count is a request, not a contract
Because nothing enforces the list, the delivered cut count can differ from the one you wrote — and it does. Independent cut-counting across Alibaba's own published showcase found single multi-shot jobs coming back with roughly half the shots the prompt specified.
That is worth internalising before you plan a seven-shot brief. The format gives the model an intention, not a timeline it is bound to. Three consequences:
- Do not put a hard requirement in shot seven. If the logo has to appear, it should not depend on the model honouring your last time range.
- Count the cuts in the result before you judge the content. A sequence that "feels rushed" is usually a sequence that merged two of your shots.
- The more shots you ask for, the more the count drifts. Four to five shots across thirty seconds is far more reliable than seven, and seven is more reliable than a shot every two seconds.
Keeping a character stable across the cuts
Consistency inside one continuous take is mostly Wan 3.0's problem. Across cuts it becomes yours, and the failure is specific: the face usually holds while the accessories wander. In Alibaba's own demo footage, observers found a hat band and a necklace pendant both changing while the face stayed steady — inside a single unbroken take, not even across a cut.
Three habits that work, in order of how much they buy you:
Change exactly one variable per shot. Most consistency complaints are a prompt that moved the location, the camera angle and the lighting at the same boundary. The model has nothing to hold on to. Move one thing, hold the rest.
Restate the full wardrobe in every shot description. It feels redundant and it is the highest-value redundancy in the whole prompt. "She" in shot five is not the same instruction as "she, in the grey apron over a white tee, hair tied back" in shot five.
Bind your references by name and reuse the name. Wan 3.0 takes 10 reference images, 5 reference videos and 5 reference audio clips, addressed in the prompt as "Image 1", "Video 2" and so on, with images and videos counted separately. Bind once at the top — Image 1 defines the baker; take only the apron from Image 2 — and then refer to the binding rather than re-describing the person.
One economic note that shapes how you compose all this: reference images cost nothing extra, reference video seconds are billed at your output rate. Ten stills are free; a fifteen-second reference clip is the most expensive input on the menu. For identity and wardrobe, use stills.
Style: 3D and cartoon hold up, if you legislate them once
Stylised looks — 3D animation, cel-shaded cartoon, claymation — hold across a multi-shot Wan 3.0 sequence about as well as photoreal does, and they fail the same way when they fail: drift at a shot boundary where too much changed.
The technique is to put style in the global rules, above the shot list, rather than repeating it per shot:
Style: 3D animation, soft rounded forms, matte surfaces, no photoreal skin
texture. Palette: warm ochre, cream, one accent of teal. Key light always
from frame left. Do not change render style between shots.Two things earn their place there. A named palette is far more stable than "warm colours", because the model can hold three named values across seven shots and cannot hold an adjective. And an explicit prohibition — do not change render style between shots — is worth writing out; the exclusions layer is where the official long-form examples are most aggressive, and style drift is exactly what it exists to prevent.
When to split the job instead
Sometimes the right answer is not a better prompt. If the sequence has a hard deliverable in it — a product that must be correct, a logo that must land, a line that must be said — the shot-count drift above makes a single seven-shot generation the wrong tool.
The alternative that works:
- Lock the look once, as a single still.
- Build one keyframe per shot from that still, changing exactly one thing each time.
- Generate each shot separately with the reference bound.
- Stitch locally.
It costs more generations and it gives you a shot count you chose. Worth it when something in the frame is non-negotiable; overkill for a mood piece.
And if exactly one detail is wrong in one shot, do not regenerate the shot. Wan 3.0 can edit an existing clip in place — change the prop, keep the actions, the wardrobe and the rest of the frame — which is a fraction of the cost of rolling the dice again on all thirty seconds.
Four claims about multi-shot that the parameter table does not support
| Circulating claim | What the documentation shows |
|---|---|
| An "AI Director" mode with per-shot parameters | No such field. Multi-shot is prompt formatting |
| A hard six-shot ceiling | No shot parameter exists, so there is nothing to cap. What limits you is thirty seconds and readability |
| Cross-session character memory, marketed as "Identity Lock" | No such field, and nothing in Wan 3.0 persists between jobs. Consistency comes from reference bindings you supply every time |
| Phoneme-level lip sync across twelve languages | Not a documented capability. Wan 3.0 generates speech with the picture; no language count is published |
If a guide describes a feature you cannot express in a request, it was not written from the docs. That is a useful filter well beyond this one model.
A seven-shot checklist
- Summary line first, then numbered shots, then tone and audio.
- Time ranges tile the full duration with no gaps.
- One primary event and one end state per shot.
- No transition words between shots.
- Style, palette and prohibitions in a global block above the list.
- Wardrobe restated in every shot that contains the character.
- Nothing load-bearing in the final shot.
Questions
Does Wan 3.0 have a multi-shot or director mode?
No. There is no shot parameter in the API and no mode to switch on. You get multiple shots by writing numbered shots with time ranges into the prompt, which means it works the same way on every provider.
How many shots can one generation have?
There is no documented limit, because there is no setting to limit. Thirty seconds is the real constraint — four to five shots hold together most reliably in Wan 3.0, and seven is workable at roughly four seconds each. Ask for more and the model starts merging them.
How do I stop Wan 3.0 from cutting?
Write "one continuous take". Left unspecified the model chooses its own coverage and will usually cut, even on short clips.
Why does my character change between shots?
Almost always because one shot boundary changed several things at once. Move one variable per shot, restate the wardrobe in full every time, and bind an identity reference image rather than re-describing the person. Watch accessories specifically — hats, jewellery and props drift long before faces do.
Does 3D or cartoon style work as well as photoreal?
Yes, and Wan 3.0 holds a stylised look across cuts about as well as it holds a photoreal one. Put the style, a named colour palette and an explicit "do not change render style between shots" in a global block above the shot list rather than repeating them shot by shot.
Written by
Editorial desk
wan-3.run


