Wan 3.0 prompt guide

A Wan 3.0 prompt guide from Alibaba's own formulas: the six layers the model reads, the rewriter that runs before every take, and dialogue it performs.
Sep 7, 2026

A Wan 3.0 prompt is read as one continuous take, in layers, and the model fills in every layer you leave blank. Which means the fix for a clip that came back wrong is almost never a longer prompt. It is naming the layer you left to chance.

Here is the whole shape. Everything further down is one of these six lines, with the rule it comes from and the date that rule was checked.

Subject     who or what, described the way you want it to stay
Scene       where it happens, and what the light is doing
Motion      what changes between the first second and the last
Camera      where the lens is, and whether it moves
Sound       a line, an effect, music — any of the three, or none
Reference   Image 1, Video 1, Audio 1, if you attached anything

Paste that into the Wan 3.0 prompt generator and it will fill the layers in Alibaba's own phrasing — free, no account, no daily cap. Or write them yourself and use this page to check each one.

What Wan 3.0 does with a layer you leave out

Same pocket watch, same seed: the clip that named the camera came back overhead, the clip missing its Camera line came back at eye level

Two matched pairs, run here on 2026-09-07. Same seed inside each pair, the rewriter switched off, five seconds at 480P, and the only edit between the two prompts was deleting the Camera line and the Sound line.

Camera and Sound writtenBoth lines deleted
Coffee cup, seed 71547Framing holds across all twenty sampled framesFraming drifts through the clip
Pocket watch, seed 88213Overhead close-up, as askedEye-level three-quarter view — a different shot of the same scene
Cup, audio levelmean −47.5 dB, peak −35.9 dBmean −16.8 dB, peak −2.2 dB
Watch, audio levelmean −31.5 dB, peak 0.0 dBmean −28.4 dB, peak −4.9 dB

Two of those held in both pairs, and both are worth writing prompts around:

  • Delete the Camera line and you do not get a neutral camera, you get somebody else's choice. In the watch pair that choice was a different angle entirely — overhead became eye level, which is a re-staged shot rather than a variation on one.
  • Delete the Sound line and you still get sound. Neither clip written without one came back silent. If an unrequested track is a problem for your delivery, the Sound layer is where you head it off.

One thing did not hold, and it is worth knowing before you lean on it. "Quiet room tone only, no music and no voice" is not a volume control. It took 31 dB off the mean level in the cup pair and 3 dB in the watch pair, and the watch clip that asked for quiet still peaked at full scale. Write the Sound layer to say what should be heard. Do not expect it to say how loud.

"The camera does not move" did not hold either. Both clips that asked for a locked frame still drift. In these two pairs the Camera line reliably bought the angle and never bought the lock.

Two pairs is two pairs — enough to report what happened, not enough to settle a rule for every subject. The seeds and settings are above, so the same test runs against your own Wan 3.0 brief in about a minute.

Your Wan 3.0 prompt is rewritten before it runs

One parameter changes how every other rule here behaves, and it is on unless you turn it off. prompt_extend defaults to true, and while it is on, a large language model rewrites what you typed before the video model sees any of it.

Two things follow, and both are worth acting on today:

  • The same prompt run twice is not the same prompt. Two rewrites went in, so two briefs went in. If your second attempt came back with a different composition and you changed nothing, this is usually why.
  • A thin brief gains the most and controls the least. Expanding a short prompt is exactly what the rewriter is for. Give it three words and it invents the other forty. Give it all six layers and there is very little left for it to invent.

So when you are comparing two versions of a shot — two camera moves, two wardrobes — turn the rewriter off first. Otherwise three things changed and you can only see two of them.

A fixed seed will not rescue the comparison either. Alibaba's own note on seed says results may not be perfectly identical even at the same value, so "same prompt, same seed" is two soft claims rather than one hard one.

How to write dialogue Wan 3.0 performs instead of narrating

Sound is generated in the same pass as the picture, which is why a spoken line belongs in the prompt and not in a file you record afterwards. Turning sound off does not make the clip cheaper, so there is nothing to save by leaving it out.

Alibaba publishes the formula:

Voice = Character's lines + Emotion + Tone + Speed + Timbre + Accent

Its own worked example uses double quotes:

He says, "Study hard and make progress every day," in a relaxed tone.

With more than one speaker, name them first and quote them second:

[Character A: Black-suited Agent]
[Black-suited Agent, angrily]: "Where is the truth?"

If you read the braces rule here before, it was too strong

Seven pages on this site used to say braces were the only thing separating a spoken line from narration, and that quotes would get a line read as voice-over. Braces do work. "Only" did not, and neither did the part about quotes — five of the dialogue demos on Alibaba's own release page use no braces anywhere. Corrected 2026-09-04.

Three things actually decide whether a line gets performed, and the bracket you pick is not one of them:

  1. Write the line verbatim. A summary of what someone says comes back narrated. The words you want heard have to appear as the words.
  2. Keep it separate from the description. A line buried inside a sentence about the room reads as part of the room.
  3. Attach it to someone visible in frame. A voice with no body on screen is the single most common way dialogue turns into voice-over.

Pointing your prompt at material you attached

In reference mode the prompt can address an attachment by number — Image 1, Video 1, Audio 1 — and the numbering runs per type, so images and videos each count from one.

Get the form wrong and nothing errors. The request runs, it bills, and the model invents the subject you thought you had supplied. That silence is the expensive part.

One request holds up to 10 reference images, 5 video clips totalling 15 seconds, and 5 audio clips totalling 15 seconds, plus either one document of up to 50 pages or one public web link.

Reference material and first/last frames cannot travel together. A request carrying both is rejected — and it is rejected after queueing, so you wait for the failure. Pick the family before you write: keyframes when you know exactly how the shot opens and closes, reference material when you need a face, a product or a colourway to survive the whole clip.

Three controls Wan 3.0 does not have

Guides currently ranking for this model describe all three. None of them is in the request schema, and writing a prompt that depends on one is how an afternoon disappears.

What you may have readWhat the schema has
A six-shot director modeNo shot-count parameter and no AI Director mode exist. Multi-shot is something the prompt asks for, never something a setting guarantees
Identity Lock, holding a character between sessionsNothing is remembered between generations. There is no cross-session lock; reference material is resent with every request
A camera track you can keyCamera movement is described in words, in the Camera layer, and interpreted per generation

The practical version: the shot count you ask for is a request, not a contract. Describe distinct beats and it will usually cut between them. Ask for one continuous take and it will usually hold. Both are tendencies you are steering, and the full parameter table behind them is on the Wan 3.0 specifications page.

When Wan 3.0 needs thinking turned on

If your prompt points at a document or a web page, enable_thinking has to be true or the model never opens it. The request still succeeds and still bills — it simply generates from your text and ignores the attachment, which looks exactly like a bad prompt and is not one.

It is described as reasoning about composition, staging and motion before the first frame renders, which also makes it worth trying for a sequence of events inside one continuous action. Alibaba's own recommendation is to leave it off when there is no document or link attached. This is documented consistently by four independent API gateways and appears in Alibaba's own request examples, but not yet in the English reference — so treat it as reliable rather than certain. Read 2026-09-02.

Which line of the prompt to change, by what came back

What you gotThe layer that was blankThe edit
The subject drifted by the third beatSubjectRepeat the appearance clause at the end of the prompt, naming hair, clothing and colour
The line was narrated, not spokenSoundQuote it verbatim and attach it to a person on screen
It cut to an angle you never asked forMotionDescribe one continuous action instead of a sequence of beats
Your document or link had no effectThinking was off. Nothing else in the prompt caused this
Nothing resembles your reference photoReferenceCite it as Image 1 and count per type
The bill was four times what you expectedResolution defaults to 1080P when unset. Draft at 480P and the picture is a quarter of the price
Two runs of one prompt look unrelatedThe rewriter was on for both. Turn it off before comparing

Wan 3.0 prompts that already produced a clip

Reading a rule is slower than reading a prompt that worked. Every entry in the Wan 3.0 prompt examples library is printed in full beside the clip it produced, split into these same six layers, with the levers you can move and the edits known to break it. Start there when the shape of your brief is what you are unsure about, and come back here when you know which layer is wrong.

What it costs to get a prompt right

Getting a prompt right takes attempts, and the cheapest place to take them is 480P, where a second costs a quarter of what it costs at 1080P and the sound is identical. Draft the wording at 480P, move to 720P once the beats hold, and spend 1080P only on the take you are keeping.

You can test a corrected prompt here without a card: the first clip is free — 3 seconds at 480P, with sound, on the same wan3.0-video every paid plan runs. Not a trial engine and not a watermark. When one prompt is worth running at full length, plans start at $12.90 a month billed yearly and every one of them reaches all three resolutions and the whole thirty seconds. Take the corrected prompt for a run.

Questions

How do you write a good Wan 3.0 prompt?

Name six things: subject, scene, motion, camera, sound and any reference material. The model fills in whatever you leave out, so an unnamed layer is a layer it decides for you. Length is not the lever — a 40-word prompt with all six named beats a 400-word prompt missing two.

Why is my Wan 3.0 video different from the prompt?

Most often because a language model rewrote the prompt before the video model read it. prompt_extend is on by default, and it expands thin briefs the most. Write all six layers, or switch the rewriter off, and the gap between what you typed and what you got closes.

Does Wan 3.0 need braces around dialogue?

No. Braces work, and so do the double quotes used throughout Alibaba's own examples. What matters is that the line appears verbatim, sits apart from the description, and belongs to somebody visible in frame.

Can I ask Wan 3.0 for a specific number of shots?

You can ask, and there is no parameter that holds it. No shot count and no director mode exist in the schema, so describing distinct beats is a strong hint rather than a setting.

How long can a Wan 3.0 prompt be?

Up to 20,000 characters, in Chinese or English. Past that it is truncated silently rather than refused, which is worth knowing if you paste a whole brief in.

Why did Wan 3.0 ignore my PDF?

Because thinking was off. A document or a link needs enable_thinking set to true; without it the request still runs and still bills, and the attachment is never opened.

Written and maintained by the wan-3.run editorial teamPublished Last updated