What a Wan 3.0 clip actually costs, once you count the attempts
The published rate is $0.05 to $0.20 a second. Four multipliers sit between that and your bill, and one charges you for footage you uploaded, not generated.

Alibaba publishes the rate: $0.05, $0.10 and $0.20 per second of output at 480P, 720P and 1080P. Multiply by seconds and you have the price of one generation, which is a real number and almost never your bill.
Four things sit between the two. Three of them are widely documented and rarely collected in one place; one of them bills you for video you supplied rather than video the model made, and we have not found it written down for creators anywhere. Here is the whole arithmetic, using published list prices only.
| Wan 3.0 | 2 s | 5 s | 10 s | 30 s |
|---|---|---|---|---|
| 480P | $0.10 | $0.25 | $0.50 | $1.50 |
| 720P | $0.20 | $0.50 | $1.00 | $3.00 |
| 1080P | $0.40 | $1.00 | $2.00 | $6.00 |
That is the floor. Now the multipliers.
Every rate and ceiling below is from the Wan 3.0 API reference and the create-task schema, read 2026-08-25. What Alibaba charges is on its own price list; what we charge is on ours.
Multiplier one: the attempts you throw away
Three to five generations per usable clip is the working assumption across this entire category, and Wan 3.0 is not an exception to it. Nobody's published rate includes that, because a rate describes one generation and no brief has ever been finished by one generation.
So the honest unit is cost per keeper, not cost per second. At four attempts, a ten-second 720P clip is $4.00, not $1.00. That number is the one worth budgeting against, and it is also the one you have the most control over — see the loop further down, which cuts it by more than half without changing the output.
The half of this that is not in your control is whether the failures get
charged. A generation that returns FAILED produced nothing, and upstream
providers do not bill for it, so whether it lands on your invoice is a policy
decision made by whoever sits between you and the model.
We refund those automatically;
the full failure taxonomy is here.
Multiplier two: reference video seconds are billed as output
This is the one that produces the confused invoice.
When you attach a reference video, its seconds are billed at your output rate, on top of the seconds you asked for. Every provider we can read states this independently, and the formula they publish is the same one:
billable seconds = input video seconds + requested output secondsRound up to the whole second, multiply by the resolution rate. So a 4.5-second reference clip driving a 5-second 720P output is 10 billable seconds, $1.00, against the $0.50 you budgeted. Exactly double, for a clip that is still five seconds long.
Two things make this less alarming than it first reads. It is only video
references. Reference images, reference audio, documents and web links add
nothing to the bill — you can attach ten images and five audio clips for free,
and only the video family is metered. And it is visible after the fact: the
response carries input_video_duration alongside output_video_duration, which
is precisely the field you reconcile against.
The planning consequence is worth stating plainly, because it inverts the intuition: with Wan 3.0, a fifteen-second reference video is the most expensive input you can attach, and ten reference images are free. If a still frame can carry the same information as a clip — a face, a product, a location, a palette — use the still.
Multiplier three: smart duration reserves the ceiling
Wan 3.0 accepts -1 for duration, meaning you decide how long this should be.
It is a genuinely good feature and a billing hazard, because the price cannot be
known at submit time.
Providers resolve that in two incompatible ways. Some reserve the full thirty seconds, then settle against the actual output and refund the difference. Others simply charge the thirty-second ceiling and keep it. The gap between those two policies on a clip the model decides should run six seconds is $4.80 at 1080P — the difference between $1.20 and $6.00.
There is no way to tell which policy you are on except by reading the terms, so the safe habit is to name the length. Ask for an explicit duration and the number on the button is the number you pay.
Multiplier four: variants multiply, they do not bundle
Four Wan 3.0 variants of the same prompt cost four times one variant. There is no batch discount inside a request, and there is no reason to expect one — the model runs four times.
This is obvious written down and easy to forget in a UI with a "4" chip on it. It is also the fastest way to burn a budget at 1080P: four thirty-second variants is $24.00 of compute to answer a question that three 480P drafts at $4.50 would have answered better, because you would have seen three different ideas instead of four versions of one.
The one setting that quadruples the bill for nothing
If you do not set a resolution, Wan 3.0 generates at 1080P. That is the default, it does not warn you, and it is four times the 480P rate for output you were probably going to re-run anyway. It is the cheapest mistake on this page to fix and the most expensive to leave alone — the other silent ones are here.
Three things that do not change what you pay
| Setting | Effect on price |
|---|---|
| Audio on or off | None. Alibaba's parameter table says so outright: enabling or disabling the audio track does not affect pricing. Sound is generated in the same pass, so there is no compute to save by muting it. Switch it off for creative reasons only |
| Aspect ratio | None. Vertical costs what landscape costs |
| Prompt length | None. Wan 3.0 accepts 20,000 characters and bills output seconds, so a 3,000-word prompt costs exactly what a six-word one costs. Write what the shot needs |
That first row is worth a second of attention, because it removes a decision people agonise over. There is no silent discount. If you want the clip silent, make it silent because silence is right for the cut.
The loop that halves the bill
Every quality tier is the same model. 480P is not a lesser Wan 3.0 with worse prompt comprehension; it is the same generation at a quarter of the pixels and a quarter of the price. Which means drafting cheap costs you nothing except resolution, and the framing, the motion, the pacing and the sound all read fine at 480P.
Two routes to one finished ten-second clip, four attempts either way:
| Route | What you run | Cost |
|---|---|---|
| Straight to final | 4 × 10 s at 1080P | $8.00 |
| Draft, then finish | 3 × 10 s at 480P, then 1 × 10 s at 1080P | $3.50 |
Same four generations, same four learnings, 56% less money. At thirty seconds the gap widens to $24.00 against $10.50, because everything scales linearly with length and the discipline scales with it.
The failure mode people actually have is not overspending on 1080P. It is re-running at 1080P while still deciding what the shot is, which is paying four times for information that arrives identically at a quarter of the price.
What a finished minute costs, in the units the benchmarks use
Converted to the unit public leaderboards use, Wan 3.0 at Alibaba's list price is $3.00 per minute at 480P, $6.00 at 720P and $12.00 at 1080P.
That middle number is the interesting one. On the Artificial Analysis text-to-video arena, With Audio board, snapshot 2026-08-25, Wan 3.0 sits at the top of the table on Elo — and the leaderboard's own price column for it is blank. Filling it in from the published rate puts 720P at $6.00 per minute, which is level with the second-placed model and below the third at $7.80.
Two honesty notes on that, because this is the number most likely to be quoted back at us. Wan 3.0's placement rests on 5,728 votes, the smallest sample in the top five, and its margin over second place is three Elo points against a confidence interval of roughly nine. The board itself gives it a rank range of 1–2. "Joint top, at a lower price than the models around it" is supportable. "Best model, proven" is not, and anyone telling you otherwise has not read the sample column.
The costs that never appear on a rate card
A per-second table can only price the compute. Four things reliably cost more than the compute did, and not one of them shows up as a line item.
Renders you paid for and then lost. Wan 3.0 output links expire 24 hours after a job finishes. A clip that generated perfectly and was never downloaded is a total loss, and it is invisible — it is not a failure, it is a success nobody collected. Whatever you use, finished files need to land somewhere permanent the moment they exist.
Keepers you cannot get back to. Wan 3.0 accepts a seed anywhere from 0 to 2,147,483,647, and the same seed with the same prompt and settings reproduces the same clip. Write down the seed, the exact prompt and the resolution beside every result you like. This is the cheapest habit on this page and the one almost nobody has: without it, finding the good take again is a fresh round of attempts at full price.
Prompts that were quietly rewritten. Expansion is on by default, so the text
that ran is often not the text you wrote. The response hands back orig_prompt
for exactly this reason. If your tooling discards it, you are tuning against a
prompt you never saw, which is how people end up convinced the model is
inconsistent.
Jobs that expired rather than failed. A Wan 3.0 task_id is valid for 24
hours. Query it later and you get UNKNOWN, which is neither success nor failure and
is the one state that automatic refund rules tend to miss.
All four are workflow problems rather than pricing problems, which is why comparing rate cards tells you so little about what a month actually costs. Four questions worth putting to any service, this one included: where do my files live after 24 hours, can I see which prompt actually ran, what happens to a job that expires, and what happens to one that fails. Here the answers are permanent storage the moment a job completes, the model ID and task ID shown against every result, expired jobs labelled expired rather than failed, and failed jobs refunded without a ticket. The failure taxonomy has the detail.
Five questions before you press generate
- Is the resolution set, or is it about to default to the expensive one?
- Is this a draft or a final? If you are unsure, it is a draft, so 480P.
- Is there a reference video attached, and did you add its seconds to the estimate?
- Is the duration explicit, or is
-1about to reserve thirty seconds? - How many variants, and would three different prompts teach you more than four versions of one?
Questions
How much does a 30-second Wan 3.0 video cost?
At Alibaba's published rate, $1.50 at 480P, $3.00 at 720P and $6.00 at 1080P for one generation. Budget three to five generations for a clip you actually ship, and draft the early ones at 480P.
Why was I charged more than the duration I asked for?
Almost certainly a reference video. Its seconds are billed at your output rate on top of the output seconds, rounded up to the whole second. Reference images, audio, documents and links do not add anything.
Does turning audio off make Wan 3.0 cheaper?
No. The parameter table states that toggling audio does not affect pricing, because picture and sound are generated in the same pass. There is no money in muting it.
Is smart duration cheaper than picking a length?
Not usually. Because the length is unknown at submit time, providers either reserve the thirty-second ceiling and refund the difference, or charge the ceiling outright. Naming a duration removes the question.
Does a longer prompt cost more?
No. Wan 3.0 bills seconds of output, not characters of input, and the ceiling is 20,000 characters. Describe the shot properly.
Is Wan 3.0 expensive compared with other models?
Per minute of finished video with sound, 720P works out at $6.00 — the same as the model directly below it on the Artificial Analysis With Audio board and less than the one below that. It is not a budget model and it is not an outlier either.
Written by
Editorial desk
wan-3.run


