Wan 3.0 vs MiniMax H3: one question settles it before the spec table
Both make video with sound; only one you can download and run. If that matters the comparison ends on line one. If not, the numbers are 30 seconds and 30 fps.

We run a Wan 3.0 front-end, so the interest here is obvious. Three things keep this useful anyway: the H3 column comes from the year we spent running a MiniMax H3 tool rather than from anyone's comparison post, H3's advantages get their own section before ours do, and we open with the question that disqualifies us entirely.
Can you live with a model you cannot download? MiniMax H3 has open weights. Wan 3.0 does not, has never had them, and shows no sign of getting them — four Alibaba flagships in a row have shipped closed. If your work involves fine-tuning, offline rendering, a ComfyUI graph or a client who will not send frames to someone else's cloud, H3 wins and nothing below changes that.
If hosted is fine, the comparison turns on three numbers, and two of them are wrong in most tables currently ranking for this question.
The parameter tables, side by side
Every figure from the respective vendors' own documentation, read 2026-08-25: MiniMax's H3 model card and its open-source announcement, against Alibaba's Wan 3.0 API reference.
| Wan 3.0 | MiniMax H3 | |
|---|---|---|
| Duration | 2–30 s, any integer, or model-chosen | 4–15 s, any integer |
| Frame rate | 30 fps | 24 fps |
| Resolutions | 480P / 720P / 1080P | 768P / 2K |
| Aspect ratios | adaptive, 16:9, 4:3, 1:1, 3:4, 9:16 | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
| Prompt ceiling | 20,000 characters | 7,000 characters |
| Reference images | 10 | 9 |
| Reference video | 5 clips, 15 s total | 3 clips |
| Reference audio | 5 clips, 15 s total | 3 clips |
| Documents as input | Yes — ppt, pdf, doc, xls, md, ≤50 pages | No |
| Web page as input | Yes, one public URL | No |
| Edit without regenerating | Yes | No |
| Audio | Generated in the same pass, toggle, same price either way | Generated in the same pass |
| Open weights | No | Yes |
Where MiniMax H3 wins
Weights, and everything downstream of them. This is not one feature, it is a category of work: fine-tuning, LoRAs, offline batch runs, a node in a ComfyUI graph, and the ability to promise a client that nothing left the building. H3 runs on consumer hardware — community reports have it generating on 12 GB of VRAM — which puts it in a different universe from an API-only model. Read the licence before commercial work; it has terms that reach further than the weights themselves.
2K. H3's ceiling is higher than Wan 3.0's 1080P. If the deliverable specifies 2K, that is decided.
21:9. H3 offers the cinema ratio. Wan 3.0 has no 21:9 at all — its widest is 16:9, and there is no crop-and-hope setting hiding in the API. For letterboxed work this alone is disqualifying, and it is the sort of thing you discover at the wrong moment.
No per-second meter. Once the weights are local, the eleventh attempt costs electricity. That is a genuine advantage at volume, and it arrives on the far side of a GPU, an install and a tuning session — the trade is capital and your own time against a metered rate.
Where Wan 3.0 wins
Thirty seconds in one pass. Not thirty seconds assembled from six clips — one continuous generation, 900 frames, no seam to grade across. H3 tops out at fifteen, so any longer piece is a stitching job with matching problems at every join. For a narrative spot with a beginning, a turn and an ending, this is the whole ballgame.
Documents and web pages as input. Hand Wan 3.0 a deck, a PDF, a spreadsheet or a URL and it reads the source and directs a video from it. H3 has no equivalent. It reads the whole document rather than mapping slide three to shot three, so fifty pages becomes a trailer, not a page-turn.
A bigger cast, and a bigger brief. Twenty reference assets against fifteen, and a prompt ceiling of 20,000 characters against 7,000. That second number matters more than it looks: a thirty-second clip needs beats, end states and audio direction written out, and 7,000 characters starts to bind.
Editing without full regeneration. Change a line of dialogue or a detail in the frame and Wan 3.0 revises rather than rerolling the whole clip. On H3 a change is a new generation, at full price and full lottery.
A 480P draft tier. The cheapest tier costs a quarter of the most expensive one, running the identical model. H3's floor is 768P. The cost arithmetic is here.
What neither of them fixes
Both models are close to a coin flip per attempt, and both vendors quote a rate that describes one generation rather than one usable clip. Three to five attempts per keeper is the working number on either side. So the comparison that decides your month is not ten cents a second against thirteen — it is what happens to the attempts you discard.
With local weights, the answer is your electricity bill, and that is a real advantage nobody hosted can match. On any hosted service, including this one, it comes down to two policies: whether a generation that returned nothing still costs you, and whether there is a cheap tier to fail on. The 480P row is one answer to that; whether your provider refunds a failed job is the other, and it is worth checking before you pick either model. The failure modes are here.
The frame rate, since most tables get it wrong
Several comparison pages give Wan 3.0 as 24 fps. It is 30, and the value is
not a matter of interpretation — it comes back in the API response as
usage.fps on every successful job. H3 is 24.
Two practical consequences. Thirty seconds at 30 fps is 900 frames, so a long take carries genuinely more temporal information than the length alone suggests. And if your timeline runs at 30 or 60, Wan 3.0 conforms without a pulldown while 24 fps footage does not — a small thing on a social clip, an annoying thing in a mixed edit.
Neither rate is better in the abstract. 24 is the film convention and people choose it deliberately. The point is only that if you picked a model partly on frame rate, at least half the tables out there gave you the wrong number.
What a finished minute costs, and how to read the leaderboard
Converted to the unit public boards use: Wan 3.0 at Alibaba's published rate is $3.00 per minute at 480P, $6.00 at 720P and $12.00 at 1080P.
On the Artificial Analysis text-to-video arena, With Audio board, snapshot 2026-08-25, Wan 3.0 leads on Elo at 1,240 with MiniMax H3 third at 1,227. The board lists H3 at $7.80 per minute and leaves Wan 3.0's price column blank — filling it from the published rate puts 720P at $6.00, below H3's listed rate.
Now the caveats, which are load-bearing. Wan 3.0's placement rests on 5,728 votes, the smallest sample in the top five; H3 has 8,397. Thirteen Elo points separates them against confidence intervals around nine, and the board gives Wan 3.0 a rank range of 1–2 rather than a rank of 1. "Statistically level with the top of the board, at a lower listed rate" is what the data supports. Anyone writing "beats H3" from this table is reading the Elo column and skipping the sample column.
Three specs to check yourself
| Claim you will see | What the documentation says |
|---|---|
| Wan 3.0 runs at 24 fps | 30 fps, in usage.fps |
| Wan 3.0 takes 12 reference assets, as 9 images + 3 videos + 3 audio | 10 images + 5 videos + 5 audio. Those are H3's numbers, nearly — H3 is 9 + 3 + 3 |
| There is a tier above 1080P | There is no 4K tier. The list is 480P, 720P, 1080P |
That middle row is the tell. When a table gives Wan 3.0 a reference budget of 9 + 3 + 3, it has copied H3's row and changed the heading.
Four briefs, and the answer for each
| The brief | Model | Why |
|---|---|---|
| A 30-second product story, one take, with dialogue | Wan 3.0 | H3 cannot reach the length without seams |
| A six-second loop delivered at 2K | MiniMax H3 | Wan 3.0 stops at 1080P |
| A launch clip from a 20-page deck | Wan 3.0 | Document input has no H3 equivalent |
| Anything offline, fine-tuned, or 21:9 | MiniMax H3 | Weights, and a ratio Wan 3.0 does not have |
The pattern: H3 is the better model to own, Wan 3.0 is the better model to rent. They were released three weeks apart into the same conversation, which is why they get compared, but they are optimised for different relationships with the thing doing the rendering.
Questions
Which is better, Wan 3.0 or MiniMax H3?
Neither, in the abstract. Sort by one question: if you need the model on your own hardware, H3 is the only option of the two. If hosted is acceptable, Wan 3.0 gives you double the duration, document input and in-place editing; H3 gives you 2K and 21:9.
Is Wan 3.0 open source like MiniMax H3?
No. H3 publishes weights; Wan 3.0 is API-only with no checkpoint anywhere. Alibaba's last open-weight video flagship is Wan 2.2, which is a generation behind what you are comparing.
What frame rate does Wan 3.0 output?
30 fps, reported in the API response under usage.fps. MiniMax H3 is 24. Tables
giving Wan 3.0 as 24 have copied it from somewhere rather than from the docs.
Can MiniMax H3 do 30-second clips?
No. H3's range is 4 to 15 seconds. Longer pieces require stitching, which introduces joins you then have to match on grade and motion — the specific problem a 30-second single pass exists to remove.
Which is cheaper?
Per minute of finished video with sound, Wan 3.0 at 720P works out at $6.00 against the $7.80 the Artificial Analysis board lists for H3. Self-hosting H3 changes the question rather than the answer: you swap a metered rate for a GPU, an install and your own hours, which pays off at volume and does not at four clips a month. Sort by how much you actually generate, not by the rate.
Written by
Editorial desk
wan-3.run


