Wan 3.0 document to video: what it does with your deck
Fifty pages become thirty seconds — 0.6 seconds a page. What comes back is a trailer, not a walkthrough, and knowing which one you needed decides the rest.

Wan 3.0 will take a slide deck, a PDF, a spreadsheet or a public web page as input and hand back a video. No other major video model does this, and it is the single most distinctive thing about the model.
It is also the capability people misunderstand fastest, so here is the whole answer in one line: it reads the document and directs a film from what it learned. It does not turn slide three into shot three. At the ceiling that is fifty pages compressed into thirty seconds — 0.6 seconds a page — which is arithmetic that only ever produces a trailer.
If a trailer is what you wanted, Wan 3.0 is remarkable at it. If you wanted your slides on screen with a voice reading them, this is the wrong tool and the rest of this article will save you the credits.
Document and web-link input is a documented media type rather than a marketing
claim: see input.media.type in the create-task schema, and the
Wan 3.0 API reference. Read 2026-08-25.
The limits, in one table
Every figure here is from the Wan 3.0 API reference, checked 2026-08-25.
| Value | |
|---|---|
| Formats | docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers, md |
| Per request | 1 file, or 1 link — never both, never several |
| File size | 100 MB |
| Page count | 50 |
| Web link | One public page, no login required |
| Output | Up to 30 seconds, 480P / 720P / 1080P |
Two of those catch people out. The page ceiling is a page count, not a proxy for file size, so a slim 80-page PDF fails while a heavy 40-page deck passes. And one file or one link means you cannot hand it the deck and the product page together, which is exactly the pairing most people reach for first.
There is a third, less obvious constraint. The document input belongs to Wan 3.0's reference family, alongside reference images, video and audio — and that family is mutually exclusive with first-frame and last-frame inputs. You cannot supply a deck and pin the opening frame. The request fails before generating. The full family rule is here.
What it actually does with the pages
Wan 3.0 reads the whole document, works out what it is about, and directs a video from that understanding. Practically, that has three consequences worth planning around.
Structure is interpreted, not preserved. Your section order is a signal, not a storyboard. If the argument only makes sense in your order, say so in the prompt — because you can send a prompt and a document in the same request, and the prompt is where you get to steer.
Density gets flattened. Thirty seconds cannot carry fifty pages of argument, so the model selects. What it selects from a dense deck is not always what you would have selected, which is the main reason first attempts disappoint.
Numbers deserve suspicion. Data-heavy sources — spreadsheets, financial decks, anything where a digit matters — are the least predictable input here. Check every figure that appears on screen before the clip goes anywhere near a client. This is not a knock on the model; it is the same care you would apply to any generated frame containing text.
Preparing a document that produces a decent video
This is the part no announcement covers, and it makes more difference than the prompt does.
Cut it down before you upload. Fifty pages is the ceiling, not a target. A twelve-page version of the same deck reliably produces a better thirty seconds than the full one, because you have done the selecting rather than delegating it. Strip appendices, backup slides, the methodology section and the team page.
Put the message in the first third. Given a long document the model weights early material heavily. If your conclusion lives on the last slide, move it.
One document, one idea, one video. A deck that covers a product launch and a hiring pitch produces a clip that commits to neither. Split the file.
Always send a prompt with the file. Wan 3.0 accepts both in one request: the document supplies the substance, the prompt supplies the treatment. This is the highest-leverage habit in the whole workflow, and a short one works:
A 20-second product launch film built from the attached deck.
Take the product name, the three headline capabilities and the launch date.
Ignore pricing, the roadmap and anything about the team.
Style: clean studio product photography, cool neutral palette, one slow
push-in per beat. Confident voiceover, light ambient bed, no music sting.Note what that prompt is doing: it names what to take and what to ignore. Exclusions do more work than descriptions here, because your document contains far more than thirty seconds can hold and the default selection is the model's, not yours.
Give the link the same treatment. A public URL works the same way, and the same rule applies — a product page produces a better clip than a homepage, because a homepage is about eleven things.
Which source suits which output
Wan 3.0 accepts all of these, but not equally well.
| Source | Works well for | Watch out for |
|---|---|---|
| PPT / Keynote | Launch films, brand pieces, conference teasers | Speaker notes carry meaning your slides do not; paste the important ones into the prompt |
| PDF report | Narrated briefings, summary clips | Long-form prose flattens hardest — cut to the executive summary |
| Spreadsheet | Trend and headline-figure pieces | Verify every number on screen. Send the summary sheet, not the raw data |
| Markdown / TXT | Scripts, structured briefs | The cleanest input of the lot, because there is no layout to interpret |
| Web link | Product pages, launch posts, changelogs | One page only, public, no login. Pick a page with one subject |
When this is the wrong tool
Worth being direct, because the search that brings people here is often looking for something else entirely.
If what you need is your slides on screen, in your layout, with a narrator reading them — training modules, compliance decks, onboarding, anything multilingual — that is a different product category. Tools built for narrated slides preserve your design and add a voice or a presenter. Wan 3.0 does the opposite: it discards the layout and generates new footage. Neither is better; they answer different briefs, and picking the wrong one wastes a day.
The rule of thumb: if the deck's design is the deliverable, do not use a generative model. If the deck is source material for something that will look nothing like it, this is the tool that has no real competition.
What it costs
Wan 3.0 charges nothing for the document itself. You are billed for output seconds at the resolution you pick — nothing is added for the file, its page count or its size, and the same is true of a web link. Only reference video seconds get added to a bill, and a document job has none.
So a twenty-second launch clip from a fifty-page deck costs exactly what a twenty-second clip from a text prompt costs: $1.00 at 480P, $2.00 at 720P, $4.00 at 1080P on Alibaba's published rate.
Which points at the obvious first move. Run the deck at 480P first. Document jobs miss more often than prompt jobs on the first attempt, because Wan 3.0 is making editorial choices you have not watched it make before. Watching one 480P pass tells you what it took from your document, and then you write the prompt that corrects it. The cost arithmetic in full.
Six checks before you upload
- Under 50 pages and under 100 MB.
- One file or one link — not both.
- No first-frame or last-frame image in the same request.
- Appendices, backup slides and the team page removed.
- The conclusion moved into the first third.
- A prompt attached that names what to take and what to ignore.
Questions
What file types can Wan 3.0 turn into video?
docx, doc, xlsx, xls, pptx, ppt, pdf, txt, key, pages, numbers and md, up to 50 pages and 100 MB — or one public web page instead. One file or one link per request, never both.
Does Wan 3.0 turn each slide into a scene?
No, and this is the most common misunderstanding. It reads the whole document and directs a video from what it learned. Fifty pages into thirty seconds is 0.6 seconds a page, so the output is a trailer rather than a page-turn.
Can I use a document and a first-frame image together?
No. Files and links sit in the reference family, which is mutually exclusive with first-frame and last-frame inputs. The request is rejected before anything generates.
Does a bigger document cost more?
No. Wan 3.0 bills output seconds only. A fifty-page deck and a one-line prompt cost the same for the same length of clip, and a web link adds nothing either.
Why did it ignore the most important part of my deck?
Usually because it was late in the document and the prompt did not name it. Move the conclusion forward, cut the pages that do not serve the clip, and tell the prompt explicitly what to take and what to ignore.
Can it make a training video from our manual?
It can make a thirty-second piece about the manual. If you need the procedure taught step by step, in your layout, with narration, that is narrated-slide tooling rather than a generative video model — and the mismatch is worth catching before you spend anything.
The cheapest way to learn what Wan 3.0 takes from your deck is to watch it happen once at 480P. Send a document through, read what came back, then write the prompt that corrects it.
Written by
Editorial desk
wan-3.run


