AI Explainer Video

The brief is never "make it exciting". It is: explain the thing correctly, in our language, without overclaiming, and get it past the person who has to sign it off. Write the shot, quote the line, and Wan 3.0 generates the picture and the narration together. Your first clip is free — an email, no card.

Start here

Create image

JPG · PNG · BMP · WEBP · ≤20MB · 240–8000px · ratio ≤8:1

0 / 20000

The subject keeps the same face, hair, clothing and colours throughout; only the camera and the described motion change.

Model
Aspect ratio
Duration
Resolution
免费注册 · 送一条 Wan 3.0 片子

免费那条就是同一个 Wan 3.0,声音也有,上限是 480P · 3s,走的是共用队列。方案能把上限抬高到 1080P 和三十秒 —— 免费档是一个分辨率,不是另一个模型。

1 free clip on the real model · an email, no card · the narration is generated with the picture — there is no voice file to make first

Four of Wan 3.0’s own published explainer runs, each printed with the whole prompt beside the clip it produced · Try this loads the prompt

An animated quantum-mechanics explainer: a presenter beside a glowing orb under a Chinese title cardFOUR SHOTS · ASKED FOR 120s · CAME BACK 30s

A subject explainer whose timeline was four times too long

Picture and voice, one pass

What an AI explainer video is made of here

No script upload, no voice library, no timeline. You write one shot at a time and quote what is said in it, and the clip comes back with the words already spoken.

An opening frame from a generated explainer shot, title card and presenter together

One shot at a time

Picture, performance and narration generated together

Every other explainer tool asks you to pick an avatar, paste a script and wait for a render. Wan 3.0 takes a written shot — what is on screen, what the camera does, and the line in quotation marks — and generates the picture and the speech in the same pass, as one MP4. Write in English, Chinese, Japanese or Korean. A full film is that, repeated: there is no shot-count parameter, so a two-minute explainer is twelve to twenty separate generations you cut together. The absence is checkable rather than a complaint — Alibaba's API reference lists every parameter the model takes, and a shot count is not among them; the ceiling on one request is 30 seconds.

The narration, quoted
In-prompt

The narration, quoted

Prompt and spoken line
4 languages

Prompt and spoken line

Per shot, at 30 fps
2–30 s

Per shot, at 30 fps

Generations in a 2-minute film
≈ 12–20

Generations in a 2-minute film

  • No script-to-film button — you write and cut the shots yourself
  • No avatar library — the presenter is generated, or it is your own reference image
  • No subtitle track — ask for no on-screen text and add captions in your editor

Two ways in

Start from a document, or start from the shot list

Most explainers already exist somewhere — as a deck, a spec sheet or a product page. Whether you hand that over or write the shots yourself changes who is deciding the story.

  • A frame from an explainer the model laid out itself, from source material it was handedfile · link
    Hand over the sourceDocument or linkA deck, a report or a PDF up to fifty pages, or one public web page. Wan 3.0 reads it and decides what is worth showing. Fast, and the model picks the highlights — which is fine for a summary and wrong for a compliance point you need on screen.
  • A frame from a shot written by hand, with its own start and end secondone shot, one line
    Write the shotsThis pageOne generation per shot, each with its own frame, its own action and its own quoted line. Slower, and you decide every claim that appears. This is the route the accounts making certification and safety films here actually use.

Hand over the source when the film is a summary and nobody will audit it: document to video takes fifty pages, URL to video takes a public page.

Write the shots when a specific sentence has to be said, a specific mark has to appear, or a specific thing must not be shown. That is slower and it is the only route where you control the claim. If a product, mascot or certification mark has to survive the whole cut, build the clause on character consistency first.

The constraints nobody writes down

Six things that sink a corporate explainer video

These are not aesthetic problems. Each one is the reason a finished film goes back for another round with the person who has to approve it.

  • It has to warn without showing harm

    A safety film cannot depict the accident, and "do not show X" is a weak instruction — naming a thing puts it in the model's head

    Describe the test, not the failure: the rig, the gauge, the hands, the load holding. Write what is in frame rather than what is banned

    Prompt structure
  • It sounds like an advertisement

    The default register of every video model is bright, warm and upbeat, because that is what most of its material is

    State the register as a constraint alongside the picture: "documentary, matte surfaces, low-saturation palette, no music, restrained delivery". Tone is a prompt field like any other

    Prompt structure
  • The certification mark or logo comes out wrong

    Marks are small, high-contrast and legally exact — three properties generation is worst at

    Hold it with the appearance clause, keep the camera still on that shot, and composite the real mark afterwards if it must be pixel-correct. Certification marks are registered trademarks with published display rules — Japan's SG mark is a worked example — so a generated near-copy is a different object from the mark you are licensed to show, however convincing the frame looks

    Read the frame before publishing
  • The film is shorter than the timeline you wrote

    A shot list is only obeyed when its seconds add up to something the model can render. Past 30 seconds the timeline is not rejected — it is compressed, silently, and every beat gets less time than you gave it

    Make the last number in your shot list 30 or less, and say the total out loud in the first line. All four examples at the top of this page are Wan’s own, and they split two-two: the two that asked for 120s and 150s came back at 30.0s and 25.0s with four and five shots crushed into them, while the two that declared 30 seconds up front — one of them opening “This video contains 6 shots, with a total duration of 30.0 seconds” — came back at 30.0s and 30.1s with their beats intact. Play them against the prompts printed beside them; the arithmetic is on the page

    Prompt structure
  • Shot seven does not match shot two

    There is no shot-count parameter and no memory between generations. A film is separate requests that have never met

    Fix the palette, the lens and the light in a style block you paste into every prompt, and send the same reference images each time. Reference to video

  • The narration is in the wrong register for the language

    A line translated from English keeps English cadence, and a Japanese or Korean reviewer hears it immediately

    Write the spoken line in the target language yourself and keep only the appearance clause in English. That combination is what the accounts producing non-English branded work here use

A run that fails outright is refunded automatically. None of the five above fails — each comes back finished and unusable, which is billed. Draft every shot at 480P until the register and the constraints hold, then re-run the keepers.

Pricing · 1080P on every paid plan · Native audio

What a two-minute explainer video actually costs

Do the arithmetic before you promise a date. A two-minute film is roughly twelve to twenty shots, each taking three or four generations to get right — so plan for forty to eighty runs, not twenty. Draft at 480P, where an iteration costs a quarter of a keeper, and re-run only the shots that made the cut. A failed run is refunded automatically, narration costs nothing extra, and commercial use is included on every paid plan.

How this is billed

Billed by
output second
Narration
no extra cost
Document or link
no surcharge
Your first clip
free · 480P

The whole film at 480P first is a storyboard you can actually watch, and it costs a quarter of the finished one. Full pricing.

两种付法,每月积分一样

  • 免费的第一条不用卡

    在真正的 wan3.0-video 上生成一次。拿你自己的镜头看看 Wan 3.0 会做成什么样 — 480P,三秒,声音同一次出。

    $0一次性

    全程不用卡

    生成免费的那条

    一个邮箱,不用卡

    1 条,仅此一次

    480P · 最长 3 秒 · 带声音
    一个邮箱,不用卡

    和付费套餐跑的是同一个 `wan3.0-video` — 480P 是一个分辨率,不是一个缩水的模型

    • 真正的 wan3.0-video,不是 2.7
    • 声音和画面一起生成
    • 收据上有 model ID 和 task ID
    • 跑失败永远不扣你的
    • 720P 与 1080P付费
    • 超过 3 秒的片子付费
    • 参考、文档与编辑模式付费
    • 第二条以及以后付费

    只回答一个问题:Wan 3.0 会把你的想法做成什么样。1080P、三十秒和全部输入模式,从 $12.90 起。

    一眼看完

    每月条数
    1 条,仅此一次
    每次出几条
    1
    最长时长
    3s
    最高分辨率
    480P
  • Starter最多省 20%

    一周一条发布贴。够你摸清 Wan 3.0 AI 视频生成器合不合你的工作方式,而不用先对用量下判断。

    $12.90/ 月$15.90

    按年收 $154.80

    随时取消

    640 积分 / 月 · 约 16 条

    按 5 秒 480P 算 · 720P 8 条 · 1080P 4 条

    套餐积分在当月月底过期

    • 每月 640 积分 — 大约 16 条成片
    • 同一条提示词每次出 2 条
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    前两行以下的一切,$12.90 这一档就都有。更大的套餐买的是秒数和条数 — 从来不是更好的模型。

    一眼看完

    每月条数
    5 秒 480P 下约 12 条
    每次出几条
    2
    每美元积分
    基准
    退款窗口
    7 天
  • 多数人落在这里
    Pro

    每个工作日一条成片,每条还有三个备选。480P 打草稿、1080P 定稿的那个循环,按一个工作周的量配好。

    $39.90/ 月$49.90

    按年收 $478.80

    随时取消

    2,240 积分 / 月 · 约 56 条

    按 5 秒 480P 算 · 720P 28 条 · 1080P 14 条

    套餐积分在当月月底过期

    • 每月 2,240 积分 — Starter 的 3.5 倍
    • 每月约 56 条成片,或者 14 条满 1080P
    • 同一条提示词每次出 4 条 — 挑一条,而不是重摇一次2× 条数
    • 每美元积分比 Starter 多 13%更划算
    • 同一周内既够 480P 打草稿,也够 1080P 定稿
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    从 Starter 往上,涨的是量、条数和单价。模型、分辨率和那三十秒,本来就已经是你的了。

    一眼看完

    每月条数
    5 秒 480P 下约 44 条
    每次出几条
    4
    每美元积分
    比 Starter 多 21%
    退款窗口
    7 天
  • Studio最划算

    接客户量的档,也是你不再省着用 1080P 的那一档。每月三十一条满分辨率成片,单位积分是我们卖过最低的价。

    $99.90/ 月$119.90

    按年收 $1,198.80

    随时取消

    6,240 积分 / 月 · 约 156 条

    按 5 秒 480P 算 · 720P 78 条 · 1080P 39 条

    套餐积分在当月月底过期

    • 每月 6,240 积分 — Starter 的 9.75 倍,Pro 的 2.8 倍13×
    • 每月约 156 条成片,或者 39 条满 1080P
    • 每美元积分多 26% — 我们卖过最划算的最划算
    • 1080P 多到你不必先用 480P 打草稿
    • 同一条提示词每次出 4 条
    • 按一整份客户名单配的量,不是一个频道
    • 任何一个月都能加购,不用换套餐
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    这是为一整份客户名单配的档,不是为一个频道:每条提示词出四条、1080P 不用省着用,还留得出重拍一场戏的余量,不用盯着余额往下掉。

    一眼看完

    每月条数
    5 秒 480P 下约 124 条
    满 1080P 的话
    每月约 31 条
    每美元积分
    比 Starter 多 28%
    退款窗口
    7 天

The whole procedure

How to make an AI explainer video shot by shot

Four steps, repeated once per shot. The first one is written in a document, not in this generator — and skipping it is why most attempts stall at shot four.

  1. Write the shot list before you generate anything

    One line per shot: what is on screen, what is said, how long. Twelve rows for a two-minute film. This is the part the tool cannot do for you, and it is the part that makes the rest cheap.

    Do this first
  2. Fix a style block and reuse it verbatim

    Palette, lens, light, surface finish, register, and "no on-screen text". Paste the identical block into every prompt — it is the only thing holding twenty independent generations together.

    Copy into every prompt
  3. Generate one shot, with its line in quotation marks

    The narration is produced with the picture, so there is nothing to record and nothing to sync. Write the line in the language it will be heard in.

    Voice in the same pass
  4. Draft the whole film at 480P, then re-run the keepers

    Cut the drafts together and watch it end to end before you spend anything on resolution. Most re-writes are structural, and a 480P storyboard finds them for a quarter of the price.

    Up to 1080P

The two-thirds rule from the corporate films made here: the demonstration is the film. Openers and closers are cheap to generate and cheap to cut — the shots that earn the approval are the ones showing the thing actually being done.

Questions people actually ask

AI explainer video questions

Whether it writes the script, how the voice works, how long a film really takes, and what you can publish

Does it turn a script into a finished explainer automatically?

No, and any tool promising that on this model is describing something Wan 3.0 does not have. There is no shot-count parameter and no director mode. You generate one shot at a time — 2 to 30 seconds each — and cut them together. A two-minute explainer is twelve to twenty generations. If a rough summary is enough, document to video will read a fifty-page deck and decide the highlights for you.

Where does the voice-over come from?

It is generated alongside the picture, in the same pass, from the line you put in quotation marks in the prompt. There is no voice library to choose from, no audio file to upload and nothing to sync afterwards. Audio does not change the price.

Can the narration be in Japanese, Chinese or Korean?

Yes. Wan 3.0 guarantees prompts in English and Chinese and reads others. Write the spoken line in the target language rather than translating an English one afterwards — cadence survives translation badly and a native reviewer hears it immediately. Keep any appearance clause in English.

How do I stop it looking like an advert?

Say so, as a constraint next to the picture: documentary register, matte surfaces, low-saturation palette, no music, restrained delivery, no on-screen text. The default of every video model is bright and upbeat because most of its training material is. Tone is a prompt field, and it responds to one.

How do I keep twenty shots looking like one film?

A style block pasted verbatim into every prompt, plus the same reference images sent with each generation. Nothing carries between requests on its own. Matching is close rather than exact, which is usually enough at cut speed — character consistency explains what holds and what does not.

Can it show a warning without showing the accident?

That is the standard safety-film constraint, and it is solved by what you write rather than what you ban. Describe the test instead of the failure: the rig, the load, the gauge, the hands, the thing holding. Naming what must not appear tends to put it in frame.

What does a two-minute explainer cost in credits?

Plan for forty to eighty generations rather than twenty: twelve to twenty shots, three or four attempts each. Draft everything at 480P — a quarter of the price of a finished clip — cut it together, then re-run only the keepers at 720P or 1080P. A run that fails outright is refunded automatically.

Can I publish this on our company site?

Yes. Commercial use is included on every paid plan and every credit pack — you may publish, sell and license the output. Free-tier clips are for evaluation. You remain responsible for the claims in the finished film, and responsible use covers what may not be generated at all.

Do I have to disclose that the explainer was made with AI?

On your own site that is your call and your legal team's. On a platform it is a written rule, and an explainer is the format most likely to be caught by it: YouTube's disclosure requirement covers content that looks real — a realistic scene that never happened, a person who appears to say something they did not — which is exactly what a generated presenter in a generated office is. A photorealistic explainer needs the AI-use toggle set; a fully animated one does not. The penalty for consistently not disclosing is not a warning label, it is removal of the content or suspension from the Partner Program. Everything about the film can be true and the disclosure still be owed, because the rule is about how it was made rather than whether it is honest.

Write one shot. Hear it back in ninety seconds.

Before you budget a film, generate its hardest shot — the one with the mark, the constraint and the line that has to be right. First clip free at 480P, an email and no card. If the register is wrong, you will know today rather than at the review.

由 wan-3.run 编辑团队撰写与维护发布于 最后更新 工具 版本 2026.09.4