AI Explainer Video

The brief is never "make it exciting". It is: explain the thing correctly, in our language, without overclaiming, and get it past the person who has to sign it off. Write the shot, quote the line, and Wan 3.0 generates the picture and the narration together. Your first clip is free — an email, no card.

Start here

Create image

JPG · PNG · BMP · WEBP · ≤20MB · 240–8000px · ratio ≤8:1

0 / 20000

The subject keeps the same face, hair, clothing and colours throughout; only the camera and the described motion change.

Model
Aspect ratio
Duration
Resolution
무료 회원가입 · Wan 3.0 클립 1편

무료 클립은 소리까지 같은 Wan 3.0이고 480P · 3s로 제한되며 공용 대기열을 씁니다. 요금제가 상한을 올려 1080P와 30초로 갑니다 — 무료 등급은 해상도이지 다른 모델이 아닙니다.

1 free clip on the real model · an email, no card · the narration is generated with the picture — there is no voice file to make first

Four of Wan 3.0’s own published explainer runs, each printed with the whole prompt beside the clip it produced · Try this loads the prompt

An animated quantum-mechanics explainer: a presenter beside a glowing orb under a Chinese title cardFOUR SHOTS · ASKED FOR 120s · CAME BACK 30s

A subject explainer whose timeline was four times too long

Picture and voice, one pass

What an AI explainer video is made of here

No script upload, no voice library, no timeline. You write one shot at a time and quote what is said in it, and the clip comes back with the words already spoken.

An opening frame from a generated explainer shot, title card and presenter together

One shot at a time

Picture, performance and narration generated together

Every other explainer tool asks you to pick an avatar, paste a script and wait for a render. Wan 3.0 takes a written shot — what is on screen, what the camera does, and the line in quotation marks — and generates the picture and the speech in the same pass, as one MP4. Write in English, Chinese, Japanese or Korean. A full film is that, repeated: there is no shot-count parameter, so a two-minute explainer is twelve to twenty separate generations you cut together. The absence is checkable rather than a complaint — Alibaba's API reference lists every parameter the model takes, and a shot count is not among them; the ceiling on one request is 30 seconds.

The narration, quoted
In-prompt

The narration, quoted

Prompt and spoken line
4 languages

Prompt and spoken line

Per shot, at 30 fps
2–30 s

Per shot, at 30 fps

Generations in a 2-minute film
≈ 12–20

Generations in a 2-minute film

  • No script-to-film button — you write and cut the shots yourself
  • No avatar library — the presenter is generated, or it is your own reference image
  • No subtitle track — ask for no on-screen text and add captions in your editor

Two ways in

Start from a document, or start from the shot list

Most explainers already exist somewhere — as a deck, a spec sheet or a product page. Whether you hand that over or write the shots yourself changes who is deciding the story.

  • A frame from an explainer the model laid out itself, from source material it was handedfile · link
    Hand over the sourceDocument or linkA deck, a report or a PDF up to fifty pages, or one public web page. Wan 3.0 reads it and decides what is worth showing. Fast, and the model picks the highlights — which is fine for a summary and wrong for a compliance point you need on screen.
  • A frame from a shot written by hand, with its own start and end secondone shot, one line
    Write the shotsThis pageOne generation per shot, each with its own frame, its own action and its own quoted line. Slower, and you decide every claim that appears. This is the route the accounts making certification and safety films here actually use.

Hand over the source when the film is a summary and nobody will audit it: document to video takes fifty pages, URL to video takes a public page.

Write the shots when a specific sentence has to be said, a specific mark has to appear, or a specific thing must not be shown. That is slower and it is the only route where you control the claim. If a product, mascot or certification mark has to survive the whole cut, build the clause on character consistency first.

The constraints nobody writes down

Six things that sink a corporate explainer video

These are not aesthetic problems. Each one is the reason a finished film goes back for another round with the person who has to approve it.

  • It has to warn without showing harm

    A safety film cannot depict the accident, and "do not show X" is a weak instruction — naming a thing puts it in the model's head

    Describe the test, not the failure: the rig, the gauge, the hands, the load holding. Write what is in frame rather than what is banned

    Prompt structure
  • It sounds like an advertisement

    The default register of every video model is bright, warm and upbeat, because that is what most of its material is

    State the register as a constraint alongside the picture: "documentary, matte surfaces, low-saturation palette, no music, restrained delivery". Tone is a prompt field like any other

    Prompt structure
  • The certification mark or logo comes out wrong

    Marks are small, high-contrast and legally exact — three properties generation is worst at

    Hold it with the appearance clause, keep the camera still on that shot, and composite the real mark afterwards if it must be pixel-correct. Certification marks are registered trademarks with published display rules — Japan's SG mark is a worked example — so a generated near-copy is a different object from the mark you are licensed to show, however convincing the frame looks

    Read the frame before publishing
  • The film is shorter than the timeline you wrote

    A shot list is only obeyed when its seconds add up to something the model can render. Past 30 seconds the timeline is not rejected — it is compressed, silently, and every beat gets less time than you gave it

    Make the last number in your shot list 30 or less, and say the total out loud in the first line. All four examples at the top of this page are Wan’s own, and they split two-two: the two that asked for 120s and 150s came back at 30.0s and 25.0s with four and five shots crushed into them, while the two that declared 30 seconds up front — one of them opening “This video contains 6 shots, with a total duration of 30.0 seconds” — came back at 30.0s and 30.1s with their beats intact. Play them against the prompts printed beside them; the arithmetic is on the page

    Prompt structure
  • Shot seven does not match shot two

    There is no shot-count parameter and no memory between generations. A film is separate requests that have never met

    Fix the palette, the lens and the light in a style block you paste into every prompt, and send the same reference images each time. Reference to video

  • The narration is in the wrong register for the language

    A line translated from English keeps English cadence, and a Japanese or Korean reviewer hears it immediately

    Write the spoken line in the target language yourself and keep only the appearance clause in English. That combination is what the accounts producing non-English branded work here use

A run that fails outright is refunded automatically. None of the five above fails — each comes back finished and unusable, which is billed. Draft every shot at 480P until the register and the constraints hold, then re-run the keepers.

Pricing · 1080P on every paid plan · Native audio

What a two-minute explainer video actually costs

Do the arithmetic before you promise a date. A two-minute film is roughly twelve to twenty shots, each taking three or four generations to get right — so plan for forty to eighty runs, not twenty. Draft at 480P, where an iteration costs a quarter of a keeper, and re-run only the shots that made the cut. A failed run is refunded automatically, narration costs nothing extra, and commercial use is included on every paid plan.

How this is billed

Billed by
output second
Narration
no extra cost
Document or link
no surcharge
Your first clip
free · 480P

The whole film at 480P first is a storyboard you can actually watch, and it costs a quarter of the finished one. Full pricing.

어느 쪽이든 월 크레딧은 같습니다

  • 무료 첫 클립카드 없이

    진짜 wan3.0-video로 생성 한 번. 당신의 숏으로 Wan 3.0이 무엇을 하는지 보세요 — 480P, 3초, 소리도 같은 패스에서.

    $0한 번

    어느 단계에서도 카드 없음

    무료 클립 생성하기

    이메일만, 카드는 필요 없음

    클립 1편, 한 번만

    480P · 최대 3초 · 소리 포함
    이메일만, 카드는 필요 없음

    알리바바가 유료 플랜에 돌리는 것과 같은 `wan3.0-video`입니다 — 480P는 해상도이지 낮은 모델이 아닙니다

    • 2.7이 아니라 진짜 wan3.0-video
    • 소리를 그림과 함께 생성
    • 영수증에 찍히는 모델 ID와 작업 ID
    • 실패한 생성은 절대 과금되지 않음
    • 720P와 1080P유료
    • 3초보다 긴 클립유료
    • 레퍼런스·문서·편집 모드유료
    • 클립 한 편을 넘겨서, 계속유료

    여기서 답하는 질문은 하나입니다. 당신의 아이디어로 Wan 3.0이 무엇을 하는가. 1080P와 30초, 그리고 모든 입력 모드는 $12.90부터입니다.

    한눈에

    월 클립 수
    1편, 한 번만
    생성당 테이크 수
    1
    가장 긴 클립
    3s
    최상위 해상도
    480P
  • Starter최대 20% 절약

    주에 출시 게시물 하나. Wan 3.0 AI 동영상 생성기가 당신의 작업 방식에 맞는지, 분량을 결정하지 않고도 알아보기에 충분합니다.

    $12.90/ 월$15.90

    연 $154.80 청구

    언제든 해지

    월 640 크레딧 · ≈ 16편

    480P 5초 기준 · 720P면 8편 · 1080P면 4편

    플랜 크레딧은 월말에 소멸합니다

    • 월 640 크레딧 — 완성된 클립 약 16편
    • 한 번 실행에 같은 프롬프트 테이크 2개
    • 모든 해상도 — 480P, 720P, 1080P
    • 모든 길이 — 한 번에 2초부터 30초까지
    • 모든 입력 — 텍스트, 사진, 레퍼런스, 문서, 웹페이지
    • 소리는 같은 패스에서, 워터마크 없음
    • 상업적 이용 포함 — 공개하고, 팔고, 청구해도 됩니다
    • 고속 등급 Wan 3.0 Prime, 원할 때 언제든
    • 실패한 실행은 크레딧을 쓰지 않습니다
    • 크레딧을 쓰지 않았다면 7일 환불
    • 두 번 클릭으로 해지 — 가입할 때의 가격은 그대로 잠깁니다

    위 두 줄 아래에 있는 것은 전부 $12.90짜리 이 플랜에 들어 있습니다. 더 큰 플랜이 사는 것은 초와 테이크이지, 더 좋은 모델이 아닙니다.

    한눈에

    월 클립 수
    480P 5초로 ≈ 12편
    생성당 테이크 수
    2
    1달러당 크레딧
    기준선
    환불 기간
    7일
  • 대부분 여기로 옵니다
    Pro

    근무일마다 완성된 클립 하나씩, 각각 대안 세 개까지. 480P로 초안 잡고 1080P로 마무리하는 방식을, 한 주 분량으로 맞춘 크기입니다.

    $39.90/ 월$49.90

    연 $478.80 청구

    언제든 해지

    월 2,240 크레딧 · ≈ 56편

    480P 5초 기준 · 720P면 28편 · 1080P면 14편

    플랜 크레딧은 월말에 소멸합니다

    • 월 2,240 크레딧 — Starter의 3.5배
    • 월 완성 클립 ≈ 56편, 1080P로만 하면 14편
    • 생성 한 번에 같은 프롬프트 테이크 4개 — 다시 돌리는 대신 골라 쓰세요테이크 2배
    • Starter보다 1달러당 크레딧 +13%더 나은 단가
    • 같은 주 안에 480P로 초안 잡고 1080P로 마무리할 여유
    • 모든 해상도 — 480P, 720P, 1080P
    • 모든 길이 — 한 번에 2초부터 30초까지
    • 모든 입력 — 텍스트, 사진, 레퍼런스, 문서, 웹페이지
    • 소리는 같은 패스에서, 워터마크 없음
    • 상업적 이용 포함 — 공개하고, 팔고, 청구해도 됩니다
    • 고속 등급 Wan 3.0 Prime, 원할 때 언제든
    • 실패한 실행은 크레딧을 쓰지 않습니다
    • 크레딧을 쓰지 않았다면 7일 환불
    • 두 번 클릭으로 해지 — 가입할 때의 가격은 그대로 잠깁니다

    Starter에서 올라오며 달라지는 것은 수량과 테이크와 단가입니다. 모델과 해상도와 30초는 이미 당신 것이었습니다.

    한눈에

    월 클립 수
    480P 5초로 ≈ 44편
    생성당 테이크 수
    4
    1달러당 크레딧
    Starter 대비 +21%
    환불 기간
    7일
  • Studio최고 단가

    클라이언트 물량, 그리고 1080P를 아껴 쓰지 않아도 되는 플랜. 월 31편을 최고 해상도로 완성할 수 있고, 저희가 파는 크레딧 중 가장 싼 단가입니다.

    $99.90/ 월$119.90

    연 $1,198.80 청구

    언제든 해지

    월 6,240 크레딧 · ≈ 156편

    480P 5초 기준 · 720P면 78편 · 1080P면 39편

    플랜 크레딧은 월말에 소멸합니다

    • 월 6,240 크레딧 — Starter의 9.75배, Pro의 2.8배13×
    • 월 완성 클립 ≈ 156편, 또는 1080P로 39편
    • 1달러당 크레딧 +26% — 저희가 파는 것 중 가장 좋은 단가최고 단가
    • 480P로 먼저 초안 잡을 필요가 없어질 만큼의 1080P
    • 한 번 실행에 같은 프롬프트 테이크 4개
    • 채널 하나가 아니라 클라이언트 여럿에 맞춘 크기
    • 플랜을 바꾸지 않고 아무 달에나 충전
    • 모든 해상도 — 480P, 720P, 1080P
    • 모든 길이 — 한 번에 2초부터 30초까지
    • 모든 입력 — 텍스트, 사진, 레퍼런스, 문서, 웹페이지
    • 소리는 같은 패스에서, 워터마크 없음
    • 상업적 이용 포함 — 공개하고, 팔고, 청구해도 됩니다
    • 고속 등급 Wan 3.0 Prime, 원할 때 언제든
    • 실패한 실행은 크레딧을 쓰지 않습니다
    • 크레딧을 쓰지 않았다면 7일 환불
    • 두 번 클릭으로 해지 — 가입할 때의 가격은 그대로 잠깁니다

    채널 하나가 아니라 클라이언트 여럿을 위한 플랜입니다. 모든 프롬프트에 테이크 네 개, 아껴 쓰지 않는 1080P, 그리고 잔액을 보지 않고도 한 장면을 다시 찍을 여유.

    한눈에

    월 클립 수
    480P 5초로 ≈ 124편
    1080P로만 하면
    월 ≈ 31편
    1달러당 크레딧
    Starter 대비 +28%
    환불 기간
    7일

The whole procedure

How to make an AI explainer video shot by shot

Four steps, repeated once per shot. The first one is written in a document, not in this generator — and skipping it is why most attempts stall at shot four.

  1. Write the shot list before you generate anything

    One line per shot: what is on screen, what is said, how long. Twelve rows for a two-minute film. This is the part the tool cannot do for you, and it is the part that makes the rest cheap.

    Do this first
  2. Fix a style block and reuse it verbatim

    Palette, lens, light, surface finish, register, and "no on-screen text". Paste the identical block into every prompt — it is the only thing holding twenty independent generations together.

    Copy into every prompt
  3. Generate one shot, with its line in quotation marks

    The narration is produced with the picture, so there is nothing to record and nothing to sync. Write the line in the language it will be heard in.

    Voice in the same pass
  4. Draft the whole film at 480P, then re-run the keepers

    Cut the drafts together and watch it end to end before you spend anything on resolution. Most re-writes are structural, and a 480P storyboard finds them for a quarter of the price.

    Up to 1080P

The two-thirds rule from the corporate films made here: the demonstration is the film. Openers and closers are cheap to generate and cheap to cut — the shots that earn the approval are the ones showing the thing actually being done.

Questions people actually ask

AI explainer video questions

Whether it writes the script, how the voice works, how long a film really takes, and what you can publish

Does it turn a script into a finished explainer automatically?

No, and any tool promising that on this model is describing something Wan 3.0 does not have. There is no shot-count parameter and no director mode. You generate one shot at a time — 2 to 30 seconds each — and cut them together. A two-minute explainer is twelve to twenty generations. If a rough summary is enough, document to video will read a fifty-page deck and decide the highlights for you.

Where does the voice-over come from?

It is generated alongside the picture, in the same pass, from the line you put in quotation marks in the prompt. There is no voice library to choose from, no audio file to upload and nothing to sync afterwards. Audio does not change the price.

Can the narration be in Japanese, Chinese or Korean?

Yes. Wan 3.0 guarantees prompts in English and Chinese and reads others. Write the spoken line in the target language rather than translating an English one afterwards — cadence survives translation badly and a native reviewer hears it immediately. Keep any appearance clause in English.

How do I stop it looking like an advert?

Say so, as a constraint next to the picture: documentary register, matte surfaces, low-saturation palette, no music, restrained delivery, no on-screen text. The default of every video model is bright and upbeat because most of its training material is. Tone is a prompt field, and it responds to one.

How do I keep twenty shots looking like one film?

A style block pasted verbatim into every prompt, plus the same reference images sent with each generation. Nothing carries between requests on its own. Matching is close rather than exact, which is usually enough at cut speed — character consistency explains what holds and what does not.

Can it show a warning without showing the accident?

That is the standard safety-film constraint, and it is solved by what you write rather than what you ban. Describe the test instead of the failure: the rig, the load, the gauge, the hands, the thing holding. Naming what must not appear tends to put it in frame.

What does a two-minute explainer cost in credits?

Plan for forty to eighty generations rather than twenty: twelve to twenty shots, three or four attempts each. Draft everything at 480P — a quarter of the price of a finished clip — cut it together, then re-run only the keepers at 720P or 1080P. A run that fails outright is refunded automatically.

Can I publish this on our company site?

Yes. Commercial use is included on every paid plan and every credit pack — you may publish, sell and license the output. Free-tier clips are for evaluation. You remain responsible for the claims in the finished film, and responsible use covers what may not be generated at all.

Do I have to disclose that the explainer was made with AI?

On your own site that is your call and your legal team's. On a platform it is a written rule, and an explainer is the format most likely to be caught by it: YouTube's disclosure requirement covers content that looks real — a realistic scene that never happened, a person who appears to say something they did not — which is exactly what a generated presenter in a generated office is. A photorealistic explainer needs the AI-use toggle set; a fully animated one does not. The penalty for consistently not disclosing is not a warning label, it is removal of the content or suspension from the Partner Program. Everything about the film can be true and the disclosure still be owed, because the rule is about how it was made rather than whether it is honest.

Write one shot. Hear it back in ninety seconds.

Before you budget a film, generate its hardest shot — the one with the mark, the constraint and the line that has to be right. First clip free at 480P, an email and no card. If the register is wrong, you will know today rather than at the review.

wan-3.run 편집팀이 쓰고 관리합니다게시 최종 수정 도구 버전 2026.09.4