Wan 3.0 Lip Sync Video

Drop in a face, quote the line, and the mouth follows the words. Wan 3.0 lip sync generates the voice in the same pass as the picture, so there is no audio file to record first and no dubbing pass afterwards. Two to thirty seconds, sound inside the MP4, and your first clip is free.

The speaker

Create image

JPG · PNG · BMP · WEBP · ≤20MB · 240–8000px · ratio ≤8:1

0 / 20000

The speaker keeps the same face, hair, clothing and colours throughout; only the camera and the described motion change.

Model
Aspect ratio
Duration
Resolution
免费注册 · 送一条 Wan 3.0 片子

免费那条就是同一个 Wan 3.0,声音也有,上限是 480P · 3s,走的是共用队列。方案能把上限抬高到 1080P 和三十秒 —— 免费档是一个分辨率,不是另一个模型。

1 free clip on the real model · an email, no card · turning the sound on costs the same as leaving it off

Alibaba's own Wan 3.0 dialogue demos · Try this loads the prompt

A woman at a kitchen counter holding a pink portable blender, speaking straight to cameraSPOKEN TO CAMERA · 9:16 · 720P

A piece to camera, six lines long

What comes back

What Wan 3.0 does when the prompt has a line in it

Every clip here is a demo Alibaba published for Wan 3.0, and each one arrived with the prompt that produced it. Press Hear it for the voice, or Use prompt to load the whole thing into the box at the top of this page — printed in full, in the language it was written in.

  • 1 passpicture, speech and mouth movement in one generation
  • $0extra for the audio track
  • wan3.0-videothe model behind every clip on this band
  • A woman at a kitchen counter holding a pink portable blender, speaking straight to cameraReference
    Six lines to camera, the product held to a reference image
  • A singer front and centre on a night rooftop stage with a DJ behind her and a city skyline beyondReference
    One singer switching languages, the sync holding through all eight
  • A man in a long grey coat and goggles standing in a sunlit wooden dojoReference
    Two speakers, Japanese, one line each and no crosstalk
  • A ginger cat and a golden retriever in headphones at podcast microphones under a ModelTalk neon signFirst frame
    A still of two animals becomes ten rounds of dialogue

What makes a spoken line land

  • The words

    Quote the exact line. A paraphrase gets a line the model invented.

  • Speaker

    Pin who is talking — age, hair, clothing — or the wrong mouth moves.

  • Camera

    Keep it on the face for the length of the line.

  • Text

    Close with no subtitles, or the line can arrive printed on the frame.

From a text prompt

Generate AI lip sync video from simple text prompts

No script format, no timeline, no voice file. Describe the shot in ordinary sentences, quote the words the character says and attribute them to the person on camera, and Wan 3.0 renders the picture, synthesises the voice and moves the mouth to it in a single request.

One prompt, one spoken line

Medium A-roll. The woman stands behind the kitchen counter
holding the pink portable blender, speaking naturally to camera.
She says: "I get asked how I make smoothies so quickly."
No subtitles, no text overlays, no UI, no watermark.

  • Shot, camera, and who is talking
  • The quoted line, and what not to print
2484 / 20,000

One pass out

A woman at a kitchen counter holding a pink portable blender, speaking straight to camera9:16 · 720P · 30S

wan3.0-video · 16:9 · 720P · 14s · speech and mouth in the same pass

Those four lines are lifted from Alibaba's own thirty-second demo, which runs to 2,484 characters against a 20,000-character ceiling. The room is for pinning the speaker, the product and the exclusions — not for writing more. Uploading a portrait instead of describing one changes nothing about how the line itself is written.

Before you spend a credit

Improve your lip sync video prompt with AI

A run that comes back wrong is almost never a model problem. It is a prompt that named the picture and left the delivery, the camera and the subtitle rule to chance.

The prompt generator takes the line you want said and returns a full Wan 3.0 prompt around it — speaker pinned, one camera move named, the audio layers ranked, and the line delimited so it is performed rather than narrated. It checks the result against the rules this model actually enforces, so the clause people forget is added before the credit is spent rather than after.

  1. 01

    Type the line you need said

    One sentence is enough to start. Who says it and where can be a fragment.

  2. 02

    Get a prompt built around it

    Speaker, camera, sound layers and the delimited line, in the order Wan 3.0 reads them.

  3. 03

    Send it here and add the face

    The prompt arrives in the box at the top of this page with the length and ratio already set.

Open the prompt generator

Writing the line for a paid ad rather than a piece to camera? The UGC ad generator has the word count fifteen seconds actually holds, and the disclosure rules that decide whether it runs.
Already have a clip whose delivery you liked? Video to prompt transcribes the spoken line and hands it back inside braces — the form this model performs rather than narrates.

End to end

A complete AI workflow, from script to video

Four doors into the same endpoint. Whatever you are holding — a paragraph of script, a deck, a portrait, or a clip you already rendered — there is a route from it to a talking take, and none of them needs a separate voice tool bolted on.

  • A young man and an elderly woman play chess at a table in a busy city square

    Script → prompt

    Hand over the line and get a full shot description built around it, with the speaker pinned and the line already attributed to them.

    Free · no credits
  • A long-haul rig on an open highway unfolding into a walking machine

    Deck or page → script

    A PDF or a public URL becomes the story the shot is built from, rather than slides with an avatar reading them.

    50 pages · 100 MB
  • Two men in suits in a tiled office washroom, one of them mid-line

    Portrait + line → talking take

    The face goes in as frame one, the line goes in quoted and attributed, and the mouth follows it.

    2–30 s · up to 1080P
  • The same red-haired woman held across a wide, a lift interior and a close-up in a shopping atrium

    New line, same take

    Change what an existing clip says without reshooting it: the speaker, the timing and the framing stay put and the lip movement is regenerated.

    Footage you already have

Everything above lands in the same place — my creations keeps the MP4, the prompt and the task id for every run, so a take you liked can be reopened rather than reconstructed.

The whole procedure

How to lip sync a photo with Wan 3.0

The step people get wrong is the third one: describing the picture the model can already see, instead of writing the words it cannot guess.

  1. Drop in the face

    The check runs instantly, in your browser. A good file shows a green row; a bad one names the fix before you spend anything.

    Checked free
  2. Quote the line, and say who says it

    The exact words in quotation marks, attributed to the person on camera. A paraphrase — "she says something reassuring" — gets you a line the model invented.

    Exact words only
  3. Close with "no subtitles"

    There is no negative prompt field, so the instruction not to print the line is a sentence like the others. It is one clause and it saves a whole run.

    One clause
  4. Set length and resolution, then generate

    Buy the seconds the line needs first. One MP4 comes back with the speech, the room tone and the mouth movement already in it.

    Up to 1080P

Around ninety seconds for a short clip, and you can close the tab while it runs. The uploaded photograph is never modified — the output is a new file.

What lip sync means here

Your photograph is frame one. The voice is written, not uploaded.

Most tools in this category take a picture and an audio file and move the mouth to match. This one takes a picture and a sentence.

A portrait used as the opening frame of a Wan 3.0 lip sync clip

Frame 001 · your face

Spoken, and synced, from here on

Wan 3.0 lip sync puts your photograph in as the literal first frame and generates 2 to 30 seconds forward from it, at 30 fps in 480P, 720P or 1080P. The words you quoted come back as speech in that clip's own audio track, with the mouth moving to them, because the picture and the sound are rendered in the same pass rather than stitched together in two.

Your photo is frame one
001

Your photo is frame one

Audio files you have to supply
0

Audio files you have to supply

Any whole second
2–30 s

Any whole second

MP4 back, sound already inside it
1

MP4 back, sound already inside it

  • No text-to-speech step before the render
  • No dubbing pass after it
  • No second model to license for the mouth

The line

What decides whether a line is spoken or printed across the frame

Two sentences decide whether a line is heard or read. One puts the exact words in the speaker's mouth; the other keeps them off the picture. Leave either out and the failure arrives as finished, billed video.

The anatomy of a synced line

Speech and picture · one pass

…speaking naturally to camera. She says:"It blends everything smooth."

Realistic kitchen ambience, blending motor, casual room tone.No subtitles, no text overlays.

A ginger cat and a golden retriever in headphones at podcast microphones under a ModelTalk neon signTen rounds · 720P · 30s
Two speakers, ten rounds, and only the one talking opens its mouth — press for sound

What the mouth needs to land

3

  • The exact words, quoted rather than summarised
  • A speaker described closely enough to pin
  • The camera on their face while they talk

The rule people miss

There is no negative prompt field, so "do not print this" is a sentence in the prompt like any other. Leave it out and the model is free to treat your dialogue as something to render on screen as well as speak — which is why Alibaba's own thirty-second demo, printed above, ends on that exact instruction.

  • no subtitles, no text overlaysRight
  • Negative prompt: subtitlesWrong

Anything shaped like a negative prompt is read as more words to render, and the request still succeeds — so the failure arrives as a finished clip with the label "Negative prompt" somewhere in the frame.

Braces work too — {like this} is what the prompt tools on this site emit. What the model is actually looking for is a delimited line with a visible speaker attached to it, and Alibaba's own published prompts use quotation marks for it. Text to video covers the four sound layers and the order to name them in.

Before you spend

Check the portrait before you spend a credit

A face that Wan 3.0 will not read is the most annoying way to lose a generation, because you find out after the queue rather than before it. Drop the photograph here and it is checked against the real limits — format, size, dimensions, ratio and transparency — with the fix named for whichever row fails. Nothing is uploaded: the check runs in your browser.

JPG · JPEG · PNG · BMP · WEBP · ≤20.0 MB · 2408,000px · ratio ≤8:1

One face, framed so the mouth is not the smallest thing in the picture, gives the sync the most to work with. The checker never touches the model, so a portrait that was never going to be read costs nothing to find out about.

What goes in

Wan 3.0 lip sync limits — the portrait, the clip and the voice

Five of these rows govern the photograph, two govern what comes out, and the last one is the one that produces a rejection nobody expects. The upload box at the top of this page checks the image rows before you spend anything.

A locked first frame and a voice recording cannot travel together. Frames are one family of inputs and references are the other; a request carrying both is refused outright rather than downgraded, so the choice happens before you upload.

Portrait format
JPEG · JPG · PNG · BMP · WEBPHEIC is refused — an iPhone's default container is not on the list
Official
Portrait size
≤ 20 MB
Official
Dimensions
240 – 8,000 px per side
Official
Aspect ratio
8:1 – 1:8A range, not a ceiling — a 1:9 column is refused like a 9:1 banner
Official
Transparency
not acceptedA PNG alpha channel is refused rather than flattened
Official
Clip length
2 – 30 s at 30 fpsAny whole number of seconds, in 480P, 720P or 1080P
Official
Voice reference
5 tracks · ≤ 15 s in totalWAV or MP3, up to 15 MB each, 1–15 s per track
Official
Frames and references
mutually exclusive
Official

The line itself has no separate field and no separate limit: it lives in the prompt with everything else, and the prompt holds 20,000 characters. What constrains a spoken line is the clock, not the character count — around two to three seconds of screen time per short sentence, so a fifteen-word line in a four-second clip either gets rushed or gets cut off. Buy the seconds the line needs before you buy the resolution.

Pricing · 1080P on every paid plan · Native audio

What one Wan 3.0 lip sync clip costs

Length and resolution set the price. Speech does not: the audio switch changes what comes back, not what it costs, and a voice reference is not billed at all. A failed run is refunded automatically, and the checker above keeps most of the avoidable failures away from the model in the first place.

How this is billed

Billed by
output second
Speech, voice reference
no charge
Your first clip
free · 480P

Draft the line at 480P until the delivery lands, then spend one run at 720P or 1080P — the portrait does not need re-checking between runs. Full pricing.

两种付法,每月积分一样

  • 免费的第一条不用卡

    在真正的 wan3.0-video 上生成一次。拿你自己的镜头看看 Wan 3.0 会做成什么样 — 480P,三秒,声音同一次出。

    $0一次性

    全程不用卡

    生成免费的那条

    一个邮箱,不用卡

    1 条,仅此一次

    480P · 最长 3 秒 · 带声音
    一个邮箱,不用卡

    和付费套餐跑的是同一个 `wan3.0-video` — 480P 是一个分辨率,不是一个缩水的模型

    • 真正的 wan3.0-video,不是 2.7
    • 声音和画面一起生成
    • 收据上有 model ID 和 task ID
    • 跑失败永远不扣你的
    • 720P 与 1080P付费
    • 超过 3 秒的片子付费
    • 参考、文档与编辑模式付费
    • 第二条以及以后付费

    只回答一个问题:Wan 3.0 会把你的想法做成什么样。1080P、三十秒和全部输入模式,从 $12.90 起。

    一眼看完

    每月条数
    1 条,仅此一次
    每次出几条
    1
    最长时长
    3s
    最高分辨率
    480P
  • Starter最多省 20%

    一周一条发布贴。够你摸清 Wan 3.0 AI 视频生成器合不合你的工作方式,而不用先对用量下判断。

    $12.90/ 月$15.90

    按年收 $154.80

    随时取消

    640 积分 / 月 · 约 16 条

    按 5 秒 480P 算 · 720P 8 条 · 1080P 4 条

    套餐积分在当月月底过期

    • 每月 640 积分 — 大约 16 条成片
    • 同一条提示词每次出 2 条
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    前两行以下的一切,$12.90 这一档就都有。更大的套餐买的是秒数和条数 — 从来不是更好的模型。

    一眼看完

    每月条数
    5 秒 480P 下约 12 条
    每次出几条
    2
    每美元积分
    基准
    退款窗口
    7 天
  • 多数人落在这里
    Pro

    每个工作日一条成片,每条还有三个备选。480P 打草稿、1080P 定稿的那个循环,按一个工作周的量配好。

    $39.90/ 月$49.90

    按年收 $478.80

    随时取消

    2,240 积分 / 月 · 约 56 条

    按 5 秒 480P 算 · 720P 28 条 · 1080P 14 条

    套餐积分在当月月底过期

    • 每月 2,240 积分 — Starter 的 3.5 倍
    • 每月约 56 条成片,或者 14 条满 1080P
    • 同一条提示词每次出 4 条 — 挑一条,而不是重摇一次2× 条数
    • 每美元积分比 Starter 多 13%更划算
    • 同一周内既够 480P 打草稿,也够 1080P 定稿
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    从 Starter 往上,涨的是量、条数和单价。模型、分辨率和那三十秒,本来就已经是你的了。

    一眼看完

    每月条数
    5 秒 480P 下约 44 条
    每次出几条
    4
    每美元积分
    比 Starter 多 21%
    退款窗口
    7 天
  • Studio最划算

    接客户量的档,也是你不再省着用 1080P 的那一档。每月三十一条满分辨率成片,单位积分是我们卖过最低的价。

    $99.90/ 月$119.90

    按年收 $1,198.80

    随时取消

    6,240 积分 / 月 · 约 156 条

    按 5 秒 480P 算 · 720P 78 条 · 1080P 39 条

    套餐积分在当月月底过期

    • 每月 6,240 积分 — Starter 的 9.75 倍,Pro 的 2.8 倍13×
    • 每月约 156 条成片,或者 39 条满 1080P
    • 每美元积分多 26% — 我们卖过最划算的最划算
    • 1080P 多到你不必先用 480P 打草稿
    • 同一条提示词每次出 4 条
    • 按一整份客户名单配的量,不是一个频道
    • 任何一个月都能加购,不用换套餐
    • 全部分辨率 — 480P、720P 与 1080P
    • 全部时长 — 一次 2 到 30 秒
    • 全部输入 — 文字、图片、参考、文档、网页
    • 声音写在同一次生成里,无水印
    • 含商用授权 — 发布、售卖、开票都可以
    • Wan 3.0 Prime 快速档,随时可用
    • 跑失败永远不扣积分
    • 积分未动用时,7 天内可退
    • 两步取消 — 你加入时的价格锁死

    这是为一整份客户名单配的档,不是为一个频道:每条提示词出四条、1080P 不用省着用,还留得出重拍一场戏的余量,不用盯着余额往下掉。

    一眼看完

    每月条数
    5 秒 480P 下约 124 条
    满 1080P 的话
    每月约 31 条
    每美元积分
    比 Starter 多 28%
    退款窗口
    7 天

Questions

Wan 3.0 lip sync — the questions people actually arrive with

The ones worth answering before you upload a face, with the numbers taken from Alibaba's own parameter table rather than from a feature list.

Do I need an audio file to get lip sync?

No — so you can stop looking for one. Quote the words in the prompt and attribute them to the person on camera — She says: "Two minutes to set up." — and Wan 3.0 generates the speech itself, in the same pass as the picture, with the mouth moving to it. A photograph and a sentence is the whole input. A voice recording is the other route, for when the timbre has to be a specific one.

How do I write the line so it gets spoken rather than described?

Quote it and attach it to somebody. Alibaba's own prompt formula puts the spoken content beside the emotion, tone and accent it should carry, and every demo on this page follows it — She says: "…", then the delivery. Words that are only summarised get narrated or invented instead. The prompt tools on this site wrap lines in braces, {like this}, which does the same job of marking where the line starts and stops.

My line came out printed on the screen instead of spoken. Why?

Because nothing in the prompt said not to print it. There is no negative prompt field on this model, so exclusions are ordinary sentences — close with "no subtitles, no text overlays" and the line renders as audio. Writing Negative prompt: subtitles does the opposite: those words become something to draw.

Can I upload my photo and my voice recording together?

Not in one request. A photograph pinned as the first frame belongs to the frame family and an audio track belongs to the reference family, and the two families are mutually exclusive — the request is refused rather than downgraded. Pick one: the exact opening frame with a generated voice, or your recorded voice with the picture attached as a reference instead.

How long can the line be?

Long enough for the seconds you bought. The prompt field holds 20,000 characters and the clip holds 2 to 30 seconds at 30 fps, so the constraint is the clock: Alibaba's own thirty-second demo above fits six spoken lines, which is roughly one short sentence per five seconds once the B-roll beats are counted. A line that overruns gets rushed or clipped, and no amount of prompt length fixes that — buy more seconds instead.

Which languages does the lip sync work in?

More than you would guess, and the evidence is on this page: the rooftop clip above is Alibaba's own demo of one singer switching across eight languages with the sync holding, and the dojo clip runs its dialogue in Japanese. What does not exist is a published list of supported speech languages or any phoneme-lock parameter, so treat a specific number quoted elsewhere as somebody's guess. Run three seconds at 480P in the language you need — it costs a fraction of a keeper and it answers the question for your case.

Can two people speak in the same clip?

Yes, and the podcast clip above is two of them trading ten rounds. Give each speaker a unique label and enough description to be told apart, anchor each line to something that speaker is visibly doing, then write the lines in the order they are said. Alibaba's guidance is explicit that pronouns merge speakers — "he says… then he says" is how two characters end up with one voice.

Does the speech cost extra?

No. Billing is by output second and resolution only, so a clip with dialogue costs exactly what the same clip in silence costs. Reference images and reference audio are not billed either. The one input that does add to the bill is a reference video — those seconds are charged at the output rate on top of the seconds you asked for.

Does the rest of the picture stay still while they talk?

Only if you say so. Your photograph is frame one, not a locked plate — everything after it is generated, so the room, the light and the framing drift unless the prompt pins them. Both Alibaba demos above spend a paragraph on exactly that, naming what must stay identical; the appearance lock beside the prompt box writes the same clause for you.

Can I publish a talking clip commercially?

Yes. Commercial use is included on every paid plan, with no watermark and no separate licence to buy, under Alibaba's usage policy and our terms. The limit is not the licence, it is the face: the person in the photograph and the voice on the recording both have to have agreed.

One face, one sentence

Upload the portrait, quote the line you need said, and hear what comes back. One free Wan 3.0 clip at 480P, up to three seconds — an email, no card. If the delivery is not the one you imagined, it cost you nothing.

由 wan-3.run 编辑团队撰写与维护发布于 最后更新 工具 版本 2026.09.4