Wan 3.0 Lip Sync Video

Drop in a face, quote the line, and the mouth follows the words. Wan 3.0 lip sync generates the voice in the same pass as the picture, so there is no audio file to record first and no dubbing pass afterwards. Two to thirty seconds, sound inside the MP4, and your first clip is free.

The speaker

Create image

JPG · PNG · BMP · WEBP · ≤20MB · 240–8000px · ratio ≤8:1

0 / 20000

The speaker keeps the same face, hair, clothing and colours throughout; only the camera and the described motion change.

Model
Aspect ratio
Duration
Resolution
Sign up free · 1 clip on Wan 3.0

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 free clip on the real model · an email, no card · turning the sound on costs the same as leaving it off

Alibaba's own Wan 3.0 dialogue demos · Try this loads the prompt

A woman at a kitchen counter holding a pink portable blender, speaking straight to cameraSPOKEN TO CAMERA · 9:16 · 720P

A piece to camera, six lines long

What comes back

What Wan 3.0 does when the prompt has a line in it

Every clip here is a demo Alibaba published for Wan 3.0, and each one arrived with the prompt that produced it. Press Hear it for the voice, or Use prompt to load the whole thing into the box at the top of this page — printed in full, in the language it was written in.

  • 1 passpicture, speech and mouth movement in one generation
  • $0extra for the audio track
  • wan3.0-videothe model behind every clip on this band
  • A woman at a kitchen counter holding a pink portable blender, speaking straight to cameraReference
    Six lines to camera, the product held to a reference image
  • A singer front and centre on a night rooftop stage with a DJ behind her and a city skyline beyondReference
    One singer switching languages, the sync holding through all eight
  • A man in a long grey coat and goggles standing in a sunlit wooden dojoReference
    Two speakers, Japanese, one line each and no crosstalk
  • A ginger cat and a golden retriever in headphones at podcast microphones under a ModelTalk neon signFirst frame
    A still of two animals becomes ten rounds of dialogue

What makes a spoken line land

  • The words

    Quote the exact line. A paraphrase gets a line the model invented.

  • Speaker

    Pin who is talking — age, hair, clothing — or the wrong mouth moves.

  • Camera

    Keep it on the face for the length of the line.

  • Text

    Close with no subtitles, or the line can arrive printed on the frame.

From a text prompt

Generate AI lip sync video from simple text prompts

No script format, no timeline, no voice file. Describe the shot in ordinary sentences, quote the words the character says and attribute them to the person on camera, and Wan 3.0 renders the picture, synthesises the voice and moves the mouth to it in a single request.

One prompt, one spoken line

Medium A-roll. The woman stands behind the kitchen counter
holding the pink portable blender, speaking naturally to camera.
She says: "I get asked how I make smoothies so quickly."
No subtitles, no text overlays, no UI, no watermark.

  • Shot, camera, and who is talking
  • The quoted line, and what not to print
2484 / 20,000

One pass out

A woman at a kitchen counter holding a pink portable blender, speaking straight to camera9:16 · 720P · 30S

wan3.0-video · 16:9 · 720P · 14s · speech and mouth in the same pass

Those four lines are lifted from Alibaba's own thirty-second demo, which runs to 2,484 characters against a 20,000-character ceiling. The room is for pinning the speaker, the product and the exclusions — not for writing more. Uploading a portrait instead of describing one changes nothing about how the line itself is written.

Before you spend a credit

Improve your lip sync video prompt with AI

A run that comes back wrong is almost never a model problem. It is a prompt that named the picture and left the delivery, the camera and the subtitle rule to chance.

The prompt generator takes the line you want said and returns a full Wan 3.0 prompt around it — speaker pinned, one camera move named, the audio layers ranked, and the line delimited so it is performed rather than narrated. It checks the result against the rules this model actually enforces, so the clause people forget is added before the credit is spent rather than after.

  1. 01

    Type the line you need said

    One sentence is enough to start. Who says it and where can be a fragment.

  2. 02

    Get a prompt built around it

    Speaker, camera, sound layers and the delimited line, in the order Wan 3.0 reads them.

  3. 03

    Send it here and add the face

    The prompt arrives in the box at the top of this page with the length and ratio already set.

Open the prompt generator

Writing the line for a paid ad rather than a piece to camera? The UGC ad generator has the word count fifteen seconds actually holds, and the disclosure rules that decide whether it runs.
Already have a clip whose delivery you liked? Video to prompt transcribes the spoken line and hands it back inside braces — the form this model performs rather than narrates.

End to end

A complete AI workflow, from script to video

Four doors into the same endpoint. Whatever you are holding — a paragraph of script, a deck, a portrait, or a clip you already rendered — there is a route from it to a talking take, and none of them needs a separate voice tool bolted on.

  • A young man and an elderly woman play chess at a table in a busy city square

    Script → prompt

    Hand over the line and get a full shot description built around it, with the speaker pinned and the line already attributed to them.

    Free · no credits
  • A long-haul rig on an open highway unfolding into a walking machine

    Deck or page → script

    A PDF or a public URL becomes the story the shot is built from, rather than slides with an avatar reading them.

    50 pages · 100 MB
  • Two men in suits in a tiled office washroom, one of them mid-line

    Portrait + line → talking take

    The face goes in as frame one, the line goes in quoted and attributed, and the mouth follows it.

    2–30 s · up to 1080P
  • The same red-haired woman held across a wide, a lift interior and a close-up in a shopping atrium

    New line, same take

    Change what an existing clip says without reshooting it: the speaker, the timing and the framing stay put and the lip movement is regenerated.

    Footage you already have

Everything above lands in the same place — my creations keeps the MP4, the prompt and the task id for every run, so a take you liked can be reopened rather than reconstructed.

The whole procedure

How to lip sync a photo with Wan 3.0

The step people get wrong is the third one: describing the picture the model can already see, instead of writing the words it cannot guess.

  1. Drop in the face

    The check runs instantly, in your browser. A good file shows a green row; a bad one names the fix before you spend anything.

    Checked free
  2. Quote the line, and say who says it

    The exact words in quotation marks, attributed to the person on camera. A paraphrase — "she says something reassuring" — gets you a line the model invented.

    Exact words only
  3. Close with "no subtitles"

    There is no negative prompt field, so the instruction not to print the line is a sentence like the others. It is one clause and it saves a whole run.

    One clause
  4. Set length and resolution, then generate

    Buy the seconds the line needs first. One MP4 comes back with the speech, the room tone and the mouth movement already in it.

    Up to 1080P

Around ninety seconds for a short clip, and you can close the tab while it runs. The uploaded photograph is never modified — the output is a new file.

What lip sync means here

Your photograph is frame one. The voice is written, not uploaded.

Most tools in this category take a picture and an audio file and move the mouth to match. This one takes a picture and a sentence.

A portrait used as the opening frame of a Wan 3.0 lip sync clip

Frame 001 · your face

Spoken, and synced, from here on

Wan 3.0 lip sync puts your photograph in as the literal first frame and generates 2 to 30 seconds forward from it, at 30 fps in 480P, 720P or 1080P. The words you quoted come back as speech in that clip's own audio track, with the mouth moving to them, because the picture and the sound are rendered in the same pass rather than stitched together in two.

Your photo is frame one
001

Your photo is frame one

Audio files you have to supply
0

Audio files you have to supply

Any whole second
2–30 s

Any whole second

MP4 back, sound already inside it
1

MP4 back, sound already inside it

  • No text-to-speech step before the render
  • No dubbing pass after it
  • No second model to license for the mouth

The line

What decides whether a line is spoken or printed across the frame

Two sentences decide whether a line is heard or read. One puts the exact words in the speaker's mouth; the other keeps them off the picture. Leave either out and the failure arrives as finished, billed video.

The anatomy of a synced line

Speech and picture · one pass

…speaking naturally to camera. She says:"It blends everything smooth."

Realistic kitchen ambience, blending motor, casual room tone.No subtitles, no text overlays.

A ginger cat and a golden retriever in headphones at podcast microphones under a ModelTalk neon signTen rounds · 720P · 30s
Two speakers, ten rounds, and only the one talking opens its mouth — press for sound

What the mouth needs to land

3

  • The exact words, quoted rather than summarised
  • A speaker described closely enough to pin
  • The camera on their face while they talk

The rule people miss

There is no negative prompt field, so "do not print this" is a sentence in the prompt like any other. Leave it out and the model is free to treat your dialogue as something to render on screen as well as speak — which is why Alibaba's own thirty-second demo, printed above, ends on that exact instruction.

  • no subtitles, no text overlaysRight
  • Negative prompt: subtitlesWrong

Anything shaped like a negative prompt is read as more words to render, and the request still succeeds — so the failure arrives as a finished clip with the label "Negative prompt" somewhere in the frame.

Braces work too — {like this} is what the prompt tools on this site emit. What the model is actually looking for is a delimited line with a visible speaker attached to it, and Alibaba's own published prompts use quotation marks for it. Text to video covers the four sound layers and the order to name them in.

Before you spend

Check the portrait before you spend a credit

A face that Wan 3.0 will not read is the most annoying way to lose a generation, because you find out after the queue rather than before it. Drop the photograph here and it is checked against the real limits — format, size, dimensions, ratio and transparency — with the fix named for whichever row fails. Nothing is uploaded: the check runs in your browser.

JPG · JPEG · PNG · BMP · WEBP · ≤20.0 MB · 2408,000px · ratio ≤8:1

One face, framed so the mouth is not the smallest thing in the picture, gives the sync the most to work with. The checker never touches the model, so a portrait that was never going to be read costs nothing to find out about.

What goes in

Wan 3.0 lip sync limits — the portrait, the clip and the voice

Five of these rows govern the photograph, two govern what comes out, and the last one is the one that produces a rejection nobody expects. The upload box at the top of this page checks the image rows before you spend anything.

A locked first frame and a voice recording cannot travel together. Frames are one family of inputs and references are the other; a request carrying both is refused outright rather than downgraded, so the choice happens before you upload.

Portrait format
JPEG · JPG · PNG · BMP · WEBPHEIC is refused — an iPhone's default container is not on the list
Official
Portrait size
≤ 20 MB
Official
Dimensions
240 – 8,000 px per side
Official
Aspect ratio
8:1 – 1:8A range, not a ceiling — a 1:9 column is refused like a 9:1 banner
Official
Transparency
not acceptedA PNG alpha channel is refused rather than flattened
Official
Clip length
2 – 30 s at 30 fpsAny whole number of seconds, in 480P, 720P or 1080P
Official
Voice reference
5 tracks · ≤ 15 s in totalWAV or MP3, up to 15 MB each, 1–15 s per track
Official
Frames and references
mutually exclusive
Official

The line itself has no separate field and no separate limit: it lives in the prompt with everything else, and the prompt holds 20,000 characters. What constrains a spoken line is the clock, not the character count — around two to three seconds of screen time per short sentence, so a fifteen-word line in a four-second clip either gets rushed or gets cut off. Buy the seconds the line needs before you buy the resolution.

Pricing · 1080P on every paid plan · Native audio

What one Wan 3.0 lip sync clip costs

Length and resolution set the price. Speech does not: the audio switch changes what comes back, not what it costs, and a voice reference is not billed at all. A failed run is refunded automatically, and the checker above keeps most of the avoidable failures away from the model in the first place.

How this is billed

Billed by
output second
Speech, voice reference
no charge
Your first clip
free · 480P

Draft the line at 480P until the delivery lands, then spend one run at 720P or 1080P — the portrait does not need re-checking between runs. Full pricing.

Same monthly credits either way

  • Free first clipNo card

    One generation on the real wan3.0-video. See what Wan 3.0 does with your own shot — 480P, three seconds, sound in the same pass.

    $0once

    No card at any point

    Generate free clip

    An email, no card

    1 clip, once

    480P · up to 3s · with sound
    An email, no card

    The same `wan3.0-video` Alibaba runs for the paid plans — 480P is a resolution, not a lesser model

    • The real wan3.0-video, not 2.7
    • Sound generated with the picture
    • Model ID and task ID on the receipt
    • A failed run never costs you
    • 720P and 1080PPaid
    • Clips longer than 3 secondsPaid
    • Reference, document and editing modesPaid
    • More than one clip, everPaid

    One question answered: what Wan 3.0 does with your idea. 1080P, thirty seconds and every input mode start at $12.90.

    At a glance

    Clips a month
    1, once
    Takes per run
    1
    Longest clip
    3s
    Top resolution
    480P
  • StarterSave up to 20%

    A launch post a week. Enough to find out whether the Wan 3.0 AI video generator suits the way you work, without a decision about volume.

    $12.90/mo$15.90

    $154.80 billed yearly

    Cancel anytime

    640 credits / month · ≈ 16 clips

    at 5s 480P · 8 at 720P · 4 at 1080P

    Plan credits expire at the end of the month

    • 640 credits a month — about 16 finished clips
    • 2 takes of one prompt per run
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    Everything under the first two lines is on this plan at $12.90. The bigger plans buy seconds and takes — never a better model.

    At a glance

    Clips a month
    ≈ 12 at 5s 480P
    Takes per run
    2
    Credits per dollar
    Baseline
    Refund window
    7 days
  • Most people land here
    Pro

    One finished clip every working day, with three alternates each. The draft-at-480P, finish-at-1080P loop, sized for a working week.

    $39.90/mo$49.90

    $478.80 billed yearly

    Cancel anytime

    2,240 credits / month · ≈ 56 clips

    at 5s 480P · 28 at 720P · 14 at 1080P

    Plan credits expire at the end of the month

    • 2,240 credits a month — 3.5× Starter
    • ≈ 56 finished clips a month, or 14 at full 1080P
    • 4 takes of one prompt per run — pick one instead of re-rolling2× takes
    • +13% credits per dollar than StarterBetter rate
    • Room to draft at 480P and finish at 1080P in the same week
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    The jump from Starter is quantity, takes and rate. The model, the resolutions and the thirty seconds were already yours.

    At a glance

    Clips a month
    ≈ 44 at 5s 480P
    Takes per run
    4
    Credits per dollar
    +21% vs Starter
    Refund window
    7 days
  • StudioBest rate

    Client volume, and the plan where you stop rationing 1080P. Thirty-one full-resolution finished clips a month, at the lowest price per credit we sell.

    $99.90/mo$119.90

    $1,198.80 billed yearly

    Cancel anytime

    6,240 credits / month · ≈ 156 clips

    at 5s 480P · 78 at 720P · 39 at 1080P

    Plan credits expire at the end of the month

    • 6,240 credits a month — 9.75× Starter, 2.8× Pro13×
    • ≈ 156 finished clips a month, or 39 at full 1080P
    • +26% credits per dollar — the best rate we sellBest rate
    • Enough 1080P that you stop drafting at 480P first
    • 4 takes of one prompt per run
    • Sized for a client roster rather than one channel
    • Top up any month without changing plan
    • Every resolution — 480P, 720P and 1080P
    • Every length — 2 to 30 seconds in one pass
    • Every input — text, image, reference, document, web page
    • Sound written in the same pass, no watermark
    • Commercial use included — publish it, sell it, bill for it
    • Wan 3.0 Prime, the fast tier, whenever you want it
    • A failed run never costs credits
    • 7-day refund while the credits are untouched
    • Cancel in two clicks — the price you joined at is locked

    The plan for a client roster rather than one channel: four takes of every prompt, 1080P without rationing it, and room to re-shoot a scene without watching the balance drop.

    At a glance

    Clips a month
    ≈ 124 at 5s 480P
    At full 1080P
    ≈ 31 a month
    Credits per dollar
    +28% vs Starter
    Refund window
    7 days

Questions

Wan 3.0 lip sync — the questions people actually arrive with

The ones worth answering before you upload a face, with the numbers taken from Alibaba's own parameter table rather than from a feature list.

Do I need an audio file to get lip sync?

No — so you can stop looking for one. Quote the words in the prompt and attribute them to the person on camera — She says: "Two minutes to set up." — and Wan 3.0 generates the speech itself, in the same pass as the picture, with the mouth moving to it. A photograph and a sentence is the whole input. A voice recording is the other route, for when the timbre has to be a specific one.

How do I write the line so it gets spoken rather than described?

Quote it and attach it to somebody. Alibaba's own prompt formula puts the spoken content beside the emotion, tone and accent it should carry, and every demo on this page follows it — She says: "…", then the delivery. Words that are only summarised get narrated or invented instead. The prompt tools on this site wrap lines in braces, {like this}, which does the same job of marking where the line starts and stops.

My line came out printed on the screen instead of spoken. Why?

Because nothing in the prompt said not to print it. There is no negative prompt field on this model, so exclusions are ordinary sentences — close with "no subtitles, no text overlays" and the line renders as audio. Writing Negative prompt: subtitles does the opposite: those words become something to draw.

Can I upload my photo and my voice recording together?

Not in one request. A photograph pinned as the first frame belongs to the frame family and an audio track belongs to the reference family, and the two families are mutually exclusive — the request is refused rather than downgraded. Pick one: the exact opening frame with a generated voice, or your recorded voice with the picture attached as a reference instead.

How long can the line be?

Long enough for the seconds you bought. The prompt field holds 20,000 characters and the clip holds 2 to 30 seconds at 30 fps, so the constraint is the clock: Alibaba's own thirty-second demo above fits six spoken lines, which is roughly one short sentence per five seconds once the B-roll beats are counted. A line that overruns gets rushed or clipped, and no amount of prompt length fixes that — buy more seconds instead.

Which languages does the lip sync work in?

More than you would guess, and the evidence is on this page: the rooftop clip above is Alibaba's own demo of one singer switching across eight languages with the sync holding, and the dojo clip runs its dialogue in Japanese. What does not exist is a published list of supported speech languages or any phoneme-lock parameter, so treat a specific number quoted elsewhere as somebody's guess. Run three seconds at 480P in the language you need — it costs a fraction of a keeper and it answers the question for your case.

Can two people speak in the same clip?

Yes, and the podcast clip above is two of them trading ten rounds. Give each speaker a unique label and enough description to be told apart, anchor each line to something that speaker is visibly doing, then write the lines in the order they are said. Alibaba's guidance is explicit that pronouns merge speakers — "he says… then he says" is how two characters end up with one voice.

Does the speech cost extra?

No. Billing is by output second and resolution only, so a clip with dialogue costs exactly what the same clip in silence costs. Reference images and reference audio are not billed either. The one input that does add to the bill is a reference video — those seconds are charged at the output rate on top of the seconds you asked for.

Does the rest of the picture stay still while they talk?

Only if you say so. Your photograph is frame one, not a locked plate — everything after it is generated, so the room, the light and the framing drift unless the prompt pins them. Both Alibaba demos above spend a paragraph on exactly that, naming what must stay identical; the appearance lock beside the prompt box writes the same clause for you.

Can I publish a talking clip commercially?

Yes. Commercial use is included on every paid plan, with no watermark and no separate licence to buy, under Alibaba's usage policy and our terms. The limit is not the licence, it is the face: the person in the photograph and the voice on the recording both have to have agreed.

One face, one sentence

Upload the portrait, quote the line you need said, and hear what comes back. One free Wan 3.0 clip at 480P, up to three seconds — an email, no card. If the delivery is not the one you imagined, it cost you nothing.

Written and maintained by the wan-3.run editorial teamPublished Last updated Tool version 2026.09.4