P—01 / ASTRONAUT IN THE RUINS

Wan 3.0 text to video prompt — a reveal in five seconds

Two sentences of story, one sentence of look, and no camera instruction anyone could argue with.

The plainest useful shape a Wan 3.0 prompt can take: everything that happens, then everything it looks like, separated by a full stop. There is no field protocol, no shot list and no timecode in it — Wan 3.0 reads one string of prose and this is what that string is supposed to look like. It is also, honestly, asking for more story than five seconds can hold, and the clip beside it shows how the model resolved that: it cuts.

Mode
Text to Video
Model
Wan 3.0
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Ambient, generated in the same pass
Credits
80

Alibaba’s published example for this model — not generated on this site. Source: wavespeed.ai

Output reference

A ruined overgrown avenue, a child holding a light in a derelict store, and an astronaut lifting off her helmetVideo
The clip
619 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A ruined overgrown avenue, a child holding a light in a derelict store, and an astronaut lifting off her helmetVideo
The clip
The promptthe clipWan 3.0 · Text to video
619 chars

The reveal

Three narrative beats, then a single closing clause carrying genre, light, camera register and grain.

A lone astronaut walks through the ruins of a once-busy city, surrounded by abandoned cars, overgrown skyscrapers, and trees growing through cracked streets. She discovers a small child standing inside an old convenience store, holding a glowing flower. The astronaut slowly removes her helmet as birds suddenly rise into the sky and sunlight breaks through the clouds. Emotional science-fiction film, grand post-apocalyptic environment, slow cinematic camera movement, wide establishing shots, intimate facial close-ups, realistic dust particles, warm sunlight contrasting with cold ruins, epic yet hopeful atmosphere.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    A lone astronaut … a small child … holding a glowing flower

    Two people and one prop, named and never described again. No hair, no age, no wardrobe beyond the suit — the genre supplies all of it, and every adjective spent here is an adjective not spent on the light.

  2. Scene

    the ruins of a once-busy city, surrounded by abandoned cars, overgrown skyscrapers, and trees growing through cracked streets

    Four nouns build the whole world. They are chosen so that each one implies time passing, which is the thing the shot has to establish and cannot say out loud.

  3. Motion

    walks through … discovers … slowly removes her helmet as birds suddenly rise

    Three verbs in sequence. At five seconds that is more than one shot can hold, and the run resolves it by cutting — a wide, the child, the helmet — rather than by dropping a beat.

  4. Aesthetic control

    slow cinematic camera movement, wide establishing shots, intimate facial close-ups, realistic dust particles, warm sunlight contrasting with cold ruins

    Camera, coverage, particles and colour contrast in one clause. Note that the camera is given a speed but no direction — Wan 3.0 will choose, and here that is the right trade.

  5. Stylization

    Emotional science-fiction film … epic yet hopeful atmosphere

    Genre at the front of the closing clause and mood at the back, bracketing everything technical between them. Both are register instructions, not picture instructions.

  6. Sound

    Not written — left to the model.

    Unwritten. Wan 3.0 generates audio in the same pass whether or not you ask, so leaving this layer empty is accepting whatever it invents. One sentence would have pinned it.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The closing clause

    Everything after the last full stop is transferable as a block. Paste that clause onto a completely different subject and you get the same film in a different place, which is the fastest edit on this page and the one worth learning first.

  2. 02

    The number of beats

    Walks, discovers, unhelmets. Cut it to one and five seconds becomes comfortable; keep all three and raise the duration to fifteen, adding the words "then" and "finally" so the order is stated rather than implied.

  3. 03

    The light reversal

    "Warm sunlight contrasting with cold ruins" is a single instruction that sets two colour temperatures against each other. Swap the pair — cold light on warm ruins — and the same scene reads as a threat instead of a rescue.

  4. 04

    The missing sound line

    Append a sentence: wind through empty streets, distant birds, no music. Wan 3.0 writes picture and sound together, so an audio sentence is also a timing instruction — naming the birds is how the sky beat gets a moment of its own.

Three ways to break it

  • Adding a negative prompt

    Wan 3.0 has no negative prompt field. A line beginning "Negative prompt:" is read as words to render, and the request still returns 200 — so the failure arrives as a finished clip with the words in it rather than as an error. Write exclusions as sentences: "no on-screen text".

  • Numbering the beats with timecodes

    This prompt has no "0-2s" markers and should not gain any. Wan 3.0 has no timecode syntax, so a timecode is just characters; order comes from connectives — then, after that, finally — and a schedule the model cannot read is a schedule it will not keep.

  • Describing the astronaut

    Hair, face and build are the model's to choose here, and constraining them costs attention that is currently going to the dust and the light. If you need a specific person, that is a reference image and a different mode, not more adjectives.

What it does

What the astronaut in the ruins template does

This is the first template in the library because it is the least clever one, and the shape it demonstrates is the shape most Wan 3.0 prompts should take.

Read it structurally and there are exactly two parts. Everything before the final full stop is what happens: three sentences of plain narrative prose, no markup, no labels, no numbering. Everything after it is what it looks like: one long comma-separated clause carrying genre, environment, camera behaviour, coverage, particles, colour and mood. That division is not an accident of style. It maps onto how Alibaba's own prompt guidance composes a request — entity, scene and motion first, then aesthetic control and stylization — and it is the reason this reads as a template rather than as one writer's habit.

The thing worth noticing is what is not in it. There is no field protocol, no `description:` heading, no `[Shot 1]` marker, no angle-bracket asset tag and no negative prompt. Wan 3.0 accepts a single plain string up to twenty thousand characters and reads all of it as content. Anything that looks like structure is content too — which is why a prompt copied from a model that publishes a field protocol comes back with the field names rendered into the picture, and why the request succeeds while doing it.

The honest weakness is the beat count. Three things happen: she walks, she discovers, she removes the helmet while birds rise and the sun breaks. That is a thirty-second idea running in five, and it is worth watching exactly how the model absorbs it, because it is not what people predict. It does not drop beats. It cuts: a wide of the ruined avenue, a close-up of the child at the store window with the light in her hands, then the astronaut lifting the helmet off. Three shots in five seconds, one and a bit each. What actually gets dropped is everything the third sentence attaches to that beat — no birds rise, no sunlight breaks through the clouds — because a subordinate clause is the cheapest thing in a prompt for a model to skip.

That is the real lesson and it generalises. Over-writing a short Wan 3.0 request does not buy you a compressed version of your idea; it buys you a montage, assembled by the model, of whichever nouns were load-bearing enough to carry a shot. Verbs in main clauses survive. Atmosphere hung off them in subordinate clauses does not.

Two fixes, and they are different. If you want one shot, cut to one beat: she finds the child, and that is the whole thing. If you want this story, raise the duration to fifteen or thirty and write the order down — "then", "after that", "finally" — because Wan 3.0 has no timecode grammar and connectives are the entire mechanism for saying which thing happens second. An unplanned final third is where a long Wan 3.0 clip drifts, and the word "finally" is what stops it.

The camera instruction is worth arguing with. "Slow cinematic camera movement" gives a speed and no direction and no size, which leaves Wan 3.0 free to pick a move that suits whatever composition it lands on. On a shot this loose that is a reasonable trade. On a product shot it is not, and the fix is the same phrasing the prompt generator on this site emits: the camera pushes in slowly, a small move — a verb, a size and a speed in one clause.

One last thing this prompt gets right by omission: no brand, no on-screen text, no lettering. Ruined convenience stores are covered in signage in every reference image a model has ever seen, and asking for readable text is asking for the one thing generated video still renders badly.

Everything here runs in text to video.

6 questions

Astronaut in the ruins — common questions

  • 01

    Does Wan 3.0 need a structured prompt?

    No. It takes one plain string, up to twenty thousand characters, and reads every character as content. Field names, shot markers and timecodes are not parsed — they are rendered.

  • 02

    What happens when a five-second prompt has three beats in it?

    You get a montage. This run cuts three times to fit them in, and what it drops is the atmosphere hung off the last beat — the birds and the sun break never arrive. Main-clause verbs survive; subordinate clauses are the first thing skipped.

  • 03

    How do I exclude something without a negative prompt?

    Say it as a sentence inside the prompt: "no on-screen text", "no music". Wan 3.0 has no negative field, so anything formatted as one arrives as text to render.

  • 04

    Where does the sound come from if the prompt never mentions it?

    Wan 3.0 generates audio alongside the picture on every run. Silence is a thing you ask for, not the default, and an unwritten sound layer is a soundtrack you did not choose.

  • 05

    Will I get the same clip if I copy the prompt exactly?

    Close, not identical, unless you also fix the seed — and even then a different resolution changes the result. Expect the same register rather than the same frames.

  • 06

    What does one run of this cost?

    It is on the card: five seconds at 720P on the standard tier. The rate doubles at each resolution step, so the same prompt at 480P costs half as much — still past the free clip, which stops at 3 seconds.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 80 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated