P—05 / DESERT TRAIN HEIST

Wan 3.0 reference to video prompt — four beats in five seconds

A whole second act written into a five-second request. Watch which beat survives.

She rides alongside the train, leaps onto it, fights across the roof, reaches a guarded carriage and cuts it loose. That is four beats and a location change, and the duration field says five. This template is in the library because the mismatch is the most common thing wrong with a Wan 3.0 request, and because Wan 3.0 is the one model where the fix is simply to ask for thirty seconds.

Mode
Reference to Video
Model
Wan 3.0
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Action bed, generated in the same pass
Credits
80

Alibaba’s published example for this model — not generated on this site. Source: wavespeed.ai

Output reference

A masked rider on a black horse gallops alongside a steam train across a sunset desertVideo
The clip
594 / 20,000

Images 0/10 · Clips 0/5 · Audio 0/5 · 15s of reference video max

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A masked rider on a black horse gallops alongside a steam train across a sunset desertVideo
The clip
The promptthe clipWan 3.0 · Reference to video
594 chars

The heist

Three sentences of escalating action, then a closing clause that is almost entirely camera language.

A steam train races across an endless desert at sunset while a masked female outlaw rides beside it on a black horse. She leaps onto the moving train, fights her way across the roof, and reaches a guarded carriage carrying a frightened young prisoner. As soldiers surround them, she cuts the carriage loose, sending it down a different track toward a distant canyon. Epic western action film, fast-paced camera movement, dynamic horseback tracking shots, dramatic close combat, sweeping aerial views, flying dust, golden sunset, practical stunt realism, anamorphic lens flares, cinematic scale.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    a masked female outlaw … a frightened young prisoner … soldiers

    Three parties, one adjective each, and the mask is doing quiet work — a covered face is a face the model cannot lose continuity on.

  2. Scene

    an endless desert at sunset … the moving train … a guarded carriage … a distant canyon

    Four locations, all of them on or beside one train. That is the only reason this is even arguably one shot rather than a sequence.

  3. Motion

    rides beside it … leaps onto … fights her way across the roof … reaches … cuts the carriage loose

    Five verbs, each of which is a beat in its own right. This is the layer that does not fit, and it is worth counting the verbs in your own prompts for exactly this reason.

  4. Aesthetic control

    fast-paced camera movement, dynamic horseback tracking shots, dramatic close combat, sweeping aerial views, flying dust … anamorphic lens flares

    Four different coverage types named in one clause. Like the verbs, this is a shot list pretending to be a style note.

  5. Stylization

    Epic western action film … golden sunset, practical stunt realism … cinematic scale

    "Practical stunt realism" is the best phrase in the prompt. It asks for weight and consequence rather than for fluid impossible motion, and weight is what makes generated action read.

  6. Sound

    Not written — left to the model.

    Unwritten, on a shot whose entire subject is a steam train. Hooves, couplings, wind and no score would have cost one sentence.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The duration field

    This is the real lever and it is not in the text. Wan 3.0 generates any whole number of seconds from two to thirty. A prompt with five beats in it wants thirty, and thirty is a number the previous generation of models could not offer at all.

  2. 02

    The mask

    A masked protagonist is a continuity cheat and a good one. There is no face to hold consistent across a leap, a fight and a cut, so the model spends that attention on the motion instead. Use it whenever the identity is not the point.

  3. 03

    "Practical stunt realism"

    Three words that ask for weight. They are why the horse and the train read as heavy objects rather than as animated shapes, and they transfer to anything with physics in it — a fall, a door, a collision.

  4. 04

    The coverage list

    Tracking, close combat, aerials. Keep one and the five seconds has a shot; keep all three and it has a trailer it cannot cut. If you raise the duration, keep all three and attach one to each beat.

Three ways to break it

  • Leaving five seconds in the duration field

    Wan 3.0 does not truncate a prompt that is too long for its runtime — it compresses it. Every beat arrives early and rushed, or most of them are dropped and the model picks which. Either way the result is decided by the model rather than by you.

  • Adding timecodes instead of raising the duration

    Writing "0-8s" and "8-16s" into the prose does nothing except spend characters. Wan 3.0 has no timecode grammar. Order comes from connectives — then, after that, finally — and length comes from the duration parameter.

  • Attaching a first frame to give the outlaw a face

    Reference material and keyframes are mutually exclusive in Wan 3.0, and mixing them is rejected with `InvalidParameter` after the request has queued. If you need a specific face here, it belongs in `reference_image_urls` with the prompt citing `Image 1`.

What it does

What the desert train heist template does

Count the verbs. Rides, leaps, fights, reaches, cuts. Five beats, three of which involve a change of position on a moving train, in a request whose duration field says five seconds. That is one second per beat, and one second is not enough time for a viewer to register that a beat happened, let alone for the model to stage it.

What Wan 3.0 does with a prompt like this is worth understanding precisely, because it is not what people expect. It does not truncate. It does not run out of time and stop. It compresses the schedule into the runtime it was given, which means every beat is attempted and every beat is rushed — or, at this ratio, most of them are silently dropped and the model chooses which one to build the clip around. In the published run it chose the first: the ride alongside the train, in dust, at sunset. It is a good shot and it is not the shot the prompt describes.

This is the most common failure in the category and, on Wan 3.0 specifically, the easiest one to fix. The model generates any whole number of seconds from two to thirty. A five-beat action sequence wants the top of that range. Raising the number costs six times the credits and buys a clip that contains the thing you wrote — which is a better trade than five runs of a compressed version.

What raising the number does not do on its own is give the beats an order. There is no timecode syntax in Wan 3.0. Writing "0-6s: she rides alongside" spends characters and buys nothing; the model reads it as text. The grammar that does work is ordinary English connectives — then, after that, finally — and of those the last matters most. A thirty-second clip whose prompt never says how it ends will invent an ending, and the invented ending is almost always a slow drift out of the action rather than a cut on the strongest image. Here the ending is already written into the story: the carriage rolling away down a different track. It just needs the word "finally" in front of it.

Two things in this prompt are quietly excellent and worth taking whatever you do with the rest. The first is the mask. A masked protagonist removes the hardest continuity problem in the shot — there is no face to keep stable across a leap and a fight — and hands that capacity back to the motion. Cover a face any time the identity is not the point.

The second is "practical stunt realism". Generated action tends toward the weightless: bodies that float through arcs, impacts with no follow-through, horses that glide. Those three words ask for the opposite, and they do it by naming a production method rather than a visual property, which is a much stronger instruction. "Realistic" is a wish. "Practical stunt" is a way of shooting, and the model has seen a great deal of footage shot that way.

One structural note about the mode. This is reference-to-video, which on Wan 3.0 means up to ten images, five video clips and five audio clips in a single request — and a piece of real syntax to go with them. Assets are cited in the prompt as `Image 1`, `Video 1`, `Audio 1`: capitalised, with a space, numbered by their order in the array. This prompt uses none of it, which is fine for a published sample and wrong for a production request. If the outlaw, the train and the canyon are things you have pictures of, say so by number, and the model stops inventing them.

Last, the bill. Reference video is the one input Wan 3.0 charges for: a source clip is billed at the same per-second rate as the output and counted against the same thirty-second ceiling. Ten reference images are free and a fifteen-second reference video is the most expensive thing you can attach to a request. Nobody writing about this model has said so plainly, and it is the line item that surprises people.

Everything here runs in reference to video.

6 questions

Desert train heist — common questions

  • 01

    What happens when a prompt describes more than the duration allows?

    Wan 3.0 compresses rather than truncates. Every beat is attempted and rushed, or most are dropped and the model chooses which one to keep.

  • 02

    How long can a Wan 3.0 clip be?

    Any whole number of seconds from two to thirty. A five-beat sequence like this one wants the top of that range, not the bottom.

  • 03

    Do timecodes help order the beats?

    No. There is no timecode grammar. Use connectives — then, after that, finally — and put the ending in words, because an unwritten ending gets invented.

  • 04

    How do I cite an attached image in the prompt?

    As `Image 1`, capitalised with a space, numbered by position in the array. Lower case or glued-together spellings will not bind to anything.

  • 05

    Do reference materials cost extra?

    Reference video does — it is billed at the output rate and counts against the thirty-second total. Reference images, audio, documents and links do not.

  • 06

    Can I add a first frame to lock the opening composition?

    Not in the same request. Keyframes and reference materials are mutually exclusive and the combination is rejected outright.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 80 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated