Count the verbs. Rides, leaps, fights, reaches, cuts. Five beats, three of which involve a change of position on a moving train, in a request whose duration field says five seconds. That is one second per beat, and one second is not enough time for a viewer to register that a beat happened, let alone for the model to stage it.
What Wan 3.0 does with a prompt like this is worth understanding precisely, because it is not what people expect. It does not truncate. It does not run out of time and stop. It compresses the schedule into the runtime it was given, which means every beat is attempted and every beat is rushed — or, at this ratio, most of them are silently dropped and the model chooses which one to build the clip around. In the published run it chose the first: the ride alongside the train, in dust, at sunset. It is a good shot and it is not the shot the prompt describes.
This is the most common failure in the category and, on Wan 3.0 specifically, the easiest one to fix. The model generates any whole number of seconds from two to thirty. A five-beat action sequence wants the top of that range. Raising the number costs six times the credits and buys a clip that contains the thing you wrote — which is a better trade than five runs of a compressed version.
What raising the number does not do on its own is give the beats an order. There is no timecode syntax in Wan 3.0. Writing "0-6s: she rides alongside" spends characters and buys nothing; the model reads it as text. The grammar that does work is ordinary English connectives — then, after that, finally — and of those the last matters most. A thirty-second clip whose prompt never says how it ends will invent an ending, and the invented ending is almost always a slow drift out of the action rather than a cut on the strongest image. Here the ending is already written into the story: the carriage rolling away down a different track. It just needs the word "finally" in front of it.
Two things in this prompt are quietly excellent and worth taking whatever you do with the rest. The first is the mask. A masked protagonist removes the hardest continuity problem in the shot — there is no face to keep stable across a leap and a fight — and hands that capacity back to the motion. Cover a face any time the identity is not the point.
The second is "practical stunt realism". Generated action tends toward the weightless: bodies that float through arcs, impacts with no follow-through, horses that glide. Those three words ask for the opposite, and they do it by naming a production method rather than a visual property, which is a much stronger instruction. "Realistic" is a wish. "Practical stunt" is a way of shooting, and the model has seen a great deal of footage shot that way.
One structural note about the mode. This is reference-to-video, which on Wan 3.0 means up to ten images, five video clips and five audio clips in a single request — and a piece of real syntax to go with them. Assets are cited in the prompt as `Image 1`, `Video 1`, `Audio 1`: capitalised, with a space, numbered by their order in the array. This prompt uses none of it, which is fine for a published sample and wrong for a production request. If the outlaw, the train and the canyon are things you have pictures of, say so by number, and the model stops inventing them.
Last, the bill. Reference video is the one input Wan 3.0 charges for: a source clip is billed at the same per-second rate as the output and counted against the same thirty-second ceiling. Ten reference images are free and a fifteen-second reference video is the most expensive thing you can attach to a request. Nobody writing about this model has said so plainly, and it is the line item that surprises people.