The interesting problem in this prompt is coverage, not story. Three shot sizes in fifteen seconds, in a model that has no way to express a cut.
Every field-protocol video model — the ones that publish a shot list schema, a timecode grammar, angle-bracket asset tags — makes coverage explicit. Wan 3.0 does none of that. It reads one plain string, and anything that looks like a shot marker is content to be rendered rather than structure to be parsed. So the question is what plain English reliably produces a change in shot size, and the answer turns out to be surprisingly narrow: either you say what the camera does, or you say what the frame now contains.
This prompt uses both. "The camera pushes to an extreme close-up of her face behind the glass" is the first kind — an instruction to the camera, with the destination size named. "Finally the frame opens out into a wide concrete plaza where a green-skinned armoured brute twice her height is coming at her" is the second — the frame is not told to move, it is told what is in it, and a body twice her height with room to charge cannot be in a close-up. Both work. The second is more robust, because it constrains the composition by content rather than by a term the model has to interpret.
What does not work is asking for a cut. "Cut to", "next shot", "then we see" are all read as prose. The model does something with them, and what it does is usually a whip pan or a dissolve, because those are the transitions that exist inside continuous footage. If you want the feel of a cut, the honest options are to write a strong enough content change that the model has to jump, or to generate two clips and cut them yourself.
The ending is written, and on a fifteen-second request that is not optional. "The blast throws the brute back through the paving in a wall of dust, and the shot holds on her standing alone in the rubble" assigns the last beat and names the frame it stops on. Leave that off and the last third of a Wan 3.0 clip goes to whatever the model thinks a fight resolves into, which is reliably a slow pull back. Notice also where the beam comes from: the core in her chest, not a hand. A weapon needs an origin or the model puts the light wherever the pose happens to allow, and an emitter named once in the prompt is an emitter that stays in the same place for the rest of the shot.
The second thing worth studying is the sound layer, which is written as a mirror of the picture layer. Street noise, then servos, then rumble — three sounds in the same order as the three beats, with the same connective. On a model that generates audio and picture in one pass this is not decoration; it is a second, independent statement of the running order, phrased in a different modality. When the two agree, timing gets noticeably tighter. It is the cheapest reliability trick available on long Wan 3.0 requests and it is in almost nobody's prompt.
The colour instruction deserves a note of its own, because it is the most transferable clause here. "Saturated candy pink against concrete grey" names both sides. A prompt that asks only for pink tends to return a frame that is pink everywhere — the light, the walls, the grade — at which point nothing is pink, because pink is a relationship. Naming the ground against which a colour sits is how you get a colour that reads. This works for every palette instruction and it is worth making a habit of.
Finally, the honest caveat that applies to every template in this group: the clip was published without a prompt, and this text was written by reading the footage. It will get you a clip of this kind. If you need this exact character to survive a shot-size change, text mode is the wrong tool regardless of how the prompt is written — that is what reference mode and a character image are for, and there are three examples of it elsewhere in this library.