Twelve seconds is an awkward length and this prompt is the clearest demonstration on the page of how to write for one.
Below about eight seconds a Wan 3.0 prompt should describe one action. There is no room to sequence anything, connectives buy you nothing, and the best results come from a single clear verb with a single camera move attached. Above about fifteen the opposite is true: the model has more time than the prompt has instructions, and if you have not said what happens in the last third it will invent something, which in practice means a slow drift outward and a fade. Twelve sits between those, which means the prompt has to sequence without having room for a story.
The solution here is to pick one shape and describe it arriving, dispersing and re-forming. Three beats, no plot. The dancers gather around a centre, they open outward, they close back in. Nothing happens in the sense that a story happens, and yet the clip has a beginning, a middle and an end, because the prompt names all three.
The camera instruction is the one people are most likely to argue with, so it is worth defending. It is a locked wide, and it says so. The instinct on a piece like this is to write a slow push in, because a push in is what a camera does when a filmmaker wants you to feel something — but a push in on five moving bodies means the model has to decide which of them stays in frame, and every second it spends on that decision is a second it is not spending on the fabric. Naming the frame as fixed removes the question. It also removes the drift a Wan 3.0 clip supplies for itself when the camera layer is left blank, which is the actual risk: an unwritten camera is not a still camera, it is a camera the model is choosing for you.
The lighting clause is the part worth stealing. "One hard backlight and no fill" is two instructions and one of them is a subtraction. Generative video models default to flattering, legible, evenly-lit frames, because that is what most of the footage they learned from looks like. Every high-contrast image you want has to be asked for by removing something, and the removal has to be explicit. "Backlit" on its own produces a backlight plus enough fill to see faces. "No fill" is what produces this.
The closing sound sentence is the other subtraction, and on Wan 3.0 it is not optional in the way it would be on a silent model. Audio is generated in the same pass as the picture. There is no mute flag and no separate audio prompt — silence is a thing you describe. This prompt describes it precisely: three named sounds and one named exclusion. Naming the sounds is also a timing instruction, because a footfall has to happen at a moment, and that is a second, weaker way of pinning the choreography.
One thing worth being honest about: this clip was published without its prompt, and the text above was written by reading the footage. That means it is a reconstruction, and running it will give you a shot of this kind rather than this shot. What survives reconstruction is the structure — three ordered beats, a lighting clause built on a subtraction, and a sound sentence that says what not to add. Those transfer to any subject. The particular dancers do not.