Five seconds is not enough time for a story, and this prompt gets one anyway. Understanding how is worth more than the prompt itself.
The trick is where the reversal is put. A chess upset is, in reality, a board event: a move is made, a position collapses, someone realises. None of that is visible at a camera distance that also shows a city square, and all of it takes longer than five seconds. So the prompt moves the turn onto a face — "his confident smile slowly disappears" — and a face is legible in a single second at almost any shot size. The board never has to be readable. The crowd never has to react. One expression carries the entire dramatic content of the clip.
That is a general technique, not a chess technique. When a scene is too long for the duration you have, look for the beat that can be expressed as a change in someone's face, and write only that one. The rest becomes setup, and setup compresses freely.
The camera sentence is the other thing to copy. "The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer" is a sequence, and sequence is the only scheduling grammar Wan 3.0 has. There is no timecode syntax, no shot numbering, no way to say a move happens at 2.4 seconds. What there is: starts with, then, after that, finally. Those words are load-bearing, and the last one is the most load-bearing of all — a long Wan 3.0 clip that never says how it ends will invent an ending, and the invented one is usually a slow drift.
Now the part that is really about the model rather than the prompt. This example ran on `wan3.0-video-prime`, and the temptation on a page like this is to present the Prime examples as the good ones. They are not. Alibaba documents Prime as the high-speed version, and the published output specification is identical to the standard model: the same three resolutions, the same two-to-thirty-second range, the same thirty frames per second. It costs about half again as much per second and it returns sooner. That is the whole difference.
Which means the honest read of this page is: the astronaut ran on the standard tier, this one ran on Prime, and if you put them side by side you are looking at two different prompts, not two different quality levels. If a clip comes back worse than you wanted, the tier is not the dial. The duration is, and after that the prompt is.
One thing the prompt leaves out that is worth adding on a rerun: a sound sentence. A busy square generates a plausible bed on its own, but "no music, just square ambience and the pieces on the board" would have named the two sounds that matter and removed a score nobody asked for. On Wan 3.0 the audio layer is not decoration — it is generated in the same pass as the picture, so naming a sound also schedules the moment it happens.