P—02 / STREET CHESS TURNAROUND

Wan 3.0 Prime prompt — a reversal in five seconds

A whole story with a turn in it, in five seconds, told entirely through one face.

The same two-part shape as the astronaut prompt, pointed at something much harder: a reversal. Someone is winning, someone else sits down, and the winning stops. What makes it fit inside five seconds is that the turn is written as an expression rather than as an event — nothing has to happen on the board for the shot to land. This is also the library's clearest look at what the Prime tier actually buys.

Mode
Text to Video
Model
Prime
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Square ambience, generated in the same pass
Credits
120

Alibaba’s published example for this model — not generated on this site. Source: wavespeed.ai

Output reference

A young man and an elderly woman play chess at a table in a busy city squareVideo
The clip
497 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A young man and an elderly woman play chess at a table in a busy city squareVideo
The clip
The promptthe clipWan 3.0 Prime · Text to video
497 chars

The turn

Setup, arrival, reversal — then a camera sentence that is a sequence rather than a schedule, and a three-word register.

A confident young man sits at a small chess table in a busy city square, quickly defeating several challengers while a crowd watches. An elegant elderly woman quietly sits down across from him and makes her first move. His confident smile slowly disappears as she begins outplaying him. The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer. Clever street comedy, expressive reactions, warm afternoon sunlight, cinematic realism.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    A confident young man … An elegant elderly woman

    One adjective each, and both adjectives are about bearing rather than appearance. "Confident" and "elegant" are castable; "brown-haired" would only have been renderable.

  2. Scene

    a small chess table in a busy city square … while a crowd watches

    The crowd is doing structural work. It gives the reversal an audience, which is what turns a game into a scene, and it gives the camera something to move through.

  3. Motion

    quietly sits down … makes her first move … His confident smile slowly disappears

    The only physical action is sitting and moving a piece. The turn itself is facial, which is why it fits in five seconds — a face can change in one.

  4. Aesthetic control

    The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer.

    A camera sentence with an order in it. "Starts with… then…" is how you sequence a Wan 3.0 shot; the model has no timecode grammar, so connectives are the whole mechanism.

  5. Stylization

    Clever street comedy, expressive reactions, warm afternoon sunlight, cinematic realism

    Naming the genre as comedy is what licenses the performances to be readable rather than subtle. Drop it and the same beats play as drama, which at five seconds reads as nothing at all.

  6. Sound

    Not written — left to the model.

    Unwritten again. A square full of people is one of the few settings where the invented ambience is usually fine — but "no music, just the square" would have removed the coin-flip.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    Where the turn lives

    On the face, not the board. "His confident smile slowly disappears" is the entire reversal, and it works because a smile is legible at any shot size. Move the turn onto the pieces and you need a close-up, a legible board and more seconds than you have.

  2. 02

    The two adjectives

    Confident and elegant. They are the casting brief and nothing else is given. Swap them for a pair with the same opposition — brash and unhurried, loud and precise — and the whole scene recasts without another word changing.

  3. 03

    The camera sequence

    Close-ups, then a circle out to both players. Reversing that order gives you the reveal before the setup, which is a different and worse film. This is the layer to keep verbatim while you change everything above it.

  4. 04

    The tier

    This one ran on Prime. Alibaba documents Prime as the fast version with the same output specs — same resolutions, same durations, same frame rate. It costs more per second and returns sooner. Nothing in this prompt requires it.

Three ways to break it

  • Writing the dialogue as quoted speech

    If you add a line, quotation marks get it narrated rather than performed. Wan 3.0 speaks what is inside braces: {Your move.} Everything outside the braces is description, including who is speaking and how.

  • Asking for a legible board position

    Chess pieces at square-market scale are already at the edge of what renders cleanly. Requiring a specific, readable position adds a constraint the model will satisfy badly at the cost of the faces, which are the shot.

  • Buying Prime for picture quality

    Prime is the high-speed tier. Its published output specification is identical to the standard model — 480P to 1080P, two to thirty seconds, thirty frames a second — and it costs half again as much per second. Paying the premium expecting a better picture is paying for a stopwatch.

What it does

What the street chess turnaround template does

Five seconds is not enough time for a story, and this prompt gets one anyway. Understanding how is worth more than the prompt itself.

The trick is where the reversal is put. A chess upset is, in reality, a board event: a move is made, a position collapses, someone realises. None of that is visible at a camera distance that also shows a city square, and all of it takes longer than five seconds. So the prompt moves the turn onto a face — "his confident smile slowly disappears" — and a face is legible in a single second at almost any shot size. The board never has to be readable. The crowd never has to react. One expression carries the entire dramatic content of the clip.

That is a general technique, not a chess technique. When a scene is too long for the duration you have, look for the beat that can be expressed as a change in someone's face, and write only that one. The rest becomes setup, and setup compresses freely.

The camera sentence is the other thing to copy. "The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer" is a sequence, and sequence is the only scheduling grammar Wan 3.0 has. There is no timecode syntax, no shot numbering, no way to say a move happens at 2.4 seconds. What there is: starts with, then, after that, finally. Those words are load-bearing, and the last one is the most load-bearing of all — a long Wan 3.0 clip that never says how it ends will invent an ending, and the invented one is usually a slow drift.

Now the part that is really about the model rather than the prompt. This example ran on `wan3.0-video-prime`, and the temptation on a page like this is to present the Prime examples as the good ones. They are not. Alibaba documents Prime as the high-speed version, and the published output specification is identical to the standard model: the same three resolutions, the same two-to-thirty-second range, the same thirty frames per second. It costs about half again as much per second and it returns sooner. That is the whole difference.

Which means the honest read of this page is: the astronaut ran on the standard tier, this one ran on Prime, and if you put them side by side you are looking at two different prompts, not two different quality levels. If a clip comes back worse than you wanted, the tier is not the dial. The duration is, and after that the prompt is.

One thing the prompt leaves out that is worth adding on a rerun: a sound sentence. A busy square generates a plausible bed on its own, but "no music, just square ambience and the pieces on the board" would have named the two sounds that matter and removed a score nobody asked for. On Wan 3.0 the audio layer is not decoration — it is generated in the same pass as the picture, so naming a sound also schedules the moment it happens.

Everything here runs in text to video.

6 questions

Street chess turnaround — common questions

  • 01

    Is Wan 3.0 Prime better quality than the standard model?

    No. Alibaba documents it as the high-speed tier, with an identical output specification — same resolutions, same two-to-thirty-second range, same thirty frames per second. It costs more per second and finishes sooner.

  • 02

    How do I fit a story into five seconds?

    Put the turn on a face. Physical events need setup and payoff; an expression changing is legible in a single second and needs neither.

  • 03

    How do I add a spoken line to this?

    Put the words in braces — {Your move.} — and leave the speaker and the delivery outside them. Quotation marks get the line narrated instead of performed.

  • 04

    Can I schedule the camera move to a specific second?

    No. Wan 3.0 has no timecode grammar. Order comes from connectives: starts with, then, after that, finally.

  • 05

    Would this work at fifteen or thirty seconds?

    Better, and it would need rewriting rather than just a bigger number. Add the beats you want and name them in order, including how it ends.

  • 06

    Does the crowd cost anything?

    Attention, not money. Every background figure is geometry the model has to keep coherent while the camera circles, which is why the two principals are described in one adjective each.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 120 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated