P—02 / STREET CHESS TURNAROUND

Wan 3.0 Prime prompt — a reversal in five seconds

A whole story with a turn in it, in five seconds, told entirely through one face.

The same two-part shape as the astronaut prompt, pointed at something much harder: a reversal. Someone is winning, someone else sits down, and the winning stops. What makes it fit inside five seconds is that the turn is written as an expression rather than as an event — nothing has to happen on the board for the shot to land. This is also the library's clearest look at what the Prime tier actually buys.

模式
Text to Video
模型
Prime
長度
5s
比例
16:9
解析度
720P
音訊
Square ambience, generated in the same pass
點數
120

阿里巴巴為這個模型公布的示例 —— 不是在本站生成的。 來源:wavespeed.ai

Output reference

A young man and an elderly woman play chess at a table in a busy city square影片
這支片子
497 / 20,000

免費那支就是同一個 Wan 3.0,聲音也有,上限是 480P · 3s,走的是共用佇列。方案能把上限抬高到 1080P 和三十秒 —— 免費檔是一個解析度,不是另一個模型。

1 支免費 · WAN 3.0 · 480P

01 —— 內部

一條提示詞,一套流程

片子擺在做出它的那條提示詞旁邊,整條印出來。把它複製進上面的控制台,改你需要改的,然後生成。

參考產出

A young man and an elderly woman play chess at a table in a busy city square影片
片子
這條提示詞the clipWan 3.0 Prime · Text to video
497 chars

The turn

Setup, arrival, reversal — then a camera sentence that is a sequence rather than a schedule, and a three-word register.

A confident young man sits at a small chess table in a busy city square, quickly defeating several challengers while a crowd watches. An elegant elderly woman quietly sits down across from him and makes her first move. His confident smile slowly disappears as she begins outplaying him. The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer. Clever street comedy, expressive reactions, warm afternoon sunlight, cinematic realism.

6 層

每一層在做什麼

Wan 3.0 讀的是一整條白話字串,但它是順著接縫讀的。拿走一層,而不是整條提示詞 —— 美學和聲音那兩行幾乎能換到任何主體上,而它們正是大多數人寫得最差的兩行。

  1. Entity

    A confident young man … An elegant elderly woman

    One adjective each, and both adjectives are about bearing rather than appearance. "Confident" and "elegant" are castable; "brown-haired" would only have been renderable.

  2. Scene

    a small chess table in a busy city square … while a crowd watches

    The crowd is doing structural work. It gives the reversal an audience, which is what turns a game into a scene, and it gives the camera something to move through.

  3. Motion

    quietly sits down … makes her first move … His confident smile slowly disappears

    The only physical action is sitting and moving a piece. The turn itself is facial, which is why it fits in five seconds — a face can change in one.

  4. Aesthetic control

    The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer.

    A camera sentence with an order in it. "Starts with… then…" is how you sequence a Wan 3.0 shot; the model has no timecode grammar, so connectives are the whole mechanism.

  5. Stylization

    Clever street comedy, expressive reactions, warm afternoon sunlight, cinematic realism

    Naming the genre as comedy is what licenses the performances to be readable rather than subtle. Drop it and the same beats play as drama, which at five seconds reads as nothing at all.

  6. Sound

    沒寫 —— 留給模型。

    Unwritten again. A square full of people is one of the few settings where the invented ambience is usually fine — but "no music, just the square" would have removed the coin-flip.

4 個可調項

把它變成你的

哪些改動是安全的。多數庫只發布這一份清單,這也是為什麼那麼多複製來的提示詞跑出來比原作還差。

  1. 01

    Where the turn lives

    On the face, not the board. "His confident smile slowly disappears" is the entire reversal, and it works because a smile is legible at any shot size. Move the turn onto the pieces and you need a close-up, a legible board and more seconds than you have.

  2. 02

    The two adjectives

    Confident and elegant. They are the casting brief and nothing else is given. Swap them for a pair with the same opposition — brash and unhurried, loud and precise — and the whole scene recasts without another word changing.

  3. 03

    The camera sequence

    Close-ups, then a circle out to both players. Reversing that order gives you the reveal before the setup, which is a different and worse film. This is the layer to keep verbatim while you change everything above it.

  4. 04

    The tier

    This one ran on Prime. Alibaba documents Prime as the fast version with the same output specs — same resolutions, same durations, same frame rate. It costs more per second and returns sooner. Nothing in this prompt requires it.

三種把它弄壞的方式

  • Writing the dialogue as quoted speech

    If you add a line, write the exact words and mark them off from the description — {Your move.} in the form this site emits, or "Your move." in the form Alibaba's own guide uses. Everything outside them stays description, including who is speaking and how. A paraphrase is the one thing that does not work: it gets you a line the model invented.

  • Asking for a legible board position

    Chess pieces at square-market scale are already at the edge of what renders cleanly. Requiring a specific, readable position adds a constraint the model will satisfy badly at the cost of the faces, which are the shot.

  • Buying Prime for picture quality

    Prime is the high-speed tier. Its published output specification is identical to the standard model — 480P to 1080P, two to thirty seconds, thirty frames a second — and it costs half again as much per second. Paying the premium expecting a better picture is paying for a stopwatch.

它做什麼

street chess turnaround 這個範本做什麼

Five seconds is not enough time for a story, and this prompt gets one anyway. Understanding how is worth more than the prompt itself.

The trick is where the reversal is put. A chess upset is, in reality, a board event: a move is made, a position collapses, someone realises. None of that is visible at a camera distance that also shows a city square, and all of it takes longer than five seconds. So the prompt moves the turn onto a face — "his confident smile slowly disappears" — and a face is legible in a single second at almost any shot size. The board never has to be readable. The crowd never has to react. One expression carries the entire dramatic content of the clip.

That is a general technique, not a chess technique. When a scene is too long for the duration you have, look for the beat that can be expressed as a change in someone's face, and write only that one. The rest becomes setup, and setup compresses freely.

The camera sentence is the other thing to copy. "The camera starts with close-ups of the chess pieces, then gently circles around both players as the crowd gathers closer" is a sequence, and sequence is the only scheduling grammar Wan 3.0 has. There is no timecode syntax, no shot numbering, no way to say a move happens at 2.4 seconds. What there is: starts with, then, after that, finally. Those words are load-bearing, and the last one is the most load-bearing of all — a long Wan 3.0 clip that never says how it ends will invent an ending, and the invented one is usually a slow drift.

Now the part that is really about the model rather than the prompt. This example ran on `wan3.0-video-prime`, and the temptation on a page like this is to present the Prime examples as the good ones. They are not. Alibaba documents Prime as the high-speed version, and the published output specification is identical to the standard model: the same three resolutions, the same two-to-thirty-second range, the same thirty frames per second. It costs about half again as much per second and it returns sooner. That is the whole difference.

Which means the honest read of this page is: the astronaut ran on the standard tier, this one ran on Prime, and if you put them side by side you are looking at two different prompts, not two different quality levels. If a clip comes back worse than you wanted, the tier is not the dial. The duration is, and after that the prompt is.

One thing the prompt leaves out that is worth adding on a rerun: a sound sentence. A busy square generates a plausible bed on its own, but "no music, just square ambience and the pieces on the board" would have named the two sounds that matter and removed a score nobody asked for. On Wan 3.0 the audio layer is not decoration — it is generated in the same pass as the picture, so naming a sound also schedules the moment it happens.

這裡的一切都在text to video裡跑。

6 個問題

Street chess turnaround —— 常見問題

  • 01

    Is Wan 3.0 Prime better quality than the standard model?

    No. Alibaba documents it as the high-speed tier, with an identical output specification — same resolutions, same two-to-thirty-second range, same thirty frames per second. It costs more per second and finishes sooner.

  • 02

    How do I fit a story into five seconds?

    Put the turn on a face. Physical events need setup and payoff; an expression changing is legible in a single second and needs neither.

  • 03

    How do I add a spoken line to this?

    Write the exact words and mark them off — {Your move.} — with the speaker and the delivery outside them. Quotation marks do the same job and are what Alibaba's published prompts use. What does not work is summarising: "she says something reassuring" gets you a line the model wrote.

  • 04

    Can I schedule the camera move to a specific second?

    No. Wan 3.0 has no timecode grammar. Order comes from connectives: starts with, then, after that, finally.

  • 05

    Would this work at fifteen or thirty seconds?

    Better, and it would need rewriting rather than just a bigger number. Add the beats you want and name them in order, including how it ends.

  • 06

    Does the crowd cost anything?

    Attention, not money. Every background figure is geometry the model has to keep coherent while the camera circles, which is why the two principals are described in one adjective each.

複製它,改一層,跑起來。

每一個字元都在這一頁上。720P 的 5s 要 120 點。

由 wan-3.run 編輯團隊撰寫與維護發布於 最後更新