P—08 / DANCERS IN BACKLIT FOG

Wan 3.0 prompt for twelve seconds — an ensemble that holds its shape

Five bodies, one light, and a shape that gathers, opens and closes because the prompt says it in that order.

The shortest prompt on this page that needs connectives. Twelve seconds is past the point where Wan 3.0 will hold one pose for you and short of the point where you can tell a story, so the whole job is a group finding a shape, losing it and finding it again — and the only reason it lands is that the prompt says which of those three comes first. It is also the library's clearest lesson in lighting: there is exactly one light in this prompt and no fill, and that single decision does more work than every other clause combined.

モード
Text to Video
モデル
Wan 3.0
12s
縦横比
16:9
解像度
720P
音声
Breath and floor, explicitly no music
クレジット
192

参照用のレンダー — 本サイトで生成したものではありません。 プロンプトは映像から読み戻したものです。実行時のものを引用したわけではありません。 出典:wan30.co

Output reference

Five dancers in black dresses silhouetted against a hard white backlight in ground fog動画
クリップ
876 / 20,000

無料クリップは音声つきの同じ Wan 3.0 で、上限は 480P · 3s、共有の順番待ちに入ります。プランで上限が上がり、1080P と 30 秒になります — 無料ティアは解像度であって、別のモデルではありません。

無料クリップ 1 本 · WAN 3.0 · 480P

01 — 中身

プロンプト 1 本、ワークフロー 1 本

それを生んだプロンプトの全文と、その横にクリップ。上のコンソールにコピーして、必要なところを変えて、生成してください。

参照出力

Five dancers in black dresses silhouetted against a hard white backlight in ground fog動画
クリップ
プロンプトthe clipWan 3.0 · Text to video
876 chars

The phrase

Three ordered beats in plain prose, a closing clause that is almost entirely about one light, and a sound sentence that names what to leave out.

Five dancers in long black chiffon dresses stand in a shallow bank of ground fog, lit from directly behind by a single hard white source so that every body reads as a silhouette. They begin gathered tight, one figure at the centre with an arm lifted and the other four folded low around her. Then the group opens outward, each dancer extending into a line of her own while the skirts swing a beat behind the bodies. Finally they draw back in around the centre with their arms raised together and hold the shape while the fog closes over their feet. Contemporary dance film, one hard backlight and no fill, deep black foreground, heavy low-lying haze, chiffon moving a beat behind the body, a locked wide that keeps all five inside the frame, near-monochrome blue-grey grade, fine grain. The only sound is breath, bare feet on a wooden floor and the rustle of fabric, no music.

6 層

各層が何をしているか

Wan 3.0 が読むのは 1 本のただの文字列ですが、それを継ぎ目に沿って読みます。プロンプト全体ではなく層を持ち帰ってください——美的制御と音の 2 行はほとんどどんな題材にも移せて、しかもその 2 つが最も書き方を外されます。

  1. Entity

    Five dancers in long black chiffon dresses

    A count, a garment and a fabric. No faces, no ages, no casting — at silhouette scale none of it would be visible, and every adjective spent on a face here is an adjective taken away from the fog.

  2. Scene

    a shallow bank of ground fog

    Five words, and they are the set. The fog is not atmosphere here, it is the floor: it hides the feet, which is what lets the group read as one mass before it separates.

  3. Motion

    They begin gathered tight, one figure at the centre with an arm lifted … Then the group opens outward, each dancer extending into a line of her own … Finally they draw back in around the centre

    Three beats and no plot. Gather, open, close — a shape arriving, dispersing and re-forming, which is as much structure as twelve seconds can carry and considerably more than most prompts of this length attempt.

  4. Aesthetic control

    one hard backlight and no fill, deep black foreground, heavy low-lying haze, chiffon moving a beat behind the body, a locked wide that keeps all five inside the frame

    The lighting clause is doing almost all of the work. "No fill" is the instruction that makes silhouettes; without it the model lights the faces and the whole idea collapses into a rehearsal video.

  5. Stylization

    Contemporary dance film … near-monochrome blue-grey grade, fine grain

    A genre and a grade, and the grade is stated as a restriction rather than a colour. "Near-monochrome" keeps the model from finding a warm skin tone somewhere and reaching for it.

  6. Sound

    The only sound is breath, bare feet on a wooden floor and the rustle of fabric, no music.

    The first prompt in the library to write this layer at all. "No music" is a sentence rather than a field, because Wan 3.0 has no field for it, and a dance clip is the one place a generated score arrives uninvited every single time.

4 個のレバー

自分のものにする

変えて安全なのはどこか。多くのライブラリはこの一覧しか出さず、それがコピーしたプロンプトが元より悪くなる理由です。

  1. 01

    The number of dancers

    Five is the smallest group where a centre and an outside both read at once. Drop to three and "the other four folded low around her" has nowhere to happen; go past seven and the silhouettes start merging into one dark mass at this shot size.

  2. 02

    The locked frame

    A static wide is a decision, and stating it is what stops the model adding a drift of its own. On an ensemble the camera has nothing to contribute that the bodies are not already doing, and any move competes with the choreography for the same twelve seconds.

  3. 03

    The fabric lag

    "Chiffon moving a beat behind the body" is the single clause that makes the movement read as real. It transfers to anything that hangs — a coat, a flag, hair — and it is worth keeping verbatim when you change everything else.

  4. 04

    The runtime

    Twelve seconds is the length this shape needs: roughly four to gather, four to open, four to close. Take it to six and the beats compress until the held shape disappears, which is the only part anybody remembers.

壊してしまう三つの書き換え

  • Adding a fill light

    Anything that lights the front of the dancers — a rim, a practical, "soft ambient light" — turns silhouettes into people, and people at this distance in this much haze look like an under-exposed rehearsal. The absence of fill is the shot.

  • Removing the sound sentence

    Wan 3.0 generates audio on every run whether or not you ask. Delete that last sentence and this comes back with a piano score under it, because a clip full of dancers is what a model has seen scored a hundred thousand times.

  • Numbering the beats

    Rewriting the middle as a numbered list, or attaching times to it, does not schedule anything — Wan 3.0 has no timecode grammar and no list parser, so the numbers are just characters it has to make sense of. "Then" and "finally" are the entire mechanism.

このテンプレートの働き

dancers in backlit fog テンプレートがしていること

Twelve seconds is an awkward length and this prompt is the clearest demonstration on the page of how to write for one.

Below about eight seconds a Wan 3.0 prompt should describe one action. There is no room to sequence anything, connectives buy you nothing, and the best results come from a single clear verb with a single camera move attached. Above about fifteen the opposite is true: the model has more time than the prompt has instructions, and if you have not said what happens in the last third it will invent something, which in practice means a slow drift outward and a fade. Twelve sits between those, which means the prompt has to sequence without having room for a story.

The solution here is to pick one shape and describe it arriving, dispersing and re-forming. Three beats, no plot. The dancers gather around a centre, they open outward, they close back in. Nothing happens in the sense that a story happens, and yet the clip has a beginning, a middle and an end, because the prompt names all three.

The camera instruction is the one people are most likely to argue with, so it is worth defending. It is a locked wide, and it says so. The instinct on a piece like this is to write a slow push in, because a push in is what a camera does when a filmmaker wants you to feel something — but a push in on five moving bodies means the model has to decide which of them stays in frame, and every second it spends on that decision is a second it is not spending on the fabric. Naming the frame as fixed removes the question. It also removes the drift a Wan 3.0 clip supplies for itself when the camera layer is left blank, which is the actual risk: an unwritten camera is not a still camera, it is a camera the model is choosing for you.

The lighting clause is the part worth stealing. "One hard backlight and no fill" is two instructions and one of them is a subtraction. Generative video models default to flattering, legible, evenly-lit frames, because that is what most of the footage they learned from looks like. Every high-contrast image you want has to be asked for by removing something, and the removal has to be explicit. "Backlit" on its own produces a backlight plus enough fill to see faces. "No fill" is what produces this.

The closing sound sentence is the other subtraction, and on Wan 3.0 it is not optional in the way it would be on a silent model. Audio is generated in the same pass as the picture. There is no mute flag and no separate audio prompt — silence is a thing you describe. This prompt describes it precisely: three named sounds and one named exclusion. Naming the sounds is also a timing instruction, because a footfall has to happen at a moment, and that is a second, weaker way of pinning the choreography.

One thing worth being honest about: this clip was published without its prompt, and the text above was written by reading the footage. That means it is a reconstruction, and running it will give you a shot of this kind rather than this shot. What survives reconstruction is the structure — three ordered beats, a lighting clause built on a subtraction, and a sound sentence that says what not to add. Those transfer to any subject. The particular dancers do not.

ここにあるものはすべてtext to videoで動きます。

6 個の質問

Dancers in backlit fog — よくある質問

  • 01

    Why twelve seconds and not fifteen?

    Wan 3.0 takes any whole number from two to thirty, so twelve is as valid as fifteen. This shape needs roughly equal thirds to gather, open and close, and twelve is the shortest runtime where the held shape still gets a beat of its own.

  • 02

    How do I stop the model scoring it?

    Write the sound layer, including the exclusion, as a sentence: "no music". There is no audio toggle and no negative field, and audio is generated in the same pass as the picture, so an unwritten soundtrack is a soundtrack somebody else chose.

  • 03

    Is "no fill" really necessary?

    Yes, and it is the single most load-bearing phrase in the prompt. Asking only for a backlight reliably returns a backlight plus enough front light to read faces, because that is what most footage looks like. High contrast has to be requested as a subtraction.

  • 04

    Why lock the camera instead of pushing in?

    Because there is no timecode grammar to schedule a move against, and on five moving bodies a move costs the model attention it is currently spending on the fabric. Saying the frame is locked also stops it inventing a drift, which is what an unwritten camera layer gets you.

  • 05

    Was this clip made with this prompt?

    No. It was published without one, and this prompt was written by reading the footage — the same job the video-to-prompt tool on this site does. Expect a shot of this kind, not this shot.

  • 06

    What does twelve seconds cost against five?

    Wan 3.0 meters per second from the first one, so twelve seconds is a little over twice a five-second run at the same resolution. The exact credit figure is on the card and it is computed, not typed.

似ているもの

アップロードした素材から始める、ほかのプロンプト

フレームかリファレンス素材を持ち込んでください。プロンプトが書くのは、すでにあるものではなく変わるものです。

プロンプト例をすべて見る

コピーして、層を一つ変えて、走らせる。

一文字残らずこのページにあります。720P の 12s で 192 クレジット。

wan-3.run 編集チームが執筆・管理しています公開 最終更新