P—03 / THE DRAGON OPENS ITS EYES

Wan 3.0 image to video prompt — describe the change, not the picture

Two hundred and fifty-four characters, because the first frame had already said everything else.

The shortest prompt in this library, and short for a reason rather than out of laziness. A first frame already carries the creature, the forest, the light, the palette and the style — so every one of those words is either redundant or, worse, a second opinion the model has to reconcile with the picture in front of it. What is left to write is the only thing the frame cannot contain: what moves.

モード
Image to Video
モデル
Wan 3.0
5s
縦横比
16:9
解像度
720P
音声
Forest bed, generated in the same pass
クレジット
80

このモデルについてアリババが公開している作例です — 本サイトで生成したものではありません。 出典:wavespeed.ai

Output reference

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearing動画
クリップ

JPG · PNG · BMP · WebP · ≤20 MB · 240–8000px · ratio ≤8:1

254 / 20,000

無料クリップは音声つきの同じ Wan 3.0 で、上限は 480P · 3s、共有の順番待ちに入ります。プランで上限が上がり、1080P と 30 秒になります — 無料ティアは解像度であって、別のモデルではありません。

無料クリップ 1 本 · WAN 3.0 · 480P

01 — 中身

プロンプト 1 本、ワークフロー 1 本

それを生んだプロンプトの全文と、その横にクリップ。上のコンソールにコピーして、必要なところを変えて、生成してください。

参照出力

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearing動画
クリップ
プロンプトthe clipWan 3.0 · Image to video
254 chars

The wake

Four changes and one camera move. No creature description, no forest description, no style words at all.

The dragon slowly opens its eyes, leaves move from its breathing, glowing particles float around the forest, the ranger slowly steps forward and reaches out a hand. The camera slowly circles around both characters revealing the enormous scale difference.

6 層

各層が何をしているか

Wan 3.0 が読むのは 1 本のただの文字列ですが、それを継ぎ目に沿って読みます。プロンプト全体ではなく層を持ち帰ってください——美的制御と音の 2 行はほとんどどんな題材にも移せて、しかもその 2 つが最も書き方を外されます。

  1. Entity

    The dragon … the ranger

    Definite articles, no description. "The" tells the model these are the things already in the frame; "a dragon" would have invited it to produce a second one.

  2. Scene

    around the forest

    Three words, and only because the particles need somewhere to be. The frame is the scene. Re-describing it is how an image-to-video prompt ends up fighting its own reference.

  3. Motion

    slowly opens its eyes, leaves move from its breathing, glowing particles float … steps forward and reaches out a hand

    Four changes, ordered from smallest to largest, and one of them — leaves moving from breathing — is a second-order effect. Naming a consequence is how you get the model to animate the cause.

  4. Aesthetic control

    The camera slowly circles around both characters revealing the enormous scale difference.

    A move with a purpose attached. "Revealing the enormous scale difference" tells the model what the orbit is for, which constrains the radius and the height far better than a number would.

  5. Stylization

    書かれていません — モデルに任せています。

    Absent, and correctly so. The reference frame is the style sheet. A style word here competes with the picture and the picture usually loses at the edges.

  6. Sound

    書かれていません — モデルに任せています。

    Unwritten. A forest and a breathing dragon are enough for the model to invent something plausible, but the breath is the sound that matters and it went unnamed.

4 個のレバー

自分のものにする

変えて安全なのはどこか。多くのライブラリはこの一覧しか出さず、それがコピーしたプロンプトが元より悪くなる理由です。

  1. 01

    The consequence clause

    "Leaves move from its breathing" is the best sentence in this prompt. It never says the dragon breathes — it names a visible effect and lets the model work backwards to the cause. That reads as life in a way "the dragon breathes" does not, and it transfers to anything: steam bending, dust lifting, fabric settling.

  2. 02

    The purpose on the camera

    An orbit with a stated job. Change what it is revealing — the ranger's face, the depth of the clearing, what is behind the dragon — and the same verb produces a different path, without you having to specify radius, height or direction.

  3. 03

    What is deliberately missing

    No scales, no colour, no lighting, no lens, no genre. Add any of them and you are asking the model to reconcile your words with a frame that already disagrees. The most common way an image-to-video run goes wrong is a prompt that describes the picture.

  4. 04

    The order of the changes

    Eyes, then leaves, then particles, then the ranger. Smallest to largest, which lets five seconds build. Put the ranger first and the dragon waking becomes a reaction rather than the event.

壊してしまう三つの書き換え

  • Attaching reference images as well as a first frame

    Wan 3.0 treats first frame / last frame and the reference family as mutually exclusive, and mixing them is rejected outright with `InvalidParameter`. It is the single easiest way to build a request that cannot succeed, and the rejection arrives after the queue rather than before it.

  • Re-describing the subject

    Writing "a large green scaled dragon" beside a frame that already shows one gives the model two sources for the same fact. Where they differ it will average, and averaged detail is exactly the mushy look people blame on the model.

  • Asking for a cut

    An image-to-video run starts from your frame and stays continuous with it. A second location has nothing to be continuous with, so the model either ignores the instruction or dissolves — and five seconds does not have room for a dissolve.

このテンプレートの働き

the dragon opens its eyes テンプレートがしていること

Two hundred and fifty-four characters against a twenty-thousand-character ceiling. This prompt uses roughly one per cent of what it is allowed, and it is the best-written thing in this library.

The reason is that image-to-video is a different job from text-to-video, and most people write it as though it were the same job with a picture attached. In text-to-video the prompt is the only source of truth, so it has to carry entity, scene, motion, aesthetic and style. In image-to-video the frame carries entity, scene, aesthetic and style already — in far more detail than prose can, and with no ambiguity. What the frame cannot carry is time. So the prompt has exactly one job: say what changes.

Look at what this text refuses to do. It never says the dragon is green, or large, or scaled. It never describes the forest beyond the three words needed to place the particles. It gives no lens, no grade, no genre and no mood. Every one of those would have been a second opinion about something already settled, and when a prompt and a reference frame disagree the model does not pick a winner — it interpolates. Interpolated detail is soft, and soft detail is what people mean when they say a run "looks AI".

The four changes are ordered smallest to largest: eyes, leaves, particles, the ranger stepping in. That ordering is what lets five seconds feel like it builds rather than like it happens all at once. It is also, quietly, an escalation of scale — an eyelid, then foliage, then the air, then a person crossing the frame — which is the same structural move the chess prompt makes with expressions.

The best line is the second one. "Leaves move from its breathing" never asks the dragon to breathe. It names a visible consequence and leaves the cause implied, which forces the model to produce the cause in order to justify the effect. That is a reliable technique and it generalises: steam bending over a pan, dust lifting off a road, a coat settling after someone stops walking. Naming the effect gets you the motion; naming the motion gets you an animation of the word.

The camera line does the same thing in a different register. "Slowly circles around both characters revealing the enormous scale difference" attaches a purpose to a move. Wan 3.0 has no parameters for orbit radius or camera height, and writing numbers into the prose would not create any — but "revealing the scale difference" implies a wide enough arc and a low enough angle to hold both bodies in frame, which is what those numbers would have been for.

The one thing to add on a rerun is sound. A breathing dragon in a quiet clearing is a sound design brief that writes itself, and Wan 3.0 generates the audio in the same pass as the picture — which means an audio sentence is also a timing instruction. "A low slow breath under forest ambience, no music" would have put the breath somewhere specific instead of leaving it to the model, and the breath is the beat the whole clip is built on.

Finally, the constraint that catches people on this mode specifically: a first frame and a reference image cannot travel in the same request. Wan 3.0 treats the keyframe family and the reference family as mutually exclusive and rejects the combination. If you want the dragon from one picture and the ranger from another, that is reference-to-video, and it is the next two templates on this page.

ここにあるものはすべてimage to videoで動きます。

6 個の質問

The dragon opens its eyes — よくある質問

  • 01

    Why is an image-to-video prompt so much shorter?

    Because the frame already carries the subject, the setting, the light and the style. The only thing left for the text is what changes over time.

  • 02

    Should I describe what is in the picture?

    No. Where your words and the frame disagree the model interpolates, and interpolated detail is soft. Use definite articles — "the dragon" — to point at what is already there.

  • 03

    Can I attach a reference image as well as a first frame?

    No. Wan 3.0 rejects that combination with `InvalidParameter`. Keyframes and references are two families and a request may only use one.

  • 04

    How do I make something look alive rather than animated?

    Name a consequence instead of an action. "Leaves move from its breathing" produces breathing; "the dragon breathes" produces a chest moving up and down.

  • 05

    Can I add a last frame as well?

    Yes — first frame and last frame are in the same family and can travel together. That turns the prompt into a description of the route between two fixed points.

  • 06

    Does the camera move need a speed?

    It has one here — "slowly". Giving a move a speed and a purpose is enough; giving it neither is where a camera instruction stops constraining anything.

コピーして、層を一つ変えて、走らせる。

一文字残らずこのページにあります。720P の 5s で 80 クレジット。

wan-3.run 編集チームが執筆・管理しています公開 最終更新