P—03 / THE DRAGON OPENS ITS EYES

Wan 3.0 image to video prompt — describe the change, not the picture

Two hundred and fifty-four characters, because the first frame had already said everything else.

The shortest prompt in this library, and short for a reason rather than out of laziness. A first frame already carries the creature, the forest, the light, the palette and the style — so every one of those words is either redundant or, worse, a second opinion the model has to reconcile with the picture in front of it. What is left to write is the only thing the frame cannot contain: what moves.

模式
Image to Video
模型
Wan 3.0
时长
5s
比例
16:9
分辨率
720P
音频
Forest bed, generated in the same pass
积分
80

阿里巴巴为这个模型公布的示例 —— 不是在本站生成的。 来源:wavespeed.ai

Output reference

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearing视频
这条片子

JPG · PNG · BMP · WebP · ≤20 MB · 240–8000px · ratio ≤8:1

254 / 20,000

免费那条就是同一个 Wan 3.0,声音也有,上限是 480P · 3s,走的是共用队列。方案能把上限抬高到 1080P 和三十秒 —— 免费档是一个分辨率,不是另一个模型。

1 条免费 · WAN 3.0 · 480P

01 —— 内部

一条提示词,一套流程

片子摆在做出它的那条提示词旁边,整条印出来。把它复制进上面的控制台,改你需要改的,然后生成。

参考产出

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearing视频
片子
这条提示词the clipWan 3.0 · Image to video
254 chars

The wake

Four changes and one camera move. No creature description, no forest description, no style words at all.

The dragon slowly opens its eyes, leaves move from its breathing, glowing particles float around the forest, the ranger slowly steps forward and reaches out a hand. The camera slowly circles around both characters revealing the enormous scale difference.

6 层

每一层在做什么

Wan 3.0 读的是一整条白话字符串,但它是顺着接缝读的。拿走一层,而不是整条提示词 —— 美学和声音那两行几乎能换到任何主体上,而它们正是大多数人写得最差的两行。

  1. Entity

    The dragon … the ranger

    Definite articles, no description. "The" tells the model these are the things already in the frame; "a dragon" would have invited it to produce a second one.

  2. Scene

    around the forest

    Three words, and only because the particles need somewhere to be. The frame is the scene. Re-describing it is how an image-to-video prompt ends up fighting its own reference.

  3. Motion

    slowly opens its eyes, leaves move from its breathing, glowing particles float … steps forward and reaches out a hand

    Four changes, ordered from smallest to largest, and one of them — leaves moving from breathing — is a second-order effect. Naming a consequence is how you get the model to animate the cause.

  4. Aesthetic control

    The camera slowly circles around both characters revealing the enormous scale difference.

    A move with a purpose attached. "Revealing the enormous scale difference" tells the model what the orbit is for, which constrains the radius and the height far better than a number would.

  5. Stylization

    没写 —— 留给模型。

    Absent, and correctly so. The reference frame is the style sheet. A style word here competes with the picture and the picture usually loses at the edges.

  6. Sound

    没写 —— 留给模型。

    Unwritten. A forest and a breathing dragon are enough for the model to invent something plausible, but the breath is the sound that matters and it went unnamed.

4 个可调项

把它变成你的

哪些改动是安全的。多数库只发布这一份清单,这也是为什么那么多复制来的提示词跑出来比原作还差。

  1. 01

    The consequence clause

    "Leaves move from its breathing" is the best sentence in this prompt. It never says the dragon breathes — it names a visible effect and lets the model work backwards to the cause. That reads as life in a way "the dragon breathes" does not, and it transfers to anything: steam bending, dust lifting, fabric settling.

  2. 02

    The purpose on the camera

    An orbit with a stated job. Change what it is revealing — the ranger's face, the depth of the clearing, what is behind the dragon — and the same verb produces a different path, without you having to specify radius, height or direction.

  3. 03

    What is deliberately missing

    No scales, no colour, no lighting, no lens, no genre. Add any of them and you are asking the model to reconcile your words with a frame that already disagrees. The most common way an image-to-video run goes wrong is a prompt that describes the picture.

  4. 04

    The order of the changes

    Eyes, then leaves, then particles, then the ranger. Smallest to largest, which lets five seconds build. Put the ranger first and the dragon waking becomes a reaction rather than the event.

三种把它弄坏的方式

  • Attaching reference images as well as a first frame

    Wan 3.0 treats first frame / last frame and the reference family as mutually exclusive, and mixing them is rejected outright with `InvalidParameter`. It is the single easiest way to build a request that cannot succeed, and the rejection arrives after the queue rather than before it.

  • Re-describing the subject

    Writing "a large green scaled dragon" beside a frame that already shows one gives the model two sources for the same fact. Where they differ it will average, and averaged detail is exactly the mushy look people blame on the model.

  • Asking for a cut

    An image-to-video run starts from your frame and stays continuous with it. A second location has nothing to be continuous with, so the model either ignores the instruction or dissolves — and five seconds does not have room for a dissolve.

它做什么

the dragon opens its eyes 这个模板做什么

Two hundred and fifty-four characters against a twenty-thousand-character ceiling. This prompt uses roughly one per cent of what it is allowed, and it is the best-written thing in this library.

The reason is that image-to-video is a different job from text-to-video, and most people write it as though it were the same job with a picture attached. In text-to-video the prompt is the only source of truth, so it has to carry entity, scene, motion, aesthetic and style. In image-to-video the frame carries entity, scene, aesthetic and style already — in far more detail than prose can, and with no ambiguity. What the frame cannot carry is time. So the prompt has exactly one job: say what changes.

Look at what this text refuses to do. It never says the dragon is green, or large, or scaled. It never describes the forest beyond the three words needed to place the particles. It gives no lens, no grade, no genre and no mood. Every one of those would have been a second opinion about something already settled, and when a prompt and a reference frame disagree the model does not pick a winner — it interpolates. Interpolated detail is soft, and soft detail is what people mean when they say a run "looks AI".

The four changes are ordered smallest to largest: eyes, leaves, particles, the ranger stepping in. That ordering is what lets five seconds feel like it builds rather than like it happens all at once. It is also, quietly, an escalation of scale — an eyelid, then foliage, then the air, then a person crossing the frame — which is the same structural move the chess prompt makes with expressions.

The best line is the second one. "Leaves move from its breathing" never asks the dragon to breathe. It names a visible consequence and leaves the cause implied, which forces the model to produce the cause in order to justify the effect. That is a reliable technique and it generalises: steam bending over a pan, dust lifting off a road, a coat settling after someone stops walking. Naming the effect gets you the motion; naming the motion gets you an animation of the word.

The camera line does the same thing in a different register. "Slowly circles around both characters revealing the enormous scale difference" attaches a purpose to a move. Wan 3.0 has no parameters for orbit radius or camera height, and writing numbers into the prose would not create any — but "revealing the scale difference" implies a wide enough arc and a low enough angle to hold both bodies in frame, which is what those numbers would have been for.

The one thing to add on a rerun is sound. A breathing dragon in a quiet clearing is a sound design brief that writes itself, and Wan 3.0 generates the audio in the same pass as the picture — which means an audio sentence is also a timing instruction. "A low slow breath under forest ambience, no music" would have put the breath somewhere specific instead of leaving it to the model, and the breath is the beat the whole clip is built on.

Finally, the constraint that catches people on this mode specifically: a first frame and a reference image cannot travel in the same request. Wan 3.0 treats the keyframe family and the reference family as mutually exclusive and rejects the combination. If you want the dragon from one picture and the ranger from another, that is reference-to-video, and it is the next two templates on this page.

这里的一切都在image to video里跑。

6 个问题

The dragon opens its eyes —— 常见问题

  • 01

    Why is an image-to-video prompt so much shorter?

    Because the frame already carries the subject, the setting, the light and the style. The only thing left for the text is what changes over time.

  • 02

    Should I describe what is in the picture?

    No. Where your words and the frame disagree the model interpolates, and interpolated detail is soft. Use definite articles — "the dragon" — to point at what is already there.

  • 03

    Can I attach a reference image as well as a first frame?

    No. Wan 3.0 rejects that combination with `InvalidParameter`. Keyframes and references are two families and a request may only use one.

  • 04

    How do I make something look alive rather than animated?

    Name a consequence instead of an action. "Leaves move from its breathing" produces breathing; "the dragon breathes" produces a chest moving up and down.

  • 05

    Can I add a last frame as well?

    Yes — first frame and last frame are in the same family and can travel together. That turns the prompt into a description of the route between two fixed points.

  • 06

    Does the camera move need a speed?

    It has one here — "slowly". Giving a move a speed and a purpose is enough; giving it neither is where a camera instruction stops constraining anything.

复制它,改一层,跑起来。

每一个字符都在这一页上。720P 的 5s 要 80 积分。

由 wan-3.run 编辑团队撰写与维护发布于 最后更新