P—03 / THE DRAGON OPENS ITS EYES

Wan 3.0 image to video prompt — describe the change, not the picture

Two hundred and fifty-four characters, because the first frame had already said everything else.

The shortest prompt in this library, and short for a reason rather than out of laziness. A first frame already carries the creature, the forest, the light, the palette and the style — so every one of those words is either redundant or, worse, a second opinion the model has to reconcile with the picture in front of it. What is left to write is the only thing the frame cannot contain: what moves.

Mode
Image to Video
Model
Wan 3.0
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Forest bed, generated in the same pass
Credits
80

Alibaba’s published example for this model — not generated on this site. Source: wavespeed.ai

Output reference

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearingVideo
The clip

JPG · PNG · WebP · HEIC · 30 MB

254 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A small ranger reaches toward an enormous green dragon lying in a sunlit forest clearingVideo
The clip
The promptthe clipWan 3.0 · Image to video
254 chars

The wake

Four changes and one camera move. No creature description, no forest description, no style words at all.

The dragon slowly opens its eyes, leaves move from its breathing, glowing particles float around the forest, the ranger slowly steps forward and reaches out a hand. The camera slowly circles around both characters revealing the enormous scale difference.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    The dragon … the ranger

    Definite articles, no description. "The" tells the model these are the things already in the frame; "a dragon" would have invited it to produce a second one.

  2. Scene

    around the forest

    Three words, and only because the particles need somewhere to be. The frame is the scene. Re-describing it is how an image-to-video prompt ends up fighting its own reference.

  3. Motion

    slowly opens its eyes, leaves move from its breathing, glowing particles float … steps forward and reaches out a hand

    Four changes, ordered from smallest to largest, and one of them — leaves moving from breathing — is a second-order effect. Naming a consequence is how you get the model to animate the cause.

  4. Aesthetic control

    The camera slowly circles around both characters revealing the enormous scale difference.

    A move with a purpose attached. "Revealing the enormous scale difference" tells the model what the orbit is for, which constrains the radius and the height far better than a number would.

  5. Stylization

    Not written — left to the model.

    Absent, and correctly so. The reference frame is the style sheet. A style word here competes with the picture and the picture usually loses at the edges.

  6. Sound

    Not written — left to the model.

    Unwritten. A forest and a breathing dragon are enough for the model to invent something plausible, but the breath is the sound that matters and it went unnamed.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The consequence clause

    "Leaves move from its breathing" is the best sentence in this prompt. It never says the dragon breathes — it names a visible effect and lets the model work backwards to the cause. That reads as life in a way "the dragon breathes" does not, and it transfers to anything: steam bending, dust lifting, fabric settling.

  2. 02

    The purpose on the camera

    An orbit with a stated job. Change what it is revealing — the ranger's face, the depth of the clearing, what is behind the dragon — and the same verb produces a different path, without you having to specify radius, height or direction.

  3. 03

    What is deliberately missing

    No scales, no colour, no lighting, no lens, no genre. Add any of them and you are asking the model to reconcile your words with a frame that already disagrees. The most common way an image-to-video run goes wrong is a prompt that describes the picture.

  4. 04

    The order of the changes

    Eyes, then leaves, then particles, then the ranger. Smallest to largest, which lets five seconds build. Put the ranger first and the dragon waking becomes a reaction rather than the event.

Three ways to break it

  • Attaching reference images as well as a first frame

    Wan 3.0 treats first frame / last frame and the reference family as mutually exclusive, and mixing them is rejected outright with `InvalidParameter`. It is the single easiest way to build a request that cannot succeed, and the rejection arrives after the queue rather than before it.

  • Re-describing the subject

    Writing "a large green scaled dragon" beside a frame that already shows one gives the model two sources for the same fact. Where they differ it will average, and averaged detail is exactly the mushy look people blame on the model.

  • Asking for a cut

    An image-to-video run starts from your frame and stays continuous with it. A second location has nothing to be continuous with, so the model either ignores the instruction or dissolves — and five seconds does not have room for a dissolve.

What it does

What the the dragon opens its eyes template does

Two hundred and fifty-four characters against a twenty-thousand-character ceiling. This prompt uses roughly one per cent of what it is allowed, and it is the best-written thing in this library.

The reason is that image-to-video is a different job from text-to-video, and most people write it as though it were the same job with a picture attached. In text-to-video the prompt is the only source of truth, so it has to carry entity, scene, motion, aesthetic and style. In image-to-video the frame carries entity, scene, aesthetic and style already — in far more detail than prose can, and with no ambiguity. What the frame cannot carry is time. So the prompt has exactly one job: say what changes.

Look at what this text refuses to do. It never says the dragon is green, or large, or scaled. It never describes the forest beyond the three words needed to place the particles. It gives no lens, no grade, no genre and no mood. Every one of those would have been a second opinion about something already settled, and when a prompt and a reference frame disagree the model does not pick a winner — it interpolates. Interpolated detail is soft, and soft detail is what people mean when they say a run "looks AI".

The four changes are ordered smallest to largest: eyes, leaves, particles, the ranger stepping in. That ordering is what lets five seconds feel like it builds rather than like it happens all at once. It is also, quietly, an escalation of scale — an eyelid, then foliage, then the air, then a person crossing the frame — which is the same structural move the chess prompt makes with expressions.

The best line is the second one. "Leaves move from its breathing" never asks the dragon to breathe. It names a visible consequence and leaves the cause implied, which forces the model to produce the cause in order to justify the effect. That is a reliable technique and it generalises: steam bending over a pan, dust lifting off a road, a coat settling after someone stops walking. Naming the effect gets you the motion; naming the motion gets you an animation of the word.

The camera line does the same thing in a different register. "Slowly circles around both characters revealing the enormous scale difference" attaches a purpose to a move. Wan 3.0 has no parameters for orbit radius or camera height, and writing numbers into the prose would not create any — but "revealing the scale difference" implies a wide enough arc and a low enough angle to hold both bodies in frame, which is what those numbers would have been for.

The one thing to add on a rerun is sound. A breathing dragon in a quiet clearing is a sound design brief that writes itself, and Wan 3.0 generates the audio in the same pass as the picture — which means an audio sentence is also a timing instruction. "A low slow breath under forest ambience, no music" would have put the breath somewhere specific instead of leaving it to the model, and the breath is the beat the whole clip is built on.

Finally, the constraint that catches people on this mode specifically: a first frame and a reference image cannot travel in the same request. Wan 3.0 treats the keyframe family and the reference family as mutually exclusive and rejects the combination. If you want the dragon from one picture and the ranger from another, that is reference-to-video, and it is the next two templates on this page.

Everything here runs in image to video.

6 questions

The dragon opens its eyes — common questions

  • 01

    Why is an image-to-video prompt so much shorter?

    Because the frame already carries the subject, the setting, the light and the style. The only thing left for the text is what changes over time.

  • 02

    Should I describe what is in the picture?

    No. Where your words and the frame disagree the model interpolates, and interpolated detail is soft. Use definite articles — "the dragon" — to point at what is already there.

  • 03

    Can I attach a reference image as well as a first frame?

    No. Wan 3.0 rejects that combination with `InvalidParameter`. Keyframes and references are two families and a request may only use one.

  • 04

    How do I make something look alive rather than animated?

    Name a consequence instead of an action. "Leaves move from its breathing" produces breathing; "the dragon breathes" produces a chest moving up and down.

  • 05

    Can I add a last frame as well?

    Yes — first frame and last frame are in the same family and can travel together. That turns the prompt into a description of the route between two fixed points.

  • 06

    Does the camera move need a speed?

    It has one here — "slowly". Giving a move a speed and a purpose is enough; giving it neither is where a camera instruction stops constraining anything.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 80 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated