Two hundred and fifty-four characters against a twenty-thousand-character ceiling. This prompt uses roughly one per cent of what it is allowed, and it is the best-written thing in this library.
The reason is that image-to-video is a different job from text-to-video, and most people write it as though it were the same job with a picture attached. In text-to-video the prompt is the only source of truth, so it has to carry entity, scene, motion, aesthetic and style. In image-to-video the frame carries entity, scene, aesthetic and style already — in far more detail than prose can, and with no ambiguity. What the frame cannot carry is time. So the prompt has exactly one job: say what changes.
Look at what this text refuses to do. It never says the dragon is green, or large, or scaled. It never describes the forest beyond the three words needed to place the particles. It gives no lens, no grade, no genre and no mood. Every one of those would have been a second opinion about something already settled, and when a prompt and a reference frame disagree the model does not pick a winner — it interpolates. Interpolated detail is soft, and soft detail is what people mean when they say a run "looks AI".
The four changes are ordered smallest to largest: eyes, leaves, particles, the ranger stepping in. That ordering is what lets five seconds feel like it builds rather than like it happens all at once. It is also, quietly, an escalation of scale — an eyelid, then foliage, then the air, then a person crossing the frame — which is the same structural move the chess prompt makes with expressions.
The best line is the second one. "Leaves move from its breathing" never asks the dragon to breathe. It names a visible consequence and leaves the cause implied, which forces the model to produce the cause in order to justify the effect. That is a reliable technique and it generalises: steam bending over a pan, dust lifting off a road, a coat settling after someone stops walking. Naming the effect gets you the motion; naming the motion gets you an animation of the word.
The camera line does the same thing in a different register. "Slowly circles around both characters revealing the enormous scale difference" attaches a purpose to a move. Wan 3.0 has no parameters for orbit radius or camera height, and writing numbers into the prose would not create any — but "revealing the scale difference" implies a wide enough arc and a low enough angle to hold both bodies in frame, which is what those numbers would have been for.
The one thing to add on a rerun is sound. A breathing dragon in a quiet clearing is a sound design brief that writes itself, and Wan 3.0 generates the audio in the same pass as the picture — which means an audio sentence is also a timing instruction. "A low slow breath under forest ambience, no music" would have put the breath somewhere specific instead of leaving it to the model, and the breath is the beat the whole clip is built on.
Finally, the constraint that catches people on this mode specifically: a first frame and a reference image cannot travel in the same request. Wan 3.0 treats the keyframe family and the reference family as mutually exclusive and rejects the combination. If you want the dragon from one picture and the ranger from another, that is reference-to-video, and it is the next two templates on this page.