This is the longest and most complete prompt in the library, and it is worth reading twice — once for what it does well, which is most of it, and once for the one thing it leaves out.
Start with the structure. Five paragraphs, separated by blank lines: staging, the exchange, the turn, the camera plan, the lock. Nothing is numbered, no beat has a timecode, and the order is carried entirely by the sequence in which the paragraphs appear. That is the correct shape for Wan 3.0. The model has no schedule grammar — there is no way to say a beat lands at 3.5 seconds — so the only ordering available is the order of the words. Paragraphs make that order visible to the person editing the prompt as well as to the model, which is why this reads more clearly than the run-on comma lists elsewhere on this page.
The camera paragraph is the best-constructed piece of writing in the library. It has three moves in order — begins with a wide, tracks beside her, settles into a two-shot — and then a fourth sentence that most prompts never write: "End with both characters illuminated only by the cyan star map and the orange energy core." That is the ending, stated. Wan 3.0 will produce an ending whether or not you specify one, and the unspecified version is almost always a drift: the camera keeps easing, the action peters out, the last frame is nobody's choice. Writing the final image down is one sentence and it is the highest-value sentence in any long prompt.
The lock paragraph is doing the job the mode exists for. "Maintain the exact identities, hairstyles, facial features, outfits, equipment, colors, and proportions of both reference characters" is an enumeration rather than an adjective, and that matters. "Keep the characters consistent" gives the model a goal it can claim to have met; a list of seven named attributes gives it seven things it can be measured against. Enumerating what must not change is the reliable form.
And now the gap. Wan 3.0 has exactly one piece of syntax that it parses rather than reads: an asset citation. Materials attached to a reference request are addressed in the prompt as `Image 1`, `Video 1`, `Audio 1` — capitalised, with a space, numbered by their position in the array. That is how a sentence gets bound to a specific uploaded file. This prompt never uses it. It says "both reference characters" and trusts the model to work out which description belongs to which upload.
With two visually distinct characters that mostly works. With four, or with two people who share a hair colour, it does not — the model averages the descriptions across the references, and the result is a cast who all look faintly like each other. The fix is two sentences at the top: `Image 1 is the silver-haired navigator. Image 2 is the black-coated operative.` And then every later mention becomes "the navigator from Image 1". It costs about forty characters and it converts an inference problem into a lookup.
The limits are worth having in mind while you write. Ten reference images, five reference videos, five reference audio clips, with the video and audio families each capped at fifteen seconds in total. Reference material and keyframes cannot travel together. And the one that shows up on the invoice rather than in the response: reference video is billed at the same per-second rate as the output and counted against the same thirty-second ceiling, while images, audio, documents and links are free. A fifteen-second reference video is the most expensive thing you can attach to a Wan 3.0 request, and it is the fact this page would most like you to know before you attach one.
This example ran on Prime. As on the other two Prime templates: that is the high-speed tier, with the same published output specification as the standard model, at roughly half again the cost per second. It came back sooner. The picture is not the reason to pick it.