P—07 / MEADOW SOFA PORTRAIT

Wan 3.0 reference to video prompt — citing an image by number

The only prompt here that uses the one piece of syntax Wan 3.0 actually parses.

Both published reference examples describe their characters in prose and never cite the attached files at all, which works and also throws away the mechanism. This one cites. `Image 1`, capitalised, with a space, followed by an explicit instruction to hold the wardrobe — and then almost nothing else happens, because the point of a reference shot is usually that the subject stays exactly as delivered while the world moves around it.

Mode
Reference to Video
Model
Wan 3.0
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Meadow ambience, explicitly no music
Credits
80

Reference render — not generated on this site. Prompt read back out of the footage, not quoted from the run. Source: wan30.co

Output reference

A woman in white sits on a lime-green sofa in a flower meadow as a butterfly lands on her handVideo
The clip
944 / 20,000

Images 0/10 · Clips 0/5 · Audio 0/5 · 15s of reference video max

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A woman in white sits on a lime-green sofa in a flower meadow as a butterfly lands on her handVideo
The clip
The promptthe clipWan 3.0 · Reference to video
944 chars

The held portrait

A numbered citation, one small movement, and an explicit instruction that the reference is not to be reinterpreted.

Use Image 1 as the performer. She is reclining across a lime-green inflatable modular sofa in the middle of a wildflower meadow, wearing a white hooded coverall with a ruffled panel down the front, white trousers and pale platform shoes, with a small pale-blue object tucked against her hip. She looks straight down the lens and does not move, and only her head turns slightly as a large orange butterfly settles on her outstretched arm. More butterflies drift through the flowers around the sofa. Keep her face, hair and the whole outfit exactly as they appear in Image 1. Fashion editorial film, flat frontal composition with the sofa centred, soft overcast light with no visible shadows, a heavily saturated grade that pushes the grass and the sky toward the same green, an almost imperceptible slow push in, shallow depth of field on the foreground flowers, fine grain. Wind through grass and one distant bird, no music and nobody speaking.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    Use Image 1 as the performer … Keep her face, hair and the whole outfit exactly as they appear in Image 1.

    The citation appears twice: once to assign the role and once to lock the appearance. The second is not redundant — assigning a reference tells the model who this is, and it will still restyle the wardrobe unless told not to.

  2. Scene

    a lime-green inflatable modular sofa in the middle of a wildflower meadow

    One incongruous object in one natural setting. The whole composition is that contrast, and it is why the shot survives having almost no action in it.

  3. Motion

    She looks straight down the lens and does not move, and only her head turns slightly as a large orange butterfly settles on her outstretched arm.

    A stated stillness plus one exception. "Does not move" is an instruction rather than an omission — leave it out and a five-second clip fills with fidgeting, blinking and hair drift.

  4. Aesthetic control

    flat frontal composition with the sofa centred, soft overcast light with no visible shadows … an almost imperceptible slow push in, shallow depth of field on the foreground flowers

    Symmetry, flat light and a move small enough to be deniable. Each one removes a variable, which is what a portrait needs and what an action shot cannot afford.

  5. Stylization

    Fashion editorial film … a heavily saturated grade that pushes the grass and the sky toward the same green

    The grade is described as a relationship between two things rather than as a colour. Naming what the saturation does to the sky is what produces the flat poster-like frame instead of a bright ordinary meadow.

  6. Sound

    Wind through grass and one distant bird, no music and nobody speaking.

    Two named sounds and two refusals. A near-still fashion frame is the most reliable way in the whole library to get an ambient score, so the exclusion has to be written even though nothing in the picture suggests audio at all.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The second citation

    "Keep her face, hair and the whole outfit exactly as they appear in Image 1." Assigning a reference is not the same as locking it. Drop this sentence and the face usually survives while the clothes get reinterpreted, which on a wardrobe-led shot is the whole loss.

  2. 02

    The stated stillness

    "Does not move" plus one named exception. This is the pattern for any shot where a person has to hold — a portrait, a product held in a hand, a piece to camera. Without it the model animates, because animating is what it is for.

  3. 03

    The one moving element

    A butterfly landing. A five-second clip with nothing moving at all reads as a still image with grain on it; one small, soft, unhurried motion is enough to make it a shot.

  4. 04

    The colour relationship

    Grass and sky pushed toward the same green. Describing a grade as what it does to two named things is far more reliable than naming a look, and it is the clause most worth reusing here.

Three ways to break it

  • Writing the citation as `image1`

    The syntax is case-sensitive and takes a space. `image1`, `Image_1` and "the first picture" all bind to nothing, and the failure is silent — the request succeeds and the model invents a person, which is much worse than an error.

  • Numbering past what you attached

    Reference mode accepts up to ten images, five video clips and five audio clips. Citing `Image 3` when you attached two is a citation with no referent, and again there is no error — the model fills the gap with something plausible.

  • Adding action to a reference portrait

    Every additional movement is another opportunity for the reference to drift. If the shot exists to show a specific person or a specific garment, the correct amount of action is the minimum that stops it reading as a still.

What it does

What the meadow sofa portrait template does

Reference mode is the part of Wan 3.0 that has no equivalent in most of the models it competes with, and it is also the part people use least well. This template exists to show the syntax in use, because neither of the two published examples in this library uses it at all.

The mechanism is simple and easy to get wrong. Attached materials are an ordered array, and the prompt refers to them positionally: `Image 1`, `Video 2`, `Audio 1`. Capitalised, with a space, numbered from one by position in the array. `image1` binds to nothing. `Image_1` binds to nothing. "The first reference picture" binds to nothing. And the important part: none of those produce an error. The request returns 200 and the model, having found no binding, invents a subject. You get a finished clip of the wrong person.

The second thing this prompt demonstrates is that assigning a reference and locking a reference are two different instructions. "Use Image 1 as the performer" tells the model whose face this is. It does not tell the model that the white hooded jacket, the ruffled skirt and the platform shoes are also part of what it is being given. In practice a single citation reliably preserves identity and unreliably preserves wardrobe, which for a fashion frame is the entire subject. The explicit second sentence — keep the face, the hair and the whole outfit exactly as they appear — is what closes that gap.

The third thing is the stillness, and it is the least obvious. Video models animate. That is what they are for, and given five seconds and a seated figure they will produce five seconds of small movement: a blink, a breath, a hand adjusting, hair settling. Every one of those is a moment where the reference can drift. So a reference portrait wants an explicit "does not move", and it wants exactly one named exception so that the clip is not a still with grain on it. Here that exception is a butterfly landing on an already-outstretched arm — soft, slow, and involving no change of pose.

It is worth naming what this shot gives up. There is no story, no reversal, no camera work to speak of. Five seconds and one butterfly. The trade is deliberate: reference mode is at its most reliable when the subject is asked to persist rather than to perform, and a great deal of commercial work — lookbooks, product-in-hand, character sheets, talking-head intros — is exactly that shape. Understanding that reference mode rewards restraint is more valuable than any individual clause in this prompt.

On the grade: "a heavily saturated grade that pushes the grass and the sky toward the same green" is the most transferable sentence here. It describes the look as an operation on two named elements rather than as a style word. Style words are ambiguous — "dreamy", "editorial", "cinematic" all mean six things — whereas an instruction that says which two things should end up the same colour has exactly one reading. Whenever a look is hard to name, try describing what it does to two objects in the frame instead.

The honest caveat, same as everywhere in this library that the label says reference render: the clip was published without a prompt and this text was written by reading it. What it demonstrates about citation syntax, wardrobe locking and stated stillness is true of Wan 3.0 regardless. What it will not do is reproduce this particular woman, because we do not have the image that was attached.

Everything here runs in reference to video.

6 questions

Meadow sofa portrait — common questions

  • 01

    What is the exact syntax for citing a reference?

    Capitalised, with a space, numbered by position in the array you attached: Image 1, Video 2, Audio 1. Lower case, underscores and prose descriptions all bind to nothing, and the request still succeeds.

  • 02

    Does citing an image keep the clothes as well as the face?

    Not reliably. Identity usually survives one citation; wardrobe often does not. Add a sentence that names what has to stay — face, hair, outfit — and say it should match the reference exactly.

  • 03

    How much can I attach?

    Up to ten images, five video clips and five audio clips in one request. Reference video is billed at the output rate and counts against the total runtime; images and audio are not billed.

  • 04

    Why tell the model the subject does not move?

    Because otherwise it animates, and every extra movement is a chance for the reference to drift. State the stillness and give it one small exception so the clip still reads as a shot.

  • 05

    Can I mix reference images with a first and last frame?

    No. The reference family and the first-and-last-frame family are mutually exclusive in one request. Sending both is a request that cannot succeed, and it is worth knowing before you build the payload.

  • 06

    Is this clip a Wan 3.0 render?

    It was published without a prompt or a model attribution, so the card labels it a reference render and the prompt above is a reconstruction. The syntax it demonstrates is Wan 3.0's regardless.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 80 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated