P—10 / SCHOOLGIRL TO MECH SUIT

Wan 3.0 prompt for a scale change — street level to city block

The same person in three shot sizes, and the prompt says which size each beat is.

Fifteen seconds is enough for three shot sizes, and getting three shot sizes is a different problem from getting three beats. Wan 3.0 has no cut syntax and no shot list, so the only way to change coverage mid-clip is to write it as something the camera does or as a change in what the frame contains. This prompt does both, and it is the clearest example on the page of moving from a wide to an extreme close-up and back out without a single field name.

Mode
Text to Video
Model
Wan 3.0
Duration
15s
Ratio
16:9
Resolution
720P
Audio
Street, servos, then a low rumble
Credits
240

Reference render — not generated on this site. Prompt read back out of the footage, not quoted from the run. Source: wan30.co

Output reference

A pink mech suit firing a beam at a green armoured brute across a concrete plazaVideo
The clip
1,198 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A pink mech suit firing a beam at a green armoured brute across a concrete plazaVideo
The clip
The promptthe clipWan 3.0 · Text to video
1,198 chars

The suit-up

Four beats, each one carrying its own shot size, and a colour instruction that is really a contrast instruction.

A teenage girl in an oversized grey sweater and a pleated skirt walks across a campus forecourt in the middle of the day, other students in uniform passing without looking at her; her right arm is a glossy pink prosthetic. Then she stops and armour plates unfold from that arm and travel across her body until she is inside a full pink mech suit, the visor closing last. After that the camera pushes to an extreme close-up of her face behind the glass, calm, a lollipop stick at the corner of her mouth. Finally the frame opens out into a wide concrete plaza where a green-skinned armoured brute twice her height is coming at her, she plants both feet and fires a hot pink beam from the core in her chest, the blast throws the brute back through the paving in a wall of dust, and the shot holds on her standing alone in the rubble. Anime realism, bright overcast daylight, saturated candy pink against concrete grey, glossy moulded plastic and painted metal, handheld at street level for the walk and locked off for the stand-off, shallow depth of field on the close-up, no lens flare. Street noise and footsteps first, then servo whine and plate locks, then a low rumble under the brute, no music.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    A teenage girl in an oversized grey sweater and a pleated skirt … her right arm is a glossy pink prosthetic

    The prosthetic is planted in the first sentence so the transformation has somewhere to come from. Without it the armour arrives from nowhere and the beat reads as a cut to a different character.

  2. Scene

    a campus forecourt in the middle of the day, other students in uniform passing without looking at her

    The indifferent students are the scene. They establish that this is ordinary before anything stops being ordinary, which is the entire setup done in one subordinate clause.

  3. Motion

    armour plates unfold from that arm and travel across her body until she is inside a full pink mech suit, the visor closing last

    "The visor closing last" is an ordering instruction inside a single beat. It gives the transformation a finish line, which is what stops it looking like a fade between two costumes.

  4. Aesthetic control

    handheld at street level for the walk and locked off for the stand-off, shallow depth of field on the close-up, no lens flare

    Camera behaviour assigned per beat rather than to the clip. This is how you get three shot sizes out of a model with no cut syntax: attach the coverage to the moment, not to the request.

  5. Stylization

    Anime realism, bright overcast daylight, saturated candy pink against concrete grey, glossy moulded plastic and painted metal

    The colour instruction names both sides of the contrast. Asking for pink alone gets a pink-tinted frame; asking for pink against grey gets pink that reads as pink.

  6. Sound

    Street noise and footsteps first, then servo whine and plate locks, then a low rumble under the brute, no music.

    The word "then" appears twice in the sound layer, mirroring the picture layer exactly. Audio is generated in the same pass, so a running order written twice is a running order the model is twice as likely to keep.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The three shot sizes

    Wide, extreme close-up, wide again. They are written as "walks across", "the camera pushes to an extreme close-up" and "the frame opens out into a wide". Those three phrasings are the whole vocabulary you need for coverage in a model with no shot list.

  2. 02

    The planted detail

    A pink prosthetic in sentence one, a pink suit in sentence two. Any transformation reads better when the thing it grows out of is visible beforehand, and it costs six words.

  3. 03

    The colour pair

    Candy pink against concrete grey. Naming what the colour is set against is worth more than naming the colour twice, and it transfers to every palette instruction you will ever write.

  4. 04

    The lollipop

    One small, specific, slightly wrong detail in the close-up. It is the only thing in the prompt that is not doing structural work, and it is why the close-up reads as a character rather than as a helmet.

Three ways to break it

  • Asking for a cut

    There is no cut syntax. "Cut to" and "next shot" are read as words, and the model does its best with them — which is usually a whip pan or a dissolve. Change coverage by saying what the camera does or what the frame now contains.

  • Describing the face in detail

    Wan 3.0 has no seed-independent identity mechanism in text mode, so a face specified in adjectives will still drift between the street beat and the close-up. If the same face across shot sizes matters, that is a reference image and a different mode.

  • Putting a spoken line in quotation marks

    If you add dialogue here — and the close-up is the obvious place — the words have to go inside braces. Quotation marks after a speech verb get the line narrated by an unseen voice instead of performed by the character.

What it does

What the schoolgirl to mech suit template does

The interesting problem in this prompt is coverage, not story. Three shot sizes in fifteen seconds, in a model that has no way to express a cut.

Every field-protocol video model — the ones that publish a shot list schema, a timecode grammar, angle-bracket asset tags — makes coverage explicit. Wan 3.0 does none of that. It reads one plain string, and anything that looks like a shot marker is content to be rendered rather than structure to be parsed. So the question is what plain English reliably produces a change in shot size, and the answer turns out to be surprisingly narrow: either you say what the camera does, or you say what the frame now contains.

This prompt uses both. "The camera pushes to an extreme close-up of her face behind the glass" is the first kind — an instruction to the camera, with the destination size named. "Finally the frame opens out into a wide concrete plaza where a green-skinned armoured brute twice her height is coming at her" is the second — the frame is not told to move, it is told what is in it, and a body twice her height with room to charge cannot be in a close-up. Both work. The second is more robust, because it constrains the composition by content rather than by a term the model has to interpret.

What does not work is asking for a cut. "Cut to", "next shot", "then we see" are all read as prose. The model does something with them, and what it does is usually a whip pan or a dissolve, because those are the transitions that exist inside continuous footage. If you want the feel of a cut, the honest options are to write a strong enough content change that the model has to jump, or to generate two clips and cut them yourself.

The ending is written, and on a fifteen-second request that is not optional. "The blast throws the brute back through the paving in a wall of dust, and the shot holds on her standing alone in the rubble" assigns the last beat and names the frame it stops on. Leave that off and the last third of a Wan 3.0 clip goes to whatever the model thinks a fight resolves into, which is reliably a slow pull back. Notice also where the beam comes from: the core in her chest, not a hand. A weapon needs an origin or the model puts the light wherever the pose happens to allow, and an emitter named once in the prompt is an emitter that stays in the same place for the rest of the shot.

The second thing worth studying is the sound layer, which is written as a mirror of the picture layer. Street noise, then servos, then rumble — three sounds in the same order as the three beats, with the same connective. On a model that generates audio and picture in one pass this is not decoration; it is a second, independent statement of the running order, phrased in a different modality. When the two agree, timing gets noticeably tighter. It is the cheapest reliability trick available on long Wan 3.0 requests and it is in almost nobody's prompt.

The colour instruction deserves a note of its own, because it is the most transferable clause here. "Saturated candy pink against concrete grey" names both sides. A prompt that asks only for pink tends to return a frame that is pink everywhere — the light, the walls, the grade — at which point nothing is pink, because pink is a relationship. Naming the ground against which a colour sits is how you get a colour that reads. This works for every palette instruction and it is worth making a habit of.

Finally, the honest caveat that applies to every template in this group: the clip was published without a prompt, and this text was written by reading the footage. It will get you a clip of this kind. If you need this exact character to survive a shot-size change, text mode is the wrong tool regardless of how the prompt is written — that is what reference mode and a character image are for, and there are three examples of it elsewhere in this library.

Everything here runs in text to video.

6 questions

Schoolgirl to mech suit — common questions

  • 01

    How do I change shot size mid-clip without a cut?

    Say what the camera does — "the camera pushes to an extreme close-up" — or say what the frame now contains. A body twice the height of your character with room to charge cannot be in a close-up, so naming it forces the wide.

  • 02

    Does Wan 3.0 understand "cut to"?

    Only as words. There is no cut syntax; the model reads the phrase and usually gives you a whip pan or a dissolve. Real cuts come from generating two clips and editing them.

  • 03

    Will the face stay the same across the three shots?

    Approximately, not reliably. Text mode has no identity mechanism. If a specific face has to survive a shot-size change, use reference mode with a character image.

  • 04

    Where should a beam or a muzzle flash come from?

    Somewhere you name. "From the core in her chest" fixes the emitter for the whole shot; leaving it unsaid lets the model attach the light to whichever limb the pose makes convenient, and it will move.

  • 05

    Why write the sound in the same order as the picture?

    Because Wan 3.0 generates both in one pass. A running order stated twice, once visually and once audibly, holds together better than one stated once.

  • 06

    Is this a real Wan 3.0 example?

    The clip is a render published elsewhere with no prompt attached. The prompt above was reconstructed from the footage, which is what the card says and what the video-to-prompt tool on this site does for your own clips.

Copy it, change one layer, run it.

Every character is on this page. 15s at 720P costs 240 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated