P—04 / THE MAN AT THE CLOCK TOWER

Wan 3.0 image to video prompt — a story hook without a word of dialogue

The prompt calls him "he" and never says another thing about him. The frame had already introduced them.

A tourist lowers his camera, looks up at a clock tower, and sees himself standing on it. The whole thing runs five seconds, has no dialogue, and never describes its own subject — it opens on the pronoun and trusts the reference frame for everything else. It is also the only prompt in the library that asks the camera to alternate between two points of view, which is more than five seconds can comfortably do.

Mode
Image to Video
Model
Prime
Duration
5s
Ratio
16:9
Resolution
720P
Audio
Square ambience, generated in the same pass
Credits
120

Alibaba’s published example for this model — not generated on this site. Source: wavespeed.ai

Output reference

A man with a camera stands in a busy European square looking toward a distant clock towerVideo
The clip

JPG · PNG · WebP · HEIC · 30 MB

374 / 20,000

The free clip is the same Wan 3.0 with its sound, capped at 480P · 3s, and it queues on shared capacity. A plan raises the ceiling to 1080P and thirty seconds — the free tier is a resolution, not a different model.

1 FREE CLIP · WAN 3.0 · 480P

01 — Inside

One prompt, one workflow

The clip beside the prompt that produced it, printed whole. Copy it into the console above, change what you need, generate.

Reference output

A man with a camera stands in a busy European square looking toward a distant clock towerVideo
The clip
The promptthe clipWan 3.0 Prime · Image to video
374 chars

The double

Three actions, one reveal, and a camera instruction that asks for two viewpoints in five seconds.

He lowers the camera, looks at the busy square around him, then raises it again to compare the scene. He slowly turns toward a clock tower visible in the distance, suddenly noticing someone standing there who appears to be himself. The camera alternates between his viewpoint and a slow push toward his stunned face. Urban suspense, subtle sci-fi hook, clear story movement.

6 layers

What each layer is doing

Wan 3.0 reads one plain string, but it reads it along seams. Take a layer rather than the whole prompt — the aesthetic and sound lines transfer to almost any subject, and they are the two most people write worst.

  1. Entity

    He … someone standing there who appears to be himself

    One pronoun and one doppelganger defined only by its relation to him. The frame supplies the face, and the second figure inherits it — which is cheaper and more reliable than describing the same person twice.

  2. Scene

    the busy square around him … a clock tower visible in the distance

    Both already in the frame. They are named to be pointed at, not to be built, which is why neither gets an adjective.

  3. Motion

    lowers the camera … raises it again … slowly turns toward a clock tower … suddenly noticing

    Ordered with "then", which is the only sequencing grammar Wan 3.0 has. Note the speed contrast: three slow actions and one sudden one, and the sudden one is the beat.

  4. Aesthetic control

    The camera alternates between his viewpoint and a slow push toward his stunned face.

    The most ambitious line here and the weakest. Alternating viewpoints means at least two cuts, and five seconds gives each of them under two seconds to establish.

  5. Stylization

    Urban suspense, subtle sci-fi hook, clear story movement

    "Subtle" is doing real work — it keeps the double a person on a tower rather than a glowing effect. "Clear story movement" is closer to a wish than an instruction.

  6. Sound

    Not written — left to the model.

    Unwritten. A square is safe to leave to the model, but the beat here is a realisation, and a sudden drop in the ambience is how that reads in every film it has ever been done in.

4 levers

Make it yours

What is safe to change. Most libraries publish only this list, which is why so many copied prompts come back worse than the original.

  1. 01

    The pronoun

    Opening on "he" is the point. The frame introduced him; the prompt starts mid-scene and describes only what he does. Try it on your own first frame — delete every word about who the subject is and see how much shorter the prompt gets.

  2. 02

    The speed contrast

    Lowers, looks, raises, slowly turns — then "suddenly noticing". Four calm verbs make one sudden one land. This is the cheapest way to give a five-second clip a beat, and it costs one adverb.

  3. 03

    The relational double

    "Someone standing there who appears to be himself" gets a second instance of a face without describing it. Any time you need two of something that must match, define the second by its relation to the first rather than repeating the description.

  4. 04

    The viewpoint alternation

    This is the line to cut if the run comes back muddled. Pick one: either we are behind his eyes or we are pushing into his face. At five seconds, choosing is better than alternating, and the push into the face is the shot that has the reaction in it.

Three ways to break it

  • Keeping both viewpoints and shortening further

    Two viewpoints need two establishing moments. Under five seconds there is not enough time for the second one to read, so the model resolves it by dropping one — and it does not always drop the one you would have.

  • Describing the double

    The moment you give the figure on the tower its own description it becomes a different person who happens to look similar. "Appears to be himself" is what makes it the same face; a wardrobe list is what breaks it.

  • Making the sci-fi loud

    "Subtle sci-fi hook" is the constraint holding the shot in the real world. Remove "subtle", or add a glow, a shimmer or a portal, and a suspense beat turns into a visual effect — which is a much easier thing to generate and a much less interesting thing to watch.

What it does

What the the man at the clock tower template does

This prompt does something the other five do not: it starts with a pronoun. Not "a young man with a camera" — just "he". That is only possible because a first frame is attached, and it is the clearest demonstration in the library of what image-to-video actually changes about how you write.

The saving is bigger than it looks. A text-to-video version of this shot would need a paragraph before anything could happen: age, build, clothing, the camera he is holding, the square, the time of day, the light. All of it exists in the reference frame already, in more detail and with no risk of disagreeing with itself. Deleting it does not just shorten the prompt — it removes every place where the words and the picture could conflict, and conflict is where image-to-video runs go soft.

The second technique is the double. The prompt needs two instances of one face, which is normally the hardest thing to ask a generative model for, and it gets there by never describing the second one: "someone standing there who appears to be himself". The figure is defined by its relationship to a face the model can already see. There is nothing for it to average, because there is only one description in play. Contrast this with what happens if you write out the double's appearance — now there are two descriptions of what is meant to be one person, and the model treats them as two people who happen to be similar.

The pacing is worth stealing wholesale. Four verbs go by at a calm speed — lowers, looks, raises, slowly turns — and then one adverb changes register: "suddenly noticing". A five-second clip has room for exactly one beat, and this is how you spend it. The calm verbs are not filler; they are what makes the sudden one legible as sudden. Remove them and the realisation has nothing to be a change from.

Then there is the line this template exists to argue with. "The camera alternates between his viewpoint and a slow push toward his stunned face" is asking for two distinct coverage patterns in five seconds. Each needs a moment to establish before it means anything, and there are not two such moments available. The honest advice is to choose, and to choose the push — the point of view shot shows a tower, the push shows a reaction, and the reaction is the story. Keep the alternation only when you raise the duration.

This ran on Prime, and it is worth saying again what that did and did not buy. Prime is the high-speed model. Alibaba publishes the same output specification for both tiers — the same resolutions, the same two-to-thirty-second range, thirty frames per second — and prices Prime at about half again per second. It returned faster. It did not return a better picture, and nothing about this prompt needed it.

What is missing is the sound, and here it is more of a loss than on the other templates. The beat is a realisation, and the standard film grammar for a realisation is that the world goes quiet. Wan 3.0 generates audio in the same pass as the picture, so a sentence like "square ambience that drops away as he turns, no music" would have scheduled the picture beat as well as the sound one. Leaving it unwritten means the model picked, and a busy square that stays busy is the version that reads as nothing happening.

Everything here runs in image to video.

6 questions

The man at the clock tower — common questions

  • 01

    Can a prompt really start with "he"?

    In image-to-video, yes. The reference frame has already introduced the subject, so the text can open mid-scene and describe only what he does.

  • 02

    How do I get two of the same person in one shot?

    Define the second by its relation to the first — "someone who appears to be himself" — rather than describing it again. Two descriptions of one person produce two people.

  • 03

    Why does the viewpoint alternation not fully land?

    Two coverage patterns need two establishing moments and five seconds only has one. Pick the shot with the reaction in it, or raise the duration.

  • 04

    What does "subtle" do in the style line?

    It keeps the science fiction inside the real world. Without it the double tends to arrive with a glow or a shimmer, which is easier to generate and much less unsettling.

  • 05

    Is Prime worth it for a shot like this?

    Only if you want it sooner. Prime is the high-speed tier with an identical published output specification, at roughly half again the cost per second.

  • 06

    How would I extend this past five seconds?

    Add beats and name their order — then, after that, finally — and say how it ends. A longer Wan 3.0 clip with no written ending invents one.

Copy it, change one layer, run it.

Every character is on this page. 5s at 720P costs 120 credits.

Written and maintained by the wan-3.run editorial teamPublished Last updated