P—04 / THE MAN AT THE CLOCK TOWER

Wan 3.0 image to video prompt — a story hook without a word of dialogue

The prompt calls him "he" and never says another thing about him. The frame had already introduced them.

A tourist lowers his camera, looks up at a clock tower, and sees himself standing on it. The whole thing runs five seconds, has no dialogue, and never describes its own subject — it opens on the pronoun and trusts the reference frame for everything else. It is also the only prompt in the library that asks the camera to alternate between two points of view, which is more than five seconds can comfortably do.

Modo
Image to Video
Modelo
Prime
Duración
5s
Relación
16:9
Resolución
720P
Audio
Square ambience, generated in the same pass
Créditos
120

Ejemplo publicado por Alibaba para este modelo: no se generó en este sitio. Fuente: wavespeed.ai

Output reference

A man with a camera stands in a busy European square looking toward a distant clock towerVídeo
El clip

JPG · PNG · BMP · WebP · ≤20 MB · 240–8000px · ratio ≤8:1

374 / 20,000

El clip gratis es el mismo Wan 3.0 con su sonido, limitado a 480P · 3s, y se pone en cola en capacidad compartida. Un plan sube el techo a 1080P y treinta segundos: la capa gratuita es una resolución, no un modelo distinto.

1 CLIP GRATIS · WAN 3.0 · 480P

01 — Por dentro

Un prompt, un flujo de trabajo

El clip junto al prompt que lo produjo, impreso entero. Cópialo en la consola de arriba, cambia lo que necesites y genera.

Salida de referencia

A man with a camera stands in a busy European square looking toward a distant clock towerVídeo
El clip
El promptthe clipWan 3.0 Prime · Image to video
374 chars

The double

Three actions, one reveal, and a camera instruction that asks for two viewpoints in five seconds.

He lowers the camera, looks at the busy square around him, then raises it again to compare the scene. He slowly turns toward a clock tower visible in the distance, suddenly noticing someone standing there who appears to be himself. The camera alternates between his viewpoint and a slow push toward his stunned face. Urban suspense, subtle sci-fi hook, clear story movement.

6 capas

Qué hace cada capa

Wan 3.0 lee una sola cadena llana, pero la lee por costuras. Coge una capa en lugar del prompt entero: las líneas de estética y de sonido se transfieren a casi cualquier sujeto, y son las dos que casi todo el mundo escribe peor.

  1. Entity

    He … someone standing there who appears to be himself

    One pronoun and one doppelganger defined only by its relation to him. The frame supplies the face, and the second figure inherits it — which is cheaper and more reliable than describing the same person twice.

  2. Scene

    the busy square around him … a clock tower visible in the distance

    Both already in the frame. They are named to be pointed at, not to be built, which is why neither gets an adjective.

  3. Motion

    lowers the camera … raises it again … slowly turns toward a clock tower … suddenly noticing

    Ordered with "then", which is the only sequencing grammar Wan 3.0 has. Note the speed contrast: three slow actions and one sudden one, and the sudden one is the beat.

  4. Aesthetic control

    The camera alternates between his viewpoint and a slow push toward his stunned face.

    The most ambitious line here and the weakest. Alternating viewpoints means at least two cuts, and five seconds gives each of them under two seconds to establish.

  5. Stylization

    Urban suspense, subtle sci-fi hook, clear story movement

    "Subtle" is doing real work — it keeps the double a person on a tower rather than a glowing effect. "Clear story movement" is closer to a wish than an instruction.

  6. Sound

    Sin escribir: se deja al modelo.

    Unwritten. A square is safe to leave to the model, but the beat here is a realisation, and a sudden drop in the ambience is how that reads in every film it has ever been done in.

4 palancas

Hazlo tuyo

Lo que se puede cambiar sin riesgo. Casi todas las bibliotecas publican solo esta lista, y por eso tantos prompts copiados vuelven peor que el original.

  1. 01

    The pronoun

    Opening on "he" is the point. The frame introduced him; the prompt starts mid-scene and describes only what he does. Try it on your own first frame — delete every word about who the subject is and see how much shorter the prompt gets.

  2. 02

    The speed contrast

    Lowers, looks, raises, slowly turns — then "suddenly noticing". Four calm verbs make one sudden one land. This is the cheapest way to give a five-second clip a beat, and it costs one adverb.

  3. 03

    The relational double

    "Someone standing there who appears to be himself" gets a second instance of a face without describing it. Any time you need two of something that must match, define the second by its relation to the first rather than repeating the description.

  4. 04

    The viewpoint alternation

    This is the line to cut if the run comes back muddled. Pick one: either we are behind his eyes or we are pushing into his face. At five seconds, choosing is better than alternating, and the push into the face is the shot that has the reaction in it.

Tres formas de romperlo

  • Keeping both viewpoints and shortening further

    Two viewpoints need two establishing moments. Under five seconds there is not enough time for the second one to read, so the model resolves it by dropping one — and it does not always drop the one you would have.

  • Describing the double

    The moment you give the figure on the tower its own description it becomes a different person who happens to look similar. "Appears to be himself" is what makes it the same face; a wardrobe list is what breaks it.

  • Making the sci-fi loud

    "Subtle sci-fi hook" is the constraint holding the shot in the real world. Remove "subtle", or add a glow, a shimmer or a portal, and a suspense beat turns into a visual effect — which is a much easier thing to generate and a much less interesting thing to watch.

Qué hace

Qué hace la plantilla the man at the clock tower

This prompt does something the other five do not: it starts with a pronoun. Not "a young man with a camera" — just "he". That is only possible because a first frame is attached, and it is the clearest demonstration in the library of what image-to-video actually changes about how you write.

The saving is bigger than it looks. A text-to-video version of this shot would need a paragraph before anything could happen: age, build, clothing, the camera he is holding, the square, the time of day, the light. All of it exists in the reference frame already, in more detail and with no risk of disagreeing with itself. Deleting it does not just shorten the prompt — it removes every place where the words and the picture could conflict, and conflict is where image-to-video runs go soft.

The second technique is the double. The prompt needs two instances of one face, which is normally the hardest thing to ask a generative model for, and it gets there by never describing the second one: "someone standing there who appears to be himself". The figure is defined by its relationship to a face the model can already see. There is nothing for it to average, because there is only one description in play. Contrast this with what happens if you write out the double's appearance — now there are two descriptions of what is meant to be one person, and the model treats them as two people who happen to be similar.

The pacing is worth stealing wholesale. Four verbs go by at a calm speed — lowers, looks, raises, slowly turns — and then one adverb changes register: "suddenly noticing". A five-second clip has room for exactly one beat, and this is how you spend it. The calm verbs are not filler; they are what makes the sudden one legible as sudden. Remove them and the realisation has nothing to be a change from.

Then there is the line this template exists to argue with. "The camera alternates between his viewpoint and a slow push toward his stunned face" is asking for two distinct coverage patterns in five seconds. Each needs a moment to establish before it means anything, and there are not two such moments available. The honest advice is to choose, and to choose the push — the point of view shot shows a tower, the push shows a reaction, and the reaction is the story. Keep the alternation only when you raise the duration.

This ran on Prime, and it is worth saying again what that did and did not buy. Prime is the high-speed model. Alibaba publishes the same output specification for both tiers — the same resolutions, the same two-to-thirty-second range, thirty frames per second — and prices Prime at about half again per second. It returned faster. It did not return a better picture, and nothing about this prompt needed it.

What is missing is the sound, and here it is more of a loss than on the other templates. The beat is a realisation, and the standard film grammar for a realisation is that the world goes quiet. Wan 3.0 generates audio in the same pass as the picture, so a sentence like "square ambience that drops away as he turns, no music" would have scheduled the picture beat as well as the sound one. Leaving it unwritten means the model picked, and a busy square that stays busy is the version that reads as nothing happening.

Todo lo de aquí se ejecuta en image to video.

6 preguntas

The man at the clock tower — preguntas frecuentes

  • 01

    Can a prompt really start with "he"?

    In image-to-video, yes. The reference frame has already introduced the subject, so the text can open mid-scene and describe only what he does.

  • 02

    How do I get two of the same person in one shot?

    Define the second by its relation to the first — "someone who appears to be himself" — rather than describing it again. Two descriptions of one person produce two people.

  • 03

    Why does the viewpoint alternation not fully land?

    Two coverage patterns need two establishing moments and five seconds only has one. Pick the shot with the reaction in it, or raise the duration.

  • 04

    What does "subtle" do in the style line?

    It keeps the science fiction inside the real world. Without it the double tends to arrive with a glow or a shimmer, which is easier to generate and much less unsettling.

  • 05

    Is Prime worth it for a shot like this?

    Only if you want it sooner. Prime is the high-speed tier with an identical published output specification, at roughly half again the cost per second.

  • 06

    How would I extend this past five seconds?

    Add beats and name their order — then, after that, finally — and say how it ends. A longer Wan 3.0 clip with no written ending invents one.

Cópialo, cambia una capa, ejecútalo.

Todos los caracteres están en esta página. 5s en 720P cuesta 120 créditos.

Escrito y mantenido por el equipo editorial de wan-3.runPublicado el Actualizado el