P—04 / THE MAN AT THE CLOCK TOWER

Wan 3.0 image to video prompt — a story hook without a word of dialogue

The prompt calls him "he" and never says another thing about him. The frame had already introduced them.

A tourist lowers his camera, looks up at a clock tower, and sees himself standing on it. The whole thing runs five seconds, has no dialogue, and never describes its own subject — it opens on the pronoun and trusts the reference frame for everything else. It is also the only prompt in the library that asks the camera to alternate between two points of view, which is more than five seconds can comfortably do.

模式
Image to Video
模型
Prime
时长
5s
比例
16:9
分辨率
720P
音频
Square ambience, generated in the same pass
积分
120

阿里巴巴为这个模型公布的示例 —— 不是在本站生成的。 来源:wavespeed.ai

Output reference

A man with a camera stands in a busy European square looking toward a distant clock tower视频
这条片子

JPG · PNG · BMP · WebP · ≤20 MB · 240–8000px · ratio ≤8:1

374 / 20,000

免费那条就是同一个 Wan 3.0,声音也有,上限是 480P · 3s,走的是共用队列。方案能把上限抬高到 1080P 和三十秒 —— 免费档是一个分辨率,不是另一个模型。

1 条免费 · WAN 3.0 · 480P

01 —— 内部

一条提示词,一套流程

片子摆在做出它的那条提示词旁边,整条印出来。把它复制进上面的控制台,改你需要改的,然后生成。

参考产出

A man with a camera stands in a busy European square looking toward a distant clock tower视频
片子
这条提示词the clipWan 3.0 Prime · Image to video
374 chars

The double

Three actions, one reveal, and a camera instruction that asks for two viewpoints in five seconds.

He lowers the camera, looks at the busy square around him, then raises it again to compare the scene. He slowly turns toward a clock tower visible in the distance, suddenly noticing someone standing there who appears to be himself. The camera alternates between his viewpoint and a slow push toward his stunned face. Urban suspense, subtle sci-fi hook, clear story movement.

6 层

每一层在做什么

Wan 3.0 读的是一整条白话字符串,但它是顺着接缝读的。拿走一层,而不是整条提示词 —— 美学和声音那两行几乎能换到任何主体上,而它们正是大多数人写得最差的两行。

  1. Entity

    He … someone standing there who appears to be himself

    One pronoun and one doppelganger defined only by its relation to him. The frame supplies the face, and the second figure inherits it — which is cheaper and more reliable than describing the same person twice.

  2. Scene

    the busy square around him … a clock tower visible in the distance

    Both already in the frame. They are named to be pointed at, not to be built, which is why neither gets an adjective.

  3. Motion

    lowers the camera … raises it again … slowly turns toward a clock tower … suddenly noticing

    Ordered with "then", which is the only sequencing grammar Wan 3.0 has. Note the speed contrast: three slow actions and one sudden one, and the sudden one is the beat.

  4. Aesthetic control

    The camera alternates between his viewpoint and a slow push toward his stunned face.

    The most ambitious line here and the weakest. Alternating viewpoints means at least two cuts, and five seconds gives each of them under two seconds to establish.

  5. Stylization

    Urban suspense, subtle sci-fi hook, clear story movement

    "Subtle" is doing real work — it keeps the double a person on a tower rather than a glowing effect. "Clear story movement" is closer to a wish than an instruction.

  6. Sound

    没写 —— 留给模型。

    Unwritten. A square is safe to leave to the model, but the beat here is a realisation, and a sudden drop in the ambience is how that reads in every film it has ever been done in.

4 个可调项

把它变成你的

哪些改动是安全的。多数库只发布这一份清单,这也是为什么那么多复制来的提示词跑出来比原作还差。

  1. 01

    The pronoun

    Opening on "he" is the point. The frame introduced him; the prompt starts mid-scene and describes only what he does. Try it on your own first frame — delete every word about who the subject is and see how much shorter the prompt gets.

  2. 02

    The speed contrast

    Lowers, looks, raises, slowly turns — then "suddenly noticing". Four calm verbs make one sudden one land. This is the cheapest way to give a five-second clip a beat, and it costs one adverb.

  3. 03

    The relational double

    "Someone standing there who appears to be himself" gets a second instance of a face without describing it. Any time you need two of something that must match, define the second by its relation to the first rather than repeating the description.

  4. 04

    The viewpoint alternation

    This is the line to cut if the run comes back muddled. Pick one: either we are behind his eyes or we are pushing into his face. At five seconds, choosing is better than alternating, and the push into the face is the shot that has the reaction in it.

三种把它弄坏的方式

  • Keeping both viewpoints and shortening further

    Two viewpoints need two establishing moments. Under five seconds there is not enough time for the second one to read, so the model resolves it by dropping one — and it does not always drop the one you would have.

  • Describing the double

    The moment you give the figure on the tower its own description it becomes a different person who happens to look similar. "Appears to be himself" is what makes it the same face; a wardrobe list is what breaks it.

  • Making the sci-fi loud

    "Subtle sci-fi hook" is the constraint holding the shot in the real world. Remove "subtle", or add a glow, a shimmer or a portal, and a suspense beat turns into a visual effect — which is a much easier thing to generate and a much less interesting thing to watch.

它做什么

the man at the clock tower 这个模板做什么

This prompt does something the other five do not: it starts with a pronoun. Not "a young man with a camera" — just "he". That is only possible because a first frame is attached, and it is the clearest demonstration in the library of what image-to-video actually changes about how you write.

The saving is bigger than it looks. A text-to-video version of this shot would need a paragraph before anything could happen: age, build, clothing, the camera he is holding, the square, the time of day, the light. All of it exists in the reference frame already, in more detail and with no risk of disagreeing with itself. Deleting it does not just shorten the prompt — it removes every place where the words and the picture could conflict, and conflict is where image-to-video runs go soft.

The second technique is the double. The prompt needs two instances of one face, which is normally the hardest thing to ask a generative model for, and it gets there by never describing the second one: "someone standing there who appears to be himself". The figure is defined by its relationship to a face the model can already see. There is nothing for it to average, because there is only one description in play. Contrast this with what happens if you write out the double's appearance — now there are two descriptions of what is meant to be one person, and the model treats them as two people who happen to be similar.

The pacing is worth stealing wholesale. Four verbs go by at a calm speed — lowers, looks, raises, slowly turns — and then one adverb changes register: "suddenly noticing". A five-second clip has room for exactly one beat, and this is how you spend it. The calm verbs are not filler; they are what makes the sudden one legible as sudden. Remove them and the realisation has nothing to be a change from.

Then there is the line this template exists to argue with. "The camera alternates between his viewpoint and a slow push toward his stunned face" is asking for two distinct coverage patterns in five seconds. Each needs a moment to establish before it means anything, and there are not two such moments available. The honest advice is to choose, and to choose the push — the point of view shot shows a tower, the push shows a reaction, and the reaction is the story. Keep the alternation only when you raise the duration.

This ran on Prime, and it is worth saying again what that did and did not buy. Prime is the high-speed model. Alibaba publishes the same output specification for both tiers — the same resolutions, the same two-to-thirty-second range, thirty frames per second — and prices Prime at about half again per second. It returned faster. It did not return a better picture, and nothing about this prompt needed it.

What is missing is the sound, and here it is more of a loss than on the other templates. The beat is a realisation, and the standard film grammar for a realisation is that the world goes quiet. Wan 3.0 generates audio in the same pass as the picture, so a sentence like "square ambience that drops away as he turns, no music" would have scheduled the picture beat as well as the sound one. Leaving it unwritten means the model picked, and a busy square that stays busy is the version that reads as nothing happening.

这里的一切都在image to video里跑。

6 个问题

The man at the clock tower —— 常见问题

  • 01

    Can a prompt really start with "he"?

    In image-to-video, yes. The reference frame has already introduced the subject, so the text can open mid-scene and describe only what he does.

  • 02

    How do I get two of the same person in one shot?

    Define the second by its relation to the first — "someone who appears to be himself" — rather than describing it again. Two descriptions of one person produce two people.

  • 03

    Why does the viewpoint alternation not fully land?

    Two coverage patterns need two establishing moments and five seconds only has one. Pick the shot with the reaction in it, or raise the duration.

  • 04

    What does "subtle" do in the style line?

    It keeps the science fiction inside the real world. Without it the double tends to arrive with a glow or a shimmer, which is easier to generate and much less unsettling.

  • 05

    Is Prime worth it for a shot like this?

    Only if you want it sooner. Prime is the high-speed tier with an identical published output specification, at roughly half again the cost per second.

  • 06

    How would I extend this past five seconds?

    Add beats and name their order — then, after that, finally — and say how it ends. A longer Wan 3.0 clip with no written ending invents one.

复制它,改一层,跑起来。

每一个字符都在这一页上。720P 的 5s 要 120 积分。

由 wan-3.run 编辑团队撰写与维护发布于 最后更新