Dusk Street Dialogue in Japanese · Three Lines, One Continuous Take

Wan 3.0AnimationLip SyncUrbanVerticalImage to VideoJapanese StyleStructured Prompt10s
A Japanese-language prompt that locks the characters with a first-frame reference image for a ten-second, one-continuous-take shot: two people walk side by side down a dusk street, she asks "have you been to that cafe up ahead," he answers "not yet, actually I've always wanted to," and she finishes with "let's go in, then." Each line is assigned to a speaker, and even the listener is required to show eye contact and a nod in response.
PROMPT · Video Prompt
Create with AI
Use the reference image directly as the first frame. 10 seconds, vertical 9:16, with audio, one continuous shot. [VISUALS & SCENE] Polished Pixar-style 3D animation. Keep the facial features, hairstyle, build, and outfit of the young man and woman from the reference image. The woman is on screen left, the man on screen right. The two walk side by side down a dusk-lit street sidewalk, chatting naturally in Japanese like close friends. Keep the bluish dusk ambient light and the warm light spilling from the cafe on the right side of frame. A soft warm rim light on hair and shoulders. Natural shading on the faces while keeping both people's eyes and mouths clearly visible. The streetlights in the background stay softly blurred. A tone and texture like a calm drama scene. [CAMERA] Start from the reference image's composition. The camera smoothly tracks backward at the pace of the two walking forward, keeping the characters roughly the same size. Over the first 7 seconds, the camera moves in a small arc from its current diagonal-front position toward a frontal one. The cafe and streetlights in the background drift slowly, creating a natural sense of depth. In the last 3 seconds, it keeps retreating while very slightly closing the distance, moving in toward the two people's smiles. Always keep both people in the same frame. Maintain their left-right positions and keep faces, mouths, and both hands in frame. Horizontal, stable, understated camera work. [DIALOGUE & PERFORMANCE] 0-4s: The woman turns her face slightly toward the man beside her and asks, in a casual, bright voice: 「この先のカフェ、もう行った?」("Have you been to that cafe up ahead?") The man listens with his mouth closed, naturally shifting his gaze to her. Both keep walking. 4-7s: The man turns his face slightly toward the woman and answers, in a calm, friendly voice: 「まだ。実は気になってたんだよ!」("Not yet - actually, I've been curious about it!") The woman listens with her mouth closed, smiling softly. 7-10s: The woman says, a little happily: 「じゃあ、寄ってこ!」("Then let's stop in!") The man smiles and nods once without speaking. Both turn their gaze back forward, ending while still walking. [MOVEMENT] Natural stride length, foot contact, and weight transfer. Understated arm swing and shoulder movement matched to the walking pace. Hair, clothes, and the woman's bag sway slightly with each step. When they look at each other, only a slight turn of the face, keeping an angle where the camera can still see both people's mouths. Not just the speaker - the listener should also show natural reactions through gaze, eyebrows, and small nods. Performance is calm and everyday. [AUDIO] Two clearly distinct, natural Japanese voices - a young adult woman and adult man. Only the assigned person speaks the assigned line. No overlapping speech; the conversation's pauses connect naturally. Sync each speaker's lip and jaw movement precisely to Japanese pronunciation. The listener does not speak. Quiet footsteps and subdued street ambience. Dialogue volume stays clearly audible. [CONSTRAINTS] No subtitles, no on-screen text, no BGM, no narration. No cuts, no sudden zooms, no large orbiting moves, no camera shake. No looking at the camera, no exaggerated gestures, no stopping mid-walk. No changes to the characters' faces or outfits, no foot slipping, no hand deformation. No pedestrians blocking the shot, no sudden shifts in brightness or color temperature.
✍️ Editor’s Notes

This prompt divides the work into six bracketed sections: the visuals/scene section locks in the characters and ambient light, the camera section spells out two distinct arcs of movement for the first 7 seconds and the last 3, the dialogue section assigns each of the three lines to a speaker across 0-4/4-7/7-10 seconds with its own tone, and the movement section separately insists the listener also shows eye contact and a nod - the detail most dialogue prompts skip, and the one that most often makes a scene look wooden. The audio section further demands precise lip-sync to Japanese pronunciation and no overlapping speech, heading off the usual lip-sync failure mode in advance. The closing constraints section locks in the "one continuous take" realism with a list of bans - no cuts, no looking at camera, no stopping mid-walk. The whole structure can be dropped directly onto any two-person dialogue clip.

Create with AI
Reference images
Dusk Street Dialogue in Japanese · Three Lines, One Continuous Take · 参考图
Related Works
×