Onsen Ryokan Vlog · Even the Pitch Accent Is Written Into the Prompt

Seedance 2.5RealisticVlogTravelFoodLip SyncVerticalJapanese StyleStructured Prompt7s
Two beats in seven seconds: a travel vlogger checks in at a ryokan in Atami, greeting the camera at the entrance, then cuts to dinner where she takes a bite of simmered kinmedai and says it's delicious. What makes this Japanese prompt unusual is that it marks the pronunciation and pitch accent of the dialogue syllable by syllable — down to how the fictional ryokan's name should be read and where to breathe.
PROMPT · Video Prompt
Create with AI
[Assets] @画像1 = Reference for the protagonist's face, hairstyle, build, casual clothes, and yukata @画像2 = Ryokan exterior, entrance, noren curtain, landscaping @画像3 = Guest room (table, window, furniture layout) @画像5 = Simmered kinmedai (splendid alfonsino), tableware, plating (no separate background — served on the table from @画像3) [In a Nutshell] 7 seconds | Vertical 9:16 | 24fps. A 34-year-old female travel influencer arrives at the hot-spring ryokan "Shionagi-an" in Atami and talks to the camera while filming a selfie, then time skips ahead to dinner in the guest room, where she takes a bite of kinmedai and savors it — a live-action travel vlog. [Overall Setup] Base environment and texture: September. The first half is soft natural afternoon light at the entrance; the second half is warm indoor lighting during dinner in the guest room. Push realistic physical texture to the maximum. Visual: Live-action, shallow depth of field, natural light transitioning naturally into indoor light. Camera language: Only one camera movement per shot. In the first half she holds the camera herself for a natural, slightly shaky selfie shot; in the second half, a locked-off camera centered on her face and shoulders. Character: Strictly lock the face, hairstyle, and build from @画像1 throughout. Casual clothes in the first half, yukata in the second half. Retain authentic, fine pore and skin texture. Core of the performance: A candid, easygoing travel influencer. Not raising her voice, speaking soft and natural Japanese. The distance of talking to a single friend. No ad-read delivery. Prohibited: No subtitles, no BGM, no filming equipment visible on screen, no added walking or door-opening actions, no posing while holding food with chopsticks. [Timestamps] 0-3s [Arrival & Greeting] Scene: In front of the ryokan entrance from @画像2. Action: The protagonist in casual clothes stops and takes a selfie with one hand. Framing from the chest up, face slightly above center, the entrance and noren curtain behind her. She speaks with a natural smile while looking at the lens. Line: 「きょうの おやどは、しおなぎあん。みて、いいかんじ」 (Have her speak it in hiragana as written, not in kanji.) Delivery: Say 「きょうの」 lightly, at a normal pace. Leave a small pause after 「おやどは、」 — take one breath here. 「しおなぎあん」 is a proper noun; say it smoothly in one breath without splitting it. Switch to a bright, bouncy tone for 「みて、いいかんじ」. On 「みて」, glance lightly toward the entrance to share the background. Pronunciation / Pitch: KYOU[H]->NO[L] / きょ↑う→の→ ("today" — a real word, head-high pattern) O[L]->YA[H]->DO[H] / お→や↑ど→ ("inn/lodging" — a real word, flat pattern; the particle 「は」 also stays high and continues) SHI[H]->O[L]->NA[L]->GI[L]->AN[L] / し↑お↓な→ぎ→あん→ ("Shionagi-an" — a coined name with no standard accent. Assigned a head-high pattern for directorial rhythm: rise on 「しお」, then keep 「なぎあん」 low and flat without creating a break) MI[H]->TE[L] / み↑て→ ("look" — a real word, head-high pattern) I[L]->I[H]->KAN[H]->JI[H] / い→い↑かん→じ→ ("feels nice" — a real word, flat pattern) Intent: An expression of arrival-joy that comes through naturally. The cheeks and eye area move softly in sync with the mouth. Keep breathing, blinking, and small eye movements. Camera: Fixed selfie framing, natural perspective. Don't let a wide angle distort the face or arms. [Between Cuts] Skip the several hours between the guest room, the hot spring, freshening up after the bath, and dinner being served. No hard cuts — bridge with a natural cut transition. Time moves from midday to dusk and never reverses. 3-7s [One Bite & Reaction] Scene: The table in the guest room from @画像3. The protagonist, now in a yukata, is seated. The simmered kinmedai from @画像5 has already been served. Action: Start with a small, separated piece of kinmedai being brought to her mouth with chopsticks in her right hand. It goes into her mouth; she lowers the chopsticks. A gesture of eating one bite of fish so tender it melts. She doesn't bite down hard — she lets it naturally fall apart in her mouth. Food texture: Simmered white fish. Flesh cooked so delicately it flakes apart at the touch of chopsticks. Full of moisture, melt-in-the-mouth tender. Sound: The flesh softly comes apart in her mouth — a small, muffled chewing sound carrying moist fat and cooking broth. Only a faint, moist 「もぐ、もぐ」 sound is audible. Intent: Chew slowly and naturally, swallow without rushing. Once she's swallowed, her gaze returns from the food to the lens, and she speaks as her mouth naturally relaxes into a smile. Line: 「うん、おいしい」 Intent: A relaxed smile tinged with satisfaction. Camera: Locked, no movement. [Closing] Throughout: Keep voice quality, speaking distance, and volume consistent. Deliver the lines exactly as specified, synced to the mouth movement. Ambient sound: Only outdoor wind in the first half. Only quiet indoor sound, small tableware clinks, and soft chewing in the second half. No BGM. Pure footage. No subtitles. Retain authentic, fine pore and skin texture. Don't change the character's face or clothing.
✍️ Editor’s Notes

The prompt is organized in five blocks: asset list → one-line brief → overall setup → timestamped shots → closing notes, each with a clear job. The real trick is how it handles dialogue: the lines are written in hiragana instead of kanji so the model doesn't default to kanji readings, and a separate 'Pronunciation/Pitch' block breaks each word into syllables with romanization, H/L pitch marks, and native Japanese accent notation (↑↓→). Even the invented ryokan name '汐凪庵,' which has no standard accent, is assigned a director's-choice head-high pattern, with exact instructions on where to breathe and where to run the words together in one breath. This 'write the pronunciation script before the line is spoken' approach turns dialogue from plain text into a precisely controllable sound event — useful whenever you need a character to speak a foreign language and don't want the model guessing at pitch or pronunciation.

Create with AI
Related Works
×