The prompt is just three sentences: use the two characters from Kling MCP Elements to make a one-minute Japanese music video, generate the song in one pass with no splitting into segments, and the agent is free to search online for a better approach. Everything else — shot breakdown, scenes, editing rhythm — is left entirely to the tool-equipped agent to decide. It's a sample of handing creation over to the workflow instead of the prompt.
帮我用 Kling MCP Elements 里的 Girl A 和 Girl B 制作一支一分钟的日语音乐视频。完整生成一首一分钟的日语歌曲,不要拆分成多段。你也可以上网搜索更适合制作音乐视频的技巧或方法。使用 Kling MCP。
請幫我用 Kling MCP Elements 裡的 Girl A 和 Girl B,製作一支一分鐘的日語MV。完整生成一首一分鐘的日語歌曲,不要拆分成多段生成。你也可以上網搜尋更適合用來製作MV的技巧或方法。使用 Kling MCP。
Help me create a 1-minute Japanese music video using Girl A and Girl B from the Kling MCP Elements. Generate a complete 1-minute Japanese song without splitting it into segments. You may also search online for any skills or approaches better suited to making music videos. Use the Kling MCP.
This isn't a prompt written for a video model — it's a work order handed to a tool-equipped agent. It specifies the two characters already stored in Kling MCP Elements, asks for a one-minute Japanese MV, hard-requires the whole song to be generated in a single pass with no splitting and stitching, and closes with 'you may also search online for a better approach yourself' — which effectively hands over every decision about shot breakdown, camera setup, and editing rhythm. What actually determines the final cut is the character assets pre-stored in Elements and the chain of models sitting behind the MCP; the prompt only supplies the goal and the constraints. Replicating this approach isn't about wording — it comes down to whether you have reusable character assets on hand and a workflow that can call the generation pipeline on its own.
Want tighter control? Editor-expanded reference version (not the original prompt):
Using the characters Girl A and Girl B from Kling MCP Elements, create a 60-second Japanese rock MV, with both characters' appearance, hairstyles, and school uniforms staying consistent throughout the video.
First generate the complete 60-second Japanese song in a single pass, not in segments stitched together afterward: mid-tempo rock, guitar and bass driven, with the chorus lifting, female Japanese vocals.
The visuals are split into four segments:
0:00–0:15 Rehearsal room, side lighting, medium shot, both girls standing with their instruments, slight handheld sway;
0:15–0:35 Stage with a black background, alternating between a full-frontal wide shot and push-in close-ups, spotlight on both of them, cuts follow the drumbeat;
0:35–0:50 Transition through a hallway and stairwell, tracking shot from behind, natural light, tempo slows down;
0:50–1:00 Rooftop, backlit by the sunset, the two stand side by side looking at the city, ending on a wide shot.
Shots are mainly medium and wide, each shot no shorter than 3 seconds; no subtitles, no dialogue, no logos or watermarks; the instrument-playing motions must be synced to the beat.