Black-and-White Depth Video Driven · A Dress Showcase Workflow With Your Own Model
The prompt itself is thin — just one sentence — but the real cleverness is in splitting the workflow: pose/motion data and final-frame identity are fed to the model in two separate steps. The depth video first extracts a pure motion skeleton, stripping out the original footage's person and scene, so the third-step prompt only has to do one narrowing pass of 'positive inclusion plus negative exclusion.' Packing a whitelist and a blacklist into a single sentence — 'only reference A, not B or C' — is the cheapest way to keep an image-to-video model from drifting, and it works better than piling on shot details, because the motion is already locked in by the depth video; the prompt's only job is to name what shouldn't be learned.
This block is the three-step workflow itself, and the real division of labor happens in the first two steps: step one converts a reference video into a black-and-white depth map, keeping only the motion skeleton while depth information erases the original person and scene entirely — this step can even be handed off to codex to script and automate, no manual work needed. Step two separately generates your own model (a character card is preferred, though a single image works), completely decoupling 'who performs' from 'how they move.' Step three is the only part that's actually fed to the video model as a prompt, and it's just one sentence, because the first two steps have already locked down everything that needs locking. This layered approach — do the data prep inside the workflow, let the prompt only handle the finishing touch — is more stable than trying to cram every constraint into a single sentence.



