Start with this order: subject and setting; action from opening to final state; framing and one camera move; motivated light and visual finish; sound; then continuity constraints. The labels are a drafting aid, not commands your model needs. Remove the labels after every field describes one compatible result.
For example: a florist closes a street stall at dusk; she folds the awning, lifts one last bucket, and exits frame; a medium-wide camera tracks slowly left; warm shop light meets cool rain; canvas snaps, buckets scrape, and distant traffic remains low; preserve her red coat and the wet street layout throughout.
Turn the template into visible direction
Replace emotion labels with behavior the camera can record. Instead of “she feels nervous,” write that she checks the empty street twice, tightens her grip, and pauses before stepping into the rain. Replace “cinematic” with the framing, lens distance, light source, contrast, or movement that creates the intended look.
Adapt the video prompt template by model
Keep the scene facts stable, then change emphasis. Kling needs legible physical motion. Veo benefits from sourced, timed audio. Seedance needs shot-to-shot continuity. Sora benefits from clear spatial and world rules. Runway usually works best when one action and one camera move remain concise.
If you begin from an image, treat the still as the opening frame. Do not spend half the prompt repeating what is already visible. State which details must stay fixed, what begins to move, how the environment responds, and whether the camera follows or stays locked.
Revise once for contradictions and once for timing. A five-second shot cannot comfortably contain several locations, four camera moves, long dialogue, and a complete transformation. Split that concept into separate shots, or keep the one change that carries the idea. A useful prompt is specific because its decisions agree, not because it contains the most words.
Test one prompt decision at a time
Keep a small version history when you test the prompt. In the first generation, confirm the subject and main action. In the second, change only camera behavior. In the third, refine light, texture, or audio without rewriting the action that already works. Record which model, duration, aspect ratio, source image, and seed produced each result. A video prompt generator can organize the written direction, but side-by-side tests reveal how a particular model interprets it. Controlled variants also make failures useful: you can tell whether a weak result came from an unclear verb, conflicting motion, overloaded timing, or a style instruction that changed the scene more than expected. Keep the version that solves the main visual problem during the next generation test, even when a longer draft sounds more impressive on the page.