
A video prompt needs to explain change. A beautiful location description tells the model where the scene happens, but not what viewers should watch. For Wan 3.0 on VisionDir, make the central action and the final state clear before adding decorative detail.
Open Wan 3.0 text-to-video, or use image-to-video when a source image should guide the opening composition. Available inputs and durations are shown in the current form.
Give the shot one visible job
Choose a single action that can be understood without an explanation: a curtain moves to reveal a window, a cyclist slows beside a cafe, or a hand places a cup on a table. Describe the beginning and end of that action. Avoid asking the same short shot to establish a city, introduce several people and complete a complicated chase.
Medium shot of a small ceramic cup on a wooden cafe table. A hand enters from the right, turns the handle toward the camera and leaves the frame. The cup stays upright and in the same place. Soft window light, a quiet background and a steady camera. End with a clean unobstructed view of the cup.
This is a proposed starting prompt, not a report of a tested output. Use the actual generated result to decide which instruction needs revision.
Keep camera movement separate from subject movement
A moving subject can be filmed with a fixed camera. A still object can be revealed with a slow camera move. Decide which motion carries the story; asking for both at maximum intensity makes the result harder to assess.
For the cafe example, a steady camera helps you judge the hand and cup interaction. For a landscape reveal, a slow forward movement might be the main event. Do not combine “locked camera” and “rapid orbit” in the same continuous shot.
Use references with a clear role
For image-to-video, choose a source image whose framing already supports the intended action. Tell the prompt what should remain stable, such as the subject's clothing, product shape or overall composition. Describe the movement you want instead of restating every visible detail.
Where the tool exposes additional reference inputs, assign each one a purpose. Do not assume that an input type described on another platform is available in every VisionDir mode.
Diagnose a weak result
| Observation | Revision to try |
|---|---|
| The action is unclear | Describe its starting and ending positions |
| The subject changes identity | Reduce scene changes and reinforce a few key traits |
| The camera is distracting | Use one restrained camera instruction |
| The ending is unusable | Ask for a stable final composition with no new action |
Start with a supported short duration that is long enough for the action. Review the current credit estimate before generating. Reuse a prompt only when its action makes sense for the new subject.
Common questions
Will timestamps guarantee exact timing?
No. They communicate intent, but the result still needs review. Match your described beats to the duration selected in the tool and leave enough time for the ending.
What should I do about audio?
Use the controls actually available in the selected workflow. If audio is supported, describe speech, ambience and effects separately; otherwise plan sound in your editor.
Continue with the Seedance 2.5 storyboard guide and MiniMax H3 ad planning. Model reference: Wan 3.0.
