How to Prepare Reference Images for Image-to-Video
Learn how to choose and prepare source images before using image-to-video generation.
Image-to-video generation works best when the source image already contains the composition you want to preserve. The model can add motion, but it cannot reliably fix every problem in the starting image.
Choose an image with one clear subject, readable edges, and enough background for motion. Cropped faces, cut-off products, tiny objects, and cluttered collages often lead to drifting identity or unstable movement.
Match the image shape to the target output. If you plan to publish a 9:16 short, start with a vertical reference when possible. If you need a 16:9 product demo, use a wide frame that leaves room for camera movement.
Avoid images with text-heavy layouts unless the text is not important. AI video models can distort small lettering while animating the frame. For ads or explainers, add text later in an editor instead of relying on the generation step.
Write a prompt that respects the image. Do not ask for a completely different character, setting, or lighting direction unless you expect visual changes. Better prompts describe what should move: hair fluttering, clouds drifting, product rotating slightly, water rippling, or a slow camera push.
For people or branded products, confirm that you have rights to use the image. Do not upload private, copyrighted, misleading, or sensitive images unless you have permission and a legitimate use case.
After generation, compare the output with the original reference. Check identity, proportions, logo integrity, background artifacts, and whether the added motion improves the clip. If the result changes too much, simplify the motion and use a more stable source image.