Text-to-video, image-to-video, reference and first-last frame explained
Four ways to drive a video model, from a pure text description to matching specific subjects using reference media.
Text-to-video
Describe the shot and the model generates it from nothing. Maximum freedom, least control over exactly who or what appears. Best when the subject does not need to match anything specific.
Image-to-video
Give the model a starting image and it animates from there. This is the most reliable way to control what your video actually looks like, because you can perfect the frame in the Image Generator first — where iterations cost 2 to 30 credits instead of hundreds — and only then animate it.
Reference-to-video
Attach reference images, videos or audio and the model matches the subjects and style from them rather than inventing its own. Use this to keep a specific person, product or look consistent across clips. Supported on Seedance 2.0, 2.0 Fast, 2.0 mini and 2.5, Hailuo 03, Gemini Omni Flash, the Kling O3 family and Happy Horse.
First-last frame
Provide a start image and an end image, and the model generates the motion between them. Ideal for controlled transitions and transformations. Supported on Kling 2.6, the Kling O3 family, the Seedance 2.x family and Hailuo 03. Kling 2.6 at 10 credits per second is the cheapest option, which is why Montage Maker uses it.
The mode selector only offers what your chosen model supports, so if a mode is missing, switch models.
Last reviewed against the live product in August 2026.