Master video generation across all 24 models on Fauxto Labs — Sora 2, the VEO 3.1 family, Seedance, Kling, Gemini Omni Flash, Hailuo, WAN 2.6 and LTX 2. Learn the four generation modes, what each model costs per second, and how to prompt for cinematic results.
By Fauxto Labs•22 min read•Updated August 2026
What You'll Master in This Guide
This guide covers everything you need to go from zero to confident with AI-generated video. By the end you will understand the strengths of all 24 video models on the platform, know exactly what each one costs, and have a clear workflow for text-to-video, image-to-video, first-last-frame and reference-to-video generation. Specifically, you will learn:
How the 24 models compare in quality, speed and credit cost
How per-second pricing works, and the two models that break the rule
The four generation modes and when each one is the right tool
Advanced prompting for cinematic results
Camera movement vocabulary that unlocks professional-looking footage
Workflow optimisation and audio integration strategies
The AI Video Revolution
AI video generation has reached a breakthrough moment. With models like Sora 2, VEO 3.1, Seedance 2.5 and the Kling O3 family, creators can generate professional-quality video from a text description or a single still image. Fauxto Labs does not train these models — we aggregate them from ByteDance, Kuaishou, Google DeepMind, OpenAI, MiniMax, Alibaba and Lightricks, which means you get to pick the best engine for each shot instead of living with one vendor’s strengths and weaknesses.
What makes this moment different from earlier attempts is the sheer fidelity of the output. Modern models produce smooth motion, coherent lighting and believable physics, and most of them now generate synchronised audio in the same pass. No camera crew, no equipment, no location scouting — just a well-crafted prompt and a few minutes of generation time.
Text-to-Video
Create Videos from a Single Prompt
Text-to-video is the most intuitive way to start. Describe a scene in natural language — subjects, environment, lighting, camera movement — and the model generates a complete clip. It is ideal for conceptual exploration, storyboarding and rapid prototyping when you do not yet have reference imagery.
Supported by every model on the platform, with generation times ranging from one to eight minutes depending on the model you choose.
Image-to-Video
Animate Any Still Image
Image-to-video lets you upload a photograph or AI-generated image and bring it to life. Because you control the starting frame, the output is far more predictable — perfect for product shots, hero banners, or extending existing visual assets with motion.
All 24 models support image-to-video. Pair it with our image generator for an end-to-end pipeline that locks the frame before you spend video credits.
Native Audio
Sound That Matches the Scene
Most models now generate synchronised audio alongside the picture — the Seedance 2.x family, Seedance 1.5 Pro, Kling O3 and Kling 3, Gemini Omni Flash, VEO 3.1, Sora 2, Hailuo 03, Happy Horse, WAN 2.6 and LTX 2 among them. Footsteps match the ground, ambience fits the environment, and dialogue can be generated in context.
On Gemini Omni Flash you direct the audio in the prompt itself. Otherwise our Voice Generation and Music Generation tools let you add custom audio after the fact.
Before the model table, one rule that governs everything below it: video is priced per second of output on almost every model. A 10-second clip costs roughly double a 5-second clip, and where a model offers several resolutions the higher ones cost more per second. The maths is deliberately simple — Kling O3 Pro is 18 credits per second, so a 5-second clip is 90 credits and a 15-second clip is 270.
VEO 3.1 and VEO 3.1 Fast are the only exceptions: they bill per clip. VEO 3.1 is 225 credits without audio or 338 with, and VEO 3.1 Fast is 120 or 190, charged as a flat rate for a 5 to 8-second generation. Because the price does not move with duration, you should always request the full 8 seconds on those two models. VEO 3.1 Lite is not an exception — it is 10 credits per second, roughly 80 for an 8-second clip.
The exact total is always displayed on the generation page before you commit, so you never have to do this arithmetic under pressure.
Choosing the Right Video Model
Twenty-four models are available on the video generator, each with distinct strengths. The table below summarises the key differences; the prose that follows explains which to actually pick. Credit figures are per second of output unless the row says per clip.
Model
Provider
Credits
Duration
Resolution
Best For
Seedance 2.5
ByteDance
60/sec at 720p, 30/sec at 480p
4–30 sec
480p or 720p
The longest single clip on the platform
Seedance 2.0
ByteDance
35/sec
4–15 sec
720p
Cinematic clips with audio and 21:9 ultrawide
Seedance 2.0 Fast
ByteDance
30/sec
4–15 sec
720p
Same features, quicker turnaround
Seedance 2.0 mini
ByteDance
15/sec
4–15 sec
480p or 720p
Best value in the Seedance 2 family
Seedance 1.5 Pro
ByteDance
7/sec
3–12 sec
720p
The cheapest route to native audio
Seedance 1.5 Lite
ByteDance
5/sec
3–12 sec
Up to 1080p
Drafts and high-volume work
Kling O3 Pro
Kuaishou
18/sec
3–15 sec
1080p
Best all-round value
Kling O3 4K
Kuaishou
50/sec
3–15 sec
4K
Big-screen hero shots
Kling O3
Kuaishou
14/sec
3–15 sec
—
A strong default for animating stills
Kling 3 Pro
Kuaishou
25/sec, 40/sec with audio
3–15 sec
—
Multi-shot sequences
Kling 3
Kuaishou
20/sec, 30/sec with audio
3–15 sec
—
Kling V3 quality at a lower rate
Kling 2.6
Kuaishou
15/sec with audio, 10/sec first-last frame
5 or 10 sec
—
The cheapest first-last-frame option
Gemini Omni Flash
Google
15/sec
3–10 sec
16:9 or 9:16
Affordable Google video with prompt-driven audio
VEO 3.1
Google DeepMind
225 per clip, 338 with audio
5–8 sec
720p or 1080p
Premium cinematic quality
VEO 3.1 Fast
Google DeepMind
120 per clip, 190 with audio
5–8 sec
720p or 1080p
VEO quality at roughly half the cost
VEO 3.1 Lite
Google DeepMind
10/sec (about 80 for 8 sec)
Up to 8 sec
1080p
The most affordable route into VEO
Sora 2
OpenAI
12/sec
4–12 sec
—
Rapid iteration with strong physical realism
Sora 2 Pro
OpenAI
35/sec, 60/sec at 1024p, 85/sec at 1080p
4–12 sec
Up to 1080p
Production-quality cinematic footage
Hailuo 03
MiniMax
35/sec at 2K
5–15 sec
Up to 2K
Sharp 2K output and character handling
Hailuo Pro
MiniMax
12/sec
6 or 10 sec
1080p
Fast 1080p with prompt optimisation
Hailuo Lite
MiniMax
8/sec
6 or 10 sec
768p
Budget image animation, fastest turnaround
Happy Horse
Alibaba
32/sec at 1080p
3–15 sec
720p or 1080p
Character-driven, expressive motion
WAN 2.6
Alibaba Cloud
12/sec at 720p, 18/sec at 1080p
5, 10 or 15 sec
480p, 720p or 1080p
Motion-heavy and multi-shot scenes
LTX 2
Lightricks
6/sec at 1080p, 10/sec at 1440p, 18/sec at 4K
6–20 sec
Up to 4K
Long clips and the cheapest 4K
Start with Kling O3 Pro. At 18 credits per second it gives you 1080p, native audio, reference-to-video and first-last-frame across 3 to 15 seconds — a 5-second clip for 90 credits. Nothing else on the platform combines that resolution, that feature set and that price, which is why it is the best overall value pick. Kling O3 at 14 credits per second is the same idea slightly cheaper and is a great default for animating stills, while Kling O3 4K at 50 credits per second is there when you genuinely need 4K.
Gemini Omni Flash is the other value pick, at 15 credits per second for 3 to 10 seconds in 16:9 or 9:16 — 150 credits for a full 10-second clip with audio you direct through the prompt. It is a featured model for good reason: it is the cheapest way to get Google-quality video with sound.
Sora 2 from OpenAI is 12 credits per second, so 4 seconds is 48 credits, 8 seconds is 96 and the 12-second maximum is 144. Its strength is physical realism, which makes it a good choice for shots involving impacts, liquids or objects interacting. Sora 2 Pro is the production tier at 35 credits per second standard, 60 at 1024p and 85 at 1080p — expensive enough that you should only run final selects through it.
The VEO 3.1 family from Google DeepMind is the premium naturalistic option, and the one place where per-clip pricing applies. VEO 3.1 is 225 credits without audio and 338 with; VEO 3.1 Fast is 120 or 190 for close to the same quality. If you have been avoiding VEO on cost, VEO 3.1 Lite changes the picture entirely — 10 credits per second at 1080p with audio, about 80 credits for a full 8-second clip.
The Seedance family from ByteDance spans the widest range on the platform. Seedance 2.5 reaches 30 seconds in a single generation at 60 credits per second at 720p or 30 at 480p. Seedance 2.0 is 35 credits per second at 720p with 21:9 ultrawide, 2.0 Fast is 30 with quicker turnaround, and 2.0 mini is 15 with the same modes. At the value end, Seedance 1.5 Pro is 7 credits per second with native audio and Seedance 1.5 Lite is 5, reaching up to 1080p — a 10-second Lite clip is 50 credits, which is what makes it the drafting model of choice.
The Hailuo models from MiniMax cover three price points: Hailuo 03 at 35 credits per second for up to 2K and 15 seconds, Hailuo Pro at 12 for 1080p, and Hailuo Lite at 8 for 768p with a 1 to 2 minute turnaround. Pro and Lite have fixed 6 or 10-second durations, so the totals are simply 72 or 120 credits and 48 or 80.
WAN 2.6 from Alibaba Cloud is 12 credits per second at 720p or 18 at 1080p with fixed 5, 10 or 15-second durations, multi-shot support and audio — 180 credits for a 15-second 720p clip is exceptional value for that length, and it is one of the strongest models for motion-heavy scenes. Happy Horse, also from Alibaba, is 32 credits per second at 1080p and specialises in expressive, character-driven motion. LTX 2 from Lightricks is 6 credits per second at 1080p, 10 at 1440p and 18 at 4K across 6 to 20 seconds — both the cheapest long clips and the cheapest 4K on the platform.
All 24 models, one platform
Access Sora 2, VEO 3.1, Seedance, Kling, Gemini Omni Flash, Hailuo, WAN 2.6, Happy Horse and LTX 2 — compare results side by side and find the right model for your project.
Supply a start frame and an end frame and the model builds the movement between them. It is the most controllable mode there is, because you have fixed both ends of the shot and only handed over the interpolation. Supported on the Kling O3 family, Kling 2.6, the Seedance 2.x models and Hailuo 03.
Kling 2.6 is the cheapest route at 10 credits per second without audio, so a 5-second interpolation is 50 credits.
Reference Mode
Feed It References, Not Just Prompts
Reference-to-video accepts images, existing video clips and audio files as creative inputs, so the model matches your subject, palette and camera language instead of inferring them from text. It is the most reliable way to keep a character or product consistent across a set of clips. Available on Seedance 2.0, 2.0 Fast, 2.0 mini and 2.5, Hailuo 03, Gemini Omni Flash, the Kling O3 family and Happy Horse.
Including camera movement descriptions in your prompts dramatically improves video quality. Models have been trained on footage labeled with standard cinematography terms, so using the right vocabulary is one of the simplest ways to level up your results.
Basic Movements
Static shot — the camera does not move; good for dialogue or product showcases. Pan — horizontal rotation, useful for revealing a wide scene. Tilt — vertical rotation, great for revealing tall subjects like buildings or waterfalls. Zoom in / out — changes the focal length to draw attention toward or away from a subject. Push in — physically moves the camera closer, creating a sense of intimacy. Pull back — moves the camera away, often used for dramatic reveals.
Advanced Techniques
Dolly shot — smooth forward or backward movement on a track. Tracking shot — follows a moving subject laterally, keeping pace with the action. Crane shot — sweeps upward or downward from a high angle. Handheld — produces natural, slightly shaky movement for a documentary or found-footage feel. Steadicam — smooth, floating motion often used for walking-and-talking scenes. Drone shot — aerial perspective, perfect for landscapes and establishing shots.
Crafting Effective Video Prompts
A well-structured prompt is the difference between a generic clip and a cinematic one. Think of it as five layers stacked together:
"A person walking down a busy street, natural lighting, realistic." — This works, but the model has a lot of room to improvise. You will get a usable clip, but the framing and mood will be unpredictable.
Advanced Example
"Slow push-in shot of a barista crafting latte art in a cozy coffee shop, warm golden hour lighting streaming through large windows, steam rising from the cup, cinematic depth of field, 24fps, professional color grading." — Every layer of the prompt framework is present: camera movement (slow push-in), subject (barista + latte art), environment (coffee shop + windows), mood (warm golden hour), and technical details (depth of field, 24fps, color grading). The result is dramatically more controlled and visually rich.
Draft on a Cheap Model First
The most valuable habit in AI video is separating the prompt problem from the quality problem. Prompt structure, pacing and composition all read perfectly well on an inexpensive model, so iterate on Seedance 1.5 Lite at 5 credits per second — a 5-second test is 25 credits — until the shot is right, then regenerate the keeper on whichever premium model suits it. The same 5 seconds on Sora 2 Pro at 1080p is 425 credits, so getting the prompt wrong there is seventeen times more expensive.
Once you have your clips, stitch them together in the Video Editor, which is free to use and does not cost credits to export. If you would rather have the structure planned for you, Storyboard and Scene Builder both build multi-shot sequences, and The Lab’s Marketing Room and Video Agent handle guided end-to-end briefs.
Which video model should I use for the best quality?
Sora 2 Pro, VEO 3.1 and Kling O3 Pro lead on cinematic realism. Sora 2 Pro is 35 credits per second standard, 60 at 1024p and 85 at 1080p. VEO 3.1 bills per clip rather than per second at 225 credits without audio and 338 with. Kling O3 Pro is the cheapest of the three at 18 credits per second for 1080p, so a 5 second clip is 90 credits.
What's the most cost-effective video model?
Seedance 1.5 Lite at 5 credits per second is the cheapest solid option — a 5 second clip is 25 credits. LTX 2 is 6 credits per second at 1080p, Seedance 1.5 Pro is 7 and includes native audio, and Hailuo Lite is 8. If you care about quality per credit rather than the lowest absolute price, Kling O3 Pro at 18 credits per second and Gemini Omni Flash at 15 are the best value picks on the platform.
How do I add audio to my AI videos?
Most models generate audio natively, including the Seedance 2.x family, Seedance 1.5 Pro, the Kling O3 family, Kling 3 and 2.6, Gemini Omni Flash, the VEO 3.1 family, Sora 2 and Sora 2 Pro, Hailuo 03, Happy Horse, WAN 2.6 and LTX 2. On Gemini Omni Flash you direct the audio through the prompt itself. For models without native audio, or when you want a specific track, use our Voice Generation and Music Generation tools and lay the audio over the clip in the Video Editor.
What's the difference between text-to-video and image-to-video?
Text-to-video creates a clip from a description alone — good for conceptual work. Image-to-video animates a still you provide, which makes the output far more predictable because you control the opening frame. Two further modes are available on some models: first-last-frame, where you supply a start and end image and the model animates between them, and reference-to-video, where you attach images, clips or audio so the model matches your subject and style.
What aspect ratios can I use?
Supported ratios across the platform are 21:9, 16:9, 4:3, 3:2, 1:1, 2:3, 3:4, 9:16 and 9:21, though support varies by model — Gemini Omni Flash, for example, is 16:9 or 9:16 only. Set the ratio in the prompt card before generating.
How long does video generation take?
It varies by model. Hailuo Lite is the quickest at 1 to 2 minutes, Sora 2 takes 2 to 4 minutes, Sora 2 Pro 3 to 5, VEO 3.1 Fast 4 to 6, and VEO 3.1 is the longest wait at 6 to 8 minutes.
Ready to Create Stunning AI Videos?
Access all 24 video models — Sora 2, VEO 3.1, Seedance, Kling, Gemini Omni Flash, Hailuo, WAN 2.6, Happy Horse and LTX 2. Start creating professional video from text and images today, with no subscription and a commercial licence included.