Seedance 2.5 is ByteDance's newest video generation model, and the headline change is duration: it produces up to 30 seconds of video in a single generation, double what Seedance 2.0 allowed. It runs in three modes — text-to-video, image-to-video and reference-to-video — generates synchronized audio natively, and accepts reference sets of up to 30 images, 10 videos and 10 audio files at once.
That duration jump sounds incremental and is not. Almost every AI video workflow in production today is really a stitching workflow: generate five seconds, generate another five, hope the character's face and the lighting survive the cut, assemble in an editor. Seedance 2.5 removes the seams for anything under half a minute. One generation, one continuous take, one consistent audio bed.
Resolution runs 480p, 720p or 1080p, and the rate climbs steeply with it — 1080p costs two and a half times 720p, because the provider meters by pixels rather than by seconds. For a long take that adds up fast, so it is often cheaper to finish at 720p and run the result through our video upscaler than to pay the 1080p rate for the full 30 seconds.
Duration
One continuous take
A 24-second unbroken crane into a mountain village. No stitch. No cut. This is the shot that used to be six generations and a prayer.
Showcase
Every clip generated with Seedance 2.5. Click any of them for the prompt.
Native audio
Audio is born with the shot
Dialogue, cutlery, rain on the glass. The bed is generated in the same pass as the pixels — matched to the action, not glued on after.
Seedance 2.0 is still on the platform and still the right pick for plenty of jobs. Three differences decide which one you should reach for:
Seedance 2.0
Seedance 2.5
Duration
4–15 seconds
4–30 seconds
Resolution
720p
480p, 720p or 1080p
Cost per second
35 credits
30 (480p) / 60 (720p) / 150 (1080p)
Reference images
9
30
Reference videos
3
10 (30s total)
Reference audio
3
10 (30s total)
Native audio
Yes
Yes
Read that pricing row carefully, because it cuts both ways. At 720p, Seedance 2.5 is nearly twice the per-second cost of 2.0 — a 10-second 720p clip is 600 credits against 350. But at 480p it is actually cheaper than 2.0 at 30 credits per second, which makes it the best drafting model in the family. Iterate at 480p until the prompt is right, then run the keeper at 720p, and reserve 1080p for shots that genuinely need to hold up on a big screen.
Pricing & Specifications
Seedance 2.5 bills per second of output at 30 credits per second at 480p, 60 at 720p and 150 at 1080p. The tiers are not evenly spaced because the provider prices by pixels: a second of 1080p carries more than twice the pixels of 720p and costs more per pixel on top. There is no minimum charge beyond the 4-second floor and no per-generation fee.
Duration
480p
720p
1080p
4 seconds
120
240
600
6 seconds
180
360
900
10 seconds
300
600
1,500
15 seconds
450
900
2,250
20 seconds
600
1,200
3,000
25 seconds
750
1,500
3,750
30 seconds
900
1,800
4,500
One billing detail that catches people out: reference videos are charged too. Seedance 2.5 meters the runtime of any reference clip you attach alongside the output, because the model reads that footage as input. Attaching a video does drop the per-second rate to 36 credits at 720p, though, so a 10-second generation with 8 seconds of reference video is billed as 18 seconds at the lower rate — 648 credits rather than 600 for the clip on its own. Reference images and reference audio are not metered.
Supported aspect ratios are 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, plus an “auto” option that lets the model choose based on your prompt and references.
Text-to-Video
Thirty seconds from a sentence
Describe the shot and Seedance 2.5 returns a finished clip with audio. Give it three or four beats with timings and it will pace them, rather than looping one idea for 30 seconds.
Duration from 4 to 30 seconds. Audio is on by default and can be toggled off if you plan to lay your own track.
Image-to-Video
Animate a still, and keep going
Upload a photograph or a generated still and Seedance 2.5 animates it while respecting the original composition, grade and subject. Thirty seconds is long enough that a single still can carry a whole scene.
Pairs with our image generator for a text → still → video pipeline.
Reference-to-Video
Using a 12-image reference set of one product plus two short reference clips for camera language, generate a 20-second spot that matches the reference lighting and motion style exactly while introducing a new environment.
Reference-to-Video
Fifty files of direction
Attach up to 30 images, 10 video clips and 10 audio files, then point at them in the prompt with @Image1, @Video2 and @Audio1. Images set subject and palette; video teaches camera language; audio sets the bed.
Video and audio references each get a 30-second total runtime budget. Images have no runtime limit.
What thirty seconds unlocks
A complete spot, not a fragment
Thirty seconds is the slot agencies brief against. Hook, demonstration, benefit, call to action — write those as timed beats and the whole thing arrives as one file.
Six things thirty seconds unlocks
Duration limits shape what people even attempt. Here is work that was awkward or impossible at 15 seconds and is straightforward now.
1. A complete ad spot, not a fragment of one
Thirty seconds is the standard broadcast and pre-roll slot, and it is no accident that it is also the length agencies brief against. Previously you built a 30-second spot from six generations and prayed the grade matched. Now the whole thing arrives as one file with a continuous audio bed. If you need structure — hook, demonstration, benefit, call to action — write those as four timed beats in the prompt and the model will pace them.
2. Process and instruction video
Anything that teaches a sequence — a recipe, an assembly, a technique, a repair — falls apart when cut into five-second chunks, because the whole value is watching one continuous action. A 30-second unbroken shot of hands completing a task, with synchronized foley, is genuinely useful content rather than a stylish loop.
3. Property and space walkthroughs
Real estate, hospitality, retail and venue marketing all live on continuous movement through space. Viewers read a cut as a hidden edit, so a stitched walkthrough of a room actively undermines trust in what it shows. One unbroken 30-second glide through a floor plan reads as a real place.
4. A full verse for music and lyric video
A verse or a chorus is typically 15 to 30 seconds. Generate against the beat structure in your prompt, attach the audio as a reference so the visual rhythm keys off the real track, and you get a section you can drop into a longer edit without fighting continuity at every bar.
5. Episodic content with a consistent cast
The 30-image reference set matters more than it sounds for serial content. Build a reference library for a character — face from multiple angles, wardrobe, lighting, environment — and reuse it across every episode. Nine images was enough to suggest a character; thirty is enough to pin one down.
6. Establishing shots and B-roll beds
Editors do not want five-second B-roll; they want thirty seconds they can trim anywhere. Long atmospheric plates — a city waking up, weather crossing a landscape, a factory floor running — are cheap to generate at 480p, and give an edit room actual coverage to work with.
The mistake almost everyone makes on their first 30-second generation is reusing a five-second prompt. A single-sentence description gives the model one idea and 30 seconds to fill, and what comes back drifts, loops or slows to nothing. Long generations need structure.
Write beats, with timings
Break the runtime into three to five beats and give each one a duration, a camera move and an action: “0–8s: slow push in on the workbench as hands lay out tools. 8–18s: camera rises and tracks left following the assembly. 18–30s: pull back to a wide as the finished piece is lifted into the light.” This is the highest-leverage change you can make.
State what must not change
Over 30 seconds there is far more room for drift than over five. Spell out the constants — subject appearance, wardrobe, time of day, colour palette, lens character, whether the audio bed is continuous. A line like “same woman in the red wool coat throughout, overcast daylight, muted palette, no cuts” prevents the most common failure.
Say “single continuous shot” if that is what you want
Left to its own devices, the model will sometimes introduce cuts across a long generation. If you want an unbroken take, ask for one in those words. If you want cuts, describe them as cuts and the model will place them at your beats.
Describe the audio
Native audio is generated from your prompt. Naming the sonic elements — “grinder, chatter, door chime, no music” — produces a markedly better bed than leaving it to inference.
Draft at 480p
At 30 credits per second, 480p is the cheapest way to test whether a long prompt holds together. Prompt structure, pacing and continuity all read at 480p; only fine detail does not. Lock the prompt, then regenerate the keeper at 720p.
Put a 30-second prompt to the test
Draft at 480p, finish at 720p, and let native audio come along for the ride. Seedance 2.5 is live on the video generator now.
Seedance 2.5 does not replace anything on the platform. It owns one axis — length and reference depth — and gives up ground on others.
vs. Seedance 2.0: Use 2.0 for short 720p clips, where its flat 35 credits per second beats 2.5's 60. Use 2.5 when you need more than 15 seconds, a reference set larger than 9 images, or a cheap 480p draft.
vs. VEO 3.1: VEO leads on resolution and on naturalistic, documentary-feeling footage, and it is the better choice for a polished 8-second hero shot. It cannot give you 30 continuous seconds, and it has nothing comparable to a 50-file reference set.
vs. Sora 2: Sora 2 remains the value pick for exploring ideas in short clips. Once a project needs long takes or tight visual consistency against references, that cost advantage stops being the deciding factor.
vs. Kling and the O3 family: Kling is stronger on stylized and anime-leaning work and offers motion-control editing that Seedance does not. They complement each other — Kling for stylized shorts, Seedance 2.5 for long-form realism.
Up to 30 seconds in a single generation, with a 4-second minimum. That is double Seedance 2.0’s 15-second ceiling and the longest single-shot output of any video model on Fauxto Labs. Because it is one generation rather than several stitched together, lighting, wardrobe, character appearance and audio stay continuous for the full 30 seconds.
How much does Seedance 2.5 cost?
Pricing is per second of output and depends on resolution: 30 credits per second at 480p, 60 at 720p and 150 at 1080p. A 10-second 720p clip is 600 credits; a full 30-second 720p clip is 1,800 credits. At 480p the same clips are 300 and 900 credits, and at 1080p they are 1,500 and 4,500.
What resolution does Seedance 2.5 output?
Seedance 2.5 outputs 480p, 720p or 1080p. The rate rises steeply with resolution because the provider prices by pixels rather than by seconds — 1080p costs two and a half times 720p. For a long take it is often cheaper to finish at 720p and run the result through our video upscaler than to pay the 1080p rate for the full duration.
Does Seedance 2.5 generate audio?
Yes. Native synchronized audio is generated by default across text-to-video, image-to-video and reference-to-video. It covers ambience, effects and dialogue that lines up with the on-screen action, so a 30-second clip arrives as a finished scene rather than a silent plate.
How many reference files can Seedance 2.5 take?
Up to 30 images, 10 videos and 10 audio files in a single reference-to-video generation. Video and audio references each have their own 30-second total runtime budget — 30 seconds of video across all clips, and 30 seconds of audio across all files. Images have no runtime limit.
What is the difference between Seedance 2.5 and Seedance 2.0?
Three things matter most. Duration doubles from 15 to 30 seconds. The reference set grows from 9 images / 3 videos / 3 audio to 30 / 10 / 10. And pricing moves to a resolution-based rate — 30 credits per second at 480p, 60 at 720p and 150 at 1080p, versus a flat 35 at 720p for 2.0. Seedance 2.0 remains the cheaper choice for short 720p clips; 2.5 is what you use for long takes, big reference sets, cheap 480p drafting, and anything that needs 1080p.
Are reference videos charged as well as the output?
Yes. Seedance 2.5 meters the runtime of any reference video you attach alongside the output, because the model processes that footage as input. Attaching a video also drops the per-second rate to 36 credits at 720p, so a 10-second generation with 8 seconds of reference video is billed as 18 seconds at that lower rate — 648 credits. The cost estimate on the generation page includes attached reference footage, so the number you see before generating is the number you pay.
How should I prompt a 30-second Seedance 2.5 clip?
Write it as a shot list with timing, not a single description. Break the 30 seconds into three to five beats, give each one a duration, a camera move and an action, and state what carries across all of them (subject, wardrobe, lighting, palette). Seedance accepts very long prompts, so specificity costs you nothing — a vague prompt is the main reason long generations drift.
Is Seedance 2.5 better than VEO 3.1 or Sora 2?
They win on different things. VEO 3.1 leads on naturalistic, documentary-style footage; all three reach 1080p. Sora 2 is the value pick for short exploratory clips. Seedance 2.5 is the only one of the three that produces a continuous 30-second take with native audio and a 30-image reference set, which makes it the right tool for long-form single-shot work and for anything that has to stay visually consistent against a large reference library.
Can I use Seedance 2.5 for commercial projects?
Yes. Videos you generate on Fauxto Labs can be used commercially, including in paid advertising, client work and monetized social content. See our terms for the full detail.
Ready to try Seedance 2.5?
Thirty seconds of continuous video, native audio, and reference sets fifty files deep. Free credits to start — no card required.