Seedance 2.5 Is Here: 30 Seconds, Native Audio, and What It Costs

Seedance 2.5 landed on 31 July 2026, and the headline number is the one everyone noticed first: thirty seconds in a single generation. Until now, anything longer than about fifteen seconds meant stitching clips and praying the character survived the cut. That constraint shaped how everyone wrote prompts. It is gone.
But the duration is not the interesting part. The interesting part is what it costs and why.
What actually changed
Duration: 4 to 30 seconds, natively
The API accepts any duration from 4 to 30 seconds. Not "up to 30 if you chain it" — one request, one clip, one continuous take. For narrative work this is the difference between a shot and a scene. A thirty-second product film no longer needs an edit; a thirty-second explainer can hold a single camera move from beginning to end.
The practical ceiling is still your prompt. A model that can run for thirty seconds will happily spend twenty of them drifting if you have not told it what happens in the second half. Write the beat, not just the look.
Audio in the same latent space
Earlier models bolted audio on afterwards, which is why footsteps never quite landed on the frame where the foot hit the ground. Seedance 2.5 generates picture and sound together. Ambience, impacts and room tone line up because they were never separate.
This matters most for anything with physical action — a bottle set down on a counter, a blade drawn, a door closing. Those are the moments where post-hoc audio always sounded like post-hoc audio.
Resolution: 480p and 720p only
There is no 1080p tier at launch. If you need higher, generate at 720p and upscale — the upscale path costs less than most people assume and preserves more than a native 1080p run at the same budget would suggest.
| Model | Length | Resolution | Aspect ratios | Native audio |
|---|---|---|---|---|
| Seedance 2.5 30s, native audio | 4–30s | 480p, 720p | 16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 21:9 | Yes |
Live from the Katama catalogue. Current credit costs are on the pricing page — they change with the app, not with this article.
The pricing model is different, and it is worth understanding
Most video models bill per second at a given resolution. Seedance 2.5 bills by pixels: height × width × seconds, converted to tokens. The formula the provider publishes works out to roughly (h × w × s × 24) / 1024 tokens.
The consequence is easy to miss: aspect ratio changes the price. At the same 720-pixel height, a 21:9 frame carries about 31% more pixels than 16:9. On a per-second model that difference is invisible. Here it is real money.
If you are building your own billing on top of a model like this, check whether it prices by time or by area before you decide what to charge. We found this in our own pricing and had to fix it.
Where it wins, and where it does not
Use it for: single-take scenes longer than ten seconds, anything where sound and picture have to agree, dialogue-free narrative beats, and product films where a cut would break the illusion of one continuous look.
Do not reach for it when: you need 1080p natively, you are making a five-second social hook (a faster model costs less for the same result), or your shot is fundamentally a still with a slow push — that is an image model plus a motion pass, and it will look better for a fraction of the cost.
Prompt structure that survives thirty seconds
Short clips forgive vague prompts because there is not enough time for drift to show. Thirty seconds does not forgive anything. The structure that holds up:
- Camera first. State the move before the subject. "Slow dolly in, then hold" gives the model a spine for the whole duration.
- One subject, described once. Repeating a description mid-prompt invites the model to re-interpret it halfway through.
- Beat the time. "First ten seconds: … then: …" Models respond to explicit temporal structure far better than to a list of adjectives.
- End state. Say where the shot lands. Without it, the last five seconds are a coin flip.
That last point is the single biggest quality difference we see between a good thirty-second generation and an expensive one that gets thrown away.
Doing this in Katama: Cinema Lab
Everything above is a method. If you want the method without assembling it by hand every time, that is what Cinema Lab is for — the infinite-canvas studio at katama.ai/cinema.
It is built around the four steps this article keeps coming back to:
- Canvas — lay your references and key frames side by side. This is where drift becomes visible, because you are looking at the shots together rather than one after another.
- Board — the shot list. Each card is a beat, and the beats stay in order while you change what is inside them.
- Compose — the prompt for each shot, next to the frame it produces. Change one, see the other.
- Timeline — the assembly, with the durations you actually generated rather than the ones you meant to.
Where consistency comes from
The part that matters for repeatable results is not the canvas — it is what sits behind it. A locked reference set plus a saved prompt structure means the tenth video in a series is built the same way as the first. That is the difference between a good clip and a body of work that looks like it came from one place.
And once a sequence works, it does not have to be rebuilt by hand: Workflows chains the steps into one pipeline, and Autopilot runs that pipeline on a schedule. Same references, same prompt structure, same look — on Tuesday and again three weeks later.
Start in Cinema Lab when the job is more than one shot. For a single clip, the Video Studio is faster and there is nothing to keep consistent.