← BlogKatama Blog
NewsModelsAugust 5, 20269 min read

Seedance 2.5 Is Here: 30 Seconds, Native Audio, and What It Costs

Seedance 2.5 landed on 31 July 2026, and the headline number is the one everyone noticed first: thirty seconds in a single generation. Until now, anything longer than about fifteen seconds meant stitching clips and praying the character survived the cut. That constraint shaped how everyone wrote prompts. It is gone.

But the duration is not the interesting part. The interesting part is what it costs and why.

Run the same prompt on every model and compare the result.
Open Video Studio

What actually changed

Duration: 4 to 30 seconds, natively

The API accepts any duration from 4 to 30 seconds. Not "up to 30 if you chain it" — one request, one clip, one continuous take. For narrative work this is the difference between a shot and a scene. A thirty-second product film no longer needs an edit; a thirty-second explainer can hold a single camera move from beginning to end.

The practical ceiling is still your prompt. A model that can run for thirty seconds will happily spend twenty of them drifting if you have not told it what happens in the second half. Write the beat, not just the look.

Audio in the same latent space

Earlier models bolted audio on afterwards, which is why footsteps never quite landed on the frame where the foot hit the ground. Seedance 2.5 generates picture and sound together. Ambience, impacts and room tone line up because they were never separate.

This matters most for anything with physical action — a bottle set down on a counter, a blade drawn, a door closing. Those are the moments where post-hoc audio always sounded like post-hoc audio.

Resolution: 480p and 720p only

There is no 1080p tier at launch. If you need higher, generate at 720p and upscale — the upscale path costs less than most people assume and preserves more than a native 1080p run at the same budget would suggest.

ModelLengthResolutionAspect ratiosNative audio
Seedance 2.5
30s, native audio
4–30s480p, 720p16:9 · 9:16 · 1:1 · 4:3 · 3:4 · 21:9Yes

Live from the Katama catalogue. Current credit costs are on the pricing page — they change with the app, not with this article.

The pricing model is different, and it is worth understanding

Most video models bill per second at a given resolution. Seedance 2.5 bills by pixels: height × width × seconds, converted to tokens. The formula the provider publishes works out to roughly (h × w × s × 24) / 1024 tokens.

The consequence is easy to miss: aspect ratio changes the price. At the same 720-pixel height, a 21:9 frame carries about 31% more pixels than 16:9. On a per-second model that difference is invisible. Here it is real money.

If you are building your own billing on top of a model like this, check whether it prices by time or by area before you decide what to charge. We found this in our own pricing and had to fix it.
Models — generated with Katama

Where it wins, and where it does not

Use it for: single-take scenes longer than ten seconds, anything where sound and picture have to agree, dialogue-free narrative beats, and product films where a cut would break the illusion of one continuous look.

Do not reach for it when: you need 1080p natively, you are making a five-second social hook (a faster model costs less for the same result), or your shot is fundamentally a still with a slow push — that is an image model plus a motion pass, and it will look better for a fraction of the cost.

Models — same prompt, different seed

Prompt structure that survives thirty seconds

Short clips forgive vague prompts because there is not enough time for drift to show. Thirty seconds does not forgive anything. The structure that holds up:

  1. Camera first. State the move before the subject. "Slow dolly in, then hold" gives the model a spine for the whole duration.
  2. One subject, described once. Repeating a description mid-prompt invites the model to re-interpret it halfway through.
  3. Beat the time. "First ten seconds: … then: …" Models respond to explicit temporal structure far better than to a list of adjectives.
  4. End state. Say where the shot lands. Without it, the last five seconds are a coin flip.

That last point is the single biggest quality difference we see between a good thirty-second generation and an expensive one that gets thrown away.

Doing this in Katama: Cinema Lab

Everything above is a method. If you want the method without assembling it by hand every time, that is what Cinema Lab is for — the infinite-canvas studio at katama.ai/cinema.

It is built around the four steps this article keeps coming back to:

  • Canvas — lay your references and key frames side by side. This is where drift becomes visible, because you are looking at the shots together rather than one after another.
  • Board — the shot list. Each card is a beat, and the beats stay in order while you change what is inside them.
  • Compose — the prompt for each shot, next to the frame it produces. Change one, see the other.
  • Timeline — the assembly, with the durations you actually generated rather than the ones you meant to.

Where consistency comes from

The part that matters for repeatable results is not the canvas — it is what sits behind it. A locked reference set plus a saved prompt structure means the tenth video in a series is built the same way as the first. That is the difference between a good clip and a body of work that looks like it came from one place.

And once a sequence works, it does not have to be rebuilt by hand: Workflows chains the steps into one pipeline, and Autopilot runs that pipeline on a schedule. Same references, same prompt structure, same look — on Tuesday and again three weeks later.

Start in Cinema Lab when the job is more than one shot. For a single clip, the Video Studio is faster and there is nothing to keep consistent.

Seedance 2NewsAI ModelsComparison
0 comments
Make it with KatamaRun the same prompt on every model and compare the result.Open Video Studio