What an AI Video Actually Costs: A Breakdown Nobody Publishes

The price of a generation is easy to look up. The price of a finished video is not, because it includes everything you generated and did not use.
Here is how the arithmetic actually works, with the parts people leave out.
1. The listed price is per attempt, not per result
A realistic hit rate for a first-time prompt on a new subject is somewhere between one in three and one in six. That is not a criticism of the models — it is what iteration looks like. But it means a clip with a listed cost of X has a real cost of three to six times X unless you are re-running something you have already dialled in.
The single biggest lever on your actual spend is not model choice. It is how many attempts you need before you accept a take, and that is a function of prompt discipline.
2. Duration is not always linear
Per-second models are linear: ten seconds costs twice five. Pixel-priced models are not — they scale with area × time, so a wider aspect ratio at the same height costs more for the same duration. If you are comparing two models by "price per second", you are comparing two different things.
3. Resolution costs more than the tier suggests
The jump from 720p to 1080p is usually more than the 2.25× the pixel count implies, because higher tiers often run on different hardware. Check the actual per-second rate at each tier rather than assuming a multiplier.
The practical advice: generate your exploration at the lowest resolution the model offers, and only run the take you have chosen at the resolution you need. Exploring at 1080p is burning money on frames you will delete.
4. Audio may or may not be included
Some models generate sound in the same pass at no extra cost. Others charge separately, and a few do not do audio at all — meaning a separate generation and a sync pass. When you compare two models on price, check whether the cheaper one leaves you with a silent clip.
5. The costs that are not credits
- Storage. Every take you keep. Not large per clip, relentless in aggregate.
- Time. A model that takes four minutes per generation and needs five attempts has cost you twenty minutes, whatever the credits say.
- Failed runs that still charge. If a job is cut off by a timeout after the provider has done the work, someone pays for it. Ask whether that is you.
A worked example
A thirty-second product film, 720p, synced audio, with a hit rate of one in four:
- Four generations at thirty seconds each — not one.
- Two of the four are close but wrong in a fixable way, so two more at the same length.
- One upscale pass on the winner.
Six generations plus an upscale for one deliverable. If you budgeted for one, you are out by a factor of six, and no pricing page told you that — because no pricing page can.
How to actually reduce it
Storyboard in images first. An image generation costs a fraction of a video generation. Find your framing, your lighting and your subject as a still, then animate the still you already like. This is the single highest-leverage habit in the whole workflow.
Lock what works. When a prompt structure produces a good take, save it. Re-deriving a prompt you already solved is the most common invisible cost in this work.
Explore short, deliver long. Test the idea at four seconds. Only extend once the idea is right.
Doing this in Katama: Cinema Lab
Everything above is a method. If you want the method without assembling it by hand every time, that is what Cinema Lab is for — the infinite-canvas studio at katama.ai/cinema.
It is built around the four steps this article keeps coming back to:
- Canvas — lay your references and key frames side by side. This is where drift becomes visible, because you are looking at the shots together rather than one after another.
- Board — the shot list. Each card is a beat, and the beats stay in order while you change what is inside them.
- Compose — the prompt for each shot, next to the frame it produces. Change one, see the other.
- Timeline — the assembly, with the durations you actually generated rather than the ones you meant to.
Where consistency comes from
The part that matters for repeatable results is not the canvas — it is what sits behind it. A locked reference set plus a saved prompt structure means the tenth video in a series is built the same way as the first. That is the difference between a good clip and a body of work that looks like it came from one place.
And once a sequence works, it does not have to be rebuilt by hand: Workflows chains the steps into one pipeline, and Autopilot runs that pipeline on a schedule. Same references, same prompt structure, same look — on Tuesday and again three weeks later.
Start in Cinema Lab when the job is more than one shot. For a single clip, the Video Studio is faster and there is nothing to keep consistent.