Make it with Shishō
Type your idea — Shishō builds it and drops you straight into the studio.
What it is
Multi-character lip sync extends single-face sync to full scenes. Upload a group shot or a sequence with several people, assign a voice to each face, and Katama drives every mouth independently. It is built for dialogue-heavy content — short films, explainer skits, podcast visualizers and ad creatives where more than one person talks.
How it works
- 01
Upload or generate the scene
Start from a group photo, a multi-person video clip, or a scene you generated inside Katama.
- 02
Assign voices to faces
Tag each person in the frame and pair them with a voice — upload audio or generate speech with AI TTS.
- 03
Sync and export
Katama drives each face independently, rendering accurate lip motion for every speaker in a single pass.
Why creators choose it
Frequently asked
How many faces can I sync at once?
The system handles scenes with several speakers — enough for dialogue, panels and group conversations.
Does each face get its own voice?
Yes. You assign a separate audio track or AI voice to each tagged face, and they move independently.
Can I use this for short films?
Absolutely — multi-character sync is designed for narrative scenes where realistic dialogue between characters matters.