Make it with Shishō
Type your idea — Shishō builds it and drops you straight into the studio.
What it is
AI lip sync matches a face to an audio track so the lips move exactly with the words. Katama takes a portrait or a video of a person and a voiceover, then drives realistic mouth and jaw motion frame by frame. It is how you build talking avatars, dub footage into new languages, and make any character speak.
How it works
- 01
Choose your face
Upload a portrait or a clip of a person, or pick an AI avatar you want to bring to life.
- 02
Add the voice
Provide an audio file or generate a voiceover with AI text-to-speech in the voice and language you want.
- 03
Sync and export
Katama aligns the mouth to the audio and renders a clean talking clip ready to publish or dub.
Why creators choose it
Frequently asked
Can I lip-sync a single photo?
Yes — a single portrait is enough to produce a talking avatar, and you can also drive an existing video clip of a person.
Can I use it to dub videos?
Absolutely. Pair lip sync with a translated voiceover and the speaker's mouth matches the new language, so your footage looks natively dubbed.
Where does the voice come from?
Bring your own audio or generate one with ElevenLabs-grade text-to-speech right inside Katama, then sync it to the face.