How do I make a person speak in a video?
Pair an image or video with audio in Lipsync, or let UGC Builder handle the whole flow from script to talking video.
The manual route
- 1Generate or upload a photo of the person, or use an existing video.
- 2Make the audio — generate it in Voice Generation, or use the text-to-speech built into the Lipsync page.
- 3Open Lipsync, pair the visual with the audio, and generate.
Which model appears depends on what you upload
Upload an image and you get the image models; upload a video and you get the video models. Everything is priced per second of audio.
| Input | Model | Credits/sec |
|---|---|---|
| Image | Kling Standard — best value | 7 |
| Image | Veed Fabric — 480p output | 10 |
| Image | Kling Pro — premium | 13 |
| Video | Pixverse — cheapest, and the default | 4 |
| Video | HeyGen Speed | 7 |
| Video | HeyGen Precision — highest accuracy | 12 |
HeyGen needs the video and audio lengths to match closely. If they do not, the page offers to auto-trim for you.
The guided route
UGC Builder in The Lab does all of this in one flow — it picks a creator, writes the script, generates the scene, records the voiceover and produces the talking video. Use it when you want a finished product review rather than a single lipsynced clip.
Last reviewed against the live product in August 2026.