Inputs
Outputs
- lip-synced video URL
How it works
1
Host your assets at a publicly accessible URL
Upload your video, photo, and audio files so our servers can retrieve them.
2
Send an API request
Provide the required inputs and any optional settings needed for the job.
3
Wait or query status
Use our webhook callback or poll the API with your job ID until processing is complete.
4
Download video output
Retrieve the finished talking photo or lip‑synced video from the provided URL.
Usage Limitation:
- The default concurrency limit is 10 jobs. The limit is shared by all members of a team. See Concurrency limits for details.
- Only single‑face videos or photos are supported.
- Estimated queue time: 1–120 minutes, depending on system load.
- Standard Mode processing time: ~10 minutes.
- Precision Mode processing time: ~20 minutes.
If a video or photo contains multiple faces, only the largest detected face will be lip‑synced.