Skip to main content

How it works

1

Host your assets at a publicly accessible URL

Upload your video, photo, and audio files so our servers can retrieve them.
2

Send an API request with the appropriate parameters

Reference your hosted assets and specify your desired mode (Standard or Precision).
3

Wait or query status

Use our webhook callback or poll the API with your job ID until processing is complete.
4

Download video output

Retrieve the finished talking photo or lip‑synced video from the provided URL.

Usage Limitation:

  • You may have up to 10 concurrent jobs (including queued requests).
  • Only single‑face videos or photos are supported.
  • Estimated queue time: 1–120 minutes, depending on system load.
  • Standard Mode processing time: ~10 minutes.
  • Precision Mode processing time: ~20 minutes.
If a video or photo contains multiple faces, only the largest detected face will be lip‑synced.

API Error Codes

Job Error Codes

Last modified on May 21, 2026