Skip to main content
The LipSync API creates a lip-synced video or talking photo by synchronizing facial movements in a reference video or image with the driving audio. Choose Standard mode for faster processing or Precision mode for higher-quality results.

Inputs

Outputs

  • lip-synced video URL

How it works

1

Host your assets at a publicly accessible URL

Upload your video, photo, and audio files so our servers can retrieve them.
2

Send an API request

Provide the required inputs and any optional settings needed for the job.
3

Wait or query status

Use our webhook callback or poll the API with your job ID until processing is complete.
4

Download video output

Retrieve the finished talking photo or lip‑synced video from the provided URL.

Usage Limitation:

  • The default concurrency limit is 10 jobs. The limit is shared by all members of a team. See Concurrency limits for details.
  • Only single‑face videos or photos are supported.
  • Estimated queue time: 1–120 minutes, depending on system load.
  • Standard Mode processing time: ~10 minutes.
  • Precision Mode processing time: ~20 minutes.
If a video or photo contains multiple faces, only the largest detected face will be lip‑synced.

API Error Codes

Job Error Codes

Last modified on August 26, 2026