HeyGen Avatar
HeyGen Avatar turns a photo, or a HeyGen avatar, into a lip-synced talking-avatar video. You supply the face and a speech source: an audio file that you host, or a script that HeyGen reads in a voice that you choose. The endpoint is asynchronous. Submit the job, poll until the status is terminal, then download the video. The request and response shapes mirror HeyGen’s API.Models
HeyGen offers two engines. They take the same speech sources, share every other field, and have the same price. They differ only in where the face comes from.The endpoint takes no
model field, so the ids above are not values that you send. The request body selects the engine. "type": "image" with an image object selects Avatar 4 Photo. "type": "avatar" with an avatar_id and "engine": {"type": "avatar_v"} selects Avatar 5 Digital.Endpoints
All endpoints share the base URLhttps://api.nunchux.ai.
Request
Send your API key in theX-API-Key header (see Authentication) and set Content-Type: application/json. Also send a User-Agent header that identifies your application, for example YourApp/1.0. A request that keeps the default User-Agent of your HTTP client can be rejected.
Send one face and one speech source in the body. The face fields select the engine. The speech fields are the same on both engines.
string
required
Selects the engine.
image selects Avatar 4 Photo and animates the photo in image. avatar selects Avatar 5 Digital and animates the avatar in avatar_id.Options: image, avatarobject
Avatar 4 Photo only. The photo to animate. Required when
type is image.string
Avatar 5 Digital only. The avatar to animate, from HeyGen’s avatar catalog. Required when
type is avatar.object
Avatar 5 Digital only. Selects the Avatar 5 renderer. Omit it when
type is image, because that body already selects Avatar 4 Photo.string
Both engines. Public HTTPS URL of a speech audio file (WAV or MP3), hosted by you. Send this or
script with voice_id. One speech source is required.string
Both engines. Text for the avatar to speak. Requires
voice_id. Use it instead of audio_url.string
Both engines. The voice that reads
script. Required when script is set. List the voices with GET /v1/heygen/v3/voices.string
default:"1080p"
Both engines. Output pixel tier: the size, not the shape. Resolution does not change the price or the render time.Options:
720p, 1080p, 4kstring
default:"16:9"
Both engines. Output canvas shape, independent of
resolution. See Resolution and aspect ratio.Options: auto, 16:9, 9:16, 4:5, 5:4, 1:1Resolution and aspect ratio
resolution sets the pixel tier and aspect_ratio sets the canvas. The dimensions below apply at 16:9. At any other ratio, the same tier is fitted to that canvas. Read the shape from aspect_ratio, never from the tier name.
Set
aspect_ratio explicitly. The default is 16:9, so a portrait face sent without it comes back inside a landscape frame with bars down both sides. auto has a different meaning on each engine.
- Avatar 5 Digital:
autofollows the shape of the avatar. Use it. - Avatar 4 Photo:
autodoes not read your photo. A 1080 × 1920 portrait comes back as 1280 × 720, cropped to the face, with no bars to signal it. Send the canvas nearest to the pixels of your photo:9:16for a 1080 × 1920 photo,4:5for a 3:4 phone portrait. A mismatched ratio is padded, not reframed. 1:1re-canvases every public-catalog avatar, since those are all landscape or portrait. A private or custom avatar can differ.
Avatar 4 Photo, photo and audio
Avatar 4 Photo, inline base64 photo
Avatar 5 Digital, catalog avatar and script
Voices
When you drive from ascript, choose a voice_id from the voices endpoint.
Response
Submit returns avideo_id. Poll it until status is completed or failed. Then download video_url.
Submit
object
required
The submitted job.
Poll
object
required
The job status.
Example
Submit, poll, then download. The example uses Avatar 4 Photo with a hosted photo and a hosted audio file. For Avatar 5 Digital, send"type": "avatar" with an avatar_id and "engine": {"type": "avatar_v"} instead of the image object. To drive from text, send script and voice_id instead of audio_url.
Tips
- Start from a strong face. On Avatar 4 Photo, send a clear, front-facing portrait with the face well lit and unobstructed. On Avatar 5 Digital, pick the
avatar_idwhose look fits the message. - You host the inputs. There is no upload endpoint. HeyGen downloads a URL during the job, not at submit, so the URL must stay publicly reachable until the job completes. An inline base64 photo needs no hosting.
- Match the speech source to the job. Point
audio_urlat polished narration that you host. Usescriptwithvoice_idfor fast iterations, for language swaps, and when you have nowhere to host an audio file. - Write scripts in natural, spoken phrasing. It lip-syncs better than dense written prose.
- Set
resolutionexplicitly. The default is1080p. - Set
aspect_ratioexplicitly. On Avatar 4 Photo, send the canvas nearest to the pixels of your photo. On Avatar 5 Digital, sendauto. See Resolution and aspect ratio. - Poll every 10 s at first, then back off to 30 s. Polls are free and take no job slot. Render time scales with clip length, so allow 10 minutes or more and set your client timeout to match.
- Branch on
failed. It is terminal, and the body carries the reason. - Download the video as soon as
statusiscompleted. The URL expires within hours. - Store
video_idas an unbounded string. It is about 230 characters long, and you must send it verbatim on the poll path. - A 5xx or a timeout on submit does not tell you whether the job was created. Do not resubmit blindly, because a second job that completes is billed too. If the submit returned a
video_id, poll it. If it returned nothing, wait, then read your credit balance before you submit again.
Errors and limits
- A 200 on submit means that HeyGen accepted the job. It does not mean that the job succeeded. A failure appears later as
failedin the poll response, with the reason in the body. - A failed job is never charged.
resolutionaccepts720p,1080pand4k. Any other value returns 400 with the codeunsupported_resolution. Nothing is charged.- A submit with no face (
imagewhentypeisimage,avatar_idwhentypeisavatar), with no speech source, or withscriptbut novoice_idreturns 400. - A
video_idis scoped to your account. Poll only a handle that your own submit returned. A handle that belongs to another account returns 409. - Neither engine has a duration control. The clip is as long as the audio or the script.
invalid_parameter. Host the file publicly, or send the photo inline as base64.