Skip to main content

Kling

Kling is Kuaishou’s video generation model family. Every Kling endpoint is asynchronous: submit a job, poll it until it reaches a terminal status, then download the output. The routes mirror Kling’s own API, so the request bodies, response shapes and error codes follow Kling’s documentation. You call them with your Nunchux API key and need no Kling account.

Models

std is the standard tier. It is the fastest and the most economical. pro gives sharper detail and stronger motion. 4k gives 4K-resolution output. Higher tiers take longer. Start with kling-v3. It supports the std, pro and 4k modes, and it is the only model with optional native audio. Use kling-v3-omni on the omni video endpoint for multi-reference or storyboard work. For cost-sensitive jobs, use std. It is the cheapest per second.

Endpoints

The base URL is https://api.nunchux.ai. Each submit returns a task_id immediately. Poll the same capability that you submitted to, with the task_id appended to the route.

Common fields

Send your API key in the X-API-Key header (see Authentication). Send Content-Type: application/json on every submit. Send a User-Agent header that identifies your application, such as YourApp/1.0. The request fields below appear on more than one endpoint.
string
required
The model to generate with. See Models. Defaults to kling-v3 on text-to-video, image-to-video and motion control. Defaults to kling-v3-omni on omni video.Options: kling-v3 on text-to-video, image-to-video and motion control. kling-v3-omni on omni video.
string
default:"std"
Quality and speed tier. Higher tiers take longer. Motion control accepts std and pro only.Options: std, pro, 4k
string
required
Description of the video to generate. Max 2500 characters. On image-to-video, describe the motion or the transformation. On omni video, refer to an input with a placeholder such as <<<image_1>>>. Not sent on motion control. Ignored in storyboard mode.
string
What to exclude from the video. Max 2500 characters. Text-to-video and image-to-video only.
string
default:"5"
Output length in seconds, sent as a string. Billed per second. Common values are "5" and "10". Not sent on motion control, where the output length follows the reference video. Ignored in storyboard mode.Range: 3-15.
string
default:"off"
Native audio generation. Only on kling-v3. Text-to-video and image-to-video only. In 4k mode, audio is always included and this flag is ignored.Options: on, off
string
If set, Kling POSTs status updates to this URL, and you do not need to poll. Text-to-video and image-to-video only.

Text-to-video

POST https://api.nunchux.ai/v1/klingai/videos/text2video Generate a video from a text prompt. Send model_name and prompt. Every other field is optional and falls back to its default. Use kling-v3.
string
default:"16:9"
Output aspect ratio.Options: 16:9, 9:16, 1:1
To plan a sequence of shots instead of one prompt, see Storyboard mode.

Image-to-video

POST https://api.nunchux.ai/v1/klingai/videos/image2video Animate a starting image into a video clip. Send model_name, image and a prompt that describes the motion. Use kling-v3.
string
required
The starting frame. Pass a public URL or raw base64, with no data:image/...;base64, prefix. Kling rejects the prefixed form with code 1201.

Omni video

POST https://api.nunchux.ai/v1/klingai/videos/omni-video Compose one clip from reference images, videos and elements. Use kling-v3-omni. At least one of image_list, video_list or element_list must be non-empty. You can send up to 7 reference inputs across the three lists. Refer to an input from the prompt with a 1-indexed placeholder in submission order: <<<image_1>>>, <<<video_1>>>, <<<element_1>>>.
array
Reference images.
array
Reference videos to continue or to composite from.
array
Reference elements, such as segmented objects or characters, to compose into the video.
To plan a sequence of shots, see Storyboard mode.

Motion control

POST https://api.nunchux.ai/v1/klingai/videos/motion-control Drive a still character with the motion of a reference video. Kling transfers the action from the driver clip and preserves the character’s identity. Use kling-v3. Do not send prompt or duration. The output length follows the reference video.
string
required
Public URL or raw base64 of the still character to animate, with no data:image/...;base64, prefix.
string
required
Public URL of the driver clip whose motion is transferred onto the character.
string
required
How the character reference is provided. image expects a still image in image_url (cap: 10 s). video expects a video in image_url (cap: 30 s).Options: image, video
integer
required
Length of the driver clip in seconds. This value drives per-second billing, so declare it accurately.Range: 1-10.
BillingKling bills per second of generated output. If you under-declare input_video_seconds, a follow-up charge is applied after the job completes. If you over-declare, the difference is not refunded automatically.

Storyboard mode

Text-to-video and omni video accept a multi-shot storyboard. Set multi_shot to true and describe each shot in multi_prompt. The top-level prompt and duration are ignored. The total length is the sum of the per-shot durations. All constraints are validated before billing, so a malformed request returns a 400 with no charge.
boolean
default:false
Enable multi-shot storyboard mode.
string
How shots are planned. Required when multi_shot is true.Options: customize, intelligence
array
Per-shot instructions. Text-to-video accepts up to 6 shots and 15 seconds in total. Omni video caps the total at 30 seconds.

Poll a task

GET https://api.nunchux.ai/v1/klingai/videos/{capability}/{task_id} {capability} is one of text2video, image2video, omni-video and motion-control. Poll the same capability that you submitted to. A task_id polled at another capability’s route returns 404. Poll every 5 to 10 seconds at first, then back off to 15 to 30 seconds. Polls do not use credits and do not count toward your concurrency limit. task_status moves from submitted to processing, then to succeed or failed. A task handle expires after 7 days. Output URLs expire within hours, so download the output as soon as the task succeeds.
integer
required
0 on success. A non-zero value indicates a Kling error.
string
required
Human-readable status, such as SUCCEED.
object
required
The task payload.
The submit response has the same shape, with task_status set to submitted. A poll response when the task has succeeded:

Example

Submit a text-to-video job, poll it until it reaches a terminal status, then download the clip.

Tips

  • Poll; do not busy-wait. Poll every 5 to 10 seconds at first, then back off to 15 to 30 seconds.
  • Expect text-to-video and image-to-video to finish in 60 to 120 seconds. A kling-v3 std 5-second clip takes about 45 seconds. Motion control and omni video run 2 to 6 minutes, and omni video takes longer with more shots.
  • Set a client timeout of at least 5 minutes for std text-to-video and image-to-video. Use at least 10 minutes for pro, 4k, motion control and omni video.
  • Download outputs as soon as the task succeeds. Output URLs expire within hours.
  • Keep prompts concrete and visual. Describe what the camera sees, not abstract concepts.
  • Use negative_prompt to rule out unwanted elements such as blur, text overlays and watermarks.
  • Set sound: "on" on kling-v3 to add ambient audio with no post-processing. In 4k mode, audio is always included.
  • Send base64 images as raw bytes with no data:image/...;base64, prefix. Public URLs must be reachable by Kling’s servers: no auth walls and no localhost.
  • Use a clear, well-composed starting frame for image-to-video. Kling extends what it sees, so a blurry or cluttered input produces inconsistent motion.
  • Use a clean, well-lit character image for motion control. Identity preservation is only as good as the source.
  • Pass a callback_url on text-to-video or image-to-video to receive status updates without polling.

Errors and limits

A 200 means that Kling accepted the job and billing started, not that the job succeeded. Do not resubmit after a 200; a second submit creates and bills a second job. Retry only on a network or timeout error that arrives before the submit response. Retry a 5xx or a transient 429 with backoff. Never retry another 4xx; the request is invalid as sent. For the general error contract, see Error codes. For the concurrency caps and request limits, see Rate limits. Every job counts toward your plan’s simultaneous-jobs cap; polls do not. Credits are deducted at submit and refunded when a job fails; see Credits & pricing and the Partner models overview. Kling-specific limits:
  • prompt and negative_prompt: up to 2500 characters.
  • duration: 3 to 15 seconds. In storyboard mode, text-to-video accepts up to 6 shots and 15 seconds in total, and omni video accepts up to 30 seconds in total.
  • Omni video: up to 7 reference inputs across image_list, video_list and element_list.
  • Motion control: input_video_seconds from 1 to 10, and mode is std or pro.
  • Voice control is not supported on kling-v3. Sending voice_control or voice_list returns a 400.

Migrating from /v1/video/kling/*

Kling used to live under /v1/video/kling/*. It now lives under /v1/klingai/videos/*, a faithful mirror of Kling’s own API: same request bodies, same response shapes, same error codes. If you call Kling directly today, you can point at us by swapping the base URL and using your Nunchux key as the Bearer token. Existing /v1/video/kling/* integrations keep working; migrate when convenient.
Two things that are not a find-and-replace
  • Polling moved from one URL to one per capability. There is no /v1/klingai/videos/tasks/{task_id}. Instead of a single GET /v1/video/kling/tasks/{task_id}, poll the endpoint you submitted to. For example, a job from POST /v1/klingai/videos/text2video is polled at GET /v1/klingai/videos/text2video/{task_id}. A shared poll helper needs to carry the submit path through; a task_id polled at another capability’s URL returns 404.
  • One path was renamed beyond the prefix change: video-effects → effects. The rest keep their trailing segment.