> ## Agent Instructions
> Authenticate every request with your API key in the `X-API-Key` header or `Authorization: Bearer`. Keys start with `sk-nunchux-`.
> Image generation is synchronous: the image is in the response.
> Video generation is asynchronous: submit the job, poll its task until it reaches a terminal status, then download the output promptly, because output URLs expire.
# Authentication
Source: https://docs.nunchux.ai/authentication
Learn how to authenticate with the Nunchux API using API keys.
# Authentication
### How to authenticate your API requests
## Requirements for API calls
All API requests to call any model on [Model APIs](/nunchux-optimized/overview) require authentication with your API key. Pass it in either the `X-API-Key` header or the `Authorization: Bearer` header. Both are accepted on every endpoint, and the same key works for both.
Which header should I use?
Either one. The examples in these docs use `X-API-Key`. The [OpenAI SDK](/nunchux-optimized/openai-compatibility) sends `Authorization: Bearer` automatically. If you send both headers, `Authorization: Bearer` takes precedence.
[Sign up](https://nunchux.ai/sign-up) for a Nunchux account.
Go to your [Dashboard](https://nunchux.ai/dashboard) and navigate to the API Keys section. Click "Create New Key" to generate a new API key, which you'll use to securely [access the API Reference](/nunchux-optimized/overview).
Store your key securely — your API key is shown only once when created. Copy and store it in a secure location. If you lose it, you'll need to generate a new one.
Once you've generated an API key, set it as an [environment variable](https://en.wikipedia.org/wiki/Environment_variable).
```bash Shell theme={"system"}
export NUNCHUX_API_KEY=xxxxx
```
Include your API key in every request using the `X-API-Key` header.
```text Header format theme={"system"}
X-API-Key: YOUR_API_KEY
```
## Example Request
```bash cURL theme={"system"}
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nunchux-flux.2-klein-4b",
"prompt": "a serene landscape",
"width": 1024,
"height": 1024
}'
```
```python Python theme={"system"}
import os
import requests
api_key = os.environ.get('NUNCHUX_API_KEY')
headers = {
'X-API-Key': api_key,
'Content-Type': 'application/json'
}
response = requests.post(
'https://api.nunchux.ai/v1/images/generations',
headers=headers,
json={
'model': 'nunchux-flux.2-klein-4b',
'prompt': 'a serene landscape',
'width': 1024,
'height': 1024
}
)
```
```javascript JavaScript theme={"system"}
const apiKey = process.env.NUNCHUX_API_KEY;
const response = await fetch("https://api.nunchux.ai/v1/images/generations", {
method: "POST",
headers: {
"X-API-Key": apiKey,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "nunchux-flux.2-klein-4b",
prompt: "a serene landscape",
width: 1024,
height: 1024,
}),
});
const data = await response.json();
```
## Best Practices
### Store API keys securely
* Use environment variables to store API keys
* Never hardcode keys in your source code
* Use secret management systems in production
* Rotate keys periodically
```bash theme={"system"}
# Store in environment variable
export NUNCHUX_API_KEY="your-api-key-here"
# Use in requests
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nunchux-flux.2-klein-4b", "prompt": "test", "width": 1024, "height": 1024}'
```
### Use environment-specific keys
Maintain separate API keys for different environments:
* Development key for testing
* Staging key for pre-production
* Production key for live applications
### Rotate a key
You can rotate a key from the dashboard without downtime. Open API Keys and select the rotate icon next to the key:
* A new key is created with the same name and shown one time
* The old key continues to work for 24 hours, then it stops
* Update your applications to the new key in that window
### Monitor API usage
Track your API usage to:
* Detect unauthorized access
* Optimize costs
* Identify usage patterns
* Plan capacity needs
### Rate Limiting
Every key is limited per plan on requests per minute and simultaneous jobs; over a cap you receive a 429. The numbers and how to back off are on the [Rate limits](https://nunchux.ai/docs/rate-limits) page.
# Changelog
Source: https://docs.nunchux.ai/changelog
Models available in the Nunchux playground and when they arrived: image, video, and avatar models, newest first.
Each date is the day the model became generally available on Nunchux, per the earliest release evidence on record — a recorded enable outranks route presence, and the launch catalog is dated to the site launch.
* Support Nano Banana 2.1 — text-to-image and image editing
* Support HappyHorse 1.0 — text-to-video, image-to-video, and reference-to-video
* Support HappyHorse 1.1 — reference-to-video
* Support MiniMax H3 — reference-to-video
* Support Seedance 2.5, 2.0, 1.5 Pro, and 1.0 Pro — text-to-video and image-to-video
* Support Wan 2.7 and Wan 2.6 — text-to-video and image-to-video
* Support HappyHorse 1.1 — text-to-video and image-to-video
* Support MiniMax H3 — text-to-video and image-to-video
* Support Qwen Image 2512 and Qwen Image Edit 2511 — text-to-image and image editing
* Support HeyGen Avatar V — talking-avatar video
* Support Nano Banana — text-to-image and image editing
* Support Gemini 3 Pro Image — text-to-image and image editing
* Support Veo 3.1 — text-to-video and image-to-video
* Support HeyGen Avatar IV — talking-avatar video
* Support FLUX.1 Schnell — text-to-image
* Support Kling V1 — text-to-video and image-to-video
* Support Kling V1.5 — image-to-video
* Support Kling V1.6 — text-to-video, image-to-video, and multi-image
* Support Kling V2 — text-to-video and image-to-video
* Support Kling V2.1 — image-to-video
* Support Kling V2.1 Master — text-to-video and image-to-video
* Support Kling V2.5 Turbo — text-to-video and image-to-video
* Support Kling V2.6 — text-to-video, image-to-video, and motion control
* Support Kling V3 — text-to-video, image-to-video, omni video, and motion control
* Support Kling Video O1 — omni video
* Support FLUX.2 Klein 4B and 9B — text-to-image and image editing
* Support Qwen Image — text-to-image and image editing
* Support Wan 2.2 Lightning — text-to-video and image-to-video
# Use Nunchux with coding agents
Source: https://docs.nunchux.ai/coding-agents
Connect Claude Code, Cursor and other AI coding tools to the Nunchux docs through MCP, llms.txt, and Markdown versions of every page.
# Use Nunchux with coding agents
### Give your AI coding tool the docs it needs to write working Nunchux code
AI coding tools write better API code when they read the docs directly rather than guessing. Every page of these docs is available in forms they can read, and the menu at the top of each page sends it to them in one click.
## Connect the docs MCP server
The docs MCP server lets your coding tool search these docs and read any page while it works, without you pasting anything in.
* **Server URL:** `https://nunchux.ai/docs/mcp`
* **One click:** choose **Add MCP** in the menu at the top of any page.
```bash Claude Code theme={"system"}
claude mcp add --transport http nunchux https://nunchux.ai/docs/mcp
```
```json Cursor (~/.cursor/mcp.json) theme={"system"}
{
"mcpServers": {
"nunchux": {
"url": "https://nunchux.ai/docs/mcp"
}
}
}
```
## Read any page as Markdown
Add `.md` to the end of any page URL to get the page as plain Markdown, for example [`/docs/authentication.md`](https://nunchux.ai/docs/authentication.md). The page menu does the same: **Copy page** copies it, and **Open in Claude** or **Open in ChatGPT** starts a chat with the page loaded.
Each Markdown page starts with a short **Agent Instructions** block: how to authenticate, and which calls are synchronous and which you poll.
## Give an agent the whole site
* [`llms.txt`](https://nunchux.ai/docs/llms.txt) is an index of every page with a one-line summary, so an agent can find the page it needs.
* [`llms-full.txt`](https://nunchux.ai/docs/llms-full.txt) is the full text of every page in one file, for tools that load everything up front.
* [`skill.md`](https://nunchux.ai/docs/skill.md) is a skill file for agents that support skills: a summary of the API and when to use it.
Your API key still goes in your code or environment, never in a chat. Set it as `NUNCHUX_API_KEY` and have the agent read it from there. See [Authentication](/authentication).
# Credits & Pricing
Source: https://docs.nunchux.ai/credits-pricing
Understand how credits work for the Nunchux API.
# Credits & Pricing
### How billing works for your API requests
## How Credits Work
Nunchux uses a prepaid credit system. One credit costs \$1.00.
Credits are deducted per successful request, based on the model, the [performance tier](/performance-tiers), and the output size. How a model bills depends on what it generates.
| Model type | Billing basis |
| - | - |
| Image models | Most image models price per megapixel of the output. Ideogram 4 prices per image, by step count. |
| Video models | Most are priced by the duration of the output, some also by resolution or mode. Seedance is priced per million video tokens of the finished clip, billed when it completes. |
The [pricing page](https://nunchux.ai/pricing) lists the current rate for each combination.
## Tier Pricing
The price of a Nunchux Optimized model depends on the tier. The FLUX, Qwen and HiDream O1 models have the Radical Speed and Radical Value tiers. Ideogram 4 has the Turbo, Balanced and Quality tiers, billed per image. The LTX video models have the 720p and 1080p tiers, billed per second of output video. Partner models do not use the tiers, and each one has its own price basis, such as output size, resolution, or mode. See [Performance Tiers](/performance-tiers) for a description of each tier, and the [pricing page](https://nunchux.ai/pricing) for current rates.
## Discounts
Negotiated and trial discounts apply to those rates. See [Discounts](/discounts) to see how to check which apply to your account.
## Checking Your Balance
Check your credit balance in the [Dashboard](https://nunchux.ai/dashboard). You can also call `GET /v1/credits`, which takes no parameters and returns your own account only. It uses the same API key for [authentication](/authentication) as every other endpoint.
```bash curl theme={"system"}
curl https://api.nunchux.ai/v1/credits \
-H "X-API-Key: $NUNCHUX_API_KEY"
```
```python python theme={"system"}
import os
import requests
response = requests.get(
'https://api.nunchux.ai/v1/credits',
headers={'X-API-Key': os.environ['NUNCHUX_API_KEY']},
timeout=30,
)
response.raise_for_status()
print(response.json())
```
A `200` carries the balance and the caps that apply to your plan. The API returns [standard error codes](/errors) with a JSON error body.
```json theme={"system"}
{
"credits_remaining": 96.49,
"plan_level": "pro",
"rpm_limit": 200,
"concurrent_limit": 10
}
```
| Field | Type | Description |
| - | - | - |
| `credits_remaining` | number | Credits left on the account. One credit is \$1.00. |
| `plan_level` | string | The plan the account is on. Each plan has its own caps. |
| `rpm_limit` | number | Requests per minute this plan allows. |
| `concurrent_limit` | number | Simultaneous jobs this plan allows. |
Polling
This poll is free and it never takes a job slot. It still counts toward your [requests per minute](https://nunchux.ai/docs/rate-limits).
## Credit Packages
Add credits from the [Billing tab](https://nunchux.ai/dashboard/billing) of your dashboard. Select a preset package, or enter a custom whole-dollar amount between $5 and $5,000.
| Package | Credits | \~Images |
| - | - | - |
| \$10 | 10 | \~15,900 |
| \$20 | 20 | \~31,800 |
| \$50 | 50 | \~79,500 |
Image estimates assume FLUX.2 Klein 4B on the Radical Value tier at 1024×1024.
This is our lowest published rate, 0.0006 credits per megapixel. Other models
and tiers have different costs per image.
## Insufficient Credits
If you do not have enough credits, the API returns a 402 status code. The response gives the credits required and your current balance:
```json theme={"system"}
{
"error": {
"code": "insufficient_credits",
"message": "Your account does not have enough credits for this request.",
"creditsRequired": 0.005,
"creditsBalance": 0.002
}
}
```
See [Error Handling](/errors) for a full list of API error codes.
## Refunds
Sometimes a request fails after Nunchux deducts the credits, for example because of a server error. Nunchux refunds those credits automatically.
# Discounts
Source: https://docs.nunchux.ai/discounts
A multiplier on list price, applied per entrypoint.
# Discounts
### A multiplier on list price, applied per entrypoint.
## How discounts work
Your charge is list price × multiplier. The multiplier is specific to your account and to one entrypoint. An entrypoint is a group of models, not a single model.
| Entrypoint | Covers |
| - | - |
| `kling` | Every Kling model, mode, and duration |
| `nunchux` | Nunchux-optimized image and video models |
| `google` | Every Google model. For example, Nano Banana image and Veo 3.1 video. |
| `heygen` | Every HeyGen avatar video model |
One multiplier covers every model behind its entrypoint. You hold at most one active discount per entrypoint, and entrypoints you pay list price for have no row at all.
## Checking your discounts
`GET /v1/discounts` takes no parameters and returns only your own rows. It uses the same API key for [authentication](/authentication) as every other endpoint.
```bash curl theme={"system"}
curl https://api.nunchux.ai/v1/discounts \
-H "X-API-Key: $NUNCHUX_API_KEY"
```
```python python theme={"system"}
import os
import requests
response = requests.get(
'https://api.nunchux.ai/v1/discounts',
headers={'X-API-Key': os.environ['NUNCHUX_API_KEY']},
timeout=30,
)
response.raise_for_status()
print(response.json())
```
A `200` carries a single `discounts` array, sorted by entrypoint. The API returns [standard error codes](/errors) with a JSON error body.
```json theme={"system"}
{
"discounts": [
{
"entrypoint": "kling",
"multiplier": 0.9,
"effective_from": "2026-06-10T00:00:00+00:00",
"effective_to": null
},
{
"entrypoint": "nunchux",
"multiplier": 0.9,
"effective_from": "2026-07-01T00:00:00+00:00",
"effective_to": "2026-09-30T00:00:00+00:00"
}
]
}
```
| Field | Type | Description |
| - | - | - |
| `entrypoint` | string | Which surface the discount applies to. One of the four entrypoints above. |
| `multiplier` | number | The fraction of list price you pay. |
| `effective_from` | string | ISO-8601 timestamp with UTC offset. When the discount began. Always present. |
| `effective_to` | string \| null | ISO-8601 timestamp, or null for an open-ended discount. When present, the discount ends at this instant (exclusive). |
No active discount is a normal `200`, not an error. An empty array means list price applies everywhere.
```json theme={"system"}
{
"discounts": []
}
```
The response carries `multiplier` and nothing precomputed, so derive any percentage you want to display yourself.
Polling
This is a free poll and it never takes a job slot. It still counts toward your [requests per minute](/rate-limits), and a discount changes rarely, so caching the response for minutes is fine. There is no webhook for discount changes.
## Reading the multiplier
The `multiplier` field is the fraction of list price you pay.
| Value | Meaning | How to display |
| - | - | - |
| 0 \< m \< 1 | Discount | discount = (1 − m) × 100 %, so m = 0.90 means a 10% discount |
| m = 1 | List price | No discount |
## Time windows
Every discount has a start date and may have an end date.
| Field | Meaning |
| - | - |
| `effective_from` | When the discount starts. |
| `effective_to` | When it ends. A null value means it has no end date. |
The endpoint returns only the discounts that are active now. One that has not started yet, or that has already ended, is left out entirely rather than returned as an inactive row.
# Errors
Source: https://docs.nunchux.ai/errors
API error codes, response formats, and retry guidance for the Nunchux API.
# Errors
### Validation and content errors returned
The Nunchux API uses standard HTTP status codes. All error responses include a JSON body describing what went wrong.
## Client Error Response Format
Client errors (4xx) return a JSON body. Auth, rate-limit, and billing errors carry an `error` object; request-validation failures carry a plain `detail` string.
| Field | Type | Description |
| - | - | - |
| `error.code` | string | A machine-readable error code (e.g. invalid\_api\_key, rpm\_limit\_exceeded). Use this for conditional logic. |
| `error.message` | string | A human-readable description of the error. Display this to end-users. |
| `detail` | string | Request-validation failures (unknown model, bad parameters) return this plain string instead of an error object. |
## Client Error Codes
### 400 Bad Request
The request is missing required parameters or contains invalid values such as missing prompt, invalid model ID, invalid dimensions. Returned as a plain detail string.
```json theme={"system"}
{
"detail": "unknown model 'invalid-model'; allowed values: nunchux-qwen-image-2512, nunchux-flux.2-klein-9b, ..."
}
```
Retry Guide — Do not retry. Fix the request parameters before resending.
### 401 Unauthorized
The API key in the X-API-Key header is missing, malformed, unknown, revoked, or expired (codes: missing\_api\_key, invalid\_key\_format, invalid\_api\_key, key\_revoked, key\_expired).
```json theme={"system"}
{
"error": {
"code": "invalid_api_key",
"message": "Invalid or revoked API key."
}
}
```
Retry Guide — Do not retry. Check that your API key is correct and active.
### 403 Forbidden
The account does not have access to the requested model. Restricted models need an explicit grant. This response carries `detail`, not the error envelope.
```json theme={"system"}
{
"detail": {
"code": "model_not_available",
"message": "Model nunchux-flux.2-klein-9b is restricted on this account",
"model": "nunchux-flux.2-klein-9b",
"entrypoint": "nunchaku"
}
}
```
Retry Guide — Do not retry. Use a public model, or contact support to request
access.
### 429 Rate Limit
Too many requests in a short period. Code rpm\_limit\_exceeded for the per-minute cap, concurrent\_limit\_exceeded for the simultaneous-job cap. Back off and retry.
```json theme={"system"}
{
"error": {
"code": "rpm_limit_exceeded",
"message": "Rate limit exceeded. See the rate-limit response headers for your limit and reset time."
}
}
```
Retry Guide — Implement exponential backoff. Start with a 1-second delay, then
double on each retry (1s → 2s → 4s → 8s). Cap at 3–5 retries.
### 402 Insufficient Credits
Your account does not have enough credits for this request. The response includes your current balance and the cost so you know exactly how many credits to add.
Credit metadata
The 402 body is the standard error object. `code`, `message` and `request_id` are always present. The credit amounts are not always included. When they are, they ride in `error.details`, with the deprecated `creditsRequired` and `creditsBalance` mirrors alongside.
| Field | Type | Description |
| - | - | - |
| `error.code` | string | Always "insufficient\_credits". |
| `error.message` | string | A human-readable description of the error. Display this to end-users. |
| `error.request_id` | string | Correlation id for this request. Always present. |
| `error.details.credits_required` | number | The credits this request needs. Not always included. |
| `error.details.credits_balance` | number | Your credit balance. Not always included. |
| `error.creditsRequired` | number | Deprecated mirror of error.details.credits\_required. |
| `error.creditsBalance` | number | Deprecated mirror of error.details.credits\_balance. |
```json theme={"system"}
{
"error": {
"code": "insufficient_credits",
"message": "Your account does not have enough credits for this request.",
"request_id": "",
"details": {
"credits_required": 0.005,
"credits_balance": 0.002
},
"creditsRequired": 0.005,
"creditsBalance": 0.002
}
}
```
Retry — Do not retry. Check your credit balance and purchase more credits if
needed. See [Credits & Pricing](/credits-pricing).
## Server Error Response Format
Server errors (5xx) return a JSON body with a `detail` string. When the inference backend supplies a typed envelope, an `error` object with a machine-readable `code` rides alongside it.
```json theme={"system"}
{
"detail": "Inference failed: "
}
```
## Server Error Codes
### 500 Internal Server Error
An unexpected error occurred on the server.
Retry Guide — Safe to retry with exponential backoff. Credits are
automatically refunded on failure.
### 502 Bad Gateway
The inference service returned no data. The request was processed but no output was generated.
Retry Guide — Safe to retry with exponential backoff. Credits are
automatically refunded on failure.
### 503 Service Unavailable
The server is temporarily unavailable.
Retry Guide — Wait a few seconds and retry.
### 504 Gateway Timeout
A request exceeded its time budget (code: `engine_timeout`). On an endpoint that takes a source image, fetching that image is the usual cause.
Retry Guide — Safe to retry with exponential backoff. Credits are
automatically refunded on failure.
## Retrying safely
Only `429` and the `5xx` family should be retried. Use exponential backoff: start at one second and double on each attempt (1s → 2s → 4s → 8s), capping at 3–5 tries. If a response includes a `Retry-After` header, honor it instead. Never retry a `4xx` other than `429` — the request is malformed and will fail the same way again.
## Asynchronous endpoints
Asynchronous partner models return 200 when they accept a job. A job that fails later does not return an HTTP error. The failure appears in the poll response. See [Errors and limits](/partner-models/overview#errors-and-limits) on the Partner models overview, and the model page for the error codes of each provider.
## Credit Safety
If an error occurs after credits have been deducted, they are automatically refunded. You will never lose credits due to a server-side failure.
# Image-to-image
Source: https://docs.nunchux.ai/nunchux-optimized/image-to-image
Edit an image with a text prompt with a Nunchux Optimized model. The edited image is in the response.
# Image-to-image
Send a source image and a text prompt that describes the edit. The response contains the edited image. The request is synchronous. There is no task to poll.
## Models
| Model | Model ID | Use it for |
| - | - | - |
| FLUX.2 Klein 9B Edit | `nunchux-flux.2-klein-9b-edit` | Edits that must keep the identity of the subject. |
| FLUX.2 Klein 4B Edit | `nunchux-flux.2-klein-4b-edit` | The fastest edits. High-volume edit passes. The output follows the input image. |
| Qwen Image Edit 2511 Lightning | `nunchux-qwen-image-edit-2511` | Targeted edits: swap elements, change the style, edit text in the image. The recommended editing model. |
See [Choosing a model](/nunchux-optimized/overview#choosing-a-model) for a comparison, and [Performance tiers](/performance-tiers) for the tiers that each model serves.
## Endpoint
```text theme={"system"}
POST https://api.nunchux.ai/v1/images/edits
```
## Request
Every request must include your API key and a JSON content type. See [Authentication](/authentication).
Your API key.
Must be `application/json`.
The model that edits the image.
Options: `nunchux-flux.2-klein-4b-edit`, `nunchux-flux.2-klein-9b-edit`, `nunchux-qwen-image-edit-2511`
The source image, as a data URI (for example `data:image/png;base64,...`) or as a public URL. Supported formats: PNG, JPEG, WebP. Maximum size: 10 MB. Each side must be at least 64 px. The aspect ratio must be no more extreme than 8:1 in either orientation. A remote URL must resolve to a public host. Nunchux does not follow redirects.
A text description of the edit. Say what to change, add or remove.
The balance of speed and cost. See [Performance tiers](/performance-tiers).
Options: `radical_speed`, `radical_value`
Output width in pixels. Send it together with `height`. Range: 1 to 8192.
Output height in pixels. Send it together with `width`. Range: 1 to 8192.
The number of images to generate. Only 1 is supported.
The format of the returned image.
Options: `url`, `b64_json`
In both formats, Nunchux keeps the generated image for 7 days, so that it appears in your request history. Then Nunchux deletes it.
The random seed. The same seed with the same parameters gives the same result.
Dimensions
* Send `width` and `height` together. A request with only one of them returns a 400.
* Every image-to-image model is priced per megapixel today and charges from the output area. A request with no dimensions returns a 400. The API does not use the source image size or a default size.
* Both values must be integers. The string `"1024"` is rejected, not converted.
* The source image is a reference, not a canvas. The model does not crop or stretch it to `width` and `height`. When the source ratio is different from the output ratio, the model re-frames the content.
* Smaller source images process faster. Larger output images take longer to generate.
A minimal request body:
```json theme={"system"}
{
"model": "nunchux-qwen-image-edit-2511",
"url": "https://example.com/example.jpg",
"prompt": "Transform into a watercolor painting",
"tier": "radical_speed",
"width": 1024,
"height": 1024,
"response_format": "b64_json"
}
```
## Response
The response is JSON. The `data` array contains the edited image, as base64 data or as a URL.
The Unix timestamp of the generation.
The generated items.
The URL of the image. Returned when `response_format` is `url`.
The image as base64 data. Returned when `response_format` is `b64_json`.
The model that edited the image.
```json Base64 format theme={"system"}
{
"created": 1234567890,
"data": [
{
"b64_json": "iVBORw0KGgoAAAANSUhEUgAABAAAAAQACAYAAAB/HSuDAAAACXBIWXMAAA7EAAAOxAGVKw4b..."
}
],
"model": "nunchux-qwen-image-edit-2511"
}
```
```json URL format theme={"system"}
{
"created": 1234567890,
"data": [
{
"url": "https://example.com/edited-image.png"
}
],
"model": "nunchux-qwen-image-edit-2511"
}
```
To save a base64 result from cURL, add this pipe to the command:
```bash theme={"system"}
# Replace output.jpg with the file name that you want.
| jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.jpg
```
## Example
```bash cURL theme={"system"}
# The media URL below is a placeholder. Point it at your own publicly
# reachable file before running this.
curl -X POST https://api.nunchux.ai/v1/images/edits \
-H "Content-Type: application/json" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-d '{
"model": "nunchux-qwen-image-edit-2511",
"url": "https://example.com/example.jpg",
"prompt": "Transform into a watercolor painting",
"tier": "radical_speed",
"width": 1024,
"height": 1024,
"response_format": "b64_json"
}' | jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.jpg
```
```python Python theme={"system"}
# The media URL below is a placeholder. Point it at your own publicly
# reachable file before running this.
import os
import base64
import requests
response = requests.post(
"https://api.nunchux.ai/v1/images/edits",
headers={
"Content-Type": "application/json",
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
},
json={
"model": "nunchux-qwen-image-edit-2511",
"url": "https://example.com/example.jpg",
"prompt": "Transform into a watercolor painting",
"tier": "radical_speed",
"width": 1024,
"height": 1024,
},
)
# Decode the base64 image and save it. Change output.jpg to any file name.
image_b64 = response.json()["data"][0]["b64_json"]
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(image_b64))
```
```javascript JavaScript theme={"system"}
// The media URL below is a placeholder. Point it at your own publicly
// reachable file before running this.
import { writeFile } from "node:fs/promises";
const response = await fetch("https://api.nunchux.ai/v1/images/edits", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-API-Key": process.env.NUNCHUX_API_KEY,
},
body: JSON.stringify({
model: "nunchux-qwen-image-edit-2511",
url: "https://example.com/example.jpg",
prompt: "Transform into a watercolor painting",
tier: "radical_speed",
width: 1024,
height: 1024,
}),
});
const data = await response.json();
// Decode the base64 image and save it. Change output.jpg to any file name.
await writeFile("output.jpg", Buffer.from(data.data[0].b64_json, "base64"));
```
```python Python (OpenAI) theme={"system"}
# Replace photo.jpg with the path to your source image.
import os
import requests
from openai import OpenAI
client = OpenAI(
base_url="https://api.nunchux.ai/v1",
api_key=os.environ["NUNCHUX_API_KEY"],
)
response = client.images.edit(
model="nunchux-qwen-image-edit-2511",
image=open("photo.jpg", "rb"),
prompt="Transform into a watercolor painting",
response_format="url",
extra_body={"width": 1024, "height": 1024},
)
# This tab returns a URL. Download it and save it. Change output.jpg to any file name.
image_url = response.data[0].url
with open("output.jpg", "wb") as f:
f.write(requests.get(image_url).content)
```
## Tips
* Say exactly what to change, add or remove. The models can follow instructions with more than one change.
* Use Qwen Image Edit 2511 when the edit must keep the identity of the subject, or when it changes text in the image.
* Use FLUX.2 Klein 4B Edit for fast, high-volume edit passes.
* Use `radical_speed` when response time is important. Use `radical_value` for large batches where cost per image is more important.
* Send smaller source images for faster processing.
## Errors and limits
| Status | Code | Cause |
| - | - | - |
| 400 | none | A required field is missing, the model ID is not valid, or a value is not valid. The body is a plain `detail` string. |
| 400 | `engine_error` | The model cannot render the request, or the source image is rejected: the URL cannot be fetched, the data URI cannot be decoded, or the image is outside the size and ratio limits. |
| 403 | `model_not_available` | Your account does not have access to this model. |
| 504 | `engine_timeout` | Nunchux could not fetch the source image in time. Make sure that the URL is reachable and fast, or send the image as a data URI. |
Refused requests
* Nunchux charges only after a successful generation. A refused request is never billed.
* Do not branch on a dimension-specific code. A refusal always carries the `engine_error` code.
For authentication, credit and rate-limit errors, see [Error codes](/errors). Every request counts toward the requests-per-minute cap of your plan. See [Rate limits](/rate-limits).
# Image-to-video
Source: https://docs.nunchux.ai/nunchux-optimized/image-to-video
Animate a still image into a video clip with LTX. The clip is in the response.
# Image-to-video
Send one image and a text prompt that describes the motion. The image becomes the first frame of the clip. The response contains the generated video clip. The request is synchronous. There is no task to poll. A clip is 1 to 8 seconds long, at 1080p or 720p, with an optional soundtrack.
## Models
| Model | Model ID | Use it for |
| - | - | - |
| LTX 2.5 | `nunchux-ltx-2.5-video` | New work. The newer release, with the same controls as LTX 2.3. |
| LTX 2.3 | `nunchux-ltx-2.3-video` | Continuity with clips that you already made with LTX 2.3. |
Both models also serve [Text-to-video](/nunchux-optimized/text-to-video). See [Choosing a model](/nunchux-optimized/overview#choosing-a-model).
## Endpoint
```text theme={"system"}
POST https://api.nunchux.ai/v1/video/animations
```
The path `/v1/videos/animations` is an alias. It gives the same result and the same billing.
## Request
Every request must include your API key and a JSON content type. See [Authentication](/authentication).
Your API key.
Must be `application/json`.
The model that generates the clip.
Options: `nunchux-ltx-2.5-video`, `nunchux-ltx-2.3-video`
A text description of the motion. Describe what moves, the camera movement and the light. Do not describe again what the image already shows.
The input image. Send one message with one `image_url` part. The model accepts one image.
Must be `user`.
The parts of the message.
`image_url` for the image part.
The image, as a data URI (for example `data:image/jpeg;base64,...`) or as a public `https` URL. Use a PNG or JPEG image. The clip keeps the framing of the image.
You can also send the image as a top-level `image_url` string.
Output width in pixels. Send it together with `height`. The size selects the tier, 1080p or 720p. 1080p costs more per second.
Options: `1920` (1080p), `1280` (720p)
Output height in pixels. Send it together with `width`.
Options: `1080` (1080p), `720` (720p)
The clip length in seconds. Range: 1 to 8. The model renders at 24 frames per second, so the length of the clip can differ from your value by a fraction of a second. The response gives the length of the clip.
Model options.
When `true`, the model generates a soundtrack with the picture. Describe the sound in the prompt.
The random seed. When you do not send a seed, the model selects one and returns it in the response. The same seed with the same parameters gives the same result.
The number of clips to generate. Only 1 is supported.
The format of the returned clip.
Options: `b64_json`, `url`
Unsupported fields
LTX does not accept `steps`, `guidance_scale`, `negative_prompt` or `num_frames`. A request that contains one of them returns a 400.
A minimal request body:
```json theme={"system"}
{
"model": "nunchux-ltx-2.3-video",
"prompt": "The scene comes to life with gentle motion",
"width": 1280,
"height": 720,
"duration": 6,
"extra": { "audio": false },
"response_format": "b64_json",
"messages": [
{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "https://example.com/first-frame.jpg" } }
]
}
]
}
```
## Response
The response is JSON. The `data` array contains one MP4 clip, as base64 data or as a URL.
The Unix timestamp of the generation.
The model that generated the clip.
Always `radical_speed`. This field does not show the LTX tier. The `width` and `height` of the clip show it.
The seed of the clip. Send it again to get the same result.
The generated clip.
The MP4 clip as base64 data. Returned when `response_format` is `b64_json`.
The URL of the MP4 clip. Returned when `response_format` is `url`.
The width of the clip in pixels.
The height of the clip in pixels.
The length of the clip in seconds.
The number of frames in the clip.
The frame rate of the clip.
```json theme={"system"}
{
"created": 1234567890,
"model": "nunchux-ltx-2.3-video",
"tier": "radical_speed",
"seed": 1873420651,
"data": [
{
"b64_json": "AAAAIGZ0eXBpc29tAAACAGlzb21pc28yYXZjMW1wNDEAAAAIZnJlZQ...",
"width": 1280,
"height": 720,
"duration": 6.0417,
"frames": 145,
"frame_rate": 24
}
]
}
```
A base64 clip is large. A 6 second 1080p clip can be more than 10 MB of base64 data. Use `response_format: "url"` when you do not need the data in the response.
## Example
This example reads a local JPEG file and sends it as a data URI. A clip takes some time to render. Set the timeout of your HTTP client to 120 seconds.
```bash cURL theme={"system"}
# Encode input.jpg as base64 on one line.
IMG_B64=$(base64 < input.jpg | tr -d '\n')
curl -X POST https://api.nunchux.ai/v1/video/animations \
--max-time 120 \
-H "Content-Type: application/json" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-d '{
"model": "nunchux-ltx-2.3-video",
"prompt": "The scene comes to life with gentle motion",
"width": 1280,
"height": 720,
"duration": 6,
"extra": {"audio": false},
"response_format": "b64_json",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,'"$IMG_B64"'"}}
]
}]
}' | jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.mp4
```
```python Python theme={"system"}
import os
import base64
import requests
with open("input.jpg", "rb") as f:
image_uri = "data:image/jpeg;base64," + base64.b64encode(f.read()).decode()
response = requests.post(
"https://api.nunchux.ai/v1/video/animations",
headers={
"Content-Type": "application/json",
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
},
json={
"model": "nunchux-ltx-2.3-video",
"prompt": "The scene comes to life with gentle motion",
"width": 1280,
"height": 720,
"duration": 6,
"extra": {"audio": False},
"response_format": "b64_json",
"messages": [
{
"role": "user",
"content": [{"type": "image_url", "image_url": {"url": image_uri}}],
}
],
},
timeout=120,
)
response.raise_for_status()
# Decode the base64 clip and save it. Change output.mp4 to any file name.
clip_b64 = response.json()["data"][0]["b64_json"]
with open("output.mp4", "wb") as f:
f.write(base64.b64decode(clip_b64))
```
```javascript JavaScript theme={"system"}
import { readFile, writeFile } from "node:fs/promises";
const imageUri =
"data:image/jpeg;base64," + (await readFile("input.jpg")).toString("base64");
const response = await fetch("https://api.nunchux.ai/v1/video/animations", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-API-Key": process.env.NUNCHUX_API_KEY,
},
body: JSON.stringify({
model: "nunchux-ltx-2.3-video",
prompt: "The scene comes to life with gentle motion",
width: 1280,
height: 720,
duration: 6,
extra: { audio: false },
response_format: "b64_json",
messages: [
{
role: "user",
content: [{ type: "image_url", image_url: { url: imageUri } }],
},
],
}),
signal: AbortSignal.timeout(120_000),
});
if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
const data = await response.json();
// Decode the base64 clip and save it. Change output.mp4 to any file name.
await writeFile("output.mp4", Buffer.from(data.data[0].b64_json, "base64"));
```
## Tips
* Write the motion, not the image. Say what moves, how the camera moves, and how the light changes.
* Use an image with the framing that you want in the clip. The clip keeps the framing of the image.
* When `extra.audio` is `true`, describe the sound in the prompt.
* Start at 1080p and 6 seconds. Use 720p for faster and lower-cost previews.
* Keep the `seed` when you like a result. Then change one parameter at a time.
## Errors and limits
| Status | Code | Cause |
| - | - | - |
| 400 | none | `model` is missing or not valid, or a size value is not valid. The body is a plain `detail` string. |
| 400 | `invalid_request` | The request sets `tier` to a value other than `radical_speed`. On LTX, `width` and `height` select the tier. Do not send the `tier` field. |
| 400 | `unsupported_response_format` | `response_format` is not `b64_json` or `url`. |
| 400 | `engine_error` | The model cannot render the request, for example a size, a duration, an image that it cannot read, more than one image, or a field that it does not accept. |
| 422 | `request_too_large` | The requested size is larger than 1920x1080. `details` gives the requested value and the maximum. |
| 502 | `output_too_large` | The clip is too large to return. You are not charged. Request a shorter clip or a smaller size. |
| 504 | none | The clip did not complete in 120 seconds. Send the request again. |
A clip is billed per second of output video, at the rate of its tier: 1080p costs more per second than 720p. See the [pricing page](https://nunchux.ai/pricing) for current rates. Nunchux charges only for a successful clip. A request that fails is never billed.
Every request counts toward the requests-per-minute cap of your plan. A request also holds one simultaneous-jobs slot until the response returns. See [Rate limits](/rate-limits). For authentication and credit errors, see [Error codes](/errors).
# OpenAI Compatibility
Source: https://docs.nunchux.ai/nunchux-optimized/openai-compatibility
Use the OpenAI Python or JavaScript SDK with the Nunchux API for image generation.
# OpenAI Compatibility
### Use the OpenAI SDK with the Nunchux API
## Overview
The Nunchux API is compatible with the [OpenAI Python SDK](https://github.com/openai/openai-python). Point the client at the [Nunchux base URL](/nunchux-optimized/overview#endpoints) and use your [Nunchux API key](/authentication). Your existing OpenAI code keeps working.
## Installation
Install the OpenAI Python SDK:
```bash theme={"system"}
pip install openai
```
## Client Setup
Point the OpenAI client at the Nunchux API by setting the `base_url` and `api_key`.
```python theme={"system"}
from openai import OpenAI
client = OpenAI(
base_url="https://api.nunchux.ai/v1",
api_key="YOUR_API_KEY",
)
```
The SDK sends your key as an `Authorization: Bearer` header, which the Nunchux API accepts on every endpoint. See [Authentication](/authentication) for details.
## Text-to-Image
Generate images from text descriptions using `client.images.generate()`. Nunchux-specific parameters like `tier` and `seed` are passed via `extra_body`.
```python theme={"system"}
import base64
from pathlib import Path
response = client.images.generate(
model="nunchux-flux.2-klein-4b",
prompt="A beautiful sunset over mountains",
size="1024x1024",
response_format="b64_json",
n=1,
extra_body={
"tier": "radical_speed",
"seed": 42,
},
)
image_b64 = response.data[0].b64_json
Path("output.png").write_bytes(base64.b64decode(image_b64))
```
## Image-to-Image
Edit and transform existing images using `client.images.edit()`. Pass the source image as a file object opened in binary mode.
```python theme={"system"}
response = client.images.edit(
model="nunchux-flux.2-klein-4b-edit",
image=open("photo.jpg", "rb"),
prompt="Transform into a watercolor painting",
size="1024x1024",
response_format="url",
extra_body={
"tier": "radical_speed",
},
)
# URL format requested above
print(response.data[0].url)
```
## Nunchux-Specific Parameters
The OpenAI SDK does not natively support all Nunchux parameters. Use the `extra_body` argument to pass Nunchux-specific fields.
Performance tier controlling the tradeoff between speed and cost. See [Performance Tiers](/performance-tiers) for details.
Options: `radical_speed`, `radical_value`
Random seed for reproducible generation. Use the same seed with identical
parameters for consistent results.
Output width in pixels. Send it together with `height`. The two fields take
precedence over the standard `size` argument. Models priced per megapixel
require `width` and `height`, or `size`. Range: 1-8192.
Output height in pixels. Send it together with `width`. Models priced per
megapixel require `width` and `height`, or `size`. Range: 1-8192.
Sizing
The examples on this page use `size` because it is the standard OpenAI argument and it works. Elsewhere in these docs you will see `width` and `height`, which are the preferred spelling. The API accepts either form and resolves both to the same dimensions. Pass them through `extra_body` if you want to match the rest of the platform.
Models priced per megapixel need dimensions. The [pricing page](https://nunchux.ai/pricing) shows each model's unit. Today, every image model except Ideogram 4 is priced per megapixel. The OpenAI SDK sends no `size` unless you set it. On a per-megapixel model, a request with neither `size` nor `width` and `height` returns a 400.
## Using Other SDK Format
Nothing about the API is Python-specific. The official [OpenAI Node SDK](https://github.com/openai/openai-node) works against the same base URL and the same key.
| | Python | Node | TypeScript |
| - | - | - | - |
| Base URL | `base_url` | `baseURL` | `baseURL` |
| API key | `api_key` | `apiKey` | `apiKey` |
| Methods | `client.images.generate()` | `client.images.generate()` | `client.images.generate()` |
| Nunchux params | `extra_body={...}` | inline in the request object | inline, assigned to a variable |
| Input image | `open("photo.jpg", "rb")` | a stream, or `toFile()` | a stream, or `toFile()` |
TypeScript
TypeScript refuses undeclared properties in an object literal, but applies that check only to literals. Assign the parameters to a variable first. This rule is TypeScript excess property checking, not an API limit. See the [TypeScript handbook](https://www.typescriptlang.org/docs/handbook/2/objects.html#excess-property-checks) for the full rule.
## Limitations
The API speaks the OpenAI image shape, but it is not OpenAI. Two differences matter when you port existing code:
1. `n` must be 1. One image per request. Any other value returns `400`. Send concurrent requests to generate a batch.
2. Errors do not use the OpenAI error shape. See [Error Handling](#error-handling) below.
## Error Handling
Two error shapes are in use. Most endpoints return `detail` and `request_id`. Migrated endpoints add a catalog envelope under `error`.
```jsonc theme={"system"}
// Rejecting n=4
{
"detail": "only n=1 is currently supported for image requests, got n=4",
"request_id": "9f3c1e2a4b5d4f6e8a7b9c0d1e2f3a4b"
}
```
On `/v1/images/generations`, a non-integer `n` also returns HTTP 400 with a plain `detail` string. The message text depends on the model. This example is from a per-megapixel model:
```jsonc theme={"system"}
// Rejecting n="not-a-number" on a per-megapixel model (HTTP 400)
{
"detail": "n must be a positive integer for per-megapixel billing, got 'not-a-number'",
"request_id": "9f3c1e2a4b5d4f6e8a7b9c0d1e2f3a4b"
}
```
Neither shape carries OpenAI's `type` or `param`, so `.type` and `.param` on the raised exception are always `None`. Branch on the HTTP status code, not on the error body. `request_id` is present on both shapes. Quote it when you contact support. It is how we find your request. See [Errors](/errors) for the full code catalog.
## Supported models
Every Nunchux Optimized image model works with the OpenAI SDK. See the [models table](/nunchux-optimized/overview#models) for the model IDs, and [Text-to-image](/nunchux-optimized/text-to-image) and [Image-to-image](/nunchux-optimized/image-to-image) for every request field.
# Overview
Source: https://docs.nunchux.ai/nunchux-optimized/overview
Models optimized by Nunchux, served through Nunchux's own endpoints. One request returns the output.
# Nunchux Optimized models
Nunchux Optimized models run on Nunchux's proprietary Model Optimizer and Inference Engine. You call them through Nunchux's own endpoints. Every request is synchronous: you send one POST, and the image or the video clip is in the response. There is no task to poll.
All Nunchux Optimized models accept the same request shape. The `model` field selects the model. On the FLUX, Qwen and HiDream O1 models, the `tier` field selects the tier: Radical Speed or Radical Value. Ideogram 4 has three other tiers, Turbo, Balanced and Quality. On Ideogram 4, the `steps` field selects the tier. The LTX models have two other tiers, 720p (1280x720) and 1080p (1920x1080). On LTX, the output size (`width` and `height`) selects the tier.
## Models
### Image
| Model | Model ID | Task | Tiers |
| - | - | - | - |
| FLUX.2 Klein 4B | `nunchux-flux.2-klein-4b` | Text-to-image | Radical Speed, Radical Value |
| FLUX.2 Klein 9B | `nunchux-flux.2-klein-9b` | Text-to-image | Radical Speed, Radical Value |
| FLUX.1 Schnell | `nunchux-flux.1-schnell` | Text-to-image | Radical Speed, Radical Value |
| Qwen Image 2512 Lightning | `nunchux-qwen-image-2512` | Text-to-image | Radical Speed, Radical Value |
| FLUX.2 Klein 4B Edit | `nunchux-flux.2-klein-4b-edit` | Image-to-image | Radical Speed, Radical Value |
| FLUX.2 Klein 9B Edit | `nunchux-flux.2-klein-9b-edit` | Image-to-image | Radical Speed, Radical Value |
| Qwen Image Edit 2511 Lightning | `nunchux-qwen-image-edit-2511` | Image-to-image | Radical Speed, Radical Value |
| HiDream O1 | `nunchux-hidream-o1-image` | Text-to-image | Radical Speed, Radical Value |
| Ideogram 4 | `nunchux-ideogram-4` | Text-to-image | Turbo, Balanced, Quality |
### Video
| Model | Model ID | Tasks | Tiers |
| - | - | - | - |
| LTX 2.5 | `nunchux-ltx-2.5-video` | Text-to-video, image-to-video | 720p, 1080p |
| LTX 2.3 | `nunchux-ltx-2.3-video` | Text-to-video, image-to-video | 720p, 1080p |
Text-to-image models go to the [Text-to-image](/nunchux-optimized/text-to-image) endpoint. Image-to-image models go to the [Image-to-image](/nunchux-optimized/image-to-image) endpoint. The LTX models go to the [Text-to-video](/nunchux-optimized/text-to-video) and [Image-to-video](/nunchux-optimized/image-to-video) endpoints.
## Choosing a model
FLUX.2 Klein 9B or 4B. The 9B models give higher quality, stronger text rendering, and better results on complex prompts. The 4B models are faster. Use 9B for production output. Use 4B when speed matters more than detail.
FLUX or Qwen. FLUX Klein is faster. Qwen Image renders text best, in English and in Chinese, and suits typography and layout work. Qwen Image Edit 2511 is the recommended editing model when the edit must keep the identity of the subject.
HiDream O1. It makes large, photographic images, up to four megapixels, with fine texture in skin, fabric and surfaces. It takes more time per image than FLUX and Qwen. It has a guidance control.
Ideogram 4. It renders the words in a prompt legibly, for posters, labels and covers. It takes more time per image than FLUX and Qwen. For Chinese text, use Qwen Image.
FLUX.1 Schnell. The earlier FLUX generation. It stays available for work that already depends on it.
LTX 2.5 or 2.3. LTX 2.5 is the newer release, with the same controls as LTX 2.3. Use LTX 2.3 only for continuity with clips you already made.
## Endpoints
The base URL is `https://api.nunchux.ai`. Each page describes one endpoint.
| Page | Method and path |
| - | - |
| [Text-to-image](/nunchux-optimized/text-to-image) | `POST /v1/images/generations` |
| [Image-to-image](/nunchux-optimized/image-to-image) | `POST /v1/images/edits` |
| [Text-to-video](/nunchux-optimized/text-to-video) | `POST /v1/video/generations` |
| [Image-to-video](/nunchux-optimized/image-to-video) | `POST /v1/video/animations` |
The image endpoints also work with the OpenAI SDK. See [OpenAI compatibility](/nunchux-optimized/openai-compatibility).
## How a request works
1. Send a POST with a JSON body to the endpoint. Put your API key in the `X-API-Key` header. See [Authentication](/authentication).
2. Set `model` to a model ID from the tables above. Set `prompt` to the text that describes the output. Image-to-image and image-to-video requests also carry the input image. Then select the tier. Each model uses a different field:
* FLUX, Qwen and HiDream O1: set `tier` to `radical_speed` or `radical_value`.
* Ideogram 4: set the inference steps with `steps`. The step count selects the tier: 12 for Turbo, 20 for Balanced, or 48 for Quality.
* LTX: set the size with `width` and `height`. The size selects 720p or 1080p.
Do not send `tier` to Ideogram 4 or LTX.
3. Read the output from the response. Images return as base64 data or as a URL, selected by `response_format`. The video endpoint pages describe the video response.
Each endpoint page lists every field and shows a full example.
## Billing
Each image request is billed on its tier. Radical Speed gives the lowest latency. Radical Value gives the lowest cost per output. Ideogram 4 bills a fixed price per image, at the rate of its tier: Turbo costs the least and Quality the most. Each LTX request is billed per second of output video, at the rate of its tier: 1080p costs more per second than 720p. Nunchux charges only for successful requests. When a request fails after the credits are deducted, Nunchux refunds them.
See [Performance tiers](/performance-tiers) for the tier comparison and [Credits & pricing](/credits-pricing) for how credits work.
## Errors and limits
A failed request returns an HTTP error status with a JSON body that names the error. See [Error codes](/errors) for the codes and the retry guidance.
Every request counts toward the requests-per-minute cap of your plan. See [Rate limits](/rate-limits).
# Text-to-image
Source: https://docs.nunchux.ai/nunchux-optimized/text-to-image
Generate an image from a text prompt with a Nunchux Optimized model. The image is in the response.
# Text-to-image
Send a text prompt, and the response contains the generated image. The request is synchronous. There is no task to poll.
## Models
| Model | Model ID | Use it for |
| - | - | - |
| FLUX.2 Klein 9B | `nunchux-flux.2-klein-9b` | Production output. Strong prompt adherence and clean text. |
| FLUX.2 Klein 4B | `nunchux-flux.2-klein-4b` | The fastest output. Fast iteration and batch jobs. |
| FLUX.1 Schnell | `nunchux-flux.1-schnell` | The earlier FLUX generation, for work that already depends on it. |
| Qwen Image 2512 Lightning | `nunchux-qwen-image-2512` | The best text rendering, in English and Chinese. Typography and layout. |
| HiDream O1 | `nunchux-hidream-o1-image` | Large, photographic images up to four megapixels. Slower than FLUX and Qwen. |
| Ideogram 4 | `nunchux-ideogram-4` | Legible words in the image: posters, labels and covers. Slower than FLUX and Qwen. |
See [Choosing a model](/nunchux-optimized/overview#choosing-a-model) for a comparison, and [Performance tiers](/performance-tiers) for the tiers that each model serves.
## Endpoint
```text theme={"system"}
POST https://api.nunchux.ai/v1/images/generations
```
## Request
Every request must include your API key and a JSON content type. See [Authentication](/authentication).
Your API key.
Must be `application/json`.
The model that generates the image.
Options: `nunchux-flux.2-klein-4b`, `nunchux-flux.2-klein-9b`, `nunchux-flux.1-schnell`, `nunchux-qwen-image-2512`, `nunchux-hidream-o1-image`, `nunchux-ideogram-4`
A text description of the image. Specific, descriptive prompts give better results.
The balance of speed and cost. See [Performance tiers](/performance-tiers).
Options: `radical_speed`, `radical_value`
Do not send this field to Ideogram 4. On Ideogram 4, `steps` selects the tier.
Output width in pixels. Send it together with `height`. Required on models priced per megapixel. Today, that is every model except Ideogram 4, which uses 1024x1024 when you send no dimensions. Range: 1 to 8192. On Ideogram 4: 1024 to 2048, in steps of 16.
Output height in pixels. Send it together with `width`. Required on models priced per megapixel. Today, that is every model except Ideogram 4, which uses 1024x1024 when you send no dimensions. Range: 1 to 8192. On Ideogram 4: 1024 to 2048, in steps of 16.
Ideogram 4 and HiDream O1 only. The number of denoising steps. Range: 1 to 50. Default: 48 on Ideogram 4, 50 on HiDream O1. More steps give more detail, but take longer.
On Ideogram 4, the step count selects the tier. Send 12 for Turbo, 20 for Balanced, or 48 for Quality. Fewer steps are faster and cost less, but give less detail. A request without `steps` runs 48 steps and is billed as Quality.
HiDream O1 only. How closely the image follows the prompt. Range: 0.0 to 20.0. Higher values follow the prompt more closely. Lower values give the model more freedom.
Ideogram 4 only. When `standard`, a language model rewrites the prompt into a detailed scene description before the image renders. The response returns the rewritten prompt as `revised_prompt`. If the rewrite fails, the image renders from your prompt, and the response has no `revised_prompt`.
Options: `off`, `standard`
The number of images to generate. Only 1 is supported.
The format of the returned image.
Options: `url`, `b64_json`
In both formats, Nunchux keeps the generated image for 7 days, so that it appears in your request history. Then Nunchux deletes it.
The random seed. The same seed with the same parameters gives the same result.
Dimensions
* Send `width` and `height` together. A request with only one of them returns a 400.
* Models priced per megapixel charge from the output area, so they require `width` and `height`. A request with no dimensions returns a 400. The API does not choose a default size for these models. The [pricing page](https://nunchux.ai/pricing) shows each model's unit. Today, every text-to-image model except Ideogram 4 is priced per megapixel.
* Both values must be integers. The string `"1024"` is rejected, not converted.
* Larger images take longer to generate.
* HiDream O1 renders these sizes: 1024x1024, 2048x2048, 2304x1728, 1728x2304, 2560x1440, 1440x2560, 2496x1664, 1664x2496, 3104x1312, 1312x3104, 2304x1792, 1792x2304.
A minimal request body:
```json theme={"system"}
{
"model": "nunchux-qwen-image-2512",
"tier": "radical_speed",
"prompt": "A beautiful sunset over mountains",
"width": 1024,
"height": 1024,
"response_format": "b64_json",
"seed": 42
}
```
## Response
The response is JSON. The `data` array contains the image, as base64 data or as a URL.
The Unix timestamp of the generation.
The generated items.
The URL of the image. Returned when `response_format` is `url`.
The image as base64 data. Returned when `response_format` is `b64_json`.
Ideogram 4 only. The rewritten prompt that the model rendered. Returned when `prompt_expansion` is `standard` and the rewrite succeeds. To make the same image again, send it as `prompt` with `prompt_expansion` set to `off`, the same `seed`, the same `steps`, and the same size.
The model that generated the image.
```json Base64 format theme={"system"}
{
"created": 1234567890,
"data": [
{
"b64_json": "iVBORw0KGgoAAAANSUhEUgAABAAAAAQACAYAAAB/HSuDAAAACXBIWXMAAA7EAAAOxAGVKw4b..."
}
],
"model": "nunchux-qwen-image-2512"
}
```
```json URL format theme={"system"}
{
"created": 1234567890,
"data": [
{
"url": "https://example.com/generated-image.png"
}
],
"model": "nunchux-qwen-image-2512"
}
```
To save a base64 result from cURL, add this pipe to the command:
```bash theme={"system"}
# Replace output.jpg with the file name that you want.
| jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.jpg
```
## Example
```bash cURL theme={"system"}
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "Content-Type: application/json" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-d '{
"model": "nunchux-flux.2-klein-9b",
"tier": "radical_speed",
"prompt": "cozy coffee shop interior with vintage furniture",
"width": 1024,
"height": 1024,
"response_format": "b64_json",
"seed": 42
}' | jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.jpg
```
```python Python theme={"system"}
import os
import base64
import requests
response = requests.post(
"https://api.nunchux.ai/v1/images/generations",
headers={
"Content-Type": "application/json",
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
},
json={
"model": "nunchux-flux.2-klein-9b",
"tier": "radical_speed",
"prompt": "cozy coffee shop interior with vintage furniture",
"width": 1024,
"height": 1024,
"response_format": "b64_json",
"seed": 42,
},
)
# Decode the base64 image and save it. Change output.jpg to any file name.
image_b64 = response.json()["data"][0]["b64_json"]
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(image_b64))
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const response = await fetch("https://api.nunchux.ai/v1/images/generations", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-API-Key": process.env.NUNCHUX_API_KEY,
},
body: JSON.stringify({
model: "nunchux-flux.2-klein-9b",
tier: "radical_speed",
prompt: "cozy coffee shop interior with vintage furniture",
width: 1024,
height: 1024,
response_format: "b64_json",
seed: 42,
}),
});
const data = await response.json();
// Decode the base64 image and save it. Change output.jpg to any file name.
await writeFile("output.jpg", Buffer.from(data.data[0].b64_json, "base64"));
```
```python Python (OpenAI) theme={"system"}
import os
import base64
from openai import OpenAI
client = OpenAI(
base_url="https://api.nunchux.ai/v1",
api_key=os.environ["NUNCHUX_API_KEY"],
)
# The OpenAI SDK has no width/height parameter. Send them in `extra_body`,
# which the SDK merges into the request body.
response = client.images.generate(
model="nunchux-flux.2-klein-9b",
prompt="cozy coffee shop interior with vintage furniture",
n=1,
response_format="b64_json",
extra_body={"width": 1024, "height": 1024},
)
# Decode the base64 image and save it. Change output.jpg to any file name.
image_b64 = response.data[0].b64_json
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(image_b64))
```
## Tips
* Write detailed prompts. Include the style, lighting, composition, colors, materials and atmosphere.
* Use FLUX.2 Klein 9B for production output. Use FLUX.2 Klein 4B when speed is more important than detail.
* Use Qwen Image when the image must contain text, and for Chinese text. Use Ideogram 4 when the words must be legible on a poster, a label or a cover. Put the exact words in quotes or capitals, and say where they go.
* Use HiDream O1 for large, photographic images. Describe the light and the materials. A larger size costs more, because the price is per megapixel.
* On Ideogram 4, draft at Turbo (12 steps). Use Quality (48 steps) for the final image.
* On FLUX, Qwen and HiDream O1, use `radical_speed` when response time is important. Use `radical_value` for large batches where cost per image is more important.
* Keep the `seed` when you like a result. Then change one parameter at a time.
## Errors and limits
| Status | Code | Cause |
| - | - | - |
| 400 | none | A required field is missing, the model ID is not valid, or a value is not valid. The body is a plain `detail` string. |
| 400 | `engine_error` | The model cannot render the request, for example the dimensions. |
| 403 | `model_not_available` | Your account does not have access to this model. |
Refused requests
* Nunchux charges only after a successful generation. A refused request is never billed.
* Do not branch on a dimension-specific code. A refusal always carries the `engine_error` code.
For authentication, credit and rate-limit errors, see [Error codes](/errors). Every request counts toward the requests-per-minute cap of your plan. See [Rate limits](/rate-limits).
# Text-to-video
Source: https://docs.nunchux.ai/nunchux-optimized/text-to-video
Generate a video clip from a text prompt with LTX. The clip is in the response.
# Text-to-video
Send a text prompt, and the response contains the generated video clip. The request is synchronous. There is no task to poll. A clip is 1 to 8 seconds long, at 1080p or 720p, with an optional soundtrack.
## Models
| Model | Model ID | Use it for |
| - | - | - |
| LTX 2.5 | `nunchux-ltx-2.5-video` | New work. The newer release, with the same controls as LTX 2.3. |
| LTX 2.3 | `nunchux-ltx-2.3-video` | Continuity with clips that you already made with LTX 2.3. |
Both models also serve [Image-to-video](/nunchux-optimized/image-to-video). See [Choosing a model](/nunchux-optimized/overview#choosing-a-model).
## Endpoint
```text theme={"system"}
POST https://api.nunchux.ai/v1/video/generations
```
The path `/v1/videos/generations` is an alias. It gives the same result and the same billing.
## Request
Every request must include your API key and a JSON content type. See [Authentication](/authentication).
Your API key.
Must be `application/json`.
The model that generates the clip.
Options: `nunchux-ltx-2.5-video`, `nunchux-ltx-2.3-video`
A text description of the clip. Describe the shot: the subject, the camera movement, the light, and what changes during the clip.
Output width in pixels. Send it together with `height`. The size selects the tier, 1080p or 720p. 1080p costs more per second.
Options: `1920` (1080p), `1280` (720p)
Output height in pixels. Send it together with `width`.
Options: `1080` (1080p), `720` (720p)
The clip length in seconds. Range: 1 to 8. The model renders at 24 frames per second, so the length of the clip can differ from your value by a fraction of a second. The response gives the length of the clip.
Model options.
When `true`, the model generates a soundtrack with the picture. Describe the sound in the prompt.
The random seed. When you do not send a seed, the model selects one and returns it in the response. The same seed with the same parameters gives the same result.
The number of clips to generate. Only 1 is supported.
The format of the returned clip.
Options: `b64_json`, `url`
Unsupported fields
LTX does not accept `steps`, `guidance_scale`, `negative_prompt` or `num_frames`. A request that contains one of them returns a 400.
A minimal request body:
```json theme={"system"}
{
"model": "nunchux-ltx-2.3-video",
"prompt": "A golden retriever running on a beach at sunset, cinematic",
"width": 1280,
"height": 720,
"duration": 6,
"extra": { "audio": false },
"response_format": "b64_json"
}
```
## Response
The response is JSON. The `data` array contains one MP4 clip, as base64 data or as a URL.
The Unix timestamp of the generation.
The model that generated the clip.
Always `radical_speed`. This field does not show the LTX tier. The `width` and `height` of the clip show it.
The seed of the clip. Send it again to get the same result.
The generated clip.
The MP4 clip as base64 data. Returned when `response_format` is `b64_json`.
The URL of the MP4 clip. Returned when `response_format` is `url`.
The width of the clip in pixels.
The height of the clip in pixels.
The length of the clip in seconds.
The number of frames in the clip.
The frame rate of the clip.
```json theme={"system"}
{
"created": 1234567890,
"model": "nunchux-ltx-2.3-video",
"tier": "radical_speed",
"seed": 1873420651,
"data": [
{
"b64_json": "AAAAIGZ0eXBpc29tAAACAGlzb21pc28yYXZjMW1wNDEAAAAIZnJlZQ...",
"width": 1280,
"height": 720,
"duration": 6.0417,
"frames": 145,
"frame_rate": 24
}
]
}
```
A base64 clip is large. A 6 second 1080p clip can be more than 10 MB of base64 data. Use `response_format: "url"` when you do not need the data in the response.
## Example
A clip takes some time to render. Set the timeout of your HTTP client to 120 seconds.
```bash cURL theme={"system"}
curl -X POST https://api.nunchux.ai/v1/video/generations \
--max-time 120 \
-H "Content-Type: application/json" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-d '{
"model": "nunchux-ltx-2.3-video",
"prompt": "A golden retriever running on a beach at sunset, cinematic",
"width": 1280,
"height": 720,
"duration": 6,
"extra": {"audio": false},
"response_format": "b64_json"
}' | jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.mp4
```
```python Python theme={"system"}
import os
import base64
import requests
response = requests.post(
"https://api.nunchux.ai/v1/video/generations",
headers={
"Content-Type": "application/json",
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
},
json={
"model": "nunchux-ltx-2.3-video",
"prompt": "A golden retriever running on a beach at sunset, cinematic",
"width": 1280,
"height": 720,
"duration": 6,
"extra": {"audio": False},
"response_format": "b64_json",
},
timeout=120,
)
response.raise_for_status()
# Decode the base64 clip and save it. Change output.mp4 to any file name.
clip_b64 = response.json()["data"][0]["b64_json"]
with open("output.mp4", "wb") as f:
f.write(base64.b64decode(clip_b64))
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const response = await fetch("https://api.nunchux.ai/v1/video/generations", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-API-Key": process.env.NUNCHUX_API_KEY,
},
body: JSON.stringify({
model: "nunchux-ltx-2.3-video",
prompt: "A golden retriever running on a beach at sunset, cinematic",
width: 1280,
height: 720,
duration: 6,
extra: { audio: false },
response_format: "b64_json",
}),
signal: AbortSignal.timeout(120_000),
});
if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
const data = await response.json();
// Decode the base64 clip and save it. Change output.mp4 to any file name.
await writeFile("output.mp4", Buffer.from(data.data[0].b64_json, "base64"));
```
## Tips
* Write the shot, not only the subject. Include the camera movement, the light, and what changes during the clip.
* When `extra.audio` is `true`, describe the sound in the prompt. For example: rain on a window, a crowd, or one piano line.
* Start at 1080p and 6 seconds. Use 720p for faster and lower-cost previews.
* Keep the `seed` when you like a result. Then change one parameter at a time.
## Errors and limits
| Status | Code | Cause |
| - | - | - |
| 400 | none | `model` is missing or not valid, or a size value is not valid. The body is a plain `detail` string. |
| 400 | `invalid_request` | The request sets `tier` to a value other than `radical_speed`. On LTX, `width` and `height` select the tier. Do not send the `tier` field. |
| 400 | `unsupported_response_format` | `response_format` is not `b64_json` or `url`. |
| 400 | `engine_error` | The model cannot render the request, for example a size, a duration or a field that it does not accept. |
| 422 | `request_too_large` | The requested size is larger than 1920x1080. `details` gives the requested value and the maximum. |
| 502 | `output_too_large` | The clip is too large to return. You are not charged. Request a shorter clip or a smaller size. |
| 504 | none | The clip did not complete in 120 seconds. Send the request again. |
A clip is billed per second of output video, at the rate of its tier: 1080p costs more per second than 720p. See the [pricing page](https://nunchux.ai/pricing) for current rates. Nunchux charges only for a successful clip. A request that fails is never billed.
Every request counts toward the requests-per-minute cap of your plan. A request also holds one simultaneous-jobs slot until the response returns. See [Rate limits](/rate-limits). For authentication and credit errors, see [Error codes](/errors).
# Overview
Source: https://docs.nunchux.ai/overview
Generate images and videos with one API key and one credit balance.
# Nunchux docs
Nunchux gives you one API for image and video generation. One API key and one credit balance cover every model.
The models are in two groups. The groups use different request shapes.
Image and video models that run on Nunchux's proprietary Model Optimizer and Inference Engine. One request returns the output. Performance tiers set the balance of speed and cost.
Third-party models from Google, Kling AI, Alibaba Cloud, ByteDance, MiniMax and HeyGen. Each provider keeps its own routes and request shapes. Video jobs are asynchronous.
## Start here
1. Make your first request. See [Quickstart](/quickstart).
2. Learn how to send your API key. See [Authentication](/authentication).
3. Find a model. See the two overviews above.
4. Connect Claude Code, Cursor and other tools to these docs. See [Coding agents](/coding-agents).
# HappyHorse
Source: https://docs.nunchux.ai/partner-models/happyhorse
Alibaba Cloud's HappyHorse video models for text-to-video, image-to-video and reference-to-video, with a soundtrack on every clip.
# HappyHorse
HappyHorse is Alibaba Cloud's second video model line, beside Wan. Each version generates a clip from a text prompt, from a still image, or from a set of reference images, and always scores the clip itself. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors Alibaba Cloud's own API, so the request and response shapes are the provider's.
Best for
* Sound without a switch. Every clip arrives with a generated soundtrack, so there is no audio decision to make.
* Reference-driven scenes. Up to nine reference images carry a subject into new footage.
* Draft passes. HappyHorse 1.1 sells a 480P tier that 1.0 does not.
* Steady camera work. Tracking shots, slow push-ins and pans, with the motion described in the prompt.
## Models
| Model | Model ID | Tasks | Resolution | Duration | Reference images |
| - | - | - | - | - | - |
| HappyHorse 1.1 | `happyhorse-1.1-t2v`, `happyhorse-1.1-i2v`, `happyhorse-1.1-r2v` | t2v, i2v, r2v | 480P, 720P, 1080P | 3 to 15 s | Up to 9 |
| HappyHorse 1.0 | `happyhorse-1.0-t2v`, `happyhorse-1.0-i2v`, `happyhorse-1.0-r2v` | t2v, i2v, r2v | 720P, 1080P | 3 to 15 s | Up to 9 |
Durations are whole seconds. Each version has one model ID per task. Both versions take the same request shape. Only the `model` value changes.
HappyHorse 1.1 is the current line, and the only one with a 480P tier. Use it for drafts and for the final render.
HappyHorse 1.0 is the earlier line at 720P and 1080P.
Image-to-video takes a first frame only. There is no end frame. Reference-to-video takes images only. There is no container for a reference clip or a reference track.
## Endpoints
| Step | Route |
| - | - |
| Submit | `POST /v1/alibaba/services/aigc/video-generation/video-synthesis` |
| Poll | `GET /v1/alibaba/tasks/{task_id}` |
HappyHorse shares the two routes and the body shape of [Wan](/partner-models/wan). The tier goes in `parameters.resolution`, and frames and references go in `input.media[]`.
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. The examples also send a `User-Agent` header that names your application.
The model ID from the table above. The ID names both the version and the task.
The prompt and the media items.
The text prompt. All tasks. Describe the action and the camera movement. On reference-to-video, name a reference by its position in `media[]`: `[Image 1]`, `[Image 2]` and so on, with the brackets.
The media items for image-to-video and reference-to-video. Omit it for text-to-video. A first frame and a reference set cannot appear in the same body.
What the item is.
* `first_frame`: the image the clip starts on. Image-to-video. One item.
* `reference_image`: a reference image. Reference-to-video. 1 to 9 items.
The URL of the image. The provider fetches it, so the URL must be reachable from the internet.
The generation controls. All tasks read all three fields.
The output tier. `480P`, `720P` or `1080P` on HappyHorse 1.1. `720P` or `1080P` on HappyHorse 1.0. The default is the most expensive tier, so name the one you want.
The clip length in whole seconds, 3 to 15. You are billed per second of output.
An omitted flag yields a clip with the text Happy Horse in the lower-right corner. Send `false` for a clean frame.
There is no audio parameter. The model scores every clip, and no field turns that off.
## Response
### Submit
A 200 status means that the provider accepted the task. It does not mean that the clip is ready.
The task ID. Poll `GET /v1/alibaba/tasks/{task_id}` with it.
`PENDING` on a new task.
### Poll
`PENDING` or `RUNNING` while the task runs. `SUCCEEDED`, `FAILED` or `CANCELED` when it ends. Poll until you read one of the three terminal values.
The URL of the clip. Present only when `task_status` is `SUCCEEDED`. The task ID and the URL expire after 24 hours. Download the clip as soon as the task succeeds.
The failure code. Present when the task failed.
The failure message. Read it before you resubmit.
## Example
Submit, poll, download. Swap the `model` value for the version and task you want.
```bash cURL theme={"system"}
BASE=https://api.nunchux.ai
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/alibaba/services/aigc/video-generation/video-synthesis" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "happyhorse-1.1-t2v",
"input": { "prompt": "a fox trotting through a snowy forest at dawn" },
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
TASK=$(jq -er '.output.task_id' <<<"$SUBMIT") || { echo "no task_id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until the status is terminal
DELAY=5
for _ in $(seq 1 90); do
RESP=$(curl -sS "$BASE/v1/alibaba/tasks/$TASK" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.output.task_status // ""' <<<"$RESP")
case "$STATUS" in SUCCEEDED|FAILED|CANCELED) break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
# 3. Check for failure
[ "$STATUS" = "SUCCEEDED" ] || { echo "task $STATUS: $(jq -r '.output.message // "no message"' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o happyhorse.mp4 "$(jq -r '.output.video_url' <<<"$RESP")"
```
```python Python theme={"system"}
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
f"{BASE}/v1/alibaba/services/aigc/video-generation/video-synthesis",
headers=HEADERS,
json={
"model": "happyhorse-1.1-t2v",
"input": {"prompt": "a fox trotting through a snowy forest at dawn"},
"parameters": {"resolution": "720P", "duration": 5, "watermark": False},
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
task = submit.json()["output"]["task_id"]
# 2. Poll until the status is terminal
delay = 5
for _ in range(90):
data = requests.get(f"{BASE}/v1/alibaba/tasks/{task}", headers=HEADERS).json()
status = data.get("output", {}).get("task_status", "")
if status in ("SUCCEEDED", "FAILED", "CANCELED"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("HappyHorse task did not reach a terminal status")
# 3. Check for failure
if status != "SUCCEEDED":
raise RuntimeError(f"HappyHorse task {status}: {data['output'].get('message')}")
# 4. Download
with requests.get(data["output"]["video_url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("happyhorse.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
To animate a still, add a `first_frame` item to `input.media[]` and send an `-i2v` model ID. To carry a subject into a new scene, add `reference_image` items and send an `-r2v` model ID. A first frame and a reference set are separate tasks. The bodies below replace the submit body in the example.
```json Image-to-video theme={"system"}
{
"model": "happyhorse-1.1-i2v",
"input": {
"prompt": "the fox turns and trots toward the camera",
"media": [
{ "type": "first_frame", "url": "https://example.com/frame.jpg" }
]
},
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}
```
```json Reference-to-video theme={"system"}
{
"model": "happyhorse-1.1-r2v",
"input": {
"prompt": "the person from [Image 1] walks through the doorway from [Image 2], slow tracking shot",
"media": [
{ "type": "reference_image", "url": "https://example.com/ref-1.jpg" },
{ "type": "reference_image", "url": "https://example.com/ref-2.jpg" }
]
},
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}
```
## Tips
* Say what moves and how the camera follows it. The prompt drives both.
* Send `watermark: false` on every request unless you want the mark. The vendor's default adds it.
* Send `resolution` and `duration` on every request. The defaults are 1080P and 5 seconds, and 1080P is the most expensive tier.
* Draft at 480P on HappyHorse 1.1, then re-run the prompt at the tier and length you need. You are billed per second of output, and 1080P costs more per second than 720P.
* Give a reference set images that agree on the subject. Images that disagree pull the result in different directions.
* Name the references in the prompt when the set holds different subjects, so each image has a stated role in the shot.
* Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.
## Errors and limits
A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with `task_status` set to `FAILED` or `CANCELED`, and `output.message` explains why. Read the message before you resubmit.
A failed submit returns an HTTP error status. The body carries `code`, `message` and `request_id`. See [Error codes](/errors).
A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned a `task_id`, poll it. If it returned nothing, check your credits before you try again.
Alibaba Cloud applies these limits to the images you attach:
* A first frame on image-to-video: at least 300 pixels on each side, an aspect ratio between 1:2.5 and 2.5:1, JPEG, JPG, PNG or WEBP, up to 20 MB.
* A reference image on reference-to-video: at least 400 pixels on the short side, JPEG, JPG, PNG or WEBP, up to 20 MB, one to nine images per request.
Every task counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits). Rates are per second of output and depend on the tier. See [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
# HeyGen Avatar
Source: https://docs.nunchux.ai/partner-models/heygen
Turn a photo or a HeyGen avatar into a lip-synced talking-avatar video with HeyGen Avatar 4 Photo and Avatar 5 Digital, driven by audio or a script in a chosen voice.
# HeyGen Avatar
HeyGen Avatar turns a photo, or a HeyGen avatar, into a lip-synced talking-avatar video. You supply the face and a speech source: an audio file that you host, or a script that HeyGen reads in a voice that you choose. The endpoint is asynchronous. Submit the job, poll until the status is terminal, then download the video. The request and response shapes mirror HeyGen's API.
## Models
HeyGen offers two engines. They take the same speech sources, share every other field, and have the same price. They differ only in where the face comes from.
| Engine | Id | Face input | Speech input | Output |
| - | - | - | - | - |
| Avatar 4 Photo | `avatar_iv` | A photo that you supply, by public URL or inline as base64 | `audio_url`, or `script` with `voice_id` | Lip-synced MP4 video of the photo |
| Avatar 5 Digital | `avatar_v` | A ready-made avatar from HeyGen's avatar catalog (`avatar_id`) | `audio_url`, or `script` with `voice_id` | Lip-synced MP4 video of the avatar |
The endpoint takes no `model` field, so the ids above are not values that you send. The request body selects the engine. `"type": "image"` with an `image` object selects Avatar 4 Photo. `"type": "avatar"` with an `avatar_id` and `"engine": {"type": "avatar_v"}` selects Avatar 5 Digital.
Use Avatar 4 Photo when you have a photo of the face that you need. Use Avatar 5 Digital when you want a ready-made look from HeyGen's avatar catalog. Neither engine has a duration control. The clip is as long as the audio or the script. Common uses are presenter clips, narrated explainers, and re-voiced versions of the same face.
## Endpoints
All endpoints share the base URL `https://api.nunchux.ai`.
| Purpose | Method | Path |
| - | - | - |
| Submit a video | POST | `/v1/heygen/v3/videos` |
| Poll a job | GET | `/v1/heygen/v3/videos/{video_id}` |
| List voices | GET | `/v1/heygen/v3/voices` |
## Request
Send your API key in the `X-API-Key` header (see [Authentication](/authentication)) and set `Content-Type: application/json`. Also send a `User-Agent` header that identifies your application, for example `YourApp/1.0`. A request that keeps the default `User-Agent` of your HTTP client can be rejected.
Send one face and one speech source in the body. The face fields select the engine. The speech fields are the same on both engines.
Selects the engine. `image` selects Avatar 4 Photo and animates the photo in `image`. `avatar` selects Avatar 5 Digital and animates the avatar in `avatar_id`.
Options: `image`, `avatar`
Avatar 4 Photo only. The photo to animate. Required when `type` is `image`.
`url` for a link that you host, `base64` for inline bytes.
Options: `url`, `base64`
Required when `image.type` is `url`. Public HTTPS URL of the portrait, hosted by you.
Required when `image.type` is `base64`. MIME type of the bytes, for example `image/jpeg` or `image/png`.
Required when `image.type` is `base64`. The raw base64-encoded image bytes, without a `data:` URI prefix.
Avatar 5 Digital only. The avatar to animate, from HeyGen's avatar catalog. Required when `type` is `avatar`.
Avatar 5 Digital only. Selects the Avatar 5 renderer. Omit it when `type` is `image`, because that body already selects Avatar 4 Photo.
Must be `avatar_v`.
Both engines. Public HTTPS URL of a speech audio file (WAV or MP3), hosted by you. Send this or `script` with `voice_id`. One speech source is required.
Both engines. Text for the avatar to speak. Requires `voice_id`. Use it instead of `audio_url`.
Both engines. The voice that reads `script`. Required when `script` is set. List the voices with `GET /v1/heygen/v3/voices`.
Both engines. Output pixel tier: the size, not the shape. Resolution does not change the price or the render time.
Options: `720p`, `1080p`, `4k`
Both engines. Output canvas shape, independent of `resolution`. See [Resolution and aspect ratio](#resolution-and-aspect-ratio).
Options: `auto`, `16:9`, `9:16`, `4:5`, `5:4`, `1:1`
### Resolution and aspect ratio
`resolution` sets the pixel tier and `aspect_ratio` sets the canvas. The dimensions below apply at `16:9`. At any other ratio, the same tier is fitted to that canvas. Read the shape from `aspect_ratio`, never from the tier name.
| Resolution | Dimensions at 16:9 |
| - | - |
| `720p` | 1280 × 720 |
| `1080p` | 1920 × 1080 |
| `4k` | 3840 × 2160 |
Set `aspect_ratio` explicitly. The default is `16:9`, so a portrait face sent without it comes back inside a landscape frame with bars down both sides. `auto` has a different meaning on each engine.
* Avatar 5 Digital: `auto` follows the shape of the avatar. Use it.
* Avatar 4 Photo: `auto` does not read your photo. A 1080 × 1920 portrait comes back as 1280 × 720, cropped to the face, with no bars to signal it. Send the canvas nearest to the pixels of your photo: `9:16` for a 1080 × 1920 photo, `4:5` for a 3:4 phone portrait. A mismatched ratio is padded, not reframed.
* `1:1` re-canvases every public-catalog avatar, since those are all landscape or portrait. A private or custom avatar can differ.
### Avatar 4 Photo, photo and audio
```bash theme={"system"}
# The media URLs below are placeholders. Point them at your own publicly
# reachable files before you run this.
curl https://api.nunchux.ai/v1/heygen/v3/videos \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"type": "image",
"image": { "type": "url", "url": "https://example.com/example-1080x1920.jpg" },
"audio_url": "https://example.com/example.mp3",
"resolution": "1080p",
"aspect_ratio": "9:16"
}'
```
### Avatar 4 Photo, inline base64 photo
```bash theme={"system"}
# The media URL below is a placeholder. Point it at your own publicly
# reachable file before you run this.
# 1. Encode the photo (raw base64, no "data:" prefix)
IMG=$(base64 -w0 portrait.jpg)
# 2. Submit
curl https://api.nunchux.ai/v1/heygen/v3/videos \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"type": "image",
"image": { "type": "base64", "media_type": "image/jpeg", "data": "'"$IMG"'" },
"audio_url": "https://example.com/example.mp3",
"resolution": "1080p",
"aspect_ratio": "9:16"
}'
```
### Avatar 5 Digital, catalog avatar and script
```bash theme={"system"}
curl https://api.nunchux.ai/v1/heygen/v3/videos \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"type": "avatar",
"avatar_id": "your_avatar_id",
"engine": { "type": "avatar_v" },
"script": "Hi! Here is a quick walkthrough of what we shipped this week.",
"voice_id": "en_us_female_01",
"resolution": "1080p",
"aspect_ratio": "auto"
}'
```
## Voices
When you drive from a `script`, choose a `voice_id` from the voices endpoint.
```bash theme={"system"}
curl https://api.nunchux.ai/v1/heygen/v3/voices \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0"
```
The response lists each voice with its language and gender.
```json theme={"system"}
{
"data": {
"voices": [
{ "voice_id": "en_us_female_01", "language": "English", "gender": "female" },
{ "voice_id": "en_us_male_01", "language": "English", "gender": "male" }
]
}
}
```
## Response
Submit returns a `video_id`. Poll it until `status` is `completed` or `failed`. Then download `video_url`.
### Submit
The submitted job.
The job handle. It is an opaque, variable-length token of about 230 characters. Store it as an unbounded string and send it verbatim on the poll path.
The initial job state: `waiting` or `pending` while the job is queued.
Always `mp4`.
```json theme={"system"}
{
"data": {
"video_id": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1byI6IjNmOWMyYTdlODFiNGQ2MDUiLCJ2IjoiN2QxZTlmMmM0YjhhNGUwZjljM2IyYTFkNmU1ZjRjM2IiLCJtIjoidmlkZW8iLCJleHAiOjE3NTY2OTU2MDAsImlhdCI6MTc1NjA5MDgwMH0.vYKUo0L4vIUnsO4HNHxI5lPoDaYX7gY_dwslw3nWGoI",
"status": "waiting",
"output_format": "mp4"
}
}
```
### Poll
The job status.
The job handle, echoed back.
`waiting` or `pending` while the job is queued, then `processing`, then `completed` or `failed`. Only `completed` and `failed` are terminal. Treat any other value as still running and keep polling.
Output length in seconds. Present when `status` is `completed`. This is the number of output-seconds that you are billed for.
URL of the output video. Present when `status` is `completed`. The URL expires within hours, so download it promptly.
```json theme={"system"}
{
"data": {
"id": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1byI6IjNmOWMyYTdlODFiNGQ2MDUiLCJ2IjoiN2QxZTlmMmM0YjhhNGUwZjljM2IyYTFkNmU1ZjRjM2IiLCJtIjoidmlkZW8iLCJleHAiOjE3NTY2OTU2MDAsImlhdCI6MTc1NjA5MDgwMH0.vYKUo0L4vIUnsO4HNHxI5lPoDaYX7gY_dwslw3nWGoI",
"status": "completed",
"duration": 7.02,
"video_url": "https://resource.heygen.ai/video/...mp4"
}
}
```
## Example
Submit, poll, then download. The example uses Avatar 4 Photo with a hosted photo and a hosted audio file. For Avatar 5 Digital, send `"type": "avatar"` with an `avatar_id` and `"engine": {"type": "avatar_v"}` instead of the `image` object. To drive from text, send `script` and `voice_id` instead of `audio_url`.
```bash cURL theme={"system"}
# The media URLs below are placeholders. Point them at your own publicly
# reachable files before you run this.
BASE=https://api.nunchux.ai
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/heygen/v3/videos" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"type": "image",
"image": { "type": "url", "url": "https://example.com/example-1080x1920.jpg" },
"audio_url": "https://example.com/example.mp3",
"resolution": "1080p",
"aspect_ratio": "9:16"
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
VIDEO_ID=$(jq -er '.data.video_id' <<<"$SUBMIT") || { echo "no video id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until status is completed or failed
DELAY=10
for _ in $(seq 1 60); do
RESP=$(curl -sS "$BASE/v1/heygen/v3/videos/$VIDEO_ID" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.data.status' <<<"$RESP")
case "$STATUS" in completed|failed) break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
# 3. Stop on failure
[ "$STATUS" = "completed" ] || { echo "not completed ($STATUS): $(jq -c '.data' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o avatar.mp4 "$(jq -r '.data.video_url' <<<"$RESP")"
```
```python Python theme={"system"}
# The media URLs below are placeholders. Point them at your own publicly
# reachable files before you run this.
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
f"{BASE}/v1/heygen/v3/videos",
headers=HEADERS,
json={
"type": "image",
"image": {"type": "url", "url": "https://example.com/example-1080x1920.jpg"},
"audio_url": "https://example.com/example.mp3",
"resolution": "1080p",
"aspect_ratio": "9:16",
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
video_id = submit.json()["data"]["video_id"]
# 2. Poll until status is completed or failed
delay = 10
for _ in range(60):
data = requests.get(f"{BASE}/v1/heygen/v3/videos/{video_id}", headers=HEADERS).json()["data"]
if data["status"] in ("completed", "failed"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
# 3. Stop on failure
if data["status"] != "completed":
raise RuntimeError(f"HeyGen job not completed ({data['status']}): {data}")
# 4. Download
with requests.get(data["video_url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("avatar.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
```javascript JavaScript theme={"system"}
// The media URLs below are placeholders. Point them at your own publicly
// reachable files before you run this.
import { writeFile } from "node:fs/promises";
const BASE = "https://api.nunchux.ai";
const HEADERS = {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
};
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// 1. Submit
const submitRes = await fetch(`${BASE}/v1/heygen/v3/videos`, {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
type: "image",
image: { type: "url", url: "https://example.com/example-1080x1920.jpg" },
audio_url: "https://example.com/example.mp3",
resolution: "1080p",
aspect_ratio: "9:16",
}),
});
if (!submitRes.ok) throw new Error(`submit failed: HTTP ${submitRes.status} ${await submitRes.text()}`);
const videoId = (await submitRes.json()).data.video_id;
// 2. Poll until status is completed or failed
let data;
let delay = 10_000;
for (let i = 0; i < 60; i++) {
data = (
await fetch(`${BASE}/v1/heygen/v3/videos/${videoId}`, { headers: HEADERS }).then((r) =>
r.json()
)
).data;
if (data.status === "completed" || data.status === "failed") break;
await sleep(delay);
delay = Math.min(delay * 2, 30_000);
}
// 3. Stop on failure
if (data?.status !== "completed") {
throw new Error(`HeyGen job not completed (${data?.status}): ${JSON.stringify(data)}`);
}
// 4. Download
const res = await fetch(data.video_url);
if (!res.ok) throw new Error(`download failed: HTTP ${res.status}`);
await writeFile("avatar.mp4", Buffer.from(await res.arrayBuffer()));
```
## Tips
* Start from a strong face. On Avatar 4 Photo, send a clear, front-facing portrait with the face well lit and unobstructed. On Avatar 5 Digital, pick the `avatar_id` whose look fits the message.
* You host the inputs. There is no upload endpoint. HeyGen downloads a URL during the job, not at submit, so the URL must stay publicly reachable until the job completes. An inline base64 photo needs no hosting.
* Match the speech source to the job. Point `audio_url` at polished narration that you host. Use `script` with `voice_id` for fast iterations, for language swaps, and when you have nowhere to host an audio file.
* Write scripts in natural, spoken phrasing. It lip-syncs better than dense written prose.
* Set `resolution` explicitly. The default is `1080p`.
* Set `aspect_ratio` explicitly. On Avatar 4 Photo, send the canvas nearest to the pixels of your photo. On Avatar 5 Digital, send `auto`. See [Resolution and aspect ratio](#resolution-and-aspect-ratio).
* Poll every 10 s at first, then back off to 30 s. Polls are free and take no job slot. Render time scales with clip length, so allow 10 minutes or more and set your client timeout to match.
* Branch on `failed`. It is terminal, and the body carries the reason.
* Download the video as soon as `status` is `completed`. The URL expires within hours.
* Store `video_id` as an unbounded string. It is about 230 characters long, and you must send it verbatim on the poll path.
* A 5xx or a timeout on submit does not tell you whether the job was created. Do not resubmit blindly, because a second job that completes is billed too. If the submit returned a `video_id`, poll it. If it returned nothing, wait, then read your credit balance before you submit again.
## Errors and limits
* A 200 on submit means that HeyGen accepted the job. It does not mean that the job succeeded. A failure appears later as `failed` in the poll response, with the reason in the body.
* A failed job is never charged.
* `resolution` accepts `720p`, `1080p` and `4k`. Any other value returns 400 with the code `unsupported_resolution`. Nothing is charged.
* A submit with no face (`image` when `type` is `image`, `avatar_id` when `type` is `avatar`), with no speech source, or with `script` but no `voice_id` returns 400.
* A `video_id` is scoped to your account. Poll only a handle that your own submit returned. A handle that belongs to another account returns 409.
* Neither engine has a duration control. The clip is as long as the audio or the script.
A URL that HeyGen cannot download fails the submit at once with 400 and the code `invalid_parameter`. Host the file publicly, or send the photo inline as base64.
```json theme={"system"}
{
"error": {
"code": "invalid_parameter",
"message": "Invalid URL in files[0]: Could not download the file. Ensure the URL is publicly accessible.",
"doc_url": "https://developers.heygen.com/docs/error-codes#invalid-parameter"
}
}
```
The general error contract, the rate limits and the credit rules apply to every endpoint. See [Errors](/errors), [Rate limits](/rate-limits), [Credits and pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
# Kling
Source: https://docs.nunchux.ai/partner-models/kling
Kling V3 and Kling V3 Omni video generation with your Nunchux key: text-to-video, image-to-video, omni video and motion control.
# Kling
Kling is Kuaishou's video generation model family. Every Kling endpoint is asynchronous: submit a job, poll it until it reaches a terminal status, then download the output. The routes mirror Kling's own API, so the request bodies, response shapes and error codes follow Kling's documentation. You call them with your Nunchux API key and need no Kling account.
## Models
| Model | Model ID | Tasks | Modes/resolution | Duration |
| - | - | - | - | - |
| Kling V3 | `kling-v3` | Text-to-video, image-to-video, motion control | `std`, `pro`, `4k` (motion control: `std`, `pro`) | 3 to 15 s (motion control: follows the reference video) |
| Kling V3 Omni | `kling-v3-omni` | Omni video | `std`, `pro`, `4k` | 3 to 15 s (storyboard: up to 30 s) |
`std` is the standard tier. It is the fastest and the most economical. `pro` gives sharper detail and stronger motion. `4k` gives 4K-resolution output. Higher tiers take longer.
Start with `kling-v3`. It supports the `std`, `pro` and `4k` modes, and it is the only model with optional native audio. Use `kling-v3-omni` on the [omni video](#omni-video) endpoint for multi-reference or storyboard work. For cost-sensitive jobs, use `std`. It is the cheapest per second.
## Endpoints
| Capability | Route |
| - | - |
| [Text-to-video](#text-to-video) | `POST /v1/klingai/videos/text2video` |
| [Image-to-video](#image-to-video) | `POST /v1/klingai/videos/image2video` |
| [Omni video](#omni-video) | `POST /v1/klingai/videos/omni-video` |
| [Motion control](#motion-control) | `POST /v1/klingai/videos/motion-control` |
| [Poll a task](#poll-a-task) | `GET /v1/klingai/videos/{capability}/{task_id}` |
The base URL is `https://api.nunchux.ai`. Each submit returns a `task_id` immediately. Poll the same capability that you submitted to, with the `task_id` appended to the route.
## Common fields
Send your API key in the `X-API-Key` header (see [Authentication](/authentication)). Send `Content-Type: application/json` on every submit. Send a `User-Agent` header that identifies your application, such as `YourApp/1.0`. The request fields below appear on more than one endpoint.
The model to generate with. See [Models](#models). Defaults to `kling-v3` on text-to-video, image-to-video and motion control. Defaults to `kling-v3-omni` on omni video.
Options: `kling-v3` on text-to-video, image-to-video and motion control. `kling-v3-omni` on omni video.
Quality and speed tier. Higher tiers take longer. Motion control accepts `std` and `pro` only.
Options: `std`, `pro`, `4k`
Description of the video to generate. Max 2500 characters. On image-to-video, describe the motion or the transformation. On omni video, refer to an input with a placeholder such as `<<>>`. Not sent on motion control. Ignored in [storyboard mode](#storyboard-mode).
What to exclude from the video. Max 2500 characters. Text-to-video and image-to-video only.
Output length in seconds, sent as a string. Billed per second. Common values are `"5"` and `"10"`. Not sent on motion control, where the output length follows the reference video. Ignored in [storyboard mode](#storyboard-mode).
Range: 3-15.
Native audio generation. Only on `kling-v3`. Text-to-video and image-to-video only. In `4k` mode, audio is always included and this flag is ignored.
Options: `on`, `off`
If set, Kling POSTs status updates to this URL, and you do not need to poll. Text-to-video and image-to-video only.
## Text-to-video
`POST https://api.nunchux.ai/v1/klingai/videos/text2video`
Generate a video from a text prompt. Send `model_name` and `prompt`. Every other field is optional and falls back to its default. Use `kling-v3`.
Output aspect ratio.
Options: `16:9`, `9:16`, `1:1`
To plan a sequence of shots instead of one prompt, see [Storyboard mode](#storyboard-mode).
```bash cURL theme={"system"}
curl https://api.nunchux.ai/v1/klingai/videos/text2video \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model_name": "kling-v3",
"prompt": "a paper crane unfolding into a real bird, cinematic slow motion",
"mode": "std",
"duration": "5",
"aspect_ratio": "16:9"
}'
```
```python Python theme={"system"}
import os, requests
resp = requests.post(
"https://api.nunchux.ai/v1/klingai/videos/text2video",
headers={"X-API-Key": os.environ["NUNCHUX_API_KEY"], "User-Agent": "YourApp/1.0"},
json={
"model_name": "kling-v3",
"prompt": "a paper crane unfolding into a real bird, cinematic slow motion",
"mode": "std",
"duration": "5",
"aspect_ratio": "16:9",
},
)
task_id = resp.json()["data"]["task_id"]
```
```javascript JavaScript theme={"system"}
const submit = await fetch("https://api.nunchux.ai/v1/klingai/videos/text2video", {
method: "POST",
headers: {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
},
body: JSON.stringify({
model_name: "kling-v3",
prompt: "a paper crane unfolding into a real bird, cinematic slow motion",
mode: "std",
duration: "5",
aspect_ratio: "16:9",
}),
}).then((r) => r.json());
const taskId = submit.data.task_id;
```
## Image-to-video
`POST https://api.nunchux.ai/v1/klingai/videos/image2video`
Animate a starting image into a video clip. Send `model_name`, `image` and a `prompt` that describes the motion. Use `kling-v3`.
The starting frame. Pass a public URL or raw base64, with no `data:image/...;base64,` prefix. Kling rejects the prefixed form with code 1201.
```bash cURL theme={"system"}
IMG_B64=$(base64 < input.jpg | tr -d '\n')
curl https://api.nunchux.ai/v1/klingai/videos/image2video \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d "{
\"model_name\": \"kling-v3\",
\"prompt\": \"camera pushes in slowly, subtle ambient motion\",
\"mode\": \"pro\",
\"duration\": \"5\",
\"image\": \"$IMG_B64\"
}"
```
```python Python theme={"system"}
import base64, os, requests
with open("input.jpg", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
resp = requests.post(
"https://api.nunchux.ai/v1/klingai/videos/image2video",
headers={"X-API-Key": os.environ["NUNCHUX_API_KEY"], "User-Agent": "YourApp/1.0"},
json={
"model_name": "kling-v3",
"prompt": "camera pushes in slowly, subtle ambient motion",
"mode": "pro",
"duration": "5",
"image": img_b64,
},
)
task_id = resp.json()["data"]["task_id"]
```
```javascript JavaScript theme={"system"}
import { readFile } from "node:fs/promises";
const imgB64 = (await readFile("input.jpg")).toString("base64");
const submit = await fetch("https://api.nunchux.ai/v1/klingai/videos/image2video", {
method: "POST",
headers: {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
},
body: JSON.stringify({
model_name: "kling-v3",
prompt: "camera pushes in slowly, subtle ambient motion",
mode: "pro",
duration: "5",
image: imgB64,
}),
}).then((r) => r.json());
const taskId = submit.data.task_id;
```
## Omni video
`POST https://api.nunchux.ai/v1/klingai/videos/omni-video`
Compose one clip from reference images, videos and elements. Use `kling-v3-omni`. At least one of `image_list`, `video_list` or `element_list` must be non-empty. You can send up to 7 reference inputs across the three lists. Refer to an input from the prompt with a 1-indexed placeholder in submission order: `<<>>`, `<<>>`, `<<>>`.
Reference images.
Public URL or raw base64 of the reference image, with no `data:image/...;base64,` prefix.
Optional sequence anchor. An entry with no type acts as an additional reference.
Options: `first_frame`, `end_frame`
Reference videos to continue or to composite from.
Public URL of the reference video.
Reference elements, such as segmented objects or characters, to compose into the video.
To plan a sequence of shots, see [Storyboard mode](#storyboard-mode).
```bash cURL theme={"system"}
# Placeholder URL: point it at your own publicly reachable file.
curl https://api.nunchux.ai/v1/klingai/videos/omni-video \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model_name": "kling-v3-omni",
"prompt": "<<>> walks through a neon-lit market, cinematic",
"mode": "std",
"duration": "5",
"image_list": [{ "image_url": "https://example.com/example.jpg" }]
}'
```
```python Python theme={"system"}
# Placeholder URL: point it at your own publicly reachable file.
import os, requests
resp = requests.post(
"https://api.nunchux.ai/v1/klingai/videos/omni-video",
headers={"X-API-Key": os.environ["NUNCHUX_API_KEY"], "User-Agent": "YourApp/1.0"},
json={
"model_name": "kling-v3-omni",
"prompt": "<<>> walks through a neon-lit market, cinematic",
"mode": "std",
"duration": "5",
"image_list": [{"image_url": "https://example.com/example.jpg"}],
},
)
task_id = resp.json()["data"]["task_id"]
```
```javascript JavaScript theme={"system"}
// Placeholder URL: point it at your own publicly reachable file.
const submit = await fetch("https://api.nunchux.ai/v1/klingai/videos/omni-video", {
method: "POST",
headers: {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
},
body: JSON.stringify({
model_name: "kling-v3-omni",
prompt: "<<>> walks through a neon-lit market, cinematic",
mode: "std",
duration: "5",
image_list: [{ image_url: "https://example.com/example.jpg" }],
}),
}).then((r) => r.json());
const taskId = submit.data.task_id;
```
## Motion control
`POST https://api.nunchux.ai/v1/klingai/videos/motion-control`
Drive a still character with the motion of a reference video. Kling transfers the action from the driver clip and preserves the character's identity. Use `kling-v3`. Do not send `prompt` or `duration`. The output length follows the reference video.
Public URL or raw base64 of the still character to animate, with no `data:image/...;base64,` prefix.
Public URL of the driver clip whose motion is transferred onto the character.
How the character reference is provided. `image` expects a still image in `image_url` (cap: 10 s). `video` expects a video in `image_url` (cap: 30 s).
Options: `image`, `video`
Length of the driver clip in seconds. This value drives per-second billing, so declare it accurately.
Range: 1-10.
Billing
Kling bills per second of generated output. If you under-declare `input_video_seconds`, a follow-up charge is applied after the job completes. If you over-declare, the difference is not refunded automatically.
```bash cURL theme={"system"}
# Placeholder URLs: point them at your own publicly reachable files.
curl https://api.nunchux.ai/v1/klingai/videos/motion-control \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model_name": "kling-v3",
"mode": "std",
"image_url": "https://example.com/example.jpg",
"video_url": "https://example.com/example.mp4",
"character_orientation": "image",
"input_video_seconds": 5
}'
```
```python Python theme={"system"}
# Placeholder URLs: point them at your own publicly reachable files.
import os, requests
resp = requests.post(
"https://api.nunchux.ai/v1/klingai/videos/motion-control",
headers={"X-API-Key": os.environ["NUNCHUX_API_KEY"], "User-Agent": "YourApp/1.0"},
json={
"model_name": "kling-v3",
"mode": "std",
"image_url": "https://example.com/example.jpg",
"video_url": "https://example.com/example.mp4",
"character_orientation": "image",
"input_video_seconds": 5,
},
)
task_id = resp.json()["data"]["task_id"]
```
```javascript JavaScript theme={"system"}
// Placeholder URLs: point them at your own publicly reachable files.
const submit = await fetch("https://api.nunchux.ai/v1/klingai/videos/motion-control", {
method: "POST",
headers: {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
},
body: JSON.stringify({
model_name: "kling-v3",
mode: "std",
image_url: "https://example.com/example.jpg",
video_url: "https://example.com/example.mp4",
character_orientation: "image",
input_video_seconds: 5,
}),
}).then((r) => r.json());
const taskId = submit.data.task_id;
```
## Storyboard mode
Text-to-video and omni video accept a multi-shot storyboard. Set `multi_shot` to `true` and describe each shot in `multi_prompt`. The top-level `prompt` and `duration` are ignored. The total length is the sum of the per-shot durations. All constraints are validated before billing, so a malformed request returns a 400 with no charge.
Enable multi-shot storyboard mode.
How shots are planned. Required when `multi_shot` is `true`.
Options: `customize`, `intelligence`
Per-shot instructions. Text-to-video accepts up to 6 shots and 15 seconds in total. Omni video caps the total at 30 seconds.
Sequential from 1, with no gaps or duplicates.
Non-empty description for this shot.
Integer seconds for this shot, sent as a string. Range 1-15 on text-to-video and 1-10 on omni video.
```bash theme={"system"}
curl https://api.nunchux.ai/v1/klingai/videos/text2video \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model_name": "kling-v3",
"mode": "std",
"multi_shot": true,
"shot_type": "customize",
"multi_prompt": [
{ "index": 1, "prompt": "wide establishing shot at dawn, fog over a mountain lake", "duration": "5" },
{ "index": 2, "prompt": "close-up on a ripple spreading across the water surface", "duration": "5" }
]
}'
```
## Poll a task
`GET https://api.nunchux.ai/v1/klingai/videos/{capability}/{task_id}`
`{capability}` is one of `text2video`, `image2video`, `omni-video` and `motion-control`. Poll the same capability that you submitted to. A `task_id` polled at another capability's route returns 404.
Poll every 5 to 10 seconds at first, then back off to 15 to 30 seconds. Polls do not use credits and do not count toward your concurrency limit. `task_status` moves from `submitted` to `processing`, then to `succeed` or `failed`. A task handle expires after 7 days. Output URLs expire within hours, so download the output as soon as the task succeeds.
0 on success. A non-zero value indicates a Kling error.
Human-readable status, such as `SUCCEED`.
The task payload.
The task handle. Poll the same capability that you submitted to with it.
`submitted`, `processing`, `succeed` or `failed`.
Explains why a job failed. Show it to your users.
Present once `task_status` is `succeed`.
URL of the output. It expires within hours, so download promptly.
Length of the generated clip in seconds.
The submit response has the same shape, with `task_status` set to `submitted`. A poll response when the task has succeeded:
```json theme={"system"}
{
"code": 0,
"message": "SUCCEED",
"data": {
"task_id": "eyJhbGciOiJIUzI1NiIs...",
"task_status": "succeed",
"task_result": {
"videos": [
{ "url": "https://v16-kling-fdl.klingai.com/...",
"duration": "5.0" }
]
}
}
}
```
## Example
Submit a text-to-video job, poll it until it reaches a terminal status, then download the clip.
```bash cURL theme={"system"}
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body https://api.nunchux.ai/v1/klingai/videos/text2video \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model_name": "kling-v3",
"prompt": "A cinematic drone shot sweeping over snow-capped mountains at sunrise",
"mode": "std",
"duration": "5",
"aspect_ratio": "16:9"
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
TASK=$(jq -er '.data.task_id' <<<"$SUBMIT") || { echo "no task id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until the task reaches a terminal status
DELAY=5
TERMINAL=0
for _ in $(seq 1 90); do
RESP=$(curl -sS https://api.nunchux.ai/v1/klingai/videos/text2video/$TASK \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.data.task_status' <<<"$RESP")
case "$STATUS" in succeed|failed) TERMINAL=1; break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
[ "$TERMINAL" = 1 ] || { echo "timed out (last status: $STATUS)" >&2; exit 1; }
# 3. Check for failure
[ "$STATUS" = "succeed" ] || { echo "Kling: status $STATUS, not succeed: $(jq -c '.data.task_status_msg // empty' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o kling.mp4 "$(jq -r '.data.task_result.videos[0].url' <<<"$RESP")"
```
```python Python theme={"system"}
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
f"{BASE}/v1/klingai/videos/text2video",
headers=HEADERS,
json={
"model_name": "kling-v3",
"prompt": "A cinematic drone shot sweeping over snow-capped mountains at sunrise",
"mode": "std",
"duration": "5",
"aspect_ratio": "16:9",
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
task_id = submit.json()["data"]["task_id"]
# 2. Poll until the task reaches a terminal status
TERMINAL = ("succeed", "failed")
delay = 5
for _ in range(90):
task = requests.get(f"{BASE}/v1/klingai/videos/text2video/{task_id}", headers=HEADERS).json()
if task["data"]["task_status"] in TERMINAL:
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("Kling task did not finish")
# 3. Check for failure
if task["data"]["task_status"] != "succeed":
raise RuntimeError(f"Kling: status {task['data']['task_status']}, not succeed: {task['data'].get('task_status_msg')}")
# 4. Download
with requests.get(task["data"]["task_result"]["videos"][0]["url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("kling.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const BASE = "https://api.nunchux.ai";
const HEADERS = {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
};
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// 1. Submit
const submitRes = await fetch(`${BASE}/v1/klingai/videos/text2video`, {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
model_name: "kling-v3",
prompt: "A cinematic drone shot sweeping over snow-capped mountains at sunrise",
mode: "std",
duration: "5",
aspect_ratio: "16:9",
}),
});
if (!submitRes.ok) throw new Error(`submit failed: HTTP ${submitRes.status} ${await submitRes.text()}`);
const taskId = (await submitRes.json()).data.task_id;
// 2. Poll until the task reaches a terminal status
const TERMINAL = ["succeed", "failed"];
let task;
let delay = 5_000;
for (let i = 0; i < 90; i++) {
task = await fetch(`${BASE}/v1/klingai/videos/text2video/${taskId}`, { headers: HEADERS }).then((r) =>
r.json()
);
if (TERMINAL.includes(task?.data?.task_status)) break;
await sleep(delay);
delay = Math.min(delay * 2, 30_000);
}
if (!TERMINAL.includes(task?.data?.task_status)) throw new Error("Kling task did not finish");
// 3. Check for failure
if (task.data.task_status !== "succeed") {
throw new Error(`Kling: status ${task.data.task_status}, not succeed: ${JSON.stringify(task.data.task_status_msg)}`);
}
// 4. Download
const res = await fetch(task.data.task_result.videos[0].url);
if (!res.ok) throw new Error(`download failed: HTTP ${res.status}`);
await writeFile("kling.mp4", Buffer.from(await res.arrayBuffer()));
```
## Tips
* Poll; do not busy-wait. Poll every 5 to 10 seconds at first, then back off to 15 to 30 seconds.
* Expect text-to-video and image-to-video to finish in 60 to 120 seconds. A `kling-v3` `std` 5-second clip takes about 45 seconds. Motion control and omni video run 2 to 6 minutes, and omni video takes longer with more shots.
* Set a client timeout of at least 5 minutes for `std` text-to-video and image-to-video. Use at least 10 minutes for `pro`, `4k`, motion control and omni video.
* Download outputs as soon as the task succeeds. Output URLs expire within hours.
* Keep prompts concrete and visual. Describe what the camera sees, not abstract concepts.
* Use `negative_prompt` to rule out unwanted elements such as blur, text overlays and watermarks.
* Set `sound: "on"` on `kling-v3` to add ambient audio with no post-processing. In `4k` mode, audio is always included.
* Send base64 images as raw bytes with no `data:image/...;base64,` prefix. Public URLs must be reachable by Kling's servers: no auth walls and no localhost.
* Use a clear, well-composed starting frame for image-to-video. Kling extends what it sees, so a blurry or cluttered input produces inconsistent motion.
* Use a clean, well-lit character image for motion control. Identity preservation is only as good as the source.
* Pass a `callback_url` on text-to-video or image-to-video to receive status updates without polling.
## Errors and limits
A 200 means that Kling accepted the job and billing started, not that the job succeeded. Do not resubmit after a 200; a second submit creates and bills a second job. Retry only on a network or timeout error that arrives before the submit response.
| Response | Meaning | What to do |
| - | - | - |
| `400` | Invalid request: a bad model name, an unsupported mode or duration, or a missing field. | Read the message and fix the request. A 400 is rejected before billing, so it is never charged. |
| `401` | Authentication failed, or the task handle expired. | Check your API key and the `task_id`. Task handles expire after 7 days. |
| `200` with `code` `1201` | Kling could not read an asset. Usually an `image_url` or `video_url` that is not publicly reachable, or an image sent with a `data:` prefix. | The submit was billed. Fix the asset (public URL or raw base64) and submit again. Open a support ticket for a refund. |
| `200` with `task_status` `failed` | The job failed after Kling accepted it. | No action is needed. Credits are refunded automatically. Show `task_status_msg` to your users. |
| Other `4xx` or `5xx` | Forwarded from Kling: content moderation, bad parameters or another upstream reason. The response body explains the cause. | Retry only when the cause is transient. |
Retry a `5xx` or a transient `429` with backoff. Never retry another `4xx`; the request is invalid as sent. For the general error contract, see [Error codes](/errors). For the concurrency caps and request limits, see [Rate limits](/rate-limits). Every job counts toward your plan's simultaneous-jobs cap; polls do not. Credits are deducted at submit and refunded when a job fails; see [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
Kling-specific limits:
* `prompt` and `negative_prompt`: up to 2500 characters.
* `duration`: 3 to 15 seconds. In storyboard mode, text-to-video accepts up to 6 shots and 15 seconds in total, and omni video accepts up to 30 seconds in total.
* Omni video: up to 7 reference inputs across `image_list`, `video_list` and `element_list`.
* Motion control: `input_video_seconds` from 1 to 10, and `mode` is `std` or `pro`.
* Voice control is not supported on `kling-v3`. Sending `voice_control` or `voice_list` returns a 400.
## Migrating from /v1/video/kling/\*
Kling used to live under `/v1/video/kling/*`. It now lives under `/v1/klingai/videos/*`, a faithful mirror of Kling's own API: same request bodies, same response shapes, same error codes. If you call Kling directly today, you can point at us by swapping the base URL and using your Nunchux key as the Bearer token. Existing `/v1/video/kling/*` integrations keep working; migrate when convenient.
Two things that are not a find-and-replace
* Polling moved from one URL to one per capability. There is no `/v1/klingai/videos/tasks/{task_id}`. Instead of a single `GET /v1/video/kling/tasks/{task_id}`, poll the endpoint you submitted to. For example, a job from `POST /v1/klingai/videos/text2video` is polled at `GET /v1/klingai/videos/text2video/{task_id}`. A shared poll helper needs to carry the submit path through; a `task_id` polled at another capability's URL returns 404.
* One path was renamed beyond the prefix change: `video-effects` → `effects`. The rest keep their trailing segment.
| Legacy (superseded) | Current |
| - | - |
| `POST /v1/video/kling/text2video` | `POST /v1/klingai/videos/text2video` |
| `POST /v1/video/kling/image2video` | `POST /v1/klingai/videos/image2video` |
| `POST /v1/video/kling/omni-video` | `POST /v1/klingai/videos/omni-video` |
| `POST /v1/video/kling/motion-control` | `POST /v1/klingai/videos/motion-control` |
| `POST /v1/video/kling/video-effects` | `POST /v1/klingai/videos/effects` |
| `POST /v1/video/kling/image-recognize` | `POST /v1/klingai/videos/image-recognize` |
| `GET /v1/video/kling/tasks/{task_id}` | `GET /v1/klingai/videos/{capability}/{task_id}` |
# MiniMax H3
Source: https://docs.nunchux.ai/partner-models/minimax
MiniMax's H3 video model for text-to-video, image-to-video and reference-to-video through one asynchronous endpoint.
# MiniMax H3
MiniMax H3 is an omni-modal model that reads text, images, video and audio in one context and generates a video clip. One model ID serves all three tasks. The `content[]` items you send decide which task you get. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors MiniMax's own API, so the request and response shapes are the provider's.
Best for
* Multi-reference scenes. Up to nine reference images, three reference clips and a reference track carry a subject or a style into new footage.
* 2K output. A higher tier than the other partner video models sell.
* Animating a still. A first frame fixes the composition and the prompt drives the motion.
* Longer prompts. The vendor accepts up to 7,000 characters per text item.
## Models
| Model | Model ID | Tasks | Resolution | Duration | Reference media |
| - | - | - | - | - | - |
| MiniMax H3 | `MiniMax-H3` | t2v, i2v, r2v | 768P, 2K | 4 to 15 s | Up to 9 images, 3 clips and 1 audio track |
Durations are whole seconds. All three tasks send `MiniMax-H3` as the `model` value. 2K costs more per second than 768P.
| Task | Resolution | Duration | Aspect ratio |
| - | - | - | - |
| Text-to-video | 768P, 2K | 4 to 15 s | Required: 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 |
| Image-to-video | 768P, 2K | 4 to 15 s | Follows the source image |
| Reference-to-video | 768P, 2K | 4 to 15 s | Follows the reference set |
The `content[]` items select the task. A `text` item alone gives text-to-video. A `first_frame` image item gives image-to-video. `reference_image`, `reference_video` and `reference_audio` items give reference-to-video.
A first frame and a reference set cannot appear in the same body. The two use different `role` values, and a body that mixes the role sets is rejected. Pick one task per request.
## Endpoints
| Step | Route |
| - | - |
| Submit | `POST /v1/minimax/v2/video_generation` |
| Poll | `GET /v1/minimax/v2/query/video_generation/{task_id}` |
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. The examples also send a `User-Agent` header that names your application.
The model ID: `MiniMax-H3`.
The prompt and the media items. Put the `text` item first. Pair each role with its own container: an image rides `image_url`, a clip rides `video_url`, a track rides `audio_url`. A role on the wrong container is accepted without complaint and then does not do what you meant.
`text` for the prompt. `image_url` for an image. `video_url` for a clip. `audio_url` for a track.
The prompt, up to 7,000 characters. Required on a `text` item. All tasks.
What a media item is for.
* `first_frame`: the image the clip starts on. Image-to-video. One item.
* `last_frame`: the image the clip ends on. Image-to-video, optional. One item.
* `reference_image`: a reference image. Reference-to-video. Up to 9 items.
* `reference_video`: a reference clip. Reference-to-video. Up to 3 items, 2 to 15 seconds each, 15 seconds in total.
* `reference_audio`: a reference track. Reference-to-video, optional. One item.
On an `image_url` item. Holds `url`, the URL of the image.
On a `video_url` item. Holds `url`, the URL of the clip.
On an `audio_url` item. Holds `url`, the URL of the track.
The output tier: `768P` or `2K`. All tasks.
The clip length in whole seconds, 4 to 15. All tasks. You are billed per second of output.
The output shape. Required on a text-only body: `21:9`, `16:9`, `4:3`, `1:1`, `3:4` or `9:16`. `adaptive` is rejected there. Leave it out when the body carries an image, a clip or a track. The output then follows the source.
## Response
### Submit
A 200 status means that the provider accepted the task. It does not mean that the clip is ready.
The task ID. Poll `GET /v1/minimax/v2/query/video_generation/{task_id}` with it.
### Poll
`queued` or `running` while the task runs. `succeeded`, `failed` or `cancelled` when it ends. Poll until you read one of the three terminal values.
The URL of the clip. Present only when `status` is `succeeded`. The URL is time-limited. Download the clip as soon as the task succeeds. A later poll returns a fresh URL. A task can be polled for 7 days.
The failure detail. Present when the task did not succeed. Read it before you resubmit.
The billed quantities: output seconds, reference-video seconds and the billed image count.
## Example
Submit, poll, download. The submit and the poll use different paths.
```bash cURL theme={"system"}
BASE=https://api.nunchux.ai
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/minimax/v2/video_generation" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "a fox trotting through a snowy forest at dawn" }
],
"resolution": "768P",
"duration": 6,
"ratio": "16:9"
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
TASK=$(jq -er '.task_id' <<<"$SUBMIT") || { echo "no task_id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until the status is terminal
DELAY=5
for _ in $(seq 1 90); do
RESP=$(curl -sS "$BASE/v1/minimax/v2/query/video_generation/$TASK" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.task.status // ""' <<<"$RESP")
case "$STATUS" in succeeded|failed|cancelled) break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
# 3. Check for failure
[ "$STATUS" = "succeeded" ] || { echo "task $STATUS: $(jq -r '.task.error // "no error"' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o minimax.mp4 "$(jq -r '.task.content.url' <<<"$RESP")"
```
```python Python theme={"system"}
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
f"{BASE}/v1/minimax/v2/video_generation",
headers=HEADERS,
json={
"model": "MiniMax-H3",
"content": [
{"type": "text", "text": "a fox trotting through a snowy forest at dawn"}
],
"resolution": "768P",
"duration": 6,
"ratio": "16:9",
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
task = submit.json()["task_id"]
# 2. Poll until the status is terminal
delay = 5
for _ in range(90):
data = requests.get(
f"{BASE}/v1/minimax/v2/query/video_generation/{task}", headers=HEADERS
).json()
status = data.get("task", {}).get("status", "")
if status in ("succeeded", "failed", "cancelled"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("MiniMax task did not reach a terminal status")
# 3. Check for failure
if status != "succeeded":
raise RuntimeError(f"MiniMax task {status}: {data['task'].get('error')}")
# 4. Download
with requests.get(data["task"]["content"]["url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("minimax.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
To animate a still, append an `image_url` item with `role` set to `first_frame` and leave `ratio` out. To carry a subject into a new scene, append `reference_image`, `reference_video` or `reference_audio` items instead, and leave `ratio` out. The bodies below replace the submit body in the example.
```json Image-to-video theme={"system"}
{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "the fox turns and trots toward the camera" },
{
"type": "image_url",
"role": "first_frame",
"image_url": { "url": "https://example.com/frame.jpg" }
}
],
"resolution": "768P",
"duration": 6
}
```
```json Reference-to-video theme={"system"}
{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "the subject walks down a rain-soaked street at night, slow tracking shot" },
{
"type": "image_url",
"role": "reference_image",
"image_url": { "url": "https://example.com/ref-1.jpg" }
},
{
"type": "image_url",
"role": "reference_image",
"image_url": { "url": "https://example.com/ref-2.jpg" }
},
{
"type": "video_url",
"role": "reference_video",
"video_url": { "url": "https://example.com/clip.mp4" }
},
{
"type": "audio_url",
"role": "reference_audio",
"audio_url": { "url": "https://example.com/track.mp3" }
}
],
"resolution": "768P",
"duration": 6
}
```
## Tips
* Name a shape on text-to-video. `ratio` is required there, so decide the frame before you run instead of accepting whatever the first attempt gives you.
* Give at least one image or one clip on reference-to-video. A reference set describes the subject. The prompt describes the motion.
* Keep the reference images consistent in subject and style. Images that disagree pull the result in different directions.
* Draft at 768P. 2K costs more per second, so settle the prompt at the lower tier first.
* Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.
## Errors and limits
A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with `task.status` set to `failed` or `cancelled`, and `task.error` explains why. Read it before you resubmit.
A failed submit returns an HTTP error status. See [Error codes](/errors). Two request shapes are rejected: a text-only body without `ratio`, or with `ratio` set to `adaptive`, and a body that mixes the frame roles with the reference roles.
A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned a `task_id`, poll it. If it returned nothing, check your credits before you try again.
MiniMax applies these limits to the media you attach, and caps the whole set at 12 files:
* Images: 256 to 5,760 pixels on each side, JPG, JPEG, PNG, WEBP, HEIC or HEIF, up to 30 MB each
* Video: MP4 or MOV, up to 50 MB each, 2 to 15 seconds per clip and 15 seconds in total
* Audio: WAV or MP3, up to 15 MB each, 2 to 15 seconds
A file you pass by URL is not checked at submit. A file over a cap is refused by the provider after the request was accepted.
Reference media changes the charge. Reference video is billed per second at the same rate as the output. The first five reference images are included in the rate, and each image after that adds a charge. Reference audio is free. Every task counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits), [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
# Nano Banana
Source: https://docs.nunchux.ai/partner-models/nano-banana
Generate and edit images with Google's Nano Banana models in one synchronous call.
# Nano Banana
Nano Banana is Google's Gemini image model. It generates an image from a text prompt, and it edits images that you send with the prompt. The call is synchronous: one POST returns the finished image in the response body, so there is nothing to poll. The request and response shapes mirror Google's own API.
## Models
| Model | Model ID | Tasks | Sizes | Aspect ratios |
| - | - | - | - | - |
| Nano Banana 2 | `gemini-3.1-flash-image` | Generate, edit | `512`, `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |
| Nano Banana 2.1 | `gemini-nano-banana-2.1` | Generate, edit | `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |
| Nano Banana Pro | `gemini-3-pro-image` | Generate, edit | `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |
All three models generate and edit on the same endpoint. All three accept up to 14 reference images per edit. The size is a pixel budget that you set with `image_size`. The shape is a separate axis that you set with `aspect_ratio`.
Use Nano Banana 2 for fast everyday generation and edits. It is the default model, and it is the only one that accepts the `512` size.
Use Nano Banana 2.1 for the same everyday work when you do not need the `512` size. It is Google's update to Nano Banana 2, and it starts at `1K`. It is a thinking model, so it inserts a `thought` step before the output step in the response.
Use Nano Banana Pro when you need the highest fidelity. It starts at `1K`. It is a thinking model, so it inserts a `thought` step before the output step in the response.
## Endpoints
The base URL is `https://api.nunchux.ai`.
| Step | Method | Path |
| - | - | - |
| Generate or edit | `POST` | `/v1/google/v1beta/interactions` |
There is no poll route. The image is in the response.
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. Send a `User-Agent` header that names your application, for example `YourApp/1.0`.
The JSON body names the model and an `input` array of content parts. A text-only input generates a new image. Image parts turn the request into an edit. See [Editing](#editing).
The Nano Banana model to run. See the [Models](#models) table.
Options: `gemini-3.1-flash-image`, `gemini-nano-banana-2.1`, `gemini-3-pro-image`
Ordered content parts. A text-only input generates. Add `{ "type": "image" }` parts (up to 14) and the text becomes an edit instruction across them. Image parts also carry `data` and `mime_type`. See [Editing](#editing).
The part kind.
Options: `text`, `image`
The prompt or the edit instruction. Required on `type: "text"` parts.
Output controls. A request that omits `response_format` or its `type` is rejected with `no_image_format` or `invalid_response_format`.
Must be `"image"`.
Options: `image`
Output resolution, as a pixel budget. `512` is accepted on `gemini-3.1-flash-image` only. Google's documentation calls that tier 0.5K, but `0.5K` is rejected as a value.
Options: `1K`, `2K`, `4K`
Output shape, independent of `image_size`. Omitted on a generation, the output is `1:1` on Nano Banana 2 and Nano Banana Pro. On Nano Banana 2.1, set the ratio when the shape matters. Omitted on an edit, the output keeps the shape of the first input image.
Options: `1:1`, `2:3`, `3:2`, `3:4`, `4:3`, `4:5`, `5:4`, `9:16`, `16:9`, `21:9`, `1:8`, `8:1`, `1:4`, `4:1`
Output size and shape
* `image_size` is a pixel budget, not a shape. `1K` is 1024×1024 at `1:1`, 1584×672 at `21:9` and 768×1376 at `9:16`.
* Set `image_size` explicitly. When you omit it, the call is billed at the largest tier (`4K`), on generations and on edits.
* The 14 ratios in the list are accepted on every model. `1:1`, `21:9` and `9:16` are the ones that have been run end to end.
* There is no `auto` value for `aspect_ratio`. To keep the shape of the input on an edit, omit the field.
### Editing
Add one or more `{ "type": "image" }` parts to `input`, up to 14. The text part becomes an instruction across all of them. The result is always one new image, not a batch and not a mechanical merge. `image_size` and `aspect_ratio` work exactly as on a generation.
Raw base64 image bytes on `type: "image"` parts, with no `data:` prefix. Images are sent inline. There is no URL form.
The MIME type of the image on `type: "image"` parts, for example `image/png` or `image/jpeg`.
## Response
The image comes back in the same response, as base64 JPEG bytes.
Identifier for this interaction.
The envelope kind.
The model that ran, echoed back.
The status of the interaction. It is terminal on arrival. The call is synchronous, so there is no in-progress state to poll for.
Unix timestamp when the interaction was created.
Unix timestamp of the last update, in practice when the image finished.
The service tier the call was served on.
Accounting for the call.
The steps of the interaction. The image is on the first entry whose `content[]` carries a `data` field.
Opaque vendor blob on a step of its own. It can be absent. It is not the image, so skip it.
The output parts. Take the first step that has one. A `thought` step can sit ahead of it.
Base64-encoded JPEG bytes. Decode and save.
Reading the response
* Select the output step by shape, not by index: the first `steps[]` entry whose `content[]` carries a `data` field. The `signature` step is optional, and a thinking model such as Nano Banana 2.1 or Nano Banana Pro inserts a `thought` step ahead of the output.
* The bytes are JPEG whatever the request asked for. Name the file accordingly.
* There is no top-level `data` array and no `b64_json` field. If you port a parser from `/v1/images`, this is the line to change.
```json theme={"system"}
{
"id": "int_01k9v0f2m8e0",
"object": "interaction",
"model": "gemini-3.1-flash-image",
"status": "completed",
"created": 1786636800,
"updated": 1786636809,
"service_tier": "default",
"usage": { "...": "accounting for the call" },
"steps": [
{ "signature": "CtwBAdHtim8yq0pQ7f..." },
{
"content": [
{ "data": "/9j/4AAQSkZJRgABAQAAAQABAAD..." }
]
}
]
}
```
## Example
### Generate
One POST returns the image inline. The example takes the first `steps[]` content part that has a `data` field and saves it as a JPEG.
```bash cURL theme={"system"}
# The image is the first steps[] content part with a data field. There is no
# data[].b64_json here, and no fixed index: the step layout varies by model.
curl -s https://api.nunchux.ai/v1/google/v1beta/interactions \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "gemini-3.1-flash-image",
"input": [{ "type": "text", "text": "a nano banana dessert, studio lighting" }],
"response_format": { "type": "image", "image_size": "1K", "aspect_ratio": "16:9" }
}' | jq -r '[.steps[].content[]? | select(.data)][0].data' | base64 -d > output.jpg
```
```python Python theme={"system"}
# pip install requests
import base64, os, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# Synchronous: the image is in the response body, so there is no polling.
resp = requests.post(
f"{BASE}/v1/google/v1beta/interactions",
headers=HEADERS,
json={
"model": "gemini-3.1-flash-image",
"input": [{"type": "text", "text": "a nano banana dessert, studio lighting"}],
"response_format": {"type": "image", "image_size": "1K", "aspect_ratio": "16:9"},
},
)
resp.raise_for_status()
# The envelope is {id, status, usage, created, updated, service_tier, steps, ...},
# not data[].b64_json. The steps[] layout varies (the signature step is optional,
# thinking models add a thought step), so take the first part that has data.
steps = resp.json()["steps"]
b64 = next(p["data"] for s in steps for p in s.get("content", []) if p.get("data"))
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(b64))
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const BASE = "https://api.nunchux.ai";
const HEADERS = {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
};
// Synchronous: the image is in the response body, so there is no polling.
const res = await fetch(`${BASE}/v1/google/v1beta/interactions`, {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
model: "gemini-3.1-flash-image",
input: [{ type: "text", text: "a nano banana dessert, studio lighting" }],
response_format: { type: "image", image_size: "1K", aspect_ratio: "16:9" },
}),
});
if (!res.ok) throw new Error(`request failed: HTTP ${res.status} ${await res.text()}`);
const body = await res.json();
// The envelope is { id, status, usage, created, updated, service_tier, steps, ... },
// not data[].b64_json. The steps[] layout varies (the signature step is optional,
// thinking models add a thought step), so take the first part that has data.
const b64 = body.steps.flatMap((s) => s.content ?? []).find((p) => p.data).data;
await writeFile("output.jpg", Buffer.from(b64, "base64"));
```
### Edit
The same call with image parts. The prompt places the person from the first image into the room from the second. Omit `aspect_ratio` to keep the shape of the first input. `image_size` still applies, and it still bills at `4K` when you omit it.
```bash cURL theme={"system"}
# Images are sent inline as base64. There is no URL form.
PERSON=$(base64 < person.jpg | tr -d '\n')
ROOM=$(base64 < room.jpg | tr -d '\n')
curl -s https://api.nunchux.ai/v1/google/v1beta/interactions \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "gemini-3.1-flash-image",
"input": [
{ "type": "text", "text": "put the person from the first image into the room from the second, keep the lighting natural" },
{ "type": "image", "data": "'"$PERSON"'", "mime_type": "image/jpeg" },
{ "type": "image", "data": "'"$ROOM"'", "mime_type": "image/jpeg" }
],
"response_format": { "type": "image", "image_size": "2K", "aspect_ratio": "3:4" }
}' | jq -r '[.steps[].content[]? | select(.data)][0].data' | base64 -d > output.jpg
```
```python Python theme={"system"}
# pip install requests
import base64, os, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
def b64_file(path):
with open(path, "rb") as f:
return base64.b64encode(f.read()).decode()
# Images are sent inline as base64. There is no URL form.
resp = requests.post(
f"{BASE}/v1/google/v1beta/interactions",
headers=HEADERS,
json={
"model": "gemini-3.1-flash-image",
"input": [
{"type": "text", "text": "put the person from the first image into the room from the second, keep the lighting natural"},
{"type": "image", "data": b64_file("person.jpg"), "mime_type": "image/jpeg"},
{"type": "image", "data": b64_file("room.jpg"), "mime_type": "image/jpeg"},
],
"response_format": {"type": "image", "image_size": "2K", "aspect_ratio": "3:4"},
},
)
resp.raise_for_status()
steps = resp.json()["steps"]
b64 = next(p["data"] for s in steps for p in s.get("content", []) if p.get("data"))
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(b64))
```
```javascript JavaScript theme={"system"}
import { readFile, writeFile } from "node:fs/promises";
const BASE = "https://api.nunchux.ai";
const HEADERS = {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
};
const b64File = async (path) => (await readFile(path)).toString("base64");
// Images are sent inline as base64. There is no URL form.
const res = await fetch(`${BASE}/v1/google/v1beta/interactions`, {
method: "POST",
headers: HEADERS,
body: JSON.stringify({
model: "gemini-3.1-flash-image",
input: [
{ type: "text", text: "put the person from the first image into the room from the second, keep the lighting natural" },
{ type: "image", data: await b64File("person.jpg"), mime_type: "image/jpeg" },
{ type: "image", data: await b64File("room.jpg"), mime_type: "image/jpeg" },
],
response_format: { type: "image", image_size: "2K", aspect_ratio: "3:4" },
}),
});
if (!res.ok) throw new Error(`request failed: HTTP ${res.status} ${await res.text()}`);
const body = await res.json();
const b64 = body.steps.flatMap((s) => s.content ?? []).find((p) => p.data).data;
await writeFile("output.jpg", Buffer.from(b64, "base64"));
```
## Tips
* Write the prompt as an instruction. Name what must change and what must stay.
* When you send several reference images, say which element comes from which image.
* Set `image_size` explicitly. An omitted size is billed at the `4K` tier, on edits as well as on generations.
* Use a larger `image_size` only when you need the detail. Larger sizes cost more and take longer.
* Set `aspect_ratio` when you want a shape other than `1:1`. Size and shape are separate axes.
* Use Nano Banana Pro for the highest fidelity. Use Nano Banana 2 for fast everyday work.
* Read the image from the first `steps[]` entry that has a `data` part. Do not hardcode a step index.
* Decode and save the image on receipt. The response carries the bytes, not a URL.
## Errors and limits
* A request takes up to 14 image parts. Images are sent inline as base64, with no URL form.
* A request that omits `response_format` or `response_format.type` is rejected with `no_image_format` or `invalid_response_format`. A rejected request is not charged.
* `512` is accepted on `gemini-3.1-flash-image` only. `0.5K` is rejected on every model.
* Nano Banana is priced per image. The price rises with `image_size`, and an omitted `image_size` bills at the `4K` tier. The rates are on the [pricing page](https://nunchux.ai/pricing). See [Credits and pricing](/credits-pricing).
* A Nano Banana call counts toward your requests per minute. It does not take a simultaneous-job slot. See [Rate limits](/rate-limits).
* A `502` after the request was sent is refunded automatically. It is safe to retry.
* A `503` with the code `google_unreachable` means that Google could not be reached before anything was sent. Nothing was charged. Retry with backoff.
* A `503` whose message says `google pass-through not enabled` means that the route is not available. Do not retry it in a loop.
The full error contract, the retry rules and the per-plan caps are on the [Errors](/errors) and [Rate limits](/rate-limits) pages. How every partner model bills and refunds is on the [Partner models overview](/partner-models/overview).
# Overview
Source: https://docs.nunchux.ai/partner-models/overview
Third-party models served through your Nunchux key. Each provider keeps its own routes, request shapes and job lifecycle.
# Partner models
Partner models are third-party models that you call with your Nunchux API key. Each provider has its own route family on `api.nunchux.ai`. The request and response shapes follow the provider's own API. Code written against the provider works after you change the base URL and the key.
One key and one credit balance cover every provider. You do not need an account with the provider.
Partner models do not use the Nunchux performance tiers. Each model is priced on its own terms. See the model page.
## Models
| Model | Provider | Modality | Tasks | Request type |
| - | - | - | - | - |
| [Nano Banana 2, Nano Banana 2.1, Nano Banana Pro](/partner-models/nano-banana) | Google | Image | Text-to-image, image-to-image | Synchronous |
| [Veo 3.1, Veo 3.1 Fast, Veo 3.1 Lite](/partner-models/veo) | Google | Video | Text-to-video, image-to-video | Asynchronous |
| [Kling V3, Kling V3 Omni](/partner-models/kling) | Kling AI | Video | Text-to-video, image-to-video, omni video, motion control | Asynchronous |
| [Wan 3.0, Wan 3.0 Prime, Wan 2.7](/partner-models/wan) | Alibaba Cloud | Video | Text-to-video, image-to-video, reference-to-video | Asynchronous |
| [HappyHorse 1.1, HappyHorse 1.0](/partner-models/happyhorse) | Alibaba Cloud | Video | Text-to-video, image-to-video, reference-to-video | Asynchronous |
| [Seedance 2.5](/partner-models/seedance) | ByteDance | Video | Text-to-video, image-to-video | Asynchronous |
| [MiniMax H3](/partner-models/minimax) | MiniMax | Video | Text-to-video, image-to-video, reference-to-video | Asynchronous |
| [Avatar 4 Photo, Avatar 5 Digital](/partner-models/heygen) | HeyGen | Avatar video | Talking avatar from a photo or a HeyGen avatar | Asynchronous |
Each model page lists the model IDs to send.
## Choosing a model
Text-to-video. Veo 3.1, Seedance 2.5 and HappyHorse generate a soundtrack with the picture. Wan 3.0 holds one shot for up to 30 seconds. Kling V3 and MiniMax H3 cover the same task.
Image-to-video. Every video family animates a starting image. Wan 2.7 also takes a driving audio track. Wan 3.0 and Wan 2.7 take an optional end frame.
Reference-to-video. Wan, HappyHorse and MiniMax H3 carry a subject from reference images into new footage. Wan 3.0 takes up to 10 reference images.
Multi-reference and storyboard. Kling V3 Omni composes one clip from reference images, videos and elements, or a multi-shot sequence.
Motion transfer. Kling V3 motion control drives a still character with the motion of a reference video.
Talking avatar. HeyGen turns a photo or a HeyGen avatar into a lip-synced video, driven by audio or by a script in a chosen voice.
Image generation and editing. Nano Banana generates and edits images in one synchronous call.
## Endpoints
Each family has one submit route. The asynchronous families also have one poll route.
| Family | Submit | Poll |
| - | - | - |
| [Nano Banana](/partner-models/nano-banana) | `POST /v1/google/v1beta/interactions` | None. The image is in the response. |
| [Veo](/partner-models/veo) | `POST /v1/google/v1beta/models/{model}:predictLongRunning` | `GET /v1/google/v1beta/operations/{handle}` |
| [Kling](/partner-models/kling) | `POST /v1/klingai/videos/{capability}` | `GET /v1/klingai/videos/{capability}/{task_id}` |
| [Wan](/partner-models/wan), [HappyHorse](/partner-models/happyhorse) | `POST /v1/alibaba/services/aigc/video-generation/video-synthesis` | `GET /v1/alibaba/tasks/{task_id}` |
| [Seedance](/partner-models/seedance) | `POST /v1/bytedance/contents/generations/tasks` | `GET /v1/bytedance/contents/generations/tasks/{task_id}` |
| [MiniMax](/partner-models/minimax) | `POST /v1/minimax/v2/video_generation` | `GET /v1/minimax/v2/query/video_generation/{task_id}` |
| [HeyGen](/partner-models/heygen) | `POST /v1/heygen/v3/videos` | `GET /v1/heygen/v3/videos/{video_id}` |
Kling's `{capability}` is one of `text2video`, `image2video`, `omni-video` and `motion-control`. Poll the same capability that you submitted to.
## How a request works
Nano Banana is synchronous. Send one POST and read the image from the response.
Every video family is asynchronous. A request has three steps.
1. Submit. Send a POST to the submit route with your API key in the `X-API-Key` header. The response contains a task identifier. A 200 status means that the provider accepted the job. It does not mean that the job is complete.
2. Poll. Send a GET to the poll route with the task identifier. Repeat until the status is terminal. Wait a few seconds between polls. Each provider uses its own status field and its own values.
3. Download. When the status is the success value, the poll response contains the URL of the output. Output URLs expire, some within an hour. Download the output as soon as the job succeeds.
| Family | Status field | Success | Failure |
| - | - | - | - |
| Veo | `done` | `true`, with `response` | `true`, with `error` |
| Kling | `task_status` | `succeed` | `failed` |
| Wan, HappyHorse | `output.task_status` | `SUCCEEDED` | `FAILED`, `CANCELED` |
| Seedance | `status` | `succeeded` | `failed`, `cancelled`, `expired` |
| MiniMax | `task.status` | `succeeded` | `failed`, `cancelled` |
| HeyGen | `data.status` | `completed` | `failed` |
Any other value means that the job is still running. Keep polling. Each model page shows a full submit, poll and download example.
## Billing
Each partner model is priced on its own terms, such as resolution, duration or mode. The rates are on the model page and on the [pricing page](https://nunchux.ai/pricing).
Nunchux deducts the credits when it accepts a job. When the job fails, Nunchux refunds the credits. See [Credits & pricing](/credits-pricing).
## Errors and limits
A failed submit returns an HTTP error status. See [Error codes](/errors).
A failed job does not return an HTTP error. The submit returns 200, and the failure appears later as the failure status in the poll response. The poll response also carries a message that explains the failure. Read it before you resubmit.
When the provider cannot read an input that you referenced, such as an image URL that is not publicly reachable, the submit can return 200 with a provider error code in the body. The submit was billed. Correct the input, then submit again.
Every job counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits).
# Seedance
Source: https://docs.nunchux.ai/partner-models/seedance
ByteDance's Seedance video model for text-to-video and image-to-video with generated audio.
# Seedance
Seedance is ByteDance's video model line. Each version generates a clip from a text prompt or from a starting image, and can score the clip with generated audio. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors ByteDance's own API, so the request and response shapes are the provider's.
Best for
* Prompt-driven scenes. A usable clip from a written description, with no shoot and no source footage.
* Animating a still. A first frame fixes the composition and the prompt drives the motion.
* Clips with sound. The model generates a track with the picture unless you turn it off.
## Models
| Model | Model ID | Tasks | Resolution | Duration |
| - | - | - | - | - |
| Seedance 2.5 | `dreamina-seedance-2-5-260628` | t2v, i2v | 480p, 720p | Up to 15 s |
One model ID serves both tasks. The items in `content[]` select the task. A `text` item alone gives text-to-video. A `text` item plus a `first_frame` image item gives image-to-video.
On image-to-video the output shape follows the first frame. Crop the image to the shape you want before you upload it.
## Endpoints
| Step | Route |
| - | - |
| Submit | `POST /v1/bytedance/contents/generations/tasks` |
| Poll | `GET /v1/bytedance/contents/generations/tasks/{task_id}` |
The poll route is the submit route with the task ID appended.
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. The examples also send a `User-Agent` header that names your application.
The model ID: `dreamina-seedance-2-5-260628`.
The prompt and, on image-to-video, the first frame. Put the `text` item first.
`text` for the prompt. `image_url` for an image.
The prompt. Required on a `text` item. All tasks. Describe the subject, the action and the camera move.
On an `image_url` item. Set `first_frame` for the image the clip starts on. Image-to-video.
On an `image_url` item. The image container.
The URL of the image. The provider fetches it, so the URL must be reachable from the internet.
The output tier: `480p` or `720p`. All tasks. A larger frame uses more video tokens.
Send `false` for a silent clip. All tasks. The flag is a billing dimension: a clip with sound and a silent clip are priced differently.
## Response
### Submit
A 200 status means that the provider accepted the task. It does not mean that the clip is ready.
The task ID. It starts with `cgt-`. Poll `GET /v1/bytedance/contents/generations/tasks/{task_id}` with it.
### Poll
`queued` or `running` while the task runs. `succeeded`, `failed`, `cancelled` or `expired` when it ends. Poll until you read one of the four terminal values.
The URL of the clip. Present only when `status` is `succeeded`. The URL expires. Download the clip as soon as the task succeeds.
The failure detail. Present when the task did not succeed. Read it before you resubmit.
The number of video tokens the clip used. Present when `status` is `succeeded`. Seedance is billed by this count, so the cost is known when the task completes.
## Example
Submit, poll, download. The submit and the poll use the same path. The poll appends the task ID.
```bash cURL theme={"system"}
BASE=https://api.nunchux.ai
TASKS=$BASE/v1/bytedance/contents/generations/tasks
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$TASKS" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "dreamina-seedance-2-5-260628",
"content": [
{ "type": "text", "text": "a fox trotting through a snowy forest at dawn" }
],
"resolution": "720p",
"generate_audio": false
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
TASK=$(jq -er '.id' <<<"$SUBMIT") || { echo "no id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until the status is terminal
DELAY=5
for _ in $(seq 1 90); do
RESP=$(curl -sS "$TASKS/$TASK" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.status // ""' <<<"$RESP")
case "$STATUS" in succeeded|failed|cancelled|expired) break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
# 3. Check for failure
[ "$STATUS" = "succeeded" ] || { echo "task $STATUS: $(jq -r '.error // "no error"' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o seedance.mp4 "$(jq -r '.content.video_url' <<<"$RESP")"
```
```python Python theme={"system"}
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
TASKS = f"{BASE}/v1/bytedance/contents/generations/tasks"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
TASKS,
headers=HEADERS,
json={
"model": "dreamina-seedance-2-5-260628",
"content": [
{"type": "text", "text": "a fox trotting through a snowy forest at dawn"}
],
"resolution": "720p",
"generate_audio": False,
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
task = submit.json()["id"]
# 2. Poll until the status is terminal
delay = 5
for _ in range(90):
data = requests.get(f"{TASKS}/{task}", headers=HEADERS).json()
status = data.get("status", "")
if status in ("succeeded", "failed", "cancelled", "expired"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("Seedance task did not reach a terminal status")
# 3. Check for failure
if status != "succeeded":
raise RuntimeError(f"Seedance task {status}: {data.get('error')}")
# 4. Download
with requests.get(data["content"]["video_url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("seedance.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
To animate a still, append an item with `type` set to `image_url` and `role` set to `first_frame`, with the image URL under `image_url.url`. The body below replaces the submit body in the example.
```json Image-to-video theme={"system"}
{
"model": "dreamina-seedance-2-5-260628",
"content": [
{ "type": "text", "text": "the fox turns and trots toward the camera" },
{
"type": "image_url",
"role": "first_frame",
"image_url": { "url": "https://example.com/frame.jpg" }
}
],
"resolution": "720p",
"generate_audio": false
}
```
## Tips
* Describe motion and camera, not just the scene. A static description gives the model nothing to animate.
* Keep the prompt to one clear action. Busy, multi-event prompts are harder to render cleanly.
* Draft at 480p to settle the wording, then re-run the same prompt at 720p. A smaller frame uses fewer video tokens.
* Decide about audio before you run. The flag changes the price, so a silent draft is the cheaper way to test a prompt.
* Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.
## Errors and limits
A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with `status` set to `failed`, `cancelled` or `expired`, and `error` explains why. Read it before you resubmit.
A failed submit returns an HTTP error status. See [Error codes](/errors).
A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned an `id`, poll it. If it returned nothing, check your credits before you try again.
ByteDance applies these limits to each first or last frame image you attach:
* 300 to 6,000 pixels on each side
* An aspect ratio, width over height, between 0.4 and 2.5
* Under 30 MB per file
Seedance is billed by the number of video tokens the finished clip uses. The charge is known when the task completes, not before it starts. A longer or larger clip uses more tokens, and the audio flag changes the rate. Every task counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits), [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
# Veo 3.1
Source: https://docs.nunchux.ai/partner-models/veo
Generate video with native audio from a text prompt or a starting image with Google's Veo 3.1 models.
# Veo 3.1
Veo 3.1 is Google's video model. It generates a clip with native audio from a text prompt or from a starting image. The call is asynchronous: submit a job, poll the operation until `done` is `true`, then download the clip. The request and response shapes mirror Google's own API.
## Models
Veo 3.1 ships in three tiers. Each tier is a model ID that you name in the submit path. The request body is identical across tiers.
| Model | Tier | Model ID | Tasks | Resolution | Duration |
| - | - | - | - | - | - |
| Veo 3.1 | Standard | `veo-3.1-generate-preview` | Text-to-video, image-to-video | 720p, 1080p, 4K | 4, 6 or 8 s |
| Veo 3.1 | Fast | `veo-3.1-fast-generate-preview` | Text-to-video, image-to-video | 720p, 1080p, 4K | 4, 6 or 8 s |
| Veo 3.1 | Lite | `veo-3.1-lite-generate-preview` | Text-to-video, image-to-video | 720p, 1080p | 4, 6 or 8 s |
Every tier generates native audio with the picture.
Use Standard when quality matters most. It has the highest fidelity.
Use Fast for iteration. It is quicker and cheaper than Standard, and it still reaches 4K.
Use Lite for the cheapest runs. It is the most economical tier. It stops at 1080p, so it does not accept a 4K request.
## Endpoints
The base URL is `https://api.nunchux.ai`. `{model}` is the model ID of the tier. `{handle}` is the `name` that the submit returns.
| Step | Method | Path |
| - | - | - |
| Submit | `POST` | `/v1/google/v1beta/models/{model}:predictLongRunning` |
| Poll | `GET` | `/v1/google/v1beta/operations/{handle}` |
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json` on the submit. Send a `User-Agent` header that names your application, for example `YourApp/1.0`.
The body carries a single `instances` entry (the prompt, plus a starting `image` for image-to-video) and optional `parameters`.
One generation request. Veo takes a single instance.
Description of the video to generate.
The starting frame, for image-to-video. Omit it for text-to-video.
Base64-encoded image bytes, with no `data:` prefix.
The MIME type of the image, for example `image/jpeg` or `image/png`.
Generation controls.
Output resolution. `4k` is available on Standard and Fast only. Lite stops at `1080p`.
Options: `720p`, `1080p`, `4k`
Clip length in whole seconds. Billed per output-second.
Options: `4`, `6`, `8`
Image input
Send `bytesBase64Encoded` with its `mimeType`. Google's own Veo REST documentation shows an `inlineData` wrapper. Do not copy it here. Send `bytesBase64Encoded`.
## Response
The submit returns only the operation handle. Poll it until `done` is `true`. On success, the video URL is under `response.generateVideoResponse.generatedSamples[0].video.uri`.
### Submit
The operation handle. Poll `GET /v1/google/v1beta/operations/{handle}` with it. The handle is opaque and variable in length (about 260 characters), with no `operations/` prefix. Store it as an unbounded string, not in a fixed-width column sized from the sample.
```json theme={"system"}
{
"name": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1byI6IjNmOWMyYTdlODFiNGQ2MDUiLCJvcCI6Im1vZGVscy92ZW8tMy4xLWdlbmVyYXRlLXByZXZpZXcvb3BlcmF0aW9ucy85bG5rY2Y5enZ6d20iLCJtIjoidmlkZW8iLCJleHAiOjE3NTY2OTU2MDAsImlhdCI6MTc1NjA5MDgwMH0.QmyRdv9C0cALtiL7GKEn619WTM3IGg1hpwkQzJSLJtc"
}
```
### Poll
The operation handle, echoed back.
`false` while the job runs. `true` once the video is ready or the job failed. Check `error` before you read `response`.
Present once `done` is `true` and the job failed. `response` is then absent. Carries a `message` that describes the failure.
Present once `done` is `true` and the job succeeded.
The generated clips.
URL of the output. It expires about an hour after the operation completes. Download it promptly.
The status of the job follows from `done` and from which of `error` and `response` is present.
| `done` | Present | Meaning |
| - | - | - |
| `false` | neither | The job is still running. Keep polling. |
| `true` | `response` | The job succeeded. Download `video.uri` now. |
| `true` | `error` | The job failed. `error.message` explains why. |
```json theme={"system"}
{
"name": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJ1byI6IjNmOWMyYTdlODFiNGQ2MDUiLCJvcCI6Im1vZGVscy92ZW8tMy4xLWdlbmVyYXRlLXByZXZpZXcvb3BlcmF0aW9ucy85bG5rY2Y5enZ6d20iLCJtIjoidmlkZW8iLCJleHAiOjE3NTY2OTU2MDAsImlhdCI6MTc1NjA5MDgwMH0.QmyRdv9C0cALtiL7GKEn619WTM3IGg1hpwkQzJSLJtc",
"done": true,
"response": {
"generateVideoResponse": {
"generatedSamples": [
{ "video": { "uri": "https:///.mp4?" } }
]
}
}
}
```
## Example
Submit, poll, download. Swap the model ID in the submit path for the tier that you want. The poll starts at 5 s and backs off to 30 s.
```bash cURL theme={"system"}
BASE=https://api.nunchux.ai
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/google/v1beta/models/veo-3.1-generate-preview:predictLongRunning" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"instances": [{ "prompt": "a fox trotting through a snowy forest at dawn" }],
"parameters": { "resolution": "1080p", "durationSeconds": 8 }
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
OP=$(jq -er '.name' <<<"$SUBMIT") || { echo "no operation name in: $SUBMIT" >&2; exit 1; }
# 2. Poll until done is true
DELAY=5
for _ in $(seq 1 60); do
RESP=$(curl -sS "$BASE/v1/google/v1beta/operations/$OP" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
[ "$(jq -r '.done // false' <<<"$RESP")" = "true" ] && break
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
[ "$(jq -r '.done // false' <<<"$RESP")" = "true" ] || { echo "timed out" >&2; exit 1; }
# 3. Check for failure. done is true on failure too.
jq -e '.error' >/dev/null 2>&1 <<<"$RESP" && { echo "failed: $(jq -c '.error' <<<"$RESP")" >&2; exit 1; }
# 4. Download. The URL expires about an hour after completion.
curl -sSL -o veo.mp4 "$(jq -r '.response.generateVideoResponse.generatedSamples[0].video.uri' <<<"$RESP")"
```
```python Python theme={"system"}
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
model = "veo-3.1-generate-preview"
submit = requests.post(
f"{BASE}/v1/google/v1beta/models/{model}:predictLongRunning",
headers=HEADERS,
json={
"instances": [{"prompt": "a fox trotting through a snowy forest at dawn"}],
"parameters": {"resolution": "1080p", "durationSeconds": 8},
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
op = submit.json()["name"]
# 2. Poll until done is true
delay = 5
for _ in range(60):
data = requests.get(f"{BASE}/v1/google/v1beta/operations/{op}", headers=HEADERS).json()
if data.get("done"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("Veo operation did not finish")
# 3. Check for failure. done is true on failure too.
if "error" in data:
raise RuntimeError(f"Veo failed: {data['error']}")
# 4. Download. The URL expires about an hour after completion.
url = data["response"]["generateVideoResponse"]["generatedSamples"][0]["video"]["uri"]
with requests.get(url, stream=True, timeout=300) as r:
r.raise_for_status()
with open("veo.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const BASE = "https://api.nunchux.ai";
const HEADERS = {
"X-API-Key": process.env.NUNCHUX_API_KEY,
"Content-Type": "application/json",
"User-Agent": "YourApp/1.0",
};
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
// 1. Submit
const model = "veo-3.1-generate-preview";
const submitRes = await fetch(
`${BASE}/v1/google/v1beta/models/${model}:predictLongRunning`,
{
method: "POST",
headers: HEADERS,
body: JSON.stringify({
instances: [{ prompt: "a fox trotting through a snowy forest at dawn" }],
parameters: { resolution: "1080p", durationSeconds: 8 },
}),
}
);
if (!submitRes.ok) throw new Error(`submit failed: HTTP ${submitRes.status} ${await submitRes.text()}`);
const op = (await submitRes.json()).name;
// 2. Poll until done is true
let data;
let delay = 5_000;
for (let i = 0; i < 60; i++) {
data = await fetch(`${BASE}/v1/google/v1beta/operations/${op}`, {
headers: HEADERS,
}).then((r) => r.json());
if (data.done) break;
await sleep(delay);
delay = Math.min(delay * 2, 30_000);
}
if (!data?.done) throw new Error("Veo operation did not finish");
// 3. Check for failure. done is true on failure too.
if (data.error) throw new Error(`Veo failed: ${JSON.stringify(data.error)}`);
// 4. Download. The URL expires about an hour after completion.
const uri = data.response.generateVideoResponse.generatedSamples[0].video.uri;
const res = await fetch(uri);
if (!res.ok) throw new Error(`download failed: HTTP ${res.status}`);
await writeFile("veo.mp4", Buffer.from(await res.arrayBuffer()));
```
## Tips
* Poll, do not busy-wait. Veo usually takes 1 to 4 min. Poll every 5 to 10 s at first, then back off to 30 s.
* Polls are free, and they never take a simultaneous-job slot.
* Set client timeouts to 10 min or more. Render time scales with clip length.
* Check `error` before you read `response`. `done` is `true` on failure too.
* Download the clip as soon as `done` is `true`. The output URL expires about an hour after the operation completes, so do not store the URL.
* Store `name` as an unbounded string, and poll with that exact value.
* Pick `durationSeconds` deliberately. You are billed per output-second, so 4, 6 and 8 s each cost differently.
* Use Lite for cheap iteration and Standard when fidelity matters. 4K is available on Standard and Fast only.
* Do not double-submit. A submit reserves credits, so resubmitting a running job charges twice. If the submit returned a `name`, poll it.
## Errors and limits
* A `200` on the submit is acceptance, not success. A failed job does not return an HTTP error. It appears in the poll as `done: true` with `error`, and `error.message` explains the failure. Read it before you resubmit.
* `durationSeconds` is `4`, `6` or `8`. `resolution` is `720p`, `1080p` or `4k`, and `4k` is available on Standard and Fast only. A parameter that is out of range, such as a Lite request for `4k`, is rejected with a `400`. A `400` is never charged.
* Veo is priced per output-second times the duration, reserved at submit. A job that ends in failure is refunded automatically. The rates are on the [pricing page](https://nunchux.ai/pricing). See [Credits and pricing](/credits-pricing).
* A submit takes a simultaneous-job slot. Polls do not. See [Rate limits](/rate-limits).
* A `5xx` or a timeout on the submit does not tell you whether the job was created, so a bare resubmit can bill twice. If the submit returned a `name`, poll it. If it returned nothing, check `GET /v1/credits` before you try again. A charge with no delivered job is refunded automatically.
* A `5xx` on a poll is safe to retry with backoff. Polls change nothing.
* A `503` with the code `google_unreachable` means that Google could not be reached before anything was sent. Nothing was charged. Retry with backoff.
* A `503` whose message says `google pass-through not enabled` means that the route is not available. Do not retry it in a loop.
The full error contract, the retry rules and the per-plan caps are on the [Errors](/errors) and [Rate limits](/rate-limits) pages. How every asynchronous partner model submits, polls, bills and refunds is on the [Partner models overview](/partner-models/overview).
# Wan
Source: https://docs.nunchux.ai/partner-models/wan
Alibaba Cloud's Wan video models for text-to-video, image-to-video and reference-to-video through one asynchronous endpoint.
# Wan
Wan is Alibaba Cloud's video model line. Each version generates a clip with a soundtrack from a text prompt, from a still image, or from a set of reference images. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors Alibaba Cloud's own API, so the request and response shapes are the provider's.
Best for
* Longer takes. Wan 3.0 holds a shot for up to 30 seconds in one clip.
* Reference-driven scenes. Up to ten reference images carry a subject into new footage instead of animating one frame.
* Draft passes. Wan 3.0 and Wan 3.0 Prime sell a 480P tier that the earlier versions do not.
* Speed grades. Wan 3.0 Prime returns the same capabilities sooner than the standard grade.
## Models
| Model | Model ID | Tasks | Resolution | Duration | Reference images |
| - | - | - | - | - | - |
| Wan 3.0 | `wan3.0-video` | t2v, i2v, r2v | 480P, 720P, 1080P | 2 to 30 s | Up to 10 |
| Wan 3.0 Prime | `wan3.0-video-prime` | t2v, i2v, r2v | 480P, 720P, 1080P | 2 to 30 s | Up to 10 |
| Wan 2.7 | `wan2.7-t2v`, `wan2.7-i2v`, `wan2.7-r2v` | t2v, i2v, r2v | 720P, 1080P | 2 to 15 s | Up to 5 |
Durations are whole seconds. Wan 3.0 and Wan 3.0 Prime use one model ID for all three tasks. The items in `input.media[]` select the task. No media item gives text-to-video. A `first_frame` item gives image-to-video. `reference_image` items give reference-to-video. Wan 2.7 uses one model ID per task.
Wan 3.0 gives the longest clips and the widest reference set. Use it when a shot has to develop, or when a subject must carry across a new scene. It takes an optional end frame on image-to-video.
Wan 3.0 Prime is the accelerated grade of Wan 3.0. Alibaba Cloud describes it as the high-speed version with capabilities aligned to the standard version. It has the same tasks, tiers, lengths and reference ceiling as Wan 3.0. Use it while you iterate.
Wan 2.7 has a shorter ceiling and a smaller reference set. On image-to-video it takes an optional end frame and a driving audio track.
## Endpoints
| Step | Route |
| - | - |
| Submit | `POST /v1/alibaba/services/aigc/video-generation/video-synthesis` |
| Poll | `GET /v1/alibaba/tasks/{task_id}` |
[HappyHorse](/partner-models/happyhorse) uses the same two routes and the same body shape.
## Request
Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. The examples also send a `User-Agent` header that names your application.
The model ID from the table above. On Wan 3.0 and Wan 3.0 Prime the same ID serves every task.
The prompt and the media items.
The text prompt. All tasks. Describe what moves and how the camera follows it. On reference-to-video, name a reference by its position in `media[]`: `Image 1`, `Image 2` and so on, with a space and a capital letter.
The media items for image-to-video and reference-to-video. Omit it for text-to-video. A first frame and a reference set cannot appear in the same body.
What the item is.
* `first_frame`: the image the clip starts on. Image-to-video. One item.
* `last_frame`: the image the clip ends on. Image-to-video, optional. One item.
* `driving_audio`: a WAV or MP3 track for the clip. Wan 2.7 image-to-video only, optional.
* `reference_image`: a reference image. Reference-to-video. 1 to 10 items on Wan 3.0 and Wan 3.0 Prime, 1 to 5 items on Wan 2.7.
The URL of the file. The provider fetches it, so the URL must be reachable from the internet.
The generation controls. All tasks read all three fields.
The output tier. `480P`, `720P` or `1080P` on Wan 3.0 and Wan 3.0 Prime. `720P` or `1080P` on Wan 2.7. The default is the most expensive tier, so name the one you want.
The clip length in whole seconds. 2 to 30 on Wan 3.0 and Wan 3.0 Prime. 2 to 15 on Wan 2.7. You are billed per second of output.
Send `true` to burn the vendor's mark into the lower-right corner of every frame.
## Response
### Submit
A 200 status means that the provider accepted the task. It does not mean that the clip is ready.
The task ID. Poll `GET /v1/alibaba/tasks/{task_id}` with it.
`PENDING` on a new task.
### Poll
`PENDING` or `RUNNING` while the task runs. `SUCCEEDED`, `FAILED` or `CANCELED` when it ends. Poll until you read one of the three terminal values.
The URL of the clip. Present only when `task_status` is `SUCCEEDED`. The task ID and the URL expire after 24 hours. Download the clip as soon as the task succeeds.
The failure code. Present when the task failed.
The failure message. Read it before you resubmit.
## Example
Submit, poll, download. Swap the `model` value for the version you want.
```bash cURL theme={"system"}
BASE=https://api.nunchux.ai
# 1. Submit
SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/alibaba/services/aigc/video-generation/video-synthesis" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "Content-Type: application/json" \
-H "User-Agent: YourApp/1.0" \
-d '{
"model": "wan3.0-video",
"input": { "prompt": "a fox trotting through a snowy forest at dawn" },
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
TASK=$(jq -er '.output.task_id' <<<"$SUBMIT") || { echo "no task_id in: $SUBMIT" >&2; exit 1; }
# 2. Poll until the status is terminal
DELAY=5
for _ in $(seq 1 90); do
RESP=$(curl -sS "$BASE/v1/alibaba/tasks/$TASK" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-H "User-Agent: YourApp/1.0")
STATUS=$(jq -r '.output.task_status // ""' <<<"$RESP")
case "$STATUS" in SUCCEEDED|FAILED|CANCELED) break ;; esac
sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
done
# 3. Check for failure
[ "$STATUS" = "SUCCEEDED" ] || { echo "task $STATUS: $(jq -r '.output.message // "no message"' <<<"$RESP")" >&2; exit 1; }
# 4. Download
curl -sSL -o wan.mp4 "$(jq -r '.output.video_url' <<<"$RESP")"
```
```python Python theme={"system"}
# pip install requests
import os, time, requests
BASE = "https://api.nunchux.ai"
HEADERS = {
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
"User-Agent": "YourApp/1.0",
}
# 1. Submit
submit = requests.post(
f"{BASE}/v1/alibaba/services/aigc/video-generation/video-synthesis",
headers=HEADERS,
json={
"model": "wan3.0-video",
"input": {"prompt": "a fox trotting through a snowy forest at dawn"},
"parameters": {"resolution": "720P", "duration": 5, "watermark": False},
},
)
if not submit.ok:
raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
task = submit.json()["output"]["task_id"]
# 2. Poll until the status is terminal
delay = 5
for _ in range(90):
data = requests.get(f"{BASE}/v1/alibaba/tasks/{task}", headers=HEADERS).json()
status = data.get("output", {}).get("task_status", "")
if status in ("SUCCEEDED", "FAILED", "CANCELED"):
break
time.sleep(delay)
delay = min(delay * 2, 30)
else:
raise TimeoutError("Wan task did not reach a terminal status")
# 3. Check for failure
if status != "SUCCEEDED":
raise RuntimeError(f"Wan task {status}: {data['output'].get('message')}")
# 4. Download
with requests.get(data["output"]["video_url"], stream=True, timeout=300) as r:
r.raise_for_status()
with open("wan.mp4", "wb") as f:
for chunk in r.iter_content(1 << 14):
f.write(chunk)
```
To animate a still, add a `first_frame` item to `input.media[]`. To carry a subject into a new scene, add `reference_image` items instead. A first frame and a reference set are separate tasks. The bodies below replace the submit body in the example.
```json Image-to-video theme={"system"}
{
"model": "wan3.0-video",
"input": {
"prompt": "the fox turns and trots toward the camera",
"media": [
{ "type": "first_frame", "url": "https://example.com/frame.jpg" }
]
},
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}
```
```json Reference-to-video theme={"system"}
{
"model": "wan3.0-video",
"input": {
"prompt": "the person from Image 1 walks through the doorway from Image 2, slow tracking shot",
"media": [
{ "type": "reference_image", "url": "https://example.com/ref-1.jpg" },
{ "type": "reference_image", "url": "https://example.com/ref-2.jpg" }
]
},
"parameters": { "resolution": "720P", "duration": 5, "watermark": false }
}
```
## Tips
* Say what moves and how the camera follows it. The prompt drives both, and motion is what the model has to generate.
* Name the sounds you want in the prompt when audio matters. The model scores what it reads.
* Send `resolution` and `duration` on every request. The defaults are 1080P and 5 seconds, and 1080P is the most expensive tier.
* Draft at 480P on Wan 3.0, then re-run the prompt at the tier and length you need. You are billed per second of output, so a long 1080P draft is the expensive way to test wording.
* Match the model to the job. A first frame fixes the composition. A reference set fixes the subject and lets the scene change.
* Upload the frame at the tier you intend to render. A small source limits what the model can resolve.
* Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.
## Errors and limits
A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with `task_status` set to `FAILED` or `CANCELED`, and `output.message` explains why. Read the message before you resubmit.
A failed submit returns an HTTP error status. The body carries `code`, `message` and `request_id`. See [Error codes](/errors).
A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned a `task_id`, poll it. If it returned nothing, check your credits before you try again.
Alibaba Cloud applies these limits to each reference image:
* 240 to 8,000 pixels on each side
* An aspect ratio no wider than 8:1
* JPEG, JPG, PNG, BMP or WEBP. A transparent channel is not supported.
* Up to 20 MB per file
Every task counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits). Rates are per second of output and depend on the tier. See [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).
# Performance Tiers
Source: https://docs.nunchux.ai/performance-tiers
Balance speed and cost with the performance tiers of the Nunchux Image API.
# Performance Tiers
### Balance speed and cost with performance tiers
## Available Tiers
### Nunchux Optimized models
The Radical Speed tier delivers the lowest latency, while Radical Value gives the lowest cost per image. Both tiers run on Nunchux's proprietary Model Optimizer and Inference Engine. Together they give the best tradeoff between speed, cost, and quality on the market.
These tiers apply to the FLUX, Qwen and HiDream O1 models. Ideogram 4 has three other tiers: Turbo (12 steps), Balanced (20 steps) and Quality (48 steps). The step count selects the tier, and the price is per image. The LTX video models have two other tiers, 720p and 1080p. The output size selects the tier, and the price is per second of video.
### Partner models
Partner models do not use the Nunchux tiers. Each one prices on its own terms, such as output size, resolution, or mode. See the [pricing page](https://nunchux.ai/pricing) for current per-model rates.
## Tier Comparison
| Tier | Speed | Cost | Quality | Use Case |
| - | - | - | - | - |
| `radical_speed` | Fastest | Low | High | Real-time apps, previews |
| `radical_value` | Fast | Lowest | High | High-volume batch, cost-sensitive jobs |
What tier should I choose?
Use Radical Speed if the response time is important. Use Radical Value if cost per image is more important across large batches.
## Send a Tier in a Request
Set the `tier` field in your API request:
```bash theme={"system"}
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nunchux-qwen-image-2512",
"tier": "radical_speed",
"prompt": "a beautiful mountain landscape",
"width": 1024,
"height": 1024
}'
```
## Tier and Model Combinations
Each tier and model combination has a different use. The Qwen model with the Radical Speed tier gives the lowest latency. Use this combination for real-time previews and interactive applications.
```bash theme={"system"}
# Fastest generation: Qwen Image 2512 + radical_speed
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nunchux-qwen-image-2512",
"tier": "radical_speed",
"prompt": "real-time preview render",
"width": 1024,
"height": 1024
}'
```
A small model such as Klein 4B with the Radical Value tier gives the lowest cost per image. Use this combination for high-volume batch jobs.
```bash theme={"system"}
# Lowest cost: Klein 4B + radical_value
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "X-API-Key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nunchux-flux.2-klein-4b",
"tier": "radical_value",
"prompt": "high-volume batch icon design",
"width": 1024,
"height": 1024
}'
```
Workflow Recommendations
* For development and test, use Radical Speed. It has the lowest latency.
* For high-volume batch jobs, use Radical Value. It has the lowest cost per image.
* For production, use Radical Speed if latency is important. Use Radical Value if cost per image is more important.
## Error Handling
If you send a tier that a model does not support, the API returns an error.
```json theme={"system"}
{
"detail": "unknown tier 'turbo'; allowed values: radical_speed, radical_value"
}
```
Check the model's page to make sure that it supports the tier you send. See the [Error Handling](/errors) guide for more information.
# QuickStart
Source: https://docs.nunchux.ai/quickstart
Get started generating images and videos with the Nunchux API in minutes.
# Quickstart
### Generate your first image in under 5 minutes.
## 1. Create your API key
Before you begin, [sign up](https://nunchux.ai/sign-up) for an account to get an API key.
Once you've registered, create an API key from your [dashboard](https://nunchux.ai/dashboard) and set it as an environment variable.
```bash Export an environment variable theme={"system"}
export NUNCHUX_API_KEY=xxxxx
```
## 2. Make your first request
Generate your first image with a simple API request and save your result.
```bash cURL theme={"system"}
# Replace $NUNCHUX_API_KEY with your API key.
# Change output.jpg to the filename you want.
curl -X POST https://api.nunchux.ai/v1/images/generations \
-H "Content-Type: application/json" \
-H "X-API-Key: $NUNCHUX_API_KEY" \
-d '{
"model": "nunchux-flux.2-klein-4b",
"prompt": "A beautiful sunset over mountains",
"tier": "radical_speed",
"width": 1024,
"height": 1024,
"response_format": "b64_json"
}' | jq -r '.data[0].b64_json // error(.error.message // .detail // tostring)' | base64 -d > output.jpg
```
```python Python theme={"system"}
import base64
import os
import requests
response = requests.post(
"https://api.nunchux.ai/v1/images/generations",
headers={
"Content-Type": "application/json",
"X-API-Key": os.environ["NUNCHUX_API_KEY"],
},
json={
"model": "nunchux-flux.2-klein-4b",
"prompt": "A beautiful sunset over mountains",
"tier": "radical_speed",
"width": 1024,
"height": 1024,
"response_format": "b64_json",
},
)
# Decode the base64 image and save it. Change output.jpg to any filename.
image_b64 = response.json()["data"][0]["b64_json"]
with open("output.jpg", "wb") as f:
f.write(base64.b64decode(image_b64))
```
```javascript JavaScript theme={"system"}
import { writeFile } from "node:fs/promises";
const response = await fetch("https://api.nunchux.ai/v1/images/generations", {
method: "POST",
headers: {
"Content-Type": "application/json",
"X-API-Key": process.env.NUNCHUX_API_KEY,
},
body: JSON.stringify({
model: "nunchux-flux.2-klein-4b",
prompt: "A beautiful sunset over mountains",
tier: "radical_speed",
width: 1024,
height: 1024,
response_format: "b64_json",
}),
});
const data = await response.json();
// Decode the base64 image and save it. Change output.jpg to any filename.
await writeFile("output.jpg", Buffer.from(data.data[0].b64_json, "base64"));
```
Congratulations – you've just made your first call to Nunchux! From here:
* [Nunchux Optimized models](/nunchux-optimized/overview): the models in this example, and the video models.
* [Partner models](/partner-models/overview): third-party image and video models.
* [Performance Tiers](/performance-tiers): what the tier field does, and when to change it.
* [Text-to-Image reference](/nunchux-optimized/text-to-image): all parameters and response formats.
* [Image-to-Image reference](/nunchux-optimized/image-to-image): edit an input image with a prompt.
* [Video](https://nunchux.ai/docs/video-generation/overview): the async submit, poll, download flow — Kling, Veo and HeyGen.
# Rate Limits
Source: https://docs.nunchux.ai/rate-limits
Per-plan caps on requests per minute and simultaneous jobs, which requests count toward each, and how to back off from a 429.
# Rate Limits
### Understand and manage how many requests you can run simultaneously on Nunchux
## Plan Limits
Every API key belongs to a plan. The plan sets two independent caps: requests per minute, and simultaneous jobs.
| Plan | Requests per Minute | Simultaneous Jobs |
| - | - | - |
| Free | 60 | 5 |
| Pro | 200 | 10 |
| Enterprise | 600 | 20 |
## Checking Your Caps
Read your own caps at any time from `GET /v1/credits`, which takes no parameters and returns your own account only. It uses the same API key for [authentication](/authentication) as every other endpoint.
```bash curl theme={"system"}
curl https://api.nunchux.ai/v1/credits \
-H "X-API-Key: $NUNCHUX_API_KEY"
```
```python python theme={"system"}
import os
import requests
response = requests.get(
'https://api.nunchux.ai/v1/credits',
headers={'X-API-Key': os.environ['NUNCHUX_API_KEY']},
timeout=30,
)
response.raise_for_status()
print(response.json())
```
A `200` carries the balance and the caps that apply to your plan. The API returns [standard error codes](/errors) with a JSON error body.
```json theme={"system"}
{
"credits_remaining": 96.49,
"plan_level": "pro",
"rpm_limit": 200,
"concurrent_limit": 10
}
```
| Field | Type | Description |
| - | - | - |
| `credits_remaining` | number | Credits left on the account. One credit is \$1.00. |
| `plan_level` | string | The plan the account is on. Each plan has its own caps. |
| `rpm_limit` | number | Requests per minute this plan allows. |
| `concurrent_limit` | number | Simultaneous jobs this plan allows. |
Polling
This poll is free and it never takes a job slot. It still counts toward your requests per minute.
## What Counts
Every authenticated request counts toward requests per minute, polls included. Only requests that start a generation take a job slot, and how long they hold it depends on the kind of request.
| Request | Requests per Minute | Job slot |
| - | - | - |
| Async submit: Kling, Veo 3.1, HeyGen | Counts | Takes a slot. A rejected submit does not. |
| Synchronous generation: `/v1/images/*`, `/v1/videos/*` | Counts | Holds a slot for the duration of the request. |
| Nano Banana: `/v1/google/v1beta/interactions` | Counts | No slot. The call is synchronous and returns in seconds. |
| Status polls | Counts | No slot. |
| Account reads: `GET /v1/credits`, `GET /v1/discounts`, `GET /v1/heygen/v3/voices` | Counts | No slot. |
## When You Hit a Limit
The API returns a `429` for both caps. The error code says which cap you hit.
### Requests per Minute
```json theme={"system"}
{
"error": {
"code": "rpm_limit_exceeded",
"message": "Rate limit exceeded. See the rate-limit response headers for your limit and reset time."
}
}
```
What to do
Wait, then retry with exponential backoff: 1s, 2s, 4s, 8s. Cap at five attempts. Spreading submits evenly over the minute avoids it altogether.
### Simultaneous Jobs
```json theme={"system"}
{
"error": {
"code": "concurrent_limit_exceeded",
"message": "Concurrent job limit reached. Wait for a running job to finish, then retry."
}
}
```
What to do
Retrying quickly will not help. Wait, poll your running jobs, and submit again once one has completed.
The full error contract is on the [Errors](/errors) page.