> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nunchux.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Authenticate every request with your API key in the `X-API-Key` header or `Authorization: Bearer`. Keys start with `sk-nunchux-`.
> Image generation is synchronous: the image is in the response.
> Video generation is asynchronous: submit the job, poll its task until it reaches a terminal status, then download the output promptly, because output URLs expire.

# MiniMax H3

> MiniMax's H3 video model for text-to-video, image-to-video and reference-to-video through one asynchronous endpoint.

# MiniMax H3

MiniMax H3 is an omni-modal model that reads text, images, video and audio in one context and generates a video clip. One model ID serves all three tasks. The `content[]` items you send decide which task you get. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors MiniMax's own API, so the request and response shapes are the provider's.

<Note>
  Best for

  * Multi-reference scenes. Up to nine reference images, three reference clips and a reference track carry a subject or a style into new footage.
  * 2K output. A higher tier than the other partner video models sell.
  * Animating a still. A first frame fixes the composition and the prompt drives the motion.
  * Longer prompts. The vendor accepts up to 7,000 characters per text item.
</Note>

## Models

| Model | Model ID | Tasks | Resolution | Duration | Reference media |
| - | - | - | - | - | - |
| MiniMax H3 | `MiniMax-H3` | t2v, i2v, r2v | 768P, 2K | 4 to 15 s | Up to 9 images, 3 clips and 1 audio track |

Durations are whole seconds. All three tasks send `MiniMax-H3` as the `model` value. 2K costs more per second than 768P.

| Task | Resolution | Duration | Aspect ratio |
| - | - | - | - |
| Text-to-video | 768P, 2K | 4 to 15 s | Required: 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16 |
| Image-to-video | 768P, 2K | 4 to 15 s | Follows the source image |
| Reference-to-video | 768P, 2K | 4 to 15 s | Follows the reference set |

The `content[]` items select the task. A `text` item alone gives text-to-video. A `first_frame` image item gives image-to-video. `reference_image`, `reference_video` and `reference_audio` items give reference-to-video.

A first frame and a reference set cannot appear in the same body. The two use different `role` values, and a body that mixes the role sets is rejected. Pick one task per request.

## Endpoints

| Step | Route |
| - | - |
| Submit | `POST /v1/minimax/v2/video_generation` |
| Poll | `GET /v1/minimax/v2/query/video_generation/{task_id}` |

## Request

Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. The examples also send a `User-Agent` header that names your application.

<ParamField body="model" type="string" required>
  The model ID: `MiniMax-H3`.
</ParamField>

<ParamField body="content" type="array" required>
  The prompt and the media items. Put the `text` item first. Pair each role with its own container: an image rides `image_url`, a clip rides `video_url`, a track rides `audio_url`. A role on the wrong container is accepted without complaint and then does not do what you meant.

  <Expandable title="item properties">
    <ParamField body="type" type="string" required>
      `text` for the prompt. `image_url` for an image. `video_url` for a clip. `audio_url` for a track.
    </ParamField>

    <ParamField body="text" type="string">
      The prompt, up to 7,000 characters. Required on a `text` item. All tasks.
    </ParamField>

    <ParamField body="role" type="string">
      What a media item is for.

      * `first_frame`: the image the clip starts on. Image-to-video. One item.
      * `last_frame`: the image the clip ends on. Image-to-video, optional. One item.
      * `reference_image`: a reference image. Reference-to-video. Up to 9 items.
      * `reference_video`: a reference clip. Reference-to-video. Up to 3 items, 2 to 15 seconds each, 15 seconds in total.
      * `reference_audio`: a reference track. Reference-to-video, optional. One item.
    </ParamField>

    <ParamField body="image_url" type="object">
      On an `image_url` item. Holds `url`, the URL of the image.
    </ParamField>

    <ParamField body="video_url" type="object">
      On a `video_url` item. Holds `url`, the URL of the clip.
    </ParamField>

    <ParamField body="audio_url" type="object">
      On an `audio_url` item. Holds `url`, the URL of the track.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField body="resolution" type="string">
  The output tier: `768P` or `2K`. All tasks.
</ParamField>

<ParamField body="duration" type="integer" required>
  The clip length in whole seconds, 4 to 15. All tasks. You are billed per second of output.
</ParamField>

<ParamField body="ratio" type="string">
  The output shape. Required on a text-only body: `21:9`, `16:9`, `4:3`, `1:1`, `3:4` or `9:16`. `adaptive` is rejected there. Leave it out when the body carries an image, a clip or a track. The output then follows the source.
</ParamField>

## Response

### Submit

A 200 status means that the provider accepted the task. It does not mean that the clip is ready.

<ResponseField name="task_id" type="string" required>
  The task ID. Poll `GET /v1/minimax/v2/query/video_generation/{task_id}` with it.
</ResponseField>

### Poll

<ResponseField name="task.status" type="string" required>
  `queued` or `running` while the task runs. `succeeded`, `failed` or `cancelled` when it ends. Poll until you read one of the three terminal values.
</ResponseField>

<ResponseField name="task.content.url" type="string">
  The URL of the clip. Present only when `status` is `succeeded`. The URL is time-limited. Download the clip as soon as the task succeeds. A later poll returns a fresh URL. A task can be polled for 7 days.
</ResponseField>

<ResponseField name="task.error" type="object">
  The failure detail. Present when the task did not succeed. Read it before you resubmit.
</ResponseField>

<ResponseField name="task.usage" type="object">
  The billed quantities: output seconds, reference-video seconds and the billed image count.
</ResponseField>

## Example

Submit, poll, download. The submit and the poll use different paths.

<CodeGroup>
  ```bash cURL theme={"system"}
  BASE=https://api.nunchux.ai

  # 1. Submit
  SUBMIT=$(curl -sS --fail-with-body "$BASE/v1/minimax/v2/video_generation" \
    -H "X-API-Key: $NUNCHUX_API_KEY" \
    -H "Content-Type: application/json" \
    -H "User-Agent: YourApp/1.0" \
    -d '{
      "model": "MiniMax-H3",
      "content": [
        { "type": "text", "text": "a fox trotting through a snowy forest at dawn" }
      ],
      "resolution": "768P",
      "duration": 6,
      "ratio": "16:9"
    }') || { echo "submit failed: $SUBMIT" >&2; exit 1; }
  TASK=$(jq -er '.task_id' <<<"$SUBMIT") || { echo "no task_id in: $SUBMIT" >&2; exit 1; }

  # 2. Poll until the status is terminal
  DELAY=5
  for _ in $(seq 1 90); do
    RESP=$(curl -sS "$BASE/v1/minimax/v2/query/video_generation/$TASK" \
      -H "X-API-Key: $NUNCHUX_API_KEY" \
      -H "User-Agent: YourApp/1.0")
    STATUS=$(jq -r '.task.status // ""' <<<"$RESP")
    case "$STATUS" in succeeded|failed|cancelled) break ;; esac
    sleep "$DELAY"; DELAY=$(( DELAY * 2 > 30 ? 30 : DELAY * 2 ))
  done

  # 3. Check for failure
  [ "$STATUS" = "succeeded" ] || { echo "task $STATUS: $(jq -r '.task.error // "no error"' <<<"$RESP")" >&2; exit 1; }

  # 4. Download
  curl -sSL -o minimax.mp4 "$(jq -r '.task.content.url' <<<"$RESP")"
  ```

  ```python Python theme={"system"}
  # pip install requests
  import os, time, requests

  BASE = "https://api.nunchux.ai"
  HEADERS = {
      "X-API-Key": os.environ["NUNCHUX_API_KEY"],
      "User-Agent": "YourApp/1.0",
  }

  # 1. Submit
  submit = requests.post(
      f"{BASE}/v1/minimax/v2/video_generation",
      headers=HEADERS,
      json={
          "model": "MiniMax-H3",
          "content": [
              {"type": "text", "text": "a fox trotting through a snowy forest at dawn"}
          ],
          "resolution": "768P",
          "duration": 6,
          "ratio": "16:9",
      },
  )
  if not submit.ok:
      raise RuntimeError(f"submit failed: HTTP {submit.status_code} {submit.text}")
  task = submit.json()["task_id"]

  # 2. Poll until the status is terminal
  delay = 5
  for _ in range(90):
      data = requests.get(
          f"{BASE}/v1/minimax/v2/query/video_generation/{task}", headers=HEADERS
      ).json()
      status = data.get("task", {}).get("status", "")
      if status in ("succeeded", "failed", "cancelled"):
          break
      time.sleep(delay)
      delay = min(delay * 2, 30)
  else:
      raise TimeoutError("MiniMax task did not reach a terminal status")

  # 3. Check for failure
  if status != "succeeded":
      raise RuntimeError(f"MiniMax task {status}: {data['task'].get('error')}")

  # 4. Download
  with requests.get(data["task"]["content"]["url"], stream=True, timeout=300) as r:
      r.raise_for_status()
      with open("minimax.mp4", "wb") as f:
          for chunk in r.iter_content(1 << 14):
              f.write(chunk)
  ```
</CodeGroup>

To animate a still, append an `image_url` item with `role` set to `first_frame` and leave `ratio` out. To carry a subject into a new scene, append `reference_image`, `reference_video` or `reference_audio` items instead, and leave `ratio` out. The bodies below replace the submit body in the example.

<CodeGroup>
  ```json Image-to-video theme={"system"}
  {
    "model": "MiniMax-H3",
    "content": [
      { "type": "text", "text": "the fox turns and trots toward the camera" },
      {
        "type": "image_url",
        "role": "first_frame",
        "image_url": { "url": "https://example.com/frame.jpg" }
      }
    ],
    "resolution": "768P",
    "duration": 6
  }
  ```

  ```json Reference-to-video theme={"system"}
  {
    "model": "MiniMax-H3",
    "content": [
      { "type": "text", "text": "the subject walks down a rain-soaked street at night, slow tracking shot" },
      {
        "type": "image_url",
        "role": "reference_image",
        "image_url": { "url": "https://example.com/ref-1.jpg" }
      },
      {
        "type": "image_url",
        "role": "reference_image",
        "image_url": { "url": "https://example.com/ref-2.jpg" }
      },
      {
        "type": "video_url",
        "role": "reference_video",
        "video_url": { "url": "https://example.com/clip.mp4" }
      },
      {
        "type": "audio_url",
        "role": "reference_audio",
        "audio_url": { "url": "https://example.com/track.mp3" }
      }
    ],
    "resolution": "768P",
    "duration": 6
  }
  ```
</CodeGroup>

## Tips

* Name a shape on text-to-video. `ratio` is required there, so decide the frame before you run instead of accepting whatever the first attempt gives you.
* Give at least one image or one clip on reference-to-video. A reference set describes the subject. The prompt describes the motion.
* Keep the reference images consistent in subject and style. Images that disagree pull the result in different directions.
* Draft at 768P. 2K costs more per second, so settle the prompt at the lower tier first.
* Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.

## Errors and limits

A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with `task.status` set to `failed` or `cancelled`, and `task.error` explains why. Read it before you resubmit.

A failed submit returns an HTTP error status. See [Error codes](/errors). Two request shapes are rejected: a text-only body without `ratio`, or with `ratio` set to `adaptive`, and a body that mixes the frame roles with the reference roles.

A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned a `task_id`, poll it. If it returned nothing, check your credits before you try again.

MiniMax applies these limits to the media you attach, and caps the whole set at 12 files:

* Images: 256 to 5,760 pixels on each side, JPG, JPEG, PNG, WEBP, HEIC or HEIF, up to 30 MB each
* Video: MP4 or MOV, up to 50 MB each, 2 to 15 seconds per clip and 15 seconds in total
* Audio: WAV or MP3, up to 15 MB each, 2 to 15 seconds

A file you pass by URL is not checked at submit. A file over a cap is refused by the provider after the request was accepted.

Reference media changes the charge. Reference video is billed per second at the same rate as the output. The first five reference images are included in the rate, and each image after that adds a charge. Reference audio is free. Every task counts toward the simultaneous-jobs cap of your plan. See [Rate limits](/rate-limits), [Credits & pricing](/credits-pricing) and the [Partner models overview](/partner-models/overview).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.