Skip to main content

MiniMax H3

MiniMax H3 is an omni-modal model that reads text, images, video and audio in one context and generates a video clip. One model ID serves all three tasks. The content[] items you send decide which task you get. Generation is asynchronous. You submit a task, poll it until the status is terminal, then download the clip. The route mirrors MiniMax’s own API, so the request and response shapes are the provider’s.
Best for
  • Multi-reference scenes. Up to nine reference images, three reference clips and a reference track carry a subject or a style into new footage.
  • 2K output. A higher tier than the other partner video models sell.
  • Animating a still. A first frame fixes the composition and the prompt drives the motion.
  • Longer prompts. The vendor accepts up to 7,000 characters per text item.

Models

Durations are whole seconds. All three tasks send MiniMax-H3 as the model value. 2K costs more per second than 768P. The content[] items select the task. A text item alone gives text-to-video. A first_frame image item gives image-to-video. reference_image, reference_video and reference_audio items give reference-to-video. A first frame and a reference set cannot appear in the same body. The two use different role values, and a body that mixes the role sets is rejected. Pick one task per request.

Endpoints

Request

Send your API key in the X-API-Key header. See Authentication. Set Content-Type: application/json. The examples also send a User-Agent header that names your application.
string
required
The model ID: MiniMax-H3.
array
required
The prompt and the media items. Put the text item first. Pair each role with its own container: an image rides image_url, a clip rides video_url, a track rides audio_url. A role on the wrong container is accepted without complaint and then does not do what you meant.
string
The output tier: 768P or 2K. All tasks.
integer
required
The clip length in whole seconds, 4 to 15. All tasks. You are billed per second of output.
string
The output shape. Required on a text-only body: 21:9, 16:9, 4:3, 1:1, 3:4 or 9:16. adaptive is rejected there. Leave it out when the body carries an image, a clip or a track. The output then follows the source.

Response

Submit

A 200 status means that the provider accepted the task. It does not mean that the clip is ready.
string
required
The task ID. Poll GET /v1/minimax/v2/query/video_generation/{task_id} with it.

Poll

string
required
queued or running while the task runs. succeeded, failed or cancelled when it ends. Poll until you read one of the three terminal values.
string
The URL of the clip. Present only when status is succeeded. The URL is time-limited. Download the clip as soon as the task succeeds. A later poll returns a fresh URL. A task can be polled for 7 days.
object
The failure detail. Present when the task did not succeed. Read it before you resubmit.
object
The billed quantities: output seconds, reference-video seconds and the billed image count.

Example

Submit, poll, download. The submit and the poll use different paths.
To animate a still, append an image_url item with role set to first_frame and leave ratio out. To carry a subject into a new scene, append reference_image, reference_video or reference_audio items instead, and leave ratio out. The bodies below replace the submit body in the example.

Tips

  • Name a shape on text-to-video. ratio is required there, so decide the frame before you run instead of accepting whatever the first attempt gives you.
  • Give at least one image or one clip on reference-to-video. A reference set describes the subject. The prompt describes the motion.
  • Keep the reference images consistent in subject and style. Images that disagree pull the result in different directions.
  • Draft at 768P. 2K costs more per second, so settle the prompt at the lower tier first.
  • Poll every 5 seconds at first, then back off to 30 seconds. Polls are free and never use a concurrency slot.

Errors and limits

A 200 status on submit means that the provider accepted the task. It does not mean that the task succeeded. A task that fails later ends with task.status set to failed or cancelled, and task.error explains why. Read it before you resubmit. A failed submit returns an HTTP error status. See Error codes. Two request shapes are rejected: a text-only body without ratio, or with ratio set to adaptive, and a body that mixes the frame roles with the reference roles. A 5xx status or a timeout on submit does not tell you whether the task was created. Never resubmit blindly. If the submit returned a task_id, poll it. If it returned nothing, check your credits before you try again. MiniMax applies these limits to the media you attach, and caps the whole set at 12 files:
  • Images: 256 to 5,760 pixels on each side, JPG, JPEG, PNG, WEBP, HEIC or HEIF, up to 30 MB each
  • Video: MP4 or MOV, up to 50 MB each, 2 to 15 seconds per clip and 15 seconds in total
  • Audio: WAV or MP3, up to 15 MB each, 2 to 15 seconds
A file you pass by URL is not checked at submit. A file over a cap is refused by the provider after the request was accepted. Reference media changes the charge. Reference video is billed per second at the same rate as the output. The first five reference images are included in the rate, and each image after that adds a charge. Reference audio is free. Every task counts toward the simultaneous-jobs cap of your plan. See Rate limits, Credits & pricing and the Partner models overview.