> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nunchux.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Authenticate every request with your API key in the `X-API-Key` header or `Authorization: Bearer`. Keys start with `sk-nunchux-`.
> Image generation is synchronous: the image is in the response.
> Video generation is asynchronous: submit the job, poll its task until it reaches a terminal status, then download the output promptly, because output URLs expire.

# Nano Banana

> Generate and edit images with Google's Nano Banana models in one synchronous call.

# Nano Banana

Nano Banana is Google's Gemini image model. It generates an image from a text prompt, and it edits images that you send with the prompt. The call is synchronous: one POST returns the finished image in the response body, so there is nothing to poll. The request and response shapes mirror Google's own API.

## Models

| Model | Model ID | Tasks | Sizes | Aspect ratios |
| - | - | - | - | - |
| Nano Banana 2 | `gemini-3.1-flash-image` | Generate, edit | `512`, `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |
| Nano Banana 2.1 | `gemini-nano-banana-2.1` | Generate, edit | `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |
| Nano Banana Pro | `gemini-3-pro-image` | Generate, edit | `1K`, `2K`, `4K` | 14 values, `1:1` to `21:9` |

All three models generate and edit on the same endpoint. All three accept up to 14 reference images per edit. The size is a pixel budget that you set with `image_size`. The shape is a separate axis that you set with `aspect_ratio`.

Use Nano Banana 2 for fast everyday generation and edits. It is the default model, and it is the only one that accepts the `512` size.

Use Nano Banana 2.1 for the same everyday work when you do not need the `512` size. It is Google's update to Nano Banana 2, and it starts at `1K`. It is a thinking model, so it inserts a `thought` step before the output step in the response.

Use Nano Banana Pro when you need the highest fidelity. It starts at `1K`. It is a thinking model, so it inserts a `thought` step before the output step in the response.

## Endpoints

The base URL is `https://api.nunchux.ai`.

| Step | Method | Path |
| - | - | - |
| Generate or edit | `POST` | `/v1/google/v1beta/interactions` |

There is no poll route. The image is in the response.

## Request

Send your API key in the `X-API-Key` header. See [Authentication](/authentication). Set `Content-Type: application/json`. Send a `User-Agent` header that names your application, for example `YourApp/1.0`.

The JSON body names the model and an `input` array of content parts. A text-only input generates a new image. Image parts turn the request into an edit. See [Editing](#editing).

<ParamField body="model" type="string" required default="gemini-3.1-flash-image">
  The Nano Banana model to run. See the [Models](#models) table.

  Options: `gemini-3.1-flash-image`, `gemini-nano-banana-2.1`, `gemini-3-pro-image`
</ParamField>

<ParamField body="input" type="array" required>
  Ordered content parts. A text-only input generates. Add `{ "type": "image" }` parts (up to 14) and the text becomes an edit instruction across them. Image parts also carry `data` and `mime_type`. See [Editing](#editing).

  <Expandable title="Item properties">
    <ParamField body="type" type="string" required>
      The part kind.

      Options: `text`, `image`
    </ParamField>

    <ParamField body="text" type="string">
      The prompt or the edit instruction. Required on `type: "text"` parts.
    </ParamField>
  </Expandable>
</ParamField>

<ParamField body="response_format" type="object" required>
  Output controls. A request that omits `response_format` or its `type` is rejected with `no_image_format` or `invalid_response_format`.

  <Expandable title="Item properties">
    <ParamField body="type" type="string" required>
      Must be `"image"`.

      Options: `image`
    </ParamField>

    <ParamField body="image_size" type="string">
      Output resolution, as a pixel budget. `512` is accepted on `gemini-3.1-flash-image` only. Google's documentation calls that tier 0.5K, but `0.5K` is rejected as a value.

      Options: `1K`, `2K`, `4K`
    </ParamField>

    <ParamField body="aspect_ratio" type="string">
      Output shape, independent of `image_size`. Omitted on a generation, the output is `1:1` on Nano Banana 2 and Nano Banana Pro. On Nano Banana 2.1, set the ratio when the shape matters. Omitted on an edit, the output keeps the shape of the first input image.

      Options: `1:1`, `2:3`, `3:2`, `3:4`, `4:3`, `4:5`, `5:4`, `9:16`, `16:9`, `21:9`, `1:8`, `8:1`, `1:4`, `4:1`
    </ParamField>
  </Expandable>
</ParamField>

<Note>
  Output size and shape

  * `image_size` is a pixel budget, not a shape. `1K` is 1024×1024 at `1:1`, 1584×672 at `21:9` and 768×1376 at `9:16`.
  * Set `image_size` explicitly. When you omit it, the call is billed at the largest tier (`4K`), on generations and on edits.
  * The 14 ratios in the list are accepted on every model. `1:1`, `21:9` and `9:16` are the ones that have been run end to end.
  * There is no `auto` value for `aspect_ratio`. To keep the shape of the input on an edit, omit the field.
</Note>

### Editing

Add one or more `{ "type": "image" }` parts to `input`, up to 14. The text part becomes an instruction across all of them. The result is always one new image, not a batch and not a mechanical merge. `image_size` and `aspect_ratio` work exactly as on a generation.

<ParamField body="input[].data" type="string">
  Raw base64 image bytes on `type: "image"` parts, with no `data:` prefix. Images are sent inline. There is no URL form.
</ParamField>

<ParamField body="input[].mime_type" type="string">
  The MIME type of the image on `type: "image"` parts, for example `image/png` or `image/jpeg`.
</ParamField>

## Response

The image comes back in the same response, as base64 JPEG bytes.

<ResponseField name="id" type="string" required>
  Identifier for this interaction.
</ResponseField>

<ResponseField name="object" type="string" required>
  The envelope kind.
</ResponseField>

<ResponseField name="model" type="string" required>
  The model that ran, echoed back.
</ResponseField>

<ResponseField name="status" type="string" required>
  The status of the interaction. It is terminal on arrival. The call is synchronous, so there is no in-progress state to poll for.
</ResponseField>

<ResponseField name="created" type="integer" required>
  Unix timestamp when the interaction was created.
</ResponseField>

<ResponseField name="updated" type="integer" required>
  Unix timestamp of the last update, in practice when the image finished.
</ResponseField>

<ResponseField name="service_tier" type="string" required>
  The service tier the call was served on.
</ResponseField>

<ResponseField name="usage" type="object" required>
  Accounting for the call.
</ResponseField>

<ResponseField name="steps" type="array" required>
  The steps of the interaction. The image is on the first entry whose `content[]` carries a `data` field.

  <Expandable title="Item properties">
    <ResponseField name="signature" type="string">
      Opaque vendor blob on a step of its own. It can be absent. It is not the image, so skip it.
    </ResponseField>

    <ResponseField name="content" type="array">
      The output parts. Take the first step that has one. A `thought` step can sit ahead of it.

      <Expandable title="Item properties">
        <ResponseField name="data" type="string">
          Base64-encoded JPEG bytes. Decode and save.
        </ResponseField>
      </Expandable>
    </ResponseField>
  </Expandable>
</ResponseField>

<Note>
  Reading the response

  * Select the output step by shape, not by index: the first `steps[]` entry whose `content[]` carries a `data` field. The `signature` step is optional, and a thinking model such as Nano Banana 2.1 or Nano Banana Pro inserts a `thought` step ahead of the output.
  * The bytes are JPEG whatever the request asked for. Name the file accordingly.
  * There is no top-level `data` array and no `b64_json` field. If you port a parser from `/v1/images`, this is the line to change.
</Note>

```json theme={"system"}
{
  "id": "int_01k9v0f2m8e0",
  "object": "interaction",
  "model": "gemini-3.1-flash-image",
  "status": "completed",
  "created": 1786636800,
  "updated": 1786636809,
  "service_tier": "default",
  "usage": { "...": "accounting for the call" },
  "steps": [
    { "signature": "CtwBAdHtim8yq0pQ7f..." },
    {
      "content": [
        { "data": "/9j/4AAQSkZJRgABAQAAAQABAAD..." }
      ]
    }
  ]
}
```

## Example

### Generate

One POST returns the image inline. The example takes the first `steps[]` content part that has a `data` field and saves it as a JPEG.

<CodeGroup>
  ```bash cURL theme={"system"}
  # The image is the first steps[] content part with a data field. There is no
  # data[].b64_json here, and no fixed index: the step layout varies by model.
  curl -s https://api.nunchux.ai/v1/google/v1beta/interactions \
    -H "X-API-Key: $NUNCHUX_API_KEY" \
    -H "Content-Type: application/json" \
    -H "User-Agent: YourApp/1.0" \
    -d '{
      "model": "gemini-3.1-flash-image",
      "input": [{ "type": "text", "text": "a nano banana dessert, studio lighting" }],
      "response_format": { "type": "image", "image_size": "1K", "aspect_ratio": "16:9" }
    }' | jq -r '[.steps[].content[]? | select(.data)][0].data' | base64 -d > output.jpg
  ```

  ```python Python theme={"system"}
  # pip install requests
  import base64, os, requests

  BASE = "https://api.nunchux.ai"
  HEADERS = {
      "X-API-Key": os.environ["NUNCHUX_API_KEY"],
      "User-Agent": "YourApp/1.0",
  }

  # Synchronous: the image is in the response body, so there is no polling.
  resp = requests.post(
      f"{BASE}/v1/google/v1beta/interactions",
      headers=HEADERS,
      json={
          "model": "gemini-3.1-flash-image",
          "input": [{"type": "text", "text": "a nano banana dessert, studio lighting"}],
          "response_format": {"type": "image", "image_size": "1K", "aspect_ratio": "16:9"},
      },
  )
  resp.raise_for_status()

  # The envelope is {id, status, usage, created, updated, service_tier, steps, ...},
  # not data[].b64_json. The steps[] layout varies (the signature step is optional,
  # thinking models add a thought step), so take the first part that has data.
  steps = resp.json()["steps"]
  b64 = next(p["data"] for s in steps for p in s.get("content", []) if p.get("data"))
  with open("output.jpg", "wb") as f:
      f.write(base64.b64decode(b64))
  ```

  ```javascript JavaScript theme={"system"}
  import { writeFile } from "node:fs/promises";

  const BASE = "https://api.nunchux.ai";
  const HEADERS = {
    "X-API-Key": process.env.NUNCHUX_API_KEY,
    "Content-Type": "application/json",
    "User-Agent": "YourApp/1.0",
  };

  // Synchronous: the image is in the response body, so there is no polling.
  const res = await fetch(`${BASE}/v1/google/v1beta/interactions`, {
    method: "POST",
    headers: HEADERS,
    body: JSON.stringify({
      model: "gemini-3.1-flash-image",
      input: [{ type: "text", text: "a nano banana dessert, studio lighting" }],
      response_format: { type: "image", image_size: "1K", aspect_ratio: "16:9" },
    }),
  });
  if (!res.ok) throw new Error(`request failed: HTTP ${res.status} ${await res.text()}`);
  const body = await res.json();

  // The envelope is { id, status, usage, created, updated, service_tier, steps, ... },
  // not data[].b64_json. The steps[] layout varies (the signature step is optional,
  // thinking models add a thought step), so take the first part that has data.
  const b64 = body.steps.flatMap((s) => s.content ?? []).find((p) => p.data).data;
  await writeFile("output.jpg", Buffer.from(b64, "base64"));
  ```
</CodeGroup>

### Edit

The same call with image parts. The prompt places the person from the first image into the room from the second. Omit `aspect_ratio` to keep the shape of the first input. `image_size` still applies, and it still bills at `4K` when you omit it.

<CodeGroup>
  ```bash cURL theme={"system"}
  # Images are sent inline as base64. There is no URL form.
  PERSON=$(base64 < person.jpg | tr -d '\n')
  ROOM=$(base64 < room.jpg | tr -d '\n')

  curl -s https://api.nunchux.ai/v1/google/v1beta/interactions \
    -H "X-API-Key: $NUNCHUX_API_KEY" \
    -H "Content-Type: application/json" \
    -H "User-Agent: YourApp/1.0" \
    -d '{
      "model": "gemini-3.1-flash-image",
      "input": [
        { "type": "text", "text": "put the person from the first image into the room from the second, keep the lighting natural" },
        { "type": "image", "data": "'"$PERSON"'", "mime_type": "image/jpeg" },
        { "type": "image", "data": "'"$ROOM"'", "mime_type": "image/jpeg" }
      ],
      "response_format": { "type": "image", "image_size": "2K", "aspect_ratio": "3:4" }
    }' | jq -r '[.steps[].content[]? | select(.data)][0].data' | base64 -d > output.jpg
  ```

  ```python Python theme={"system"}
  # pip install requests
  import base64, os, requests

  BASE = "https://api.nunchux.ai"
  HEADERS = {
      "X-API-Key": os.environ["NUNCHUX_API_KEY"],
      "User-Agent": "YourApp/1.0",
  }

  def b64_file(path):
      with open(path, "rb") as f:
          return base64.b64encode(f.read()).decode()

  # Images are sent inline as base64. There is no URL form.
  resp = requests.post(
      f"{BASE}/v1/google/v1beta/interactions",
      headers=HEADERS,
      json={
          "model": "gemini-3.1-flash-image",
          "input": [
              {"type": "text", "text": "put the person from the first image into the room from the second, keep the lighting natural"},
              {"type": "image", "data": b64_file("person.jpg"), "mime_type": "image/jpeg"},
              {"type": "image", "data": b64_file("room.jpg"), "mime_type": "image/jpeg"},
          ],
          "response_format": {"type": "image", "image_size": "2K", "aspect_ratio": "3:4"},
      },
  )
  resp.raise_for_status()

  steps = resp.json()["steps"]
  b64 = next(p["data"] for s in steps for p in s.get("content", []) if p.get("data"))
  with open("output.jpg", "wb") as f:
      f.write(base64.b64decode(b64))
  ```

  ```javascript JavaScript theme={"system"}
  import { readFile, writeFile } from "node:fs/promises";

  const BASE = "https://api.nunchux.ai";
  const HEADERS = {
    "X-API-Key": process.env.NUNCHUX_API_KEY,
    "Content-Type": "application/json",
    "User-Agent": "YourApp/1.0",
  };
  const b64File = async (path) => (await readFile(path)).toString("base64");

  // Images are sent inline as base64. There is no URL form.
  const res = await fetch(`${BASE}/v1/google/v1beta/interactions`, {
    method: "POST",
    headers: HEADERS,
    body: JSON.stringify({
      model: "gemini-3.1-flash-image",
      input: [
        { type: "text", text: "put the person from the first image into the room from the second, keep the lighting natural" },
        { type: "image", data: await b64File("person.jpg"), mime_type: "image/jpeg" },
        { type: "image", data: await b64File("room.jpg"), mime_type: "image/jpeg" },
      ],
      response_format: { type: "image", image_size: "2K", aspect_ratio: "3:4" },
    }),
  });
  if (!res.ok) throw new Error(`request failed: HTTP ${res.status} ${await res.text()}`);
  const body = await res.json();

  const b64 = body.steps.flatMap((s) => s.content ?? []).find((p) => p.data).data;
  await writeFile("output.jpg", Buffer.from(b64, "base64"));
  ```
</CodeGroup>

## Tips

* Write the prompt as an instruction. Name what must change and what must stay.
* When you send several reference images, say which element comes from which image.
* Set `image_size` explicitly. An omitted size is billed at the `4K` tier, on edits as well as on generations.
* Use a larger `image_size` only when you need the detail. Larger sizes cost more and take longer.
* Set `aspect_ratio` when you want a shape other than `1:1`. Size and shape are separate axes.
* Use Nano Banana Pro for the highest fidelity. Use Nano Banana 2 for fast everyday work.
* Read the image from the first `steps[]` entry that has a `data` part. Do not hardcode a step index.
* Decode and save the image on receipt. The response carries the bytes, not a URL.

## Errors and limits

* A request takes up to 14 image parts. Images are sent inline as base64, with no URL form.
* A request that omits `response_format` or `response_format.type` is rejected with `no_image_format` or `invalid_response_format`. A rejected request is not charged.
* `512` is accepted on `gemini-3.1-flash-image` only. `0.5K` is rejected on every model.
* Nano Banana is priced per image. The price rises with `image_size`, and an omitted `image_size` bills at the `4K` tier. The rates are on the [pricing page](https://nunchux.ai/pricing). See [Credits and pricing](/credits-pricing).
* A Nano Banana call counts toward your requests per minute. It does not take a simultaneous-job slot. See [Rate limits](/rate-limits).
* A `502` after the request was sent is refunded automatically. It is safe to retry.
* A `503` with the code `google_unreachable` means that Google could not be reached before anything was sent. Nothing was charged. Retry with backoff.
* A `503` whose message says `google pass-through not enabled` means that the route is not available. Do not retry it in a loop.

The full error contract, the retry rules and the per-plan caps are on the [Errors](/errors) and [Rate limits](/rate-limits) pages. How every partner model bills and refunds is on the [Partner models overview](/partner-models/overview).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.