Nano Banana 2.1 API Guide

Nano Banana 2.1 is generally available in the Gemini API as gemini-nano-banana-2.1. This is the short version of Google's docs: the request shape, working examples in curl, Python and JavaScript, every parameter that matters, and the limits that bite in production.

Checked against Google's Gemini API and Google Cloud documentation on · Sources

This site doesn't offer its own API. Everything below calls Google's Gemini API directly, with your own key and your own billing.

Nano Banana 2.1 API at a glance

Model IDgemini-nano-banana-2.1
StatusGenerally available (stable) since October 6, 2026
InterfaceInteractions API — POST /v1beta/interactions. Google's default since June 2026; generateContent is now labeled legacy.
InputsText, images, video, PDF
OutputsImages and text
Sizes1K (default), 2K, 4K — no 512px
Aspect ratios1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9
Reference imagesUp to 14 per request: up to 4 for characters, up to 10 for objects. No style references.
Thinkingminimal, medium (default), high
GroundingGoogle Web Search and Image Search
Context window131,072 input tokens · 32,768 output tokens
Batch APISupported. Flex and Priority inference are not.
Not supportedFunction calling, structured outputs, code execution, URL context, Live API
WatermarkSynthID on every generated image
On Google Cloudglobal region only

1. Get an API key

Create a key in Google AI Studio and export it as GEMINI_API_KEY — the official SDKs read that variable automatically. Check the model's row on the pricing page for free-tier availability and rates before you build on it.

Shell
export GEMINI_API_KEY="your-api-key"

pip install -U google-genai     # Python SDK
npm install @google/genai       # JavaScript SDK

Upgrade old installs. In May 2026 Google moved image settings out of generation_config.image_config and into response_format, and older SDK versions break on Interactions API calls.

2. Generate an image

One request: the model ID, a prompt as input, and a response_format of type image with the aspect ratio and size.

curl

Shell
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-nano-banana-2.1",
    "input": "Editorial product photo of a matte black coffee bag on a white marble counter, soft window light. The label reads \"ROAST NO.7\" in bold serif type.",
    "response_format": {
      "type": "image",
      "mime_type": "image/png",
      "aspect_ratio": "1:1",
      "image_size": "2K"
    }
  }' > response.json

The image comes back base64-encoded inside the JSON. The SDKs below decode it for you; with raw REST, find the content block with "type": "image" in the model's output and base64-decode its data field.

Python

Python
import base64
from google import genai

client = genai.Client()  # reads GEMINI_API_KEY

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=(
        "Editorial product photo of a matte black coffee bag on a white marble counter, "
        'soft window light. The label reads "ROAST NO.7" in bold serif type.'
    ),
    response_format={
        "type": "image",
        "mime_type": "image/png",
        "aspect_ratio": "1:1",
        "image_size": "2K",
    },
)

if not interaction.output_image:
    raise SystemExit(interaction.output_text or "No image returned")

with open("coffee.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

JavaScript (Node.js)

JavaScript — save as generate.mjs
import { GoogleGenAI } from "@google/genai";
import { writeFileSync } from "node:fs";

const ai = new GoogleGenAI({}); // reads GEMINI_API_KEY

const interaction = await ai.interactions.create({
  model: "gemini-nano-banana-2.1",
  input:
    "Editorial product photo of a matte black coffee bag on a white marble counter, " +
    'soft window light. The label reads "ROAST NO.7" in bold serif type.',
  response_format: {
    type: "image",
    mime_type: "image/png",
    aspect_ratio: "1:1",
    image_size: "2K",
  },
});

const image = interaction.output_image;
if (!image) throw new Error(interaction.output_text ?? "No image returned");
writeFileSync("coffee.png", Buffer.from(image.data, "base64"));

Interactions API field names are snake_case in every SDK, including JavaScript: response_format, output_image, mime_type. Older generateContent samples use camelCase.

3. Edit an image

Send input as a list: a text block with the instruction, then the image as base64 data.

Python
import base64
from google import genai

client = genai.Client()

with open("portrait.jpg", "rb") as f:
    photo = base64.b64encode(f.read()).decode("utf-8")

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {"type": "text", "text": "Turn this into a 1990s film photo with warm grain. "
                                 "Keep the face, hairstyle and expression unchanged."},
        {"type": "image", "data": photo, "mime_type": "image/jpeg"},
    ],
    response_format={"type": "image", "mime_type": "image/png", "image_size": "2K"},
)

with open("portrait-film.png", "wb") as f:
    f.write(base64.b64decode(interaction.output_image.data))

Without an aspect_ratio, Google's docs say the output follows the input image's shape, or 1:1 when there is no input image. Set it explicitly when the format matters.

4. Multi-turn editing

Pass the previous interaction's id as previous_interaction_id and the model sees the whole history — this is where 2.1's improved character consistency pays off.

Python — continues the example above
follow_up = client.interactions.create(
    model="gemini-nano-banana-2.1",
    previous_interaction_id=interaction.id,
    input="Now make the jacket red. Keep everything else exactly the same.",
    response_format={"type": "image", "mime_type": "image/png", "image_size": "2K"},
)

# Upscale the keeper as one more turn. Re-running the original prompt
# at 4K would give a new composition: the model has no seed.
final = client.interactions.create(
    model="gemini-nano-banana-2.1",
    previous_interaction_id=follow_up.id,
    input="Render this exact image again at higher resolution. Change nothing.",
    response_format={"type": "image", "mime_type": "image/png", "image_size": "4K"},
)

5. Multiple reference images

Up to 14 images per request: up to 4 for keeping characters consistent and up to 10 for objects. Refer to them by order in the prompt. Each input image counts as 1,120 input tokens.

Python
import base64
from google import genai

client = genai.Client()

def image_part(path, mime_type="image/png"):
    with open(path, "rb") as f:
        data = base64.b64encode(f.read()).decode("utf-8")
    return {"type": "image", "data": data, "mime_type": mime_type}

interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input=[
        {"type": "text", "text": "The woman from image 1 wearing the jacket from image 2, "
                                 "holding the bottle from image 3, on a rooftop at sunset. "
                                 "Keep her face exactly as in image 1."},
        image_part("woman.png"),
        image_part("jacket.png"),
        image_part("bottle.png"),
    ],
    response_format={"type": "image", "mime_type": "image/png",
                     "aspect_ratio": "4:5", "image_size": "2K"},
)

Style reference images aren't supported on 2.1 — Google lists them only for Nano Banana Pro. Describe the style in the prompt instead.

6. Grounding with Google Search

Add the google_search tool to ground the image in live search results. Include image_search to let the model use web images as visual context.

Python
interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input="A travel poster of the Taipei 101 observatory entrance as it looks today, daylight, wide angle.",
    tools=[{"type": "google_search", "search_types": ["web_search", "image_search"]}],
    response_format={"type": "image", "mime_type": "image/png",
                     "aspect_ratio": "3:4", "image_size": "2K"},
)
  • Off by default. Without the tool there is no grounding. With web_search alone, image results are not passed to the model.
  • No real people. Google says Image Search grounding can't use real-world images of people.
  • Display rules. If you show grounded results to users, Google requires you to display the search suggestions returned with them. Read the grounding docs before you ship.
  • Billed separately. Search requests are priced apart from image tokens — see the pricing page.

7. Thinking level

2.1 supports minimal, medium (default) and high. Use minimal for fast drafts and simple prompts, high for infographics and dense layouts.

Python
interaction = client.interactions.create(
    model="gemini-nano-banana-2.1",
    input="A dense infographic comparing five coffee brewing methods by time, grind size and flavor.",
    generation_config={"thinking_level": "high"},
    response_format={"type": "image", "mime_type": "image/png",
                     "aspect_ratio": "3:4", "image_size": "2K"},
)

The model may render up to two interim "thought images" to test the composition. Those images aren't charged; thinking tokens are.

Output sizes by aspect ratio

Pixel dimensions per Google's image generation guide. Every aspect ratio uses the same number of output tokens at a given size.
Ratio 1K 2K 4K
1:11024 × 10242048 × 20484096 × 4096
16:91376 × 7682752 × 15365504 × 3072
9:16768 × 13761536 × 27523072 × 5504
21:91584 × 6723168 × 13446336 × 2688
4:12048 × 5124096 × 10248192 × 2048
8:13072 × 3846144 × 76812288 × 1536
Output tokens1,1201,6802,520

Pricing

Nano Banana 2.1 is billed per token, like the rest of the Gemini API:

  • Output images are billed as image output tokens: 1,120 at 1K, 1,680 at 2K and 2,520 at 4K, the same at every aspect ratio.
  • Input images count as 1,120 input tokens each.
  • Thinking tokens are billed; the interim "thought images" are not.
  • Search grounding is billed separately from tokens.
  • Batch API is supported for asynchronous jobs.

Google describes 2.1 as keeping Flash-level cost. The per-token rate is on Google's pricing page — multiply it by the token counts above for a per-image price. For reference, here are the per-image list prices of the other Nano Banana models on the same page:

Standard tier, per output image, checked October 7, 2026. For Nano Banana 2 and Pro, Batch prices are half of these.
Model 512px 1K 2K 4K
Nano Banana 2 Lite
gemini-3.1-flash-lite-image
—$0.0336——
Nano Banana 2
gemini-3.1-flash-image
$0.045$0.067$0.101$0.151
Nano Banana Pro
gemini-3-pro-image
—$0.134$0.134$0.24

Limits and gotchas

  • No 512px. image_size: "512" works on Nano Banana 2 only. Use "1K".
  • Uppercase sizes. "1K", "2K", "4K" — lowercase values like "1k" are rejected.
  • No sampling parameters. Google Cloud's model page says setting seed, temperature, topP, topK or logprobs returns an error. There is no way to reproduce an image exactly.
  • Old code breaks. Since May 2026, image settings go in response_format, not generation_config.image_config. Upgrade the SDK and update the request shape.
  • Don't mix interfaces. generateContent, now labeled legacy, nests image settings differently. A snippet from an older tutorial won't work inside an Interactions request.
  • Thinking is on by default. The default level is medium. Set minimal on latency-sensitive paths.
  • 4K token count. The Gemini API docs list 2,520 tokens for a 4K image; Google Cloud's model page lists 3,780. Check the usage figures in a real response before budgeting 4K at scale.
  • Region. On Google Cloud, 2.1 is served from the global region only.
  • Watermark. Every image includes SynthID; there is no parameter to turn it off.

Call the API, or use this site?

Gemini API directly This site
SetupAPI key, billing, codeNone — type a prompt
PricingGoogle's per-token ratesCredit packs from $9, no subscription
ModelsEvery Nano Banana modelNano Banana 2.1 only
Best forProducts, pipelines, volumeIndividual images and quick edits

Generating at volume? Call the API — it is cheaper. Need a few images without writing code? Use the generator.

FAQ

What is the Nano Banana 2.1 model ID?

gemini-nano-banana-2.1. It is the stable, generally available model ID, released on October 6, 2026.

Which endpoint does Nano Banana 2.1 use?

Google's documentation uses the Interactions API: POST https://generativelanguage.googleapis.com/v1beta/interactions. The older generateContent API is now labeled legacy in Google's docs.

How do I generate a 4K image with the Nano Banana 2.1 API?

Set response_format to type image with image_size: "4K". The K must be uppercase, and 1K is the default.

Is the Nano Banana 2.1 API free?

Free-tier availability is listed per model on Google's Gemini API pricing page, so check it before you build. Some image models, such as Nano Banana Pro, have no free API tier. To try 2.1 without code, use the Gemini app or AI Mode in Google Search.

Can I turn off the SynthID watermark?

No. Google states that all generated images include a SynthID watermark, and there is no parameter to remove it.

Does this site offer its own Nano Banana 2.1 API?

No. This site is a web interface only. For programmatic access, call Google's Gemini API directly as shown on this page.

Rather not write code?

Use Nano Banana 2.1 in the browser. Credit packs from $9, no subscription.

Open the generator