Nano Banana 2.1 API Guide
Nano Banana 2.1 is generally available in the Gemini API as gemini-nano-banana-2.1. This is the short version of Google's docs: the request shape, working examples in curl, Python and JavaScript, every parameter that matters, and the limits that bite in production.
This site doesn't offer its own API. Everything below calls Google's Gemini API directly, with your own key and your own billing.
Nano Banana 2.1 API at a glance
| Model ID | gemini-nano-banana-2.1 |
|---|---|
| Status | Generally available (stable) since October 6, 2026 |
| Interface | Interactions API — POST /v1beta/interactions. Google's default since June 2026; generateContent is now labeled legacy. |
| Inputs | Text, images, video, PDF |
| Outputs | Images and text |
| Sizes | 1K (default), 2K, 4K — no 512px |
| Aspect ratios | 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9 |
| Reference images | Up to 14 per request: up to 4 for characters, up to 10 for objects. No style references. |
| Thinking | minimal, medium (default), high |
| Grounding | Google Web Search and Image Search |
| Context window | 131,072 input tokens · 32,768 output tokens |
| Batch API | Supported. Flex and Priority inference are not. |
| Not supported | Function calling, structured outputs, code execution, URL context, Live API |
| Watermark | SynthID on every generated image |
| On Google Cloud | global region only |
1. Get an API key
Create a key in Google AI Studio and export it as GEMINI_API_KEY — the official SDKs read that variable automatically. Check the model's row on the pricing page for free-tier availability and rates before you build on it.
export GEMINI_API_KEY="your-api-key"
pip install -U google-genai # Python SDK
npm install @google/genai # JavaScript SDK
Upgrade old installs. In May 2026 Google moved image settings out of generation_config.image_config and into response_format, and older SDK versions break on Interactions API calls.
2. Generate an image
One request: the model ID, a prompt as input, and a response_format of type image with the aspect ratio and size.
curl
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-nano-banana-2.1",
"input": "Editorial product photo of a matte black coffee bag on a white marble counter, soft window light. The label reads \"ROAST NO.7\" in bold serif type.",
"response_format": {
"type": "image",
"mime_type": "image/png",
"aspect_ratio": "1:1",
"image_size": "2K"
}
}' > response.json
The image comes back base64-encoded inside the JSON. The SDKs below decode it for you; with raw REST, find the content block with "type": "image" in the model's output and base64-decode its data field.
Python
import base64
from google import genai
client = genai.Client() # reads GEMINI_API_KEY
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=(
"Editorial product photo of a matte black coffee bag on a white marble counter, "
'soft window light. The label reads "ROAST NO.7" in bold serif type.'
),
response_format={
"type": "image",
"mime_type": "image/png",
"aspect_ratio": "1:1",
"image_size": "2K",
},
)
if not interaction.output_image:
raise SystemExit(interaction.output_text or "No image returned")
with open("coffee.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))
JavaScript (Node.js)
import { GoogleGenAI } from "@google/genai";
import { writeFileSync } from "node:fs";
const ai = new GoogleGenAI({}); // reads GEMINI_API_KEY
const interaction = await ai.interactions.create({
model: "gemini-nano-banana-2.1",
input:
"Editorial product photo of a matte black coffee bag on a white marble counter, " +
'soft window light. The label reads "ROAST NO.7" in bold serif type.',
response_format: {
type: "image",
mime_type: "image/png",
aspect_ratio: "1:1",
image_size: "2K",
},
});
const image = interaction.output_image;
if (!image) throw new Error(interaction.output_text ?? "No image returned");
writeFileSync("coffee.png", Buffer.from(image.data, "base64"));
Interactions API field names are snake_case in every SDK, including JavaScript: response_format, output_image, mime_type. Older generateContent samples use camelCase.
3. Edit an image
Send input as a list: a text block with the instruction, then the image as base64 data.
import base64
from google import genai
client = genai.Client()
with open("portrait.jpg", "rb") as f:
photo = base64.b64encode(f.read()).decode("utf-8")
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{"type": "text", "text": "Turn this into a 1990s film photo with warm grain. "
"Keep the face, hairstyle and expression unchanged."},
{"type": "image", "data": photo, "mime_type": "image/jpeg"},
],
response_format={"type": "image", "mime_type": "image/png", "image_size": "2K"},
)
with open("portrait-film.png", "wb") as f:
f.write(base64.b64decode(interaction.output_image.data))
Without an aspect_ratio, Google's docs say the output follows the input image's shape, or 1:1 when there is no input image. Set it explicitly when the format matters.
4. Multi-turn editing
Pass the previous interaction's id as previous_interaction_id and the model sees the whole history — this is where 2.1's improved character consistency pays off.
follow_up = client.interactions.create(
model="gemini-nano-banana-2.1",
previous_interaction_id=interaction.id,
input="Now make the jacket red. Keep everything else exactly the same.",
response_format={"type": "image", "mime_type": "image/png", "image_size": "2K"},
)
# Upscale the keeper as one more turn. Re-running the original prompt
# at 4K would give a new composition: the model has no seed.
final = client.interactions.create(
model="gemini-nano-banana-2.1",
previous_interaction_id=follow_up.id,
input="Render this exact image again at higher resolution. Change nothing.",
response_format={"type": "image", "mime_type": "image/png", "image_size": "4K"},
)
5. Multiple reference images
Up to 14 images per request: up to 4 for keeping characters consistent and up to 10 for objects. Refer to them by order in the prompt. Each input image counts as 1,120 input tokens.
import base64
from google import genai
client = genai.Client()
def image_part(path, mime_type="image/png"):
with open(path, "rb") as f:
data = base64.b64encode(f.read()).decode("utf-8")
return {"type": "image", "data": data, "mime_type": mime_type}
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input=[
{"type": "text", "text": "The woman from image 1 wearing the jacket from image 2, "
"holding the bottle from image 3, on a rooftop at sunset. "
"Keep her face exactly as in image 1."},
image_part("woman.png"),
image_part("jacket.png"),
image_part("bottle.png"),
],
response_format={"type": "image", "mime_type": "image/png",
"aspect_ratio": "4:5", "image_size": "2K"},
)
Style reference images aren't supported on 2.1 — Google lists them only for Nano Banana Pro. Describe the style in the prompt instead.
6. Grounding with Google Search
Add the google_search tool to ground the image in live search results. Include image_search to let the model use web images as visual context.
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A travel poster of the Taipei 101 observatory entrance as it looks today, daylight, wide angle.",
tools=[{"type": "google_search", "search_types": ["web_search", "image_search"]}],
response_format={"type": "image", "mime_type": "image/png",
"aspect_ratio": "3:4", "image_size": "2K"},
)
- Off by default. Without the tool there is no grounding. With
web_searchalone, image results are not passed to the model. - No real people. Google says Image Search grounding can't use real-world images of people.
- Display rules. If you show grounded results to users, Google requires you to display the search suggestions returned with them. Read the grounding docs before you ship.
- Billed separately. Search requests are priced apart from image tokens — see the pricing page.
7. Thinking level
2.1 supports minimal, medium (default) and high. Use minimal for fast drafts and simple prompts, high for infographics and dense layouts.
interaction = client.interactions.create(
model="gemini-nano-banana-2.1",
input="A dense infographic comparing five coffee brewing methods by time, grind size and flavor.",
generation_config={"thinking_level": "high"},
response_format={"type": "image", "mime_type": "image/png",
"aspect_ratio": "3:4", "image_size": "2K"},
)
The model may render up to two interim "thought images" to test the composition. Those images aren't charged; thinking tokens are.
Output sizes by aspect ratio
| Ratio | 1K | 2K | 4K |
|---|---|---|---|
| 1:1 | 1024 × 1024 | 2048 × 2048 | 4096 × 4096 |
| 16:9 | 1376 × 768 | 2752 × 1536 | 5504 × 3072 |
| 9:16 | 768 × 1376 | 1536 × 2752 | 3072 × 5504 |
| 21:9 | 1584 × 672 | 3168 × 1344 | 6336 × 2688 |
| 4:1 | 2048 × 512 | 4096 × 1024 | 8192 × 2048 |
| 8:1 | 3072 × 384 | 6144 × 768 | 12288 × 1536 |
| Output tokens | 1,120 | 1,680 | 2,520 |
Pricing
Nano Banana 2.1 is billed per token, like the rest of the Gemini API:
- Output images are billed as image output tokens: 1,120 at 1K, 1,680 at 2K and 2,520 at 4K, the same at every aspect ratio.
- Input images count as 1,120 input tokens each.
- Thinking tokens are billed; the interim "thought images" are not.
- Search grounding is billed separately from tokens.
- Batch API is supported for asynchronous jobs.
Google describes 2.1 as keeping Flash-level cost. The per-token rate is on Google's pricing page — multiply it by the token counts above for a per-image price. For reference, here are the per-image list prices of the other Nano Banana models on the same page:
| Model | 512px | 1K | 2K | 4K |
|---|---|---|---|---|
Nano Banana 2 Litegemini-3.1-flash-lite-image | — | $0.0336 | — | — |
Nano Banana 2gemini-3.1-flash-image | $0.045 | $0.067 | $0.101 | $0.151 |
Nano Banana Progemini-3-pro-image | — | $0.134 | $0.134 | $0.24 |
Limits and gotchas
- No 512px.
image_size: "512"works on Nano Banana 2 only. Use"1K". - Uppercase sizes.
"1K","2K","4K"— lowercase values like"1k"are rejected. - No sampling parameters. Google Cloud's model page says setting
seed,temperature,topP,topKorlogprobsreturns an error. There is no way to reproduce an image exactly. - Old code breaks. Since May 2026, image settings go in
response_format, notgeneration_config.image_config. Upgrade the SDK and update the request shape. - Don't mix interfaces.
generateContent, now labeled legacy, nests image settings differently. A snippet from an older tutorial won't work inside an Interactions request. - Thinking is on by default. The default level is
medium. Setminimalon latency-sensitive paths. - 4K token count. The Gemini API docs list 2,520 tokens for a 4K image; Google Cloud's model page lists 3,780. Check the usage figures in a real response before budgeting 4K at scale.
- Region. On Google Cloud, 2.1 is served from the
globalregion only. - Watermark. Every image includes SynthID; there is no parameter to turn it off.
Call the API, or use this site?
| Gemini API directly | This site | |
|---|---|---|
| Setup | API key, billing, code | None — type a prompt |
| Pricing | Google's per-token rates | Credit packs from $9, no subscription |
| Models | Every Nano Banana model | Nano Banana 2.1 only |
| Best for | Products, pipelines, volume | Individual images and quick edits |
Generating at volume? Call the API — it is cheaper. Need a few images without writing code? Use the generator.
FAQ
What is the Nano Banana 2.1 model ID?
gemini-nano-banana-2.1. It is the stable, generally available model ID, released on October 6, 2026.
Which endpoint does Nano Banana 2.1 use?
Google's documentation uses the Interactions API: POST https://generativelanguage.googleapis.com/v1beta/interactions. The older generateContent API is now labeled legacy in Google's docs.
How do I generate a 4K image with the Nano Banana 2.1 API?
Set response_format to type image with image_size: "4K". The K must be uppercase, and 1K is the default.
Is the Nano Banana 2.1 API free?
Free-tier availability is listed per model on Google's Gemini API pricing page, so check it before you build. Some image models, such as Nano Banana Pro, have no free API tier. To try 2.1 without code, use the Gemini app or AI Mode in Google Search.
Can I turn off the SynthID watermark?
No. Google states that all generated images include a SynthID watermark, and there is no parameter to remove it.
Does this site offer its own Nano Banana 2.1 API?
No. This site is a web interface only. For programmatic access, call Google's Gemini API directly as shown on this page.
Use Nano Banana 2.1 in the browser. Credit packs from $9, no subscription.
Sources
- Gemini API — Gemini Nano Banana 2.1 model page
- Gemini API — Nano Banana image generation guide (Interactions API)
- Gemini API — Interactions API reference
- Gemini API — Interactions API breaking changes (May 2026)
- Gemini API — Grounding with Google Search
- Gemini API — Pricing
- Google Cloud — Gemini Nano Banana 2.1
Google updates these pages without notice. If anything here disagrees with them, Google's page is right.