Nano Banana 2.1 Prompt Guide
Nano Banana 2.1 reads a prompt the way a designer reads a brief: full sentences, specific details, exact text in quotes. This guide covers the patterns that work for generating, editing and multi-turn changes — plus 24 prompts you can copy and run.
The prompt formula
Google's own Gemini guidance starts from a simple structure: action, subject, setting, atmosphere. For images you will actually use, fill in a few more slots. You rarely need all of them — you always need the subject.
| Slot | What to write | Example |
|---|---|---|
| Subject | Who or what, with specifics | a matte black coffee bag with a kraft paper label |
| Action or pose | What it is doing, how it is placed | standing upright, turned slightly toward the camera |
| Setting | Where and when | on a white marble counter in a sunlit kitchen |
| Composition | Shot type, angle, lens | eye-level close-up, shallow depth of field |
| Lighting | Source, direction, mood | soft window light from the left |
| Style | Medium or genre | editorial product photography |
| Text | Exact words in quotes, plus placement | the label reads "ROAST NO.7" in bold serif type |
coffee bag, marble, product photo, 4k, high quality, masterpiece
Editorial product photo of a matte black coffee bag with a kraft paper label, standing upright on a white marble counter in a sunlit kitchen. Eye-level close-up, shallow depth of field, soft window light from the left. The label reads "ROAST NO.7" in bold serif type.
Words like "4k", "masterpiece" or "high quality" don't change the output size — resolution is a setting, not a prompt word. Spend those words on the subject and the light instead.
Getting text to render exactly
Text rendering and infographic layouts are among the improvements Google lists for 2.1. You still get the best results by being explicit.
- Quote every exact string.
the sign reads "OPEN LATE", nota sign saying open late. - Say where it goes and how it looks. "Centered at the top, bold condensed sans-serif, white on red."
- Keep strings short. Headlines, labels and short lines render far more reliably than paragraphs. Set body copy in a design tool.
- Limit the number of text elements. Two or three per image is comfortable; a dozen invites a typo.
- Name the language for non-English text.
the menu title reads "早餐" in Traditional Chinese. - Fix typos with a targeted edit.
Change the headline to "SUMMER SALE". Keep everything else exactly the same.
Editing a photo
Upload the image, then describe the change — and what must stay the same. The pattern is: the change, what to keep, and how it should blend.
Change + keep + blend
Replace the background with a sunny beach at golden hour. Keep the person's face, hair, pose and clothing unchanged, and match the lighting on them to the new scene.
Upload a photo first- Name what to preserve. Face, expression, pose, product shape, logo, existing text — anything you don't list is fair game.
- Point at the area. "The mug on the left", "only the sky", "the text on the sign".
- One precise change per request. Combine changes only when each one is simple.
- Restyle, don't replace. For a new look on the same photo, say "restyle" and list what stays: "Restyle as a watercolor painting; keep the composition and everyone's position."
Multi-turn edits
Multi-turn character consistency is one of the four areas Google improved in 2.1. To get the most out of it:
- Take small steps. Several small, sequential edits beat one giant rewrite.
- Refer back. "Same character as before, now sitting at a desk."
- Re-anchor when something drifts. "Keep her face, red hair and freckles exactly as in the first image."
- Start fresh when a chain gets long. Begin a new session from the best image so far rather than stacking edit after edit.
- Upscale last. Ask for the final image at 4K as one more follow-up edit. Re-running the original prompt at 4K produces a new composition — the model has no seed.
A three-turn chain
Turn 1: A small orange robot with round eyes watering a plant on a windowsill, soft morning light, 3D render. Turn 2: Same robot, now holding an umbrella in the rain outside the same window. Turn 3: Make the rain heavier and add puddle reflections — keep the robot's design identical.
Send each turn as a follow-upReference images
2.1 accepts up to 14 reference images in one request — up to 4 for keeping characters consistent and up to 10 for objects you want reproduced faithfully.
- Refer to images by order. "The woman in image 1, wearing the jacket from image 2, holding the bottle from image 3."
- Use clean references. One subject per image, well lit, plain background.
- Style references are a Pro feature. Google's docs list style-reference slots only for Nano Banana Pro (up to 3). On 2.1, describe the style in words — medium, palette, lighting, era — or use Nano Banana Pro when matching a look matters.
Combining three references
The woman from image 1 wearing the denim jacket from image 2, holding the perfume bottle from image 3, on a rooftop at sunset, fashion editorial photo. Keep her face exactly as in image 1.
Upload 3 images · 4:5 · 2KAspect ratio and resolution
Pick the ratio for where the image will be used, and set it in the controls rather than in the prompt. 2.1 supports 14 ratios: 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9 and 21:9.
| Where it's going | Ratio | Size |
|---|---|---|
| Marketplace product photo, avatar | 1:1 | 1K–2K |
| Instagram feed post | 4:5 | 2K |
| Stories, Reels, TikTok, Shorts | 9:16 | 2K |
| YouTube thumbnail, blog hero, slides | 16:9 | 1K–2K |
| Printed poster | 2:3 or 3:4 | 4K |
| Cinematic banner | 21:9 | 2K–4K |
| Website header strip | 4:1 or 8:1 | 2K–4K |
| Tall infographic, vertical banner | 1:4 or 1:8 | 2K–4K |
1K is the default and there is no 512px option. Iterate at 1K, then ask for the keeper at 4K as a follow-up edit.
Grounded prompts
2.1 can ground a prompt in Google Search: web results, and with Image Search, web images used as visual context. It helps most with real places, current products and recent events.
- Ask for the real thing explicitly. "The Taipei 101 observatory entrance as it looks today", "this year's packaging".
- Ask for details grounding can pin down. Signage, season, materials, colors.
- Not for real people. Google says Image Search grounding can't use real-world images of people.
- Still check the result. Grounding reduces guessing; it does not turn the output into a photograph of the real place.
In the API, grounding stays off until you add the google_search tool — see the API guide.
Thinking level (API)
2.1 can reason through a prompt before it renders, producing up to two interim "thought images" to test the composition. You aren't charged for those images, but thinking adds latency and thinking tokens.
| Level | Use it for |
|---|---|
minimal | Simple prompts, fast drafts, high volume |
medium (default) | Most work |
high | Infographics, dense layouts, many subjects, tight constraints |
Mistakes to avoid
- Keyword soup. Write sentences. The model reads descriptions better than tag lists.
- Negative lists. "No cars, no people" works worse than describing the scene you want: "an empty street at dawn". Google's own prompting guidance recommends this positive phrasing.
- Clashing styles. "Photorealistic watercolor anime" — pick one medium.
- Paragraphs of text. Keep strings short and add long copy in a design tool.
- Upscaling by re-prompting. The same prompt at 4K gives a new image. Ask for 4K as a follow-up edit instead.
- Real people through grounding. Image Search grounding doesn't use real-world images of people.
24 Nano Banana 2.1 prompts to copy
Each card lists the aspect ratio and size it was written for. Swap in your own product, place or text.
Product and e-commerce
Hero product shot
Editorial product photo of a matte black coffee bag with a kraft paper label, standing on a white marble counter. Soft window light from the left, shallow depth of field, eye-level close-up. The label reads "ROAST NO.7" in bold serif type.
1:1 · 2KLifestyle scene
A frosted glass skincare bottle on a bathroom shelf next to a folded white towel and a small plant, morning light through a window, calm and clean, magazine-style photo. The bottle label reads "DEW SERUM".
4:5 · 2KPackshot on color
A pair of white running sneakers floating at a slight angle over a solid tangerine background, crisp studio lighting, soft shadow below, clean commercial look.
1:1 · 2KPackaging mockup
Top-down photo of a cardboard tea box on a linen tablecloth, botanical line illustrations on the box. The front reads "LOW TIDE TEA" with "Jasmine Green · 20 bags" underneath.
4:3 · 2KPosters and text-heavy designs
Travel poster
Vintage screen-print travel poster of Taipei at sunset, Taipei 101 silhouetted against an orange sky, limited palette of four colors, visible paper texture. The headline reads "VISIT TAIPEI" in bold condensed letters at the top.
2:3 · 4KEvent flyer
A modern flyer for a jazz night: dark navy background, a gold saxophone illustration, clean geometric layout. The title reads "BLUE NOTE FRIDAYS" and the date line reads "OCT 24 · 9 PM".
3:4 · 2KShop sign
A hand-painted wooden sign above a small bakery door on an early-morning street, photographed straight on. The sign reads "MORNING LIGHT BAKERY" in cream lettering with a gold outline.
3:2 · 2KMenu board
A hand-lettered chalkboard café menu on a brick wall. The title reads "TODAY'S BREW", followed by three lines: "Flat White 4.50", "Cold Brew 5.00", "Matcha Latte 5.50".
4:5 · 2KPhoto editing
Background swap
Replace the background with a sunny beach at golden hour. Keep the person's face, hair, pose and clothing unchanged, and match the lighting on them to the new scene.
Upload a photo · keep its ratioFilm look
Turn this portrait into a 1990s film camera photo: warm color cast, soft grain, slight flash falloff. Keep the same face, hairstyle and expression.
Upload a photo · keep its ratioClean up a scene
Remove the people in the background and fill the space with the same street. Leave everything in the foreground exactly as it is.
Upload a photo · keep its ratioRestore an old photo
Restore this old photo: remove scratches and dust, bring back the faded colors and sharpen the faces slightly. Keep it looking like an authentic photo from its era.
Upload a photo · 2KCharacters and mascots
Mascot sheet
A friendly banana mascot with big round eyes and a small smile, flat vector style, thick outlines, white background, shown in three poses side by side: waving, thumbs up, jumping.
16:9 · 2KSame mascot, new scene
Same mascot as before, now sitting at a laptop in a cozy home office. Keep its shape, colors and face identical.
Follow-up editConsistent person
A woman in her 30s with short curly red hair and freckles, green wool coat, standing on a train platform, candid street photo. Next turn: same woman, same coat, now reading in a café by the window.
4:5 · 1K · two turnsWordless comic strip
A four-panel comic strip of the same small robot: it finds a seed, plants it, waters it in the rain, then sits under the grown tree. Wordless panels, consistent robot design, soft watercolor style.
4:1 · 2KInfographics and diagrams
Step-by-step infographic
A clean infographic titled "How Coffee Is Made" with five numbered steps — Harvest, Process, Roast, Grind, Brew — each with a simple flat icon and a one-line caption, beige background, one terracotta accent color.
3:4 · 2K · high thinkingComparison chart
A simple two-column comparison graphic titled "Tea vs Coffee", with rows for Caffeine, Flavor and Ritual, short labels only, flat icons, generous white space.
4:5 · 2KProcess diagram
An isometric flowchart of how a parcel is delivered: Order, Warehouse, Truck, Doorstep, connected by arrows, soft blue palette, labels in clean sans-serif.
16:9 · 2KRecipe card
A recipe card for "Garlic Butter Noodles": a top-down photo of the dish, an ingredients list of five short items and three numbered steps, warm kitchen palette.
2:3 · 2K · high thinkingBanners and social formats
Website header
A wide website header of a mountain range at sunrise, soft fog in the valleys, pastel sky, empty space on the left third for a headline.
8:1 · 2KYouTube thumbnail
A YouTube thumbnail: a surprised chef holding a giant golden croissant in a bright kitchen, bold text on the right reads "1000 LAYERS?".
16:9 · 1KStory ad
A vertical story ad for a running app: a runner on a coastal path at sunrise with slight motion blur. The top reads "RUN THE MORNING".
9:16 · 2KTall vertical illustration
A tall cross-section illustration of a tree from roots to canopy: roots and soil layers at the bottom, a bird's nest in the branches at the top, earthy palette, clean editorial style.
1:8 · 2KFAQ
What is the best prompt format for Nano Banana 2.1?
Full sentences that describe the scene: subject, action, setting, composition, lighting and style, with any exact text in double quotes. Set the aspect ratio and size in the controls, not in the prompt.
How do I get Nano Banana 2.1 to spell text correctly?
Put the exact words in double quotes, say where and how they should appear, keep each string short and limit the number of text elements. If one word comes out wrong, fix it with a follow-up edit that changes only that text.
Does Nano Banana 2.1 support negative prompts?
There is no separate negative prompt field. Describe what you want instead: "an empty street at dawn" works better than "no cars, no people".
How do I keep the same character across images?
Edit in one multi-turn session, restate the identity details such as face, hair and outfit when something drifts, and pass up to 4 character reference images when you start a new session.
Do Nano Banana 2 prompts work on 2.1?
Yes. 2.1 is an update to Nano Banana 2 and reads the same prompts. Text-heavy prompts and long edit chains are where Google says 2.1 improved, so re-test those first. What changed in 2.1.
1K, 2K or 4K output. Credit packs from $9, no subscription.
Sources
- Gemini API — Nano Banana image generation guide
- Gemini API — Gemini Nano Banana 2.1 model page
- Gemini app — Image generation overview and prompting tips
Prompt advice beyond Google's documentation is our own, based on how the model responds in practice. Google's pages are the authority on what the model supports.