If you've already looked into how to generate images in Gemini, you probably ran into a confusing mess of names: Nano Banana, Nano Banana 2, Nano Banana Pro, Gemini 3.1 Flash Image. They sound like different products, and they aren't — Nano Banana is simply the name of Gemini's native image generation, and the rest are versions of it.

This clarification matters because choosing the model is the decision that most affects the result. An infographic with dense text made in the fast model comes out with crooked letters; the same request in the Pro model comes out legible. Knowing which one to use in each case saves attempts and daily quota.

This guide covers the three models and when to use each one, the step-by-step generation process, how to write prompts the way Gemini makes the most of, the conversational editing that's the tool's main advantage, and its real limitations.

What Nano Banana is

Nano Banana is the name of Gemini's native image generation and editing capability. Native means the image model is built into Gemini itself, and not called as a separate system — it understands the request within the same conversation context, including images you've uploaded and what's already been said before.

In February 2026, Nano Banana 2 became the standard path for image generation across Google's main surfaces, including the Gemini app. Nano Banana Pro still exists as a high-fidelity option, mainly for people with a paid plan.

The three models and when to use each one

Inside Gemini, the model menu shows up with behavior-based names — Fast, Reasoning, and Pro. Under the hood, they correspond to three distinct models.

ModelProfileUse when
Nano Banana 2 Lite (Fast)The fastest and most economicalTests, quick variations, simple single-scene images
Nano Banana 2 (Reasoning)The all-purpose one, standard for almost everythingMost cases: scenes, portraits, products, illustrations
Nano Banana ProHigh fidelity, dense text and complex layoutsInfographics, posters with dense text, production pieces

A useful feature of the workflow: you can generate first with the standard model and, if the result needs more detail, redo the same image in Pro. On paid plans, the option appears in the generated image's three-dot menu, as "Redo with Pro." The limitation is that once the standard model's daily quota runs out, redoing in Pro also stops being available.

💡 Quota strategy: use the fast model to find the composition that works and only then redo it in Pro. Spending the more expensive model's generations on tests is the mistake that wastes the most daily quota.

How to generate your first image

  1. Open Gemini. On a computer, access it through the browser; on a phone, through the app.
  2. Go to the images area. There's an Images section in the sidebar; alternatively, the tools menu has the option to create images.
  3. Choose the model. Fast, Reasoning, or Pro, according to the table above. If unsure, start with the standard one.
  4. Write the request in natural language. Complete sentences work better than comma-separated word lists.
  5. Ask for adjustments in the same conversation. "Make the background darker," "change it to late afternoon," "remove the plant in the corner." It keeps the image and only changes what was asked.
  6. Download at full resolution. Use the download button, never a screenshot.

Resolution: what you can download

Download resolution depends on the plan. Without a paid Google AI plan, downloads come at 1K. With a plan, you can download at 2K. Nano Banana Pro is the path to higher-fidelity output, including material with text and layout that needs to hold up under upscaling.

In practice this means that for common digital use — a post, a thumbnail, a blog image, a presentation — what Gemini delivers is already enough. For large-format printing or a high-resolution monitor, you'll need to upscale after generation. You can increase an image's resolution without losing quality, and the result tends to be good when the image is clean and not overly detailed.

How to write prompts for Gemini

Google itself suggests a simple starting formula: ask for the creation of an image stating the subject, the action, and the scene. It's a good starting point — and the next step is exactly to expand on it.

Start with the formula, then add detail

A minimal request would look like this:

Create an image of a cat napping in a sunbeam on a windowsill.

That already gives a result. But each additional layer brings the result closer to what you imagined:

Create an image of a black cat napping curled up in a sunbeam on a wooden windowsill, a linen curtain beside it swaying gently, visible suspended dust in the light, side medium shot, warm late-afternoon light, beige and amber tones, contemplative domestic photography.

Write in prose, not tags

This is the most important difference compared to tools like Midjourney. Gemini is conversational and makes better use of complete sentences than comma-separated term lists. Write as if you were explaining the scene to someone.

State the aspect ratio

If you don't indicate it, the model chooses. Say "vertical format for stories," "square for Instagram," "widescreen horizontal," or the aspect ratio directly.

Put text in quotes

When the image needs to contain words, write them literally in quotes and state the position. Without this, the model may paraphrase or place it wherever it thinks is best.

Iterate instead of rewriting

If something came out wrong, ask for the specific fix instead of rewriting everything. "Keep the composition and only change the coat color to green" preserves what was already good. Rewriting the entire prompt generates a new image unrelated to the previous one.

Conversational editing: the main advantage

This is where Gemini stands out. You upload an image — generated by it or your own — and describe the change in words, with no mask, layer, or selection.

Swap the background

Keep the subject and replace the setting behind them. Useful for portraits, products, and marketing pieces.

Keep the person exactly as they are and replace the background with a light exposed-concrete wall, with diffuse natural light coming from the left. Preserve the lighting on the face.

Change the mood and time of day

Turn a sunny day into night, add rain, swap autumn for winter. The same scene with a different atmosphere.

Keep the composition and all the elements. Change from a sunny day to a rainy night, with the wet street reflecting the lights and low fog.

Change the camera angle

Asking for the same scene seen from another point. It's not perfect, because the model needs to invent what wasn't visible, but it works well in scenes with room to spare.

Show the same scene seen from a lower angle, close to the ground, keeping the same objects, the same lighting, and the same style.

Transfer style from a reference

One of the most useful and least-known features: uploading two images and asking for the texture, color, or style of one to be applied to the other. It's the fastest way to test aesthetics without starting from scratch.

Apply the color palette and texture of the second image to the first, keeping the composition, objects, and framing of the first exactly the same.

Add or remove elements

Removing an unwanted object, adding a plant, including a person in the background. Works well with clearly delimited elements; removals in areas of complex texture sometimes leave marks.

Remove the parked car on the right and naturally complete the street, keeping the rest of the image unchanged.

Working with reference images

Uploading images along with the request is where Gemini shows the most strength, and it's worth understanding the three ways to use it.

Subject reference. You upload a photo of a person, product, or character and ask for them to appear in a new scene. The model tries to preserve the subject's identity and swap everything else.

Style reference. You upload an image whose aesthetic you want to reproduce and ask for the palette, texture, or visual language to be applied to another image or to a scene described in text.

Composition reference. You upload a sketch, a poorly taken photo, or a rough draft and ask for it to be turned into a finished image while keeping the arrangement of the elements.

One practical tip worth noting: the cleaner the reference, the better the result. Images with a cluttered background make the model carry over elements from the setting that you didn't ask for. If your reference has that problem, removing the background before uploading improves the result a lot.

Text and infographics

Legible text rendering is one of the most visible advances in this generation of models, and it's exactly where Nano Banana Pro justifies itself. A poster with a title and subtitle, an infographic with labels, a diagram from notes, a cover with a brand name — all of this used to come out crooked until recently and today comes out clean on most attempts.

Three things greatly improve the result with text:

Even so, for pieces where the typography needs to be the brand's own, the more reliable path is still to generate the image with no text and add the text afterward, with the right font.

Eight ready-made everyday requests

Templates that cover the most common needs. Swap what's in brackets and keep the rest.

Header image for an article

Create a widescreen horizontal image about [topic], with a clean composition, one central element in focus, and the left side emptier to receive text. Soft natural lighting, palette in [color] tones and neutrals, no text at all in the image.

Square social media post

Create a square image of [subject] with a minimalist aesthetic, plain [color] background, soft shadow, and plenty of negative space around the object. Uniform diffuse light, high contrast between object and background, no text.

Poster with a title

Create a vertical promotional poster for [event]. Title "[exact text]" in large letters at the top and, smaller, below: "[exact text]". [Color] background with a simple graphic element. Editorial style, sans-serif typography, plenty of empty space.

Swap a photo's background

Keep the person exactly as they are, without changing the face, hair, or clothing. Replace only the background with [scene description], preserving the same light direction hitting them.

Style variation

Keep the same composition, the same elements, and the same framing, but redo the image in [watercolor / vector illustration / realistic photography] style, with a [description] palette.

Product mockup

Create a product photo of [product] on [surface], soft gradient [color] background, studio lighting with a side softbox, slightly elevated frontal angle, controlled reflection, high sharpness, no text or label.

Simple infographic

Create a horizontal infographic with three equal columns explaining [topic]. Each column has a simple icon at the top, a short title, and a supporting sentence. The titles are "[text]", "[text]" and "[text]". Palette in [two colors], clean editorial style, no shadows.

Clean up a photo

Remove [unwanted element] from the image and naturally complete the area, keeping everything else unchanged, with the same lighting and the same level of detail.

Character consistency in Gemini

Since editing is conversational, the most effective method is also the simplest: stay in the same conversation thread. After generating the character, ask for the next scenes without re-describing them — "now show the same character sitting in a café" — and the accumulated context does the work.

When the series is long or you need to pick it back up another day, use the image as a subject reference and keep a fixed description of the permanent traits, repeated word for word. The guide on AI character consistency details the four methods and when each one pays off.

Gemini or ChatGPT: what changes in practice

AspectGeminiChatGPT
Prompt styleConversational proseConversational prose
Conversational editingStrong, with good preservation of the originalStrong, with good preservation of the original
Multiple input imagesStrong point, including style transferSupported
Text in the imageVery good, even better in the Pro modelVery good
Model choiceExplicit: Fast, Reasoning, or ProInstant or reasoning, depending on the plan
Download resolution1K free, 2K with a planUp to 2K

The two tools have converged quite a bit. The practical criterion is to test the same request in both and compare — and use whichever delivers better results for your type of image. The ChatGPT guide for creating images covers the other side in detail, and the comparison of the best AI tools to create images puts the main options side by side.

What Gemini doesn't do well

Seven mistakes that hurt the most

  1. Using the Pro model for testing. Spends quota that would be needed for the final version.
  2. Writing in tag format. A word list is the language of other tools; here, prose performs better.
  3. Not stating the aspect ratio. With no indication, the format is almost never what you need.
  4. Rewriting the prompt on every attempt. You lose what you'd already gotten right. Ask for specific fixes.
  5. Uploading a reference with a cluttered background. The model carries over elements from the setting that you didn't ask for.
  6. Saving via screenshot. Loses resolution and adds recompression.
  7. Publishing without optimizing. 2K images are heavy and tank the loading time of any page.

After generating: preparing the image

The image that comes out of Gemini is ready to be viewed, not to be published. Four steps solve it.

Convert the format. PNG is great for transparency and bad for photos on the web. For a website, WebP tends to cut the file size in half with no visible difference; for general sharing, JPG works. The image converter does the swap in the browser, and the guide on image formats helps you choose.

Reduce the file size. A 2K image easily goes over several megabytes. The image compressor cuts a good chunk of that with no visible loss.

Adjust size and framing. For a specific measurement, use the resizer; if the aspect ratio came out different from what you need, cropping preserves more quality than stretching.

Isolate the subject when needed. For products, logos, and elements that will go over a colored background, removing the background gives you the transparent PNG the generation doesn't produce.

💡 Order that works: download at full resolution → crop or resize → convert the format → compress last. Compressing before resizing wastes quality.

Checklist before generating

Convert your images to the right format

Turn Gemini's heavy PNG into WebP or JPG and get the image ready to publish, right in your browser.

Convert image

Frequently asked questions

What is Nano Banana and how does it relate to Gemini?
Nano Banana is the name of Gemini's native image generation — it's not a separate product or a standalone app. Today there are three models in the family: Lite, focused on speed and cost, Nano Banana 2, which has been the standard all-purpose model since February 2026, and Nano Banana Pro, high-fidelity, aimed at dense text and complex layouts. In the app, they appear in the menu as Fast, Reasoning, and Pro.
Can you create images in Gemini for free?
Yes, with limitations. Free use has a daily generation quota and downloads come at 1K resolution. With a paid Google AI plan, downloads go up to 2K and the option to redo the image in Nano Banana Pro becomes available. A practical detail that catches a lot of people: when the standard model's daily quota runs out, redoing in Pro also stops working, even with a paid plan.
Which model should I choose for each type of image?
For most cases — scenes, portraits, products, illustrations — the standard model works well. The fast model is good for exploring compositions and generating variations without spending quota on the more expensive one. Pro pays off when the request involves dense text, an infographic, a diagram, a poster with typographic hierarchy, or a piece headed for production. The most economical strategy is to test on the fast model and redo only the chosen version on Pro.
How do I edit my own photo in Gemini?
Upload the image in the conversation and describe the change in natural language — swap the background, change the lighting, remove an object, apply a different style. No mask or selection needed. It's worth knowing that the result is a new image generated from yours, not a pixel-by-pixel edit: the composition is preserved, but fine details from the original change. For changes that require absolute fidelity, a traditional editor is still more appropriate.
How do I get a transparent background?
Even when explicitly requested, the image tends to come with a solid background instead of real transparency. The fastest path is to ask for a plain white background, which is the easiest to cut out, and create the transparency afterward by removing the background and saving as PNG. Insisting with the model tends to result in a drawn checkerboard background imitating transparency, which isn't useful for anything.
Can I use the generated images commercially?
The rules depend on the plan, the country, and Google's current terms, and they can change. Also, permission for commercial use isn't the same as legal protection: in several jurisdictions, works created entirely by a machine face restrictions on copyright registration. Before using it in client material or a product for sale, check the current terms and avoid requests involving registered trademarks, licensed characters, or real people.