If you've already looked into how to generate images in Gemini, you probably ran into a confusing mess of names: Nano Banana, Nano Banana 2, Nano Banana Pro, Gemini 3.1 Flash Image. They sound like different products, and they aren't — Nano Banana is simply the name of Gemini's native image generation, and the rest are versions of it.
This clarification matters because choosing the model is the decision that most affects the result. An infographic with dense text made in the fast model comes out with crooked letters; the same request in the Pro model comes out legible. Knowing which one to use in each case saves attempts and daily quota.
This guide covers the three models and when to use each one, the step-by-step generation process, how to write prompts the way Gemini makes the most of, the conversational editing that's the tool's main advantage, and its real limitations.
What Nano Banana is
Nano Banana is the name of Gemini's native image generation and editing capability. Native means the image model is built into Gemini itself, and not called as a separate system — it understands the request within the same conversation context, including images you've uploaded and what's already been said before.
In February 2026, Nano Banana 2 became the standard path for image generation across Google's main surfaces, including the Gemini app. Nano Banana Pro still exists as a high-fidelity option, mainly for people with a paid plan.
The three models and when to use each one
Inside Gemini, the model menu shows up with behavior-based names — Fast, Reasoning, and Pro. Under the hood, they correspond to three distinct models.
| Model | Profile | Use when |
|---|---|---|
| Nano Banana 2 Lite (Fast) | The fastest and most economical | Tests, quick variations, simple single-scene images |
| Nano Banana 2 (Reasoning) | The all-purpose one, standard for almost everything | Most cases: scenes, portraits, products, illustrations |
| Nano Banana Pro | High fidelity, dense text and complex layouts | Infographics, posters with dense text, production pieces |
A useful feature of the workflow: you can generate first with the standard model and, if the result needs more detail, redo the same image in Pro. On paid plans, the option appears in the generated image's three-dot menu, as "Redo with Pro." The limitation is that once the standard model's daily quota runs out, redoing in Pro also stops being available.
💡 Quota strategy: use the fast model to find the composition that works and only then redo it in Pro. Spending the more expensive model's generations on tests is the mistake that wastes the most daily quota.
How to generate your first image
- Open Gemini. On a computer, access it through the browser; on a phone, through the app.
- Go to the images area. There's an Images section in the sidebar; alternatively, the tools menu has the option to create images.
- Choose the model. Fast, Reasoning, or Pro, according to the table above. If unsure, start with the standard one.
- Write the request in natural language. Complete sentences work better than comma-separated word lists.
- Ask for adjustments in the same conversation. "Make the background darker," "change it to late afternoon," "remove the plant in the corner." It keeps the image and only changes what was asked.
- Download at full resolution. Use the download button, never a screenshot.
Resolution: what you can download
Download resolution depends on the plan. Without a paid Google AI plan, downloads come at 1K. With a plan, you can download at 2K. Nano Banana Pro is the path to higher-fidelity output, including material with text and layout that needs to hold up under upscaling.
In practice this means that for common digital use — a post, a thumbnail, a blog image, a presentation — what Gemini delivers is already enough. For large-format printing or a high-resolution monitor, you'll need to upscale after generation. You can increase an image's resolution without losing quality, and the result tends to be good when the image is clean and not overly detailed.
How to write prompts for Gemini
Google itself suggests a simple starting formula: ask for the creation of an image stating the subject, the action, and the scene. It's a good starting point — and the next step is exactly to expand on it.
Start with the formula, then add detail
A minimal request would look like this:
That already gives a result. But each additional layer brings the result closer to what you imagined:
Write in prose, not tags
This is the most important difference compared to tools like Midjourney. Gemini is conversational and makes better use of complete sentences than comma-separated term lists. Write as if you were explaining the scene to someone.
State the aspect ratio
If you don't indicate it, the model chooses. Say "vertical format for stories," "square for Instagram," "widescreen horizontal," or the aspect ratio directly.
Put text in quotes
When the image needs to contain words, write them literally in quotes and state the position. Without this, the model may paraphrase or place it wherever it thinks is best.
Iterate instead of rewriting
If something came out wrong, ask for the specific fix instead of rewriting everything. "Keep the composition and only change the coat color to green" preserves what was already good. Rewriting the entire prompt generates a new image unrelated to the previous one.
Conversational editing: the main advantage
This is where Gemini stands out. You upload an image — generated by it or your own — and describe the change in words, with no mask, layer, or selection.
Swap the background
Keep the subject and replace the setting behind them. Useful for portraits, products, and marketing pieces.
Change the mood and time of day
Turn a sunny day into night, add rain, swap autumn for winter. The same scene with a different atmosphere.
Change the camera angle
Asking for the same scene seen from another point. It's not perfect, because the model needs to invent what wasn't visible, but it works well in scenes with room to spare.
Transfer style from a reference
One of the most useful and least-known features: uploading two images and asking for the texture, color, or style of one to be applied to the other. It's the fastest way to test aesthetics without starting from scratch.
Add or remove elements
Removing an unwanted object, adding a plant, including a person in the background. Works well with clearly delimited elements; removals in areas of complex texture sometimes leave marks.
Working with reference images
Uploading images along with the request is where Gemini shows the most strength, and it's worth understanding the three ways to use it.
Subject reference. You upload a photo of a person, product, or character and ask for them to appear in a new scene. The model tries to preserve the subject's identity and swap everything else.
Style reference. You upload an image whose aesthetic you want to reproduce and ask for the palette, texture, or visual language to be applied to another image or to a scene described in text.
Composition reference. You upload a sketch, a poorly taken photo, or a rough draft and ask for it to be turned into a finished image while keeping the arrangement of the elements.
One practical tip worth noting: the cleaner the reference, the better the result. Images with a cluttered background make the model carry over elements from the setting that you didn't ask for. If your reference has that problem, removing the background before uploading improves the result a lot.
Text and infographics
Legible text rendering is one of the most visible advances in this generation of models, and it's exactly where Nano Banana Pro justifies itself. A poster with a title and subtitle, an infographic with labels, a diagram from notes, a cover with a brand name — all of this used to come out crooked until recently and today comes out clean on most attempts.
Three things greatly improve the result with text:
- Writing the exact text in quotes. Without this, the model interprets the word as a theme and may paraphrase it.
- Describing the hierarchy. "Big title at the top, smaller subtitle below, three blocks of short text at the bottom" gives much more control than leaving it open.
- Using the Pro model. Dense layout and small text are exactly the case where the difference between models shows up.
Even so, for pieces where the typography needs to be the brand's own, the more reliable path is still to generate the image with no text and add the text afterward, with the right font.
Eight ready-made everyday requests
Templates that cover the most common needs. Swap what's in brackets and keep the rest.
Header image for an article
Square social media post
Poster with a title
Swap a photo's background
Style variation
Product mockup
Simple infographic
Clean up a photo
Character consistency in Gemini
Since editing is conversational, the most effective method is also the simplest: stay in the same conversation thread. After generating the character, ask for the next scenes without re-describing them — "now show the same character sitting in a café" — and the accumulated context does the work.
When the series is long or you need to pick it back up another day, use the image as a subject reference and keep a fixed description of the permanent traits, repeated word for word. The guide on AI character consistency details the four methods and when each one pays off.
Gemini or ChatGPT: what changes in practice
| Aspect | Gemini | ChatGPT |
|---|---|---|
| Prompt style | Conversational prose | Conversational prose |
| Conversational editing | Strong, with good preservation of the original | Strong, with good preservation of the original |
| Multiple input images | Strong point, including style transfer | Supported |
| Text in the image | Very good, even better in the Pro model | Very good |
| Model choice | Explicit: Fast, Reasoning, or Pro | Instant or reasoning, depending on the plan |
| Download resolution | 1K free, 2K with a plan | Up to 2K |
The two tools have converged quite a bit. The practical criterion is to test the same request in both and compare — and use whichever delivers better results for your type of image. The ChatGPT guide for creating images covers the other side in detail, and the comparison of the best AI tools to create images puts the main options side by side.
What Gemini doesn't do well
- Identifiable real people. Generating public figures is blocked, and rephrasing the request doesn't get around it.
- Brands and licensed characters. Logos of existing companies and protected characters are refused.
- True vector. The output is always rasterized. For an editable file in curves, the result works as a reference to redraw.
- Reliable transparent background. Even when requested, it tends to come with a solid background. Transparency is faster to solve afterward.
- Absolute fidelity when editing your photo. When changing an uploaded image, the model regenerates the scene and fine details change.
- High volume without a plan. The daily quota for free use is limited.
Seven mistakes that hurt the most
- Using the Pro model for testing. Spends quota that would be needed for the final version.
- Writing in tag format. A word list is the language of other tools; here, prose performs better.
- Not stating the aspect ratio. With no indication, the format is almost never what you need.
- Rewriting the prompt on every attempt. You lose what you'd already gotten right. Ask for specific fixes.
- Uploading a reference with a cluttered background. The model carries over elements from the setting that you didn't ask for.
- Saving via screenshot. Loses resolution and adds recompression.
- Publishing without optimizing. 2K images are heavy and tank the loading time of any page.
After generating: preparing the image
The image that comes out of Gemini is ready to be viewed, not to be published. Four steps solve it.
Convert the format. PNG is great for transparency and bad for photos on the web. For a website, WebP tends to cut the file size in half with no visible difference; for general sharing, JPG works. The image converter does the swap in the browser, and the guide on image formats helps you choose.
Reduce the file size. A 2K image easily goes over several megabytes. The image compressor cuts a good chunk of that with no visible loss.
Adjust size and framing. For a specific measurement, use the resizer; if the aspect ratio came out different from what you need, cropping preserves more quality than stretching.
Isolate the subject when needed. For products, logos, and elements that will go over a colored background, removing the background gives you the transparent PNG the generation doesn't produce.
💡 Order that works: download at full resolution → crop or resize → convert the format → compress last. Compressing before resizing wastes quality.
Checklist before generating
- Does the chosen model match the complexity of the request?
- Is the aspect ratio stated?
- Is the prompt in prose, with subject, action, and scene?
- If there's text, is it in quotes with a defined position?
- Do the uploaded references have a clean background?
- Are you asking for a fix instead of rewriting everything?
- Are you going to download at full resolution, not via screenshot?
Convert your images to the right format
Turn Gemini's heavy PNG into WebP or JPG and get the image ready to publish, right in your browser.
Convert image