Creating images in ChatGPT has stopped being a side feature and become one of the things the tool does best. Since April 2026, image generation runs on ChatGPT Images 2.0, a native model that fully replaced DALL·E — retired in May of the same year — and that fixed the most persistent problem in all image AI: writing text inside a picture without getting the letters wrong.
This changes the kind of thing you can do. A poster with a legible title, an infographic with correct numbers, a cover with the brand name, a comic book page with dialogue, a packaging mockup with a label — all of it used to come out crooked before, and now comes out clean on most attempts.
This guide covers the complete path: how to generate your first image, how to write the prompt the way ChatGPT understands best, how to edit without regenerating everything, how to keep the same character across images, what changes between the free and paid plans, and what the tool still doesn't do well.
What changed in ChatGPT's image generation
Until 2025, ChatGPT generated images by calling DALL·E as a separate system: you'd write something, it would translate your request into an internal prompt and send it to another model. The result was inconsistent, because there was a translation step in the middle that you didn't control.
Today the image model is native and shares context with the language model. In practice this means it understands the request the same way it understands a text question: if you say "make a poster for a café that opens at 7am in downtown," it doesn't need you to describe every visual element — it infers layout, hierarchy and content.
Three practical consequences matter more than the others:
- Text inside the image works. Titles, paragraphs, labels, captions, and even text in multiple languages come out legible. This used to be image AI's biggest Achilles' heel.
- Complex layout is possible. Infographics, magazine spreads, character sheets, and multi-panel pages stopped being a matter of luck.
- Consistency between images improved. You can generate a set of images while keeping the same character or the same product.
💡 Worth knowing: the image-generation limits on the free plan aren't officially published and change with platform demand. Treat any number you see out there as an estimate, not a fixed rule.
How to generate your first image, step by step
The workflow is as simple as it gets — there's no separate panel or mandatory pre-configuration.
- Open a new conversation. There's no need to select any special mode; just ask.
- Describe the image in natural language. Write as if you were explaining it to a designer, in complete sentences. Don't use a comma-separated list of words: that format belongs to other tools and performs worse here.
- State the aspect ratio. "Vertical format for stories," "square for Instagram," "horizontal 16:9." If you don't say, it chooses on its own.
- Wait for the generation. Images take longer than text answers, especially ones with dense layouts.
- Ask for adjustments in the same conversation. "Make the background darker," "make the title bigger," "change the coat color to red." It keeps the image and changes only what you asked for.
- Download at full resolution. Use the download button, not a screenshot — a screenshot loses quality and comes out at the size of your window.
That last point matters more than it seems. A lot of people save the image by right-clicking the thumbnail and end up with a small, recompressed file.
The two modes: instant and reasoning
ChatGPT Images 2.0 operates in two modes, and understanding the difference avoids frustration.
Instant mode generates directly from your request. It's fast, available on every plan including the free one, and covers simple images well: a scene, a portrait, an illustration, a background.
Reasoning mode — the "thinking" mode — makes the model plan before drawing: it checks the number of objects, verifies the constraints you asked for, organizes the layout, and, when needed, searches the web to bring current information into the image. This is the mode that delivers a correct infographic, a poster with real data, and a composition with many aligned elements. It's restricted to paid plans.
In practice: if your request is a single scene, instant mode is enough. If it involves counting ("exactly five icons"), text hierarchy, or data that needs to be accurate, reasoning makes a real difference.
How to write prompts ChatGPT understands best
Here's the biggest difference between ChatGPT and tools like Midjourney: it prefers prose to a list of tags. Writing "woman, 35 years old, gray blazer, office, 85mm, bokeh" works, but performs worse than the same information written in sentences.
Start with the image's purpose, not the description
Saying what it's for changes the entire result. "A promotional poster for the photography workshop on the 12th" produces something far more usable than "an image with a camera and some text." The model uses the purpose to decide hierarchy, spacing and style.
Write the exact text in quotes
When the image needs to contain words, write them literally and in quotes: the title "Morning Coffee," the subtitle "every Saturday, 8am to 11am." Without quotes, the model may paraphrase or invent. And say where each piece of text goes: top, footer, bottom-right corner.
Describe the structure before the style
The order that works is: what appears and where → then what it looks like. "Three equal columns, an icon at the top of each, a title below, and a short paragraph underneath; minimalist style with a beige background and sans-serif typography." Structure first means much less rework.
Be explicit about what you don't want
ChatGPT doesn't have a negative-prompt field. The restriction goes inside the request itself, and works better stated affirmatively: "clean composition, with no text beyond the title" performs better than "no text."
Iterate instead of rewriting
This is the biggest advantage of the conversational interface. Don't rewrite the prompt from scratch when something comes out wrong: ask for the fix. "Keep everything, but change the background to navy blue" preserves what was already good. Rewriting the entire prompt generates a new image unrelated to the previous one.
💡 Useful shortcut: ask ChatGPT itself to improve your prompt before generating. "Rewrite this request as a detailed image prompt, without generating anything yet" often reveals details you hadn't thought of.
Aspect ratio and resolution: what to ask for
The model accepts a wide range of formats, from ultra-wide to ultra-tall, and generates up to 2K. Setting the aspect ratio before generating avoids cropping afterward — and cropping afterward always costs composition.
| Destination | How to ask for it | Aspect ratio |
|---|---|---|
| Instagram feed | square format | 1:1 |
| Vertical / portrait feed | portrait format for feed | 4:5 |
| Stories and Reels | vertical stories format | 9:16 |
| YouTube thumbnail | widescreen horizontal format | 16:9 |
| Website banner | wide panoramic format | 3:1 |
| E-book cover | vertical book-cover format | 2:3 |
| Desktop wallpaper | widescreen horizontal format | 16:9 |
If you need a specific resolution — 1920 by 1080, for example — generate at the right aspect ratio and adjust afterward with the image resizer. Asking for exact pixel dimensions inside the prompt is rarely followed to the letter.
Editing images: the commands that work
Editing is where ChatGPT stands out compared to pure generation tools. You don't need a mask, a layer, or a selection — you describe it in words.
Change a specific element
"Keep the composition and change only the wall color to sage green." The more explicit you are about what should stay the same, the less the model touches the rest. The word "keep" does a surprisingly large amount of work here.
Add or remove objects
"Remove the plant in the right corner and leave the wall plain." Works well with clearly delimited objects. Removals in areas with heavy background texture sometimes leave marks — in that case, generating again tends to work out better than insisting on the edit.
Change the style while keeping the scene
"Same composition, same elements, but in watercolor style." This is useful for testing different visual languages without losing the framing you already approved.
Editing from your own image
You can upload a photo and ask for changes: swap the background, adjust the lighting, turn it into an illustration, create variations. It's worth remembering that the result is a new image generated from yours, not a pixel-by-pixel edit of the original.
Expanding the framing
"Expand this image to the horizontal format, naturally continuing the scenery to the sides." This is the conversational equivalent of outpainting, and it solves the classic case of having a vertical image that needs to become a banner.
How to keep the same character across multiple images
This used to be the hardest problem in image AI, and it's become much more manageable. Three approaches, from simplest to most reliable.
Continue in the same conversation. This is the easiest method and the most effective for everyday use. After generating the character, ask for the next scene without re-describing them: "now show the same character sitting in a café." The model carries the conversation's context and keeps the traits.
Describe fixed traits in writing. Build a short, unchanging description — hair, eyes, characteristic clothing, a distinctive accessory — and repeat the exact same words in every request. Synonyms get in the way: "beige coat" and "sand-colored overcoat" produce different pieces.
Ask for a reference sheet first. Generate a character sheet with the same face at several angles and expressions, and use that image as a reference in the following generations. This is the most laborious path and the most consistent, especially for long series.
The same logic applies to products and brand identity: the more fixed the description, the more stable the result.
Free or paid: what actually changes
Image generation is available on every plan, including the free one. What varies is volume, speed, and access to reasoning mode.
| Aspect | Free plan | Paid plans |
|---|---|---|
| Image generation | Available, instant mode | Available, instant and reasoning |
| Volume | Limited, quota not published | Much higher |
| Speed | Slower, subject to queuing | Priority at peak hours |
| Web search during generation | No | Yes, in reasoning mode |
| Complex layouts and infographics | Inconsistent result | Much more reliable |
| Consistent image sets | Limited | Yes |
For occasional personal use — a wallpaper, an illustration for a post, a supporting image — the free plan works. For recurring work with layout, text and volume, the difference shows up fast. If you want to compare it with other options before subscribing, the guide to the best AI tools to create images puts the main names side by side, and the Gemini guide for creating images covers the main alternative.
What ChatGPT still doesn't do well
Knowing the limitations saves time. None of them are prompt flaws: they're limits of the tool.
- Real people. Generating images of identifiable people, especially public figures, is blocked. Rephrasing the request doesn't help.
- Brands and licensed characters. Logos of existing companies and copyright-protected characters are also refused.
- True vector. The output is always rasterized, in pixels. If you need an editable file in curves, the result works as a reference to redraw, not as the final file.
- Reliable transparent background. Even when requested, the image tends to come with a solid background. Transparency is faster to solve afterward, by removing the background.
- Resolution above 2K. For large-format printing or 4K monitors, you need to upscale afterward.
- Exact reproduction of your own photo. When editing an uploaded image, the model regenerates the scene. Fine details from the original change.
Seven mistakes that hurt the most
- Writing in tag format. A comma-separated list of words is the language of other tools. Here, prose performs better.
- Not stating the aspect ratio. With no indication, the model chooses — and almost always picks the wrong format for your use.
- Rewriting the prompt on every attempt. You lose what you'd already gotten right. Ask for specific fixes instead.
- Leaving the text loose in the request. Without quotes and without a position, the model paraphrases or places it wherever it wants.
- Saving via screenshot. Loses resolution and adds recompression. Use the download.
- Asking for everything at once. Prompts with ten simultaneous requirements tend to fail on two or three of them. Generate the base and adjust in steps.
- Publishing without optimizing. A 2K PNG image weighs several megabytes and tanks the loading time of any page.
What to do with the image after generating
This last point deserves detail, because it's where most people go wrong. The image that comes out of ChatGPT is ready to be viewed, not to be published.
Convert to the right format. PNG is great for transparency and terrible for photos on the web. For a website, WebP tends to cut the file size in half with no visible difference; for general sharing, JPG works. The image converter does the swap right in the browser, and the guide on image formats helps you choose.
Reduce the file size. A 2K image can go over 5 MB. The image compressor cuts a good chunk of that with no visible loss — and if the file is really large, the walkthrough on how to reduce an image from MB to KB covers the details.
Isolate the object when needed. For logos, products, and elements that will go over a colored background, removing the background and saving as a transparent PNG is the missing step.
Adjust resolution and framing. If the image came out smaller than needed, you can increase the resolution without losing quality. If it came out in the wrong aspect ratio, cropping preserves more than stretching.
Reuse the palette. To keep visual consistency across images generated in different sessions, extract the colors from the one that came out best and reference those tones in your next prompts.
💡 Order that works: download at full resolution → crop or resize → convert the format → compress last. Compressing before resizing wastes quality.
Three complete example requests
To close, three requests written the way ChatGPT makes the most of — notice that all of them state the purpose, the exact text, and the aspect ratio.
Checklist before generating
- Did you say what the image is for, not just what it shows?
- Is the aspect ratio stated?
- Is the text that should appear in quotes, with a defined position?
- Was the structure described before the style?
- Are the constraints written affirmatively?
- Are you asking for a fix instead of rewriting everything?
- Are you going to download at full resolution, not via screenshot?
Get your images ready to publish
Images generated at 2K easily go over 5 MB. Reduce the file size with no visible quality loss, right in your browser.
Compress image