Almost everyone who starts generating images with AI spends months trying to improve results through the prompt alone: swapping "beautiful" for "stunning", piling on "8K", "ultra detailed", "masterpiece". And the output barely changes. The reason is simple: much of what determines the final image isn't in the text at all — it's in the parameters.
Parameters are the numeric controls that tell the model how to draw, not what to draw. Canvas proportions, how strictly it should follow your text, how many refinement steps it runs, which random seed it starts from, how much aesthetic freedom it gets. They're what separates someone generating ten random variations from someone who can repeat, adjust and evolve an image on purpose.
This guide is a technical reference for the five parameters that show up in practically every tool — aspect ratio, seed, CFG, stylize and steps — plus a few neighbouring controls worth knowing. The idea is that you can come back whenever you need a reminder of what each value does, which range to adjust, and what breaks when you push too far.
Why parameters matter more than adjectives
A diffusion model starts from pure noise and progressively "cleans" that noise until an image emerges. The prompt works like a map of the destination; the parameters define the road, the speed, and how much the driver is allowed to improvise along the way.
That explains why two identical prompts produce completely different images, and why adding more adjectives often changes nothing. If the model already has low adherence to the text (low CFG), piling on words won't help — it simply isn't paying much attention to them. If steps are too low, no adjective will rescue the texture of the image, because the model didn't have enough iterations to resolve the detail.
The habit that separates beginners from experienced users is this: change one parameter at a time. If you swap the prompt, the seed and the CFG in the same attempt, you won't know what caused the improvement. Lock the seed and adjust only the CFG, and you isolate the variable and actually learn how that model behaves.
💡 Tip: before optimising anything, generate four images with the same prompt and the default parameters. They show you the model's baseline. Only then start changing values — and always one at a time.
Aspect ratio: the proportion that defines the framing
Aspect ratio is the relationship between the width and the height of the image. In Midjourney it's written as --ar 16:9; in Stable Diffusion and its derivatives you set width and height in pixels; in ChatGPT and Gemini you can usually just ask in plain language ("portrait 9:16 format").
It looks like the most trivial parameter on the list, but it's the one with the biggest influence on composition. The model doesn't simply crop the same scene into different shapes — it recomposes. A portrait requested in 16:9 tends to become a medium shot with a wide background; the same prompt in 2:3 becomes a classic vertical photographic close-up. The proportion changes what the AI understands as "natural framing" for that scene.
The most common values and what they're for
| Ratio | Typical use | What the model tends to do |
|---|---|---|
| 1:1 | Avatar, Instagram feed, icon, product thumbnail | Centres the subject, simplifies the background |
| 4:5 | Vertical Instagram and Facebook posts | Medium shot, strong subject presence |
| 2:3 | Photographic portrait, poster, book cover | Close-up or American shot, editorial feel |
| 9:16 | Stories, Reels, TikTok, phone wallpaper | Vertical composition, lots of space above and below |
| 16:9 | YouTube thumbnail, banner, cinematic scene | Wide shot, landscape, surrounding context |
| 21:9 | Ultrawide cinematic, website hero | Panoramic, subject small within the frame |
| 3:1 | Header banner, LinkedIn cover | Horizontal strip; usually needs a very specific prompt |
One frequent trap: very extreme ratios (like 3:1 or 1:3) confuse models trained mostly on images closer to square. Duplication is common — two faces, two heads, elements repeated along the edges — because the model tries to "fill" a space it has no strong reference for. In those cases the safer route is to generate in 16:9 and then use cropping or outpainting to reach the final shape.
Another practical point: generating at the right ratio from the start always beats cropping afterwards. Cropping discards pixels and reduces usable resolution. If you need the same image in several formats, generate in the most restrictive one (usually the vertical) and adapt the rest with controlled resizing.
Seed: the number that makes results repeatable
The seed is the integer that initialises the random noise the image grows out of. Same prompt + same seed + same parameters + same model = the same image, pixel for pixel. Change the seed and everything changes.
This is probably the most underused parameter among beginners, and the most important one for anyone working with AI professionally. Without controlling the seed you can't iterate — you can only roll the dice again.
There are three practical uses:
- Refining an almost-good result. You generated something with perfect composition but the wrong colour. Lock the seed, change only the colour word in the prompt and generate again. Most of the structure holds, and only the altered element changes.
- Comparing parameters honestly. Want to know whether CFG 7 beats CFG 11 for your prompt? Lock the seed and generate both. Without a fixed seed you're comparing two different images, not two CFG values.
- Consistency across a series. In a product catalogue, a sequence of illustrations or a visual identity, keeping the seed helps preserve the "signature" of the set.
One important limitation: the seed is specific to the model and the version. Seed 42 in SDXL produces nothing resembling seed 42 in Flux. Switch models and your old seeds become meaningless numbers. So always record model + version + seed together, never the seed on its own.
💡 Tip: when a generation comes out well, immediately save the full set: prompt, negative prompt, seed, CFG, steps, sampler, model and ratio. A simple spreadsheet with those columns is worth more than any folder of loose images — without that data, a great result is irreproducible.
CFG scale: how strictly the AI follows your prompt
CFG stands for Classifier-Free Guidance, and the CFG scale controls how much weight the model gives your text versus what it would "prefer" to draw on its own. In practice, it's an obedience dial.
Low values leave the model creative and loose, but it starts ignoring parts of the prompt. High values force literal adherence, but the image turns saturated, with hard outlines, exaggerated contrast and an artificial look — the classic "overcooked" result.
| CFG range | Behaviour | When to use it |
|---|---|---|
| 1–3 | Nearly ignores the prompt; unpredictable, ethereal output | Artistic exploration, abstract textures |
| 4–6 | Free interpretation, soft colours, more natural look | Photorealism, portraits, organic scenes |
| 7–9 | Balance between obedience and naturalness — the default range | General use; recommended starting point |
| 10–14 | Strong adherence to the text, more intense colours | Illustration, graphic design, specific compositions |
| 15+ | Saturation, artefacts, hard edges, distortion | Rarely useful; only in very specific cases |
One detail that confuses a lot of people: newer models use different ranges. Flux, for example, works on its own guidance scale where values between 2 and 4 are normal — setting 7 there is already high. Turbo and distilled models, built to run in very few steps, usually want CFG between 1 and 2. In other words: there's no universal number, only the right number for that model. Always check the recommendation from whoever published the checkpoint before assuming 7 is the default.
Practical symptom of CFG being too high: plastic-looking skin, neon colours, a dark halo around outlines, text or patterns repeated aggressively. Symptom of CFG too low: the image is beautiful, but half of what you asked for isn't in it.
Stylize: the model's own aesthetic
The --stylize parameter (or --s) is specific to Midjourney and controls how much of its own artistic sensibility the model applies on top of your request. It's almost the conceptual opposite of CFG: instead of measuring obedience to the text, it measures aesthetic freedom.
The scale runs from 0 to 1000, with 100 as the default. In practice:
- 0–50: literal. The model tries to draw exactly what you described, without "improving" anything. Useful for diagrams, technical references, layouts and anything where fidelity matters more than beauty.
- 100–250: balanced. Good composition and palette without drifting from the brief. This range handles most cases.
- 250–750: the model takes over. Dramatic lighting, bold palettes, more artistic compositions. Great for posters, concept art and covers.
- 750–1000: style above all. The image comes out beautiful and frequently far from what you asked for. Good for creative exploration, bad for a fixed brief.
There's a sibling parameter that often gets confused with it: --chaos (or --c, from 0 to 100), which controls the variety between the four images in a batch, not the style of each one. High chaos returns four radically different interpretations of the same prompt — excellent for early brainstorming, terrible once you already know what you want. There's also --weird, which injects deliberate strangeness and belongs more to experimentation than to client work.
Outside Midjourney there's no direct stylize equivalent. A similar effect comes from other routes: style LoRAs, aesthetic keywords in the prompt, or fine-tuning the CFG. The concept is still worth knowing, because the logic of "how much personality the model may add" shows up in different forms in nearly every tool.
Steps: how many iterations the model runs
Steps (or sampling steps) is the number of denoising stages. Each step brings the image a little closer to the final result. More steps means more processing time — and, up to a point, more detail.
The key phrase is up to a point. This parameter has very clear diminishing returns: the difference between 10 and 25 steps is enormous; between 30 and 60 it's subtle; above 80 it's practically imperceptible, and you're only burning time and energy.
| Steps | Result | Note |
|---|---|---|
| 1–8 | Only works on turbo/distilled models | On normal models it comes out blurry and incomplete |
| 10–20 | Quick draft, defined shapes, weak detail | Good for testing composition and prompt |
| 25–35 | Production quality in most cases | Recommended default range |
| 40–60 | Marginal gain on very complex scenes | Only worth it with many elements in frame |
| 80+ | No perceptible gain | Wasted time and processing |
Steps and sampler are directly linked. Samplers like DPM++ 2M Karras converge quickly and deliver good results around 25 steps. Samplers like Euler a (ancestral) never fully converge — the image keeps changing as you raise the step count, which can be good (more variation to explore) or bad (it hurts reproducibility). If consistency is your goal, favour non-ancestral samplers.
A practical workflow decision: run early exploration at low steps (15–20) to test dozens of prompts quickly, and only regenerate the final version, with the winning seed, at 30–40 steps. That saves a lot of time compared to running everything at maximum quality from the start.
The neighbouring parameters also worth knowing
Beyond the main five, a few controls come up often and solve specific problems:
- Negative prompt: a list of what you don't want in the image (deformed hands, watermark, text, low quality). Present in Stable Diffusion and derivatives; in Midjourney it exists as
--no. It's one of the highest-impact, lowest-effort adjustments — worth reading the dedicated guide on negative prompts. - Denoising strength: in image-to-image, it defines how much the original will be altered. Values near 0.2 make subtle retouches; above 0.7 the original becomes only a vague suggestion of composition.
- Reference image weight (
--iwin Midjourney): how much the uploaded image weighs against the text. Essential for character consistency. - Sampler / scheduler: the algorithm that performs the denoising. It changes convergence speed and final texture; DPM++ 2M Karras and DDIM are safe choices.
- Tile / seamless: generates textures that repeat without visible seams, for patterns and backgrounds.
- Model version (
--vin Midjourney): different versions have different aesthetics and parameter ranges. Changing version invalidates your previous seeds.
Which parameters exist in each tool
Naming varies a lot between platforms, which causes much of the confusion. This table maps the equivalents:
| Tool | Aspect ratio | Seed | CFG / guidance | Steps | Stylize |
|---|---|---|---|---|---|
| Midjourney | --ar 16:9 | --seed 1234 | Not exposed | Not exposed | --s 100 |
| Stable Diffusion / ComfyUI | Width × height | Seed field | CFG Scale | Sampling steps | Via LoRA or prompt |
| Flux | Width × height | Seed field | Guidance (2–4) | 20–30 typical | Via LoRA or prompt |
| ChatGPT / GPT Image | Plain language | Not exposed | Not exposed | Not exposed | Via prompt |
| Gemini / Nano Banana | Plain language | Not exposed | Not exposed | Not exposed | Via prompt |
| Leonardo AI | Presets + custom | Seed field | Guidance Scale | Yes | Style presets |
Notice the pattern: the more consumer-facing the tool (ChatGPT, Gemini), the fewer parameters it exposes — the platform decides for you. The closer to the raw model (ComfyUI, Automatic1111), the more control you get. Neither approach is superior; they serve different goals. If your work depends on repeating results, you'll eventually need a tool that exposes seed and CFG.
How parameters interact: recipes by goal
Parameters don't work in isolation. These combinations are tested starting points — adjust them to your model:
| Goal | Ratio | CFG | Steps | Note |
|---|---|---|---|---|
| Photorealistic portrait | 2:3 or 4:5 | 4–6 | 30 | Low CFG avoids plastic skin |
| Product photo | 1:1 | 7–9 | 30–35 | Neutral background in the prompt |
| YouTube thumbnail | 16:9 | 9–12 | 25 | High CFG for strong colours |
| Illustration / concept art | 16:9 or 21:9 | 7–10 | 30 | High stylize in Midjourney |
| Logo and branding | 1:1 | 8–12 | 30 | Low stylize for fidelity |
| Phone wallpaper | 9:16 | 6–9 | 35 | Leave clear space for icons |
| Quick prompt test | 1:1 | 7 | 15 | Just to validate the idea |
The logic behind these choices is consistent: photorealism wants low CFG (the model already knows what a photo looks like; forcing it ruins the effect); graphic design and promotional material want high CFG (you need exactly those elements, in those colours); and steps only go up when there's a lot of fine detail competing for space in the scene.
Common mistakes and how to fix them
- Changing everything at once. New prompt, new seed, new CFG. The result improved, but you don't know why — and you won't be able to repeat it. One variable per test.
- Raising steps to fix a problem that isn't about steps. Deformed hands, wrong anatomy or a missing object won't be solved by more iterations; those are prompt, model or extreme-ratio problems.
- Copying CFG from an old tutorial. A value recommended for SD 1.5 back in 2023 can be terrible in Flux or SDXL. Every model family has its own range.
- Ignoring the seed. Without it, every attempt is a fresh roll of the dice and you never build on what already worked.
- Starting with extreme ratios. Instead of 3:1 straight away, generate 16:9 and expand with outpainting — fewer duplications and artefacts.
- Confusing resolution with quality. Generating far above the resolution the model was trained on tends to create duplicated subjects. Better to generate at native resolution and upscale afterwards.
- Recording nothing. The perfect image you can't reproduce is an accident, not a result.
Quick checklist before generating
- Choose the ratio based on the image's final destination, not the tool's default.
- Start from the model's default values — and confirm what they are, because they vary.
- Run a cheap test: low steps, four variations, just to validate the direction.
- Pick the best variation and write down the seed.
- With the seed locked, adjust one parameter at a time until it lands.
- Regenerate the final version at production steps (30–40).
- Record prompt, negative, seed, CFG, steps, sampler, model and ratio.
- Do the post-production: upscale, crop, sharpening and compression for the web.
After generating: the step almost everyone skips
An AI-generated image is rarely ready to publish. It usually arrives as a heavy PNG, at the model's native resolution, without the exact crop your channel requires. That final treatment stage is what separates an image that reads as "AI output" from an image that simply works on your site or profile.
The basic post-production routine is short: adjust the framing with cropping when the final ratio doesn't match exactly; bring it to the exact target dimensions with resizing; recover micro-detail with sharpening if the image was enlarged; and finally convert to WebP and compress it with the compressor, because a 4 MB PNG at the top of the page drags down PageSpeed no matter how good the artwork is.
If the image is going to become part of a visual identity, the path continues into other tools: pull the palette with the colour palette generator to keep the pieces coherent, and use the AI logo maker when the goal is branding.
Get your AI images into the right format
Once the parameters are dialled in, bring the image to the exact dimensions of its destination — free, straight in your browser, with nothing uploaded to a server.
Resize an image now