AI Image Prompts: A Practical Guide to Getting the Shot You Want

AI Image Prompts: A Practical Guide to Getting the Shot You Want

The gap between a mediocre AI image and a genuinely good one is almost never the tool. It is that one prompt made decisions and the other one did not.

“A beautiful woman in a city” gives the model nothing to act on, so it produces the statistical average of every similar image it has seen. That average is what people mean when they say AI images look generic.

This is a practical guide to making those decisions, in the order that matters most.

Start with the subject, specifically

Most prompts fail at the first word. “A woman”, “a man”, “a dog” are placeholders, not subjects.

Give the model something to hold onto: age, build, expression, clothing, one distinguishing detail. “A woman in her sixties with short grey hair and deep laugh lines, wearing a worn denim jacket” is a person. The model now has constraints to work within instead of a blank to fill.

The test: could two people read your subject description and picture noticeably different people? If yes, the model is choosing for you.

Then say what is happening

This is the step people skip most often, and it is the one that separates images that feel alive from images that feel like a catalogue.

A subject standing and looking at the camera is a portrait. A subject mid-action in a specific place is a scene. “Sitting at a cluttered kitchen table, hands wrapped around a mug, mid-conversation” gives the image a moment, and moments are what make pictures interesting.

Environment does the same work. “In a city” is nothing. “On a narrow street after rain, neon reflecting off wet tarmac” is a place.

One subject, four style decisions

The same subject rendered four ways. Style choice changes more than any other single word.

35mm film Grain, soft roll-off, natural color, shallow depth of field Reads as photographed Oil painting Visible brushwork, canvas texture, richer shadow, looser edges Reads as painted Vector illustration Flat color, clean shapes, no texture, strong silhouette Reads as designed 3D clay render Matte surfaces, soft studio light, rounded forms, subtle shadow Reads as modelled Add these next, in order 1   Distance and angle — close-up, wide shot, low angle, overhead 2   Light source and quality — soft window light, harsh midday sun, single candle 3   Time of day — golden hour, blue hour, overcast noon 4   Technical detail — lens, aspect ratio, depth of field. Smallest effect, added last.

Naming a medium is the biggest single lever

If you change one thing about how you prompt, change this. A medium carries an entire visual grammar with it — lighting behavior, color, texture, edge quality — and applying a different one to the same subject transforms the result completely.

“Artistic” and “high quality” mean nothing. “Charcoal on textured paper, heavy smudging, high contrast” means something executable.

On naming living artists: it raises genuine ethical and legal questions, and some platforms restrict it. Describing the qualities you want — the brushwork, the palette, the era — gets you most of the way without appropriating a working artist’s name.

Light and framing do the emotional work

Most people describe what is in the image and never how it is seen. That is why so much generated work feels flat and centred — nothing told the model otherwise.

  • Distance. Extreme close-up, medium shot, wide establishing shot. Changes the emotional register entirely.
  • Angle. A low angle makes a subject imposing without you writing “imposing”. An overhead shot makes them small.
  • Light source and quality. Soft window light from the left, hard midday sun, a single candle, neon signage. This is the most reliable way to stop an image looking generated.
  • Time of day. Golden hour, blue hour, overcast noon — each brings its own color temperature and shadow behavior.

The words that do nothing

A set of terms circulate endlessly in prompt guides and mostly function as reassurance rather than instruction.

  • “Masterpiece”, “award-winning”, “best quality”. These describe your hopes. Modern models already aim for coherent output.
  • “8k”, “ultra HD”, “highly detailed”. Resolution comes from your generation settings. These sometimes just push toward a glossy digital-art look you may not want.
  • “Trending on ArtStation”. A relic of earlier models that mostly pulls toward one era of digital fantasy illustration.
  • Twenty comma-separated adjectives. Attention gets diluted. Five specific decisions beat twenty vague ones.

Every one of these occupies space a real decision could occupy.

Negative prompts, where supported

Some tools offer a separate field for what you do not want. It works well for concrete objects and badly for abstract qualities.

“No text, no watermark, no extra fingers” is actionable. “Not ugly, not boring” is not, because the model has no stable representation of either.

Where there is no negative field, phrase positively in the main prompt. “Empty street” beats “street with no people”, because naming people at all raises the chance they appear.

Change one thing at a time

The most common workflow error is rewriting the entire prompt when a result disappoints. You then have no idea which change helped.

  1. Start with subject, action and context only. Generate.
  2. Add the medium. Generate. Keep it if it moved toward what you wanted.
  3. Add light and framing. Generate.
  4. Only then touch technical parameters.

Where the tool allows it, fix the seed while testing. A fixed seed means the only thing changing is your word, which is the entire point of the exercise.

Where models still fail

  • Text in images. Improving quickly, still unreliable beyond a few words. Add important text afterwards in an editor.
  • Counting. “Exactly five birds” is a request, not a guarantee.
  • Spatial relationships. “The red cup left of the blue book” frequently comes out reversed.
  • Consistent characters. Getting the same face twice is genuinely hard without dedicated features.
  • Hands. Much improved, still the first thing to check before calling an image finished.

When a prompt keeps failing on one of these, change the composition so the problem disappears rather than rewording endlessly. Framing a portrait with the hands out of shot solves the hands problem entirely.

Frequently asked questions

How long should an image prompt be?

Twenty to sixty words of specific description usually outperforms both a five-word prompt and a two-hundred-word one. Long enough to remove ambiguity, short enough that every word is doing work.

Do prompts work across different tools?

Descriptive language transfers well. Tool-specific parameters and syntax do not. Writing in plain description keeps your prompts portable.

Can I use AI images commercially?

It depends on the tool’s terms and on your jurisdiction, and the legal position on copyright in AI-generated work is still developing in several countries. Check the specific platform’s license and confirm your local position before using generated images in client work.

Why do I get a different image every time?

Generation is random by design, seeded differently each run. Fixing the seed where possible gives you a reproducible base for testing changes.

Build a library, not a keyword list

The specific tricks age quickly — half the advice from two years ago is now useless. What does not age is knowing what you want and describing it precisely.

Keep the prompts that worked. A personal collection of proven starting points beats any list of magic words, because it is calibrated to the tools you use and the results you actually like.

If you would rather start from prompts already tested, our Android app GenZ AI Prompts is a curated library organized by category — copy something proven and adjust from there. For prompting text models rather than image models, see how to write AI prompts that actually work.