AI AI images, video and voice: how they are made
Writing image prompts: subject, style, light and framing
A structure for image prompts that works across tools, the words that reliably change results, and how to fix a picture that is close but not quite right.
The short answer
- A reliable image prompt has five slots: subject, action or pose, setting, style and medium, and lighting with camera detail.
- Put the most important element first, because words near the front of a prompt generally carry more weight than words near the end.
- Concrete craft vocabulary such as 35mm lens, overcast light or linocut changes output far more than praise words like beautiful or masterpiece.
- Fix one variable per revision, and keep the seed fixed while you do it, so you can tell which word caused the change.
- Name a living artist and you may get their style, but you inherit an unresolved legal and ethical problem along with it.
The difference between a mediocre AI image and a good one is almost never effort, it is structure. Most disappointing results come from prompts that describe a subject and nothing else, leaving the model to invent the setting, the light, the medium and the camera position. It will invent them toward the bland average of the pictures it was trained on. Fill those slots yourself and the output changes immediately, on every tool, because they all respond to the same kind of description.
The five slots
Write your prompt as five parts, in this order.
- Subject. Who or what, with the two or three attributes that matter. "An elderly fisherman with a white beard" beats "a man".
- Action or pose. What they are doing, or how they are arranged. "Mending a net, seated on an upturned crate."
- Setting. Where, and what surrounds them. "On a stone harbor wall, fishing boats blurred behind."
- Style and medium. What kind of picture this is. "Documentary photograph" or "oil painting on canvas" or "flat vector illustration."
- Lighting and camera. "Overcast late afternoon light, 85mm lens, shallow depth of field."
Assembled: "An elderly fisherman with a white beard, mending a net while seated on an upturned crate, on a stone harbor wall with fishing boats blurred behind, documentary photograph, overcast late afternoon light, 85mm lens, shallow depth of field."
That is 40 words and every one of them is doing a job. Order matters because most models weight earlier words more heavily, so whatever you would be most annoyed to lose belongs in the first clause. If you are new to prompting in general, the same discipline that works for writing a good text prompt applies here: specifics beat adjectives.
One prompt, four revisions
Start: "a coffee shop."
You get a generic bright cafe, shot straight on, nobody in it.
Revision 1, add subject and action: "a barista pouring latte art into a white cup, in a small coffee shop." Now there is a focal point, but the room is still bland and the pour looks staged.
Revision 2, add setting detail and medium: "a barista pouring latte art into a white cup, in a narrow wood paneled coffee shop with a steamed up window, 35mm documentary photograph." The room has character and the flat commercial look is gone.
Revision 3, add light and camera: "... warm tungsten light from above, cool daylight through the steamed window behind, 35mm at f/2, focus on the cup, slight motion blur in the pouring milk." This is where most of the improvement happens. Two light sources of different color temperature is one of the single most effective things you can specify, because it is how real interiors actually look.
Revision 4, fix what is wrong: the hands came out badly and there is unreadable text on a chalkboard, the two flaws that give a generated image away fastest. Crop tighter with an aspect ratio change to 4:5, add "no signage" or remove the chalkboard from the description, and inpaint the hands rather than rerolling the whole scene.
Vocabulary that actually moves the output
| Instead of | Say |
|---|---|
| Good lighting | Golden hour backlight, or overcast soft light, or single hard key light from the left |
| Close up | 85mm portrait lens, head and shoulders, shallow depth of field |
| Wide shot | 24mm, low angle, full body in frame |
| Vintage | Kodachrome slide film, or 1970s color print with faded cyan shadows |
| Artistic | Charcoal on toned paper, or gouache, or linocut in two colors |
| High quality | Nothing. Spend the words on a real detail instead |
| Cinematic | Anamorphic lens flare, 2.39:1 framing, practical lights in shot |
Aspect ratio deserves its own line because it changes composition more than any adjective. A 1:1 square pushes the model toward centered portraits, 16:9 invites landscapes and environmental context, 9:16 forces a vertical subject and usually a tighter crop, and 4:5 is the classic portrait frame. Set it deliberately rather than accepting the default.
Time of day is similarly powerful: "midday" gives hard overhead shadows, "blue hour" gives flat cool ambience with warm windows, "night, lit by a single streetlamp" gives you dramatic falloff for free.
Iteration mechanics: seeds, variations and local fixes
A seed is the starting random number for the image. Same prompt plus same seed plus same settings equals the same picture, every time. That makes the seed your control variable. Lock it and change one word, and any difference you see is caused by that word. Reroll it and you get a different composition of the same idea. The mechanism behind this is explained in how AI image generators turn noise into a picture.
A working loop that wastes very little time:
- Generate four images from a rough prompt with random seeds, at low resolution if the tool offers it.
- Pick the one whose composition you like and note its seed.
- Lock that seed and refine the wording, one change at a time, until the content is right.
- Upscale the winner.
- Inpaint the remaining flaws: hands, eyes, a merged strap, unreadable text.
Variations sit between steps 2 and 3 in some tools: they keep the overall composition and nudge the details, which is useful when the layout is right but a pose is stiff.
Style, artists and the line to be careful about
Naming a living artist is the fastest style shortcut available and the one most likely to cause you trouble. It often works, because their work was scraped along with everything else, and it puts you in a position you may not want: a picture that is recognizably in someone's signature style, used commercially, without their involvement. Several image tools now block or discourage the practice, and some jurisdictions are actively legislating around style and likeness.
Describe the qualities instead. Rather than a name, say what you actually want: "thick impasto brushwork, limited palette of ochre and slate blue, visible canvas texture." You get the look, it is more controllable, and nobody's name is attached to your marketing material. Artists who have been dead for a long time are a different matter legally, though platform rules still apply. The full picture on ownership, commercial use and likeness is in using AI images legally.
What to check before your next session
Write your next prompt with the five slots visibly separated by commas, then read it back and ask which slot is empty. That empty slot is where the model is currently guessing on your behalf. Fill it, fix the seed, and change one thing at a time. Keep a short file of prompts that worked, with their seeds and settings, because the wording that produced a result you liked is far harder to reconstruct from memory than you expect.
Common questions
How long should an image prompt be?
Somewhere between 15 and 60 words suits most tools. Shorter than that and the model fills the gaps with its own defaults, longer and later details get diluted, though newer tools that accept full sentences handle long descriptions better than older keyword based ones.
Do words like masterpiece and 8k actually help?
Rarely. They were useful on older models trained on image sites where those tags described popular uploads. Modern models are tuned toward attractive output already, so the slot is better spent on a real detail such as the time of day or the lens.
How do I get the same character in several pictures?
Fix the seed, keep the character description word for word identical, and change only the setting or pose. If the tool supports a character reference image or a trained character model, use that instead, because prompt wording alone drifts after a few images.
What is a negative prompt for?
It lists what you do not want, such as text, watermark or extra limbs, and the model steers away from those concepts. Some tools have a dedicated field for it while others expect you to just describe the scene positively, which often works better.