How to art direct AI images
Make the calls you would make on set, then write each one as words a model can act on: subject, action, ground, light, palette, lens, framing, texture and mood. Write a short take for the whole set before any prompt. Give each shot only the reference photos it needs, reuse the sentences that worked word for word, and change one thing per render so you know what fixed it.
Your decisions, in words a model can use
An image model fills every decision you leave open with its own guess. Each row below is a call you already make on set, with a line that carried it in a prompt behind a photo our agency kept in 2026. The lines come from two prompts: a store photo of trousers on a model, and a fabric close-up of a henley.
| Decision | What to write | A line that worked |
|---|---|---|
| Subject | Which reference picture is the product or person, and what must match it | "the exact same man as INPUT IMAGE 1, same face, same hair, same skin tone, same build" |
| Action | What the subject is doing, in physical terms | "a subtle weight shift resting smoothly on the right leg, arms relaxed at the sides" |
| Ground | The surface or backdrop, with a hex code and how it falls off | "a softly out-of-focus Sea Foam White #F2F5F8 seamless cyclorama falling completely away" |
| Light | Source, direction, hardness, and what it does to a surface | "Hard raking natural daylight enters at a 45-degree angle across the chest, carving the herringbone weave" |
| Palette | Color names with hex codes, tied to objects in the frame | "the muted dusty mauve and Faded Sky Blue #B0C4DE threads inside the stripes" |
| Lens | One sentence: body, focal length, aperture, camera height | "a full frame mirrorless body with an 85mm lens at f8, camera at chest height, level, no tilt" |
| Framing | Where the frame starts and ends | "The frame begins above the head and ends below the shoes" |
| Texture | Grain, sharpness, contrast | "clean near-zero grain, balanced contrast preserving honest shadow detail inside the fabric folds" |
| Mood | A feeling attached to something visible | "natural saturation emphasizing the washed, lived-in heavyweight cotton texture" |
| Fidelity | What may not change about the product | "nothing added and nothing removed" |
Write the take before the prompts
A take is a short paragraph that sets the look for the whole set before anyone writes a shot: the idea, the light family, the ground, the palette, the lenses, the grain, the cast, and what must never appear. Every prompt borrows from it, so ten shots read as one shoot. Write it the way you would brief a photographer the night before.
Idea: [one sentence on what the set says about the product] Light family: [for example, one large soft key from camera left, low fill, soft contact shadows] Ground: [surface or backdrop with a hex code, and how it falls off] Palette: [3 to 5 color names with hex codes, and which objects carry them] Lenses: [for example, 85mm for people, 100mm macro for details, 50mm top-down for flat lays] Texture: [grain level, sharpness, color grade reference] Cast: [who appears, from which approved character sheet] Never: [4 to 8 things that must not appear, for the reviewer]
The take sits on top of the brand’s standing look, its visual DNA, and changes with each campaign. Keep the never list for the reviewer and the brief, and write the prompts themselves as positive instructions. See how to write a "never do this" list.
Vague direction, and what to write instead
Words like "premium" describe a result. A model needs the choices that produce it. Translate each adjective into the light, ground, lens and palette you would pick on set.
| If you would say | Write the decision instead |
|---|---|
| "Premium" | Matte black ground (#111111), two strip lights behind the bottle, one on each side, drawing a thin highlight down each edge, no fill. |
| "Moody" | Low key: one soft source from behind at camera left, most of the frame in shadow, highlights only on the rim of the product. |
| "Natural light" | Soft window light from camera right on an overcast day, falling across the tabletop and fading toward the back wall. |
| "Fresh" | High key: white seamless (#FFFFFF), large soft sources in front and above, a pale palette, cold droplets on the glass. |
| "Hero shot" | Camera just below the product’s mid-height, looking slightly up, 85mm, product filling about 60 percent of the frame height, plain ground. |
| "Lifestyle" | On a kitchen counter at 8 a.m., a half-drunk coffee at the frame edge and out of focus, one hand entering from the right. |
| "On-brand colors" | Backdrop in [color name] (#hex), props only in [color name] (#hex), no other saturated color in frame. |
| "Energy" | The product mid-fall, crumbs frozen in the air, hard flash from above to freeze them. |
| "Make it look real" | A soft contact shadow where the base meets the table, visible film grain, crumbs or dust where they would fall. |
Change one thing per render
When a render misses, change one variable and render again. If you move the light and swap the lens together and the next render works, you will not know which fix to carry to the next shot. OpenAI’s prompting guide for its image models gives the same advice: start from a clean base prompt and make small, single-change follow-ups.
Consistency across a set is the same rule turned around. Hold every sentence that worked, and change only what the shot needs. On a trouser campaign our agency made in 2026 on GPT Image 2, the backdrop and lighting sentences were reused word for word across the store photos. The camera sentence changed only for the fabric close-up, to a 100mm macro at f5.6. The framing and pose sentences changed from photo to photo. This is one of those prompts, about 350 words long.
This is a photograph of the exact same man as INPUT IMAGE 1, same face, same hair, same skin tone, same build, wearing exactly the same wardrobe head to toe as the character sheet with only the trousers replaced by the product trouser shown in INPUT IMAGE 2. The frame begins above the head and ends below the shoes, the whole figure inside the frame with clean margin, head and shoes fully visible. Full body front view, standing composed and upright with a subtle weight shift resting smoothly on the right leg, arms relaxed at the sides, gaze soft toward camera. The trouser must match INPUT IMAGE 2 exactly, navy smoke houndstooth pattern in a textured woolen weave with a matte finish and substantial drape, flat front with no pleats, angled slash front pockets, standard waistband with belt loops and an extended tab closure, a single dark button at the waist closure, no rivets, minimal tonal visible stitching, sharp center front crease, straight to mildly tapered leg with a plain finished hem, nothing added and nothing removed. If a fabric close-up reference is provided its weave and texture must be matched exactly. Studio setting with the backdrop exactly as follows, Light Gray, a seamless mid-gray studio backdrop with a soft cool tone that grades slightly darker toward the edges and blends into a matching gray floor with no visible seam, lit evenly with diffused soft light so the model casts only a faint soft shadow, no props. Lighting is a large soft key from the upper left at forty five degrees with even wraparound fill, a soft natural contact shadow under the shoes, no hard shadows, no colored light. Camera is a full frame mirrorless body with an 85mm lens at f8, camera at chest height, level, no tilt. Colors stay exactly true to the flat-lay, never shifted, never saturated. Crisp commercial e-commerce finish. Portrait orientation. The garment is never altered from the reference. No text, no graphics, no logos other than what exists on the garment, no watermarks. Natural believable hands and body proportions. Photorealistic only, never illustrated, never stylized.
INPUT IMAGE 1 was the approved character sheet and INPUT IMAGE 2 the flat lay of the trousers. The sentences that start "Studio setting" and "Lighting is" stayed fixed across the set. The frame and pose sentences near the top are the ones that changed per shot. The long trouser sentence lists every construction detail, and it doubles as the reviewer’s checklist.
Use references on purpose
Every picture you attach is an instruction. Give each shot only the pictures it needs, say in the prompt what each one is for, and leave out anything the model could copy by mistake.
- Name each input by number and job, as the prompt above does. OpenAI’s guide recommends labeling each input by index and description, such as "Image 1: product photo".
- Send the picture that shows the part in frame. A back view needs the back photo and a close-up needs the texture photo, or the model invents what it has not seen. See why AI changes your product.
- Send fewer, better references. The character sheet our agency kept on a surf apparel job came on the 20th attempt, from 5 reference images and a four-line prompt. The longest prompt and the attempts with up to 8 references were not kept.
- Know the limit. As of September 2026, Google lists Nano Banana 2 at up to 10 object images and 4 character images per request, and Nano Banana Pro at up to 6 object images, 5 character images and 3 style references.
- Use a mood board for its feeling only: light, palette, energy. Objects on a board can end up in the photo. See how to use a mood board with AI.
Write for the model you render on
The same direction needs a different shape on different models. In our 2026 campaigns, kept prompts ran about 140 to 200 words on Nano Banana models and about 300 on GPT Image. The Nano Banana prompts opened with the subject and its reference and described light as something falling on a surface. The GPT Image prompts opened "This is a photograph of", named each input image, and closed with rules. How to write a product prompt takes the structure apart line by line, and the vocabulary is in lighting words and camera and lens words.
Questions people also ask
- How many renders should I expect per shot?
- In four campaigns our agency made in 2026, 361 renders gave 45 keepers, about 8 per photo, with a range of about 3 to 14 by job. A person who must stay the same person, and motion such as falling crumbs, took the most tries. See how many renders per usable photo.
- Do negative prompts work?
- It depends on the model. Black Forest Labs says FLUX.2 does not support negative prompts and suggests writing "sharp focus" instead of "no blur". OpenAI’s guide suggests stating exclusions such as "no watermark" and "no extra text" outright. Write what you want first, and keep a short rules list only on models that follow one.
- Should the art director write the prompts?
- Someone who can see the photo in their head should write or approve every prompt, because each sentence is a creative decision. One split that works is for a producer to draft prompts from the take and the shot list, and for the art director to edit the light, lens and framing lines before anything renders.
Where Overs fits
In Overs, art direction is its own step, after research, brand rules and visual style and before shot planning and prompt writing. Each photo gets only the reference pictures its plan names, and you can remake one photo with a note, up to four versions at once, to change one thing at a time.
Free for 40 photos a month. The AI that makes the photos is billed separately, on your own key, with no markup from Overs.