How to Direct ChatGPT to Create Images at an Expert Level
The Art Director Method
Creating strong AI images is not primarily about discovering a magical prompt.
It is much closer to
directing a photographer, illustrator, graphic designer, set designer, lighting director, and Photoshop artist at the same time.
ChatGPT does most of the actual rendering. Your job is to provide the creative judgment.
The basic loop is:
Visualize → Describe → Generate → Critique → Correct → Refine
The people who become good at this are usually not the people who write the longest prompts. They are the people who get good at
seeing what is wrong with an image and explaining exactly how it should change.
1. Start With the Idea, Not the Prompt
Before asking ChatGPT to generate anything, decide what you want the viewer to experience.
Ask yourself:
- What is the image trying to communicate?
- What should the viewer notice first?
- What should they notice second?
- Is the image funny, disturbing, dramatic, satirical, cinematic, elegant, absurd, documentary-like, etc.?
- Is there a visual punchline?
- Is there a specific person, object, sign, document, building, or expression that tells the story?
You do not need every detail figured out.
You do need the
central visual idea.
For example, instead of starting with:
Make an image about someone receiving a subpoena.
Think:
I want the viewer to instantly understand that this person has just discovered they have been subpoenaed. The laptop or document needs to make that fact visually unmistakable, and the person's reaction should reinforce it.
That is direction.
2. Describe the Scene as Though You Were Directing a Film Crew
A powerful prompt usually answers several different questions.
SUBJECT
Who or what is in the image?
ACTION
What is happening?
ENVIRONMENT
Where is it happening?
COMPOSITION
Where are the important objects located?
CAMERA
What viewpoint are we seeing?
LIGHTING
What kind of illumination creates the mood?
EXPRESSION
What emotion should a person communicate?
STYLE
Photorealistic? Editorial? Vintage comic? Movie poster? Documentary? Advertising photography?
STORY
What should someone understand within two seconds of seeing it?
Instead of:
A worried businessman at a computer.
Try:
Create a photorealistic cinematic scene of a businessman seated at his desk staring at his laptop after receiving a deposition subpoena. His expression should show alarm and disbelief. The laptop screen should clearly communicate that a subpoena has been received. Keep the desk realistic and uncluttered enough that the laptop remains visually important. Use dramatic but believable office lighting.
The second version gives the model a scene to stage.
3. Think in Layers
One of the easiest ways to control a complicated image is to mentally divide it into layers.
Background
Building, room, landscape, wall, sky, curtains, city, etc.
Middle ground
Furniture, vehicles, architectural elements, secondary people.
Foreground
Main person, sign, document, computer, prop, visual joke.
Graphic layer
Headlines, labels, captions, logos, speech bubbles, warning panels.
For a poster, for example, you might specify:
TOP: Large headline
CENTER: Museum entrance with red carpet
LEFT: Information placard
RIGHT: Valet beside black luxury car holding key fob
BOTTOM: Event information and warning panels
Spatial instructions like
top, center, left, right, foreground, background, behind, beside, above, below are extremely useful.
You are blocking the scene.
4. Establish Visual Hierarchy
Everything cannot be equally important.
Tell ChatGPT what should dominate.
For example:
The headline should be the first thing the eye sees.
The man's facial expression is more important than the background.
The subpoena notice on the laptop must remain readable.
The speech bubble is the punchline, so make it large enough to read immediately on a phone.
This is one of the biggest differences between merely generating an image and directing one.
You are controlling the viewer's eye.
5. Give Concrete Instructions Instead of Vague Adjectives
AI responds much better to observable details.
Weak:
Make him look really angry.
Better:
Give him a visibly angry expression: furrowed brow, narrowed eyes, tense jaw, lips pressed together.
Weak:
Better:
Use strong directional lighting, deep shadows, a bright focal area around the subject, and a darker background.
Weak:
Make the poster look professional.
Better:
Use a symmetrical advertising-poster composition, restrained typography, strong visual hierarchy, generous spacing, and a polished theatrical lighting treatment.
Whenever possible, convert an adjective into something the camera could actually see.
6. Use Reference Images When Accuracy Matters
If you want something to resemble a real object, person, package, room, drawing, or previous image, upload the reference.
Then explain
what the reference controls.
For example:
Use the uploaded cigarette package as the reference for the package design and proportions.
Or:
Keep the composition from the existing image, but replace the person on the right.
Or:
Preserve the artwork exactly except for the dialogue bubble.
This is important because the model otherwise has to invent details.
A reference image converts some of those guesses into instructions.
7. Do Not Try to Perfect Everything in One Generation
This is probably the most important lesson.
The first image is a
draft.
Treat it the same way a designer treats a first comp.
Look at it and ask:
- What worked?
- What is wrong?
- What is missing?
- What is distracting?
- What needs to be larger?
- What needs to move?
- What needs to disappear?
- Is the facial expression right?
- Is the story instantly understandable?
- Is the text readable?
- Is the visual joke landing?
Then issue the next direction.
8. Correct One Problem at a Time When Possible
Suppose an image is 90% right but the character's expression is wrong.
Do NOT describe the entire picture again from scratch.
Say:
Keep everything else exactly the same. Change only his facial expression so that he looks furious rather than surprised.
Or:
Preserve the existing composition, lighting, clothing, background, camera angle, and text. Only replace the object in his hand with a key fob.
This reduces creative drift.
A useful formula is:
KEEP: what is already correct
CHANGE: what is wrong
DO NOT CHANGE: anything important that might accidentally be altered
Example:
Keep the existing image and composition. Keep the woman's pose, background, lighting, clothing, and camera angle unchanged. Change only the sign she is holding. Replace the existing wording with: "[TEXT]."
That is essentially an AI version of a Photoshop change request.
9. Learn to Diagnose Why an Image Feels Wrong
Sometimes an image looks wrong even when all the requested objects are technically present.
That usually means the problem is one of these:
Composition
Objects are badly positioned.
Scale
Something is too large or too small.
Hierarchy
The wrong element attracts attention.
Expression
The face communicates the wrong emotion.
Lighting
The subject does not visually belong in the scene.
Perspective
Objects appear to exist on different planes.
Density
There is too much competing information.
Readability
Important wording is too small or visually buried.
Storytelling
The elements exist, but the viewer cannot immediately understand what is happening.
Do not simply say:
Diagnose it.
For example:
The problem is that the laptop is visually insignificant. Enlarge it slightly and angle the screen toward the viewer so the subpoena becomes one of the main storytelling elements.
That gives ChatGPT something actionable.
10. Treat Text as a Separate Design Problem
AI-generated text has historically been one of the trickier parts of image creation, although it continues improving dramatically.
For important wording:
- Specify the exact text.
- Put it in quotation marks.
- State where it belongs.
- Describe its size and hierarchy.
- Check spelling carefully after generation.
Example:
At the top, in large bold lettering, write exactly:
"GRAND OPENING"
Then inspect it.
If one word is wrong:
Keep the image exactly as it is. Correct only the headline. It must read exactly: "GRAND OPENING."
For posters containing lots of information, build them in sections rather than asking the model to magically organize twenty unrelated pieces of copy.
11. Direct Facial Expressions Carefully
Faces can change the entire meaning of an image.
Instead of simply naming an emotion, describe its visible characteristics.
For anger:
Furrowed brow, narrowed eyes, tense jaw, stern mouth.
For panic:
Eyes widened, eyebrows raised, mouth slightly open, body visibly tense.
For smugness:
Slight asymmetric smile, relaxed posture, chin marginally elevated.
For suspicion:
Eyes slightly narrowed and directed sideways, restrained expression, subtle tension in the brow.
Think like a director giving an actor notes.
12. Control Camera Position
Camera instructions can radically change an image.
Experiment with:
- eye level
- low angle
- high angle
- close-up
- medium shot
- wide establishing shot
- over-the-shoulder
- first-person viewpoint
- centered symmetrical composition
- three-quarter angle
For example:
Show this from the driver's point of view with the steering wheel visible in the foreground.
That instruction does far more than saying:
The viewpoint itself tells part of the story.
13. Use Lighting to Control Emotion
Lighting is storytelling.
Examples:
Serious / cinematic
Directional lighting, shadows, controlled highlights.
Luxury
Glossy highlights, warm architectural lighting, polished reflective surfaces.
Ominous
Dark environment with isolated pools of light.
Documentary
Natural, believable ambient illumination.
Comic
Bright, clean, evenly illuminated forms.
Nightlife
Neon reflections, wet pavement, contrasting light sources.
You don't need to know professional cinematography terminology.
Simply describe
where the light seems to come from and what feeling it should create.
14. Protect the Parts That Already Work
As images improve, this becomes increasingly important.
Suppose version six finally has:
- the perfect background
- the correct face
- the right camera angle
- good lighting
But one label is wrong.
Tell ChatGPT:
Do not redesign the image. Preserve everything except the specified correction.
The more specific the preservation instructions, the less likely you are to lose something good while fixing something else.
15. Iterate Through Smaller and Smaller Corrections
Early iterations might involve:
Completely change the composition.
Later iterations might involve:
Move the speech bubble slightly upward.
Then:
Increase the speech bubble text approximately 20%.
Then:
This is normal.
Image creation often progresses from
architecture → composition → objects → expressions → typography → polish.
Trying to solve all six simultaneously can create chaos.
16. Design for Where the Image Will Be Seen
An image that looks great full-screen may fail completely when posted on a message board or viewed on a phone.
Before finishing, ask:
- Will people see this primarily on phones?
- Is the headline readable at thumbnail size?
- Are speech bubbles large enough?
- Is the punchline obvious without zooming?
- Are important faces recognizable?
- Is there unnecessary empty space?
- Is the file unnecessarily huge?
Sometimes the final stage is not artistic improvement.
It is
communication improvement.
A slightly larger speech bubble may matter more than another hour spent perfecting background lighting.
17. Use ChatGPT as a Critic Too
You do not always need to identify the problem yourself.
Ask:
Analyze this image as an art director. What are the three biggest problems with composition, visual hierarchy, readability, or storytelling?
Or:
The image doesn't quite work for me. Diagnose why before changing anything.
Or:
Which element does the viewer notice first, second, and third?
Then decide whether you agree.
The model can help generate the image and help critique its own work.
18. Separate "Creative Exploration" From "Precision Editing"
These are two different modes.
Exploration Mode
Use this when you do not yet know exactly what you want.
Say:
Give me a dramatically different interpretation.
Try a darker cinematic approach.
Reimagine this as a 1950s comic-book advertisement.
Make the visual metaphor more obvious.
You WANT variation.
Precision Mode
Use this once the image is close.
Say:
Preserve everything else.
Do not alter the composition.
Keep the person's identity, pose, clothing, lighting, and background unchanged.
You want variation to STOP.
Knowing which mode you are in is a major skill.
19. Don't Be Afraid to Give Tiny Corrections
You can direct details that might seem absurdly specific:
Make the key fob more visible.
Open the package instead of showing it closed.
Put cigarette butts in the ashtray.
Make the document clearly identifiable as a subpoena.
Increase the contrast behind the speech bubble.
Remove the person entirely.
Make the building symmetrical.
Keep everything else unchanged.
Those tiny corrections often make the difference between "AI-generated picture" and an image that actually communicates the intended idea.
20. Develop the Most Important Skill: Visual Criticism
Prompt writing gets most of the attention.
But the real superpower is:
Knowing what to ask for next.
After every generation, mentally complete this sentence:
"This would be much better if..."
Then turn the answer into the next instruction.
For example:
This would be much better if the viewer could actually read the joke.
So:
Increase the speech bubble and its lettering while preserving the illustration.
Or:
This would be much better if he looked angry rather than vaguely concerned.
So:
Change only his expression. Make the anger unmistakable.
Or:
This would be much better if the viewer understood immediately that this is a deposition subpoena.
So:
Make the subpoena wording visually prominent enough to communicate the situation at a glance.
That cycle is where the real expertise develops.
21. A Useful Master Prompt Structure
When starting a complex image, this framework works well:
PURPOSE
What the image should communicate.
SCENE
What is happening.
SUBJECTS
People and important objects.
COMPOSITION
Where everything belongs.
CAMERA
Viewing angle and framing.
LIGHTING
Mood and illumination.
STYLE
Photorealistic, comic, poster, painting, etc.
TEXT
Exact wording and placement.
PRIORITY
What absolutely must work.
For example:
Create a cinematic promotional poster.
PURPOSE: The image should feel like the grand opening of a strange, prestigious museum.
COMPOSITION: Use a vertical poster layout. Place a dramatic museum entrance in the center with a red carpet leading toward it. Put a valet scene with a black luxury car on the right. Place an informational placard on the left.
LIGHTING: Nighttime setting with theatrical architectural lighting and reflections on the pavement.
STYLE: Polished, photorealistic luxury advertising campaign.
TEXT: Place the main headline prominently at the top.
PRIORITY: The museum entrance and headline should dominate the image. Secondary information must remain readable without competing with them.
You can add or remove sections depending on the project.
22. The Expert Workflow
For complicated work, use this sequence:
1. Concept
Decide what the picture means.
2. Rough generation
Get the basic idea onto the screen.
3. Composition correction
Move, enlarge, remove, or reorganize major elements.
4. Story correction
Make sure someone understands what is happening.
5. Character correction
Fix pose, expression, clothing, interaction, etc.
6. Prop correction
Fix objects that carry the story.
7. Typography correction
Fix wording, placement, hierarchy, and readability.
8. Visual polish
Lighting, realism, texture, reflections, atmosphere.
9. Viewing-size test
Make sure it still works small.
10. Final surgical edits
Change only what still needs changing.
You may go through two generations.
You may go through twenty.
There is no prize for accepting version one.
23. The Golden Rule
Never assume that because the AI generated something, you have to accept its creative decisions.
You are the director.
If the car is wrong, change the car.
If the joke cannot be read, enlarge it.
If a person's expression is wrong, redirect the performance.
If the composition is cluttered, remove things.
If the visual metaphor isn't obvious, strengthen it.
If something excellent appears unexpectedly, keep it.
The model supplies enormous visual capability.
You supply taste, intention, judgment, and persistence.
That combination is where the interesting work happens.
The One-Sentence Version
Picture what you want, describe it concretely, examine what the AI actually produced rather than what you hoped it produced, and keep giving increasingly precise art-direction notes until the image communicates exactly what you intended.