Taste Engineering · 9 / 12

Generating Images: The Language of Composition, Light and Shadow, Color

Generating images follows the same rule as generating UI: vocabulary sets the ceiling. The trio — composition, light and shadow, color — three words each. Train your eye on three real-image sets, then assemble an image prompt you can use right away.

CompositionLight and shadowColorSide-by-sideFour-pick acceptance
Image quality is capped by your vocabulary

Say “a nice coffee poster,” and AI gives the training-data average. Say “rule-of-thirds composition, golden-hour rim light, low-saturation Morandi tones,” and AI suddenly has a blueprint. Photographers spent a century on this language — it's ready to use. The aesthetic trio: composition places the subject, light and shadow build volume, color sets mood. Three words each is enough. The drills below put two images side by side — Tufte's old move: comparison creates context; good and bad only show when they sit together.

Try it · Round 1: Composition
Composition: where should the subject stand Pick one
Tap the one you think is better.
ALighthouse coast photo; lighthouse dead-centered, horizon cutting mid-frame
BSame lighthouse coast; rule-of-thirds composition, lighthouse on the right intersection, sea whitespace on the left
Same lighthouse coast, two compositions. Image: GPT Image 2 demo
Rule of thirdsDivide the frame into thirds both ways; put the subject on an intersection — more durable than dead center.
WhitespaceEmpty space yields to the subject — the frame doesn't need to be full.
Leading linesUse lines like a coastline or road to pull the eye to the subject.
Try it · Round 2: Light and shadow
Light and shadow: where volume comes from Pick one
Again — tap the better one.
ACeramic coffee-cup still life; rim light, highlight on the rim, elongated shaped shadow
BSame ceramic cup; frontal flat light, no shadow, little volume
Same coffee cup, two lighting setups. Image: GPT Image 2 demo
Side lightLight from the side — volume rides on the shadow.
Golden hourThe hour around sunrise and sunset — soft light with clear direction.
ContrastPull lights and darks apart so the subject pops from the background.
Try it · Round 3: Color
Color: who's fighting, who's speaking Pick one
Last round — tap the better one.
ADesk scene; high-saturation red/yellow/blue/green clash, no unified palette
BSame desk scene; low saturation, unified palette
Same desk scene, two palettes. Image: GPT Image 2 demo
Unified paletteOne color family across the frame — fewer stray hues, longer lasting.
Low saturationTurn the color volume down — material quality rises at once.
Accent colorLeave one high-saturation spot as the focal accent — more and they fight.
Why pick-one works: Tufte's two principles

Information-design elder Edward Tufte left seven principles in The Visual Display of Quantitative Information; About Face 4, Chapter 17, carried them into interface design. Two of them explain exactly how this page plays.

First: strengthen visual contrast — comparison creates context. Alone, a lighthouse shot won't tell you if the composition works; side by side, rule of thirds vs dead center is instant. You answered the first three rounds so fast because comparison magnified the gap. That's also why good composition and good light “read at a glance”: the eye is a contrast machine — it just needs a side-by-side chance.

Second: show side by side in adjacent space, better than stacking over time. Flip one by one and you lean on short-term memory — by the fourth, the first is already fuzzy. Side by side, comparison rides the eye. So when AI outputs candidates, don't delete as you go — gather them and compare together.

The right way to ask AI for images: request four at once, lay them side by side, then pick — don't forget as you flip one by one.
Try it · Carousel vs side-by-side

Doubt that side-by-side is faster? Run a trial with the six images from the first three rounds. Task: find the “golden-hour rim light” shot. Hunt in carousel first, then switch to side-by-side — feel the gap.

Find the rim-light shot 0 / 2 modes
Find it once in each mode: in carousel tap “Pick this one”; side by side, tap the image.
1 / 6 Current carousel image
The six images from the first three rounds. Image: GPT Image 2 demo
Try it · Acceptance: pick one of four to deliver

Three rounds done — real work: you asked AI for four coffee-brand poster candidates; the client wants one tomorrow. Use the trio you just learned to pick a deliverable. Note the posture: lay all four side by side, then pick — Tufte's adjacent-space comparison.

Coffee-brand poster acceptance Pick one of four
Tap the one you'd send the client; after that, verdicts reveal card by card.
Candidate 1Coffee poster candidate 1: kettle hugging the edge, large empty side with no content — composition unbalanced
Candidate 2Coffee poster candidate 2: rule-of-thirds composition, shaped rim light, low-saturation unified palette
Candidate 3Coffee poster candidate 3: frontal flat light, kettle lacks volume
Candidate 4Coffee poster candidate 4: stacked high-saturation colors — color overload
Four coffee-brand poster candidates. Image: GPT Image 2 demo
Try it · Assemble an image prompt

Pick one block from each of the trio — what you assemble is a prompt you can send straight to an image model. Swap the subject line for your own scene and it's ready.

Image-prompt block builder Live assemble
Pick one per group; the prompt below updates. When it looks right, copy and go generate.
Composition
Light and shadow
Color
Key Takeaways

The aesthetic trio directs image gen: composition for placement, light and shadow for volume, color for mood — three words each.

Picking and generating share one vocabulary: if you can say “rule of thirds,” “rim light,” “low saturation,” a four-pick has grounds — delivery isn't luck.

View candidates side by side: comparison creates context; adjacent space beats flipping one by one. Tufte's two principles are the method under this page.

Next image gen, write all three lines — composition, light and shadow, color — before you send; when the output is off, check which line was unclear.

Source: Original to Xiaoshan Academy's Taste Engineering series; demo images on this page generated with GPT Image 2; some design principles adapted from About Face 4: The Essentials of Interaction Design, Chapter 17.