Describe depth, or get cardboard
Ask for a remote control and you get a flat rectangle with dots on it. Not because the model is weak — because you described a diagram and it drew one. Depth is something you have to say out loud.
Why it comes back flat
The model has no idea what a remote control is. It has no geometry, no notion of a solid object sitting in a room. What it has is an enormous number of pictures that were captioned “remote control”, and it produces something statistically typical of that pile.
So ask what that pile actually looks like. Icons. App screenshots. Vector illustrations. Silhouettes on white. Instruction-manual diagrams. Product listings shot dead-on with the shadows lit out of them. The centre of that distribution is genuinely flat, and the model landed exactly where you pointed it.
This is the part worth internalising: when you name the shadow under the device, the chamfer on its edge and the doming of the buttons, you are not adding detail. You are moving to a different subpopulation of the training data — the pictures that were lit and photographed rather than drawn. Depth language is a selector, not a decoration.
The same law, on the other side of the house
We paid for this lesson once already, compositing rather than generating. Putting a contractor's logo on a claymation truck door with flat ink reads as a sticker — every time, at any resolution, however much you blur or blend the edge. The fix was never softer edges. It was giving the mark its own highlight, its own shadow and its own contact shadow, so it sits in the surface instead of on top of it.
Different pipeline, identical failure. Which makes the rule general:
Depth has to be stated. Nothing in the pipeline — generative or compositing — will infer it for you.
The vocabulary that actually moves it
Vague requests for “realistic” or “3D” or “high quality” do very little, because those words appear on flat pictures too. Words describing how light behaves do the work, because they only appear on pictures where light behaved.
- Contact — where the object touches the surface. Contact shadow, occlusion tightening in the seams, the direction the cast shadow falls. An object with no contact shadow floats.
- Edges — chamfer, fillet, a raised lip, a recessed well. “Bezel” on its own is weaker than saying the bezel is raised and catches a thin highlight along its top edge.
- Buttons and small forms — domed, concave, the gap around each key, how far they sit proud of the body.
- Surface — soft specular rolloff along the edge, satin versus matte, a broad highlight rather than a hotspot. This is what separates moulded plastic from printed paper.
And the counter-cues — words that quietly pull you back toward the flat pile, worth removing rather than balancing: icon, flat, vector, logo, clean product shot on white.
How to tell which failure you have
This matters because the two failures look similar and have opposite fixes. If the object is flat — no contact shadow, uniform lighting, edges with no highlight — that is a conditioning failure. More resolution will not touch it. A larger model will not touch it. You will get a crisper cardboard remote.
If the object has believable light and shadow but the details are mushy — button legends unreadable, textures smeared — that is a capacity failure, and resolution or a better model genuinely helps.
Getting this diagnosis right saves the most expensive mistake in AI production: re-rolling the same flat prompt at higher settings and paying more each time for the same wrong picture.
Applying it
Before you render any hero object — a remote, a phone, a tool, a sign, a product in someone's hand — say four things about it: where its shadow lands, what its edges do to light, which of its forms are raised or recessed, and how glossy the surface is.
Four clauses. It is the difference between a prop that sits in the scene and a cutout someone pasted into it, and it costs one render rather than five.