Why the best e-commerce interfaces know when not to show one.
A picture is worth a thousand words. In e-commerce, this is treated as settled law. Product grids are images. Category tiles are images. Recommendation carousels are images. When a design question comes up, the reflex answer is to add a visual, and that reflex is usually right.
But it isn’t always right, and the places where it fails are more interesting than the places where it works.
Consider search refinements — the filters and chips that help a shopper narrow a broad query into something specific. A shopper searching for tops sees options like crop, peplum, tunic, wrap. These words do real work for people who already know the vocabulary. For everyone else they’re noise. You cannot picture a peplum top from the word alone unless you already know what one looks like, and if you already know, you didn’t need the filter explained. A small image next to each label converts an opaque list into an obvious one. This is the case for visual refinements, and it’s a strong one.
Now consider a shopper searching for a television, offered 55 inch, 65 inch, 77 inch. Show images and something odd happens. The pictures are nearly identical — three black rectangles. They convey nothing the numbers didn’t already convey perfectly. Worse, they take up several times the space, so fewer options fit on screen. Worse still, a photo of a specific television makes an implicit promise about what the shopper will find on the other side of the click, and that promise may not hold.
Or take a shopper buying paper towels, choosing between 12–24 count, 24–48 count, 48–72 count. The same collapse. A number is already the most precise possible representation of a quantity. No image improves on it.
The asymmetry nobody prices in
Here’s the thing that makes this more than a style debate. The benefit and the cost of a visual land on different populations.
The benefit of an image is received by a subset — the shoppers who couldn’t decode the label. Not everyone. Probably not most people, for most attributes.
The cost is paid by everyone. A visual element is a larger fixation target. It slows the scan. It occupies real estate that could have held more options. It pulls attention toward itself whether or not attention was needed there. Every shopper who renders that interface pays this, including the ones who knew exactly what the word meant and just wanted to tap it and move on.
So the honest question isn’t “does this image help someone?” It’s “how many people does it help, weighted against a cost that everyone absorbs?” For a jargon-heavy apparel attribute, plenty of people benefit and the trade is clearly worth it. For a pack count, essentially nobody benefits and the trade is strictly negative. Same interface pattern, opposite verdicts, and the deciding factor is not aesthetic preference — it’s the distribution of who already understands the word.
Text, viewed this way, is not the boring fallback. It’s the dense, fast, low-tax option. It’s what lets a shopper see eight choices instead of three and finish in one glance. A text label is the most compressed representation available of anything that can be fully named — and quantities, dimensions, and counts can always be fully named.
Two questions, not one
The pattern that emerges is that visual presentation earns its place when two conditions hold together.
Is the attribute visually distinguishable at the size it will actually render? Not at full resolution — at chip size. Patterns, silhouettes, and styles separate cleanly at a glance. Screen sizes and capacities do not; they render as the same shape at any thumbnail scale. This is the condition that eliminates the television case.
Is the word insufficient on its own? Some attributes are perfectly specified by their labels. 65 inch is complete. 48 count is complete. Nothing visual is left to add. This is the condition that eliminates the paper towel case.
When an attribute clears both — visually distinct and linguistically opaque — an image is doing genuine work. When it clears neither, an image is pure tax. And when it clears one but not the other, the answer is usually still text, because the tax is certain while the benefit is not.
There’s a fourth quadrant worth naming: attributes that are neither distinguishable nor self-explanatory. Display panel technologies, fabric weights, sun protection ratings. Here an image actively misleads, because it implies a visible difference that doesn’t exist. What these need is explanation — a tooltip, a definition — not a picture.
Restraint is the feature
The instinct in a lot of organizations is to treat “we added images” as the achievement. But an interface that shows images everywhere it can is not a well-designed interface — it’s an undecided one. It has declined to make the judgment call and pushed the cost onto the shopper.
The systems worth building make that call deliberately, and they make it per attribute rather than globally, because the answer genuinely differs between pattern and screen size even within the same store. They treat text as the default and visuals as something an attribute has to earn. They cap how much of any single page can go visual, because even good visual elements compete with each other for the same finite attention. And they degrade to text on any uncertainty, because a plain label is never actively wrong, while a poorly chosen image can be.
There’s a version of this that sounds like an argument against images. It isn’t. Images are enormously powerful for the attributes that need them, and stripping them out wholesale would make a lot of categories meaningfully harder to shop. The argument is narrower: their power comes from being used where they carry information, and gets diluted everywhere else.
The best interfaces aren’t the ones that show the most. They’re the ones that clearly know the difference.
Leave a comment