Entry № 82Capsule Design

Steam Capsule Genre Signal: Data From 373 Store Pages

88 of 373 analyzed Steam pages have a capsule that never says what the game is. They score 10.5 points lower, and the gap survives every control we ran.

12 min readBy Steam Page Analyzer Team

Between February 10 and July 26, 2026 our analyzer scored 373 unique public Steam store pages and logged 390 separate complaints about their capsule art. Those complaints are not evenly interesting. On 88 of the 373 pages (23.6%), the verdict was that the capsule never tells you what kind of game it is. Those pages carry a median overall score of 59.5 against 70.0 for every other page in the set, a 10.5-point gap at p less than 0.0001, and it is the only defect pattern in this dataset that both survives controlling for everything else wrong with the page and points the way you would expect.

Two caveats before any number. This is a convenience sample: pages somebody chose to paste into a free tool, not a random draw from Steam. And the scores are our own AI rubric, not a Valve metric and not an objective measure of quality. Everything below is what that rubric found, including the parts that argue against us.

Capsule advice is everywhere and it is almost entirely taste. “Make it pop.” “Show a face.” “Read at thumbnail size.” What has been missing is a count. This is the count.

What we flagged, and what the label actually means

Every analysis emits a list of issues tagged by category and severity. We took the 390 capsule-category issues and grouped them into recurring defect patterns with rule-based text matching over each issue’s title, description and evidence text. Two patterns swallow almost all of them:

Capsule verdictIssuesPages carrying it
Illegible or cluttered at thumbnail size266265
Does not communicate what the game is8888
Composition problems1414
Logo or mascot with no subject77
Art contradicts the actual game55

Be precise about that second label, because the cluster is broader than its name. It merges two things the rubric says in the same breath: no genre signal (nothing in the image tells you this is a builder, a shooter, a farming sim) and no visual hook (nothing in the image is interesting enough to stop a scroll). We could not cleanly separate them, so the honest label is the wider one: the capsule fails to communicate the game. Where the post says “genre signal” below, read it as that.

53 pages (14.2%) drew no capsule complaint at all. 46 of the 88 flagged pages also picked up the thumbnail-legibility flag, so the two are not mutually exclusive.

How far behind do these pages score?

MetricFlagged (n=88)Not flagged (n=285)Differencep
Overall, median59.5 (CI 56 to 62)70.0 (CI 69 to 71)-10.5<0.0001
Overall, mean58.368.8-10.6<0.0001
Capsule subscore, median58 (CI 45 to 58)72 (CI 72 to 72)-14<0.0001
Capsule subscore, mean54.871.6-16.9<0.0001

Confidence intervals are bootstrapped over 20,000 resamples; p-values come from two-sided permutation tests on the median difference, also 20,000 resamples, fixed seed. The capsule subscore is worth reading as a mean as well as a median, because the rubric piles 152 of 373 pages onto exactly 72. When a median sits on a mass point like that it stops moving, so the mean carries the real signal there.

Severity says the same thing louder. 34 of the 88 issues in this cluster are rated critical (38.6%), against a corpus-wide critical rate of 11.8% and a capsule-category rate of 12.8%. All 15 issues in the dataset titled exactly “Capsule lacks visual impact and genre clarity” are rated critical, without exception.

The rubric’s own repair estimate agrees. It attaches a predicted score gain to issues it considers fixable: mean +11.57 for this cluster against +7.70 across all 1,578 estimated issues (p less than 0.0001), and +6.85 for the thumbnail-legibility flag. Read those as predicted movements in our score and nothing else. No sales, wishlist or traffic data is joined to any page in this set, so nothing here can be converted into money.

Share of pages whose capsule was flagged for not communicating the game, by overall score
Score under 55 (n=46)65.2%
Score 55-64 (n=86)37.2%
Score 65-74 (n=161)12.4%
Score 75 or higher (n=80)7.5%
Source: Steam Page Analyzer, 373 unique Steam store pages analyzed Feb 10 - Jul 26 2026

Split at deciles, 37 pages each side: the flag appears on 2 of the 37 best-scoring pages and 25 of the 37 worst (p less than 0.0001).

Does the gap survive controlling for everything else wrong with the page?

This is the check that matters, and it is where most of our candidate findings died.

The problem is obvious once you see it. Flagged pages carry a median of 8 issues each; unflagged pages carry 6. Issue count correlates -0.537 with the overall score all by itself. So a raw comparison is partly just counting defects: pages with more problems score worse, and this is one of the problems.

The fix is to compare a flagged page only against unflagged pages carrying a similar number of other defects, and to test by shuffling the flag within those bins rather than across the whole sample. Four versions, because the binning choice should not decide the answer:

ControlEffect on overall scorep
Bin on total issue count, one bin per count-9.00<0.0001
Bin on total issue count, quartiles-9.01<0.0001
Bin on other-issue count (flag excluded), quartiles-9.70<0.0001
Bin on other-issue count (flag excluded), one bin per count-11.77<0.0001

The raw 10.5-point gap becomes roughly a 9-point gap. It does not become zero.

Here is why that is the interesting sentence. We ran the identical stratified test on 24 recurring defect patterns across all four page sections, with a Bonferroni threshold of 0.05/24 = 0.00208. Twenty-two of the 24 died. “Weak opening hook”, which is the largest description cluster in the dataset at 226 pages, comes out at -0.14 points (p = 0.85). “Wall of text” description: +0.09 (p = 0.92). “Irrelevant or diluting tags” on 233 pages: +0.15 (p = 0.86). Once you hold defect count constant, almost every recurring complaint stops predicting anything.

Two survived. One is this capsule cluster. The other is the thumbnail-legibility flag, and it survived pointing the wrong way, which is the next section.

Does it hold before launch and across genres?

208 of the 373 pages (55.8%) were unreleased Coming Soon pages, and release status confounds nearly everything in this dataset, so both strata get tested separately.

StratumFlaggedNot flaggedMedian differencep
Coming Soonn=50, 58.5n=158, 69.0-10.5<0.0001
Releasedn=38, 60.0n=127, 70.0-10.0<0.0001

The flag rate is nearly identical either side of launch: 50 of 208 unreleased pages, 38 of 165 released ones. This is not a pre-launch placeholder problem that resolves itself on release day.

By genre, with Holm correction across the seven tested:

GenreFlaggedNot flaggedMedian differenceHolm-adjusted p
Actionn=39, 56.0n=126, 70.0-14.00.0003
Indien=65, 56.0n=192, 70.0-14.00.0002
Adventuren=32, 56.5n=114, 70.0-13.50.0003
Strategyn=33, 61.0n=68, 71.5-10.50.0002
Simulationn=24, 60.0n=86, 70.0-10.00.0004
Casualn=25, 60.0n=98, 70.0-10.00.0008
RPGn=13, 63.0n=58, 70.0-7.00.0170

Seven for seven. Games carry multiple genre labels, so these groups overlap heavily and are not independent tests, but there is no genre here where the pattern reverses or disappears.

What does a capsule that fails to communicate look like?

The rubric records an evidence string with each issue, describing what it saw in the image. These are the most concrete artifacts in the dataset. Developer and game names are redacted below, and one dash was normalized to a comma, because the point is instructive rather than punitive. All are small teams.

A strategy game, scoring 54 overall:

“Capsule shows barren landscape with no units, UI, or strategic elements visible”

A space strategy game, 47 overall:

“Header capsule shows only spaceships over a forest/mountain landscape at sunset with no strategy-game visual cues (no UI, no colony icons, no fleet composition)”

A casual puzzle game, 44 overall, capsule subscore 45:

“Header and small capsule both show only the [game] logo, a cute blue octopus peeking, and a pastoral background, no puzzle grid or gameplay elements visible”

An action-adventure game, 43 overall, capsule subscore 22, the lowest capsule score in the set:

“Header capsule: plain cornflower-blue background with black text '[game title]' and two nearly invisible dark sprites at bottom corners; small capsule is an even more compressed version of the same nearly-empty composition.”

A free-to-play action game, 46 overall:

“Both header and small capsules show only the title text with forest background”

A tower defense game, 56 overall:

“Capsule shows static character in landscape with no towers or enemies visible”

The shape repeats: a logo, a mood, a background, a character standing still. Nothing that tells a stranger what they would be doing. The tower defense capsule with no towers and the strategy capsule with no units are the same failure as the puzzle capsule with no puzzle.

A related defect worth naming: on 9 pages the rubric found the header and the small capsule were effectively the same image at two sizes. Those pages have a median capsule subscore of 45 against 72 for everyone else (p = 0.0018). The small capsule is not a resize job; it is a separate composition problem with a fraction of the pixels.

Why the other capsule complaint is worth ignoring

219 pages drew the thumbnail-legibility complaint and nothing else about their capsule: text too small, detail lost at browse scale, hierarchy unclear. It is by far the most common capsule verdict in the dataset.

It also carries no information whatsoever.

Those 219 pages have a median capsule subscore of 72 and a median overall of 70. The 53 pages that drew no capsule complaint at all have a median capsule subscore of 72 and a median overall of 70. The difference is exactly zero on both, at p = 1.0000.

In the stratified sweep, the thumbnail flag across all 265 pages carrying it came out at +5.00 points, p less than 0.0001, in the wrong direction. A page whose worst capsule problem is legibility is a page with nothing worse to say about its capsule. Our own rubric reaches for it as a default when the art is basically fine, and we would not have found that without running the control.

So: if your report comes back saying your text is a bit small at thumbnail scale, that is close to a clean bill of health. If it comes back saying the capsule does not say what the game is, that is the 88.

One note on quoting our own evidence back at you. The rubric sometimes writes capsule sizes as 460x215 and 231x87. Those are render sizes, not upload specs. Steam’s current small capsule upload size is 462x174 and the header is 920x430, which our capsule sizes reference covers in full after Valve’s August 2024 asset change.

What this data cannot tell you

The flag is not capsule-local. This is the honest weakness and you should have it before you act on anything above. Pages carrying it also score lower on screenshots (-13, p less than 0.0001) and on description (-5.5, p = 0.0014). It does not predict tags at all (-4, p = 0.14). So the flag marks a page that fails to communicate in several places at once, not a page with one bad JPEG.

What it does not mark is a page with fewer assets. Flagged and unflagged pages have an identical median screenshot count of 9 (p = 1.00) and an identical median trailer count of 1. The material is there. The rubric’s verdict is that it does not say anything.

There is no human inter-rater check. One model looked at one image and wrote a sentence. Nobody independently re-rated a sample, so we cannot report agreement. The evidence strings above are legible and specific, which is encouraging, and it is not the same as validated.

The rubric is not Valve. A capsule that scores 58 here may be converting fine. We have no click-through, wishlist or sales data joined to any of these pages, and no result in this post is causal.

Two known scoring bugs. These 373 analyses were produced by a rubric with two defects we have since fixed in code: tag subscores were inflated by roughly 11 points by a padded tag count, and description subscores were depressed by roughly 11 points by a formatting check that read the wrong markup. Neither touches the capsule subscore, and because every page was distorted identically, comparisons between groups within the same category stay valid. Comparisons across categories do not, which is why this post never ranks the four subscores against each other.

How to check your own capsule

The test is cheaper than the analysis. Show your capsule to somebody who has never heard of your game, for one second, and ask them what the game is. If the answer is a genre, you pass. If the answer is “something with a fox in it”, you are in the 88.

Then run the mechanical checks: the capsule validator confirms your files match Steam’s current slots and dimensions before Steamworks rejects them, and the store page checklist covers the rest of the assets. The capsule design guide has the design process, and our tally of 98 top-seller capsules is a separate, complementary dataset: that one is a hand count of what shipping best-sellers actually do, measured from the images themselves, while this one is our rubric’s verdict on pages people asked us to audit. Different samples, different methods, different questions, so read them side by side rather than as one result.

For the rest of the picture, the full audit of all 373 pages covers every section, what the lowest-scoring pages get wrong is the bottom of this distribution up close, and the leaderboard shows public scores for pages already run. Our companion post on price and page quality works the same corpus from the pricing angle.

End of entry № 82

Field work

Put this entry into practice.

Run a free analysis on your Steam page and get specific, actionable fixes for your capsule, description, screenshots, and tags.

Continue reading

The Journal, weekly

Enjoyed this entry?

Get one actionable Steam page optimization tip in your inbox each week.