Issue 01 · Published 2026

Where games lose their players

Some games lose the battle in the first two hours, while the refund button is still warm. Others lose trust in month two, right when the studio is celebrating a solid launch. We walked the whole route — from first launch to a thousand hours — and mapped the places where trust breaks.

202,000 reviews 3,111 games 15 genre groups 4 lifecycle stages 50,500 per stage

Limit

An honest warning before you read. We compare reviews from different games at different ages — not the same players over time. The numbers describe how players write, not what caused it. Limits

And this is a sample built for fair comparison between groups of games, not a portrait of all of Steam.

The strongest findings

  1. 01 ×2 The verdict arrives before the refund window closes 37.6% of refund-window reviews are negative, against 18.3% after it
  2. 02 +9 pp Every game gets a honeymoon. It ends in month two 17.7% negative at launch, 26.9% in months two through six
  3. 03 98% Ninety-eight percent of reviews get no answer only 2% get an official response, and it goes to criticism 4.3 times more often than to praise
  4. 04 16 pp Fourteen languages, a sixteen-point spread 31.9% in the Traditional Chinese segment, 15.7% in Latin American Spanish

Field 01 · Refund window Plate 01 · Field observation CritField · Steam Reviews 202K

The verdict arrives before the refund window closes

Plotted across a game's playtime, negativity draws a U — and its highest point sits inside the two hours where Steam still allows a refund. Reviews written there are negative twice as often as reviews written after.

37.6% of reviews written inside the refund window are negative N = 36,803 refund-window reviews ×2the measure 18.3% reference — the average once the two-hour boundary is passed
37.6% of reviews written inside the refund window are negative N = 36,803 refund-window reviews

Plate 01

Negativity draws a U: the newest and the oldest players are the unhappiest

% of reviews that are negative, by playtime band

Negativity by playtime band, plotted against the 18.3% after-boundary reference 40% 30 20 10 0 BOUNDARY · REFUND WINDOW ENDS AT 2H REFERENCE · 18.3% AFTER THE BOUNDARY ×2 17% 16.7% 20.1% 29% 29.6% 0–2h 2–10h 10–50h 50–200h 200–1,000h 1,000h+ SIX PLAYTIME BANDS · CATEGORICAL INTERVALS, EQUALLY SPACED — HORIZONTAL DISTANCE IS NOT ELAPSED TIME OBSERVATIONS JOINED TO SHOW SHAPE, NOT A CONTINUOUS SERIES · DOMAIN 0–40%
02040%
0–2h 37.6%

Boundary · refund window ends at 2h

2–10h 17%
10–50h 16.7%
50–200h 20.1%
200–1,000h 29%
1,000h+ 29.6%

Six playtime bands · categorical intervals · domain 0–40%

Peak 37.6% @ 0–2h · 36,803 reviews Reference 18.3% Measure ×2 Floor 16.7% @ 10–50h Tail 29.6% @ 1,000h+ · 2,210 reviews
Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific CritField · Steam Reviews 202K

Trust runs highest in the middle: newcomers and long-time veterans are both harder to satisfy than everyone in between.

LimitPlaytime is measured at the moment the review was published, not total hours ever played.

In the full studyChapter 02 · First hours

Field 01Refund window

Field 02Review language

From when a review is written to the language it is written in.

Field 02 · Review language Plate 02 · Field calibration CritField · Steam Reviews 202K

Fourteen review languages, sixteen points apart in tone

Ranked by the share of reviews that are negative, the fourteen review-language segments in this corpus register from 31.9% down to 15.7% — a sixteen-point band around the study's 22% reference. This is the language a review was written in. The mix of games behind each segment differs, so the band describes the reviews, not the people who wrote them.

HIGH TRADITIONAL CHINESE LOW LATIN AMERICAN SPANISH DOMAIN 14–34% · VERTICAL POSITION IS THE MEASURE REFERENCE 22% STUDY NEGATIVITY 31.9% 15.7% 16 pp MEASURED SPAN

Plate 02

Fourteen review languages, ranked against one reference

% of reviews that are negative, by the language the review is written in

16 pp between the highest and the lowest review-language segment 31.9% · Traditional Chinese — 15.7% · Latin American Spanish
Fourteen review-language segments ranked by negativity and read against the 22% study reference 34% 30 26 18 14 22 STUDY REFERENCE · 22% 31.9% 15.7% 16 pp OBSERVED SPAN RANKED REGISTER · FOURTEEN REVIEW-LANGUAGE SEGMENTS 01 TRADITIONAL CHINESE 31.9% 02 GERMAN 24.4% 03 JAPANESE 24.2% 04 SIMPLIFIED CHINESE 24.1% 05 KOREAN 23.2% 06 RUSSIAN 22.7% 07 TURKISH 21.3% 08 ENGLISH 21.1% 09 FRENCH 21% 10 POLISH 18.9% 11 ITALIAN 18.7% 12 SPANISH 17.7% 13 BRAZILIAN PORTUGUESE 16.5% 14 LATIN AMERICAN SPANISH 15.7% REFERENCE 22% · 6 SEGMENTS ABOVE / 8 BELOW FOURTEEN REVIEW-LANGUAGE SEGMENTS · RANKED, EQUALLY SPACED — HORIZONTAL DISTANCE CARRIES NO MEASURE VERTICAL POSITION IS THE ONLY MEASURE · DOMAIN 14–34% · REFERENCE 22%

1422% ref34%

01Traditional Chinese31.9%
02German24.4%
03Japanese24.2%
04Simplified Chinese24.1%
05Korean23.2%
06Russian22.7%
07Turkish21.3%
08English21.1%
09French21%
10Polish18.9%
11Italian18.7%
12Spanish17.7%
13Brazilian Portuguese16.5%
14Latin American Spanish15.7%
High 31.9% · Traditional Chinese · 1,915 reviews Reference 22% study negativity Low 15.7% · Latin American Spanish · 1,627 reviews Span 16 pp Field 14 segments · 6 above the reference, 8 below
Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific CritField · Steam Reviews 202K

Six segments sit above the study reference and eight below it, and the two ends of the field are sixteen points apart. The effect is real and it is weak — enough to change which baseline you compare a number against, not enough to carry a headline on its own.

What this is notThis is the language a review was written in, not a nationality and not a “national temperament” — the game mix differs between language segments, and 22,065 reviews with no language recorded are excluded. Fourteen language segments are displayed here; the corpus itself records 29 review languages in all.

In the full studyChapter 05 · Language & scale

Enter the full study · 7 chapters

  1. 01Lifecycle
  2. 02First hours
  3. 03Community response
  4. 04Veteran reviews
  5. 05Language & scale
  6. 06Review length
  7. 07Publisher action map

Continues below with Chapter 01 · Lifecycle

01Lifecycle

Every game gets a honeymoon — and a month when the patience ends

Right after launch, players want to believe. The real cross-examination starts in month two — and only quiets down over years.

Figure 1.1

Criticism peaks in month two, not at launch

% of reviews that are negative, by stage of a game's life

Negativity keeps climbing for months after the launch-week calm, then eases only gradually.

"Strict recount" runs the same numbers on a more conservative slice of the sample — the picture holds: 17.7% → 28.2%.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

Figure 1.2

At launch, one review in three comes from the first two hours

% of reviews written inside the first two hours of play

The audience changes along with the mood — that's the key to reading any comparison between stages.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

What the numbers say

In the first 28 days, 17.7% of reviews are negative: players give the game a chance, forgive rough edges, enjoy the novelty. In the window from one to six months, criticism jumps to 26.9% — the peak of the whole life cycle. After that the tension slowly eases: 25.0% in the mature phase and 18.2% for legacy games. The worst time to go quiet is the second month after launch: launch week, with its 17.7% negativity, looks like a win, the monitoring budget gets wound down — and the criticism curve is just starting to climb.

What to do

Plan content, fixes, and communication for the one-to-six-month window as seriously as for launch itself. Don't switch off review-tone monitoring two weeks after release. For strategy and other slow-burn genres, keep a support plan for year two: their wave of criticism arrives last.

What we can't claim: we compared reviews written in different windows of different games' lives — not the same people over time. So "players grow disillusioned by month two" does not follow from this data; what follows is only that reviews published in that window are markedly more negative.

Deep notesHow the sample was built · who is writing at each stage

We took 50,500 reviews each from games at four stages of life — from the first weeks after release to old games past their second birthday — and compared the tone. Legacy games keep a remaining audience that mostly knows exactly what it came for. In full-length reviews the arc is even steeper: from 19.1% up to 30.4%.

The crowd also changes underneath you: right after release, almost one review in three is written inside the first two hours of play; for legacy games it's one in nine. It's not just the mood that shifts — it's who is writing.

Supporting evidenceWhere the arc is sharpest · median hours behind a review

Where it shows up most

The arc is sharpest in service-driven genres. In MMOs and live-service games, negativity climbs from 26.4% in the first month to 42.4% in months two through six; in battle royales, from 29.9% to 39.8% by the mature phase. Narrative indies spike too: 10.0% → 24.9%. Strategy games explode late — their peak lands in the mature phase (29.6%), when the community has played everything and is waiting for the big updates. Survival and horror are the exception: tone barely moves with game age (15–17% at every stage). Action RPGs and simulations only mellow with age, down to 13.9% and 10.6% in their legacy years.

Figure 1.3

Critics have played less, at every stage

median hours played when the review was written

At every stage of a game's life, the authors of negative reviews had played noticeably less than the authors of positive ones.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

Chapter archiveWhat else changes as a game ages
  • The writing ages with the game: median review length drops from 171 characters (first month) to 68 (legacy games), and the share of full-length reviews from 76% to 47%. A late review is more often a gesture — "a classic!", "don't bother" — than an argument.
  • "Helpful" votes age too: in the first months a negative review collects 5.6–5.8 votes on average; for legacy games, 1.9. The review page is most flammable in the first half-year.
  • The first two hours are judged hardest at games aged one month to two years: around 48% of those reviews are negative, versus 26% at brand-new releases. Newcomers forgive a fresh game more.
  • The criticism arc survived every check: a recount without the five biggest games, a recount on full-length reviews only, and the strict recount — the peak stays in months two through six in all of them.
02First hours

The first two hours are a trial. Two hundred hours in, it's a tribunal

A game is at risk twice: while the player can still get their money back — and after they've given it a piece of their life.

The Refund window field above mapped this pattern across every kind of game at once. Here the same question splits by type of game.

Figure 2.1

In live-service games, every second first-session review is a complaint

% of refund-window reviews that are negative, by type of game

For MMOs and MOBAs the first two hours are the toughest filter there is; narrative and casual games pass it far more gently.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

What the numbers say

Line up all 202,000 reviews by hours played and the negativity curve draws a letter U. The first two hours — the window where Steam usually allows a refund — produce 37.6% negativity: double everything that comes after. The 10–50 hour middle hits the minimum, 16.7%. But past 200 hours, discontent rises again, to 29.0%. On the left are people the game pushed away immediately. On the right are people who can't leave, but keep demanding more.

What to do

Test the first session as a product of its own: install, first boot, settings, tutorial, the first 30 minutes of play — with ticket priority above normal. In service genres, read the 200+ hour reviews as a separate feed: their complaints are about different things than newcomers' — and they arrive before the retention metrics move.

What we can't claim: the curve describes who writes a review and in what mood — not "stages every player goes through." And it names no causes: the data shows the language of players, not the mechanism of their discontent.

Deep notesWhy both edges of the U are expensive

Then the game earns trust: players with 2–10 hours sit at 17.0%, and past 1,000 hours negativity reaches 29.6%.

The first two hours are the only stretch where discontent converts into a refund immediately: the angry review and the refund often arrive together. And the right edge of the curve is discontent from the most expensive audience there is: someone with 200+ hours is clearly not a drive-by hater, and the community reads their words as expert testimony. This is not a handful of people, either — the sample holds more than ten thousand such reviews.

Supporting evidenceHow hard the first trial is, genre by genre

Where it shows up most

The left edge of the U is universal: newcomers are angrier than mid-hour players in all 15 genre groups, no exceptions. But the severity of that first trial varies wildly: in MMOs and live-service games, every second refund-window review is negative (57.8%); in MOBAs it's 45.8%, battle royales 43.6%, VR 43.3%. Narrative indies (20.1%) and puzzle games (21.4%) pass the same filter almost unscathed. The right edge is a property of the genre, not a law of nature: veteran negativity rises in 11 of the 14 groups with enough data, but in simulations, action RPGs, and Early Access games the veterans are actually kinder than the mid-hour crowd.

Chapter archiveHow the letter U was checked
  • The left edge (newcomers angrier than mid-hour players) repeats in all 15 genre groups. The right edge (veterans angrier) repeats in 11 of the 14 groups with sufficient data.
  • Both edges survive removing the five biggest games in the sample, and hold in full-length reviews, where veterans run 36.6% negative against 20.5% for the mid-hour group.
  • The reversed right edge in simulations (10.0% among veterans), action RPGs (14.6%), and Early Access (14.8%) is not a single-game artifact: the biggest game accounts for only 19.5% of the simulation group.
  • 220 reviews with no playtime data are excluded from this cut.
03Community response

An angry review carries 3.6 times further — and almost nobody answers it

A game's storefront is shaped not by the average player but by the loudest one. And the loudest players are the unhappy ones.

×3.6how much more often a negative review collects 5+ "helpful" votes
4.5 : 1.2average "helpful" votes: criticism versus praise
2%of reviews receive an official developer response
×4.3how much more often that response goes to a negative review

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

What the numbers say

Steam surfaces the reviews the community found helpful — and that machinery favors criticism. Five or more "helpful" votes go to 19.6% of negative reviews and only 5.4% of positive ones. The first page of reviews is the shop window, and community votes pin negativity to it for a long time — sometimes longer than the problem itself lives: the bug is long fixed, and the essay about it still sits on top. Meanwhile the response field is nearly empty: only 2% of reviews get an official developer reply, and it goes to criticism 4.3 times more often than to praise.

What to do

Answer substantive criticism publicly, calmly, and once — especially in the first half-year, while every reply is read by thousands. Arguing with emotion is pointless. When a problem ships fixed, report back under the most visible review about it: that top negative review will be read either way, so let the resolution sit next to it. At today's 2% response rate, a systematic reply practice is a competitive advantage all by itself.

What we can't claim: how many votes a review collects also depends on Steam's visibility algorithms — we see the combined outcome of "text plus storefront," not the pure effect of the text. And this data doesn't measure whether a developer's reply changes anyone's mind: it only shows who gets the attention.

Deep notesVote averages, review length, and who replies

On average, criticism collects 4.5 votes to praise's 1.2. Negativity is also more thorough: the median angry review runs 183 characters against 106 for a kind one. The result is a front page of reviews with a heavy tilt toward the most dissatisfied and most articulate voices.

The official reply rate splits by genre too: developer responses reach 4.9% of negative reviews against 1.2% of positive ones. Free-to-play studios answer most often (5.1% of reviews), then VR teams (3.9%); puzzle developers answer least (0.7%).

Supporting evidenceWhen an angry review is worth the most

Where it shows up most

The loud-negativity effect peaks in a game's first half-year: in that window an angry review gathers 5.6–5.8 votes on average, and one negative text in four collects 5 or more. On older games the storefront cools down: the average negative review there draws just 1.9 votes. In other words, the reputational price of a single angry essay is highest exactly when the game has the most eyes on it.

04Veteran reviews

Veteran negativity is not ordinary negativity — and some genres barely have it

The same life cycle looks like different lives in different kinds of games: in some, the veterans are the game's advocates; in others, its loudest prosecutors.

Figure 4.1

Past 200 hours, some genres produce advocates and others prosecutors

% of 200+ hour reviews that are negative, by type of game

The average across all veterans is 29.1% negative, but that hides a four-fold gulf between genres.

Groups shown have at least 100 veteran reviews.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

The negativity map: game age × type of game

The redder the cell, the higher the share of negative reviews. A hatched cell means we honestly have no data for that group. Cells with few reviews carry a dot — read them with care. And treat the fiery free-to-play cell in the early months as one game's story, not a genre trend: a single troubled launch accounts for more than 99% of the reviews in that cell.

CritField · Steam Reviews 202K

The negativity map

15 genres, most negative first

What the numbers say

The sharpest contrast is among veterans past 200 hours. In MOBAs, 41.8% of their reviews are negative; in MMOs and live-service, 38.0%. At the other pole sit simulations (10.0%), narrative indies (12.7%), and action RPGs (14.6%) — there, hundreds of hours turn a player into a defender, not a critic. The average across all veterans is 29.1%, but that average hides a four-fold gulf between genres — and veteran negativity is the most expensive kind, written by people whose opinion the community reads as expertise.

What to do

Benchmark your game against your genre's norm, not against a Steam-wide average: 30% veteran negativity in a shooter is a Tuesday; 30% in a simulation is an alarm. In service games, give 200+ hour reviews their own monitoring feed — that's where complaints about grind, economy, and an empty endgame surface first.

What we can't claim: genre groups are groups of games with different audiences and different business models; the data shows where each tone is more common, not whether "the genre is to blame." A few cells of the map are empty or thin — we don't hide that.

Deep notesHow the groups were cut · when to treat it as an alarm

We split the games into 15 genre groups — from shooters and strategy to MMOs, VR, and free-to-play — and watched how player sentiment ages in each. Shooters sit at 36.3% among veterans.

In service genres veteran negativity is predictably massive. For MMOs, MOBAs, and shooters it is background weather you learn to work with, not an emergency. And the reverse holds: if a simulation's veterans suddenly start writing like MMO veterans, that is an anomaly worth investigating immediately.

Supporting evidenceWhy the gap is not a single-game artifact

Where it shows up most

The gap is not a single-game artifact: the largest title makes up only 6% of the MMO veteran group and 19.5% of the simulation group. Nor does it reduce to "angry genres": strategy veterans are calmer than their mid-cycle players (24.5%), while sports and racing veterans are angrier (32.8%). The likelier story is the relationship model. Service games live in a "playing and demanding" mode — enormous playtime coexists with grievances about the economy, the balance, the support. In "finished" games — sims, RPGs, indies — hundreds of hours mean the game already delivered what the player came for.

05Language & scale

Read every other number twice: once for the review's language, once for the game's review count

Neither signal is strong enough to lead a story. Both are strong enough to change which baseline a number should be compared against.

The full language evidence is already on this page: the Review language field above sets out all fourteen displayed language segments and the sixteen-point band between them. This chapter carries that dimension forward and meets it with a second, separate weak signal — how many reviews the game has collected in its life.

Figure 5.1

Bigger games draw less criticism — mostly because they survived

% of reviews that are negative, by lifetime review volume

Giants draw less negativity (11.9% versus 25.4% for small games) — largely because a game with a million reviews has already passed natural selection.

Review-volume tier describes how many reviews a game has accumulated — not its budget and not “AAA status.” 26,722 reviews from games of unknown tier are excluded from this cut.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

What the numbers say

Both effects are real and both are weak — which is exactly why they sit here rather than in a headline. Read either signal as context for the rest of the atlas, never as a finding on its own.

What we can't claim: a language segment is not a country and not an audience temperament — the mix of games differs between segments, so the spread describes the reviews, not the people who wrote them. Review volume is a measure of how much a game has been reviewed, not of its budget or its quality.

Supporting evidenceWhy the scale effect is mostly survivorship

What review volume is actually recording

The scale effect mostly records survivorship: a game with a million reviews has already passed a filter that small games are still standing in. It is a statement about which games are left to review, not about how those games were made.

06Review length

Praise is a thumbs-up. A complaint is an essay

A happy player clicks the thumb and goes back to the game. An unhappy one sits down to write.

Figure 6.1

Two reviews in three are full-length texts

% of the sample, by length of review

63% of the sample is full-length text; only 9% is an ultra-short "+", "gg," or an empty click.

Figure 6.2

The longer the review, the likelier it is negative

% of reviews that are negative, by length of text

Negativity climbs step by step with length — 11.1% among ultra-short reviews, 25.8% among full-length ones.

Source · CritField, Steam Reviews 202K · N = 202,000 reviews, 3,111 games · correlational, corpus-specific

What the numbers say

63% of reviews in the sample are full-length texts; 9% are "+", "gg", or entirely empty. And effort tracks tone directly: among ultra-short reviews only 11.1% are negative — and among full-length texts, 25.8%. This is a trap for any review analytics. Read only the "substantive" long texts and your picture is systematically darker than reality (25.8% negativity against 22% overall); look only at the thumbs and it's systematically rosier.

What to do

Use two lenses and never mix them: measure tone across all ratings, mine topics from full-length texts — always remembering the second lens runs dark. Count ultra-short reviews as votes for or against, but never as a source of topics.

What we can't claim: length doesn't make an opinion more correct — it only shows who was motivated to write. The effort–negativity link is correlational: we don't know whether people sit down to write because they're angry, or get angrier as they write.

Deep notesThe full length ladder · who writes the short forms

The ladder runs the whole way: 11.1% negative among ultra-short reviews, 15.4% among short ones, 18.0% medium, 25.8% full-length. Negativity isn't just more frequent, it's heavier: 73.7% of angry reviews are written at full length, versus 59.9% of kind ones. The ultra-short form is practically a monopoly of praise: "+" and "gg" come from the satisfied (9.8% of positive reviews) and almost never from the dissatisfied (4.3%).

Any report built on "the 100 most helpful reviews" is biased toward angry essays by construction.

Supporting evidenceHow the writing culture ages

Where it shows up most

The effort gap holds across the whole life cycle: at every one of the four game ages, negative reviews run longer than positive ones. But the writing culture itself ages: at fresh releases three reviews in four are full-length (76%); at legacy games, fewer than half (47%). So for old games both lenses narrow — praise and criticism alike shrink into gestures.

Chapter archiveHow reviews were sorted by length
  • The split is mechanical, by length and word count: empty → ultra-short (up to 3 words) → short (up to 40 characters) → medium (up to 80 characters) → full-length (80+ characters and 12+ words).
  • For languages written without spaces between words (Chinese, Japanese, Korean, Thai) the thresholds are adjusted — otherwise their reviews would unfairly land in "short."
  • 1,204 reviews with corrupted text encoding were repaired in a working copy; the source data was never modified.
07Publisher action map

The publisher action map: what to actually do with all this

These are not universal laws of game development — they're the sensible moves if your players resemble the players in this sample. Every point leans on a specific chapter of the atlas.

QA

The first session is a product of its own

The first two hours produce double the negativity — and they are the refund window. Turn the recurring first-hour complaints into test scenarios: install and boot across configurations, first-run settings, the tutorial, the first 30 minutes of play. Every "won't start / confusing / miserable opening" complaint becomes a ticket above normal priority. For MMOs and MOBAs this filter is twice as harsh as for narrative games.

UX

First-hour signals are onboarding review triggers

A spike of negativity in the refund window is a reason to re-examine onboarding and interface now, not after the next content patch. It's especially alarming when the game is already a few months old: at games aged one month to two years, nearly every second refund-window review is negative. A newcomer arriving at a not-so-new game is the strictest judge of all.

LiveOps

Monitor tone for six months after launch, not two weeks

The peak of criticism is not launch week — it's the window from month one to month six (26.9% against 17.7%). Plan content, fixes, and communication for that window in advance. For strategy and other slow-burn genres, keep the watch running longer: their wave of criticism arrives in the mature phase, in year two.

Community

Systematic replies are rare — which makes them an edge

Angry reviews collect 3.6 times more "helpful" votes and stay on the storefront longer than the problems they describe. Answer substantive criticism publicly, calmly, and once; after the fix ships, report back under the most visible review of the problem. Today only 2% of reviews get any reply: simply doing this systematically sets a studio apart.

Economy & design

Veteran reviews are the early warning system

In MMOs, MOBAs, and shooters, the most devoted players are also the loudest critics (36–42% negativity past 200 hours). Give their reviews a separate feed: complaints about grind, economy, DLC value, and an empty endgame appear there before newcomers or metrics see them. And if the veterans of a "peaceful" genre suddenly start sounding like MMO veterans — that's a design problem signal, not "toxicity."

Marketing

Don't promise a first two hours that doesn't exist

Players charge the gap between the trailer and the first session at the refund desk — with an angry review attached, on the most visible page of the store, in the loudest period of the game's life. Check the campaign's promises against what a player actually sees before hour two. And remember the analytics lenses: long reviews run systematically darker, short ones kinder — decisions made on "the most substantive texts" will be biased.

Archives

Every number in the atlas unfolds into a table

Pick a game age, a type of game, and an hours-played band — the table is built from the same data as every chart above. Nothing here is eyeballed.

Rows with fewer than 100 reviews are flagged as low-confidence.
CritField · Steam Reviews 202K

Data explorer

Rows with fewer than 100 reviews are flagged as low-confidence.
Method

How we separated signal from noise

What the data is

  • 202,000 Steam reviews of 3,111 games, written between January 2025 and May 2026, in 29 languages.
  • We deliberately took an equal number of reviews — 50,500 — from games at each of four ages: the first month after release, one to six months, six months to two years, and older than two years. Otherwise any comparison between stages would be comparing apples to a mountain of oranges.
  • Within each stage the sample is balanced across types of games and their scale, and no single game can occupy more than 850 reviews — so no loud release can shout over the rest.
  • Every review is counted once: no duplicates, verified against each review's unique identifier.

How we kept ourselves honest

  • We compared reviews across games of different ages and across players with different hours behind them — from the first session to a thousand-plus — and separately across different types of games.
  • When you test many ideas at once, some coincidences appear by pure chance. We corrected for that statistically and kept weak effects out of the conclusions.
  • Every headline conclusion was recounted three ways: without the five biggest games, on full-length reviews only, and within each genre group separately. Anything that didn't survive the recounts was cut from the published findings — not softened, cut.
  • We separately checked whether any effect was carried by a single game. One tempting genre-wide "sensation" died exactly that way, and was never published as a finding.
  • Beautiful stories with weak numbers didn't make it in — even the most quotable ones.

What our words mean

  • Game age — days between a game's release and the review's publication.
  • Hours played — how many hours the author had in the game when they wrote the review. "First 2 hours" is the window where Steam usually allows a refund.
  • Type of game — one of 15 genre groups; three of them are business or platform models rather than genres (Early Access, free-to-play, VR).
  • Game scale — how many reviews the game has accumulated over its life. It measures visibility, not budget.
  • Full-length review — 80+ characters and 12+ words; anything shorter counts as a vote, not as a source of topics.

What we refuse to do, on principle

  • No causal claims: reviews show how players describe their experience, not what caused it.
  • No stories about "the same players a year later" — we don't have that data.
  • No projecting onto "all of Steam": the sample was built for fair comparison between groups, not as a shrunken copy of the store.
  • No judging nationalities by review language, and no confusing a game's scale with its budget.
  • No hiding the gaps: for open-world games in the mature phase we simply have no data — that cell of the negativity map is honestly hatched.
Method & limitsFor specialists: statistical details
  • Proportion comparisons use two-proportion z-tests with Wilson confidence intervals; effect size is Cohen's h. Multi-group comparisons use chi-square with Cramér's V; skewed quantities (length, votes) use Mann–Whitney tests.
  • Multiple-comparison correction is Benjamini–Hochberg (FDR) across the whole family of main tests; only effects that stayed significant after correction and meaningful in size made it into the conclusions.
  • Robustness checks: excluding the five largest games; recounting on full-length texts only; repeating per genre group; checking for single-game concentration of each effect; a strict recount of the criticism arc on a conservative slice of the sample.
  • Known limitations: one genre cell is empty (open world × mature phase); 22,065 reviews carry no language; the "games owned by author" field is unreliable due to private profiles and is not used in the analysis.
CritField · Steam Reviews 202K

Chapter route