A game drops. You pull up Metacritic. The Metascore reads 74. You close the tab and move on.
That number just made a decision for you. Maybe it was the right one. Maybe the game you just dismissed was exactly what you needed. You will never know, because 74 looked like a failing grade and you trusted the aggregate over your own curiosity.
Metacritic has been the dominant force in game review aggregation since its founding in 1999, and its Metascore has become the closest thing the industry has to an official verdict on whether a game is good. Publishers cite it in marketing campaigns. Developers have had bonuses tied to it. Retailers use it to decide shelf space. An entire infrastructure of financial and cultural decisions rests on a single number that fewer people than you might think actually understand.
This is not an argument that Metacritic is useless. It surfaces real criticism and aggregates genuine opinions. But the way most people use the Metascore — as a straightforward verdict — misunderstands what it actually measures. Understanding the gap between what the score looks like and what it genuinely represents is how you start making better decisions about the games you play.
The Black Box Behind the Number
Metacritic calculates the Metascore as a weighted average of individual critic scores, all converted to a 100-point scale. The weighting is where things get complicated — and opaque.
According to Metacritic's own explanation, they assign different weights to different publications based on their perceived quality, scale, and industry standing. A review from a major outlet counts more toward the final Metascore than a review from a smaller one. Metacritic has consistently refused to reveal which outlets receive which weights or how those weights are assigned.
Researchers at Full Sail University attempted to reverse-engineer the system and identified six distinct weighting tiers across 188 publications, as reported by Game Developer. The highest tier, assigned a 1.5x multiplier, included 26 publications such as IGN, Game Informer, and The New York Times. The lowest tier used a 0.25x multiplier on 15 smaller sites. Metacritic disputed the findings, calling them "wildly, wholly inaccurate" — but offered no transparency about what the actual weights are.
The practical consequence is this: when you look at a Metascore, you cannot know which reviews mattered most in producing it. A 74 might be driven primarily by a handful of major outlets whose preferences skew toward a particular genre, or it might reflect genuine consensus across dozens of publications. The score looks like math. The inputs feeding it are hidden.
When Scales Don't Match
The weighting is only part of the problem. Individual critics use wildly different scoring systems, and Metacritic converts them all to a 100-point scale using its own mapping rules — rules that can distort meaning in the process.
Consider a critic who uses a five-point scale where 3/5 means "good, recommended." Metacritic converts that to 60. On a 100-point scale, 60 typically reads as below average or disappointing. The critic recommended the game. Metacritic's conversion suggests they didn't. As Kotaku documented in their deep-dive on the Metascore's problems, critic Tom Chick's 3/5 reviews are consistently interpreted negatively by Metacritic's conversion despite his intention being the opposite.
This matters at scale. When 80 critics submit reviews using eight different scoring systems, and Metacritic converts them all through a single mapping process, the aggregate can end up representing something no individual critic actually said. The number that emerges is mathematically derived from the inputs but semantically divorced from what those inputs meant.
The Bonus Clause Problem
The distortions built into the Metascore would matter less if the score had purely cultural weight. The problem is that it developed financial weight too — and that financial weight feeds back into how games are made.
The most widely cited example involves Fallout: New Vegas. Developer Obsidian had a bonus clause in their contract with publisher Bethesda that would pay out a significant sum if the game hit an 85 Metascore. The game released to generally strong reviews and scored 84 on Xbox 360 and PC. One point separated Obsidian from a bonus that would have amounted to roughly $1 million split across the team. They received nothing.
That gap — 84 versus 85 — was the difference between individual converted scores and Metacritic's weighting decisions. It was not a gap that reflected any meaningful difference in quality. But it had real financial consequences for the people who made the game.
Bonus clauses tied to Metacritic thresholds have been documented across the industry. CD Projekt previously conditioned developer bonuses on hitting a 90 Metacritic average before eventually dropping the requirement. 2K Games once used Metacritic averages in job applications, requiring candidates to list the Metascore of games they had shipped. The score stopped being a reflection of how good a game is and became a metric that shaped what games get made, how developers are compensated, and who gets hired.
When bonuses depend on scores, developers face pressure to optimize for what reviewers reward. Scripted cinematic sequences, multiplayer modes added late, additional missions that bulk up apparent value — these are the fingerprints of development shaped by anticipated Metacritic performance rather than by design vision.
Review Bombing and Score Manipulation
The Metascore measures critic consensus. The User Score underneath it measures something closer to organized participation.
User scores are vulnerable to review bombing — coordinated campaigns that flood a game with negative ratings for reasons unrelated to its quality. The 2017 Firewatch incident remains the clearest case: after developer Campo Santo took a legal action against a YouTuber, angry fans targeted the game's Steam rating and drove it from "very positive" toward "mixed." The game did not become worse. The score changed because an organized group decided to make it change.
Metacritic's user review system has faced similar dynamics. Games featuring diverse characters or non-traditional narratives have attracted coordinated negative reviews aimed at cultural objections rather than gameplay assessment. The user score on Metacritic can reflect what a particular subset of the internet decided to do on a particular day as much as it reflects what players actually think of the experience.
The critic-weighted Metascore is less vulnerable to this specific problem but has its own version. Publishers who host lavish review events, provide extended access under favorable conditions, or select review outlets strategically can influence which reviews appear and when. Metacritic does not control the circumstances under which reviews are written — only the math that aggregates them.
The Patch Problem
Here is a failure mode that rarely gets discussed: most reviews are written at or near launch, and most review outlets do not update their scores when patches significantly change the game.
A game that releases with serious performance issues might earn a 65 Metascore based on a broken launch state. The developer releases three major patches over six weeks, addresses the problems, and ships a fundamentally different experience. The Metascore still reads 65. The underlying reviews — each representing a specific build of the game reviewed under specific conditions — remain static.
Metacritic locks scores once published, a policy originally implemented to protect critics from publisher pressure to revise upward after release. The protection is real. But the side effect is a record that captures a game's launch state indefinitely, even when that state is no longer the experience players encounter.
OpenCritic: A More Transparent Alternative
The transparency problems in Metacritic's methodology directly motivated the creation of OpenCritic, which launched publicly in September 2015. Founded by Matthew Enthoven from Riot Games and engineer Charles Green, OpenCritic uses a simple arithmetic mean rather than a weighted average. Every review counts equally. The formula is visible. Users can see which publications are included and which critics wrote the reviews.
As Game Developer noted at the time of OpenCritic's launch, the site also includes non-scored reviews and YouTube video reviews — content that Metacritic excludes entirely. This means games reviewed primarily by video critics, which is increasingly common for certain genres, receive fairer representation on OpenCritic than on Metacritic.
OpenCritic also displays the percentage of critics who recommend a game alongside the numeric score. This matters because "recommended" is a cleaner signal than a number. A game at 76 where 80% of critics recommend it communicates something different from a game at 76 where only 45% do.
OpenCritic is not perfect. Its simple average can be skewed by outlier reviews just as Metacritic's weighted average can be skewed by its opaque tier system. Acquired by media company Valnet in July 2024, it faces the same questions about editorial independence that Metacritic faces. But it gives you more information to work with, and the information it shows you is the information it actually used.
What a Score Actually Captures
Even a perfectly calculated aggregate score has a fundamental ceiling on what it can tell you.
A Metascore is an aggregate of professional opinions produced under specific conditions, converted through various scale-mapping rules, and weighted according to an undisclosed formula. It captures the average reaction of a specific class of critics who reviewed a game during a specific window, typically at launch. It does not capture how the game ages. It does not capture whether the experience matches your taste in particular. It does not capture community reception, modding potential, long-term playability, or how the game compares to others in its genre that you specifically like.
The score is useful as a rough signal. A 45 Metascore usually indicates a game with serious problems that most critics found frustrating. A 90 usually indicates a game that critics broadly loved. In the large middle — 65 to 85 — where most games land, the number is far less reliable as a predictor of whether you will enjoy the experience.
This is the range where reading the actual reviews matters. Two reviews that both give a game 7/10 might do so for completely opposite reasons. One critic loved the combat and found the story weak; the other loved the story and found the combat repetitive. You might be someone who plays for story. The aggregate tells you nothing about which experience awaits you.
Building Your Own Informed Take
The alternative to trusting the aggregate is building a relationship with critics whose taste you understand. This is how people engaged with game reviews before aggregation existed, and it is still the most reliable method for matching reviews to your own needs.
Find two or three critics whose taste aligns with yours — specifically on the genres you play most. Pay attention to their reasons more than their scores. When they love a game for reasons you also value, that recommendation carries more predictive power than a Metascore assembled from critics who may play entirely different things.
You can explore reviews across the full spectrum of what critics say on your games page to cross-reference opinions on titles you're considering. And when you finish a game yourself, writing down what you actually thought — in your own review space — builds a personal record that the aggregate never can: your genuine reaction to the specific experience you had.
This is the kind of record that the broader review culture evolution points toward: not more aggregation, but more personal authorship. The Metascore tells you what a weighted average of professional critics thought. Your own review history tells you who you are as a player, what you value, how your taste has changed over time. One of those records compounds in value the longer you maintain it. The other resets to zero with every game.
The Number Is Not the Verdict
Metacritic is a useful tool misused as a verdict. The Metascore aggregates real criticism and produces a number that contains genuine signal. But the signal is filtered through a weighting system nobody outside Metacritic can fully verify, a scale conversion system that distorts individual meaning, a launch-window snapshot that does not update, and a financial incentive structure that has shaped what games get made.
None of this means you should ignore Metacritic. It means you should use it the way you would use any single data point: as one input among several, not as the final word.
The final word belongs to you. Read the reviews. Watch the critics whose taste you know. Play the game your instincts tell you to play, even if the 74 says otherwise. Build your own record of what you thought. That record — not an aggregate, not a weighted mean, not a score that carries bonus clauses — is the one that actually reflects your experience as a player.