Wine scores are so embedded now that it's easy to forget how recent they are. The hundred-point scale arrived in wine criticism in the 1970s and 80s, spread rapidly because it was intuitive to American consumers familiar with school grading, and within two decades it had reorganised the entire market.
It's a genuinely useful tool that has done a considerable amount of damage, and both halves of that are worth taking seriously.
What it got right
Before scores, wine writing was descriptive and comparative, which made it nearly useless for anybody choosing between bottles in a shop. A score is instantly comparable, requires no expertise to interpret, and travels across languages.
It also democratised something. A consumer with no knowledge and no relationship with a merchant could walk in, look at numbers, and buy something decent. That was a real improvement over a system where knowing what to buy required social capital.
And it created accountability of a sort. A critic who scores publicly can be checked against outcomes, which is harder with prose.
The compression problem
Here's the first structural flaw. It's called a hundred-point scale and it functions as roughly a fifteen-point one.
In practice, almost no commercially reviewed wine scores below 80. The meaningful range is about 85 to 100, and the range that affects purchasing decisions is about 88 to 96.
That means the entire spectrum of wine quality is being expressed in a band narrow enough that the difference between adjacent scores is well within the noise of human tasting variation. A 91 and a 93 are not reliably distinguishable, by anybody, and yet that gap has significant commercial consequences.
The precision implied by an integer score is far greater than the precision the underlying judgement can support. That's the core problem and everything else follows from it.
The price effect
Scores move prices, and the relationship is nonlinear.
Research into wine pricing has repeatedly found that crossing certain thresholds — particularly 90, and again around 95 — produces disproportionate price increases. A wine scoring 89 and one scoring 90 are indistinguishable in quality and meaningfully different in market value.
That creates enormous pressure on producers, and where there's pressure there's adaptation.
How it changed the wine
This is the part that matters most and it's well documented.
Certain critics had identifiable preferences — for concentration, ripeness, power, new oak, high alcohol. Wines in that style scored well. Wines in that style sold for more. So producers made more of them.
Over roughly two decades, a lot of regions saw their wines get riper, more extracted, more oaked and higher in alcohol. Some of this was climate, which is real and independent. But a substantial part was producers chasing scores, and plenty of winemakers have said so openly since.
The consequence was homogenisation. Wines from very different places started tasting more similar, because they were all being optimised against the same taste.
The counter-movement — lower alcohol, less extraction, less new oak, more emphasis on freshness — has been substantial in the last fifteen years, and it's partly a rejection of exactly this.
The reliability question
There's a body of work on the consistency of wine scoring that's worth confronting.
Studies of wine competition judging have found that individual judges frequently give the same wine noticeably different scores when it's presented multiple times blind in a single session. Agreement between judges on the same wines is often weak. Medals awarded at one competition correlate poorly with medals at another for the same wines.
This doesn't mean tasters are incompetent. It means that fine discrimination between similar wines is genuinely hard, that palate fatigue is real, and that context effects — what you tasted immediately before — are substantial.
What it does mean is that treating a two-point difference as meaningful information is not supportable.
What I'd rather see
Some alternatives that seem better and get less traction because they're less convenient.
Bands rather than integers. Categories — outstanding, very good, good, sound — that reflect the actual resolution of the judgement. Several publications do this and it's more honest.
Scores with a stated range. "92, plus or minus 2" would be more accurate and would immediately kill the threshold effects.
Separate scores for different things. Quality now versus potential in ten years are different judgements and get collapsed into one number.
Prose with a recommendation. Which is what serious criticism in most fields does, and which requires the reader to engage rather than scan.
How to use scores anyway
Since they exist and aren't going away, some practical guidance.
Find a critic whose palate matches yours rather than trusting scores generically. A 95 from somebody who loves everything you dislike is a warning, not a recommendation.
Treat differences under about three points as noise.
Pay more attention to the note than the number, if a note is available. It'll tell you the style, which is what actually determines whether you'll like it.
And be aware that unscored wine isn't worse wine. Enormous quantities of excellent wine are never reviewed, because reviewing requires the producer to submit samples and small producers frequently don't bother.