#1Debut

Biggest or greatest? The argument I have not settled

By the person who builds #1 Debut · Last updated 7 August 2026

Every rating in the game is a number between 55 and 99. I am still not entirely sure what that number is claiming.

This is the open question in the rating pipeline, and unlike the other write-ups here it does not end with a fix. It ends with a decision nobody has made yet, and a stopgap I am not especially proud of.

What the scale currently measures

The top of the scale is ordered largely by verified streams and YouTube views. Both of those measure the same underlying thing: how many people played it. That is a popularity measure, and it has real virtues. It is reproducible, it is checkable by anyone with access to the same public numbers, and it does not quietly encode one person's taste into a catalog of several hundred albums.

Critical stature is handled separately, through a curated set of acclaim signals covering a few dozen tracks, and through chart peaks and certifications which can earn a song relief from its era ceiling. Those inputs are real, but their coverage is thin compared to streaming data, which is available for essentially everything.

The consequence is structural. In the low-to-mid nineties, songs that are enormously popular sit directly beside songs that are canonical, and the scale does not distinguish between the two. A recent single with a billion plays and a record that reorganised its genre can land on the same number, for the same stated reason.

The case for biggest

Popularity is measurable. Two people looking at the same song and the same public data will produce the same rating, and if they disagree, the disagreement is about a number rather than about taste. Across a catalog this size, that consistency is worth a great deal. It is the difference between a rating system and a very long opinion.

It also fails honestly. When the ratings are wrong in a popularity-ordered scale, they are wrong in a way you can point at: the stream count was mismatched, the age adjustment was too aggressive, the certification was missing. Every one of those is fixable by someone who is not us.

And there is a decent argument that popularity is simply the correct measure for this particular game. You are building an album that gets projected onto a chart. Charts measure commercial performance. A scale that tracks what people actually played is arguably the one that belongs underneath that fiction.

The case for greatest

The counter-argument is that nobody plays this game to simulate the sales chart. They play it to build a good record, and "good" is not a synonym for "streamed".

A purely popularity-ordered scale produces specific results that feel wrong to anyone who knows the music. Deep cuts that are widely considered among an artist's finest work sit in the seventies because they were never singles. Records from before streaming existed are systematically undervalued, since their stream counts measure residual curiosity rather than the scale of the thing at the time. Meanwhile a modest recent hit riding a playlist can outrank both.

There is also a subtler cost. If the scale is pure popularity, the drafting skill it rewards is recall of chart performance rather than judgement about music, and that is a less interesting game. The version where you have to weigh whether a song is genuinely great, and are sometimes rewarded for taking the album track over the single, is the better one to play.

The stopgap, and why it is unsatisfying

Where a rating is clearly indefensible, the pipeline supports a pin: a hand-set value that overrides the computed one. There are a few dozen, each individually approved, and they exist to fix cases where the popularity ordering produces a result nobody can defend.

Pins work and they do not scale. Every pin is a small admission that the method produced the wrong answer and a person corrected it by hand. A few dozen is manageable. A few hundred would mean the ratings are a curated list wearing a pipeline as a costume, and the reproducibility that makes the whole thing trustworthy would be gone.

The pins are also unevenly distributed by definition. They cluster on songs somebody happened to notice, which means the catalog is quietly more accurate in the corners I listen to most. That is the same coverage problem the acclaim data has, just less visible.

What would actually resolve it

The systematic fix is broad acclaim coverage from a published critical source, applied across the catalog rather than curated by hand on the tracks I thought to check. With acclaim available at something like the coverage streaming data has, it could be weighted into the blend properly and the top band would stop being ordered by popularity alone.

That is a real project rather than a tweak, and it cannot be started until the underlying question is answered, because the weighting is the answer. Decide the scale means "biggest" and acclaim stays a tiebreaker. Decide it means "greatest" and acclaim becomes a primary axis, with streams demoted to evidence of reach.

My honest position is that the scale currently means "biggest, with corrections", and that this is a description of where the data got me rather than a decision I ever defended on the merits. Saying so in public is uncomfortable, which is roughly why it is worth doing. A rating system that hides its open questions is not more trustworthy than one that publishes them. It is just harder to check.

If you have a view, it is genuinely useful. This is the question I get the least mail about and think about the most.

Play #1 Debut