← All insights
Insights

Our Numbers and Everyone Else’s

How this site’s player value compares with EPM, RAPM, RPM, RAPTOR, BPM, VORP, Win Shares and PER: what each one counts, which ones actually predict, and why the box score can’t find Marcus Smart.

By Ben StupendousSep 25, 20266 min read
Illustrated portrait of Marcus Smart

Deb, from the group chat, asked me a fair question last week, in all caps, which is how Deb asks questions: why should anybody trust your number over Win Shares? Win Shares has been around for twenty years. Everybody knows what it is. Nobody has ever had to read a footnote about it.

It’s a good question, and it deserves better than “because I built it.” So here’s the whole thing: what each of the public numbers actually counts, one test of which ones predict anything, and the players where ours disagrees with everybody else’s.

What each one counts

The rules first. Every metric here is trying to answer the same question, how much a player helps his team, and they differ mostly in what evidence they’ll accept. Box-score metrics count what the scorekeeper writes down: points, rebounds, assists, steals. Plus-minus metrics ignore all of that and watch the scoreboard while a player is on the floor, adjusted for who he played with and against. The best modern ones do both, using the box score as a starting guess and letting the scoreboard correct it.

MetricBuilt fromRate or totalWhoStatus
PERBox scoreRateJohn Hollinger; Basketball-ReferencePublished
Win SharesBox score, plus team defenseTotalBasketball-ReferencePublished
BPMBox score, fitted to plus-minusRateBasketball-ReferencePublished
VORPBPM, turned into a total above replacementTotalBasketball-ReferencePublished
RAPMPlus-minus only, adjusted for teammates and opponentsRateVariousPublished
RPMPlus-minus, with a box-score priorRateESPNNo longer published
RAPTORBox score, tracking and on/offRateFiveThirtyEightEnded after 2022-23
EPMPlus-minus, with a box-score and tracking priorRateDunks & ThreesPublished
This sitePlus-minus, with a box-score and tracking priorRate, and WAR as the totalHerePublished

The other difference is rate versus total. A rate is how good a player is per possession; a total multiplies that by how much he played. Ours is a rate, and WAR is the total, measured above a replacement level the model estimates rather than assumes.1

Which ones predict?

Here is the test, in two sentences. Take every player’s number from last season, weight it by the minutes he actually played this season, and add it up for each team. A good metric should tell you which teams were going to be good.2 We ran every metric we could get through the same test, over the same seasons, 2013-14 to 2025-26:

MetricCorrelationError (pts/100)
This site (forecast)0.8122.79
BPM0.7812.99
PER0.7063.39
Win Shares per 480.6293.72

Correlation between the predicted and actual team ratings (higher is better) and the typical miss in points per 100 possessions (lower is better), 390 team-seasons.

Ours comes out on top, and the order below it is the order you’d expect: the more a metric trusts the scoreboard over the scorekeeper, the better it predicts. RAPTOR stopped after 2022-23, so it gets its own run over the seasons it covered, and the order holds: ours 0.823, RAPTOR 0.786, BPM 0.776.

Now the honest part. The two metrics I most wanted in that table aren’t in it. ESPN no longer publishes RPM, and EPM’s history isn’t public. Dunks & Threes, who invented this test, ran it on their own version of the target, and there EPM scores 2.48, RPM 2.60 and BPM 2.71. On ours, BPM scores 2.99, so the two rulers aren’t the same length, and comparing a number from one to a number from the other would be exactly the kind of thing Professor Voss used to circle in red. The fair comparison is against the common yardstick: EPM beats BPM by 8% on their test; we beat BPM by 7% on ours. Close. EPM’s edge is a hair bigger, and on two different rulers a hair is inside the margin. I’ll call it a tie rather than claim a win I can’t check.

Where we agree, and where we don’t

Among the 269 players who logged 1,000 minutes in 2025-26, our rating lines up best with EPM (a rank correlation of 0.84), then BPM (0.72), Win Shares per 48 (0.60) and PER (0.56). Our WAR tracks VORP at 0.78 and Win Shares at 0.71. The metric we agree with most is the one that did best on the test we can’t run, which I choose to find reassuring.

The disagreements are where it gets fun. These are the players we rank far above VORP:

PlayerWAROur rankVORPVORP rank
AJ Green5.656-0.8257
Marcus Smart5.751-0.1224
Herbert Jones4.576-0.5248

And far below it:

PlayerWAROur rankVORPVORP rank
Nic Claxton-0.52581.491
Keyonte George-0.42571.399
Jerami Grant0.52410.9130

The pattern is the whole argument between the two families. The players we like more are the ones whose value doesn’t get written down: the defender who blows up a set, the shooter whose man can’t leave him. The players the box score likes more are the ones who collect what does get written down. Neither is cheating. They are measuring different things, and only one of them was on the scoreboard.

Which brings me to Marcus Smart. The box score has him 224th by VORP. We have him 51st by WAR. I have loved Marcus Smart since the day he was drafted, so I want to be very clear that this is the model’s opinion and not mine, and also that the model is right.3

So which one should you use?

Use the one that answers your question. If you want to know who filled up the box score, PER and Win Shares will tell you honestly, and they’re easy to explain at a bar. If you want to know who made his team better, use one that watches the scoreboard: EPM, or ours. Just know what each one is counting before you bring it to the group chat. Deb will ask.

Figures: 2025-26 from Basketball-Reference (Win Shares, VORP, BPM, PER) and Dunks & Threes (EPM, as of 2026-06-13); RAPTOR from FiveThirtyEight’s final release; the prediction test from metric_compare.py on this site’s harness, as of this build. Wins are on this model’s own scale and dollars at the open-market price of a win; the method explains both.

  1. VORP sets replacement at a BPM of −2.0 by convention. Ours sits at −1.77 per 100, estimated from the games; what WAR means has the arithmetic. ↩
  2. This is the Dunks & Threes “metric comparison” protocol: the target is each team’s schedule-adjusted net rating; players under 250 minutes, and rookies, get a flat replacement rating. It isolates the rating from the minutes, which are taken as known. ↩
  3. Win Shares has him 210th. Deb’s position is that she didn’t need a model to tell her. Fair. ↩
More insights
All insights →