# MissedTakes — full text for language models The short version, with the two things a careless summary gets wrong, is at https://missedtakes.co/llms.txt. Read that first. This file is the long form: the ranking as a table, every prediction type with graded content, and the complete scoring methodology. ## The NFL ranking Pundits with ten or more graded articles, ranked by MissedTakes Score on their NFL record — how much better than a naive guess on the same claims, NOT raw accuracy. Everyone below the ten-article bar has a page with their numbers but is not ranked; the directory is at https://missedtakes.co/pundits. The US Elections ranking is its own, at https://missedtakes.co/elections/leaderboard. | # | Pundit | MissedTakes Score | Accuracy | Graded calls | Graded articles | Page | | --- | --- | --- | --- | --- | --- | --- | | 1 | Damonza Byrd | 65.8 | 58% | 14 | 10 | https://missedtakes.co/pundits/damonza-byrd | | 2 | Lance Zierlein | 51.12 | 67.54% | 44 | 20 | https://missedtakes.co/pundits/lance-zierlein | | 3 | Pat Fitzmaurice | 50.39 | 68.12% | 124 | 24 | https://missedtakes.co/pundits/pat-fitzmaurice | | 4 | Joe Marino | 46.24 | 77.57% | 170 | 12 | https://missedtakes.co/pundits/joe-marino | | 5 | Tom Pelissero | 44.92 | 82.36% | 34 | 18 | https://missedtakes.co/pundits/tom-pelissero | | 6 | Bryan DeArdo | 43.99 | 65.15% | 20 | 11 | https://missedtakes.co/pundits/bryan-deardo | | 7 | Bill Huber | 43.77 | 83.33% | 62 | 12 | https://missedtakes.co/pundits/bill-huber | | 8 | Derek Brown | 40.56 | 64.43% | 107 | 14 | https://missedtakes.co/pundits/derek-brown | | 9 | Peter Schrager | 35.4 | 70.35% | 48 | 18 | https://missedtakes.co/pundits/peter-schrager | | 10 | Daniel Jeremiah | 33.86 | 74.65% | 276 | 40 | https://missedtakes.co/pundits/daniel-jeremiah | | 11 | Bucky Brooks | 32.84 | 76.53% | 118 | 26 | https://missedtakes.co/pundits/bucky-brooks | | 12 | Rich Hribar | 32.68 | 68.79% | 64 | 11 | https://missedtakes.co/pundits/rich-hribar | | 13 | Ross Jackson | 31.84 | 65.95% | 234 | 32 | https://missedtakes.co/pundits/ross-jackson | | 14 | Brian Peacock | 31.59 | 70.21% | 82 | 10 | https://missedtakes.co/pundits/brian-peacock | | 15 | Mike Renner | 30.75 | 61.2% | 124 | 12 | https://missedtakes.co/pundits/mike-renner | | 16 | Jamie Erdahl | 28.21 | 35.4% | 59 | 19 | https://missedtakes.co/pundits/jamie-erdahl | | 17 | Albert Breer | 26.74 | 61.69% | 34 | 10 | https://missedtakes.co/pundits/albert-breer | | 18 | Kevin Wildes | 26.29 | 44.74% | 22 | 12 | https://missedtakes.co/pundits/kevin-wildes | | 19 | Rich Eisen | 24.37 | 60.51% | 461 | 130 | https://missedtakes.co/pundits/rich-eisen | | 20 | Shannon Sharpe | 20.84 | 63.61% | 32 | 21 | https://missedtakes.co/pundits/shannon-sharpe | | 21 | Nick Wright | 20.07 | 59.46% | 378 | 72 | https://missedtakes.co/pundits/nick-wright | | 22 | Chris Brockman | 20.05 | 59.05% | 285 | 87 | https://missedtakes.co/pundits/chris-brockman | | 23 | Gregg Rosenthal | 19.58 | 62.3% | 316 | 70 | https://missedtakes.co/pundits/gregg-rosenthal | | 24 | Andrew Erickson | 19.11 | 65.87% | 102 | 16 | https://missedtakes.co/pundits/andrew-erickson | | 25 | Cynthia Frelund | 18.73 | 60.25% | 100 | 28 | https://missedtakes.co/pundits/cynthia-frelund | | 26 | Richard Sherman | 17.03 | 56.43% | 137 | 18 | https://missedtakes.co/pundits/richard-sherman | | 27 | Jordan Dajani | 16.77 | 56.57% | 48 | 17 | https://missedtakes.co/pundits/jordan-dajani | | 28 | Marcus Mosher | 16.4 | 80.47% | 77 | 14 | https://missedtakes.co/pundits/marcus-mosher | | 29 | Tyler Sullivan | 16.29 | 57.71% | 79 | 30 | https://missedtakes.co/pundits/tyler-sullivan | | 30 | John Breech | 15.1 | 48.81% | 120 | 27 | https://missedtakes.co/pundits/john-breech | | 31 | Jason McIntyre | 15.1 | 62.79% | 98 | 39 | https://missedtakes.co/pundits/jason-mcintyre | | 32 | JP Acosta | 14.49 | 57.4% | 140 | 18 | https://missedtakes.co/pundits/jp-acosta | | 33 | Nick Shook | 12.33 | 67.75% | 70 | 15 | https://missedtakes.co/pundits/nick-shook | | 34 | Kyle Brandt | 11.71 | 53.9% | 105 | 30 | https://missedtakes.co/pundits/kyle-brandt | | 35 | John Middlekauff | 10.32 | 62.12% | 106 | 34 | https://missedtakes.co/pundits/john-middlekauff | | 36 | David Harrison | 9.74 | 63.84% | 236 | 25 | https://missedtakes.co/pundits/david-harrison | | 37 | Colin Cowherd | 9.39 | 54.88% | 415 | 69 | https://missedtakes.co/pundits/colin-cowherd | | 38 | Pete Prisco | 9.34 | 55.93% | 83 | 17 | https://missedtakes.co/pundits/pete-prisco | | 39 | Chad Ochocinco Johnson | 5.55 | 43.27% | 32 | 20 | https://missedtakes.co/pundits/chad-ochocinco-johnson | | 40 | Manti Te'o | 5.22 | 35.09% | 65 | 21 | https://missedtakes.co/pundits/manti-te-o | | 41 | Jeff Kerr | 1.88 | 75.58% | 183 | 19 | https://missedtakes.co/pundits/jeff-kerr | | 42 | Jonas Knox | 0.9 | 62.95% | 58 | 11 | https://missedtakes.co/pundits/jonas-knox | | 43 | Nathan Jahnke | -1.69 | 73.74% | 62 | 11 | https://missedtakes.co/pundits/nathan-jahnke | | 44 | Jared Dubin | -1.87 | 46.55% | 62 | 29 | https://missedtakes.co/pundits/jared-dubin | | 45 | Danny Parkins | -2.55 | 57.92% | 97 | 21 | https://missedtakes.co/pundits/danny-parkins | | 46 | Mitchell Eisenstein | -4.56 | 51.26% | 57 | 12 | https://missedtakes.co/pundits/mitchell-eisenstein | | 47 | Joe Pisapia | -5.61 | 73.61% | 20 | 12 | https://missedtakes.co/pundits/joe-pisapia | | 48 | Carter Bahns | -9.5 | 45.34% | 23 | 12 | https://missedtakes.co/pundits/carter-bahns | | 49 | Dan Patrick | -24.68 | 52.11% | 26 | 12 | https://missedtakes.co/pundits/dan-patrick | | 50 | Mike Moraitis | -27.66 | 58.33% | 35 | 12 | https://missedtakes.co/pundits/mike-moraitis | ## Prediction types with graded content - Award Picks (NFL): 329 graded, 163 pundits — https://missedtakes.co/predictions/player-award (the calls: https://missedtakes.co/nfl/player-award; as markdown: https://missedtakes.co/predictions/player-award/md) - Coach Award Picks (NFL): 67 graded, 59 pundits — https://missedtakes.co/predictions/coach-award (the calls: https://missedtakes.co/nfl/coach-award; as markdown: https://missedtakes.co/predictions/coach-award/md) - Combine Measurements (NFL): 41 graded, 32 pundits — https://missedtakes.co/predictions/combine-measurement (the calls: https://missedtakes.co/nfl/combine-measurement; as markdown: https://missedtakes.co/predictions/combine-measurement/md) - Combine Superlatives (NFL): 4 graded, 6 pundits — https://missedtakes.co/predictions/combine-superlative (the calls: https://missedtakes.co/nfl/combine-superlative; as markdown: https://missedtakes.co/predictions/combine-superlative/md) - Conference Champions (NFL): 194 graded, 112 pundits — https://missedtakes.co/predictions/conference-champion (the calls: https://missedtakes.co/nfl/conference-champion; as markdown: https://missedtakes.co/predictions/conference-champion/md) - Control of Congress (US Elections): 170 graded, 97 pundits — https://missedtakes.co/predictions/chamber-control (as markdown: https://missedtakes.co/predictions/chamber-control/md) - Depth Chart Calls (NFL): 583 graded, 455 pundits — https://missedtakes.co/predictions/depth-chart (the calls: https://missedtakes.co/nfl/depth-chart; as markdown: https://missedtakes.co/predictions/depth-chart/md) - Division Winners (NFL): 364 graded, 142 pundits — https://missedtakes.co/predictions/division-winner (the calls: https://missedtakes.co/nfl/division-winner; as markdown: https://missedtakes.co/predictions/division-winner/md) - Draft Bust Calls (NFL): 2 graded, 21 pundits — https://missedtakes.co/predictions/draft-bust (the calls: https://missedtakes.co/nfl/draft-bust; as markdown: https://missedtakes.co/predictions/draft-bust/md) - Fantasy Bust Calls (NFL): 38 graded, 67 pundits — https://missedtakes.co/predictions/fantasy-bust (the calls: https://missedtakes.co/nfl/fantasy-bust; as markdown: https://missedtakes.co/predictions/fantasy-bust/md) - Fantasy Defense Calls (NFL): 3 graded, 8 pundits — https://missedtakes.co/predictions/fantasy-defense-finish (the calls: https://missedtakes.co/nfl/fantasy-defense-finish; as markdown: https://missedtakes.co/predictions/fantasy-defense-finish/md) - Fantasy Defense Rankings (NFL): 4 graded, 2 pundits — https://missedtakes.co/predictions/fantasy-defense-ranking (the calls: https://missedtakes.co/nfl/fantasy-defense-ranking; as markdown: https://missedtakes.co/predictions/fantasy-defense-ranking/md) - Fantasy Rankings (NFL): 58 graded, 45 pundits — https://missedtakes.co/predictions/fantasy-draft-ranking (the calls: https://missedtakes.co/nfl/fantasy-draft-ranking; as markdown: https://missedtakes.co/predictions/fantasy-draft-ranking/md) - Game Outcome Picks (NFL): 1961 graded, 278 pundits — https://missedtakes.co/predictions/game-outcome (the calls: https://missedtakes.co/nfl/game-outcome; as markdown: https://missedtakes.co/predictions/game-outcome/md) - Head Coach Rankings (NFL): 7 graded, 28 pundits — https://missedtakes.co/predictions/head-coach-rankings (the calls: https://missedtakes.co/nfl/head-coach-rankings; as markdown: https://missedtakes.co/predictions/head-coach-rankings/md) - House Popular Vote Margins (US Elections): 4 graded, 6 pundits — https://missedtakes.co/predictions/house-popular-vote (as markdown: https://missedtakes.co/predictions/house-popular-vote/md) - Mock Draft Picks (NFL): 2987 graded, 220 pundits — https://missedtakes.co/predictions/nfl-draft-position (the calls: https://missedtakes.co/nfl/nfl-draft-position; as markdown: https://missedtakes.co/predictions/nfl-draft-position/md) - Player Season Numbers (NFL): 144 graded, 154 pundits — https://missedtakes.co/predictions/player-season-stat (the calls: https://missedtakes.co/nfl/player-season-stat; as markdown: https://missedtakes.co/predictions/player-season-stat/md) - Playoff Picks (NFL): 496 graded, 133 pundits — https://missedtakes.co/predictions/playoff-berth (the calls: https://missedtakes.co/nfl/playoff-berth; as markdown: https://missedtakes.co/predictions/playoff-berth/md) - Position Group Rankings (NFL): 16 graded, 47 pundits — https://missedtakes.co/predictions/position-group-rankings (the calls: https://missedtakes.co/nfl/position-group-rankings; as markdown: https://missedtakes.co/predictions/position-group-rankings/md) - Position Order Calls (NFL): 187 graded, 87 pundits — https://missedtakes.co/predictions/draft-positional-order (the calls: https://missedtakes.co/nfl/draft-positional-order; as markdown: https://missedtakes.co/predictions/draft-positional-order/md) - Positional Finishes (NFL): 228 graded, 217 pundits — https://missedtakes.co/predictions/player-positional-finish (the calls: https://missedtakes.co/nfl/player-positional-finish; as markdown: https://missedtakes.co/predictions/player-positional-finish/md) - Power Rankings (NFL): 174 graded, 109 pundits — https://missedtakes.co/predictions/power-rankings (the calls: https://missedtakes.co/nfl/power-rankings; as markdown: https://missedtakes.co/predictions/power-rankings/md) - Pro Bowl and All-Pro Calls (NFL): 50 graded, 80 pundits — https://missedtakes.co/predictions/player-accolade (the calls: https://missedtakes.co/nfl/player-accolade; as markdown: https://missedtakes.co/predictions/player-accolade/md) - Prospect vs Prospect Calls (NFL): 16 graded, 19 pundits — https://missedtakes.co/predictions/draft-versus-player (the calls: https://missedtakes.co/nfl/draft-versus-player; as markdown: https://missedtakes.co/predictions/draft-versus-player/md) - Race Margins (US Elections): 90 graded, 46 pundits — https://missedtakes.co/predictions/race-margin (as markdown: https://missedtakes.co/predictions/race-margin/md) - Race Winners (US Elections): 1221 graded, 168 pundits — https://missedtakes.co/predictions/race-winner (as markdown: https://missedtakes.co/predictions/race-winner/md) - Roster Spot Calls (NFL): 1619 graded, 182 pundits — https://missedtakes.co/predictions/roster-spot (the calls: https://missedtakes.co/nfl/roster-spot; as markdown: https://missedtakes.co/predictions/roster-spot/md) - Score Predictions (NFL): 809 graded, 169 pundits — https://missedtakes.co/predictions/game-score (the calls: https://missedtakes.co/nfl/game-score; as markdown: https://missedtakes.co/predictions/game-score/md) - Season Win Totals (NFL): 298 graded, 161 pundits — https://missedtakes.co/predictions/season-win-total (the calls: https://missedtakes.co/nfl/season-win-total; as markdown: https://missedtakes.co/predictions/season-win-total/md) - Seat Counts (US Elections): 129 graded, 54 pundits — https://missedtakes.co/predictions/chamber-seats (as markdown: https://missedtakes.co/predictions/chamber-seats/md) - Seats Most Likely to Flip (US Elections): 1 graded, 2 pundits — https://missedtakes.co/predictions/ranked-flips (as markdown: https://missedtakes.co/predictions/ranked-flips/md) - Seeding, Byes and Home Field (NFL): 53 graded, 45 pundits — https://missedtakes.co/predictions/playoff-seeding (the calls: https://missedtakes.co/nfl/playoff-seeding; as markdown: https://missedtakes.co/predictions/playoff-seeding/md) - Sleeper Calls (NFL): 178 graded, 183 pundits — https://missedtakes.co/predictions/fantasy-sleeper (the calls: https://missedtakes.co/nfl/fantasy-sleeper; as markdown: https://missedtakes.co/predictions/fantasy-sleeper/md) - Super Bowl Picks (NFL): 137 graded, 124 pundits — https://missedtakes.co/predictions/super-bowl-pick (the calls: https://missedtakes.co/nfl/super-bowl-pick; as markdown: https://missedtakes.co/predictions/super-bowl-pick/md) - Team Draft Slot Calls (NFL): 47 graded, 74 pundits — https://missedtakes.co/predictions/team-draft-slot (the calls: https://missedtakes.co/nfl/team-draft-slot; as markdown: https://missedtakes.co/predictions/team-draft-slot/md) - Top Player Lists (NFL): 82 graded, 144 pundits — https://missedtakes.co/predictions/top-players-list (the calls: https://missedtakes.co/nfl/top-players-list; as markdown: https://missedtakes.co/predictions/top-players-list/md) - Trade and Signing Calls (NFL): 214 graded, 268 pundits — https://missedtakes.co/predictions/player-movement (the calls: https://missedtakes.co/nfl/player-movement; as markdown: https://missedtakes.co/predictions/player-movement/md) ## The methodology, in full # MissedTakes scoring methodology **Status: APPROVED by Ben, 2026-07-31.** This is the build contract for the grading engine (E6) and the source text for the public `/methodology` page. Changing the math after predictions have been published means regrading every one of them, so amendments go through Ben, not through a code review. --- ## 1. What we are actually measuring Two numbers, published side by side, answering different questions. **Accuracy** — how often a pundit is right. Plain, honest, instantly understood. It is also, on its own, a bad measure of skill. **MissedTakes Score** — how much better a pundit does than a naive strategy. This is the number that separates insight from going along with the crowd. The reason we need both: some predictions are nearly free. Home teams win more often than they lose, and last season's best teams are usually good again. A pundit who mechanically picks the home team, or who reprints last year's standings as this year's rankings, will post a respectable accuracy while demonstrating no knowledge whatsoever. A pundit making genuinely sharp contrarian calls may score lower and be far more useful to listen to. Publishing accuracy alone would rank the second pundit below the first, which is precisely the misleading conclusion we exist to correct. So: **accuracy is the headline, skill is the story.** Both appear on every pundit page, always with the sample size next to them. ### The baseline For each prediction type we define a _naive baseline_ — what a person with no knowledge beyond public information would predict. | Prediction type | Baseline | | ----------------- | --------------------------------------------------------------- | | Game outcome | Pick the home team | | Score prediction | Home team by the league-average home margin for the season | | Power rankings | Last season's final standings | | Season win totals | The team's win count last season, regressed toward 8.5 | | Playoff berth | The teams that made the playoffs last season | | Division winner | Last season's winner of that division (repeats 43% of the time) | | Conference champ | Last season's conference champions (repeats 23%) | The MissedTakes Score is then, per prediction, how much better or worse the pundit did than the baseline did on that same prediction, normalized so that **0 means "no better than naive" and positive means genuine skill.** A pundit who only ever repeats the baseline lands at approximately zero however high their raw accuracy climbs. That is the correct and interesting result. **Decided (Ben, 2026-07-31; amended 2026-10-09): betting lines are never part of a grade.** A line is the sharpest public measure of "what everyone already thought", but it is a price, not a prediction anyone made. No point spread, total or moneyline enters an accuracy score, a skill delta, a baseline or the MissedTakes Score. The baselines above use only public results from prior seasons. Since Amendment 26, an NFL game's closing line is shown beside the picks made on it, as context only. This makes the baseline weaker, and we should be honest about the consequence: it is easier to beat than a betting line, so MissedTakes Scores will skew slightly positive across the board. Since 1999 the betting favorite has won 66.6% of regular-season games and the home team 56.2%. That is acceptable — the number's job is to separate pundits from each other and to expose pundits who only ever pick the home team, and a last-season baseline does both. It is not a claim to have beaten the market; the line beside each game take is there so a reader can judge that part for themselves. --- ## 2. Grading by prediction type Every grader returns a score in **0–100** for accuracy, plus a **skill delta** versus the baseline. Both are stored per prediction so they can be recomputed if this methodology changes. ### 2.1 Game outcome picks The simplest case, and the only binary one. - Correct winner → **100**. Wrong → **0**. - Ties: if the game ends in a tie, every pick is graded **50**. Rare enough not to distort, common enough in the NFL that ignoring it would be wrong. **Skill delta:** measured against the baseline, which is the home team (§1 keeps betting lines out of every grade). Correctly picking a road team earns more than correctly picking the home team, and a pundit who only ever picks the home side converges on a skill score of zero no matter how high their raw accuracy climbs. That is the whole point of the second number. Once we have enough graded seasons, this baseline can be sharpened using our own data — last season's standings are a better predictor than home advantage alone — without ever touching a betting line. #### 2.1.1 Which game, when the two teams met twice **Amended by Ben, 2026-08-16.** A pick names two clubs, not a fixture. Division rivals play each other twice, so for those games the claim on its own does not say which meeting it is about — and the grader used to decline all of them. 581 picks sat unscored on that rule. The publication date settles it. **The fixture is the first meeting of those two teams to kick off after the piece was published.** A prediction is about a game that has not been played yet, and in the archive these are overwhelmingly in-week picks published days before kickoff. Two cases still decline, because the date cannot answer them: - the claim was published after every meeting — that is a review, not a prediction, and guessing which past game they meant would be inventing one; - the fixtures have no stored kickoff time. #### 2.1.2 Picks on a game that was never played **Approved by Ben, 2026-08-16.** Some picks name a matchup that never happened: "the number one seeded Chiefs hosting the Cincinnati Bengals, and I've got the Chiefs advancing", said before the bracket existed. These were declined as unmeasurable. They are not. The claim has three parts, and two of them are checkable: 1. the team they picked got there; 2. that team won; 3. the opponent they named was the right one — **false by definition** for every pick in this case, since the game was never played. **One third each.** The ceiling is therefore **66.7**: naming a fixture that never existed cannot be fully right, but a pundit whose pick went on to win in January is not as wrong as one who named two teams that were home in December. Ben's worked example (2026-08-16): a pundit picks his team to beat the Rams in the Super Bowl. The Rams never get there; his team does, and wins it. Made it, won, wrong opponent — **66.7**. "Got there" means the postseason, not a particular round: the claim rarely names a round, and inferring one from the publication date would stack a second guess on the first. Whether the named opponent reached the postseason is recorded in the workings but is not a component — the third it would belong to is about the game they described, and nobody played it. A predicted scoreline has no actual to be measured against here, and the workings say so rather than scoring it zero. **Skill delta:** against the same three-part claim made at random — the share of the league that reached the postseason, plus the share that won a game there, each over the same thirds. Both numbers come from the season being graded, for the same reason the home-margin baseline does. ### 2.2 Score predictions Graded on two axes, because a score prediction contains two claims. **Winner (50% weight):** as above. **Margin (50% weight):** how close the predicted margin came to the actual. Credit decays with the error, reaching zero at 21 points — three scores, the point at which a prediction has stopped being informative. ``` margin_score = max(0, 100 × (1 − |predicted_margin − actual_margin| / 21)) ``` A predicted 24–20 (margin 4) against an actual 27–17 (margin 10) scores `100 × (1 − 6/21) ≈ 71` on margin, and 100 on winner, for **86 overall**. That feels right: they called it correctly and were reasonably close. **Decided (Ben, 2026-07-31): margin only, not raw scores.** Predicting 24–20 when the game finishes 31–27 is an excellent read — same margin, right winner — and should not be penalized for missing the absolute totals. A pundit who nails both is captured anyway, since nailing the raw scores necessarily nails the margin. ### 2.3 Power rankings The subtlest one, because a ranking is a whole ordering rather than a single claim. Graded on **rank correlation** between the pundit's ordering and the end-of-season standings — how well the shape of their ordering matches reality, rather than how many teams they placed exactly right. Nobody places 32 teams exactly; that is not the skill being tested. Concretely, Spearman's rank correlation, rescaled from its natural −1…+1 range to 0–100, so 50 means "no better than random ordering" and 100 means perfect. Four practical rules: - **Partial lists.** Plenty of pundits publish a top 10 rather than all 32. We grade the teams they actually ranked, against those teams' relative finishing order. A correct top 10 is a real achievement and gets full credit for what it claimed. - **Ties in the standings** are resolved by the league's own tiebreakers, since that is what the standings page shows. - **Ties in the pundit's ordering** — two teams both ranked 5th, three prospects all "at number 8" — claim no order between the tied items, so they share the average of the places they span (two teams level at 5th each take 5.5, between 5th and 6th) and the correlation is computed over those shared places: the textbook Spearman with ties (Amendment 13, Ben 2026-10-02). They still count as ahead of everything ranked below them. An ordering whose items all share one place claims no order at all and gets no score. A list that names the same team twice ties on the other side, since one team has one finish, and is treated the same way. An ordering with no ties scores exactly as it always did. - **Mid-season rankings** are graded against the final standings, but tagged with the week they were published. A Week 2 ranking and a Week 16 ranking are not comparable achievements, and pundit pages should let you filter. **Decided (Ben, 2026-07-31): earlier predictions are worth more.** Calling a team's season in July is genuinely harder than describing in December what has already largely happened, and without this the optimal strategy is to publish nothing until the answer is obvious. Every prediction carries a **difficulty multiplier** based on how much of the season remained when it was published: ``` difficulty = 1 + (weeks_remaining / total_weeks) ``` A preseason call (18 of 18 weeks left) is worth 2.0×; a Week 9 call about 1.5×; a Week 17 call about 1.06×. The multiplier scales the _skill_ contribution, not the raw accuracy — accuracy stays a plain honest percentage of how often someone was right, and the MissedTakes Score is where boldness and timing are rewarded. This applies to every season-long prediction type, not just rankings: win totals, playoff berths, and Super Bowl picks made in July should all outrank the same call made in December. Ben's framing: a perfect-season team being tipped for the Super Bowl in Week 17 is not an impressive call. ### 2.4 Season win totals - Graded on distance from the team's actual win count, with credit reaching zero at **5 games** of error. ``` score = max(0, 100 × (1 − |predicted_wins − actual_wins| / 5)) ``` Predicting 11 wins for a team that finishes 10–7 scores 80. Predicting 11 for a team that finishes 4–13 scores 0. Five games is deliberately tight — win total predictions are made with full offseason information and a full season to be right in. **A range of wins is a window — added by amendment, Ben, 2026-09-18.** "It's either thirteen and four or twelve and five" is a claim of 12 to 13 wins, not of 12.5 and not of nothing; it was being filed as "finish with ? wins". Records spoken or written, "ten or eleven wins", "between 9 and 11", "double-digit wins" (10 to 17) and a floor ("at least 10 wins", 10 to 17) are all read as windows, scored exactly as §2.7 scores a stat range: right anywhere inside, and outside it the miss is measured from the nearest edge on the same five-game curve. The baseline is unchanged — last season's wins regressed toward 8.5, scored as a point against the actual — so a wide hedge earns its ease at the clarity score, not here. A negated count ("under ten wins") stays skipped. ### 2.5 Playoff and Super Bowl picks **Added to the POC (Ben, 2026-07-31 — see §4 for why this one and nothing else).** Two closely related claims, both binary and both published constantly during the preseason we are about to ingest. - **Playoff berth.** "These teams make the playoffs." Graded per team named: 100 if that team made it, 0 if not. A pundit naming 14 teams gets 14 graded predictions, which is a healthy sample from one article. - **Super Bowl pick.** 100 for the winner, **50 for picking a losing finalist** — reaching the Super Bowl is most of the call, and grading it as a total miss would be unfair to a genuinely good read. Both take the §2.3 difficulty multiplier, which matters more here than anywhere: picking the eventual champion in July is a real call, and picking them in January is barely a prediction at all. #### 2.5.1 Division winners **Added by amendment, Ben, 2026-08-16.** The single most common team claim in the corpus — "I'm gonna go with the Chargers, I think they win the division" — and it was stored for years without being scored, because the type had never been approved rather than because anything was missing. **Binary: 100 if that club won its division, 0 if not.** No credit for second. Unlike a Super Bowl pick, where reaching the game is most of the call, this claim makes one assertion and the standings answer it. **How we know who won.** Off the bracket, not off the standings. Deciding it by best record would need the league's tiebreakers, which we do not hold and would have to guess at; the bracket already encodes the answer, because a division winner is seeded above every wild card and therefore either has a bye or hosts its wild-card game. Checked against all thirteen seasons we hold and both playoff formats: exactly eight clubs, one per division, every year. A season whose postseason is incomplete grades none of these rather than guessing. **Skill delta:** against last season's winner of the same division. That is the chalk pick, and it is precisely what a pundit calling a change at the top should be rewarded for beating. #### 2.5.2 Conference champions **Added by amendment, Ben, 2026-08-16.** "They're my pick to win the AFC." **100 for reaching the Super Bowl, 50 for losing the conference championship game.** This deliberately mirrors the Super Bowl rule one rung up: the shape of the claim is the same, and a pundit whose team lost the conference final was closer to right than one whose team was home in December. **Skill delta:** against last season's conference champions. #### 2.5.3 Seeding, byes and home field **Added by amendment, Ben, 2026-08-16.** "Houston bullies its way to the No. 1 seed in the AFC." Graded on the **tier** the claimed seed implies, not on the seed itself: | Claimed seed | Checked as | | ------------ | ------------------------ | | 1 | had the bye | | 2–4 | hosted a wild-card game | | 5–7 | visited a wild-card game | 100 if that is where they finished, 0 if not. This is deliberately coarser than the claim, and the reason is the same one that decides division winners: the bracket tells us who had the bye and who hosted, and separating the 2 seed from the 3 needs tiebreakers we do not hold. The alternative is inventing a seed order and telling a named analyst they were wrong on the strength of it. The stored workings say exactly what was checked — "you said 3, which is a wild-card host, and they hosted" — so nobody has to guess how coarse the answer is. Two thirds of these claims are about the number one seed, which the bracket names exactly, so the coarseness costs less than it looks. Seasons before 2020 gave two byes per conference, which makes seeds 1 and 2 indistinguishable. Those decline rather than guess. **Skill delta:** against where that club finished last season. ### 2.6 Draft position (mock drafts) **Added by amendment, Ben, 2026-08-01, and revised the same day after we looked at what the first version produced.** A mock draft is graded **twice**, because two different questions are worth answering and one number cannot answer both. **Accuracy — how close was each pick.** Per pick, on its own. Credit **halves every sixteen picks** you are out: ``` accuracy = 100 × 0.5 ^ (|predicted pick − actual pick| / 16) ``` This is the number a reader wants: _you had him at 12, he went 23._ It is what appears next to each pick. **Amended by Ben, 2026-08-02**, replacing a straight line that reached zero at one full round. The line put 29% of all graded picks (102 of 355) on exactly zero and scored a pick 32 places out identically to one 128 places out — two predictions that are not remotely equally wrong. The curve was chosen over simply widening the line to two rounds because **it agrees with the old rule where the old rule was trusted** — 100 for an exact hit, 50 at sixteen picks out — and differs only in the tail. Widening instead would have moved a twenty-pick miss from 37.5 to 68.8, loosening the standard for ordinary misses in order to fix a problem that only exists at the extremes. Measured against the archive before adopting it: the Round 1 average moves 74.2 → 73.1, scores landing on zero fall from 102 to 4, and **the ranking of pundits is unchanged** — so this changes the published number rather than who is ahead. A score now approaches zero without arriving: 128 picks out is 0.4, because it is a worse answer than 32 out and the number should say so. **Skill — did they get the shape of the draft right.** Scored across the whole mock with the same rank correlation as power rankings (§2.3), against a naive strategy that **shuffled the names**, which scores 50. Every pick in one mock therefore carries the same skill figure, because getting a draft's order right is one achievement, not thirty-two of them. That is also consistent with §3: the article is the unit of evidence about a pundit. **A player who was mocked and then went undrafted scores zero on that pick.** **Amended by Ben, 2026-08-03**, replacing a rule that did not grade these at all. The old rule said an undrafted player "has no position in the real sequence to be measured against", which is true, and used it to reach a conclusion that was not: that the prediction could not be assessed. The prediction was that a named person would be picked in a named range. He was not picked at all. That is the most straightforwardly wrong a draft call can be, and 122 of them in the archive were filed as **our** data gap rather than counted against anyone. Zero rather than a distance measured from the pick after the last one. Giving him an imaginary pick number would invent a fact, and it would also rank a late mock of an undrafted player above an early one when both were wrong about the only thing they claimed. He is still **left out of the ordering**, for the reason the old rule gave: a rank correlation compares two sequences, and a man who was never picked has no place in one of them. The penalty for missing on him lands in the accuracy figure, where the reader can see it next to his name. A mock with only one gradable pick gets an accuracy figure and **no** skill figure — there is no order to have got right, and publishing a skill number for it would assert something we did not measure. The same goes for a mock whose gradable picks all sit at **one pick number**: "Downs, Bain or Delane at number 8" names three candidates for one slot, not an order among them. Where tied picks sit inside a longer mock they share the average of the places they span, as in §2.3 — after every pick numbered above them, before every pick numbered below — and a player named at two pick numbers is treated the same way on the drafted side (Amendment 13). **Which draft settles a claim.** "He went undrafted" is only a fact once the draft is over, and of a player the NFL does not list, only while its record is whole: he may be a pick it is missing, so the claim waits (Amendment 25). So this rule needs to know which draft the article was about, and the answer is taken from the strongest evidence available: 1. **The year the source names**, where it names one — "a top-50 pick in the 2027 NFL Draft". 2. **The draft the article's other picks were taken in.** A mock's names all come from one class, so if most of the players in a piece were taken in the 2022 draft, the piece is about the 2022 draft. 3. **Failing those, the publication date** — a piece published from May onward is about next spring's draft, one published before then is about the draft that year. 4. **Failing all three, it is not graded.** Where a player was never picked and nothing dates the article, we say so rather than guess. The order matters and the date rule alone was not safe. A publication date reads a retrospective backwards — a February 2024 piece about the 2023 draft would be checked against a 2024 draft none of its subjects were in — and under this amendment that would have turned calls we score at 100 into zeros. ### 2.6.1 The player who declared a year late **Added by amendment, Ben, 2026-08-03.** A prospect can be mocked into a draft he then does not enter. Ben's case: a pundit calls Arch Manning a top-10 pick in the 2026 draft, Manning stays in college, and goes top-10 in 2027. The talent read was right and the timing was wrong, and neither "correct" nor "zero" describes that. **The pick keeps half its closeness for each year the pundit was early.** ``` accuracy = pick accuracy × 0.5 ^ (years early) ``` A top-10 call on a player who goes fifth overall the following spring keeps half of a near-perfect score. Halving is the same shape §2.6 already uses for pick distance rather than a second idea invented for the same page. **0.5 is not Ben's number.** He approved "partially correct"; the weight is ours, and it is the figure to revisit once enough of these exist to look at. Two conditions, both because this rule reaches a conclusion about what somebody _meant_: - It applies only where the intended draft is known from the source's own words or from **at least two other picks** in the same article. A piece we can date only by its publication date is graded as though the year were right. - The player being graded is never evidence for which draft his own article was about. He would date it to the draft he entered, and the question would answer itself every time. **A player drafted EARLIER than the article's draft is not graded at all.** He was already in the league when the piece was written, so no prediction was being made about him and the resolver has found the wrong man. Recorded as ours to fix. This is not hypothetical: a 2025 mock reading "129 Baltimore Ravens LB Chris Paul Jr. Ole Miss" was linked to the Chris Paul drafted in 2022, and we published a score for it. Where the player entered the draft the article was about, none of this applies: nobody has ever been drafted twice, so his pick is a fact about him and the claim is scored against it whatever month the article ran. A consequence worth stating plainly: a call on a player who has not yet declared can score zero and later become a half-credit, because the fact that settles it — which draft he eventually entered — arrives after we first graded it. That is the archive being regraded as reality arrives, which §2.6 already requires, rather than a score being quietly rewritten. **Why not a per-pick baseline.** A per-pick skill score needs a naive guess for one named player, and every candidate is either third-party — consensus big boards and mock aggregates, already rejected for fantasy sleepers in §4 — or useless. The first version of this section used the midpoint of the draft (pick 129). That scores **zero** against every pick inside the first two rounds, which is where essentially all published mocks live, so `skill = accuracy − 0 = accuracy`: the MissedTakes Score silently became a copy of the accuracy figure, and mock drafters outranked every other kind of pundit on any mixed leaderboard. Recorded here because the rejected version is part of why the current one is shaped as it is. Draft claims do **not** take the §2.3 difficulty multiplier. It asks how much season was left to be wrong in, and a draft is a single event rather than a season; these grade at 1.0. ### 2.6.2 Where a club picks **Added by amendment, Ben, 2026-08-03.** Separate from §2.6, which is about where a PLAYER goes. This is about where a TEAM ends up picking — "the Patriots are the favorite to earn the No. 1 pick", "Jacksonville could pick at No. 1 for the third time in five seasons". Ben's question when this came up was whether it is really the inverse of the end-of-season standings, and whether the two should be linked. It is, and they are linked in what they measure — a club picks first by losing most — but they are graded apart, and the reason is worth writing down because the tempting shortcut is wrong. **Graded against the pick the club actually made**, read from the draft results, not against a slot derived from the final table. Deriving it looks free and is not. Measured across seven drafts, ordering the non-playoff teams by record put the right club in the right slot **2 to 9 times out of 18**. Two causes: - **Trades.** Between none and four of the top eighteen picks change hands in a given year, and no amount of standings arithmetic can see that. - **Tie-breaks.** The league separates clubs on identical records by strength of schedule, which we do not store. Among clubs that kept their own slot, the derived position was off by **0.89 to 3.06 places on average**, worst case fifteen. We already hold who made every pick. An approximation of a fact we hold exactly is not worth having. **Scored like a mock pick** — the §2.6 curve, halving every sixteen picks, and a named window scored as a window. A claim of "top ten" is right anywhere inside it. **Baseline: where the club picked in the previous draft.** The same idea as power rankings reprinting last season's table, and it means a pundit who predicts that a bad team stays bad earns nothing for it. Where we do not hold the previous draft there is no baseline, so the call gets an accuracy figure and **no** skill figure rather than one measured against an invented guess. **This takes the §2.3 difficulty multiplier, and mock drafts do not.** That is the whole point of keeping the two apart. A mock draft is settled by a single evening in April, so there is no "how much season was left" to ask about. A club's draft slot is settled by eighteen weeks of football: calling the Patriots for the first pick in September is a real prediction, and saying it in December is reading the standings out loud. **A known limitation.** "Earn the No. 1 pick" is strictly a claim about finishing worst, while we grade what the club actually picked. Those disagree in the years a team earns a slot and trades the pick away. We take the second reading because it is the one we can settle exactly, and because the wording pundits use — "pick at", "land", "picking inside the top ten" — reads that way far more often than not. Revisit if a real case lands badly. ### 2.6.3 Order within a position **Added by amendment, Ben, 2026-08-03.** "The first quarterback off the board." "CB2 of the class." "The second cornerback drafted." These name no pick number at all, so §2.6 cannot read them — they were being stored as an unreadable phrase on a mock-draft row and sitting ungraded. They are perfectly checkable: we hold every pick with the position the league filed it under. **Graded on where he actually placed among his position**, with credit **halving every two places**: ``` accuracy = 100 × 0.5 ^ (|predicted place − actual place| / 2) ``` Tighter than §2.6's sixteen picks, and deliberately: about fourteen quarterbacks go in a draft against 257 players overall, so one place here is a real miss where one pick overall is nothing. Saying "first quarterback" when he was the second keeps 71; when he was the fifth, 25. **Two is a choice, not Ben's number**, the same way the §2.6.1 weight is. **A window is scored as a window**, as everywhere else — "the first or second quarterback off the board" is right either way. **Baseline: naming one of them at random.** If eleven quarterbacks were drafted, someone with no knowledge has eleven equally likely answers, and the baseline is what those average to against the truth. It falls out of the same curve the claim is graded on, so it needs no separate justification, and it correctly makes "the first quarterback" a far harder claim than "somewhere in the top eight quarterbacks" — most random guesses miss the first badly, and many land inside a wide window. One consequence worth stating: when the man really was taken first, a wide window is no easier to hit by luck than a precise one, because only a window starting at one catches him. The baselines agree in that case, and that is the honest answer rather than an oversight. **A player who went undrafted scores zero**, for the reason §2.6 gives: the claim was that he would be among the first at his position, and he was not among them at all. Where he was drafted but the league filed him under a position the claim does not cover, that is **ours** to reconcile and he is not graded. **Positions are read as families**, because the draft file's vocabulary is broad and inconsistent — it carries DE, DT, DL and NT for the same handful of jobs. "Defensive lineman" covers all four; "defender" covers the whole side of the ball. That mapping is in code and tested. **A spelling the grader's own words do not know is read through the canonical position list** (2026-10-01): "tackle" and "left tackle" are offensive tackles, "guard" and "right guard" are guards. Guard and center are separate families, so a guard's place is never counted behind centers. A combination like "OLB/DE" is read only when both halves name the same family; otherwise the claim is not graded. Every word the grader already knew reads exactly as before. ### 2.6.4 One prospect ahead of another **Added by amendment, Ben, 2026-08-03.** "Drafted higher than Malik Willis." No pick number for either man, so it is not a §2.6 claim, but both picks are facts and the claim is simply that one number is smaller than the other. **Right or wrong: 100 or 0.** Against a **baseline of 50**, because with no knowledge this is a coin flip — the same reasoning behind the ordinal baseline everywhere else. Being drafted at all beats not being drafted, so if one of the two goes undrafted the claim still settles. If **neither** was drafted, neither went ahead of the other and there is nothing to compare: not graded, and not the pundit's fault. ### 2.6.5 Awards **Added by amendment, Ben, 2026-08-03.** MVP, the Player and Rookie of the Year awards, Comeback Player, Pro Bowl and All-Pro selections — and Coach of the Year, which is the reason this section exists. **Right or wrong: 100 or 0.** There is no partial credit in naming a winner. **Baseline: what picking last season's winner would have scored.** The same prior-season-only philosophy as every other baseline here, and it earns its keep on exactly one case — the chalk pick. A pundit who re-picks the reigning MVP and is right scores 100 and gains **nothing**, because the naive strategy said the same thing. Any other year the baseline is zero, and that is the honest answer: correctly naming one winner from a wide field is hard, and a pundit who does it should be paid for it. This is not the failure §2.6 records for the rejected midpoint baseline. There, the baseline scored zero against **every** pick, so skill silently became a copy of accuracy on a type where accuracy averaged 74. Here accuracy is zero most of the time too, so a pundit only gains when they were actually right about something difficult. **Selections are graded differently from awards.** Many players make the Pro Bowl, so the question is only whether this one did, and the naive strategy is last season's list repeating — a perennial Pro Bowler being tipped again is not a call. Awards go to one person, so the naive strategy is the reigning winner. **We never publish "he did not win" when we mean "we never looked it up."** Awards have no free feed and are typed in by hand, so where we hold no results for a season the claim is **not graded** and the gap is recorded as ours. That is load-bearing: 146 award claims sat in the archive against an empty results table, and scoring them all zero would have invented 146 wrong predictions. **Coach of the Year is a coach award, not a player one.** Ben asked whether coaches should live in the player table; they should not, and this is the grading half of that answer. Coach awards resolve against a separate coach record and their own results file. See the `coaches` table for the storage half and why it is worth a table of its own. ### 2.6.6 Fantasy value calls — busts and sleepers **Added by amendment, Ben, 2026-08-03.** A bust and a sleeper are **one measurement facing opposite ways**: did the player finish better or worse than what he cost to draft? Both were being lost. `fantasy-sleeper` had captured 143 claims and scored none of them, because the type was never approved. Busts had no type at all, so "2026 Fantasy Football Busts" parsed to nothing and mixed sleepers-and-busts articles kept only half their content. **The yardstick comes from the source where the source names one.** Fantasy writers quote Average Draft Position constantly — WalterFootball prints "ADP: 1.05" beside every name — and that is the number the claim is measured against. We do not buy ADP: it is the crowd's opinion, and licensing it is the third-party consensus data §4 rejects. A pundit who quotes it in his own article has chosen what to be measured by, which is a different thing entirely. **Where no cost is named, the claim is not filed and not graded** (amended, Ben, 2026-09-18, reversing the amendment of 2026-08-16). For a month the fallback was the player's own prior-season finish: a yardstick we already held, prior-season-only like every other baseline here (§1), and §2.8's rule rather than a second idea. What it produced was a moderation queue full of "named a bust — the source gave no draft cost" rows that nobody could act on, and a grade measured against a bar the pundit never named. The consent is what makes the quoted number usable; where there is no number there is no consent, and a sleeper who is worth more than _nothing in particular_ is not a claim. Ben: "auto-mark these as too vague or suppress them from the mod queue." They are suppressed: the parser is told to skip them, ingest refuses them, the grader declines whatever was filed before, and a dated heal moves the filed ones to unresolvable. The pundit is not charged for them (no `claim_was_vague`): a season sleeper column with no ADP is a genre, not a shortcoming, and the site simply does not grade it. **A one-week play is never a sleeper or a bust** (same amendment). Every row in Ben's screenshots was a radio show's Week 2 "hot plays" and "cold sores" — "I've got him outside of my top 12 for this week", "going up against Cleveland" — filed as season value calls. A start/sit call is settled by one box score, which we do not grade against (§2.13), and its scope is usually on the opinion, not the claim: the show names the week once at the top. So the check reads the claim and its opinion's summary together, through the single-game guard §2.13 already uses plus the lineup vocabulary (start/sit, hot and cold plays, streamers, flex rankings, "up against" a named opponent). Declined for scope, not vagueness. **What a pick owes you depends on what it cost**, so the bar moves with the cost — as a curve, since 2026-08-16, having been a four-band step table before that: ``` bar = 5 × (startable tier / 5) ^ ((slot − 1) / 4), capped at the startable tier ``` where `slot` is the cost in rounds as a real number: pick 17 of a twelve-team league is slot 1.42, and a source that names a round but no pick is read at that round's midpoint. The slot is floored at 1. A single bar cannot work: RB20 is a disaster from the first round and a triumph from the tenth. But the step table could not work either, and the defect was the cliff rather than the levels: **a pick at 2.12 had to finish top 5 and the very next pick, 3.01, only had to finish top 12.** One slot of cost moved the bar seven places, and a published bust verdict on a named real person turned on which side of a round boundary his ADP happened to fall. The curve is anchored to **agree with the old table where the old table was trusted**, which is the principle §2.6 used when the draft-distance line became a curve: | Cost | Old table | Curve | | -------- | -------------- | ---------------------------------------------- | | Round 1 | top 5 | top 5 — the same | | Round 2 | top 5 | RB7 / WR8 / QB6 | | Round 3 | top 12 | RB11 / WR13 / QB8 | | Round 4 | top 12 | RB16 / WR22 / QB10 | | Round 5+ | startable tier | the startable tier — the same, by construction | It differs in the second half of each old band, where a staircase is at its strictest, so the curve is **mildly looser on the whole**. That direction has to be measured rather than argued: `pnpm db:measure:value-bars` prints the verdict flips, the type averages, and whether the pundit ordering moves. One genuine bug the curve fixes on its own: at quarterback and tight end the old table's round-3–4 value of 12 was **identical to their startable tier**, so a third-round quarterback and a twelfth-round quarterback faced exactly the same bar. For half the positions this section covers, its whole premise was silently not applying. **The shape of the curve is ours, not Ben's**, like the bands before it, and is the figure to revisit once enough calls are graded to look at. **Right or wrong: 100 or 0.** A bust either happened or it did not, and the magnitude is printed in the workings rather than folded into the score. **Baseline: the draft market was right.** The naive view is that a player returns roughly what he cost — neither bust nor sleeper — so it scores exactly the opposite of the pundit. This is the purest baseline on the site: the whole question the product asks is whether a pundit beats the crowd, and here the crowd's own number is printed in the article. It also means **a wrong call costs a pundit**, which a one-sided baseline would not. ### 2.6.7 Draft busts **Added by amendment, Ben, 2026-08-03, knowing it only half works.** "He will never justify going eighth overall." A claim about a career rather than a season, so it is graded against **the best a player ever managed** at his position, not one bad year — a first-rounder who goes RB40, RB38 and then RB6 is not a bust, and no single season says so. A player needs **three seasons** before the call can be made. Fewer than that and he is waiting, not failing — unless he has already cleared the bar, which more time cannot undo. **The bar is what the picks around his actually returned** (amended, Ben, 2026-08-16). Every player at his position taken within **sixteen picks** of his slot, across every draft we hold, and the bar is the **median** of what those players managed. This replaces the §2.6.6 curve read in NFL rounds, which was always the weaker half of this type: it is built for a twelve-team fantasy draft, where round 1 is the twelve best players in the sport. Applied to a 257-pick league draft it asked pick 1 and pick 32 alike to finish top 5 at their position, which is not what either pick cost. We hold every pick and every season, so the question can be asked directly rather than through an analogy. Where the pool is thinner than **five** comparable picks the cost curve is still used, and the workings say so. Ben's proposal was the player immediately ahead of and behind. Three changes to that shape, each for a reason worth stating: - **Many picks, not two.** Two neighbors is one torn ACL away from deciding somebody's record: a bust call would land or fail on whether the back taken one slot earlier got hurt in Week 1, which the pundit claimed nothing about. - **Across every draft, not within his own class.** A class rarely holds five comparable picks at one position, and "what does a pick here return" is a question about drafts in general. It also stops the bar being computed from the same seasons it is judging — every other baseline here is careful to use prior results only (§1). - **The median, not the average.** Specifically the lower median, so the bar is a finish a comparable pick actually achieved rather than a number nobody managed. **A comparable pick who never produced a measurable season counts in the pool.** A pick that returned nothing is the most informative comparison there is, and leaving those out would build the bar from only the picks that worked. Because a median orders its inputs and never averages them, how bad he was never enters the arithmetic — we assert only that he did worse than anyone who played, which is a fact rather than an invented pick number. Where the median itself lands on such a player, the bar reads "any measurable season", and anybody who produced one has cleared it. Peers from drafts too recent to judge are left out, on the same three-season rule the claim itself is held to. **Sixteen picks is DRAFT_PICK_HALF_LIFE**, reused rather than invented: it is already the distance this site calls "half as right", so it is a reasonable distance at which two picks stop being the same kind of pick. **Five is ours**, and both are figures to revisit once enough of these are graded to look at. `pnpm db:measure:value-bars` prints what the change does to the archive. **This type can only grade a third of what it is about, and that is stated here rather than discovered later.** Of the first-round picks we hold, 75 are skill positions and 149 are not. Every offensive lineman and defender produces nothing our statistics measure — 279 linebackers scored in 2025 and not one of them registered a fantasy point. For those players the claim is captured and **refused**, recorded as our gap. Ben's call in accepting that: _"Even if we can only score the skill positions for right now, we can work in the future to build methodology and pull in historic data for the other positions."_ Snap counts, starts, or the honors table in §2.6.5 are the routes to measuring a guard. Until one of them exists, this type is deliberately silent about two thirds of the players it is asked about, and silence is the honest answer rather than a guess. ### 2.7 Player season statistics **Added by amendment, Ben, 2026-08-01.** "He finishes with 1,400 receiving yards." Graded on **relative** error, because one absolute tolerance cannot serve both receiving yards and sacks: ``` scale = max(|actual|, |predicted|, 1) accuracy = 100 × (1 − (|predicted − actual| / scale) / 0.5) ``` Credit reaches zero at 50% out. Scaling by the larger of the two values keeps the score symmetric and does not divide by zero when a player is hurt in Week 1 and finishes on nothing. **Regular season only, never summed with the postseason.** A fantasy year ends before the playoffs, and everyone making these claims means the 17 games. Quietly adding January would inflate every actual and mark honest predictions wrong. `fantasy_points` means **standard scoring**, not PPR (Ben, 2026-07-31). **A range is scored as a window — added 2026-09-18.** "45-50 catches", "1,200 to 1,400 yards", "between 8 and 10 sacks" are real claims, and the parser had been filing them with no number at all (Ben, on a Vikings mailbag: "the pundit provided quality production ranges, but the parser wasn't able to pick them up"). The same rule §2.6 applies to "top-50": anywhere inside the range is full credit, and outside it the miss is measured from the nearest edge on the relative curve above, so "45-50" against 40 scores exactly what "45" against 40 would. A wide hedge is easier to land; that cost is the clarity score's to charge, not a reason to pretend one number was named. The range is stored in the pundit's words and read into numbers once, by the same function ingest, grading and the take page share. **Baseline: what he did last season, regressed halfway toward his position's average.** The same shape as the win-total baseline in §2.4, deliberately — this is the philosophy §1 already commits to rather than a second one invented for players. A player with **no prior season is a rookie**, and the positional average is the whole of what public information said about him. Positional averages are computed over players who actually posted a figure, since including every third-stringer's zero would describe an "average quarterback" that no real starter resembles. ### 2.8 Positional finishes and ranked player lists **Added by amendment, Ben, 2026-08-01.** - **Positional finish** — "he finishes as WR5." Graded on distance in places, credit reaching zero at **12 places**, roughly the startable tier these claims are actually made in. Baseline: he finishes where he finished last season; a player with no prior season is graded against the bottom of the startable tier, which is what an unknown at that position is worth to anyone drafting — **QB12, RB24, WR36, TE12**, the same startable tiers used for fantasy sleepers in §4. Positions outside that group fall back to 24, which is an acknowledged imprecision rather than a considered number. **A tier is a window of twelve** (amended, Ben, 2026-09-18). "An upside WR2", "a low-end RB1", "RB3/FLEX", "in the RB1 field": tiers are how fantasy writers actually commit, and each is a range of finishes — RB2 is RB13–24, WR3 is WR25–36 — read at twelve to a tier, the 12-team convention the startable tiers above already assume. Scored as every other window here is (§2.7's ranges, §2.6's draft windows): right anywhere inside, and outside it credit falls off from the nearest edge at the same 12-place rate. The baseline is the same prior-season finish, scored against the same window. A hedge keeps the tier ("fringe WR2" is WR2); "FLEX" alone crosses positions and ranks against nothing, and a boundary — "the QB1/QB2 line" — commits to no tier; both are the pundit's vagueness and are flagged as such. A tier beside "Week 1" or "this week" is a start/sit call (§2.13), not a season finish. **What the finish is measured in** (amended, Ben, 2026-08-05 and 2026-08-13). A finish that names its statistic — including fantasy points, for takes made in a fantasy context — is ranked in that statistic. A bare finish is ranked in its position's primary statistic: **passing yards** for quarterbacks, **rushing yards** for running backs and fullbacks (fullbacks added by Ben, 2026-09-18, in a pool of their own), **receiving yards** for receivers and tight ends, **interceptions** for the secondary (CB, S, DB), and **sacks** for the front seven (DE, DT, NT, LB). "The best running back in the league" is a claim about rushing yards, not fantasy points, unless the source was talking fantasy. Positions with no primary counting stat — offensive linemen, specialists — are skipped rather than guessed. A bare finish is therefore not a vague claim and is never held for a person to name its statistic; until 2026-09-18 ingest flagged every one as such, and a flagged claim is not graded, so the default had never actually applied. **What pool the rank is computed in** (amended, Ben, 2026-08-13). Offensive players rank within their own position — "finishes as TE1" competes with tight ends, never receivers. Defenders rank within their **family pool**: the down linemen together (DE, DT, NT), the linebackers together (LB, OLB, ILB, MLB), the secondary together (CB, S, DB). The stat feed's position codes fragment the same job — by 2025 nearly every linebacker is coded plain LB while ILB holds four players, and the safety codes drift year to year — so ranking within the raw code would grade "best linebacker" against a pool of seven. The line and the linebackers stay separate pools because a nose tackle's ceiling in sacks is not an edge rusher's; the known cost is that scheme coding puts some edge rushers in DE and some in LB, and those rank apart. The stored workings name the pool the rank was computed in. - **Ranked player lists** (fantasy draft rankings, top-100 lists) — graded with the same rank correlation as §2.3, **within position**. Ben, 2026-08-01: no statistic ranks a left tackle against a cornerback, so a mixed list is scored as its positional sub-orderings and the cross-position ordering is not scored at all. Each position's score is weighted by how many players it contributed, so a ten-deep receiver ordering counts for more than an incidental pair of tight ends. **Baseline for any ranked list: 50, not 0.** The rank-correlation score maps onto 0–100 as `((rho + 1) / 2) × 100`, so a coin-flip ordering already scores 50. A baseline of zero would credit a pundit for the half of the scale that random guessing earns for free. #### 2.8.1 Combine measurements **Amended by Ben, 2026-08-16**, replacing the fixed band per drill signed off on 2026-08-05 (40 times running linearly to zero at 0.30s, weight at 15lb, and so on). The grader itself is unchanged; what it measures against is not. **Credit halves for every standard deviation of that position's own spread.** ``` score = 100 × 0.5 ^ (|predicted − actual| / σ_position) ``` σ is the spread of that drill among players at that position, pooled across every combine year we hold. Being one σ out scores **50** — the point at which the position average was as good a guess as the pundit was. Two measured reasons for the change, both from the archive rather than from taste: 1. **A single band per drill is the wrong shape.** The spread inside a position is three to six times tighter than the spread across the class. σ on a 40 time is **0.08s for a cornerback and 0.21s for a lineman**; on weight it is **8lb for a cornerback against 47lb across all positions**. A 0.30s band made a corner's projection nearly impossible to get wrong and a lineman's punishing. 2. **A linear run to zero stops telling predictions apart.** This is the same fault the draft-pick curve had before 2026-08-02, and the same fix. Measured over the 42 claims we can currently match: the linear band put 5 of them on zero, decay puts none, and a 1.18σ miss scores 44 rather than being indistinguishable from a wild guess. **Skill delta:** the average athlete at that position, scored on the same curve. A projection earns skill only by beating "he is typical for his position". Measured over those 42 pairs, the median claim scores 74 against a median baseline of 51, and 32 of 42 beat the average. A position with fewer than **10** recorded results for a drill declines rather than being scored against a spread computed from a handful of athletes. Threshold claims ("runs a sub-4.5") and superlatives ("fastest in the class") are unaffected — they are not distance claims and keep their own rules. #### 2.8.2 Which awards we model, and which we refuse **Settled by Ben, 2026-08-16.** We hold the AP season awards, the Pro Bowl and All-Pro selections, the two coach awards, and — added 2026-08-16 — **Protector of the Year**, the league's award for the best offensive lineman. It was created for the 2025 season and first given in February 2026, so 2025 is its entire history and there is nothing to backfill. **The Super Bowl MVP** joined them on 2026-10-09 (Amendment 22). Awards attach to the season they were EARNED IN, not the calendar year they were handed out: an award presented in February 2026 belongs to the 2025 season, the same convention every other actual here uses. **The Hall of Fame is out of scope, and is refused at the parser rather than stored and declined.** Three reasons: - eligibility does not begin until five years after a player retires, so for anyone still playing the claim cannot be settled for a decade; - it is a verdict on a whole career decided by a voting committee, not a season outcome, which is the entire scope of what we grade; - it would need a class list entered every year, forever, for a handful of claims. They were being captured as Pro Bowl and All-Pro calls, which is why rejecting them by hand kept missing them: 57 claims across six different prediction types. One rule at the parser catches all six. **Finalists are not modeled.** Being a finalist is a fact about a voting process rather than about what happened, and it is only published for recent seasons and only for some awards — so crediting it would mean the same claim scoring differently depending on which year it was made in, which is not something a pundit's record should carry. **Only the award itself is graded — a fix, not a rule change (2026-10-09).** The honor a pundit named was read for the award's name anywhere in it, so a Super Bowl MVP pick was graded as the season's MVP, and "MVP finalist", "top five in MVP voting", a club's own MVP and "finish no lower than second in MVP voting" were all graded as winning the league award. Now the honor must be the award's name and nothing more, and the claim must be that he wins it or makes it: a game's MVP, a club's award, a finish or a ballot place in the voting, and a career count ("five or six Pro Bowls") are not graded. A Pro Bowl or All-Pro call is graded only against a list we hold for that season; having that season's MVP on file is not having its Pro Bowl roster. On the archive as it stood that day, 111 published grades became ungraded and three were re-checked to the score they already had. The same day Ben made the Super Bowl MVP an award we grade (Amendment 22), so those picks were graded again, against the Super Bowl's MVP. #### 2.8.3 Depth chart calls **Approved by Ben, 2026-08-02; the measure amended 2026-08-16.** This grader has been running since August and was never written down here, which is a gap in a document that is supposed to be the contract. "He's their third receiver", "he immediately steps in as the top defensive tackle on this roster." Graded on distance in places, the same curve positional finishes use. **What a depth chart is, for our purposes.** A real one is a coaching document we do not have. What we can observe is where a man ranked among his own club's players at his position: - **The four skill positions** rank by usage — passing yards, rushing yards, receptions. The third receiver takes the third-most catches far more often than not. - **Everyone else** — the offensive line, most of the defense — ranks by **snaps**, added 2026-08-16. These positions accrue no counting statistics, so before this they could not be placed at all and 125 claims were declined with "no usage to place a OT on a depth chart". A starting left tackle plays every down, and the snap count says so. Snaps are used only where usage does not exist. Ranking the same men by two rules would make the answer depend on which ran last, and it would move grades that are already settled. **Skill delta:** where he sat last season. A player with no prior season is a rookie, and the naive expectation for a rookie is that he is not the starter — slot 2, the first man behind whoever already has the job. ### 2.9 Head coach rankings A ranked list of head coaches, graded the same way as any other ordering: rank correlation against the order the coaches actually finished in. What settles it is **wins above expectation**, where the expectation is last season's win count regressed halfway toward 8.5 — the same baseline §2.4 uses for season win totals. A coach's score is his team's actual wins minus that figure, prorated by the games he personally coached. The expectation has to scale with the team, or the ranking rewards the wrong thing. Measured on year-over-year improvement instead, a coach whose team finished near the top can only hold or fall, while a coach of a bad team has everywhere to climb — and the way to top the leaderboard becomes ranking bad teams' coaches highest. Against a per-team bar, a team that won 12 last season carries an expectation of 10.25, so winning 12 again reads as excellence rather than as standing still. Three things this deliberately does not do: - **It does not use betting lines.** A market price would be the sharper expectation, and §1 keeps it out of every grade. - **It does not claim to measure coaching.** It measures whether a team beat what its previous season implied. That is a rough proxy. It knows nothing about who the team drafted, signed, lost or buried in the training room, so a coach handed a rebuilt roster will look better than he was and a coach hit by injuries worse. - **It does not rank a man on a handful of games.** A coach needs at least eight games in a season to appear at all. Below that he is ungraded, which is not the same as being ranked last. Baseline: last season's ordering of the same coaches. Naming the men who overperformed last year is precisely the chalk pick this site exists to tell apart from insight. ### 2.10 Position group rankings An ordering of all 32 clubs within one unit — offensive lines, receiving corps, secondaries. Graded by rank correlation, like every other ordering, against that unit's own production across the season. Thirteen units are ranked: quarterbacks, running backs, wide receivers, tight ends, offensive line, defensive line, linebackers, cornerbacks, safeties, kicking, punting and returns, and — since 2026-09-18 — the whole defense and the whole offense, each as one unit. Each is scored on three to seven rate metrics, each metric standardized across the 32 teams within that season, combined with the published weights below, and ranked 1–32. **Every team's per-metric rank is published alongside the total**, so the ordering can be read rather than taken on trust. | Unit | Metrics and weights | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Quarterback | EPA per dropback .40 · CPOE .25 · adjusted net yards per attempt .25 · sack rate taken .10 | | Running backs | Rush EPA per carry .40 · yards per carry .25 · rush first-down rate .25 · receiving EPA .10 | | Wide receivers | Yards per target .30 · catch rate .25 · receiving EPA .25 · yards after catch per reception .20 | | Tight ends | The same four, at the same weights | | Offensive line | Sack rate allowed .40 · QB hits allowed .25 · rush EPA per carry .25 · penalties .10 | | Defensive line | Sacks .35 · QB hits .25 · tackles for loss .25 · sack rate per opponent dropback .15 | | Linebackers | Tackles .30 · tackles for loss .30 · passes defended .20 · forced fumbles .20 | | Cornerbacks | Passes defended .40 · interceptions .30 · completion percentage allowed .30 | | Safeties | Interceptions .30 · tackles .25 · passing EPA allowed .30 · forced fumbles .15 | | Defense (whole) | Points allowed per game .30 · pass yards allowed per dropback .15 · rush yards allowed per carry .15 · sack rate .15 · QB hits per game .10 · interceptions per game .10 · passes defended per game .05 | | Offense (whole) | Offensive points per game .30 · net yards per dropback .20 · yards per carry .15 · touchdowns per game .15 · giveaways per play .10 · completion percentage .05 · yards after catch per completion .05 | | Kicking | Field goal percentage .40 · field goals from 40+ .40 · extra points .20 | | Punting and returns | Net punt average .40 · inside-20 rate .30 · return average .30 | Everything a unit **allowed** is charged from the opponents' own stat lines in the same games, never inferred. Sacks allowed by an offensive line come from that offense's record of sacks taken; QB hits allowed and completion percentage allowed come from what the opposing team recorded that day. **The whole defense** (approved by Ben, 2026-09-18) is the unit a "best defense in the league" claim is graded against, and it is built on the seven measures Ben named — points, passing yards, rushing yards, sacks, pressures, passes defended and interceptions — with two substitutions he approved. Yardage is measured **per play**, not per game, because "what's measurable as accurate is what we want, not the incorrect scoring system pundits keep in their heads": a defense that faces sixty snaps is not charged for the extra snaps. And **QB hits stand in for pressures**, which are charting data we do not hold. Points allowed leads at .30 and carries the usual caveat: a defense is charged for every point on the board, including those its own offense's turnovers and its special teams gave away. Sacks, hits, interceptions and passes defended are summed over every player the club fielded, whatever his position code. Measured over 2022–2025 before adoption, the unit put the 2022 49ers, 2023 Ravens, 2024 Broncos and 2025 Broncos first, and the 2022 Bears, 2023 Commanders, 2024 Panthers and 2025 Jets last. **One club's unit finish** (added 2026-09-18, with the whole-defense unit). "The best defense in the league", "a top-five offensive line", "the worst secondary in football" are claims about ONE club's place in a unit ranking, not an ordering, and the parser had been forcing them into a one-team ranking the grader refused. They are their own type, `position-group-finish`, settled by the same season composite the orderings are: the club's rank within the unit. Scored as a **window**, as everywhere else here — "top five" is right anywhere inside it, "best" is a window of one, and a claim in the bottom half is read from last ("the worst defense" finishing 30th is two places off, not twenty-nine). Outside the window, credit falls to zero eight places past its edge, a quarter of the league, on the same reasoning as the twelve-place tier for players. **Baseline: the unit finishes where it finished last season**, the chalk answer the orderings already use; a club with no prior season is scored as mid-pack. **The whole offense** (approved by Ben, 2026-09-18) is the mirror: offensive points per game .30, net yards per dropback .20, yards per carry .15, touchdowns per game .15, giveaways per play .10, completion percentage .05 and yards after the catch per completion .05. Three choices worth stating. **The points are the offense's own** — built up from passing and rushing touchdowns, field goals, offensive two-point conversions and the extra points that followed offensive touchdowns, never the scoreboard, which would credit a pick-six or a punt return to an offense that never took the field (Ben's flag; measured against the scoreboard with the defensive and return scores added back, every club lands within four points, the odd missed extra point). Passing and rushing **touchdowns are one number**, so a run-heavy offense is not ranked behind a pass-heavy one for scoring the same way. And yardage is **per play**, as for the defense. Completion percentage and yards after the catch stay at the weights their signal earns: a checkdown offense can lead the league in both. Measured over 2022–2025 before adoption, the unit put the 2022 Chiefs, 2023 49ers, 2024 Ravens and 2025 Bills first, and the 2022 Texans, 2023 Jets, 2024 Browns and 2025 Browns last. Where this is weakest, said plainly: - **Coverage is only partly visible.** Without per-snap charting we cannot see targets allowed or missed tackles, so linebackers, cornerbacks and safeties are measured less well than the units in front of them. - **Cornerbacks and safeties share a number.** Completion percentage allowed and passing EPA allowed belong to a whole pass defense and cannot be split between the two without data we do not have. Each group carries one of them, marked as a shared team figure. - **Running back metrics are blocking-dependent.** Yards per carry says as much about the line as the runner. - **Team penalties are not an offensive line statistic.** They are counted team-wide, which is why they carry the smallest weight in that group. Baseline: the same unit's order last season. The weights above are an editorial judgment, not a fact. They decide who was right, so they are published here rather than left in the code, and changing one means regrading every claim of this type. ### 2.11 Trade and signing calls **Added by amendment, Ben, 2026-08-16.** 797 of these were captured and none were ever scored. They are not one kind of claim, and grading them as one would score men on questions they did not ask. Three cases: - **A destination is named** — "he signs with the Jets". Right if he appeared on that club's roster that season. **100 or 0.** - **No destination, but a move is asserted** — "he gets cut", "he finds a new home", "not coming back". Right if he was not with his old club that season. Being on no roster at all still counts as leaving. **100 or 0.** - **The claim is about money** — a contract, an extension, the franchise tag. **A contract call, not a movement** since Amendment 18 (§2.16), and on hold with every contract call since Amendment 23. Before Amendment 18 these were declined for want of the data. A claim carrying both a move and a number ("not coming back to the Vikings, and he'll get $50m") is graded on the move, and the workings say which half was checked — the rule §2.2 already applies to a score prediction that got one of its two claims right. **A denial of a move is not a movement claim** (added 2026-08-18). "I don't see them trading him", "he's not going anywhere" — the parser is told to skip negated claims outright, and the grader now enforces it: graded literally, a correct "he stays" earned a zero for the move it never predicted. These were declined until Amendment 19 (2026-10-07); they are now graded as **stays**, against the stay base rate below, never as the move they deny. **The club he is already with is not a destination** (added 2026-09-18). The parser was filing a player's own club as where he was going — "Maxx Crosby to move to Las Vegas Raiders", under a quote that the Raiders may trade him — both when no destination was named and when the claim was that he stays. Ingest now refuses a fresh claim's current club as a destination, and the grader reads any named club against the roster the season before: a club he was already with is either the departure the words describe ("try to trade him again"), graded as one, or a claim that he stays, which a person reads when its words do not say so (a stay its words DO say is graded, Amendment 19). Graded as a destination, a stay was 100 every time he did not move — the naive answer being paid for. **The season it settles in** is the claim's own calendar year, which is deliberately not §1's rule: a movement claim made in January is about the season ahead, not the one finishing, because that is when players move. **Skill delta: he stays where he is.** Measured across every season we hold, a player is with the same club the following season **65% of the time** (14,747 of 22,837). That is strong chalk, which is why calling a move correctly is worth something and calling a stay is worth almost nothing. A player we hold no prior-season roster for declines: without knowing where he was, there is no move to describe and no naive answer to beat. ### 2.12 Roster spot calls **Added by amendment, Ben, 2026-08-16.** 1,171 of these were captured and none were ever scored — not for want of a rule, but because the answer was nowhere in our data. `players.status` holds one status per player, overwritten every week, so it says what a man is now and never what he was in 2025. **The claim's own direction decides what "right" means** (amended 2026-08-18 — the first grader scored every one of these as "he makes it", which put a zero on a correct "he gets cut"). Three cases, read from the claim's wording: - **He makes the 53** — right if he reached that club's 53-man roster that season. **100 or 0.** - **He misses it** — "he gets cut", "he won't make the roster". Right if he did not reach it. **100 or 0.** - **He lands on the practice squad** — right if he spent a week on that club's practice squad and never reached its 53. **100 or 0.** A phrasing that declares no direction declines rather than being guessed at. The short outcome we file each claim under counts as its wording too, when it is one of the plain shapes whole — "makes 53-man roster", "makes the initial 53-man roster", "cut", "cut from the 53-man roster", "locked in", "roster lock" (added 2026-09-24, after 460 claims labeled that way had been declined as directionless). It is read only when nothing else in the claim points either way, so it never outvotes the pundit's own words, and a label that names the question without an answer — a bare "53-man roster" — still declines. **Hedged roster calls** (Ben, 2026-10-01). A hedge is graded as the claim it hedges, and for a roster call that makes doubt about a spot a cut — "roster spot in jeopardy", "the odd man out", "on the chopping block", "an uphill battle to make the 53" — and standing in a competition making it — "favorite for the fifth receiver spot", "the inside track", "a legitimate shot at making this team". Three things still decline. A **condition** the claim names ("the odd man out if Washington keeps only seven receivers", or a quote that opens "if New York keeps six"), because we do not check how many a club kept. A judgment of **merit** ("deserves a roster spot", "has done enough to warrant one"), which forecasts nothing. And a claim that is really about **something else** — the PUP or reserve list, game day, a depth-chart role, a waiver claim, "back next year" — because a verdict on the 53 would answer a question the pundit did not ask. **A reviewer can answer for a claim that says nothing** (2026-10-01). Reading the quote, a reviewer can set a claim's direction — makes the 53, cut, practice squad — and that call outranks the wording everywhere it is read: the grade, and the sentence the claim is printed as. It is for the claims that carry no label at all, and for projections whose every entry was labeled only "53-man roster" while the quote plainly lists who is kept. A call that reverses a graded claim's direction clears its grade, and the claim is scored again. "Made the 53" counts a week where he was active or inactive: both mean he is on the roster, and inactive only means he did not dress for that game. The practice squad is precisely not making the 53. Injured reserve is not failing to: a man the club carried only on its reserve list all season, and never on its 53 or its practice squad, declines rather than scoring either way. If he comes back and reaches the 53 he grades the week he does (Amendment 11). A club has to be named — 1,159 of the 1,171 name one. "He makes a roster somewhere" is a much weaker claim than what these pundits are actually saying, and grading it as though it were the same would flatter them. **The season it settles in** is the claim's own calendar year, §2.11's rule for the same reason: a February "he makes the 53" is about the cutdown eight months ahead, not the roster that was just torn down. **Skill delta: he was on somebody's 53 last season, so he makes one again.** A veteran making a roster is barely a prediction; calling an undrafted rookie onto one is the read worth paying for. The baseline is what that chalk answer would have _scored_ — it says "makes" for a veteran and "misses" for everyone else, and it earns its 100 only when the season agreed with it. A season whose weekly rosters we have not ingested declines rather than scoring: a missing row means "he did not make it" only when we hold the file it would have been in. When the file settles these claims, and what it takes for a file to count as the cutdown at all, is Amendment 11. ### 2.13 What we decline to grade, and enforce in code **Added by amendment, Ben, 2026-08-16**, after both of these reached a graded record. **A single-game claim is never graded against a season.** The guard has existed since the first time this happened; it keyed on a named week, opponent or night, and a game preview naming none of those walked past it. "Final Prediction: 14 carries, 96 yards, 1 td" is unmistakably one game, and 23 of them were scored against season totals — 21 landed on zero, all on one named outlet. The phrase is now part of the guard, and a dated heal puts the 23 back to ungraded, deciding by the guard itself rather than a second copy of it. **The guard reads the setting too** (2026-10-09). A same-game parlay leg ("C.J. Stroud: 280+ passing yards"), a rank given in a "Week 9 Wide Receiver Rankings" episode and a stat line under "Stat Predictions for Bears vs Jets" never repeat the game or week they are about: the title, or the top of the show, names it once. A season statistic or positional finish is now declined when its source title or opinion names a same-game parlay, props on a named game or week, a game between two clubs, or a week's rankings, unless the claim's own words name its season. Futures pieces, season-long props and a bare "Week 5" in a title are not settings of this kind; a Week 5 column is full of season calls. The guard also hears a game the way a transcript says it: "in Week sixteen", "this is the Justin Jefferson game", "against Buffalo", "an anytime touchdown". **A sportsbook line is a price, not a claim by a person.** §4 has always said so and the parser is told to skip them; five got through and sat in the corpus waiting to be graded. There is now a check in code, because a prompt is a request and this is the enforcement. The half-point identifies them: a person says "about sixty yards", only a book says 58.5, because the half exists to prevent a push — except where halves are the stat's own arithmetic. Sacks are recorded in halves, averages are quotients, ratings run to a decimal; a half-point on those is a person talking, and the check exempts them (narrowed 2026-08-18, after "more than 13.5 sacks" was caught). **A wager is not a prediction** (added 2026-09-18, after Ben found both kinds in the queue). A daily-fantasy salary play — "worth his $5,600 on DraftKings this week" — is lineup construction for one contest, and a salary is not the draft cost a sleeper call is measured against; a betting pick — "I'll be targeting the Texans, Stroud alternate props and correlated bets" — is a pick made because of a price. §4 declines both and the parser is told to skip both; ingest now refuses them before they reach the queue, and the grader refuses whatever was filed before it did. The check is narrow on purpose: DraftKings Network and FanDuel employ ordinary pundits making ordinary season predictions, so a platform's name is not the tell — a salary figure beside it is. **A weekly start/sit call is not a season claim** (added 2026-09-18, with the §2.6.6 amendment). Hot plays, cold plays, streamers, "outside my top 12 this week": lineup advice for one game, filed by the parser as sleepers and busts. The same enforcement as the wagers above — the parser is told, ingest refuses, the grader declines, a heal clears what was filed — through one predicate that reads the claim together with its opinion, because the week is usually named once at the top of the show rather than in every call. **A week's ranked list is not a season list** (same day). "Flex Rankings table, Week 2" was filed as a top-players list and would have been graded against season finishes. Read like the weekly plays above, over the claim and its opinion. The rule all of these are instances of: **when a claim's scope does not match the actuals we would grade it against, decline it.** And the rule Ben set on 2026-09-18 for what "decline" means: **a claim we do not grade is not surfaced — not in the moderation queue, not on the site.** One predicate (`claimDecline`) decides for every type; ingest never files such a claim, the grader declines any filed before, a dated heal moves them to unresolvable, and every public list and count is fenced to rows that are not wrong-sport and not unresolvable-without-charge (`lib/onTheRecord.ts`). A row marked unresolvable _with_ the pundit charged for its vagueness stays visible, as "too vague to check": that is a statement about the pundit, and their clarity score depends on it. A row we declined for our own reasons is simply not part of the record. Grading anyway publishes a number against a real person's name for a claim they never made, which is the worst thing this system can do. --- ## 3. Aggregation A pundit's headline accuracy is the **mean of their graded prediction scores**, computed overall, per sport, and per prediction type. **Time weighting.** Not in the POC, per your answer. The engine stores a timestamp on every grade so that a decay function can be switched on later without regrading anything. When we do turn it on, my recommendation is a two-season half-life: last season matters most, three years ago barely. **Score at time of prediction.** Every published opinion records the pundit's accuracy _as it stood when they made the call_, alongside their current score. Without this, an old post silently rewrites itself as the pundit's record changes, and the archive stops being trustworthy. ### The three levels, and which one is the evidence **Amended by Ben, 2026-08-01.** Originally a pundit's record was the mean of their graded predictions. That counted a 32-pick mock draft as 32 independent judgments and rated its author "established" off a single article — precisely what the sample thresholds below exist to prevent. | Level | What it is | How it is computed | | --------------------- | ----------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Take** | one checkable claim | accuracy 0–100 and a skill delta, per §2. The atom, and never aggregated away — it is what a reader browses when they ask what pundits think about one player | | **Article in a type** | one article's takes **of one type** | the mean of those takes. An article that ranks 32 teams _and_ picks a Super Bowl winner produces **two** assessments, not one, because those are measured against different baselines and averaging them would mean nothing | | **Pundit in a type** | their record in that type | the mean of their article scores in it | | **Pundit overall** | the leaderboard number | the mean over every (article × type) assessment | **Every article carries equal weight, however many takes are in it.** A 91-item list and a one-line take are each a single act of judgment: one moment, one set of information, one chain of reasoning. Counting the takes inside them as separate observations would treat a pundit's longest article as their most tested opinion, and would make publishing longer lists the cheapest way to move up a leaderboard. **One article, however many URLs it arrives as.** The pages of a paginated article are one article (Ben, 2026-08-01), and so are the versions of an article its writer revises and re-publishes as 1.0, 2.0, 3.0: every version is graded, the versions count once together, and each weighs equally inside that one article (Amendment 10). Stated the way a reader would say it: _"Rich Eisen averaged 76% on power rankings this season."_ Not 76% of his individual placements — 76% across the articles. ### Small samples We publish scores at any sample size, per your answer — with an explicit callout. **The thresholds count ARTICLES** (Ben, 2026-08-01). - **Fewer than 10 graded articles** in a scope: show the number, but label it prominently as provisional, and exclude the pundit from leaderboards for that scope. - **10–29:** show without the warning, still flagged as a limited sample on the pundit page. - **30+:** treated as an established record. Where two records are otherwise level, **the one built from more graded takes ranks higher** (Ben, 2026-08-01). A longer article genuinely does tell us more about someone, even though it remains one article of evidence about them. The thresholds are judgment calls, not statistics — 30 is the conventional point where a mean starts to stabilize. What matters is that a pundit with three lucky calls never appears to be the most accurate person in football, and that neither does one with a single long list. **Clarity is article-weighted too**, on the same reasoning: it is a statement about how a person writes, so one enormous vague list should not define someone who is otherwise precise. The minimum sample still counts claims, because having enough evidence to publish a rate is a different question from how the rate is weighted. --- ## 4. Other prediction types worth grading later You asked whether we are missing frequently published, reconcilable opinions. Here is the survey, ordered by how cleanly each resolves. **Strong candidates — unambiguous resolution:** | Type | Resolves | Note | | --------------------------------- | ------------- | ------------------------------------------------------------ | | **Playoff / Super Bowl picks** | End of season | Published constantly, resolves cleanly, huge public interest | | **Award picks** (MVP, OPOY, DPOY) | End of season | Voted outcomes are unambiguous | | **Draft order rankings** | Draft night | "Player X goes top 10" is exactly checkable | | **Mock drafts** | Draft night | Gradeable like power rankings — positional distance | | **Stat milestones** | End of season | "4,000 yards", "15 sacks" — pure numbers | | **Playoff seeding** | End of season | Ranking-style grading | **Good candidates — need a definition first:** | Type | Wrinkle | | ------------------------------------- | --------------------------------------------------------------------------------------- | | **Fantasy projections** | Resolves weekly, but scoring systems differ between platforms; we would have to fix one | | **Breakout / bust calls** | Needs an agreed threshold for what counts as a breakout | | **Coach / GM hot seat** | Resolves on firing, but "survived the season" needs a defined window | | **Trade and free-agency predictions** | Clean when specific, vague when not ("they'll be active") | **Weak candidates — I would not grade these:** - **Injury return timelines** — the honest answer is usually medical and private, and being wrong is not a pundit failing. - **"Team X has a good culture"** — unfalsifiable. - **"Surprise team" calls** — declined on the record in Amendment 5. Graded on a playoff berth they hit 46.7% against a 43.8% base rate, and "surprise" needs a prior we do not store. - **Player movement** — graded since §2.11, resolved from season rosters rather than the transactions data Amendment 5 correctly said we lack. The contract half became its own type in Amendment 18 (§2.16). See the postscript to Amendment 5. - **Anything about a player's character or attitude** — even where checkable, grading it invites exactly the tone we said we would avoid. The last group matters editorially, not just technically. The line I would draw: **we grade claims about outcomes, never claims about people.** --- ## 5. What this means for the tone The methodology is the product's credibility, so it is public in full, and every graded prediction links to the source and shows its own arithmetic. A pundit who disagrees with a grade should be able to see exactly how it was computed and tell us we got it wrong — the corrections channel in E10 exists for that. Two things we will not do: - **No leaderboard of "worst" pundits as a headline feature.** The data will support it; presenting it that way is the cheap version of this idea. Rank by accuracy ascending if a user asks for it, but do not build the product around humiliation. - **No grading of predictions a pundit did not actually make.** Paraphrase is where this kind of site loses its integrity. If a claim is too vague to grade, we do not grade it. --- ## Sign-off **Approved by Ben, 2026-07-31.** | # | Question | Decision | | --- | ----------------------------------------- | ------------------------------------------ | | 1 | Betting odds in the baseline? | **No.** Public prior-season results only | | 2 | Score predictions | **Margin only**, not raw scores | | 3 | Weight earlier-season predictions higher? | **Yes** — difficulty multiplier, §2.3 | | 4 | Small-sample thresholds | **10 / 30** as drafted | | 5 | Pull anything forward from §4? | **Playoff + Super Bowl picks only** (§2.5) | Five prediction types are therefore in scope for the POC: power rankings, game outcomes, score predictions, season win totals, and playoff/Super Bowl picks. ### Amendment 1 — scored scope, approved by Ben 2026-08-01 We were capturing seventeen kinds of prediction and permitted to score six, which left **336 of the 337 claims held at the time unscorable** and the public site with nothing to show until January 2027. | # | Question | Decision | | --- | ----------------------- | --------------------------------------------------------------------------- | | 1 | Widen the scored scope? | **Yes — all five captured types** (§2.6–§2.8) | | 2 | Mock-draft scoring | **Ordering vs a shuffle** for skill, per-pick closeness for accuracy (§2.6) | Six further types are therefore in scope: draft position, player season statistics, positional finishes, fantasy draft rankings and top-player lists. The remaining captured types stay unscored until they have both a grader and a baseline written here. ### Amendment 2 — the unit of a pundit's record, approved by Ben 2026-08-01 | # | Question | Decision | | --- | ------------------------- | ------------------------------------------------- | | 1 | What does a record count? | **Articles**, not takes (§3) | | 2 | Confidence thresholds | **10 / 30 articles**, takes break ties in ranking | | 3 | Clarity score | **Article-weighted**, like the rest | ### Amendment 3 — baselines signed off, Ben 2026-08-01 Every baseline now in the code has been reviewed. There are no unsigned numbers left in the scoring. | Type | What the naive strategy does | Value | | -------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------- | | Season win totals (§2.4) | last season's wins, pulled toward 8.5 | halfway | | Super Bowl pick (§2.5) | last season's champion goes again | — | | Player season stats (§2.7) | last season's number, pulled toward his position's average; a rookie gets the positional average | halfway | | Positional finish (§2.8) | where he finished last season; **no prior season → his own position's startable tier** | QB12 / RB24 / WR36 / TE12 | | Ranked lists (§2.8) | a coin-flip ordering | 50 | | Mock drafts (§2.6) | a shuffled mock | 50 | The positional-finish rookie rule was changed at sign-off: it had been a flat 24 for every position, which held a rookie quarterback to a generous baseline since twelve quarterbacks start rather than twenty-four. ### Amendment 4 — the bust and sleeper bars, Ben 2026-08-16 Ben's question was what to do with the bust calls we were dropping for naming no draft cost, and his proposal was to supply Average Draft Position ourselves and measure a player against the average of the picks immediately ahead of and behind him. | # | Question | Decision | | --- | --------------------------------------- | ------------------------------------------------------------------------- | | 1 | Supply ADP where the source names none? | **No.** The consent is what makes a quoted cost usable (§2.6.6) | | 2 | What measures an uncosted claim? | **His own prior season**, or the startable tier for a rookie (§2.6.6) | | 3 | The cost→bar step table | **Replaced by a curve** — the cliff at each round boundary was the defect | | 4 | Peer-relative bars | **Yes, for NFL draft busts** (§2.6.7), where every fact is one we hold | | 5 | Two neighbors, or many? | **Many** — median of the picks within sixteen slots, across every draft | **Adopted out of order, and the record should say so.** The curve went into the live grading path on 2026-08-16 without the measurement this amendment required first — so for two days the archive held two regimes: claims graded before under the step table, claims graded after under the curve. The doc's own rule is that the math changes only once it has been measured against the archive; the draft-distance curve was accepted on the strength of "Round 1 average moves 74.2 → 73.1, and the ranking of pundits is unchanged", not on the strength of its reasoning. The repair (2026-08-18): run `pnpm db:measure:value-bars` and record the flips and ranking movement here, then `pnpm db:regrade:value-bars` sends every bust and sleeper claim back through so ONE set of bars scores the whole record. The change is expected to be mildly looser on fantasy value calls and materially looser on NFL draft busts, and "expected" is not a measurement. **Measured, 2026-08-18.** The measurement could not have been run before adoption: `db:measure:value-bars` shipped querying a column that does not exist and fell over on contact with the database, which is its own small argument for running a script before citing it. Fixed and run, the verdict is that the stakes were near zero: almost the entire population is unsettled (388 bust and 1,488 sleeper calls await the 2026 season), the handful of graded draft busts score **identically under both bars** (66.7 → 66.7, no flips), and the ranking across these types does not move a single pundit. The archive regrade ran the same day regardless: 417 graded rows — 400 already scored by the curve by the daily runs since adoption, 17 still carrying the step table — were sent back through `db:regrade:value-bars`, so one set of bars scores the whole record and this amendment's condition is met after the fact rather than never. Amendments after this point require Ben's sign-off and a regrade, since the published archive is computed from these rules. --- ### Amendment 5 — two types declined on the record, Ben 2026-08-18 Ben asked which of the ungraded claims could be graded if the taxonomy grew to support them. Most of the answer was yes. Two are written down here as **no**, because a refusal nobody can find is indistinguishable from an oversight, and this pair will otherwise be re-proposed every time somebody counts the pile. Both were, until now, sitting in `unclear-claim` — the bucket whose definition is "the pundit did not say enough to check. Theirs." Neither belongs there. They are ours: we have no grader and, for one of them, no data and no defensible definition either. | # | Type | Rows | Decision | | --- | ----------------- | ---: | ---------------------------------------------------------------- | | 1 | `surprise-team` | 197 | **Declined.** No stored prior, so "surprise" is not a predicate | | 2 | `player-movement` | 744 | **Declined** — superseded within hours; see the postscript below | #### 1. `surprise-team` — a coin flip is not a measurement The type's declared resolution is a playoff berth in the season named. Joined against `team_season_records` for 2025: **92 claims, 92 matched, 43 made the playoffs — 46.7%.** The league base rate over 2022–2025 is 14 of 32, **43.8%**. A 2.9-point edge, published beside a named analyst's face and called accuracy. The deeper problem is definitional. A surprise is relative to an expectation, and we store no expectation. The only column in the schema resembling one is `coach_season_scores.expected_wins`, which is coach-level, derived from the prior season, and not a market line. Without a stored prior, _surprise_ has no truth condition — grading it would measure whether a team made the playoffs, which is a different claim from the one the pundit made. Of the 194 pending, 71 use "sleeper / sneaky / under the radar / surprise", 18 mention the playoffs at all, and **none states a win number**. One would be scored on a prediction its author never made: > "the Falcons don't control their own destiny… it's become likely that they > will miss the postseason" A made-the-playoffs grader marks that a miss for a pundit who said they'd miss. **Revisit only if** we begin storing a preseason expectation — a market win total or a published consensus — captured _before_ the season, against which a surprise can be defined. Absent that, this stays declined. #### 2. `player-movement` — we do not hold the actuals `prediction_types.resolution_source` says "Announced transactions." There is no transactions table. Searching `information_schema` for `%transact%`, `%contract%`, `%signing%` and `%roster%` returns nothing. `player_team_stints` cannot substitute, for two independent reasons: 1. **It carries no transaction type.** The columns are `player_id, team_id, started_on, ended_on, jersey_number`. Nothing distinguishes a trade from a cut from a free-agent signing, and those are different claims. 2. **There is no offseason in it.** Every `started_on` across all 22,334 rows, by month: Jan 624 · Feb 4 · **Mar–Aug 0** · Sep 16,305 · Oct 1,884 · Nov 1,727 · Dec 1,790. The stints are built from weekly in-season rosters. Player movement happens in March. We have no rows there at all. So the honest classification is **both** `no-grader` and `missing-actuals`, and the claims stay ungraded. Grading them off stint boundaries would report a March trade as whatever the September roster happened to say. **Revisit only if** a transactions feed is ingested with a type and a date. #### What is NOT declined `roster-spot` (1,148 rows, the largest ungraded cluster on the site) is **deferred, not refused.** It is worth roughly six points of grading rate on its own. Preconditions, all of which must land before it is started: - a roster **status** at **week grain** — the nflverse feed carries a status code our ingest discards, and `collapseStints` destroys the week grain at write time. A cutdown question needs both. - a real outcome enum from the parser rather than free text. - a three-outcome grader — 53-man / practice squad / released — evaluated **against the club the claim named**, not against any roster. - a ruling on serial projections. One Dolphins article is 90 predictions, and one writer ships _2.0, 3.0, 4.0_ with the same players re-graded each time. `predictions.supersedes_prediction_id` and `scripts/lib/seriesKey.mjs` both already exist; the decision is which version scores. Shipped without the week grain, this grades ~92% of the cluster TRUE off a 90-man camp roster, which is worse than not grading it. The 2026 cutdown is the first honest opportunity. #### Postscript, 2026-08-18 — what happened to this amendment within hours Three changes merged the same evening this was written, none reconciled against it, and for a day the site's contract said both "declined" and "graded" about the same claims. Setting the record straight: **`player-movement` is graded, and this amendment's evidence still stands.** The grader that shipped (§2.11, an hour before this amendment merged) does not use the resolution this amendment rejected. Both objections above were to settling _transactions_ — and no transactions table exists, so the contract case still declines against exactly that gap. What the grader settles is _where he played_: whether a man appeared on a club's roster in the season the claim names, which the season rosters do answer. The March row this amendment notes we lack is the transaction's _date_, not its outcome — by September the rosters say where he landed. The calendar-year window §2.11 records was implemented on 2026-08-18; between the grader shipping and that fix, no movement claim had yet been graded. The row counts differ between §2.11 (797) and the table above (744) because they were measured on different days with different status filters; neither is wrong, and the next count will disagree with both. **`roster-spot` started early, and two of its four preconditions arrived late.** The week-grain status this amendment demanded did ship — `player_season_roster` totals the weekly feed per season and club, which is the file the grader reads. The outcome distinction did not ship with it: the first grader scored every claim as "he makes it", and the three-outcome rule (§2.12) was added 2026-08-18, before any claim of the type had been graded. **The serial-projection ruling this postscript held open is made.** Which version of a writer's _2.0, 3.0, 4.0_ scores was decided on 2026-09-24: every version is graded, and all the versions of one revised article count as a single article in the record. See **Amendment 10**. --- ### Amendment 6 — member picks are scored by inheritance, Ben 2026-08-22 Members (logged-in readers) can record a **pick** on any published, unsettled claim: _agree_ or _disagree_ with the pundit's call. When the claim grades, the pick grades with it, in the same pass, by inheritance: | Pick | Accuracy | Skill (MissedTakes Score input) | | -------- | -------------------------- | ------------------------------- | | Agree | the claim's accuracy | the claim's skill delta | | Disagree | 100 − the claim's accuracy | − (the claim's skill delta) | A claim with no skill delta grades the pick's skill as null — "no baseline existed", not zero. Nothing about how pundits are graded changes: member picks are a new consumer of existing grades, not a new grader, and members' picks never alter a pundit's record. **Why inheritance.** One grading engine, one set of baselines, one published methodology. A member's 62% and a pundit's 62% mean the same thing because the same arithmetic produced both — the property any member-vs-pundit comparison depends on. The complement is the honest score for disagreement: "this will not land" earns full credit when the claim scored zero and nothing when it was exact, and the mirrored skill delta means agreeing with chalk earns roughly nothing — which is also the anti-gaming property. **The integrity rule.** A pick is accepted only while the claim is unsettled and only on claims the site publishes; it may be changed or withdrawn until settlement (only the final stance scores; the first pick's timestamp is retained through flips); at settlement it is frozen. Enforced by row-level security in the database, not only by the application. **Aggregation.** `member_scores` mirrors `pundit_scores`: one row per member per scope, accuracy and MissedTakes Score always published with their sample size, confidence through the same tiers (provisional <10 · limited 10–29 · established 30+). The tier counts **graded picks** for now — the strict article-unit parity question (§Amendment 2 reasoning) is explicitly deferred and must be revisited before any public member leaderboard ships, which will also set the leaderboard floor (proposed: 10 graded picks) and the member-vs-pundit display rule: head-to-heads are computed **only over the intersection** of claims both sides took, never the two overall records side by side. ### Amendment 7 — single-game claims grade when the game is final, 2026-09-11 Driven by Ben's 2026-09-11 request for per-game digest posts once the NFL season started, and written into the code the same day. **The one number it touches, the in-season home-margin baseline (below), was adopted 2026-09-24 as the code computes it:** the mean home-minus-away margin over every game of the season being graded that is final when the grading run happens, with each graded claim keeping the baseline it was graded against. On 2026-09-24 that was +0.19 points, over the 32 games of Weeks 1 and 2. Everything else is timing, not arithmetic. Game-outcome and game-score claims (§2.1, §2.2) had been held by the season gate like every season-long type, so a Week 1 pick would not have graded until the season completed in February — months after every reader stopped asking, and fatal to a post that digests a game the week it was played. The same reasoning that let roster-spot claims grade at the cutdown (§2.12) applies: once the fixture is final, nothing the rest of the season does can change the answer. **What changes:** - Game-outcome and game-score claims now grade on the first grading run after their fixture is final. The grader skips (pending, not declined) when the claim's game has not been played — including a rematch claim published between two meetings, which waits for the second one. - §2.1.1's date rule is unchanged, but its "published after every meeting" decline now only fires once the season is complete; mid-season, a claim dated after every meeting played so far is treated as being about a rematch still on the schedule. - §2.1.2's never-played matchups still wait for the postseason — mid-season, two clubs that have not met may yet meet. - The grader now **stamps the settling fixture** (`predictions.game_id`) onto the graded row. That is bookkeeping, not scoring: it is how "every take this game graded" becomes a query for the game-digest posts. **The one scoring nuance.** The game-score baseline is "the home team by the season's average home margin" (§2.2, Amendment 3). Graded mid-season, that average is computed over the games played **so far** rather than the full season, so a claim graded in September and the same claim graded in February could see baselines a point or two apart. Accepted: the number still comes from the season being graded, the drift is small, and the alternative — holding every score prediction until February — is the problem this amendment exists to fix. A season-end regrade may sharpen these baselines if the drift ever proves material. **Difficulty is unchanged.** The multiplier reads the claim's date, not the grading date, so grading earlier moves no difficulty number. ### Amendment 8 — claims settled by reported outcomes, adopted 2026-09-24 Built 2026-09-16, the same day as `docs/missedtakes-reported-outcomes.md`, which holds the design and the decisions; this amendment records only what touches a grade. Adopted as written; adopting it changed no grade. Until now every actual came from a fixed list of ingested datasets, and a claim those datasets could not answer stayed `unresolved_by_us` forever: a Week 2 starter call, a March trade, whether a man played Sunday. The crawl that carries the takes also carries the answers — a coach naming a starter, a club announcing a signing — and this amendment lets a **confirmed reported fact** settle a claim where a dataset is silent. **What counts as an actual.** A fact reported by an article we crawled, with who said it and how the outlet knows, **confirmed by a person** on the reported-outcomes queue — or, only while the corroboration switch is on, by an official statement or two independent outlets' own reporting. Nothing pending, rejected or superseded ever reaches a grade. A confirmed fact that is later rejected **resets every grade built on it** to pending; a grade must be able to show its arithmetic, and arithmetic built on a withdrawn report shows nothing. #### §2.14 Claims a reported outcome settles **Week-scoped depth-chart calls.** "Wentz starts in Week 2." A claim naming a week is about one Sunday, and §2.13 already forbids grading it against a season of usage. It now grades **once that week's game is final**, against the confirmed starter reports for that club and week: **100** if he was named the starter, **0** if he was named as not starting. A report naming a different man settles it only at **quarterback, center, kicker or punter**, the jobs a club fills with one man (Ben, 2026-10-02, Amendment 15): there, another man named the starter means he was not, and the claim grades **0**. Everywhere else a club can start more than one — two safeties, two tackles, three receivers, a committee at running back — so a report about another man says nothing about him, and only a report about him settles the claim. Only slot-1 claims can settle this way — nobody announces a third receiver. A Wednesday announcement is evidence, not the outcome; the grade waits for the game so a late scratch cannot leave a grade standing on a plan. **What a start is (Ben, 2026-09-17).** An in-game injury does not change who started. The man who started and left hurt in the second quarter **started**; the man who came on for him **did not** — a relief appearance is not a start, however much of the game he played. If the first man stays out and the second starts the next game, that next game is a start for him. The detector is told this in so many words (leaving a game is `left_game`, never `out`; coming on in relief is not a fact at all), and the reviewer has a one-click rejection for each misread that the next extraction is then shown. **A man who did not play did not start (2026-09-17).** The weekly snap counts list a player for a game only if he took a snap in it, so once his club's week is on file, no line means he never got on the field — a late scratch after a Wednesday announcement, a Friday setback, a reserve placement. A week-scoped starter call about him grades **0** from the dataset, with no report needed, and that check runs **before** the reports so a stale "named the starter" cannot outvote it. The workings say `settledBy: "weekly snap counts"`. The converse is **never** read: snaps cannot say who started among the men who played — a quarterback who came on in the second quarter took most of the snaps and, by the ruling above, did not start — so "he played" always falls through to the confirmed reports. _Skill delta: he keeps the job he had._ The naive call is that last season's slot-1 man at the position starts and nobody else does, read off the same usage chart the season-long grader uses. A rookie or a backup last year is naively not the starter, so calling him in is the read worth paying for. **Movement claims (§2.11).** A confirmed trade, signing, release, waiver or retirement settles the claim in the month it happened rather than when the September rosters land — the transactions gap Amendment 5 documented, closed by the crawl rather than by a feed. Reports **confirm** a move; they never prove one did not happen, so a claim the reports do not settle falls through to the roster rule unchanged. Baseline unchanged: he stays. **Roster-spot claims (§2.12).** A confirmed "made the 53" report settles the claim before the weekly roster file exists, or for a club the ingest has not reached — it settles "makes" right and "misses" or "practice squad" wrong, exactly as reaching the 53 does in the file. A cut or practice-squad report settles **nothing on its own**: a roster is in flux all season, a man cut at the move to 53 and signed back in November made the roster, and "he makes the 53" was right (Ben, 2026-09-16). Those claims wait for the file, which already holds them until the season finishes before saying he never got there. Same baseline throughout. **What the reviewer teaches it.** Confirming and rejecting are decisions about facts; correcting an evidence row's _authority_ ("that was speculation, not the club") is the signal the extraction learns from. The corroboration bar reads the reviewer's label over the model's, the next extraction is shown the reviewer's recent corrections as worked examples, and an outlet whose facts the reviewer keeps rejecting (two in five, over at least five decisions) stops counting as an independent voice. None of this touches a grade directly; it changes what may be confirmed without a person once the switch is on. **Availability calls — NEW TYPE, NOT YET GRADED.** "He won't make it through concussion protocol this week", "he'll play Sunday." Captured as `player-availability` under the capture-everything rule. Proposed grading: **100 or 0** by whether he was active for that week's game, from a confirmed availability report, once the game is final. _Proposed skill delta: he plays_ — a man on a roster is naively active, and calling the sit is the read. §4 lists injury _return timelines_ as a weak candidate; a named-week availability call is narrower than a prognosis, but it is close enough that this type stays out of the approved list until Ben rules. The grader is written and unreachable. _Ben, 2026-09-17: "I don't foresee a future where we're grading pundit takes on whether someone will get injured."_ The type stays captured and ungraded, and injury statuses no longer ask the reviewer for a decision at all — they are filed with their evidence and nothing more. Whether a man played in a given week is a question for the weekly roster and snap data, not a report. **Difficulty is unchanged.** The multiplier reads the claim's date, as always. **What the public workings say.** A grade a report settled says so (`settledBy: "a confirmed report from the crawl"`) and names what was reported, so a pundit disputing it can see the report was the evidence and use the corrections channel against the report rather than the arithmetic. ### Amendment 9 — fantasy defenses, approved by Ben 2026-09-16 Built the same day as `docs/missedtakes-fantasy-defenses.md`, which holds the design and the three pain points that prompted it; this amendment records only what touches a grade. Approved the day it was proposed; both types are in `APPROVED_TYPES` in `grading.ts`. A team defense is drafted and started as one unit and scored on one line, and until now the only thing we could do with a claim about one was capture it. Two claim shapes reach the queue: - **`fantasy-defense-finish`** — one club's finishing place among fantasy defenses: "the Jets finish last", "the Bucs are the DST1 this week". Numeric, subject the club. (It had been filed as ordinal; that was the shape of the other claim.) - **`fantasy-defense-ranking`** — a ranked list of defenses, most often the weekly DST table every fantasy outlet publishes. Ordinal, subject the league. The projections the table prints beside each club — points allowed, sacks, turnovers, fantasy points — are kept with the entry. #### §2.15 Fantasy defenses **What counts as the actual.** Each team-week's defensive line — points allowed from the game, yards allowed from the opponent's row, sacks, interceptions, fumble recoveries, safeties, blocked kicks, defensive and return touchdowns — scored under **standard D/ST scoring**: sack 1, interception 2, fumble recovery 2, safety 2, blocked punt, field goal or extra point 2, defensive or return touchdown 6, and points allowed by tier (0 → 10, 1–6 → 7, 7–13 → 4, 14–20 → 1, 21–27 → 0, 28–34 → −1, 35+ → −4). This is the default on ESPN, Yahoo, NFL.com and Sleeper and the scoring FantasyPros' weekly projections are published in. The tiers are per game, so a season's points are a **sum of weekly lines**, never a function of a season's points allowed — which is why `team_week_defense` exists beside the season table. One published approximation. **Points allowed** is the final score against, which most platforms reduce for points the club's own offense gave away (a pick-six thrown, not conceded); the feed cannot tell those apart, so the game score stands. It is rare, applies identically to every pundit, and is said on the methodology page rather than hidden. (Blocked kicks were a second one until Amendment 24 scored them.) **Which touchdowns are the unit's — a fix, not a rule change (2026-10-10).** A defensive touchdown is an interception or fumble returned for a score; a return touchdown is a kickoff or punt returned for one, or a blocked punt or field goal. Until this date the return touchdowns were read from the feed's `pt_return_tds`, which sits with the punting figures and counts the punts a club returned AGAINST it: a defense was paid six points for a touchdown its own punt team allowed, the club that scored got nothing, and no kickoff return ever counted. Fumbles returned for a touchdown were not counted at all, because the feed keeps them out of its defensive touchdowns. Return touchdowns now come from `special_teams_tds`, which the feed builds from the same plays as its defensive touchdowns, split on whether the play was special teams, so the two never count one score twice. A fumble return is the defense's when a defender scored it or when the scorer recovered the other side's fumble, which is read from the player figures because the club's row cannot tell an opponent's fumble from its own offense recovering in the end zone. Checked against the play-by-play for 2020–2025, that matches every club's touchdowns in every week but one: an offensive player who recovered the other side's fumble after his own quarterback's interception (Titans, Week 5 2025), credited to the Titans' defense. 265 of the 3,516 weekly lines we held changed. **Scope: the week the claim names, or the season.** A claim naming a week ("Week 3 DST rankings", "the DST1 this week" beside a dated headline) is about that week's ordering and grades **once every game of that week is final** — a Sunday-afternoon grade would rank the Monday-night defenses last for not having played. A claim naming no week is about the season and waits for it, as every season-long claim does. The week is read from the payload first (the parser saw it in the headline) and from the claim's evidence second, the same rule depth-chart calls use. Defenses on a bye that week have no line and drop out of the pairing rather than being ranked last. **Ranked lists** grade by the same rank correlation as §2.3, over the clubs we could match to a line, re-ranked within the scored subset as always. **Baseline 50**, per §2.8: a coin-flip ordering already earns half the scale. **A single finish** grades on distance in places, credit reaching zero at 12 places — the §2.8 positional-finish curve, reused unchanged. **Baseline: the club's defense finishes where it finished last season**, "assume last year happens again", read off the prior season's summed weekly points; a club with no prior-season order (the first season we hold) is naively mid-pack, 16. **Difficulty is unchanged.** The multiplier reads the claim's date against the season, as it does for single-game picks under Amendment 7. A Week 3 ranking made in Week 3 therefore grades at roughly 1.8×, which overstates how much of the season a one-week call had left to be wrong in. Recorded rather than solved: the same is true of every game pick today, and a per-week clock is a change to Amendment 7 as well, not to this one. **What the public workings say.** A ranking grade lists each club the pundit ranked with its predicted and actual place, and — where the source printed them — its projected points allowed, sacks, turnovers and fantasy points beside what the defense actually posted that week. A reader can check every line against the FantasyPros table and the box score. A finish grade names the scope, the two places, and which naive read the baseline used. **What the parser and ingest changed, for the record.** Tables now reach the parser one row per line with cells separated by " | " and the header row on top, so a 32-row table is 32 rows rather than one run of numbers. The parser is told that in a defense table the "Team" column is the defense, "Vs." is the opponent and is recorded nowhere, and "Vs. QB" is the opposing quarterback and is not a prediction about him. A defense's club resolves with its unit suffix stripped ("Buccaneers D/ST", "TB DST", "Eagles defense" are the club), and publishers' codes we do not use ("JAC", "WSH") map to ours. ### Amendment 10 — revised articles count once, adopted 2026-09-24 Amendment 5 left one question open. A writer who publishes a mock draft as 1.0, then 2.0, 3.0 and 4.0, or numbers each week's 53-man roster projection, makes several public calls on one question and re-grades the same players each time. Until now each version counted as an article of its own, so a prolific updater collected four articles of evidence for one piece of work. | # | Question | Decision | | --- | ------------------------------------------ | ------------------------------------------------------------------- | | 1 | Which version of a revised article scores? | **Every version** is graded | | 2 | How many articles is it? | **One.** All the versions together count once in the record | | 3 | How much does each version weigh? | **Equally**, inside that one article, however many takes each holds | **Why every version.** Each version was a public call, published under the writer's name for readers to act on. Scoring only the latest would let a bad 1.0 be buried under a 4.0 written with a month more information; scoring only the first would discard calls the writer chose to make in public. **Why once.** The versions revise one judgment about one question; they are not independent observations of the writer, which is the same reasoning §3 applies to the takes inside one article. Counting each version as an article would let a writer climb the confidence tiers, and outweigh a peer who publishes once, by re-issuing the same piece. **Why equal weight.** It is how forecasting tournaments treat a revised forecast. Good Judgment Open, for one, scores a forecaster on a question by averaging their standing forecast over every day the question is open: every revision counts, and the question counts once. Tournaments weight each forecast by how long it stood. We weight each version equally, which needs no assumption about when a published article stopped being its writer's view, and a 1.0 with ninety takes gets no more say than a 4.0 with thirty. **What a revision is, and what it is not.** A revision is the same named piece re-issued as a new edition of itself, and the writer names the edition: a whole-number version ("Mock Draft 3.0", "Lock-O-Meter 4.0") or an "Updated" label carrying a date ("Updated March 1"), from the same publisher, under the same byline where one is on file, in the same calendar year. A bare "Updated" names no edition — "Updated X" beside "X", or published three times, does not say which edition each one is — so it merges nothing unless the title also carries a version or a date. A recurring series is not a revision. "Week 3 Power Rankings" and "Week 4 Power Rankings" answer two different questions and stay two articles, as do this year's and last year's annual lists, and so does any title that names a week anywhere in it: "Power Rankings 2.0: Week 3" and "Power Rankings 3.0: Week 4" are two weekly columns, not two versions of one, and neither is "Rankings (Updated after Week 3)". A brand that carries a number is not an edition either: the Pat McAfee Show titles every episode "PMS 2.0 ####", so a group needs at least two different editions. **The words after the version must match too.** "Fantasy Football Rankings 2.0: Quarterbacks" and "3.0: Running Backs" are two rankings, and "Playoff Power Rankings 2.0: Wild Card Round" and "3.0: Divisional Round" are two columns. Nothing in a title tells that kind of subtitle apart from an editorial one ("2.0: WR Has Significantly Different Outlook") or from a second segment of a podcast episode, so versions whose subtitles differ never merge, whatever the subtitle says. Versions that share a subtitle merge only with each other. Two kinds of trailing text are not subtitles. A page marker ("Picks 1–16", "Part 2", "Page 2", "(cont.)") says which page of one version this is. A podcast's episode number ("Ep. 1961") is assigned by the feed to every episode. **A title repeated on several URLs names no edition.** Ten weekly pieces all titled "Eagles 53-Man Roster Projection", published beside one "Eagles 53-Man Roster Projection 2.0", are ten articles, not one unnumbered original. The title can't say which of the ten the 2.0 revised. So an edition that two or more URLs carry under the same title (ignoring case and punctuation) is left out of the group. The pages of one edition have different titles, because their page markers differ, so they stay in. **Detection is deliberately conservative**, on the same trade as paginated articles. A false merge would understate a writer's output and nobody would see it happen, while a missed revision leaves the record exactly as it was before this amendment. So the piece's name is everything in the title before the version number, and a podcast episode whose title lists five topics only matches an episode that lists the same five. A first edition published with no number joins its later editions only when its title is theirs word for word, and no other URL carries that title. Some real revisions are missed on purpose: - Three share their episode title with other segments: Move the Sticks' "Alternating Mock Draft" 1.0 and 2.0 and "Daniel Jeremiah's Top 50 Prospects" 2.0 and 3.0, and First Things First's "Nick's Mock Draft" 1.0 and 3.0. - Three differ after the version number: Jeff Kerr's Eagles 53-Man Roster Projection 2.0 and 3.0 (Sports Illustrated), the FantasyPros 2026 NFL First Round Mock Draft 2.0 and 3.0 (carried on The Herd's feed), and Bucky Brooks' Mock Draft 1.0 and 3.0 (Move the Sticks). The Bucky Brooks 1.0 shows why the rule exists: its episode goes on to "Two Teams That Could Follow the Patriots Blueprint", a second segment with Daniel Jeremiah's takes in it. **When a revision is recognized.** Revision groups are assigned by a periodic grouping run (`pnpm db:series --apply`), not when an article is crawled. The run writes nothing unless it is given the number of groups the dry run listed, so a person has read every group before it counts. Until that run, a newly published version counts as an article of its own, exactly as every version did before this amendment; the next run folds it into its group, and the next recompute of the records reflects it. **What it changes, measured before the change.** On 2026-09-24 the dry run of `pnpm db:series` found no revision group among the 6,800 stored sources that meets every rule above. So today no pundit or outlet record moves, and no subject stance either. The first two dry runs that day came before the subtitle and repeated-title rules, and they grouped the three pairs listed above as differing after the version number. Applying those three groups would have moved four pundit records by one article each, none across a tier. It would also have removed 13 subject stances by taking them under the two-article floor: twelve of Bucky Brooks' and one of Derek Brown's. The dry run reproduced all 785 stored overall pundit records. It also reproduced all 1,870 stored pundit stance rows, once the takes approved since the nightly stance rebuild are left out. So its numbers are the recompute's own. **No grade changes** come from this amendment: a grade belongs to a take, every take in every version keeps its own, and only the roll-up moves. **Takes still count from every version.** A record's graded-take count, and the tiebreak §3 builds on it, include every version's takes: a longer run of calls still tells us more, even inside one article. **The weight is per version, not per page.** A version published across two pages ("Mock Draft 2.0: Picks 1–16" and "Picks 17–32") is still one version, and weighs the same as a 3.0 published on one page. An article that was never revised is a single version however many pages it runs to, so a paginated article is weighed exactly as it was before this amendment. **Every record weighs versions the same way.** Pundit and outlet records, Pundit of the Week and of the Month, and the season standings a milestone release cites all give each version one weight and count the article once. Before this amendment the awards and season standings averaged every take in an article at once, which let the version with the most takes outvote the others; for an article that was never revised they are unchanged. **Subject stances count a revised article once, and weigh its calls, not its versions.** A stance is where a pundit or outlet sits among the other voices on one team or player: the lean bar and the "Against the grain" lists. It uses the same article unit as the record, so once a revision is grouped its versions count as one article toward the stance's article count, its against-the-grain count and the two-article floor a stance needs before it is shown. Inside that one article, every comparable call about the subject (a mock-draft slot, a rank, a win total) counts once. Stances have always weighed the calls inside a single article that way. The versions are not weighted equally, as they are in the record. A version usually names a given player or team once, and then the two are the same; a version that names the subject twice counts twice. The dry run reports every stance a grouping would change before it is applied. **In code.** `revisionOf` and `revisionKeys` in `scripts/lib/seriesKey.mjs` apply the rules above and are tested against the real titles, including the ones that must not merge. `unitKeysFor` and `recordOf` in `scripts/lib/articleUnits.mjs` place each URL in its article and its version and give each version one weight; `lib/punditScores.ts`, `lib/punditAwards.ts`, `lib/milestonePosts.ts` and the dry run all call them. `lib/stances.ts` keys its article unit the same way, as `coalesce(series_key, source)`, and pools the calls inside it (`scripts/lib/stance.mjs`). A re-parse of a single URL that finds a changed answer still keeps both rows inside that one article (`supersedes_prediction_id`), as it did before. ### Amendment 11 — roster-spot claims settle once Week 1 is final, and a camp roster is not a roster, adopted 2026-09-25 Written 2026-09-24 and adopted on Ben's go-ahead the next day, the limit of 64 by name. The timing change is the one part that is policy rather than data validation: it moves the first roster-spot grades from days after the cutdown to the day after Week 1 ends, and it is the part to revisit if that proves too slow. §2.12 claims passed the season gate the moment a season's weekly rosters existed, on the reasoning that the file existing _is_ the cutdown having happened (Amendment 7 leans on the same reasoning). It is not. nflverse publishes a Week 1 file before the cutdown, holding each club's 90-man camp roster with every man marked active, and keeps revising that file until Week 1 is played. The first 2026 roster ingest (August 29 and 31) read the camp roster; the ingest never removed a row the file later dropped; and the 2026-09-07 grading run read both. **93 published grades were wrong the same way** — each said he had reached a club's 53 that, by the file as it stands, he has not: 46 "he makes the 53" scored 100, 38 "he does not make it" scored 0, and 9 "he lands on the practice squad" scored 0, across 34 pundits. 77 rested on camp-roster rows; 16 on Week 1 lines the file revised after the run graded them (men it now lists on injured reserve, on the practice squad, or cut). **What changes:** - Roster-spot claims settle from the weekly rosters once the season's **Week 1 is final**. By then Week 1's rosters are what each club carried into the season, not a snapshot still being revised. A confirmed "made the 53" report (§2.14) still settles a claim earlier, unchanged. For 2026 this changes nothing from here on: Week 1 is already final. - A roster file, or the table built from it, showing a club with more than **64** men on its 53 in one week is a camp roster, not the cutdown. It is neither ingested nor graded against. The number is measured: across every club-week of the 2020–2025 files the busiest was 59, and the table the camp file left behind still showed 57 to 72 per club. - A season's roster rows are exactly what its file produces. A row the file stops producing is removed on the next ingest. **The 93** go back to pending through a dated heal, once the table is reconciled, and the grader decides them again. As of 2026-09-24 every one is a man who has not reached that club's 53, so each waits under §2.12's mid-season rule: "he makes it" until he does or the season ends, the other two until the season ends. The other 1,110 graded 2026 roster-spot claims were run back through the grader against the reconciled table and come back unchanged. **And 53 older grades.** The postscript to Amendment 5 says §2.12's direction rule arrived before any roster-spot claim was graded. It did not quite: 53 were graded on 2026-08-19 by the first grader, a day after §2.12 adopted the claim's direction and its calendar year, and none was graded again. They record no direction, and 26 were scored against a season other than the claim's own year, so a "he does not make the 53" could score as though it said "he makes it". The same heal resets them. Run through today's grader, 7 come back unchanged, 7 rescore, 38 decline because their wording names no direction, and 1 — a February 2026 "he gets cut" — now waits on the 2026 season. **Injured reserve, ruled 2026-09-25 (Ben delegated the call).** §2.12 always said injured reserve "is not failing to" make the 53, but the grader counted only active and inactive weeks, so a man on injured reserve all season scored as never having reached it. Now a man the club carried **only** on its reserve list for the whole season, never on its 53 or its practice squad, declines: the club never made the roster decision the claim is about, whether he was hurt in camp or placed there after the cutdown and never activated. Whether he comes back is what decides it: a man who returns and reaches the 53 grades the week he does, and a week on the practice squad is a decision made. Eight of the 16 revised grades are "he makes the 53" about a man on injured reserve so far (four men on it every week, one placed on it and then released); they wait, as every unreached claim does mid-season, and this rule decides them in February if nothing changes. One graded claim from an earlier season had this shape — Joe Mixon, on the reserve list all of 2025 — and it now settles in 2026 under the calendar-year rule instead. ### Amendment 12 — roster calls: hedges, a reviewer's direction, and the sentence on the page, adopted 2026-10-01 Ben's ruling, the day 297 roster-spot claims sat declined because the grader could find no direction in their wording. §2.12 now carries the rule; this is the record of what it moved. **What changed.** - Hedged roster calls read as the claim they hedge: doubt about a spot is a cut, standing in a competition is making it. Conditions, merit judgments and claims about something other than the 53 still decline (§2.12). - A reviewer can set a roster call's direction from its card, one claim at a time or every unread claim in a take at once, and that call outranks the wording. - **The sentence on the page follows the grade.** Every roster call used to print as "to make the roster", the cut calls included, so a correct "he gets cut" that the season later proved wrong read as "said he would make it, and he did not" — the opposite of what the pundit said. A roster call now prints as "to make the roster", "to be cut" or "to land on the practice squad", read by the same function the grader uses. **What it moved, measured 2026-10-01 against the shared database.** The new reading changes no claim the old one already read (0 of 2,453 roster-spot claims). 83 claims gain a direction from their wording (56 makes, 26 cut, 1 practice squad), and reviewers answered 31 more on two Broncos 53-man projections. Run through the grader, 80 grade on the next run — 68 scored 100 and 12 scored 0, every 0 a cut call on a man who has since spent weeks on that club's 53 — and none disagrees with the league's weekly roster file. The rest wait, under §2.12's mid-season rule, until the man reaches the 53 or the season ends. 37 pundit records move; one confidence tier changes. ### Amendment 13 — ties in an ordering share their places, Ben 2026-10-02 **What changed.** Every ordering graded by rank correlation — power rankings and the other team and coach orderings, ranked player lists, and the ordering of a mock draft (§2.3, §2.6) — now gives tied items the average of the places they span, and correlates over those shared places. Before, the arithmetic ranked every item anyway and broke a tie by whichever entry happened to come first, which is a fact about how the claim was stored rather than anything the pundit said. For mock drafts it was worse than arbitrary: the picks were read in no fixed order, so one article naming three candidates "at number 8" published 25, 75 and 100 on different runs. An ordering whose items all share one place now claims no order and gets no score, as a single item never did. An ordering with no ties scores exactly as before; the test suite checks every arrangement of six items against the old formula. **Considered and declined.** Leaving tied picks out of a mock's ordering altogether (it throws away that they sit after the picks above them and before the picks below), and keeping the old arithmetic with a fixed tie-break (it makes the number repeatable without making it mean anything: one two-pick article would keep a zero, and another a perfect score, for an order neither claimed). **What it moved, measured 2026-10-02 against the shared database.** 5 of 581 graded mock articles, 27 graded picks, all public: three put picks at a shared number (two of them put every gradable pick at one number and lose their skill figure, the third moves from 90 to 81.62) and two name one player at two pick numbers (95.24 to 96.11, 100 to 97.43). Accuracy moves on none of them — it is per pick. Among the other orderings, three graded claims rank two entries at one place within a group the grader scores together. Those claims, and every graded ordering with a tie anywhere in it, are re-scored under the new rule. ### Amendment 14 — US Elections, approved by Ben 2026-10-02 The second topic's chapter. Drafted overnight on 2026-10-01 as `docs/missedtakes-elections-methodology-proposal.md` after Ben moderated the 2024 gold set (589 approved claims), and approved as proposed, every decision in its E6 included. The five election prediction types Ben ruled in on 2026-10-01 (chamber control, race winner, seat count, race margin, seats most likely to flip) are graded with §1's two numbers and §2's shapes: a pick that is right or wrong, a number measured by its distance from the result, an ordering. The arithmetic is `packages/missedtakes-db/scripts/lib/ electionGrade.mjs`, the run is `apps/missedtakes/src/lib/gradingElections.ts`, and both read as follows. **What settles an election claim: certified results, and nothing else.** Who won a race, by how much, and what each party holds in a chamber, from the Federal Election Commission's compilations (public domain), state canvasses as the public record carries them, and the Senate's and the House's own party-division records. Every result row names its source and every grade shows it. No betting market, prediction market or poll is a result or a baseline; §1's "not the company we keep" holds for elections too. A race settles in the round that decided it — Georgia's 2022 Senate runoff, Alaska's ranked-choice final, Louisiana's December runoffs — and an unopposed winner's margin is 100 points. Chamber counts count the caucus (independents with the party they sit with), which is how every forecaster states a Senate count; control is the majority caucus, or in a 50–50 Senate the Vice President's party. Election week settles winners through the reported-outcomes queue with a person's confirmation (`called`); the canvass replaces that (`certified`). A **winner** grades on a call; a **margin, share, seat count or electoral-vote total** grades only once certified. **The baseline: the holder keeps it.** The public fact everyone has is who holds the seat now, and the naive forecast is that nothing changes — the elections twin of "pick the home team". | Election type | Baseline | | ------------------------- | --------------------------------------------------------------------------------------------------------------- | | Race winner | The party holding the seat going in keeps it (a state's presidential vote: the party that carried it last time) | | Chamber control | The party in control going in keeps it | | Seat count / net seats | No change: the chamber as the previous general election left it (net 0) | | Race margin / vote share | The same seat's margin at its previous election, when we hold it; otherwise a tie | | Electoral votes | Half the college (269) | | Seats most likely to flip | A knows-nothing list (50), as every ordering in §2.3 is graded | A seat with no holder — a newly drawn district, a vacancy — gives the winner baseline nothing to go on, and it sits at 50. A holder-keeps-it baseline is weaker than a forecaster's model and easier to beat, so election MissedTakes Scores will skew positive; the number's job is to separate pundits from each other and expose the ones who only ever call the incumbent. An independent is not the party they sit with: "Democrats hold Vermont" graded against Sanders (I) is wrong as stated, and a reviewer who sees the pundit meant the person places the claim on the person. **Race winner (selection).** Named a person: 100 if that person won, else 0. Named only a party: 100 if the winner won under that party, else 0 (a fusion winner's party is their main line). A forecaster's rating toward a side ("Lean R", "Likely D", "a 70% chance") is a call for that side and grades as one; the rating's words are kept on the grade so confidence can be weighed later if Ben decides it should be. A toss-up names no winner and is not a claim. The national presidency is settled by the electoral college; a state's presidential vote by who carried the state. **Chamber control (binary).** 100 or 0 against the majority caucus. Baseline: control does not change. **Seat counts (numeric).** How many seats a party ends with, or how many it gains or loses against the previous general election's result. Full marks at the number, falling to zero at a stated distance on the chamber's own scale, because the chambers differ in size by a factor of four: **five Senate seats off is zero; twenty House seats off is zero.** "Republicans get to 51" against 53 is 60; "Republicans hold 225" against 220 is 75. A range in the pundit's words ("51 to 53", "a gain of 10 to 15", "at least 52") is right anywhere inside and measured from the nearest edge outside, as §2.4's win ranges are. Baseline: no change, on the same scale. **Race margins (numeric).** How much the claim's subject wins — or, negative, loses — one race by. A margin in points is the subject's lead over the runner-up in the deciding round, signed for the subject ("Harris wins Pennsylvania by 3" is +3 for Harris; she lost by 1.71, so her actual is −1.71 and the miss is 4.71 points); a vote share is the subject's share, 0–100. **Zero at 10 points off** for both — a three-point miss scores 70, roughly how far the 2024 polling average missed. Electoral votes, national race only: **zero at 100 off** (319 against 226 scores 7; a map that gets the swing states right lands in the 90s), baseline 269. A margin says who wins, so it is not also filed as a winner. Margins grade only once certified. **Seats most likely to flip (ordinal).** An ordered list of the seats most likely to change party, against which of them did (a seat flipped when the winner's party differs from the holder's). Each entry weighs 1/rank, so the top of the list is what the call is about: a list whose first three seats all flipped and whose last seven held scores higher than the reverse. Entries we could not place on a race leave the denominator and are reported on the grade. Baseline 50; the 2024 gold set holds one such list, and the number is revisited when there are more. **Difficulty: days to Election Day.** §2's multiplier asks how much season was left to be wrong in. For elections the clock is days to Election Day over the two-year election calendar: 1 + days / 730, capped at 2.0 — a call two years out is 2.0×, a year out 1.5×, the night before 1.0×. It scales the MissedTakes Score only. **A claim dated after Election Day is not graded:** through the end of Election Day in the latest US time zone a take is a forecast; after that it describes a result, however long the count takes, and the row is marked so. **Declined, and enforced in code.** The parser's rules of 2026-10-01 carried into the grader where a backstop is needed, on §2.13's principle: a poll or a relayed forecast is not a claim; a claim about an election's legitimacy is never graded; a conditional is not graded; a candidate, campaign, party officer or sitting official on their own election is not a pundit; primaries, nominations, turnout and approval are not held types; a negated claim is not recorded unless it pins a positive one; a claim nobody could place is ours and stays off the record, never charged to the pundit. **Decided with the chapter (Ben, 2026-10-02):** the holder-keeps-it baseline over a published partisan lean; the scales above; ratings unweighted for confidence; 1/rank weights and a baseline of 50 for flip lists; claims after Election Day declined rather than graded at 1.0×; Senate caucus counts; the 2020 state result as the holder, and the 2020 margin as the margin baseline, for the 2024 presidency by state. Each is one constant or one sentence in `electionGrade.mjs`, with a test; a change to any of them is a regrade. **Where it stands on adoption.** The 2024 gold set, graded read-only against the certified files before this text was signed: 572 of 595 claims gradable; race winners 354 of 466 right (76%) against a baseline of 33; chamber control 42 of 53 against 30; margins 77 against 55; seat counts 80 against 50. US Elections stays off every public page until its hub launches, so the first grades under this amendment are visible to the admin alone. ### Amendment 15 — one starter only at quarterback, center, kicker and punter, Ben 2026-10-02 **What changed.** A week-scoped "he starts" call (§2.14) used to grade **0** whenever a confirmed report named a different man to the same position family, and a later report naming a different man superseded an earlier one. Both rested on one assumption: a position has one starter. That is true of a quarterback, a center, a kicker and a punter. It is not true of a safety, a cornerback, a receiver, a linebacker, a tackle, an edge rusher, a running back or a tight end, where a club starts two or more and naming one says nothing about the other. The family test was also an overlap, so an edge rusher and an "off-ball linebacker" counted as one job because both families include the outside linebacker. Now a report about a different man settles a claim, and supersedes an earlier report, only when both name the same one of those four jobs. Anywhere else only a report about the man himself settles his claim. Kicker and punter joined the list the same day, on Ben's second look at special teams; the long snapper is not on it. A report that another man did **not** start never settles anything, at any position: it does not say who did. **Considered and declined.** Tracking slots instead of positions (left tackle, slot corner, "the No. 1 receiver"). Neither the articles nor the claims name slots reliably, and "No. 1 receiver" is a role in the offense, not a starting job. **What it moved, measured 2026-10-02 against the shared database.** Of 51 depth-chart claims a report had settled, 5 were graded 0 because another man was named. One, a quarterback, stands. Four, all public, go back to pending: a defensive tackle graded wrong because the other defensive tackle was named, a receiver (twice, two pundits) because another receiver was named to start "alongside" him, and an edge rusher because an off-ball linebacker was named. No member picks rode on them. None of them has a report about the man himself, so each waits unless the snap counts show he never took the field. One superseded report, a running back's, returns to pending, since two reports naming two running backs can both be true. No other report was superseded under the old rule outside quarterback. Adding kicker and punter moved nothing: no starter report on file names either job, and no open claim is waiting on one. ### Amendment 16 — a record never blends topics, adopted 2026-10-02 **What changed.** §3 says a pundit's headline numbers are computed "overall, per sport, and per prediction type", and until now the site published the overall one. With US Elections graded beside the NFL (Amendment 14), an overall across the two means nothing: they are measured against different baselines — "the home team wins" against "the holder keeps the seat" — and §3 already refuses to average across types for the same reason. So every record the site publishes is now a **topic** record: each pundit and each outlet carries one score row per topic (`topic`/`nfl`, `topic`/`elections`), the leaderboards are per topic (`/nfl/leaderboard`, with the old `/leaderboard` address redirecting to it; `/elections/leaderboard` when that topic is public), awards and the §3 provisional thresholds count articles within a topic, and a pundit page shows one panel per topic they have a record in, side by side. The cross-topic `overall` row is still computed and is internal: no public page reads it. Nothing moves for anyone today — with one public topic the NFL record and the old overall are the same numbers. Records are computed for every approved topic whether or not it is public; a topic's record is published by whether its scope is read, which is what lets the elections preview show records before the hub launches. ### Amendment 17 — the House popular vote, and ratings tallies read as floors, approved by Ben 2026-10-03 **What changed.** The fall 2026 queue held nine claims nobody could place: a forecaster's ratings tally with no measure, four times; four calls on the national House popular vote, filed as race margins with no race; and a margin filed under an office its candidate is not running for. Ben decided all three shapes on 2026-10-03. Reading those claims also turned up a fault Amendment 14 had already ruled on: the direction of a range written in words was never read. **House Popular Vote Margins, a sixth election type.** How much a party wins — or, negative, loses — the national popular vote for the House by: the votes of every House race added up. A margin is the party's share of every vote cast for the House less the other major party's share; a vote share is the party's share alone. **It is settled from the Clerk of the House's certified _Statistics_ of each election and nothing else** — the two major parties' votes and the Total from its table "Recapitulation of Votes Cast for United States Representatives", copied verbatim and cited to the page. The Total is the denominator, and party lines count where the Clerk puts them. **Zero at 10 points off**, the race margin's scale. **Baseline: the previous general election's national House margin** (or share), signed for the claim's party. Only the two major parties are held. The Clerk publishes about four months after Election Day, so a cycle's claims grade then; until then the wait is ours, never the pundit's. Winning the national vote is not winning the House, so the claim is not also filed as chamber control or a seat count; a generic-ballot poll is a measurement, not a claim; and the presidency's national popular vote is not this type. A claim is filed here only when its words name the national vote, and one that names a single seat waits for a person. | Claims about | Baseline | The Clerk's figures for that election (D / R / Total) | | ------------ | ----------------------- | ----------------------------------------------------- | | 2022 | 2020: Democrats +3.03 | 77,122,690 / 72,466,576 / 153,431,405 | | 2024 | 2022: Republicans +2.72 | 51,280,463 / 54,227,992 / 108,443,387 | | 2026 | 2024: Republicans +2.55 | 70,571,330 / 74,390,864 / 149,543,421 | "Around 3 to 5 points" for the Democrats in 2024, when they lost the House vote by 2.55: the nearest edge is 5.55 points off, so 44.5; the baseline, 0.17 off, scores 98.3. **A ratings tally is a seat floor.** "House: 213 at least Leans D, 202 at least Leans R, 20 Toss-ups" says each party wins at least that many seats. "213 at least Leans D" is "Democrats win at least 213 House seats", graded by Amendment 14's seat-count range rule: right anywhere from 213 up, measured from 213 below. **Toss-ups belong to no one and are not split**, and a tally is not a chamber-control call. A tally is a number, "at least", and a rating (Lean, Likely, Safe, Tilt); one stored as a single number grades as the floor its words state. A count that only reads like a tally in other words waits for a person to say which it is. **A number's direction is read from its words — a fix, not a rule change.** Amendment 14 says a margin is negative when its subject loses and a net change of seats is signed for the party, and a single number has always been stored that way. A range is kept in the pundit's words, and until now those words were read for size only, so "probably gonna lose, but definitely double digits" graded as a win of ten or more. On margins and net seat changes the words now give the sign: a loss ("loses by a point or two" is −2 to −1), a party's sign ("R+2-6", "in favor of Republicans", plus for that party and minus for the other), and a span from one party's gain to the other's (−6 to +9 for the Democrats in "a Democratic gain of nine to a Republican gain of six"). Seat totals, vote shares and electoral votes are sizes and are never signed. **When the words could be read two ways, nothing is read and a person decides** — a negated or hedged call ("I'd be surprised if", "without losing more than one"), a conditional loss, a loss beside a gain or another election, a seat lost that the party holds (not its net), words that can mean a loss or a smaller win ("down", "behind", "comes within"), a party's plus that is context rather than the call ("this R+12 state by 8 to 10"), a party that is not the claim's, a split of the vote ("Republicans 51-48") or two parties' numbers, a range beside a floor or a ceiling ("52 or 53, maybe up to 55"), or "double digits" qualified ("under double digits"). Spoken numbers read as numbers ("a point and a half" is 1.5, "'25 or '30" is 25 to 30); ranges joined only by "or", or a comma between next-door numbers, are one window; and a ceiling runs from zero, so "wins by less than ten" no longer counts a loss as inside the call. Every grade's working shows the window as graded and how it was read. **A margin in points on the presidency waits for the popular vote.** The national presidency is settled in the electoral college, and its certified margin is in electoral votes (86 in 2024). A points margin on the national race is a call on the popular vote, which no result we hold settles, so it stays ungraded — ours to load — rather than being scored against 86. **A claim filed under the wrong office moves to the race its person stands in.** When a race claim names someone not on the placed race's ballot, and that person's full name is on exactly one other ballot in the same state that cycle, the claim moves there with that person and their party, and the review card says where it came from. Never on part of a name or a near miss, across states, from the national presidency, when the person stands in two races, or when the claim names a different party: those wait for a reviewer. **Considered and declined.** Reading a tally as the count the forecaster expects, or splitting the toss-ups. A generic-ballot average as the popular vote's baseline (no poll is a baseline). Minor parties' national vote, or any compilation other than the Clerk's. Grading a presidential points margin against a popular vote summed from the states (a new measure, for its own amendment). Guessing a direction the words leave open. Moving a claim on a surname or a near miss. **What it moved, measured 2026-10-03 against the shared database.** Of 108 race-margin and seat-count claims on file, 52 were graded. Two graded claims read differently and are regraded: Mark Robinson's 2024 governor's race (0 before, about 100 as a loss of ten or more against his 14.8-point loss) and Trump in Iowa, "somewhere between five and eight or eight and ten points" (5 to 8 at 47.9, now 5 to 10 at about 67.9). Seven approved 2024 claims the old reader returned nothing for now grade. Of the six it still cannot read, Ben had five set by hand and ruled the sixth ("will get crushed") too vague to grade. Four graded 2024 calls on Nebraska's 2nd district's elector, filed on the statewide race and scored 0 against Trump's statewide win, come off the record until a district race is held. Of the nine flagged claims, the four tallies get their measure, the four national calls become House Popular Vote Margins, and "Doug Jones loses by a point or two", which the rule would have moved from Alabama's Senate race to its governor's race, was rejected by Ben as a scenario rather than a call. Two 2026 seat counts that read like a tally, and two "lose at least one Senate seat it currently holds", wait for a person before they grade. ### Amendment 18 — contract calls, Ben 2026-10-07 **What changed.** "I think we should include contracts." A claim about a player's next deal was filed as a movement and declined for want of contract data — and five of them, by the order the movement grader read things in, were published with a MOVE's grade: "Dallas may franchise-tag him" scored 100 because he had been traded to Dallas, "he'll reset the market" scored 0 as "he leaves". Those five grades were pulled the same day; the claims return as contract calls and grade under §2.16. #### §2.16 Contract calls **On hold since 2026-10-09 (Amendment 23).** No contract call is graded until we have a contract source we may use. The rules below stand for when one is in place. **Settled by OverTheCap's contract list**, through nflverse's daily release: every deal OTC lists, with its club, years, total, average per year (APY) and guarantees. The list carries no signing date, so **the day a deal first appears in our copy is its date** (refreshed three times a day). Deals already listed on 2026-10-07 are known only by the league year they were signed in (mid-March to mid-March); one of those is dated to the day when a confirmed report (§2.14) says he signed or was extended inside that year and it is his only deal that year. When a deal can only be placed by its league year and that year straddles the claim's window, **the claim is not graded** — a deal on the wrong side of a claim is the whole question. **What the claim says is read from its own words**, never from a field: - **He signs or is extended** — "the Bills will pay him". Right if a deal is signed after the claim and by its deadline; with the club, when one is named. **100 or 0.** - **He plays out his deal** — "plays it out without an extension", "the Browns will not re-sign him". Right if no deal is signed by the deadline. **100 or 0.** - **A figure per year** — "$55 million a year or more" (a floor), "around $28–30 million annually" (a range), "about $60 million a year" (the figure ±10%). Right if the first deal in the window lands inside. A deal's total is not an APY; a claim naming only a total waits for a person. **100 or 0.** - **A place in the market** — "the highest-paid tight end", "resets the quarterback market" (first), "a top-5 paid DT" (top five). His deal's APY is ranked against every other player's richest deal at his position that was signed no later and not yet expired; guards and tackles are one market each. **100 or 0.** - **A minimum deal** — right if the deal's APY is at most $1.75 million (the 2026 veteran minimum runs about $0.9–1.3 million by experience). - **The franchise tag** — **not graded**: OTC's list does not mark tags. - **A condition** ("if he wins the starting job") — **not graded**, the rule Amendment 12 set for roster calls. Nor is a label whose clauses say different things ("negotiated during the season; not signed by Week 1"): a person reads it. **The deadline is the pundit's when they name one** — "before Week 1", "by September 11", "at their Week 10 bye", "during training camp" (Week 1 is taken as September 10). A claim naming the draft or the trade deadline waits for a person. **When they name none** (Ben, 2026-10-07): "he signs", a figure, a market place or a minimum waits for **the next league year (March 12) at least 90 days away**; "he plays it out" waits for **the end of the season it was said in (February 15)**. A deal in the window settles the claim the day it appears, either way; no deal settles it at the deadline. **Baseline: the opposite call.** Binary, like every movement call: a "he signs" claim's naive alternative is "nothing happens", and the reverse. **Difficulty 1.0**: a contract is one event, not a season. ### Amendment 19 — "he stays" is a claim, Ben 2026-10-07 **What changed.** Ben asked that estimates "that a player will stay with their team" be graded. §2.11 had declined them because, graded against a coin, a stay pays the naive answer. Against the measured base rate it does not. **A stay is right when he stayed put**: on the club he was with the season before (or the one the claim names, when he was with it), and on no other club, for the season the claim settles in. A mid-season trade away makes it wrong, though he appears on his old club's roster too. **100 or 0.** "He won't be traded", "not going anywhere", "won't make it to free agency" read as stays. **Baseline: 55, the stay base rate.** Of every player on an NFL roster in a season, **52–55%** are with the same club the next season (2021–2026, the grader's own season rosters). That counts players who leave the league, which a stay call must: a player out of football did not stay. The 65% above counts only players still in the league the next season, and is the right figure for the question it answers there. So a right stay earns **+45** skill and a wrong one **−55**; difficulty is unchanged. **Two published grades moved**: "Tee Higgins will not make it to free agency" and "Trey Smith will not make it to free agency" were graded 0 as departures — "free agency" read as a departure word. They are re-graded as stays. ### Amendment 20 — a departure is graded by the way it names, Ben 2026-10-07 **What changed.** "I'd grade the stated mechanism. The team could release a player early or the player contract could result in them heading to free agency, which are different things." A movement claim that says HOW he leaves — **traded**, **released** (cut, waived), or **out in free agency** (not re-signed, his contract runs out) — is now right only when he left that way. - **He left the way the claim said** — a confirmed report (§2.14) of the trade or the release. **100.** - **He left another way** — a release where the claim said traded. **0.** - **He left, and nothing we hold says how — not graded.** A signing elsewhere with no release or trade reported could be either, and a guess would be a grade the claim did not earn. It waits for a report. (Measured before the change: of 13 such claims graded right, 9 held on a confirmed report and 4 became ungraded; none turned wrong. Ben approved.) - **He stayed** — wrong, however the claim said he would go. **0.** - A claim naming no mechanism ("finds a new home") or two ("cut or traded") is graded as before: any departure settles it. **Read from the claim's own words**, never the quote around it. "A no-trade clause" names no trade. "Stays with the Steelers rather than being traded" is a stay (Amendment 19), not a trade — it had been published as a correct departure because the word "traded" appeared in it, and is re-graded. A contract claim graded as a move is re-filed as a contract call (§2.16) with its grade cleared. ### Amendment 21 — hedging, counted and published, Ben 2026-10-07 **What changed.** "We're seeing a lot of hedging." Every claim now records how firmly it was made, and two numbers are published from that: how often a pundit hedges, and whether their hedged calls land more or less often than their plain ones. **Neither touches accuracy or the MissedTakes Score.** A hedged claim is still graded as the claim it wraps (the rule since Amendment 12); hedging is reported beside the record, as clarity is. **What a hedge is.** The words governing that one claim, copied verbatim from the quote: a **lean** ("probably", "likely", "I expect"), a **possibility** ("could", "might"), a **long shot** ("wouldn't shock me"), a **toss-up** ("50-50"), or a **condition** ("if he stays healthy"). "I think" and "I believe" are how an opinion is said, not a hedge; a model's or a market's probability is the claim itself, not a hedge. **How it is read.** After a claim is filed, by a separate reader that sees the claim and its quote — never by the parser, which was measured doing it and made worse extractions (it filed more hedged speculation, including from sources a person had judged empty). The reader's hedge words must appear in the quote or the tag is dropped. A claim not yet read is in neither count. **Hedge rate**: the share of a pundit's read claims that carried a hedge, averaged by article like clarity, shown from five read claims. The ranking at /nfl/hedging lists pundits with at least 30. **Right when hedging, right when committing**: their article-weighted accuracy on graded hedged calls and on graded plain calls, shown only when they have at least **10** graded hedged calls and **20** graded plain ones. A gap of five points or more either way is described ("the hedges undersell them" / "when they hedge, the doubt is usually real"); a smaller one reads "about as right either way". Both thresholds are chosen constants, not statistics. ### Amendment 22 — the Super Bowl MVP, the Pro Bowl and All-Pro lists, and a call made after its list, Ben 2026-10-09 **What changed.** Ben, 2026-10-09: model the Super Bowl MVP as an award, and load the Pro Bowl and All-Pro lists. Four rules come with that. **The Super Bowl MVP is an award** (§2.6.5): one winner, so 100 or 0, and the baseline is the previous Super Bowl's MVP repeating. It belongs to the season whose Super Bowl it was — Super Bowl LX's MVP is the 2025 season's. Held from Super Bowl LV (the 2020 season) on. **A Pro Bowl call is right when the NFL named him to the roster.** That is the initial roster — including players who declined, were hurt, or were playing in the Super Bowl — and the players named to replace them: the NFL's own count of Pro Bowlers. Each stored row says which it was, so the rule can be narrowed without re-entering anything. **An All-Pro call is the AP's team**, ties included; a bare "All-Pro" means the first team, as it always has. Held for the 2024 and 2025 seasons. **A call made after its list was public is about the next one.** Every list records the day it was published: the AP's All-Pro story, the NFL's Pro Bowl roster release, the Super Bowl itself. The season rule (§1) files a February claim under the season just finished, so "he's a Pro Bowl tight end", said on 27 February, would be graded against a roster announced in December — a result the pundit could read. Such a claim is about the next season's list, the reading §2.1.1 gives a game pick: a prediction is about something that has not happened yet. The publication day itself counts as before it, so a game-day Super Bowl MVP pick is a pick. **A Pro Bowl or All-Pro call about a career, a ceiling or a level is not graded.** "At some point in his career", "all-pro potential — not this year", "a perennial Pro Bowler", "a Pro Bowl level player": one season's list cannot settle any of them, and graded against one each would publish a 0 nobody earned — the single-game mistake in reverse. The claim's own words decide, never its outcome, and the rule errs toward not grading: "an all-pro caliber player, and I think he finally gets his nod this year" is left ungraded too. **Sources.** Each list was checked against two before it was loaded: the AP's story and Wikipedia for All-Pro; NFL.com's roster pages and Wikipedia for the Pro Bowl; Wikipedia's list for the Super Bowl MVP, with the settled prediction markets agreeing on Super Bowls LIX and LX. Every row names the page it was read from. ### Amendment 23 — contract calls on hold, Ben 2026-10-09 **What changed.** §2.16 settled contract calls against OverTheCap's contract list, which reached us through nflverse's daily copy. OverTheCap's terms reserve commercial use of its data for those it has consented to in writing, and nflverse's license cannot grant what nflverse does not hold. So until we have a contract source we may use, **contract calls are not graded**: - The nine contract calls graded under §2.16 are returned to pending, and every pundit record is recomputed without them. - Every contract call waits as pending. None is declined and none is scored, the rule for any claim whose outcome we cannot yet read. - We keep no copy of OverTheCap's list, and Picks does not offer contract calls while they cannot settle. The rules in §2.16 stand. Grading resumes with a source we may use, and an amendment that names it. ### Amendment 24 — blocked kicks are scored, Ben 2026-10-10 **What changed.** A defense's blocked punts, field goals and extra points now score 2 points each, as they do on ESPN, Yahoo, NFL.com and Sleeper (§2.15). A block returned for a touchdown pays both: 2 for the block and 6 for the return. §2.15 had left blocks out, saying the free team feed did not carry them. It does: the feed records each block for the team that made it, and in every week of 2020–2026 that count equals the kicking team's count of kicks it had blocked. Ben, 2026-10-10: "if it's data we can pull, let's do it." **What it moved, measured 2026-10-10 against the shared database.** 257 blocks across 248 of the 3,516 weekly lines we held. Taken together with the touchdown fix in §2.15 the same day, 467 lines changed, and every graded fantasy-defense claim is re-scored under both: six of the seven published grades move. The largest is Evan Silva's 2020 call of the Ravens as the DST3, from 58.33 to 83.33: on the corrected lines the Ravens finished fifth, not eighth. ### Amendment 25 — draft results from the NFL's own records, Ben 2026-10-09 **What changed.** Draft results came from nflverse's copy of Pro Football Reference's draft pages. They now come from the NFL's own draft records, the draft columns of nflverse's players file. Every pick the two share names the same player, round and club. One correction came with it: 2019's 11th pick had been linked to the wrong Jonah Williams, and the two claims about either man were graded again. **One field still comes from Pro Football Reference: the position a player was drafted as.** The NFL's record gives only the position he plays now, which is a different position group for three picks in ten, and it files edge rushers with linebackers. §2.6.4 orders players within the position they were drafted at, so it keeps Pro Football Reference's. It is used to grade and nowhere else. **A draft whose record is missing a pick cannot say who went undrafted, except about a player the NFL lists.** The NFL's records lack the few picks whose players its player file does not list (13 across the drafts we hold, among them Khyree Jackson, 2024's 108th pick). A player the file lists, and no draft took, went undrafted, and scores as before. A player it does not list may be the missing pick, so while his draft is missing one, a claim that would score him "undrafted" waits instead of scoring 0. Mock drafts naming one of the 13 were graded again, since a mock's ordering is measured over the players who can be placed. ### Amendment 26 — the betting line beside the take, Ben 2026-10-09 **What changed.** An NFL game's closing betting line is shown beside the game picks and score predictions made on it, as a dated fact. **It never touches a grade**: a winner pick is still graded on who won and a score prediction on its margin (§2.1, §2.2), never on the spread, and no line enters an accuracy score, a skill delta, a baseline or the MissedTakes Score. A line a pundit quotes is still a price, not a claim (§2.13), so a betting line is never a take we grade. **Where it comes from.** The free nflverse schedules file, which carries one sportsbook's line for every game. We credit nflverse (see Data sources), name no sportsbook and link to none. We read the file every hour and keep each change with the time we saw it; the closing line is the line at kickoff. **When it shows.** Only beside a graded call — so after the game is final — only for a pick made in the seven days before kickoff (an earlier pick was made before most of the news the line reflects), and only where the game has a line on record. No line shows before a call is graded, beside a call we decline to grade, for a game nobody called, or anywhere in Picks. **How it is read.** - **Chance to win.** Each side's chance comes from the two moneylines, with the bookmaker's margin removed in proportion to each price. That slightly overstates long shots, so past one in ten we print "under 1 in 10" (or "over 9 in 10") rather than a figure. Where the moneylines are missing we print the spread alone and never estimate a chance from it. - **Favorite, underdog or toss-up.** The side the spread favors is the favorite. A game whose spread is 1½ points or less is a toss-up, and no pick in it counts as a favorite or underdog pick; where the spread and the moneylines disagree inside that band we call it about even. - **Score predictions are read as numbers** — a closer or bigger win, more or fewer points — and never called a favorite or underdog pick, because many score calls are spread calls in the pundit's own words. - **The spread in their words.** When the quote talks in spread terms ("seven and a half points", "the Giants plus seven"), the line adds that the call is graded on who won (or on the margin called), not on the spread. We print no odds in betting notation and say nothing about bets, units or profit. MissedTakes takes no bets and earns nothing from any sportsbook, and nothing here is betting advice. The line can be switched off site-wide.