A Cy Young predictor usually treats the electorate as though it has one mind. Give every pitcher a score, sort the scores, and see whether the formula identifies the winner.
That can work surprisingly well. It is also not how the award is decided.
There are 30 voters in each league. Each of them submits a five-player ranking. One voter may still care about pitcher wins and team success. Another may pay much more attention to FIP. A third may not consciously use either formula, but may end up producing a ballot that looks a lot like one of them.
Tom Tango made the better question explicit when I showed him some earlier work on the changing Cy Young vote:
We are talking about 30 individual voters. How many follow the Bill James predictor? How many follow my original Cy predictor? How many follow the FIP-based Cy predictor? How many follow the latest predictor?
So I went through the individual ballots.
Turning 840 ballots into a voter model
The Baseball Writers’ Association of America has disclosed every individual Cy Young ballot since 2012. That gives us 840 complete ballots across 28 American League and National League races through 2025, with 4,200 ranked choices in total.
I compared each five-player ballot with five different ways of ranking that season’s pitchers:
| Voter profile | What it tends to reward |
|---|---|
| Bill James/Neyer | Wins, losses, innings, team success, strikeouts, shutouts and saves |
| Tango original | Run prevention, innings, strikeouts and wins |
| FIP-adjusted Tango | Tango’s original approach with additional weight on FIP-based performance |
| Pure ERA+FIP | The lowest combined ERA and FIP among pitchers who meet the innings threshold |
| Learned hybrid | A fitted combination of ERA, FIP, workload, wins and strikeouts |
The learned model was always trained only on earlier seasons. A ballot from 2024, for example, was not used to teach the model that evaluated the 2024 ballot.
I also compared each ballot against the full objective candidate pool, not only the pitchers who received votes. Otherwise, a formula would get too much credit simply for sorting five names we already knew voters liked.
For every ballot, the model asks how likely each formula would have been to produce that complete ranking. It then divides the ballot’s weight among the five profiles. If a ballot looks equally consistent with the original and FIP-adjusted Tango models, each can receive partial credit.
This is important. The results are measures of similarity, not proof that a particular writer opened a spreadsheet and used a particular equation.
What the individual ballots show
The clearest movement is away from the older Bill James/Neyer profile and toward FIP-aware voting.
The shortened 2020 season is excluded from these period comparisons.
If we translate the recent percentages into one hypothetical 30-voter league electorate, the mix is approximately:
- 11 voters’ worth of FIP-adjusted Tango similarity
- 9 voters’ worth of original Tango similarity
- 7 voters’ worth of learned-hybrid similarity
- 1 voter’s worth of pure ERA+FIP similarity
- 1 voter’s worth of Bill James/Neyer similarity
Rounding prevents those figures from adding perfectly to 30. More importantly, these are not 30 hard labels. I cannot identify eleven writers and declare that each follows the FIP-adjusted formula. The model is saying that the combined similarity across the 30 ballots is equivalent to about eleven fully FIP-adjusted voters.
This looks more gradual than a single paradigm shift
Corbin Burnes winning in 2021 is a natural marker for a FIP-heavy era. His 2.43 ERA was excellent, but it was not the lowest in the National League. His 1.63 FIP was extraordinary. A traditional winner-based model would not have expected a pitcher with 11 wins to beat Zack Wheeler and his 14 wins.
The individual ballots suggest that 2021 was a visible result of a movement already underway, not the day the electorate suddenly changed its mind. FIP-adjusted similarity had already risen from 23.3% in 2012-16 to 34.0% in 2017-19. The Bill James/Neyer share had already fallen by more than half.
The recent electorate also does not look uniformly FIP-driven. The original Tango model still accounts for about 31% of ballot similarity, while the learned hybrid accounts for another 25%. Pure ERA+FIP describes only a small share of complete ballots, even though it may explain the winner in a particular race.
That last distinction matters. A formula can correctly pick the winner without doing a good job of reproducing the other four names or their order.
The same writers changed too
One obvious explanation is voter turnover. Perhaps older voters left the pool and more analytically inclined writers replaced them.
That happened to some extent, but it is not the entire explanation. Fifty-three writers appeared in both the 2012-16 and 2021-25 samples. Their own ballots moved in the same direction:
The composition of the electorate changed, but returning voters changed their behaviour as well. The movement toward FIP awareness was not simply one generation replacing another.
Which formulas actually reproduced the ballots?
There is no single winner across every measure. The learned hybrid most often identified a voter’s first-place choice, while Tango’s original model produced the best overall probability for the complete five-player ranking.
| Formula | First-place match | Pairwise ballot order | Top-five overlap |
|---|---|---|---|
| Learned hybrid | 79.2% | 81.9% | 78.3% |
| Tango original | 76.7% | 82.9% | 80.3% |
| FIP-adjusted Tango | 69.5% | 80.2% | 77.0% |
| Pure ERA+FIP | 51.0% | 75.2% | 75.6% |
| Bill James/Neyer | 40.3% | 70.1% | 57.3% |
"Pairwise ballot order" asks whether a formula agreed each time the voter ranked one of the five chosen pitchers above another. "Top-five overlap" asks how many of the formula’s five highest-ranked pitchers appeared anywhere on the ballot.
The original Tango predictor remains remarkably good. The evidence does not show that voters discarded it and collectively switched to one FIP formula. It shows a mixed electorate whose centre of gravity has moved toward FIP.
What this does not prove
There are several reasons to treat the exact percentages as estimates rather than a census of voter beliefs.
First, similar formulas are difficult to separate. The original Tango model and the learned hybrid had a 0.968 average rank correlation and selected the same top pitcher in almost 89% of the league-seasons. A five-name ballot often cannot tell us which of two closely related ideas a voter preferred.
Second, the answer depends on the five profiles included in the comparison. A sixth useful formula could absorb weight from the existing five. These are the most relevant profiles for the question Tom posed, not every possible way a writer might evaluate a pitcher.
Third, the public individual-ballot history begins in 2012. This analysis can describe the transition within that window, but it cannot directly classify the voters from the earlier Bill James era or the beginning of the original Tango era.
Fourth, the voter pool is not fixed. Writers rotate in and out, and a returning writer can change how they evaluate pitchers. Using the 2021-25 mix to predict the next election assumes that recent voting behaviour is a useful starting point. It should not be treated as a permanent 11-9-7-1-1 roster.
Finally, the uncertainty intervals measure sampling variation within the ballots we observed. They do not capture every modelling choice, data error or philosophy missing from the five-profile menu.
What I think the evidence supports
Cy Young voting has changed, but it does not look like one hard paradigm instantly replacing another. The older wins-and-team-success profile has nearly disappeared. FIP-aware rankings have become much more common. Tango’s original formula remains highly relevant, and the best description of recent voting is still a combination of approaches.
That changes how I think a Cy Young probability model should work. Averaging every voter into one formula creates a person who may not exist. A better model simulates 30 individual ballots, allows those voters to resemble different philosophies, and admits uncertainty about the mix itself.
The next step is to use these historical profiles as a starting distribution for a 30-voter simulation, then allow both pitcher performance and voter behaviour to vary. That still will not tell us exactly how this year’s writers think. It should get us closer than pretending the electorate has one mind.
Methodology and sources
Individual ballots were collected from the BBWAA’s published Cy Young result pages for every AL and NL race from 2012 through 2025. The dataset contains 840 complete voter ballots, 4,200 ranked choices and 368 unique writers. The five profiles are the Bill James/Neyer predictor, Tango’s original Cy Young Points, Tango’s FIP-adjusted version, Tango’s latest ERA+FIP approach, and an expanding-window learned model based on ERA, FIP, workload, wins and residual strikeout value. Tom’s discussion of the evolving formulas and the possible post-Burnes shift is available in Describing the Cy Young voter mindset.
Candidate statistics were standardized within each league-season before calculating ranked-choice probabilities. Each formula was allowed its own fitted level of ballot randomness. Trend estimates use equal-prior soft similarity weights for each ballot, and uncertainty ranges were estimated by resampling complete ballots. The main period comparisons exclude 2020 because the 60-game season changed normal workload and qualification patterns.