Corbin Burnes won the 2021 National League Cy Young Award with 151 points. Zack Wheeler finished second with 141.
That makes it sound like Burnes won a close but otherwise conventional race. The individual ballots tell a more interesting story. Burnes and Wheeler each received exactly 12 first-place votes. Neither convinced even half of the 30 voters that he was the best pitcher in the league.
Burnes won because more voters agreed that he was at least very close to the best.
The race is also useful because it sits directly inside the larger change in Cy Young voting. Wheeler offered the traditional workhorse case. Burnes offered fewer innings with almost unprecedented dominance in the outcomes most directly controlled by a pitcher. The electorate split almost perfectly over which form of value mattered more.
Burnes’ ERA led MLB. His FIP made history
Burnes won the NL ERA title with a 2.43 ERA, which was also the lowest qualified mark in MLB.
He narrowly finished ahead of Max Scherzer at 2.46 and Walker Buehler at 2.47. It was an excellent result in a season with several excellent run-prevention performances, but it was not historically unique.
Looking at every qualified pitcher-season from 1967 through 2025, Burnes’ ERA ranks 76th after adjusting for the league run environment. That places it around the top two percent, but 46 of the 117 other Cy Young winners in that span had a better league-relative ERA.
The historic number was not 2.43. It was 1.63.
Two very different Cy Young cases
| 2021 statistics | Corbin Burnes | Zack Wheeler |
|---|---|---|
| Innings | 167 | 213 1/3 |
| Wins | 11 | 14 |
| Strikeouts | 234 | 247 |
| ERA | 2.43 | 2.78 |
| FIP | 1.63 | 2.59 |
| WHIP | 0.94 | 1.01 |
| Complete games | 0 | 3 |
| Shutouts | 0 | 2 |
| FanGraphs WAR | 7.5 | 7.3 |
| Baseball-Reference WAR | 5.3 | 7.5 |
Wheeler threw 46 1/3 more innings. That is not a rounding error or a small durability bonus. It is roughly seven additional modern starts. He led MLB in innings, completed three games and threw two shutouts. He also finished with more wins and more strikeouts.
If the award is primarily about how much value a pitcher accumulated over six months, Wheeler has a compelling case. Baseball-Reference WAR, which begins with actual runs allowed and accounts for workload, agreed.
Burnes’ argument was about the quality of each inning. His ERA was 0.35 runs lower, but the much larger difference was in FIP. Burnes struck out 234 hitters, walked only 34 and allowed seven home runs. Wheeler was excellent in those areas too, but his 2.59 FIP was almost a full run behind Burnes.
Our historical calculation places Burnes’ 1.63 FIP second among all qualified pitcher-seasons since 1967, trailing only Pedro Martinez in 1999. It remains second when FIP is compared with the league scoring environment. MLB’s own review similarly identified it as the second-lowest FIP of the Divisional Era.
That is why saying Burnes won with only 167 innings leaves out half the case. Voters were not merely forgiving a light workload. They were deciding whether one of the most extreme fielding-independent seasons on record was enough to overcome Wheeler’s 46-inning advantage.
FanGraphs WAR effectively called that trade a tie, giving Burnes a narrow 7.5 to 7.3 advantage. Baseball-Reference WAR preferred Wheeler by more than two wins. Even the two major versions of WAR could not agree because they were answering the value question differently.
The first-place vote ended 12-12
Here is how all 30 ballots were distributed:
| Pitcher | 1st | 2nd | 3rd | 4th | 5th | Points |
|---|---|---|---|---|---|---|
| Corbin Burnes | 12 | 14 | 3 | 1 | 0 | 151 |
| Zack Wheeler | 12 | 9 | 4 | 4 | 1 | 141 |
| Max Scherzer | 6 | 5 | 13 | 6 | 0 | 113 |
Burnes and Wheeler tied for the most first-place votes. The ten-point difference came entirely from the rest of the rankings.
Twenty-six voters placed Burnes first or second. Only 21 did the same for Wheeler. Burnes appeared in the top three on 29 ballots and never appeared fifth. Wheeler received four fourth-place votes and one fifth-place vote.
The pattern becomes even clearer when the ballots are separated by first choice. Ten of the 12 Wheeler voters put Burnes second. Among the 12 Burnes voters, only seven put Wheeler second. The six Scherzer voters also leaned toward Burnes, placing him second four times compared with twice for Wheeler.
Burnes therefore won as the broader consensus candidate. He was not more popular in first place, but voters who preferred someone else were more reluctant to push him down their ballot.
That is precisely why reproducing only the winner or the first-place vote misses part of the story. The Cy Young is a five-name ranked election scored 7-4-3-2-1. Being the common second choice can decide a close award.
Scherzer was not a spoiler
Because every individual ranking is public, we can remove one pitcher and see which of the other two each voter ranked higher. This creates three hypothetical head-to-head elections:
| Two-pitcher matchup | Preference |
|---|---|
| Burnes over Wheeler | 16-14 |
| Burnes over Scherzer | 23-7 |
| Wheeler over Scherzer | 19-11 |
Burnes and Wheeler began with 12 first-place votes each. Among the six Scherzer-first voters, four ranked Burnes above Wheeler and two ranked Wheeler above Burnes. Removing Scherzer therefore gives Burnes a 16-14 win.
This matters because Scherzer did not act as a spoiler. Burnes still defeats Wheeler directly, although only narrowly. Scherzer loses decisively to both of them. The ballots reveal a clear order, but also an electorate that was nearly divided between Burnes and Wheeler.
How fragile was the result?
One way to measure that closeness is to treat each observed head-to-head percentage as the preference of a larger voter population, then randomly select another panel of 30. Under that assumption, the hypothetical results are:
| Matchup | First pitcher wins | Tie | Second pitcher wins |
|---|---|---|---|
| Burnes vs. Wheeler | 57.4% | 13.5% | 29.1% |
| Burnes vs. Scherzer | 99.9% | 0.1% | <0.1% |
| Scherzer vs. Wheeler | 4.6% | 4.8% | 90.6% |
This is a fragility exercise, not Burnes’ true probability before the vote. It freezes the pitchers’ completed statistics and assumes the 30 observed voters accurately represent the larger pool of plausible voters. Its narrow conclusion is that another similar panel could reasonably have preferred Wheeler, while Scherzer was a much less plausible alternate winner.
A second calculation keeps more information. Put the 30 complete ballots into a virtual hat, draw 30 with replacement, and score the resulting panel using the actual 7-4-3-2-1 system. Repeating that process many thousands of times produces approximate winner frequencies among the three leading pitchers:
| Pitcher | Simulated win frequency |
|---|---|
| Corbin Burnes | 71.1% |
| Zack Wheeler | 28.6% |
| Max Scherzer | 0.35% |
Drawing with replacement means that some observed ballot types appear more than once while others disappear. It is a simple stand-in for asking what another 30-person panel with similar voting tendencies might do.
The two calculations are not contradictory. Burnes and Wheeler were nearly tied in direct voter preference, but Burnes had a modest structural advantage under the ranked scoring system. His 14 second-place votes and almost complete absence from the bottom of ballots made him more likely to survive changes in panel composition.
Nor does 71.1 percent describe Burnes’ pre-vote odds. The exercise does not simulate uncertainty in performance or discover the preferences of writers who were not observed. It measures how sensitive this completed result was to the particular mix of 30 ballots.
The market missed the real race
The betting market did not expect the ballots to look anything like this.
| Pitcher | Oct. 3 closing odds | Raw implied chance | First-place votes | Final points |
|---|---|---|---|---|
| Corbin Burnes | +100 | 50.0% | 12 | 151 |
| Max Scherzer | +115 | 46.5% | 6 | 113 |
| Zack Wheeler | +875 | 10.3% | 12 | 141 |
The implied chances include the bookmaker margin, so they are not meant to add to 100 percent. After removing the margin from the full listed board, Wheeler’s price represented roughly a nine percent chance of winning.
This was not simply a case of a longshot getting lucky. Wheeler tied Burnes with 12 first-place votes and finished only ten points behind him. Scherzer, one of the two market co-favorites, received six first-place votes and finished 38 points behind Burnes.
The market got the winner right, but it badly misread the shape of the race. It priced Burnes versus Scherzer when the actual electorate produced Burnes versus Wheeler.
The ballot resampling does not mean Wheeler’s true pre-vote probability was precisely 28.6 percent. It measures uncertainty in panel composition after the statistics and observed voter preferences are known. It still tells us something narrower and useful: the market greatly underestimated how much first-place and overall support Wheeler actually had among the 30 voters.
The miss is especially interesting because the market had recognized Wheeler earlier. He was the +240 favorite on August 8 and remained near the top of the board through mid-August before drifting to +325 on August 22 and eventually +875. His underlying case did not disappear. The market increasingly preferred the rate-stat cases of Burnes and Scherzer, while half of the first-place votes ultimately went to the workload candidates, Wheeler and the slightly higher-volume Scherzer.
Was this a paradigm shift?
Burnes’ victory looks like an obvious marker for a FIP-heavy era. An 11-win pitcher with the lowest full-season workload ever for a starting-pitcher Cy Young winner defeated the MLB innings leader. The statistic separating him most dramatically from the field was FIP.
But the vote was not a declaration that workload no longer mattered. Eighteen of the 30 voters chose someone else first. Wheeler tied Burnes in first-place support, and six voters preferred Scherzer. If the electorate had completely adopted one new philosophy, the result would not have been this fragmented.
Our individual-ballot research suggests the movement toward FIP-aware voting had already begun before 2021. Burnes was the season in which that growing share of the electorate became visible enough to determine the winner. It was evidence of a shift already underway, not necessarily the day every voter changed.
Team context may have supplied a little extra narrative. Milwaukee won 95 games and its division, while Philadelphia finished 82-80 and missed the postseason. I would not treat that as the main explanation. The ballot was much too divided, and Burnes’ statistical case was much too strong, to reduce the result to team success.
Who should have won?
I do not think Wheeler was robbed. I also do not think Burnes was an obvious choice.
Wheeler supplied 46 more innings of elite pitching. Burnes supplied the better ERA and a historically extreme FIP. One version of WAR preferred Wheeler comfortably while another gave Burnes a tiny advantage. The voters divided their first-place support evenly.
My own choice would depend on how much responsibility the Cy Young should place on volume. If forced to cast one ballot, I would lean Wheeler because 46 high-quality innings are a substantial amount of additional value. But that is a preference between two legitimate definitions of an award season, not evidence that the Burnes result was indefensible.
The best explanation for the result is also the simplest. Twelve voters preferred Burnes’ dominance. Twelve preferred Wheeler’s combination of excellence and volume. Six preferred Scherzer. Once the first-place vote ended in a tie, Burnes won because nearly everyone could accept him as the next-best answer.
He did not beat Wheeler in a landslide. He beat him by being harder to rank lower.
Sources and methodology
The complete 2021 ballot is available from the Baseball Writers’ Association of America. MLB described it as the closest National League result since the ballot expanded to five names in 2010 and summarized the competing Burnes and Wheeler cases. The historical market prices come from Sports Betting Dime’s archived 2021 odds.
The historical ERA and FIP comparisons use the same Lahman-based pitcher dataset as my individual-ballot study. Pitchers were required to reach 162 innings in a full season. FIP uses a season-specific MLB constant, and league-relative rankings compare each pitcher’s figure with the scoring environment in his league and season. The shortened 2020 season is excluded.
The pairwise comparisons use each voter’s published ordering to determine which pitcher was ranked higher. The head-to-head panel probabilities use a 30-trial binomial calculation with the observed preference rate treated as fixed. The ranked-ballot resampling draws 30 of the observed ballots with replacement and reapplies the BBWAA’s 7-4-3-2-1 scoring. Both are sensitivity exercises based on the completed electorate, not estimates of the award probabilities that existed before voting.