There shouldn’t be a consensus pick for AL Rookie of the Year, but it seems there is

So by now (late October 2024), all votes are in for 2024 AL and NL Rookies of the Year, so at this point anything I say in this article is me ranting and complaining about something that I think has already happened, but won’t know for certain has happened until three weeks from now. Nothing I can do to change it now!

But if you read articles like this one, you get the idea that Colton Cowser is the consensus pick to be AL Rookie of the Year for 2024. Or at least it’s down to him and Luis Gil.

Then there’s this betting odds line from when there were 5 days left in the season. It had Luis Gil with a greater than 50% chance to win, with Colton Cowser and Austin Wells also having good chances, and Wilyer Abreu being an extreme longshot, with a less than 2% chance to win it.

Luis Gil and Colton Cowser. Why these two? It seemed to me like there was a crowded field of top candidates in this race.

Well, someone in that article mentioned WAR as part of their reasoning. And though WAR isn’t the be-all end-all criterion for choosing Rookies of the Year and MVPs, I think it’s a great first step in creating a pool of top candidates to ponder and consider, to compare and contrast. So I considered it. I considered it in 16 different ways, in fact.

So it was interesting to learn in this assessment of the WARs that AL rookies accumulated in 2024, that Gil maybe shouldn’t even be in the mix. I have him as just outside the 6 strong contenders, who are, in alphabetical order:

Wilyer Abreu
Colton Cowser
Wyatt Langford
Mason Miller
Cade Smith
Austin Wells

How did I come to that conclusion? Well let’s start by looking at the top WARs of AL players. Note that there are two different commonly used ways of calculating WAR, that are scaled to be a similar size, but can produce quite different answers for individuals. One of these ways comes from Baseball-Reference.com, and it is often labeled bWAR for clarity. Here are the AL rookies with the highest bWARs for 2024:

There are Gil and Cowser, but they’re in the third and fourth spots. Yet they’re both within 1 WAR of the leaders, Wyatt Langford and Wilyer Abreu, which is not considered a truly significant difference.

What does the other WAR look like? It is calculated by FanGraphs.com, and it is often labeled fWAR for short. Here are the AL rookies with the highest fWARs for 2024:

Ah, there’s Cowser at the top. But Gil is in 7th place? And of the 8 players appearing on both lists, none of them are in the same place as before. Also, as with bWAR, a lot of guys are within 1 fWAR of each other.

Some of these differences are surely indicative of the inherent inaccuracies of estimating the value of a player’s performance. That said, WAR is our best current tool for objectively attempting to measure the overall value of a player’s performance. We just must recognize the approximate size of those inaccuracies, and as others have said before me, we must use WAR as a conversation starter, not a conversation ender.

But hey, maybe we can squeeze a little more accuracy out by averaging the bWAR and the fWAR? That looks like this:

Colton Cowser on top again (thanks to his high fWAR), but the top 6 guys all within one WAR of each other! And a new top 2. This makes things look even tighter.

But maybe it would be better to look at each player’s peak assessment – their highest WAR attained between the two systems?

The same top 4 and top 5 as with the average WAR, though a different top 3. But no … maximum can be completely influenced by one system having a very inaccurate high assessment. Having a high minimum WAR between the two systems only happens if both systems rate you highly – we should therefore have more confidence in a high minimum WAR than a high maximum WAR. Here’s what that looks like:

Well that shook things up a bit! A new name on top.

But looking over all five lists, there are four names always hanging out near the top – (alphabetically) these are Wilyer Abreu, Colton Cowser, Wyatt Langford, and Austin Wells. Perhaps we should declare these four in a dead heat as far as WAR is concerned, and focus on these four in determining, using other measures, which one should be AL Rookie of the Year. But Luis Gil is not part of that group. Why is he such a frontrunner for Rookie of the Year, then? I’ll share some thoughts on that later.

But hang on. There’s an important factor we haven’t considered here: playing time. Among the highest performers, WAR acts like a counting statistic – the more you play, the more WAR you accumulate. So in comparing top players, it may be more appropriate to turn WAR into a rate stat by dividing it by some measure of playing time. So let’s do that.

For position players, I like to use plate appearances as the measure of playing time, as plate performance has an outsized effect on overall WAR.

For pitchers, I like to use batters faced instead of innings pitched, because it’s the number that is most relatable to the position player’s plate appearances. In fact, if you add up all players’ plate appearances over all of MLB for an entire year, you get exactly the same number as you get when totaling batters faced in the same way.

But pitchers as a whole don’t get as much WAR as position players do. Probably rightly so, because position players can add to the run values of their batting contributions with their defense and/or baserunning; pitchers can only add value by reducing the run contributions of the batter. Because major league players collectively will do better than replacement-level players at baserunning and fielding, on average these will add a positive amount of WAR on top of their batting contributions. Pitchers only have access to the run values available in the pitcher/batter interaction.

So since there’s less WAR to go around for pitchers, to assess how good their play was against position players, we must grant them more opportunity by allowing them to accumulate WAR over a larger number of batter/pitcher interactions. Looking at the numbers from 2023 and 2024, I found a pretty consistent trend: position players get a little over 40% more fWAR per batter/pitcher interaction than pitchers do, and a little over 50% more bWAR. To be a little conservative about it, I used the 40% and 50% adjustments. So we calculate WAR per 500 plate appearances for position players, fWAR per 700 batters faced for pitchers, and bWAR per 750 batters faced for pitchers. These numbers were chosen to get into the vicinity of one season’s worth of plate appearances.

The results of this adjustment for all five tables we made before can be found below.

Now you may object to my adjustment, and insist that position players and pitchers should be evaluated based on equal numbers of batter/pitcher interactions. Okay, to humor you, I also did the same five tables as before that way – WAR per 500 PA for position players, and WAR per 500 BF for pitchers. So now we have 10 new tables to show you.

But before we do, we have one more adjustment to make. Because we’re dealing with rate stats now, players with a lower number of chances will be able to ride a “hot streak” of good luck to achieve higher rates than a player with more chances is able to access. So we insist on a minimum number of chances – 400 for each. If a player has below 400 PA or BF, we divide their WAR total by 400 instead of by their actual, lower number. It’s like we’re assuming that a player with 300 PA would have accumulated 0 WAR over their next 100 PA.

With that adjustment, now we’re ready. Here are the five tables using WAR per 500 PA and WAR per 700 or 750 BF:

Wow, Cade Smith and Mason Miller completely took over the top of the lists!

But here are the five tables using WAR per 500p PA/BF. Will they stay atop those?

They’re not atop these, but still placing higher than before, as are David Hamilton and Parker Meadows. But the new dominant pair on these five lists are Wilyer Abreu and Austin Wells. And Luis Gil has nearly disappeared from them!

All these lists, all 15 of them looked at collectively, tell a tale of ambiguity. These 15 lists have 14 different top 3s. There are nine different players occupying top 3 positions. There are five different players occupying the top position of a list. There is no clear cut frontrunner here!

But surely some are appearing near the top moreso than others. Here’s one thing we can do to sort those out: for each player, find their highest position on any list, and their lowest position on any list. Here is what that looks like:

And there we see the two runaway favorites to win AL Rookie of the Year coming in behind five other guys. And looking at their numbers, I can see that a good case could be made to vote for Wilyer Abreu, Cade Smith, or Wyatt Langford over Colton Cowser or Luis Gil. Maybe even Austin Wells too.

So why are Cowser and Gil so favored to win?

Is it because they’re from major markets? No, because Wilyer Abreu is from a major market.

Is it because most of the others played less? I think that is a factor, especially for Cade Smith, who as a reliever doesn’t face nearly as many batters as starting pitchers do. And a little bit for Wilyer Abreu, who missed part of the middle of the season due to injury.

But I think the biggest factor is that Colton Cowser and Luis Gil played for playoff contenders, and except for Cade Smith and Austin Wells, most of the rest didn’t. And the fact that that’s probably true really makes me sad.

Why should otherwise equal players be judged differently based on the performance of the other players on their teams? That makes no sense, and should not factor at all into the voting. But I fear it does. It’s looking like Wilyer Abreu and Wyatt Langford will be cheated of Rookie of the Year votes due to not playing on contenders, and Cade Smith will be cheated out of votes for lack of playing time.

Here’s hoping the prognosticators got it wrong, and we see Cade Smith, Wilyer Abreu, and Wyatt Langford finish with strong scores in the Rookie of the Year voting.

Who was the “most .500” team of 2024?

Of course, one could say that the “most .500” Major League Baseball team in any given year is the one whose record is closest to .500, or 81-81 in a full season. But even a hypothetical team that always had exactly a 50% chance of winning would sometimes end up, by luck of the “coin flip”, a few games away from .500.

And what about a team that’s great for the first half of the season, then awful for the second half, ending up with a .500 record? They weren’t really a .500 team at any point in the season, in that their chance of winning games wasn’t actually close to 50% at any point, nor were they winning about half their games in any given week.

So here are a few different ways to measure how .500 a team was, along with the top teams by each method.

Final Record

We can just look at a team’s final record and see how many games away from .500 it was, above or below.

The Boston Red Sox had the only .500 record, but several other teams were close.

Run Differential

This was mentioned in Sarah Langs’ 2020 article What does a true .500 team look like?

The run differential of a team is the runs it scores over the entire season minus the runs it allowed in that same time. A small run differential is a good predictor of a team that will have a record near .500. (There is even a stat called Pythagorean expectation which estimates what record a team should have based on it totals of runs scored and runs allowed.)

Whose run differential was closest to 0 in 2024?

Four teams had a run differential close to 0. Of these, again, the Boston Red Sox were the closest to 0, just barely. It seems we have a frontrunner.

Number of times at .500

A team that plays “a .500 brand of baseball” throughout the season is likely to have a winning percentage of exactly .500 at several times during the course of the season. The most times this could possibly happen is 81, though even for a hypothetical team that always has a 50% chance of winning, the odds of that happening 81 times are over 2,000,000,000,000,000,000,000,000 to 1 against. The most times it’s ever been done, at least before 2020, is 35 by the 1959 Chicago Cubs.

The Tampa Bay Rays came close to that this year, tying for 3rd most. The Padres, Red Sox, and Cardinals also had a lot.

For fun: consecutive times at .500

This last one is more about the luck of streaks than anything else. But there was an interesting streak this year in this regard, so I thought I’d throw it in.

When a team is at .500 in the middle of the season, the next game they play takes them off of .500; it’s only 2 games later that they can be back at .500 again. So a streak of consecutive times at .500 means that at the end of every 2 games played after being at .500, they’re back at .500 again.

The Red Sox were at 26-26 on May 25, 2024 – 26 wins and 26 losses. Two games after that they were 27-27, then 28-28, 29-29, and so on up to 35-35. That’s ten times in a row at .500. The likelihood of that happening, once a team has reached a .500 record, is more than 500-to-1 against.

Here are the longest such streaks in the majors in 2024:

The “Winner”

It’s gotta be the Boston Red Sox as The Most .500 Team of 2024. They top every list except number of times at .500, and they did pretty well there, too. Runner up goes to the Tampa Bay Rays.

Interestingly, these two teams played each other in their last 3 games of the season, with the Rays winning the first two but losing the final game. Had they won it, they would have replaced the Red Sox atop the Final Record list, probably solidified the Red Sox hold on the Run Differential list, but strengthened their own position atop the Times At .500 list. That game was something of a battle for Most .500 Team of 2024. Congratulations, Red Sox, on your “victory”!

Who should AL Player of the Month be, Encarnacion or Bradley?

To think about who should be the American League player of the Month for August, we could start by looking at those with the highest OPS on the month (and at least 50 plate appearances):

Player Team Pos G AB R H 2B 3B HR RBI BB SO SB CS AVG OBP SLG OPS▼
 Encarnacion, E TOR 1B 23 86 23 35 11 0 11 35 9 15 0 0 0.407 0.460 0.919 1.379
 Ortiz, D BOS DH 26 91 17 32 8 0 9 22 16 17 0 0 0.352 0.432 0.736 1.169
 Bradley, J BOS CF 26 79 23 28 9 3 5 23 11 24 3 0 0.354 0.429 0.734 1.163
 Donaldson, J TOR 3B 27 105 29 34 7 1 11 35 16 25 2 0 0.324 0.408 0.724 1.132
 Gutierrez, F SEA LF 19 62 12 21 4 0 7 20 4 19 0 0 0.339 0.388 0.742 1.130

Based on offense alone, you have to pick Encarnacion, though Ortiz, Bradley, and Donaldson all show very well here. But can defense close the gap? Not for Ortiz, the DH, but maybe for Jackie Bradley Jr., the defensive wiz in the outfield. Now I haven’t seen Encarnacion’s defense this month, but I have to wonder, how likely is he to have made plays at first base in August like this catch:

Bradley Jr.’s incredible catch

or this catch:

Statcast: Bradley’s great grab

or this throw:

Statcast: Bradley Jr. gets Bird

or this catch:

Must C: Bradley Jr.’s great grab

or this throw:

Bradley Jr. nabs Sanchez

or this catch:

Bradley runs in for catch

or this throw:

Bradley Jr.’s throw nabs Infante

or this catch and throw:

Bradley’s running catch

Given the game-changing, run-saving nature of Bradley’s defense so many times in August, that has to propel him squarely into a two-person discussion for who should be AL player of the Month for August.

Do you think the pick should be Encarnacion, Bradley, or someone else?

My previous mathematically-oriented baseball posts

I’ve been away for a while.  But I’m returning.

Where have I been the past year?  Mostly over here, and sometimes commenting over here.  But during that year, I’ve done a lot of mathematics to study questions I had about the baseball I was reading about and following, and some of that has filtered into some of those posts.  I thought I’d provide a selection here of some of the more interesting ones.  Some of these contain hints to posts I plan on putting up here in the coming weeks, posts that will include discussions of the math of streaks, and just how much a small sample size actually tells us.  There will be other new topics too, not previewed in any of these posts.  Stay tuned.

Here is that selection of my mathematically-oriented posts of the last year or so:

A post from August 2014 explaining why batting 15 points above league average is sometimes actually hitting at league average. This in the context of examining one upcoming player.

A comment titled “Expected frequency of reverse platoon splits exceeds the actual numbers” to the article “Are reverse platoon splits sustainable?” on Beyond The Box Score. In this comment (scroll to the bottom of the comments section) I used binomial theory to come up with what would be the expected number of players, based on random chance alone, having a reverse platoon split in on-base percentage for the years 2012 and 2013. I show that the actual numbers were less than the numbers you’d expect by random chance, seeming to indicate that reverse platoon splits are unsustainable.

A comment titled “No, because starters face more batters” to the article “The Hidden Perfect Games of Relievers” on Beyond The Box Score. In this comment (scroll to the next-to-last comment), after making two points about the right way to compare starters and relievers for the purpose of the article, I discussed my first attempts at producing an expected number of “wrap-around” perfect games that will occur in a given season for starters and relievers, to help clarify any meaning that might be attached to the reported results. I did complete that work, which I plan to publish later on this blog, in a post about the math of streaks.

A post from September 2014 that argues that a certain young player is better than his overall numbers say he is, by analyzing his advancement as a hitter at each new level he played at.

A post from September 2013 explaining why one baseball team’s chances of making the playoffs were ridiculously close to, but not quite exactly, 100%.

A post from later in September 2013 which explains (in more detail than anyone probably cared to read) why that same baseball team’s chances of having home-field advantage were about 7 out of 11 (washing dishes at night gave me a lot of time to listen to baseball and think about this stuff).

6-team AL wildcard race now looking like a 3-team race

It’s been exciting watching the wild card race in the American League evolving these last couple of weeks, with 6 teams having a real shot.  With division leaders pulling away, making the division races relatively uninteresting, and with the National League’s 5 playoff entrants pretty much a done deal (with only positioning remaining a question), this race has provided most of the late-season playoff race drama.

But as we approach the last week of play of the regular season, 3 of those 6 contending teams now look like outside longshots.

Each of these 6 teams has either 7 or 8 games remaining in the season.  It’s not likely that any of them will lose more than 3 or 4 of these remaining games.  However, the Yankees, Orioles, and Royals, each with 73 losses, will require at least two of the Rays (now at 69 losses) , Indians, and Rangers (70 losses each) to lose 3 or 4 games just to have a chance at tying.  Were the Indians and Rangers both to lose exactly 3 of their remaining games, one of the 73-loss teams would have to win all their remaining games just to tie.  Not unheard of; the 2007 Rockies faced this sort of scenario with just over 2 weeks to go that season, needing to win their last 15 games to make a wild card berth probable; they won 14 of those 15 to tie for the wild card and force a one-game playoff for the spot (which they won).  These streaks would be half as long, and with 3 teams poised to try for it, it’s not too out-of-the-question that one may do it.

At this point, scheduled opponents can make a big difference.  The Orioles seem to have the short end of the stick here, with 2 of their remaining 8 games against the Rays (who are fighting to keep their slim wild card lead). and 3 against the Red Sox (who will likely be trying to maintain their lead for home-field advantage against the other division leaders, Detroit and Oakland).  The Yankees also have 3 games against the Rays, but otherwise have an easy schedule, with 3 games against the bottom-dwelling Astros.  The Royals seem to have the best schedule of all though, with today’s game against the Rangers their only one against a contending opponent.

Though the Rays have the best record right now by a slim margin, if the Yankees or Orioles make a charge now, the Rays’ position in the standings will fall rapidly, while the Rangers and Indians, with easier schedules, would most likely stay put at the lead of the wildcard race.  Unfortunately for the Yankees and Orioles, this would only allow them to leapfrog one of the three leading teams; not enough to take a wildcard berth.

In the end, two of the 3 leading teams must falter, and that just doesn’t seem all that likely.  The Yankees, Orioles, and Royals are all positioned to make it interesting by winning, but won’t likely catch a wild card berth even if they do.