Table of Contents

Monday, October 22, 2007

Player Value, Part 2b: Offense - Baselines

To view the complete player value series, click on the player value label on any of these posts.

Part 2b: Baselines

Ok, now that we've decided upon a runs estimator--linear weights derived from a BaseRuns model that fits our particular context--we need to think about how exactly we're going to go about assessing player value. There are four typical ways that folks use offensive runs data to assess player value or performance: absolute runs, runs per game, runs above average, and runs above replacement.

To anchor our discussion, below I've created a table that reports values for the 2007 Reds based on each of these approaches. Runs values were estimated using the linear weights for the '03-'07 National League reported in the previous article, and replacement level was defined as 73% of position player league average, with no positional adjustment (more on that later):

Absolute Runs
Runs Per Game
Runs Above Average
Runs Above Replacement
Name PA AbsRuns
Name PA R/G
Name PA RAA
Name PA RAR
Dunn 632 107.4
Dunn 632 7.2
Dunn 632 31.0
Dunn 632 51.3
Phillips 702 97.9
Votto 89 7.1
Griffey Jr. 623 15.9
Griffey Jr. 623 36.4
Griffey Jr. 623 92.8
Hamilton 337 6.8
Hamilton 337 14.0
Phillips 702 30.2
Encarnacion 560 76.9
Hatteberg 417 6.5
Hatteberg 417 13.7
Hatteberg 417 26.9
Hatteberg 417 63.5
Keppinger 276 6.4
Keppinger 276 8.3
Hamilton 337 25.2
Hamilton 337 55.9
Griffey Jr. 623 6.2
Encarnacion 560 5.8
Encarnacion 560 24.8
Gonzalez 430 55.1
Cantu 68 6.1
Phillips 702 5.6
Keppinger 276 17.1
Hopper 335 41.4
Wise 6 6.0
Votto 89 4.3
Gonzalez 430 13.3
Keppinger 276 41.3
Encarnacion 560 5.5
Cantu 68 1.6
Hopper 335 10.8
Ross 348 31.8
Phillips 702 5.4
Wise 6 0.1
Votto 89 7.3
Freel 304 30.0
Hopper 335 5.1
Hanigan 11 -0.1
Conine 242 4.6
Valentin 265 28.7
Gonzalez 430 4.9
Jorgensen 15 -0.2
Cantu 68 3.8
Conine 242 28.3
Hanigan 11 4.8
Cruz 1 -0.3
Valentin 265 3.2
Votto 89 15.5
Jorgensen 15 4.7
Hopper 335 -0.4
Jorgensen 15 0.4
Cantu 68 9.9
Conine 242 4.5
Gonzalez 430 -2.0
Wise 6 0.3
Ellison 56 4.1
Valentin 265 4.2
Bellhorn 18 -2.1
Hanigan 11 0.3
Castro 98 2.9
Freel 304 3.7
Coats 38 -2.9
Freel 304 -0.2
Coats 38 2.6
Ross 348 3.2
Ellison 56 -3.9
Cruz 1 -0.2
Jorgensen 15 2.2
Ellison 56 2.6
Conine 242 -4.1
Bellhorn 18 -1.4
Lopez 47 1.6
Coats 38 2.4
Lopez 47 -5.6
Coats 38 -1.4
Hanigan 11 1.3
Lopez 47 1.2
Valentin 265 -6.2
Ellison 56 -1.8
Wise 6 0.9
Castro 98 1.0
Moeller 49 -7.4
Lopez 47 -3.6
Moeller 49 0.6
Bellhorn 18 0.9
Freel 304 -11.3
Ross 348 -4.6
Bellhorn 18 0.5
Moeller 49 0.4
Castro 98 -12.3
Moeller 49 -5.2
Cruz 1 -0.1
Cruz 1 -2.5
Ross 348 -18.0
Castro 98 -8.2
(as an aside, and as a case study in differences among runs estimators, you might find it interesting to compare these numbers to THT's RC calculations, which use the most recent version of RC. Some players are badly undervalued by RC, especially Phillips; perhaps this is because of the degree to which his power made up his offensive value. In fact, if you use THT's RC number and convert it to runs above average, you'll find that he's actually ranked 7 runs below average!! Phillips may be overrated as a hitter, but he sure as heck isn't a below-average hitter).

Alright, let's walk through these rankings one by one.

Absolute Runs: An estimate of the absolute number of runs produced by a player.

Absolute Runs are just the direct output of our linear weights equation (see here for overview of linear weights).

The top of this scale looks appropriate, as all of the Reds' best offensive players are up near the top of this chart. And the bottom seems appropriate too, with players who weren't around much bringing up the rear. But to my eye, there are some issues with how players are ranked near the middle.

With an absolute scale, any time you do something positive on the ball field, you're rewarded. The result of this is that players that get more playing time, almost regardless of their production in that playing time, will tend to create more runs than players with less playing time. In fact, the correlation on these Reds between PA's and absolute runs is 0.98! Now, it's true that, in fact, those players did do more total good things than players that are ranked lower. The problem is that they also might have done a lot of terrible things, like create lots of outs, which aren't fully accounted for in an absolute runs statistic.

Ryan Freel, for example, had a rather poor 0.655 OPS this season. That's well below league average, and is about the minimal level of production you'd anticipate from any scrub you call up from AAA. And yet the absolute runs estimate ranks his production as virtually identical to Javier Valentin (0.715) and Jeff Conine (0.729). Now, those guys didn't exactly knock the cover off the ball, but they certainly had a better season than Freel did...and yet, Freel is rated as their equal because he had more chances, even though his tendency to create outs may have hurt the Reds more than he helped them.

A more controversial (to some) example comes at the top of the chart. Brandon Phillips is rated second on the Reds' team with just under 100 runs produced. And yet, his low OBP caused his OPS to be a fairly mundane 0.816 this season. Certainly not bad, but would you rank his offensive contributions above guys like Griffey (0.869), Hatteberg (0.868), or Hamilton (0.922)? Maybe you would; there is certainly value to playing as much as Phillips played (700+ PA's). But maybe not, as those guys were better producers on a per-PA basis than Phillips was. For example, it could be the case that combining Griffey's plate appearances with plate appearances from an AAA scrub (DeWayne Wise?) would result in more production Phillips' production alone in the same number of PA's. To assess this, we'll need to look at other approaches.

Runs Per Game: Per-game rate of runs production.

This is a pure rate stat. And as we discussed in the piece on run estimation, because absolute runs don't completely account for the impact of producing outs on a team (fewer opportunities for future batters), the most correct way to convert absolute runs to a rate is to use runs per out. In this case, we're converting it to runs per 26.25 outs, because teams will average 26.25 outs per 9-inning game on offense (27 in all away games, 27 in losing home games, 24 in winning home games):

RpG = Runs/Outs*26.25

Where outs is AB-H+SF+0.92*SH.

Because this is purely a rate statistic, it completely ignores playing time. Brandon Phillips' ranking takes a fairly big hit here, and we see just how irrelevant Juan Castro and Chad Moeller were as hitters. Those are insights we don't get from absolute runs.

But at the same time, we see Joey Votto ranked among the best hitters the Reds had. It's true that his per-out and per-PA rates were excellent, but he also didn't even get 100 PA's, so we have to figure that the uncertainty around those estimates is pretty darn high. Votto's performance might portend great things to come, but in terms of evaluating historical value to the Reds team, I prefer the next two approaches.

Runs Above Average: players get credit or are penalized for their performance relative to their competition

To calculate runs above average, you convert a player's total runs produced to runs per game, subtract from that total the league average rate of runs per game, and then convert the resulting number back to season totals for that player. Or:

RAA = (RpG - LgAvgRpG)/26.25*Outs

Note, for National League players in particular, I think it's important to use league average runs per game for position players, rather than overall league average runs per game. Otherwise, you're including pitchers among the hitters that you're comparing your position players to, which seems inappropriate to me--we're not really interested in assessing how well our players hit relative to pitchers, are we? We want to know how well they hit vs. other position players. The difference is substantial (~0.3 r/g), so it's worth doing. In 2007, NL position players averaged 5.1 r/g according to my numbers.

Now, from the perspective of this method, players are valuable in as much as they help you beat other teams. After all, if the point is to win more games than you lose, you want players that help you accomplish this task. Therefore, you should compare each player to league average production. If a player produces above league average, he's helping you win. If he's below league average, he isn't helping you win--at least on offense.

There are a lot of advantages to using this baseline. First and foremost, it's very straightforward. If a player is above or below average, you immediately know something about how their performance was tied to your team's success. For this reason, I think it's probably one of the better statistics by which one can evaluate MVP's or other top hitting awards (Dunn, for example, was clearly the Reds' best hitter in 2007). This approach also provides some balance between valuation of playing time and performance: Phillips is ranked below guys like Hatteberg, Hamilton, and Griffey, but above Cantu and Votto (though ranking him below Keppinger is a bit problematic for me given how little Keppinger played).

Still, I'm not a huge fan of this baseline for most purposes. Two reasons. The first is largely aesthetic, and has to do with how it ranks players. For example, looking at their RAA, you see that Jeff Conine and Javier Valentin are rated as below average--which ranks them "below" the value of someone like Jason Ellison, who got such a paltry number of plate appearances and that he didn't have time to contribute much of anything, positive or negative, to the team.

My second issue is the fact that an average player is given a value of "zero," which seems to ignore how hard it can be to find "average" or even slightly below average players. The Reds had a pretty good offensive team this season, and as a result, most of their position players were ranked as above average. But Alex Gonzalez, who had a respectable (and career-best!) 0.793 OPS this year in 430 PA's, was ranked as just barely below average. And yet, that sort of player doesn't just grow on trees. Average players have a lot of value--they may not be helping the team win, but they're helping the team avoid losing, and thus avoid negating the production of the above-average players.

I realize that advocates of comparison to average systems argue that the claim that average players are undervalued by this approach is a misinterpretation of the data. And that's fair enough. But even if most of us properly interpret the data, we can't be guaranteed that all folks reading our work--especially those who are less experienced in thinking about these numbers--won't misinterpret them. That's why I prefer the last approach...

Runs Above Replacement: Players get credit for production above that which you would expect from an AAA scrub.

Calculations of RAR is similar to RAA, except that here you subtract a certain percentage of league average, rather than straight-up league average, from a player's runs per game. I recommend using 73% of position player league average (more on that below). Here's my equation:

RAR = (RpG - 0.73*LgAvgRpG)/26.25*Outs

The idea behind this approach is that even an AAA scrub--often dubbed "freely available talent," because even if a GM doesn't have an appropriate player in their system, they should be able to acquire one for next to nothing--is going to do a fair number of "Good Things" on the field if you play them for a full season. But their rate of Good Things will be much lower than what you'd expect from most MLB players, be they bench players or starters. Therefore, this approach sets a minimum level of production against which all players are judged, and players gain value based on how much better they perform above this minimum.

What I like about this approach is that it recognizes something about the importance of average performers, and more closely reflects the decisions a GM has to make. If a player is below-average, that might be a spot on the roster that a GM may try to improve. But as long as he's producing above the minimum level, he's producing at a level that provides real value for his team (i.e. better than talent could be acquired for free). Such a player may not be helping the team win, but he's helping the team avoid losing to a degree greater than what you could expect from any Joe Schmo pulled off the waiver wire or from a AAA team.

From a practical standpoint, I really like how this approach ranks the '07 Reds. Phillips is given a high ranking, which recognizes how valuable it is to have a guy who can produce at an above-average clip across 700 plate appearances. And yet, it also recognizes that excellence in performance is important, as evidenced by Joey Votto's (0.907 OPS in 89 PA's) ranking above guys like Conine, Valentin, Ross, and Freel.

The latter two guys, Ross (0.670 OPS) and Freel (0.655), were identified as hitting a touch below replacement level in the NL. That's not a good thing, and it reflects how bad their performances were--you might not gain anything by sticking any old AAA veteran in for those PA's, but you wouldn't lose anything either (remember, we're talking strictly about offense). Replacement level also recognizes, perhaps more so than any of the other approaches, the absolutely disastrous seasons by Juan Castro (0.446 OPS) and Chad Moeller (0.417) at the plate (~150 PA's between the two of them!!).

There are some problems, however, with using a replacement-level approach. The most significant is that there's not really a great consensus on where, quantitatively, we should draw the line of minimum performance. Brandon Heipp discussed a number of the attempts to come up with some kind of objective measure of replacement level in his article on baselines. I myself recently completed a study that tries to look at this, and I found that there wasn't a perfectly clear number, but it probably is somewhere between 70 and 80% of league average. At this point, I recommend using 73%. There is good theoretical justification for that number, and it falls within the empirical bounds for my study... and, compared to the higher alternatives (Keith Woolner uses 80% in VORP for most positions), it runs less risk of dismissing production that may be genuinely hard to replace.

The other problem with replacement level has to do with its typical justification. One traditional argument for replacement level is that it's the production above that of whoever would replace a player if they were no longer present on a team. As Brandon Heipp points out, however, this isn't really what we're assessing when we employ a 73% or 80% baseline: if a starter goes down due to injury, he will typically be replaced in the lineup by a bench player who will hit substantially better than replacement level. If we really were trying to identify value over replacement, we should probably try to use a method like the chaining approach Heipp describes.

But in actuality, I don't think that's really what we're after. I think we're trying to compare players to a minimum level, below which a player won't be able to retain even a fringe bench job on a major league roster. Anything better than that is worth noting. Anything less than that means you don't deserve to be in the majors...unless, of course, your defense makes up for it (more on that in a future piece).

In sum, I find runs produced over replacement, as described here, to be a very useful number. It gives a low baseline above which we can reasonably expect all players to play, it recognizes the value of average ballplayers, and player rankings based on this baseline make intuitive sense. Therefore, this is the number that I prefer to use when evaluating a player's offensive value to his team. As we argued in the first piece, however, offense is just half of a player's value...we also need to evaluate his defense. And there's also the question of whether offensive production by a player at one position is equally valuable as production at another position. More on that next time...

Update (11/7/07): I updated the above table to use slightly different linear weights, this time factoring in GDP's.

Coming up next...Positional adjustments on offensive numbers, and why we shouldn't do it.

Friday, October 19, 2007

What do we mean by replacement players?

What follows is something that I originally did in the process of putting together my player value series, which will continue tomorrow night. However, this part evolved into its own little (if 3000 words can be little) stand-alone paper. It's not perfect, but it's a nice little study that can help us understand something about how players are used in the big leagues.

Introduction

Comparison to replacement level has become a common practice among statistically oriented fans and analysts when trying to understand player value. The motivation behind this technique goes something like this: we can estimate, to a great deal of accuracy, how many runs a team has generated in a season from their offensive stats (# of singles, doubles, home runs, walks, etc), using any number of run estimation stats (runs created, linear weights, etc). Furthermore, we can estimate these same runs estimates for individual players, giving us an idea of how many runs each player contributed to a particular team's offense.

However, when you try to use these runs estimates to understand player value to a team, you run into a problem: playing time has as much, if not more to do with a player's runs created as their actual performance, because players are credited for every single "Good Thing" that they do. For example, this season, Orlando Cabrera was second on the LA Angels in runs created (93) despite hitting for just a 0.742 OPS. He did rank behind Vladimir Guerrero, but ranked well ahead of Chone Figgins (0.825 OPS) and Casey Kotchman (0.840 OPS). The reason? Cabrera led the team with 701 plate appearances compared to Figgins' 503 and Kotchman's 508. It's not that Cabrera was bad this season at the plate, but he wasn't good, and certainly wasn't the Angels second most valuable hitter to most observers' eyes. And while durability certainly has value, if a durable player doesn't hit (or play defense) any better than a veteran AAA call-up would, he's not doing his teams any favors.

This is where comparisons to replacement players come in. In Cabrera's case, what we're interested in knowing is not how many total Good Things he did; we're interested in how many more Good Things he did compared to the number of Good Things a veteran AAA replacement player would be expected to do. It's this performance above replacement-level production that we're actually interested in, because that's the difference between someone who should be starting (or not) in the big leagues, and someone who should be riding the pine or playing in AAA.

The problem when it comes to actually quantifying performance above replacement level is that there's disagreement about how well replacement players perform. For example, folks typically estimate replacement level as some percentage of league-average production. However, you'll find a lot of disagreement about what percentage to use. Keith Woolner's VORP uses 80% of league average for most positions. In contrast, two well-respected amateur baseball researchers (Tom Tango and Brandon Heipp) recommend using a value closer to 73%. Each of these variants has its own reasonable justification, but most researchers also seem to agree that there's not a distinctly correct point at which we should assign the cutoff. So, given this uncertainty, I decided to take a look at some data describing player performance over the last several years to try to figure out how replacement players really do perform.

Methods

My study assumes that, in general, playing time reflects player quality, such that starters, bench players, and replacement players can be identified based on how often they play. Briefly, my methods, which are similar those of some of the previous replacement player studies, are as follows: I pulled '04-'07 stats from THT's stats pages for all players in both leagues. After generating custom linear weights for each league over the 4-year time span in question, and then calculating runs per game for each player, I sorted players within each year by plate appearances (my indication of playing time). I then examined successive "slices" of players, removing one player for each team in the league (i.e. 14 players per slice from the AL, 16 players per slice from the NL).

For the ease of description, I'm defining the first 8 (in NL) or 9 (in AL) slices as "starters," and the next 3 (in AL) or 4 (in NL) slices as top "bench" players, making for 12 slices that consist of players that are virtually guaranteed roster spots on their teams. The next 3 slices were defined as "fringe" players, who may or may not make their teams' rosters over the course of the season depending on how many pitchers a team carries and the particular needs of the team over the course of a season. Slices below these fringe slices (slice #16 and higher) were defined as "scrubs." Note, I did not keep track of teams or positions when doing this slicing; it was strictly based on plate appearances within years (there were a variety of reasons that I went this route, but theoretical and practical). Nevertheless, post hoc assessments of positional composition, at least, showed that it was fairly consistent across positions, especially in the AL.

Now, it true that this study is somewhat limited by the fact that it assigns player roles after the fact, rather than ahead of time, and therefore suffers from sampling bias--players who might have been counted on for substantial playing time may have faltered and lost their jobs to a another player who might have previously been considered a replacement player. There are other potential problems as well. For example, if a starter is injured mid-way through the season, he will be classified as a part-time player in this scheme, while his backups will be ranked higher than they probably deserve. Nevertheless, especially after seeing and looking through the data, I do not think that this is a catastrophic problem for this dataset, or its ultimate findings. Nevertheless, a study similar to magpie's (which was not published until after I'd completed this study), in which players are classified into their roles prior to considering their playing time that season, would be a very nice complement to this one. Maybe I'll try to do that someday, though don't let me hold you back if you're interested! ;)

Finally, a brief word on the standard I used for league average: in this study, all players were compared to league average production by position players only. If one uses true league average runs per game, NL players are compared to a standard that includes a significant number of plate appearances by pitchers, and thus is about 0.3 runs per game lower than I think it should be (NL rates from '03-'07: 4.6 r/g with pitchers, 4.9 r/g without pitchers). We're interested in performance relative to other hitters, not performance relative to pitchers pretending to be hitters.

Expectations

Before we look at the results, let's think about what we're looking for in terms of replacement-level production. Here are a couple of ways that we might identify replacement level, based on the slicing procedure I'm using:
  • Replacement = Fringe: Replacement level should match the production of "fringe" players, who may or may not make the big league roster depending on the needs of their team (# of pitchers, lefty/righty composition of bench, speed on bench, defensive skill on bench, etc).
  • Replacement = Scrub: Replacement level should match that of "scrub" players, who will be unlikely to make the big league roster except at those times when the big club needs to plug an emergency hole caused by injury, trade, etc.
  • Replacement = Not Starter or Bench: Replacement level should be the average production of anyone not identified as a "starter" or "bench" player.
Think about which of these you most agree with. Or, if you disagree with all of these definitions, decide how you would identify a replacement player in this study before reading further.

Ready? Ok, go ahead and read on!

Results

Here is how those slices broke down, by league. %LgAvg is the important number with respect to replacement level. "+- Field" is a composite fielding stat and is described in a section below.

American League
National League
Slice PA PA/Plyr R/G %LgAvg OBP SLG OPS +-Field
Slice PA PA/Plyr R/G %LgAvg OBP SLG OPS +-Field
Starter1 39699 709 6.05 124% 0.367 0.476 0.843 -3.4
Starter1 45132 705 6.04 123% 0.365 0.489 0.854 2.6
Starter2 37161 664 5.65 116% 0.353 0.469 0.822 -5.1
Starter2 42042 657 5.69 116% 0.358 0.476 0.835 4.0
Starter3 35277 630 5.50 113% 0.350 0.461 0.812 -4.2
Starter3 39090 611 5.68 116% 0.357 0.477 0.834 -1.3
Starter4 32803 586 5.35 110% 0.348 0.452 0.799 -1.9
Starter4 35852 560 5.03 102% 0.343 0.442 0.785 -0.4
Starter5 30676 548 5.18 106% 0.345 0.444 0.789 -3.5
Starter5 32152 502 4.83 98% 0.340 0.428 0.768 0.5
Starter6 28307 505 4.93 101% 0.335 0.437 0.772 -0.5
Starter6 29015 453 4.77 97% 0.337 0.428 0.766 1.9
Starter7 25105 448 4.64 95% 0.329 0.421 0.750 0.6
Starter7 26118 408 4.71 96% 0.338 0.418 0.756 1.1
Starter8 22194 396 4.55 93% 0.326 0.413 0.739 0.2
Starter8 22607 353 4.62 94% 0.330 0.425 0.755 -0.6
Starter9 19489 348 4.52 93% 0.326 0.413 0.738 -0.2
Bench1 19350 302 4.55 93% 0.327 0.424 0.751 0.5
Bench1 16329 292 4.18 86% 0.318 0.394 0.711 -0.7
Bench2 16727 261 4.58 93% 0.331 0.420 0.751 0.8
Bench2 13280 237 4.08 84% 0.313 0.393 0.706 -2.6
Bench3 14357 224 4.23 86% 0.322 0.403 0.725 0.8
Bench3 11067 198 3.84 79% 0.306 0.382 0.688 0.5
Bench4 12435 194 4.18 85% 0.317 0.403 0.720 -0.1
Fringe1 9555 171 3.85 79% 0.304 0.384 0.688 3.1
Fringe1 10415 163 4.03 82% 0.318 0.387 0.705 2.4
Fringe2 7771 139 3.61 74% 0.297 0.373 0.669 -1.8
Fringe2 8648 135 3.75 76% 0.309 0.373 0.681 1.0
Fringe3 5907 105 3.31 68% 0.286 0.357 0.643 2.6
Fringe3 6959 109 3.92 80% 0.317 0.377 0.694 1.5
Scrub1 4487 80 2.99 61% 0.281 0.334 0.615 4.1
Scrub1 5292 83 3.13 64% 0.290 0.343 0.633 2.6
Scrub2 3398 61 3.33 68% 0.295 0.350 0.645 -1.5
Scrub2 4004 63 3.23 66% 0.295 0.343 0.638 2.6
Scrub3 2649 47 2.82 58% 0.274 0.329 0.604 1.4
Scrub3 2920 46 2.62 53% 0.269 0.322 0.591 3.4
Scrub4 1841 33 2.38 49% 0.268 0.285 0.554 5.5
Scrub4 2100 33 2.32 47% 0.266 0.290 0.555 -22.5
Scrub5 1237 22 1.95 40% 0.254 0.265 0.519 -8.9
Scrub5 1364 21 2.31 47% 0.263 0.293 0.556 -3.0
Other 1309 23 2.04 42% 0.248 0.281 0.529 -2.9
Other 1097 17 2.31 47% 0.253 0.305 0.558 1.9

As you can see, in both leagues, the overall amount of production, as measured by percent of league average, gradually decreases as one moves down the playing time slices. This means that teams generally gave the most plate appearances to their best hitters! Nice to see. There are certainly good performances (injuries, rookie half-seasons, etc) in some of the lower slices, and bad performances in some of the upper slices (underperforming players, hidden injuries, Adam Everett-esque defensive specialists, etc). But the overall trend is clear, and consistent with what one would expect if performance is associated with playing time. In fact, given everything that I was worried about that could go wrong with this study, I think these data look remarkably clean.

Let's now take a graphical look at offensive production and try to see what these data might say about replacement-level performance. Below I've plotted player slices against offensive performance, relative to average within each league. Vertical bands identify the different groups of slices as defined above (starter, bench, fringe, and scrub players), while the gray horizontal band denotes the range of replacement values that I've seen suggested by other researchers (73% to 80%). Colored horizontal lines indicate weighted league averages within each of the four major slice categories.
This figure does a great job illustrating how smoothly production seems to drop among player slices, which indicates to me that there is useful signal here despite potential confounds. Furthermore, I think it's remarkable how consistent the two leagues were. The primary divergences between the two leagues were among starters 4-6, two of the bench players (slices 10 and 12), as well as in the 15th player slice (fringe #3). I'm not certain what is causing the divergences at those spots, although the first separation could be related to an unusual spike in catcher numbers among the NL slices 5-7 that doesn't occur in the AL slices (though it doesn't match up to that spike perfectly, either). Alternatively, because these production estimates look only at offense, perhaps this reflects some divergence in emphasis between offense and fielding between the two leagues? More on that in a bit.

Ok, based strictly on these empirical data, here's how I would define replacement players for each of the perspectives I mentioned above:
  • Replacement = Fringe: production level typically in the range of 68-82% of league average, with a weighted mean checking in at 74% of league average offense in the AL, but 81% of league average offense in the NL. Splitting the difference puts us at ~78%.
  • Replacement = Scrub: production level typically in the range of 40-68% of league average, with a weighted mean checking in at 57% of league average offense in the AL, and 59% in the AL, or 58% overall.
  • Replacement = Not Starter or Bench: the weighted mean of all players that are defined here as either fringe or scrub players was 68 to 71% (AL vs. NL) of league averages. Over the past four years, 80593, or 11% of all plate appearances across the two leagues came from players that fell into this category.
Clearly, traditional conventions related to replacement level (73 to 80% of league average), as denoted by the gray horizontal band through the figure, best match the "Fringe" player slices. Yet some of you may have picked one of the other expectations I outlined above regarding replacement players. After all, isn't replacement level supposed to represent a minimum level of MLB player production? Why is it that we're seeing so many players performing below replacement level?

Here's one explanation: replacement level tries to describe the level of production below which players will tend to be replaced by other players. After all, if someone is producing below replacement level, you should be able to get a different player at no additional cost that will perform better. Therefore, it does make sense that the group of players who did not have playing time sufficient to rank within the top 14 or so player slices will be primarily made up of guys who didn't perform up to this minimum standard. In those cases, teams apparently decided to give more playing time and roster space to players who could actually at least perform at replacement level.

Replacement-Level Fielding

To this point, I've been looking exclusively at offensive production of position players. However, players are not given playing time simply because of their offense--defense is important as well! So how to do these same slices of players fare with respect to their fielding performance?

To look at that, I've calculated fielding stats based on conversions of the THT zone rating data using a method described here. The average, in this case, is the mean across both leagues within a particular season. Also, I only pulled these data for the players' primary positions, which means that the league totals will not sum to zero. Also, because I'm comparing across all positions here, I added a rough positional adjustment based on the difficulty of fielding different positions, as estimated by Tom Tango, which is part (though not all) of the reason that AL slices have lower-rated defense below: there are DH's in that league, and they get a hefty negative positional adjustment to put them on an even playing field with other players. Because playing time was so different across slices, the runs saved estimates are standardized to represent per-season-per-player rates.

You can see the fielding data in the table above, but here are they are in graphical form. As before, slices on on the x-axis, and colored horizontal lines indicate league averages within each major slice category:
Several things to note. First, unless there are massive park factor differences across the leagues, the NL and AL seem to value defense a bit differently. In the National League, the top slices of players in terms of plate appearances tend to be not only outstanding hitters relative to their league, but they also tend to be outstanding fielders. Below the top two slices, defense is apparently given less of a premium. Among American League clubs, however, the ~5 players per team receiving the most plate appearances tend to be below-average defenders. It's not until the bottom four starters that you tend to reach average fielding performance.

Second, with respect to replacement level, it looks like we can assume that bench and fringe players tend to be fairly average defenders. There might be a slight tendency for fringe players to be slightly above average, though it's not by more than a run or two. Furthermore, as you can see, the fielding values get rather volatile on the right side of the figure (almost certainly due to sample size--fielding stats are more volatile than hitting stats), so I think it's safest to assume that replacement level fielding is essentially the same as MLB-average fielding. This finding is consistent with similar work by Tom Tango. One thing that is very clear is that replacement players are not massively below average fielders. This puts into question the relevance of systems like BPro's Fielding Runs Above Replacement (FRAR), which sets replacement-level fielding to an approximate league minimum.

Closing Thoughts

I won't pretend that this is The Definitive Study on replacement level. There are certainly a variety of concerns that one can raise with respect to sampling, player identification, etc. Nevertheless, I think it is a good study that gives us an empirical grounding on how players perform relative to the amount of playing time they are granted by their teams and their health. From these data, we can draw some conclusions about how to best estimate replacement-level performance.

On offense, the slices that seem to best fit the description of replacement players--guys who are on the bubble of making a big league team--tended to hit at levels ranging from 68% to 82% of league average, with a rough mean of ~78%. Both of the popular standards I've seen, 73% (endorsed by Tango and Patriot) and 80% (endorsed by Woolner), fall nicely within this range. And, of course, both have their own theoretical and/or empirical justifications. But if you ask me for my recommendation after doing this study, I would recommend the lower figure of 73%. It is still a level of production above that which the "scrub" players, as I defined them here, produce. And if the primary function of a replacement level paradigm is to recognize a production threshold at which any freely available talent is likely to perform, it makes sense to me to be somewhat conservative in that threshold--that way you don't ignore production by, for example, a bench player hitting at 82% of league average, who might turn out to be difficult to replace.

With defense, it seems clear based on this and other studies that replacement players play roughly league average defense. Therefore, if you are interested in describing a player's total value above replacement level, I recommend that you follow this procedure: calculate their runs on offense relative to a hitter at 73% of position player league average runs per game, and then add to that value their fielding vs. average, as well as a positional adjustment to account for the difficulty of playing that fielder's position. More on that in coming days in my player value series.

References, Resources, and Acknowledgments

All player statistics were pulled from the stats pages at The Hardball Times.

Custom linear weights were calculated for each league via base runs using a spreadsheet created by U.S. Patriot, using initial "B" coefficients pulled from this article by Tom Tango, and using league totals pulled from Doug's Stats.

This work was stimulated, in part, by a great e-mail conversation I've been having with skyking162 the past several weeks about how to value players. Sky also provided some helpful comments on a draft of this article.

Patriot has written an extremely helpful essay on baselines at his site, along with many other excellent articles on related topics.