Table of Contents

Showing posts with label Data. Show all posts
Showing posts with label Data. Show all posts

Monday, June 20, 2011

BPro's nFRAA: 2008-2011 Averages

Data-dump post!

I pulled together three+ year (2008-2011) nFRAA averages for all players who have played in 2011 and have at least 150 PA's during that time span. The left side of the spreadsheet just reports those totals, and a per-season rate (per 700 PA's instead of innings, since that's how BPro lists it in their spreadsheet).

I also did a very unscientific "regression" of sorts on the data. The values reported at BPro are already regressed within each season based on their performance. However, if we're trying to infer something about talent level--which is the point of a 3-year average in my view--there is even greater uncertainty than with a performance/value estimate. Therefore, I've required that players effectively have 2.5 seasons of data here before we use their straight up 3+ year average. If they have less, I regress toward average in proportion to how much data we do have. Why 2.5 seasons? It's just a wild-ass guess, but it feels about right based on the data. If you don't like it, that's ok--this is why I'm reporting all of the data. You can use whichever per-700 PA stat you wish, or come up with your own way of making an adjustment. Just please be very wary of the raw per-700 PA numbers for guys with less than a few seasons of data.

One qualifier: I was a little sloppy in that I assumed all PT occurred at the player's primary position, as listed in BPro's table. Obviously, this is not true, and for some players it will result in erroneous estimates. For most players, however, I think it's probably close enough for my purposes.

Another: I'd hope this goes without saying, but please don't just use this as the be-all end-all estimate of a player's fielding performance. I would tend to be fairly skeptical of Jose Lopez's performances, for example, given how different they are from other measures and his reputation.  But it's good data, and worth considering as you evaluate a player's fielding skills.

Wednesday, May 20, 2009

New Toy at FanGraphs: Pitch Type Linear Weights

Today, FanGraphs debuted what they're calling Pitch Type Linear Weights. They're only on the player pages, for now. Appelman:
What this section does is it uses linear weights by count and by event and breaks it down by each pitch type so you can see in runs the actual effectiveness of each pitch.
So, if you make an out on a pitch, you get credited for the linear weights value of an out. If you allow a single on a pitch, you get penalized the linear weights value for that event. And it's broken down by count as well. So, you get a strike on a hitter, you are credited for typical value of a strike. Allow a ball, you're penalized for the typical value of a ball.

These are really new stats, so it's hard to know how to interpret them. So, to start, I pulled career totals for all current Reds pitchers (minus Janish...). All data are average run values per 100 pitches of that type (that's what the "/C" means). My comments are in the table.

Player wFB/C wSL/C wCT/C wCB/C wCH/C Comments
Arroyo -0.76 1.44 0.07 0.79 -0.64 Breaking stuff is strong, but high-80's fastball gets hit. Has to be thrown to set up the other pitches, though.
Burton 0.71 -2.42 1.33
3.13 You see that cutter, which he throws about half the time. Should he ditch the slider?
Cordero 0.44 1.52

1.38 Probably best slider on the staff--that's not a surprise, right?
Cueto -0.09 0.4

-1.74 In 2009, Cueto's FB & SL are +1.39/C and +3.18/C, but his change is -8.13 runs/C. Wasn't that something he was working on?
Harang 0.31 0.22
-0.57 -2.22 Harang (like Cueto) is primarily a two-pitch guy, so I guess it's not surprise that his 3rd and 4th pitches don't look so good..?
Herrera 0.75 1.44
-1.58 -0.48 I have no idea how they're classifying his screwball here. Given this and the sample size, I'm ignoring this one
Lincoln -0.39
-3.17 0.85 4.71 Last year, Lincoln's curve ball was +1.29 runs/C. This year, -3.13 runs/C. Location? Any pitchf/xers want to look?
Masset -0.38 1.26
1.24 -1.27 Masset throws his fastball 70% of the time. Should he try using his breaking pitches a bit more often (60/20/20 instead of 70/15/15)?
Owings 0.44 -0.84

-0.91 Some guys throw their fastball to set up the offspreed stuff. This makes me think Owings uses to the offspeed to set up his 89 mph FB.
Rhodes 0.93 0.89

1.52 The change-up is little used (low sample), but he gets by just fine with his FB/SL combo.
Volquez -0.26 -0.72
-1.03 0.76 His early years are weighing him down--his '08-'09 data are much better.
Weathers 0.06 0.84

0.75 He's has gone from using his change-up 3% of pitches to 7% last year and 10% this year, with good results. Simultaneously, he's relying less on his slowing FB.

Your thoughts are welcome!

Sunday, May 17, 2009

American League Update - Through 5/16/09

Tonight, I'm going to take a look at the junior circuit.

Division Roundups


Quick notes on the stats below: All runs data are park adjusted using Patriot's park factors. PythWins is from Hardball Times, and so is PythagenPat. XtrapWins is the wins that the teams' current winning percentage would provide at the season's end. Wif500 is what the team's record should be if they win half of their remaining games. %for90W is the winning percentage the teams will need to get to 90 wins at the end of the season.

East
Team W L PCT GB RS* RS/G* RA RA/G* PythW XtrapW Wif500 %for90W
TOR 25 14 0.641 0 219 5.6 167 4.3 24 104 87 0.528
BOS 22 15 0.595 2 198 5.4 180 4.9 20 96 85 0.544
NYA 19 17 0.528 4.5 195 5.4 210 5.8 17 86 82 0.563
TB 18 20 0.474 6.5 203 5.3 189 5.0 20 77 80 0.581
BAL 16 21 0.432 8 186 5.0 216 5.8 16 70 79 0.592
As much as hate to admit it, this is probably the most interesting division in baseball right now. The Red Sox and Yankees should have very good teams, and the Rays--despite their record--have played every bit as tough as the rich teams. But the Blue Jays have been the class of the division thus far, sporting both the best offensive and the best defensive (pitching + fielding) performances. You have to expect some of Toronto's hitter's to fall back to earth--none of Barajas (0.820), Hill (0.927), and Scutaro (0.869) have OPS'd over 0.800 before this season. And they've gotten unreal pitching performances from Romero and Cecil. It will be interesting to see how long they can hold off their division rivals.

Central
Team W L PCT GB RS* RS/G* RA RA/G* PythW XtrapW Wif500 %for90W
DET 19 16 0.543 0 190 5.4 163 4.7 20 88 83 0.559
KC 19 18 0.514 1 166 4.5 151 4.1 20 83 82 0.568
MIN 18 19 0.486 2 181 4.9 197 5.3 17 79 81 0.576
CHA 15 20 0.429 4 135 3.8 162 4.6 15 69 79 0.591
CLE 14 24 0.368 6.5 198 5.2 218 5.7 17 60 76 0.613
Another rather evenly matched division, though perhaps without the quality of the east. The Tigers have had the best offense, but the Royals have had the best defense--thanks in no small part to the spectacular first month and a half for Zach Greinke. Nice also to see Brian Bannister having a nice bit of success in the early goings. The Tigers have their own surprise--Brandon Inge is OPSing 0.954--but they have a lot of guys who aren't really hitting yet (Granderson, Guillen, Ordonez) and thus might actually improve even as Inge cools.

West
Team W L PCT GB RS* RS/G* RA RA/G* PythW XtrapW Wif500 %for90W
TEX 22 14 0.611 0 200 5.6 174 4.8 20 99 85 0.540
LAA 18 17 0.514 3.5 177 5.1 179 5.1 17 83 82 0.567
SEA 17 20 0.459 5.5 152 4.1 179 4.8 16 74 80 0.584
OAK 13 20 0.394 7.5 145 4.4 162 4.9 15 64 78 0.597
The Angels have had as much bad luck and tragedy as any team I can remember, but still are right at 0.500 thanks their great depth. The Rangers, though, have taken advantage of the weak division, riding hot starts by Ian Kinsler, Michael Young, and Kevin Millwood to a nice little division lead. Not as impressive is the offense of Ken Griffey Jr.'s Seattle Mariners. Thank goodness they at least have Russell Branyon.

Team performance breakdown

Below is a kind of power ranking of teams based on their component statistics. I estimate team runs scored using linear weights (actually, FanGraphs does, though I park-adjust it), pitching performance based on FIP (I do park-adjust the HR's), and fielding based on UZR. I use those numbers to estimate team runs scored and runs allowed, and then use PythagenPat to estimate expected winning percentage. It's not perfect, but this should give us another look at team performance that gets beyond actual wins, losses, and runs scored or allowed.




Offense Pitching Fielding Overall
Rank Prev Team OBP SLG wOBA* wRC* ERA FIP* K/9 BB/9 HR*/9 FIPRuns* bUZR THT+/- DER ExptW%
1 - Blue Jays 0.359 0.460 0.356 225 3.97 4.23 6.8 3.1 0.9 181 6.2 24.8 0.720 0.619
2 - Rays 0.346 0.452 0.357 214 4.83 4.83 6.5 3.8 1.2 191 16.3 0 0.689 0.595
3 - Royals 0.335 0.421 0.334 174 3.63 3.76 7.5 3.4 0.7 149 -3.8 -4.8 0.685 0.560
4 - Rangers 0.335 0.496 0.357 201 4.66 5.02 5.3 3.4 1.2 192 15.2 19.2 0.712 0.560
5 - Tigers 0.340 0.426 0.337 170 4.31 4.29 7.3 3.6 1.0 160 4.7 8.8 0.704 0.544
6 - Red Sox 0.364 0.451 0.351 204 4.82 4.67 7.4 4.0 1.1 185 -12 -7.2 0.677 0.517
7 - Yankees 0.350 0.467 0.358 211 5.41 5.28 7.2 4.2 1.4 203 -4.1 4 0.691 0.509
8 - Indians 0.355 0.420 0.345 204 5.62 5.03 6.6 4.0 1.3 199 -2.8 -12.8 0.673 0.505
9 - Angels 0.344 0.414 0.339 174 4.70 4.61 6.0 3.4 1.0 171 -4.1 -10.4 0.690 0.497
10 - Twins 0.347 0.417 0.338 183 5.20 4.91 6.2 3.0 1.4 192 -0.1 -0.8 0.696 0.476
11 - Orioles 0.342 0.437 0.340 182 5.51 5.11 6.6 3.3 1.5 196 -9 -20 0.666 0.444
12 - Mariners 0.305 0.378 0.305 138 4.25 4.46 6.8 3.7 1.0 178 10.6 -4.8 0.689 0.411
13 - White Sox 0.317 0.385 0.306 132 4.71 4.05 7.0 3.8 0.7 150 -14.4 -19.2 0.667 0.398
14 - Athletics 0.309 0.338 0.294 116 4.12 4.57 6.2 3.6 1.0 167 -0.2 -11.2 0.689 0.333
The only major disparity I see is with the Indians, who have apparently hit and pitched well enough to expect a 0.500 record, and yet are somehow 10 games below 0.500. Their true pythagorean record is 17 wins, but indications here are that they've given up quite a few more runs than expected. Are they just unlucky, or are the methods missing in this case? Part of the answer might be with the fielding estimates: bUZR rates them as only slightly below average, but THT's batted-ball based plus/minus stat (as well as straight-up DER) pegs them at second-worst in the league. Perhaps bUZR is overrating their defense?

What's up with the Athletics offense? They add Giambi, Holliday, and Nomah, and yet have a wOBA under 0.300? Beane must be blowing a gasket.

  • Top hitting teams (wOBA*): Yankees, Rays, Rangers, Blue Jays, Red Sox (holy AL East!)
  • Top pitching teams (FIP*):Royals, White Sox, Blue Jays, Tigers, Mariners
  • Top fielding teams (bUZR): Rays, Rangers, Mariners (maybe), Blue Jays, Tigers
  • "Expected" leaders: Blue Jays, Royals, Rangers
  • "Expected" Wild card: Rays
Blue Jays rank in top five in hitting, pitching, and fielding. What will they have to do to keep their lead the rest of the way?

Saturday, May 16, 2009

How much should we expect from a 1st-round pick? How have the Reds done?

As a follow-up to yesterday's post...how much should we expect from a 1st-round pick anyway?

Thus far, Erik has posted a three-year sample (1994-1996). This is probably not enough to get a reliable estimate, but to get some idea, here's a quick pivot table of his data (note: I am setting all negative WAR values to zero, as Rany Jazayerli did in his study, and as Tango recommends--players who don't make it to the big leagues should never be considered more successful picks than players who do make it, even if they stink upon arrival):

Position N WAR/player
1B 5 4.5
2B 1 7.4
3B 4 6.7
3B 1 0.0
C 6 3.5
OF 22 3.7
SS 10 4.4
PosPlyr Totals 49 4.1
P 50 2.8
Grand Total 99 3.5

I am showing the position-by-position totals, though I probably should not be--many are based on just a single player! It's no accident that you see more SS's and OF's than other positions, though: that's where the best amateur players will tend to play on their teams. Another interesting tidbit is that almost exactly half of the first round picks during 1994-1996 were pitchers, even though pitchers averaged just 64% of the production of position players.

Anyway, from 1994-1996, the average player taken in the first round returned 3.5 WAR in his career. An average MLB player will produce 2 WAR/season, so this is essentially saying that the baseline for an average first round pick is one and a half seasons of MLB-average performance. Not very good. It's no wonder that baseball drafts have a reputation for being a crapshoot, even in the first round.

I'm not sure if 3.5 WAR is a reliable figure. My guess is that it's within a half-win or so of the true mean, but it's hard to know at this point without more data. Furthermore, it's questionable whether we should even use a mean here--the data are not normally distributed, and, in fact, the median WAR is actually zero!

But if we run with it for now, how have the Reds done? Let's break it down by 5-year increments. I'm including all players shown in yesterday's post, again setting anyone with negative WAR to zero:

Years Players TotalWAR WAR/Player
65-69 16 75.5 4.7
70-74 17 12.2 0.7
75-79 20 15.2 0.8
80-84 20 40.5 2.0
85-89 10 72.3 7.2
90-94 5 19.5 3.9
95-99 5 16.3 3.3
00-05* 8 4.0 0.5
The reason there were so many more players in the 60's, 70's, and 80's is that I included all of various drafts that existed in those years--January, August, etc. I was concerned that this was the cause of the rather striking awfulness of the 1970-1984 drafts, so here are the data focused exclusively on the June draft (supplemental picks are still included, as they are in Erik's studies; I'm also setting players like Sowers who did not sign to zero, as those are wasted opportunities):

Years Players TotalWAR WAR/Player
65-69 5 62.3 12.5
70-74 5 0 0.0
75-79 6 7.5 1.3
80-84 6 9.4 1.6
85-89 4 72.3 18.1
90-94 5 19.5 3.9
95-99 5 16.3 3.3
00-05* 8 4 0.5
My take: the Reds did extremely well in the first several years of the amateur draft, but then were downright awful from 1970-1984. And then Barry Larkin happened in '85. And ever since Barry, the Reds have been just about average in the performances of their first round selections. There is an apparent downturn in the 2000's, though my feeling is that Jay Bruce and (hopefully) Homer Bailey will turn that around when all is said and done. Overall, the Reds have averaged 4.3 WAR/player taken in the first round of the June draft since its inception. Ignoring Larkin (which is probably not fair), they're averaging 2.8 WAR/player.

This is comforting. I know we (or I, at least) as Reds fans tend to be pretty down on our team's ability to draft talent. But I think the message here is that they historically have been close to average, and may even be slightly above average thanks to the Larkin pick. Let's hope that in the coming years, they can improve on that a bit--as a small market team, it may be that the Reds can't afford to be merely average in the draft.

Friday, April 17, 2009

Friday Night Fungoes: Larkin, run estimators, CHONE, and the Dunning-Kruger Effect

Looks like I'm going to the Curve game tomorrow night. Whether I'll post a report may depend on how they do...they've yet to win this year! Weather should be nice, though: 72 degrees & sunny.

Does Larkin Belong in the Hall of Fame? Revisited

I can't remember if I linked to this or not, but even if I have it's worth linking again: Rally has posted season-by-season WAR estimates for all players in the Retrosheet era. He also has a top-300 ranking, so we can look at the best of the past 50+ years using these numbers.

Rally's data include offense, defense (including turning double plays, etc), baserunning, and era-specific position adjustments. This is similar to what I tried to do in my piece on Larkin, but better because of the baserunning & especially the era-specific position adjustments. Here is how the shortstops I included in my Larkin study pan out in Rally's WAR data, plus a number of others who came up in discussions following my Larkin piece:


MyWAR RallyWAR
Alex Rodriguez 82.4 96.5
Cal Ripken+ 73.8 91.2
Robin Yount+ 62.1 75.9
Barry Larkin 57.6 70.1
Ozzie Smith+ 45.3 67.6
Alan Trammell 53.7 66.8
Derek Jeter 49.5 62.4
Ernie Banks+* 54.8 59.2
Luis Aparicio+ 30.9 50.5
Toby Harrah
47.1
Omar Vizquel 28.0 45.2
Bert Campaneris
44.5
Normar Garciaparra
43.7
Tony Fernandez
43
Miguel Tejada 39.2 40.9
Jay Bell
35.5
Davey Concepcion 26.0 34
Mark Belanger 26.2 32
Edgar Renteria
31.9
Chris Speier
24.8
Bill Russell
24.4
Rick Burleson
21.3
Freddie Patek
20.4
Steve Sax
19.8
Larry Bowa
18.9
Bucky Dent
12.4
Don Kessinger
7.2
Tim Foli
3.5

Some players got a big boost in their rating, like Ozzie and Aparicio, once you include baserunning (I only included SB's & CS's) and double play turning. But as you can see, the results are more or less the same as far as Larkin is concerned: probably the 4th best overall, and 2nd-best pure shortstop in the Retrosheet ERA, at least based on total contributions to their ballclubs relative to their league.

Career-level WAR accumulation isn't the be-all, end-all of hall of fame voting, as peak performance is also important. But in Larkin's case, career WAR is crucial. No one disputes that he was brilliant player when healthy. The knock on him is that he didn't play enough due to all of his injuries. These data clearly indicate that his total contribution, including playing time, was among the best in baseball history at his position.

Rally has Larkin as the 30th-best position player of the Retrosheet era. I know he almost certainly will not be a first-ballot Hall of Famer, but he probably should be.


Why I'm trying to stop using OPS

Colin followed up his study posted last week on run estimators with an improved method. This time, instead of looking at half-inning or even game-level combinations of team offense, he instead focused on identifying the average value of particular offensive events to games. His methodology was to take matched games--games that had the same numbers of major counting events, but that differed in how many of one specific event they contained.

For example, he might have a game with 5 singles, 3 doubles, 1 homer, and 3 walks, and he'd compare that to a game with 5 singles, 3 doubles, 1 homer, and 4 walks. Finding the average difference in runs scored between pairs of games like this would tell you the average value of a walk in runs. He then compared those actual differences in runs scored to the expect difference in runs scored according to a variety of run estimation mechanisms.

The results? Linear weights-based methods did the best. This includes his "house" linear weights (which he kindly shares), as well as manipulations of linear weights like wOBA. A bit behind them were GPA (aka 1.7 OPS), Base Runs and BPro's EqR, followed a bit more distantly by Bill James' Runs Created. The worst of the bunch were the OPS-based methods, as well as the even-more-horrible Total Average (bases/outs).

This is strong evidence that we should more or less stop using OPS to evaluate hitters. It's unnecessary, given how easy wOBA is to calculate. Is it better than batting average? Sure, of course. But it misses badly enough and often enough that we should really move past it. It's a tough habit to break, but it's time to wOBA, folks.


CHONE is a really good projection system

Matt has a fairly exhaustive projection roundup here. He notes that each system seems to have its own strengths, but often also some weaknesses:

--CHONE was the best at projecting most things.

--PECOTA was very close behind but had some systematic biases, specifically for speedy players' BABIPs, which ZIPS struggled with as well.

--ZIPS is behind the other systems, except it does quite well with projecting the three true outcomes for players over 35.

--CHONE does better with older players in general, since its specialty is aging curves, but PECOTA does better at finding comparable players for younger players for whom less data is available (unless they fall into the speedster category).

--OLIVER clearly contends and even takes the lead at some things--especially at projecting hitters with lower homerun totals and other players significantly affected by park effects. However, OLIVER under-projects walks and strikeouts systematically and over-projects homeruns systematically, and could probably be improved by adjusting how those outcomes are computed.

The nice thing about this is that we can use this information to give more or less weight to a given projection system when it differs from others in predicting a given player's performance based on the sort of player we're looking at. Or, we can do what I've essentially decided to do around here, which is to just use CHONE. :)

It's worth noting that Matt's is just the latest projection roundup in which CHONE did particularly well. Whether it will continue to do so in the future is an open question, of course, but the data suggest that it's as good as they come.


The Dunning-Kruger effect

JC posted about this terrific psychological concept: that people incompetent in a particular discipline will massively overestimate their competency in that discipline. That's pretty much the definition of a baseball fan, isn't it? :)

I'm jesting, mostly. You certainly see arguments between baseball fans who really know their stuff and baseball fans who just think they know their stuff. And I tend to think that most of what you hear on talk radio (sports, or otherwise) involves people who fall into the latter category rather than the former category. And, of course, I tend to think that on at least some issues (some areas of biology, some areas of baseball research, etc), I fall into the reasonably competent category.

But the great part of this is that the Dunning-Kruger effect predicts that we'll have a very hard time being able to tell whether we're competent or not...because the more incompetent we are, the less we'll realize it! :)

Saturday, April 19, 2008

Friday Night Fungoes: Dan Fox, Projections, and the Future, man!

Pirates sign Dan Fox

The Pirates continue to their overhaul of their front office by hiring Dan Fox, aka Dan Agonistes, who has written at both Baseball Prospectus and The Hardball Times. His official title is Director of Baseball Systems Development, so I am not sure if this is the same job that Tango advertised a while back--but it might be. The responsibilities sound similar--responsible for "integrating the array of quantitative and qualitative information in a way that makes both even more instructive."

Obviously I'm thrilled for Dan. This no doubt is incredibly exciting for him, and I'm sure the Pirates made a heck of a hire to get him. But this is kind of a bummer for me in a few ways:

1. The Pirates as an organization probably just got a bit better. As a Reds fan, that's not a good thing.
2. It means that we as a community lose access to an exceptionally good, broadly trained, and very resourceful analyst.
3. Dan Fox is one of the three main reasons that I decided to continue to subscribe to Baseball Prospectus this year. So now I'm down to PECOTA and Kevin Goldstein's work on prospects, though I'll give a nod to Nate Silver's occasional high-quality article as a secondary reason.

I have to say, I might start paying a little bit more attention to the Pirates in coming years. I like what I see and hear from their front office since they hired Neal Huntington, and I will be living minutes from their AA-franchise. I don't know if I could ever really leave the Reds and root for an NL Central foe, but you never know.


Projection System Showdown

I missed this when it happened, but (with a hat tip to studes) Tom Tango recently did a rather careful study of projection accuracy for hitters in the 2007 season.

Before I get into Tango's study, I do want to point out that his results are similar to those reported by Nate Silver last year (hitters & pitchers here). There are a few subtle methodological differences, though, and Silver focused more on rank order of the systems than the actual impact of any differences seen. The latter is probably the most important point, as you'll see...

Anyway, Tango followed two basic steps:

1. Adjust forecasted league-average OPS to match actual 2007 league-average OPS.
2. Determine the average deviation between expected OPS value and actual OPS value (i.e. the average residual).

He reported overall average error, as well as four groups based on total multi-year MLB plate appearance totals: high (grizzled veteran regulars), medium, low, and rookies (made debut in 2007). The results can be hard to sift through in that very long thread (starts on #29, ends on #103), so I opted to whip up a quick graph (hope Tango doesn't mind). Note: I inverted the y-axis to read more intuitively--the "higher" the points are (vertically), the better, because they show less error.You may want to open the graph in another tab to view it. I tried to stretch it to separate the lines as much as I could, within reason. Here are the primary conclusions of Tango's study:
  • There isn't much difference between the different projection systems.
    • Overall, the best system (PECOTA) provided OPS measurements that were, on average, 0.069 off from the actual values. That means that a player projected to have an 0.800 OPS will, on average, have an OPS between 0.731 and 0.869. Not particularly accurate, really.
    • The worst system, aside from just projecting everyone to be league-average, had an average error of 0.075 OPS points. So, again, an 0.800 OPS player will, on average, have an OPS between 0.725 and 0.875. Not particularly worse than the "good" system.
    • Free systems CHONE and ZiPS were a whopping 0.001 OPS units "worse" than the much-lauded PECOTA. Marcel was only 0.002 OPS units behind. And they beat for-cost systems like Bill James' and Shandler's.
    • The best thing to do is to take an average across all nine projection systems. But even then, you only get an extra "point" of OPS accuracy over PECOTA and 3 points over Marcel. Again, BFD.
    • Despite all the additional information that systems like CHONE and PECOTA have about minor league players, they're only marginally better than Marcel when projecting players with few MLB plate appearances.... and in those cases, Marcel can only projects league average.
  • Forecasting playing time is hard.
    • The Fans' community forecast didn't do better (or worse) than the objective systems in projecting OPS. However, they DID do substantially better when projecting playing time.
    • Sal Baxamusa and Tango think that a system combining the Fans' projections of playing time and Marcel's simple projections of performance would probably outcompete anything out there in a head-to-head contest.
Overall, in terms of accuracy, we can do just as well--if not better--using CHONE, ZiPS, or Marcel (all free) as we can using the for-cost PECOTA, Bill James, or Shandler projections. I think I knew this was true in theory, but always figured I was still gaining something by using PECOTA over something like Marcel. Now, I'm feeling doubtful...

Two caveats to this study: 1) it only looks at hitter projections, and (most importantly) 2) it only looks at 2007 data.

Even so, I guess I'm not seeing a compelling reason to use PECOTA at this point. So, on my list of reasons to renew my BPro membership, I guess I'm down to Kevin Goldstein. ... I enjoy Kevin's column, but I'm not sure if that'd be enough if I were renewing today.


The Future of Sabermetrics, man!


There were two Hardball Times articles this week that I thought did a tremendous job of setting the stage for a lot of the work we're likely to see in the coming years:
  • Sal Baxamusa did a great job of putting forth an organizational scheme to help us understand baseball player valuation. I find those sorts of charts to be tremendously effective ways to organize my thoughts. The first chart he presented is basically what I worked through in my player value series last winter. What Sal does is take it the next step and show how we might better understand pitching, especially in light of the influx of pitchf/x data. We could do similar charts for hitting, baserunning, and fielding--all of which might be influenced by the "f/x" line of data (hitf/x, fieldf/x, etc).
  • Mike Fast put forth a litany of questions that remain to be answered in baseball research. What makes his list useful is that all of these may be answered, at least in part, via the use of ball tracking data. Therefore, we may see tremendous progress on most of these within the coming year or two. Exciting times are ahead...


Signing Young Players Early

Skyking has a great post breaking down the contract extension of Evan Longoria with the Rays. He finds, as did Tango, that Longoria is almost certain sacrificing a huge amount of money by signing this deal today. On the other hand, he just guaranteed that he'll make $17.5 million, even if he goes and has a career-ending injury tomorrow.

As Sky points out, that first $17.5 million would probably worth much more to most of us then the $50 million that might come after it. He made a similar point after the Granderson signing. I'm sure that I'd have a hard time not taking the $17 million today, even if I knew I'd probably make three times that if I went year-to-year, simply because the risk could be so catastrophic.

I would not be surprised to see the Reds try to work out similar deals with Joey Votto and (especially) Jay Bruce this season or offseason. In fact, I expect them to do it: it makes too much sense NOT to do it. I'm not sure about Cueto, Volquez, and Bailey, though...pitchers are so darn injury prone that it's pretty dangerous to sign them long-term. Still, the kinds of discounts that teams are getting when signing players like Granderson, Tulowitzki, and now Longoria are so extreme that at some point it will make sense to lock up young pitchers as well.


More MLB customer service issues: Blackouts


Last week, I mentioned a series of recent issues related to MLB's tendency to not necessarily keep the best interests of their fans in mind. I forgot a big one: blackouts. Maury Brown posted a BPro article (I guess that's a reason to keep my subscription) about TV blackouts last weekend that is worth a read, though I frankly still just don't understand what MLB is thinking. Here's my favorite quote from it:
Only in baseball would there be a collective head nod to the idea that it's good business practice to restrict consumers' access to your product.
Awesome.

Also, I just became aware of a blog that has been started up in protest of MLB's blackout policy. This will be a good way to keep tabs on the issue. FWIW, they report that there may be some tangible effort to at least remove the most absurd blackout restrictions in the near future.


Reds Shut Down Scout's Blog

Finally, this week saw the birth and death of a blog by one of the Reds' scouts, Butch Baccala. He apparently was shut down due to concerns about giving out information that might put the Reds at a competitive disadvantage. To my eye Butch clearly knew where that line was and would not cross it. The Reds apparently disagreed.

I see this as yet another example of the Reds' penchant for absurd secrecy. And frankly, I find that both annoying and disappointing.

Dave from Louisville thinks I feel this way because I'm an academic. Thank goodness I'm never going to have to work in the corporate world, because it sounds awful.