For all your fancy-pants statistical needs.

Praise for The Basketball Distribution:

"...confusing." - CBS
"...quite the pun master." - ESPN
The Miami Heat, projected:

Dwayne Wade -
ORTG: 110.3
unadjusted Usg: 33.2%
DRTG: 106.1

Lebron James -
ORTG: 120.1
unadjusted Usg: 32.9%
DRTG: 101.1

Chris Bosh -
ORTG: 112.9
unadjusted Usg: 27%
DRTG: 109.9

Mario Chalmers -
ORTG: 105
unadjusted usg: 17.3%
DRTG: 104

Ilgauskas -
ORTG: 106
unadj. usg: 18.1%
DRTG: 104


Best Offensive Players

This adjusts a player's Offensive Plus Minus for their teammates.
We do this based on a Linear Regression that includes:
-a player's contributions (Dean Olivers Offensive Rating and Usage%)
-his teammate's contributions (based on four estimates)

The four estimates of his teammates are as follows:

1) Total Offensive (ONcourt)= TeamORTG(ONcourt)-playerUsg%*playerORTG
2) Average Offensive (ONcourt)= (TeamORTG(ONcourt)-playerUsg%*playerORTG)/(1-usg%)
3) Total Offense (Season) = TeamORTG(season)-playerusg%*ORTG*min%
4) Average Offense (Season)=(TeamORTG(season)-playerusg%*ORTG*min%)/(1-usg%*min%)

We take these variables versus a player's On-Court Offensive Efficiency, to give us the following regression (which has an R^2 value of >.99 since the x values are based on splitting the y value up):

Source Value Intercept -5.674 USG% 33.516 ORtg 0.193 TMO-1 0.302 TMO-2 0.554 TMO-3 -0.009 TMO-4 0.011





Since Plus-Minus is the more inclusive stat, we will take that stat and adjust from there. We simply take the player's Offensive +/-, subtract the TMO variables (times the coefficients), then add the average TMO variables times the coefficients.

The results can be found here: http://dl.dropbox.com/u/241759/adjusted%20offense.pdf

Box-Out%

This is fun:
http://dl.dropbox.com/u/241759/boxoutpercent.pdf - for all NBA players in 2010 with 25+ minutes per game, who have played 40+ games.


I created a stat that shows us some representation of the % of the time a player gets a rebound, versus the % of the time their man gets the rebound.

Their man is assumed to be an average player whose rebounding percent (offensive reb% while player in question is on defense, etc) is ~80% the rebounding percent at the player's position, and 20% the average rebounding percentage of all other players. For example:

Oklahoma City's Russell Westbrook (great offensive rebounder for a point guard) collects 6% of all available rebounds while he is on offense. His 'man' is likely to be a point guard, but there is a chance (here estimated to be 20%) that his man will be a different player.
Point guards collect ~10.2% of all available defensive rebounds, while the rest of their team gets on average around 16%. So, (80%*.102)+(20%*.16)=.1016, or ~10.2%.
His 6% versus his 'man's' 10.2% gives him an offensive boxout% of 37.1% by dividing like this:

Westbrook's 6% Offensive rebounds / (His 6% Offensive Rebounds + 'Man's' 10.2% offensive rebounds) = 37.1%

Finally, the two are averaged. This gives us the total percent of rebounds the player gets, versus their 'man' (a weighted average does not do this).

Adjusted Offensive Efficiencies

Here we estimate each player's Adjusted (against average competition) Offensive Efficiencies

(The formula is simply (Raw Offensive Efficiency x Average Team Efficiency)/ Opponent's Defensive Efficiency. This is based on the assumption that Raw Team & Player Offensive Efficiency can be described as Real Offensive Efficiency x Real Opponents' Defensive Efficiency / League Efficiency average).

The results above 20% of possessions used are here: http://dl.dropbox.com/u/241759/adjusted%20offensive%20efficiencies.htm


Monte Carlo Methods

Since my blog is so ugly, I don't really like updating it. But here's a quick rundown of my monte carlo simulation method.

1) The Ken Pomeroy simulation.
https://dl.dropbox.com/u/241759/bracket_simulation.htm

Ken Pomeroy has set up his statistics in a simple way to find point margin (see one of my way old posts). For several reasons that I have mentioned on this site, I do not agree with his % chance of win statistic, and so instead I stick to Dean Oliver's (the one I learned in statistics class). Simply, in excel, I tell it to look at the normal distribution of the expected outcome of the two teams. Assuming a standard deviation of 10.9 points (which is roughly what we find from most teams in Pomeroy ratings, and the number found by the LRMC paper), we tell the computer:

=Normdist(x, 0, 10.9, 1)
where X is the expected point margin.

Then, we tell the computer to create one random tournament. For each game, the computer generates a random number between 0 and 1. If the value surpasses the better team's win%, (i.e., if it chose .91 while the better team's win% was .9) -- the worse team moves on.

Then by setting up a macro, I record the number of times each team makes it to which round. Then we simply divide the number of times each team makes it to any given round and divide it by the total number of trials to get % chance that a team will make it to whichever round of the tournament.

2) The LRMC simulation
http://dl.dropbox.com/u/241759/lrmc_montecarlo.htm

This uses the same computer program, but different statistics.
Unfortunately, the LRMC does not post anything that we can convert to point margin or win probabilities. So we have to estimate point margin from each team's ranking. I took Jeff Sagarin's predictor rating of all 347 teams by ranking, and used the LRMC's ranking order. While this certainly has some inaccuracies, I should say that this method was by far my best for the vast majority of the tournament. Then, we just convert his numbers into a win probability by subtracting one rating from another (this gives us predicted point margin).

Hope that answers any questions!

Who needs BracketScience?

By taking the simulated # of wins (via kenpom.com's numbers) and average wins for a given seed, we can rate who will have the best performance above what their seed predicts.

Here are the results for the field.

Teams in The East Likely to Give Kentucky Trouble

Quick Math on the bracket

Here is my preliminary results of my bracket simulation, based on stats from Kenpom.com

http://dl.dropbox.com/u/241759/MidwestWest.html

(not yet adjusted for teams' consistency)

Game-Changers for NCAA Tourney Teams

Here we'll be taking a look at what is likely to alter a team's predicted final score (based on Ken Pomeroy's rankings - http://kenpom.com and the LRMC's rankings - http://www2.isye.gatech.edu/~jsokol/lrmc/)


The two things we'll be doing:
1)We'll describe what part of Ken Pomeroy's Four Factors stats (of a team's opponents) affects the predicted outcome. Relies on = is ranked high in, does not rely on= is ranked low in. All these numbers can be found on the team pages at Kenpom.com
2) We'll measure a team's predictability (in terms of consistency of actual versus expected point margin with an average value of 10.9).

-Duke: (#1 Pomeroy, #2 LRMC, #2 bLRMC)
-Predictability: +1.1 points above average
-When opponents' offense relies on heavy free-throw shooting, Duke fares better. (Correlation of +.34)
-When opponents' defense DOES NOT rely on heavy defensive rebounding, Duke fares worse. (Correlation of -.28)
`
-Kansas: (#2 Pomeroy, #1 LRMC, #1 bLRMC)
-Predictability: -.2 points above average
-When opponents' offense relies on good field-goal shooting, Kansas fares worse. (Correlation of -.43)
-When opponents' offense relies on heavy offensive rebounding, Kansas fares better. (Correlation of +.33)
-When opponents' defense DOES NOT rely on field-goal percentage, Kansas fares worse. (Correlation of -.34)

-Wisconsin: (#3 Pomeroy, #13 LRMC, #9 bLRMC)
-Predictability: +.3 points above average
-When opponents' offense relies on good field-goal shooting, Wisconsin fares worse. (Correlation of -.31)
-When opponents' defense relies on field-goal %, Wisconsin fares better. (Correlation of +.24)

-Ohio St: (#4 Pomeroy, #7 LRMC, #5 bLRMC)
-Predictability: -.5 points above average
-When opponents' offense relies on good field-goal shooting, Ohio St. fares better. (Correlation of +.24)
-When opponents' defense DOES NOT rely on field-goal %, Ohio St. fares worse. (Correlation of -.30)



UPSET WATCH

NCAA tournament upset watch: Murray State and Pittsburgh are the two teams who will likely be mis-seeded the worst: http://dl.dropbox.com/u/241759/upsets.htm

Point-margin-based Chance of win.

While I think using the four factors can give us a much better picture of point-margin (and therefore, chance of win), let's just look at the 2nd step right now: deriving chance of win from expected point margin.

The Log5 formula used by many people (including Ken Pomeroy) to determine a team's chance of win is fairly accurate. It is based on fitting a model to theoretical results.

Slightly more accurate, I believe, is the LRMC (logistic regression markov chain) steady-state formula, which does the same thing, just to a much higher degree of accuracy; steady-states offer an actual theoretical explanation for the numbers based on team play rather than simply the normal distribution.

For example:
Duke's chances against Maryland, assuming a 2-pt-win-

Log5: 61%
LRMC: 59.8%

Huge difference, huh?

Finally, I must throw in my two cents: empirically, I think it is viable to say that specific teams play more consistently than others. In that way, we can alter win probabilities based on standard deviations of actual minus expected point margin (which explains the basis for this site's creation). Using those numbers (from Kenpom.com), we see that:

Duke's standard deviation of actual minus expected point margin is 9.77.
Maryland's standard deviation of actual minus expected point margin is 10.98

By Duke's numbers alone, we see their chance of win as being 58.1%
By Maryland's numbers alone, we see their chance of win as being 42.77%

By averaging these two values in their context (.581 and 1-.4277) we see that Duke's chance of winning should be around 57.7%

This allows us to solve (or at least partially resolve) Pomeroy's two prediction flaws: lack of accounting for consistency, and lack of accounting for diminishing returns. The first is obvious, the second is because team's expected play versus their actual play should reflect the error in his ratings derivations.

If I had enough time to scour through all the teams' data, I could give an adjusted Standard Deviations (or, 'Consistency') value for teams -- adjusting their consistency based on how consistent or inconsistent their opponents play.

The rules for Step One.

Let's try this jam out on UNC.

The best way to predict a team's four factors in a future game is to create a linear regression involving their four factors, and their opponent's four factors.

Unfortunately, Ken Pomeroy has not yet adjusted the Four Factors for quality of opponent play (and for good reason - it's quite complicated). So we need to estimate how strength of schedule affects actual four factors. Unfortunately, I don't have any good way to run this analysis on every team. The best theory of adjustment would apply to all teams, but since there is a good chance that individual teams affect these numbers differently, it's not entirely bad to only regress on a team-by-team basis.

The next part of this is much harder.
We need to find the standard deviation of actual versus predicted four factors stats in order to run it through a Monte Carlo simulation that takes all likely normally-distributed values for all of the four factors+pace (which is 9 variables), which in turn spits out a point margin (whose values come from the previous post).

I'll be coming up with this system pretty soon, so watch out.



Step Two of the Two-Step Process

The best way to predict point margin is to first predict a team's four factors, then convert the four factors into point margin via linear regression.

The linear regression is the 2nd step, and here are the results (with an R^2 value of about .99)

(Numbers derived from http://kenpom.com)


Step one is a bit harder in some-ways, and should probably be done on a team-by-team basis. We'll cover that soon.

UNC's Injuries

Here's how UNC's injuries have affected their play, in terms of points. The number represents how Carolina does versus their average play.

(numbers based on Actual - Expected Point Margin, taking expected point margin from Kenpom.com)




Davis Zeller Graves Ginyard
IN -0.1 0.8 -0.3 -0.2
OUT -16.0 -10.1 -12.7 -3.9
Difference 15.9 10.9 12.4 3.7


The One-Seeds

I told my friend Stephen that Kentucky will not be a 1-seed come tournament time.

That was a pretty dumb thing to say

I picked the top few teams that I thought might make #1 seeds, and did some analysis from their stats from Kenpom.com.

Anyways, here's my #1 seed bracketology: http://spreadsheets.google.com/pub?key=tdf4HIaf_vWtQhoNCxjHLdQ&single=true&gid=0&output=html

Time Left on Shot Clock

By doing some simple multiplication and division of stats from kenpom.com, we can estimate the mean/median/expected number of seconds left on the shot clock when a team's possession will end.


I expect a high standard deviation of this number for most teams, but it is interesting to look at.

Here's the results (internet explorer might be required, hopefully not)

Adjusted Player Offensive Ratings

I adjusted Ken Pomeroy's 100 most efficient college players (with a minimum of 40% minutes played) for opponents' quality of defense.

The results are here.

(EDIT: the Usage% represents how much a teams' possessions a player ends up 'using' via shots/turnovers/etc. players under 20% are below average in usage. I will soon adjust only those who are above the 20% mark)

Texas v. UNC

Ken Pomeroy's Stats predict Carolina to lose to Texas by 20 points. Here's my basic info you need to know on these 2 teams:

1) Texas' point margin vs. predicted has a standard deviation of about 8.07 points
2) North Carolina's point margin vs. predicted has a standard deviation of about 9.86

this gives us an average of 8.97 for both teams
which means that there is, according to the normal distribution:

-a 68.2% chance that Carolina's final margin is between {-11 and -29}
-a 95.4% chance that Carolina's final margin is between {-2 and -38}


Simply by using standard deviations, Carolina has a 1.29% chance of winning, less than Ken Pomeroy's estimation (using the Log5 method) of 5%

Nathan's Statistical Rankings

Here is a link to my statistical rating of college basketball teams, according to my best possible model given the stats I currently have (which is similar in nature to the LRMC model, and similar in appearance to Sagarin ratings).


http://tinyurl.com/nathansrankings

Hopefully in January I will have a model adjusted including diminishing returns, consistency, and 'game point margin' which accurately reflects the 'real score' of a game, rather than one that was altered in the last 30 seconds to a game-insignificant-degree. (To do this, we will use Bill James' "time statistically over" stat from Statsheet.com).

UNC's terrible 2nd halves

Carolina is beating their opponents by .36 points per possession in the first half of their games.
But in the second half, they average -.04 points per possession.

Not good!

Followers

About Me

I wish my heart were as often large as my hands.