Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Tuesday, July 12, 2011

Employment-Population Ratio - a scarier way to look at unemployment

The Bureau of labor Statistics compiles a stat called the Employment-Population Ratio that is the number of people working of the available workforce, defined as people 16 and over. I am sure some of those folks are not working because they wouldn't otherwise, but in the current climate many have no choice. As has been reported repeatedly, the 9.2% unemployment underestimates the number of people unemployed because after a long enough time people just leave the workforce and are not counted in the number. The chart below shows the current ratio is 58.2%, that's the lowest Employment-Population ratio since 1984.

A longer time series shows the history since 1948.

It's been lower in the past, but that is a generation ago. I personally would like Congress and the President to start working on jobs now.

(Unstable Isotope at Delaware Liberal got me thinking about this)

Thursday, March 24, 2011

March Madness Simulations

My March Madness NCAA basketball playoff simulation this year followed the schema of the Playoff Fantasy Football simulations.

I used the Sagarin ratings to get information on the teams' expected performance.

The difference between the ratings for two teams is the expected spread and I used the normal distribution with a standard deviation of 8.83 points to get the probability of a team beating another team.

I then simulate the whole tournament as many times as I like using a random number generator and the probability generated above to pick the winner of each matchup.

To make my picks I then compared a particular set of picks to the simulations, typically 1000, to determine the points for that set of picks.

Recall that in the bracket pool in which I participate not only are the rounds weighted 1,2,4,8,16,32 for rounds 1 through 6, but you multiply those points by the seed of the winning team. If Utah State makes it to the final four, for that round you get 8 points times its 12 seed or 96 points.

I compared several selections to each other to try to find the one that had the most points in competition with other selections in my simulations.

Top seed advances wins - 6% of simulations
Top Sagarin points advances wins - 23% of simulations
Top Pomeroy rating advances wins - 25% of simulations
Secret winning picks wins - 47% of simulations


Thus I was at least able to find a set of picks that outperformed Sagarin and Pomeroy in my simulations. I suspect it has something to do with the seed rating. The results reflect the expected outcomes before the start of the tournament but after the play-in games.

A chart (click for larger) showing the expected outcome of a given seed helps explain why lower seeds might yield more potential points than higher seeds. using the simulated results, the top plot shows the expected round that a team will advance to. Round 6 is the championship. The middle plot shows the expected points for a given seed assuming a standard point scheme of 1,2,4,8,16, and 32 points for winning a round. The bottom plot is the one of interest here, it includes the seed rating in the points.

Upper plot: First seeds are expected to do well advancing to the elite eight on average. Next seeds two through four generally make it to the sweet sixteen. Seeds from about 5 through twelve generally make it one round, and that uncertainty is where the fun comes in.

Middle plot: The standard point assignments don't change the expected value of a team much.

Bottom plot: Including the seeds in the expected points really shows how three and four seeds can be worth much more than a one seed. Even more interestingly, the correct 10 through 13th seed can be worth more points than a five through 9 seed, because they do about as well in the tournament, but have more points due to the seed multiplier. That is what makes the seed multiplier a fun bracket pool game.

Saturday, March 19, 2011

Hard drive prices from $1million/GB to pennies/GB in 30 years

The ever increasing pace of progress has made once incredibly expensive things available to the masses for pennies. No where has this been more evident than in the computing industry where each year more computing power becomes available for the same price as the year before. Computing equipment is one product in a deflationary spiral, and that's good.

One example of this decrease in price is in the price per GB of hard drive storage. Some have grabbed this data and then just picked some milestones. I think it is better represented in a graph.

Hard drive storage has gone from $1million/GB to pennies/GB in 30 years. That's seven orders of magnitude in 30 years.

I also took the data and put it into a spreadsheet which you can access at Google Docs so that you can make your own charts without copying and pasting from the original data which is not in tabular format.

(via BoingBoing, via isen.blog, via ns1758.ca the data is at this link)

Sunday, January 16, 2011

After the WC and DIV week - Superbowl and Fantasy Football Playoff Roster winner predictions

After today's unexpected loss of NE to the NYJ a lot of clarity was brought to the RKB playoff fantasy football pool predictions. Firstly the model probabilities for the SB matchup and winners.


The most likely outcome and matchup is PIT playing GB in the superbowl and GB winning.

Because many rosters in the playoff pool had NE players a large number of potential winning rosters (including my own) were wiped out and have no chance of winning. Also, the GB and PIT games this weekend generated a lot of points for rosters which had those players. Finally, not very many people had NYJ and CHI players in their rosters. The combination of these factors means that for the purposes of my simulation model there is not difference in the 1st and 2nd place winners among the several superbowl mathcups and outcomes.

The expected winner of the pool is roster 64, Rich V 3, built on a GB/PIT superbowl matchup with GB kicker, QB and DEF, and filled in with PIT players for the extra RB and WR. The second place winner is roster 9, Bruschi Drink 4, which contains all GB players. The roster just placing out of the money in third place is roster 58, RDK 2, my roster built around GB kicker, QB and DEF, with NE and PIT players filling out the RB and WR positions. There is no graph because these are the results in 2000 simulations no matter the team outcomes. I didn't bother with 10,000 simulations because for this model the results look pretty unequivocal.

The results show a failure in the model in that there is no link between team performance and player performance except for the number of games played. Additionally it doesn't take into account good or bad days for players since it only uses the average number of points per game for the year. There is good and bad news in this. A weird combination of player performance could let me squeak through and win the pool, but that same combination could let a 4th or 5th place through to win. The only way to get better information is to get game by game data to try to get the variation in player performance into the simulation.

Wednesday, January 12, 2011

SB winner predictions post wild card week

Before the NFL Wild Card week the Playoff fantasy football simulations showed the likely matchups were either NE or PIT vs. either ATL, CHI or GB with NO and PHI more distant possibilities on the NFC side and BAL, NYJ and IND as even more distant possibilities on the AFC side. A NE win of the superbowl vs. many different opponents figured high in the probabilities. The chart looked like this (click below for larger):



Now that the wild card games have been played and IND, KC, PHI and surprisingly, NO, have been eliminated, plugging those completed games into the simulation and running under those assumptions yields the following chart:


The earlier conclusion that NE or PIT would likely meet ATL, CHI or GB in the superbowl still holds. The NE vs. GB matchup seems to be the primary beneficiary of the NO and PHI losses. While the probability of SEA in the superbowl is now visible on the chart they are still a distant fourth place in probability vs. the other NFC teams. The balance of probabilities on the AFC side hasn't changed much, NE still leads with PIT a close second. The following chart is a Pareto chart of the likelihood of outcomes for matchups and SB winners with the wild card week results included.

It yields a more detailed view of the outcomes. Still, 79% of the time, NE or PIT meets GB, CHI or ATL in the SB.

My playoff rosters with NO and PHI concentration are eliminated now, but the roster with a GB focus expecting a NE, GB matchup is still very much alive and a contender for first place in the RKB fantasy football playoff pool.

Thursday, January 06, 2011

Superbowl winners, Simulations vs the Professionals

I took the odds of a given team winning the superbowl from Yahoo Futures for comparison with my simulation results. My results are in red below.



I match pretty well with 5dimes.com and SBGGlobal, but bodog looks too flat and I don't know what Sportsbook was thinking with the high probability for the New York Jets.

The agreement with some of the professionals lends some credibility to the results of my simulations and the roster decisions I will make based on them. I always say I don't know anything about football so I base my work on those who do and the people who set the odds need to know because they are trying to make money doing this.

Wednesday, January 05, 2011

Superbowl simulations - matchups and winners

I am struggling to find the best way to present the data from my simulations of the football playoffs. The key to picking a good roster is to figuring out which teams play multiple games in the playoffs, essentially the ones that make it to the Superbowl. Thus I compiled 10,000 simulations of the playoffs and then determined who the AFC and NFC champions would be that would meet in the Superbowl and who the winner of that game would be. The pie charts below try to capture all of the outcomes from 12.5% chance that NE will beat ATL in the Superbowl to the less than 1 in 10,000 chance that SEA and KC would meet in the Superbowl.


The chart below is a filter of the above data with only NE or PIT as AFC champion, and ATL, CHI or GB as the NFC champion.


The Pareto below shows the top 15 outcomes for matchups in the Superbowl. They represent 78% of the outcomes of the simulations.


The matchups of either NE or PIT vs. either ATL, CHI or GB represent 68% of the outcomes of the simulations. The only wrinkle left there is which team will dominate the game an and so the makeup of the rosters to cover those possibilities. The results are heavily waited to not only a NE appearance in the Superbowl but also to a NE win.

Monday, January 03, 2011

Who will win the Superbowl this year? Simulations suggest...

I am currently crunching the numbers for this year's Playoff Fantasy Football Pool. Using the Sagarin ratings and the formula discussed earlier, I have randomly simulated this year's playoffs many times to determine who will win the Superbowl. This year New England seems to be the favorite to win if we just assume wins based on the ratings. Even if we use 13.92 as the standard deviation and the spread and a normal distribution to calculated the probability of a team winning, New England wins about 33% of the time with no home advantage and 39% of the time with a home advantage of 2.11 points as given by the Sagarin ratings. The plot below shows how the winner varies with the standard deviation used in the model for either 2.11 points home advantage on the left or 0 points on the right.


Click on the chart for larger. The X-axis varies from a standard deviation from 0.001 which essentially assumes that the team with the higher spread wins through more reasonable scenarios with higher standard deviations which are more like what has occurred in previous NFL seasons.

Using the home advantage of 2.11 and a standard deviation of 13.92 simulations also show the likely teams to make it to the AFC and NFC championships. The plot below shows those results for 1000 simulations.

The likely matchup is either New England or Pittsburgh as AFC champion vs. either Chicago, Atlanta, or Green Bay as the NFC champion. Other teams have slim chances of appearing as seen down near the x-axis. Notably, New Orleans or Philadelphia have slim but noticeable chances to be the NFC champion, as well as Baltimore, New York Jets or Indianapolis having slim chances to appear as AFC champions.

Next step, collecting player data and simulations to make the perfect roster picks.

Sunday, January 02, 2011

Probability of winning an NFL game - recalculated after thirty years

It's NFL playoff time again. I am in the process of redoing my playoff football model to more accurately reflect the probability of a given team to win a game based on the spread or the Sagarin rating difference.

Stern wrote a paper called "On the Probability of Winning a Football Game" (1991) in which he collected the final scores and the spreads from 1981, 1932, 1984 to determine the relationship between the two. He found that the final score difference between the favorite and the underdog, subtracting the spread could be modeled with a normal distribution with standard deviation of 13.89. The average was 0.07 which is effectively zero for the purposes of the analysis. The probability that a team will win a given game is then the cumulative normal distribution around the spread with a standard deviation of 13.86, or normsdist(spread/13.86) using Excel functions.

I wondered if the analysis had changed in 30 years so I pulled the data for this year through week 16. The plot is below:

The standard deviation is 13.92 with an average -0.17. Hardly any difference found from the earlier analysis for a lot of work to extract the data and get it into a format for the analysis, but at least we now know it hasn't changed. The normsdist function with the spread replaced with the Sagarin difference (home+home advantage-away) divided by 13.92 is what will be used in the game simulation for the playoff fantasy football.

Saturday, July 17, 2010

Number unemployed vs. unemployment rate

While collecting the unemployment data I also ran some reports on the time that people are unemployed. The data is available showing up to 5 weeks, 5 to 14 weeks, 15 to 26 weeks and longer than 26 weeks. One of the signs of the depth of this recession is not just the number of people unemployed but also the length of time they have been unemployed. The charts below show the number of people unemployed for a given number of weeks in a bar chart which builds the categories on top of each other. I have overlaid the rate of unemployment which corresponds to the right axis. (click for larger)

The chart below just shows the faction in the various unemployment categories adding up to 100%, instead of the actual number. (click for larger)

If I have time, I want to overlay dates of unemployment extensions and see if there is any correlation. I still don't think that would be completely relevant because it is obvious that if an extension was passed then the time people receive unemployment insurance will be longer. Will I be able to see whether this in fact means they stayed unemployed longer or do extensions occur when Congress sees that people are already unemployed longer and the extension is to address that issue?

(Once again the data is available in this spreadsheet on Google Docs. - How Long unemployed vs rate unemployed )

Monday, July 05, 2010

Modeling soccer (and perhaps your work team) as a social network

Frequent readers will be aware that an occasional hobby of mine is to numerically model sports and other interesting activities yet I lack much of the time or tools to do so. I have taken on NFL playoff football, both the entire playoffs and detailed modeling of the Superbowl down to the player level. I have also attempted, unsuccessfully, to model March Madness, the NCAA basketball playoffs, performing mostly analysis as opposed building a model that helps me win a March madness pool. Thus, I love to read interesting modeling papers, especially those which model sports or games as models for other real world activities.

The contributions of individual players in sports like football, baseball and basketball are helped by the large amount to statistics collected and available for these sports. Thus the contribution for individual team members to the team success is easier to model. Science online has a report of some work done by Jordi Duch and other researchers, at Northwestern and in Spain, that attempts to model the contributions of soccer players to the success of their team.

They point out that soccer is a very fluid game compared to baseball, or football and that combined with the very low scores makes statistics like goals and assists insufficient to model the contribution of players to the performance of the team. They hypothesize that the passes and flow of the game leading up to the rare goals are important for determining the outcome of the game and they use networks to model this flow. Players are nodes in the network and the lines between the nodes, called arcs, represent passes. Much as a Facebook or Twitter can be modeled as a network with friendship and interactions or follower/following being the connections, soccer is a "social" sport.

They also include nodes for the goal and for shots wide of the goal. To each of these arcs the attach statistics and probabilities from the 2008 European football championship on play pass accuracy, and goal accuracy to the arcs. One could them follow the ball through this "ball flow" network to a goal, a miss or to the other team. Combined with more calculations the group attempts to predict the outcome of soccer games.

Even more interestingly, the authors apply this concept to a work team that is writing a paper with several co-authors. Instead of the nodes being soccer players in paper network, a node represents a co-author in the manuscript, and the lines between the nodes represent communications directed from one co-author to the others. The e-mails represent communications between coauthors and the effectiveness of the authors is measured by completion of tasks like performing a calculation or scheduling a meeting. In the diagram below, author A2 (I think A3 in the second chart is a typos) seems to be an important and strong contributor.

One of the authors, comments on how the scheme can be used to assess the contribution of individual team members.
"One of the issues with any kind of teamwork is assigning the right credit," says Amaral. "The wild, loud people get more credit, but with this analysis you can get a picture of how much an individual really contributes to an outcome."
As work continues to evolve to be more team driven and highly networked, perhaps a scheme like this can not only point out strong contributors to a team, but also help an entire team work at a higher level. Imagine it applied to the work of developing open source software or Wikipedia articles.

(via Science online, the paper at Public Library of Science, PLoS, figures above are from the paper can be found at this .pdf link)

Tuesday, June 15, 2010

Where are people moving to?


Red lines represent moves out of the county and black lines moves into the county. The connections are interesting. There is a lot of red to very populous cities, folks are leaving New Castle County to go to New York, Seattle, San Francisco, LA, Dallas, and Houston. Presumably some are retiring to the Florida counties and to Las Vegas and Arizona. There also seem to be some moves, possibly due to the chemical industry to the Texas coast, New Orleans, Buffalo and Rochester.

There is expected inflow from obvious places like New York, and Philadelphia. I also see what I think is college inflow from Penn State area, Univ of Wisconsin, Univ of Minnesota, Univ of Michigan. All of these are my suppositions. The full data is available at the links below.

How much of a coincidence is it? The Walt Disney World couple

Boing Boing points the coincidence of a woman spotting her husband's father and her husband as a toddler in her old family pictures that records the fact that they were at Walt Disney World at the same time as children long before they ever met and then married. The future wife is with the Mr. Smee character in the foreground and the future husband's dad is clearly visible pushing a stroller with the future husband in it in the center in the background.

It's a wonderful synchronicity that foreshadows that they were meant to be from long before they ever met, or is it? What is the probability of such an occurrence?

I think it breaks into two problems:

Problem one: Given that you were at Disney World in your youth on a particular day, and you took a picture, what is the probability that someone you know was there on the same day and in the picture.

Problem two:On the other hand, we heard about this on the news, so we could state it another way and ask, how often would we expect to hear about someone having a picture from there youth at Walt Disney World, that captured someone they know now but didn't know then in it.

Here are two comments on BoingBoing trying to figure out this probability (one, two). I think the two possibilities above are more common than you think and that the second one is very likely, I just have given up trying to formulate the problem in a clear manner since I can't quite wrap my brain around the probabilities because there are some dependent ones in here. What do you think? I have provided some data below to help:


Some data that might be helpful:

According to Park World the attendance at the Magic Kingdom was 17 million in 2007. Other estimates of milestones reached by Walt Disney World are at Disney By The Numbers. For instance the 600 millionth visitor entered on June 24, 1998. The Orlando convention and visitors bureau estimated about 3 million international visitors in 2009 (from here). Assume they all go to Disney World and are part of the 17 million total mentioned.


Kodak estimates that approximately 4 percent of all the amateur photographs taken in the United States are snapped at Walt Disney World Resort or Disneyland.

Looking at the picture itself shows about 15 people in it.

Photos printed per year (from here).

Rolls developed per year. Within the amateur market, 710 million rolls of film were developed in 1995. Total rolls were down slightly from the 1994 figure of 716 million, but higher than the 1993 total of 694. (from here). Assuming 24 photos/roll gets to 17 billion pictures, and that these are US figures, which is implied in the report.

"In fact, in 1960 newborn babies and young children were the object of 55 percent of the 2.2 billion photos taken that year." (from here)

There are over 2700 photographs taken every second around the world, adding up to well over 80 billion new images a year taken on over 3 billion rolls of film, according to estimates published by the United States Department of Commerce. (from here via there) See also this chart of the number of photos taken/year.

(via BoingBoing, via The Disney Blog, via WXII TV, photo from WXII TV)

Thursday, April 01, 2010

Beautiful Bracketology

Leonardo Aranda's Bracketology - NCAA 1985 - 2009 is a beautiful rendition of the results of the NCAA March Madness men's basketball tournament for the 25+ years of the 64 team format.



I myself have compiled compiled these statistics over the years (2009, 2009, 2008, 2007, 2007, 2006, 2006) in an attempt to win the NCAA March madness bracket pools in which I have participated and just for the fun of studying the statistics. But I am green with envy as well as another appropriate color with admiration when someone takes data that I have kicked around for years and makes a striking visualization from it.

Again I am Salieri* to some Mozart. I recognize genius and beauty, but I can only produce mediocrity. (*the Amadeus movie version of this story, not the real one)

(via Castro's Favorite Color)

Tuesday, March 16, 2010

March Madness Simulations - Game winning probability schemes

After my recent success with simulating Playoff Fantasy Football, I wanted to apply that success to a simulation of the NCAA Basketball playoffs known as March Madness. Given the amount of data analysis that I have done over the years (2009, 2009, 2008, 2007, 2007, 2006, 2006) that even enabled me to win one year, I figured that a simulation might help.

My simulation matches up the teams that play in the NCAA bracket and uses one of the schmes below to generate a probability for a Monte Carlo simulation of games between the teams.

Probability scheme 1: Sagarin ratings only

The first simply uses the Sagarin ratings to create a probability of the team 1 winning. Probability = team 1 Sagarin /( Team 1 Sagarin + Team 2 Sagarin). I use the Predictor Sagarin Rating because that is what he suggests for predicting the score and outcome of a game. A random number from 0 to 1 which is less than the probability above means that team 1 wins, otherwise its team 2.



I calculated every team's probability of winning vs every other team and then plotted this vs the difference in seeds. A -15 means a 1 seed played a 16 seed. This scheme results in probabilities that only vary from 58% to about 50% for matchups between seeds with up to 15 difference to even. Unfortunately no 16 seed team has even beaten a number 1 seed so this scheme leave the games too evenly matched and does not reflect the history of outcomes in the tournament.

Simulation results with this scheme show the number of simulations out of 1000 that a given seed was the champion. The actual history is here. The results in the chart show far too high a probability that low seeds are the champion in the tournament in these simulations.

A histogram of the teams with seeds and the number of times they are champions in 10,000 simulations, shows that Kansas is the most likely winner, but the spread of the data even includes the unlikely play in winner at 16 seed as a champion. This simulation is unrealistic.

Probability scheme 2: Seed difference and tournament history only

Another approach is to use the seeds of the team in the tournament. With 25 years or so of data I captured the number of times a favorite beat an underdog based on the seed difference. For instance, never has a 16 seed beaten a 1 seed, while 8 vs. 9 seeds are almost 50/50. I use the data from 25 years of round of 64, round of 32 and round of 16 and then fit a line assuming that even seeds are 50/50 and that a seed difference of 15 (1 vs. 16) will result in a favorite win 99.07% of the time. That represents 1 in 108, though this upset has never occurred in 26 years of data, it will happen someday, and that could be as soon as 1 this year. Thus (26*4+3) wins/(27*4) attempts is 99.07%.

I did not use the fitted line in the curve above because of its unrealistic probabilities at high seed difference. While this approach captures the history, I feel this approach neglects the variation between similar seeded teams as reflected in the Sagarin ratings. Additionally the history shows pretty wide variations in outcome.

Simulation results with this scheme show the number of simulations out of 1000 that a given seed was the champion. These results are more similar to the historical outcomes, but the matchups between evenly seeded teams will be tossups that ignore the differences as determined by the Sagarin ratings.

A histogram of the teams with seeds and the number of times they are champions in 10,000 simulations, shows that Kentucky is the most likely winner, with low seeds favored to be champions, but I fear that it neglects the difference in teams as represented by the Sagarin ratings. This simulation is unrealistic.


Probability scheme 3: Sagarin ratings scaled by seed difference and tournament history

The final approach combines the two by scaling the average of the Sagarin ratings probability by the expected probability due to seeds as predicted by historical performance. Thus we make sure the average for teams. In practice I add the residuals of the line fitted through the Sagarin rating probabilities to the line fitted by setting the 15 difference probability to 99.07% and the even difference to 50%.

Thus the probabilities reflect the historical data with a more realistic and very rare chance of 16 seeds beating 1 seeds but with the Sagarin ratings to sort between evenly matched teams.

Simulation results with this scheme show the number of simulations out of 1000 that a given seed was the champion. The results is similar to the seed difference with history scheme above, but now the Sagarin ratings are included.

A histogram of the teams with seeds and the number of times they are champions in 10,000 simulations, shows that Duke is the most likely winner, and low seeds are still favored as is true historically. This is the simulation scheme we will proceed with.

Tuesday, February 09, 2010

I'm in the money - Second Place in the Playoff Fantasy Football Pool

My roster (RDK 1) came in second in the the RKB Playoff Fantasy Football pool!

I predicted a significant chance (18%) that roster RDK 1 would be in the money! And it happened. I just want to take some time to gloat. My acceptance speech:
"I want to thank the Drew and the Saints for winning the Superbowl, especially their defense for that critical touchdown and Garrett Hartley - kick away Garrett. I also want to thank Joseph Addai for getting that touchdown that helped put me over the top, even though his team lost. And Adrian Peterson, you didn't even make it to the big game, but getting those touchdowns with no credit for Brett Favre really helped. Thanks to Yahoo for your player stats, and Sagarin for your ratings. And finally, I couldn't have done it without math and statistics, you guys rock!"


Here are the final results with all of the roster's points separated by position. It pays to have a good QB on the roster, but WR, RB and K's also contribute almost the same amount of points for the roster which are towards the top. Remember that there are 3 RW's and 2 RB's so the K has more point generating power as a single player. Even the defense can be significant. Probably the TE is the least useful point generating player on a roster.

The final results separated according to the game in which the points were generated reveals a truism that has been a guiding principle all along. Rosters with players that play more games generate more points. The light blue "dusting" of Superbowl points is what determined the winner this year.

A chart with the order of the roster based on the points before the Superbowl shows a little more clearly that the Superbowl points are what changed the order around. The top contenders had many or all NO and IND players left on their sheets, especially the big point positions like QB and K.

The rosters are shown above for the top twenty finishers, with just the players in the Superbowl on them. Realize that in the above some roster (like mine, RDK1) had players that did not play in the Superbowl and so are not listed above, however the correct total points are in the grand total at bottom.

The final contenders strategies were the three fold obvious ones, all NO, all IND or a mix. Give the way the game went it didn't pay to be all IND. I was able to thread my way to second place because I was a mostly NO roster, K, QB, DEF, but with enough IND to differentiate myself from others. Those that split the K and QB between IND and NO ended up not faring so well.

Finally, I simulated this outcome. Bruschi Drink 3 in first place and RDK1 in second, was the second most likely outcome in my simulations at 10% after the one with Tim G 5 in second.
The simulations above are from the prediction before the Superbowl. What happened to Tim G 5? That roster started 2 points behind RDK1 before the Superbowl. It had IND K instead of NO K for who were 5 to 11 in the Superbowl for 6 more points of deficit. RDK 1 beats Tim G 5 entirely due to the choice of kickers. Even if Matt Stover (IND) had made the field goal he missed that would only have added 3.

The simulations also picked out particular aspects of the game. About 40% of the time when New Orleans defense forces a turnover they get a touchdown. I included that in my model and lo and behold it happened during the game. Having Joseph Addai finally get a touchdown this playoff season pushed me over some of the NO rosters, but having NO do so well pushed me over the IND rosters. It also helped when Jeremy Shockey got a touchdown because no one of the top contenders had him for points. Sometimes it is just as good when no one gets the points as when your roster gets the points.

Next up, March Madness simulations. I have to go get started.

Sunday, February 07, 2010

Snow pictures #3 with more measurements 20 to 24 inches!

The measurements in the pictures below might not reflect the "official" amount of snow that fell yesterday, but they certainly represent the "official" amount of snow I had to shovel. For the purposes of my back they are the measurements that count.



20.5 inches at the top of the driveway.


...down to 18.5 in the middle...

...and back up to 24 inches at the end of the driveway.

Thursday, January 28, 2010

18% Simulated Chance of winning money in the Playoff Fantasy Football pool

This year I have a roster (RDK1) that is currently in fifth place in the RKB Playoff fantasy football and in striking distance of first or second place and winning money in the pool. The goal is to determine the chances of that happening. The focus of this simulation is to answer the question "With only the Superbowl to go, will the RDK1 roster be in the money at the end of the playoffs?" and "What combination of player results does RDK1 need to be in the money and what is the chance that such an outcome will occur?"

Developing the simulation required the following steps and assumptions.
1.) Collect the data for each player or teams games for this season. Turnovers and touchdowns for DEF (and special teams); field goals and extra points for the kickers; passing and rushing touchdowns for the quarterback, wide receivers, tight ends; rushing touchdowns for the running backs. I will randomly select from this history to generate simulations of the Superbowl.
2.) Assume a player's performance in the Superbowl will be identical to their performance in one of the games they played in this season. If a player didn't play they get a zero for that game, except for the kickers for which I only have partial season data. This may decrease the points slightly, and is potentially a bad assumption.
3.) Kickers get their own field goals, but only get extra points equal to the touchdowns their team scores (actually all the rushing and defense touchdowns, but only the QB passing TD's to avoid double counting). Typically a game with field goals has less touchdowns, so decoupling the game history so that a game with a lot of touchdowns for the QB could be paired with a game with a lot of field goals for the kicker could result in a higher than expected points. Possible another poor assumption.
4.) Everybody (RB and QB) gets their rushing touchdowns, but passing touchdowns are awarded only if the quarterback throws at least one. There are instances in the game history of the QB's not throwing any. I really should assign each passing touchdown to a WR, TE, or RB or player not on the list but that is to complicated to program in excel. This may result in excess points, and is an expedient assumption.
5.) The score of the game is the field goals, rushing RD's, and only the quarterback's passing TD's to avoid double counting, and the defense/special teams touchdowns. This is slightly inaccurate since the passing touchdowns for the receivers are not all counted or double counted. The simulation still generates widely varying scores.
6.) Simulate many games by bootstrapping (selecting TD's or outcomes from each particular player's history this season. Add the points for each player to the rosters that have the players on them.
7.) Used the RANK() function to determine the places. Ties get the same rank using this function and the next ranks down are eliminated. For instance 3 first places get rank 1 and the next rank is 4. Rank is important to determine who is "in the money".
8.) As to the money, it is a fraction of the total collected from all of the rosters: 70% for first place, and 30% for second place. However, a tie for first divides the total money (100%) and there is no second, a tie for second with only one first divides the second place money, 30%, among the second place tied rosters. To be in the money RDK1 needs to be alone in first, tie first, or be alone or tied for second with only one first place roster ahead.



The rosters above show RDK1 roster in fifth place, but with enough similarities to other rosters both ahead and behind it that winning money in the pool will require some fine threading of the outcomes.

Remember that the focus of this simulation is to answer the question "With only the Superbowl to go, will the RDK1 roster be in the money at the end of the playoffs?" and "What combination of player results does RDK1 need to be in the money and what is the chance that such an outcome will occur?"

The histogram above (click for larger) shows the rank of the two top RDK rosters, RDK1 and RDK6 after the outcome of 20,000 simulations. The first red bar highlights the fraction of simulations with RDK1 roster in first place and in the money (alone or tied) at 1.4% of 5000 simulations. The green bar highlights the fraction of simulations with RDK1 roster in second place (alone or tied, with no first place tie) and in the money at 16.3% of 5000 simulations. RDK1 is in the money in about 18% of the 5000 simulations. RDK1 starts in fifth place before the Superbowl and can climb to first or slip to 13th place according to the simulations. There was some hope that RDK6 might have the potential to be in the money but from its starting point at 13th place, it never rises above 3rd place and can slip to 30th in the simulations.

Another way to look at this data is go ahead and calculate the winnings for each outcome.



This chart shows that the most likely outcome, 80%, is that RDK1 has no winnings, but the rest of the bars which add up to about 20% are various outcomes with winnings for the RDK1 roster.



This chart expands the Y axis to zoom in on the lower probability outcomes. There is a 10% chance of being alone in second place, a 4% chance of tieing second. There is even a less than 0.2% chance of being alone in first place. The less likely outcomes include situations in which I am tied with several others, up to 5 others, for first or, up to 7 others, for second. I need about 10% of the total collected to break even for the six rosters I entered.

Of course simulation generates outcomes for all of the rosters, otherwise I couldn't perform the comparisons needed to determine what place I am in or whether RDK1 roster will earn money. A less self-centered data reporting approach yields information about all of the outcomes.

The chart above (definitely click for larger) shows the histogram of the frequencies of the final rank after the Superbowl (simulated) of the top twenty rosters as they stand now(actual) before the Superbowl. The top twenty was chosen as a cutoff because it contains the lowest ranked roster that could win money in the simulations. The legend has the roster in their current ranking order (Cara H in 1st through RDK3 in 20th place). Bruschi Drink 3 ends most of the 1000 simulations in first with Tim G5 ending most of the 1000 simulations in second. there is a small but significant fraction of RDK 1 results in second place as we showed earlier. The chart will reward closer examination for the interested.

The information above can be used to determine the fraction of simulations (in this case, 5000) in which any given roster will be "in the money". The chart above shows that Bruschi Drink 3 is more than 80% likely to win some money followed by Tim G 5a at 42%. Almost a third of the time, Cara H in first place is likely to end up with some money. More annoying is that Bruschi Drink 4, a roster currently tied for 20th place, has a small but finite chance of being in the money. The results above do not total to 100% because more than one roster can be in the money (not just 1 and 2 but multiple rosters tieing for first, or one first place with multiple 2nds).

A compilation of the actual outcomes of each of 20,000 simulations can show the most likely particular outcome instead of the probabilistic compilations further above. The outcomes above compile the rosters in first or second place. Recall that in the case of a first place tie there is no second place.

As suggested by the charts further above, but shown directly in this one, the most likely first and second outcome at 22% is that Bruschi Drink 3 will be first with Tim G 5 second. The next most likely is heartening because it has Bruschi Drink 3 in first with RDK1 in second. Even so, these top twenty outcomes represent only 82% of the outcomes generated in 20,000 simulations. There are highly unlikely but predicted outcomes of all sorts, including some interesting ones with 6 tied in first place, or a first place with 8 tied for second, both only 1 time out of 20,000.

The RDK1 roster appears in these outcomes usually as a second place winner in the 2nd, 14th, 15th ,and 18th most likely outcome. You need to go down to the 17th most likely outcome to see RDK1 in first place, though it is tied with the ever successful Bruschi Drink 3.

Thus my final prediction is that Brschi Drink 3 will be in first place with Tim G 5 in second, though I am hoping for the 18% chance of RDK1, my own roster, being "in the money".

Wednesday, January 27, 2010

Who will win Playoff Fantasy Football Pool?

My clever analysis and modeling of this years football playoffs has yielded a roster (RDK1) that is in fifth place in the RKB Playoff fantasy football results as of the NFC and AFC Championship games. With only the Superbowl to go, the question is, "Will the RDK1 roster be in the money at the end of the playoffs?"

The chart above (click the chart for larger) shows the standings as they are right now, after the conference championship games. The y- axis is total points while the x axis is the name of each of the rosters. The colors represent contributions from each week of games, Red for the wild card week, blue for the divisional week and green for the conference championship week.

Disregarding the two lowest results, the Wild card and Divisional weeks yield anywhere from 65 to 30 points in a roster. A roster can also have a great wild card week and still lose, because your players have to generate points each week and that only happens if their team progresses. Which of these rosters will win, will it be Cara H. in the lead with 119 points?

The above plot is the same data and roster, this time sorted first by the number of players a roster has out, and then by the total points. This chart is very telling because the rosters to the left with no players out or few players out have much more points potential than roster to the left with more players out. The last grouping with all nine players out on their rosters is the most pathological example; they have all the points they are going to get. Tim G 2 with a respectable 101 points is still not in the running. By the way, the RDK1 roster only has 3 players out, and since I still have my quarterback, kicker, and defense.

Taking the starting chart and plotting the contributions from each player position to the total shows the importance of the the positions to a successful roster. The colors in the chart above represent points from a particular position, green for quarterback (QB), yellow for the wide receivers (WR), orange for running backs (RB), red for kicker (K), purple for defense (DEF), and blue for tight end (TE). The greater contribution positions are at the bottom and build up to the total number of points.

QB is the most important, and while RB and WR also contribute as much, realize that there are three WR's and two RB's per roster so the contribution above should be halved for RB or divided by three for the WR's. As an individual player the kicker contributes a fair amount of points, almost always one for each touchdown, and then field goals as well. In this league the DEF gets the special teams points if kickoffs or punts are returned for touchdowns, as well as a point for each turnover after there are three. Finally tight ends rarely receive passes in comparison to wide receivers and their contributions are the smallest.

Above is the leader grid (click for larger) with the top twenty team rosters and with only the players that are left to play in the Superbowl. A grayed out square indicates that that roster doesn't have the player, numbers are the accumulated points for a given player in that roster. The grand total is the total for each roster, and the rank is as of now. The red highlight is first, and green is second, yellow are the rest of the top ten. I included the top twenty because I have evidence that one of them can come in first, though it would be very unlikely (less than one in a thousand)

What combination of player results does RDK1 need to be in the money (70% first or 30% second place, a tie for first divides the money and there is no second, a tie for second with one first divides the second place money), and what is the chance that such an outcome will occur? That is the topic for the next analysis.