How to Track and Compile State Running Back Histories for Draft Analysis

I've spent years building out state-by-state running back databases for collegiate and professional scouting purposes. The concept sounds straightforward but the execution has a lot of gotchas that most people new to this overlook entirely. Let me walk you through the practical approach I use, the tools involved, and where it tends to break down. When scouts and analysts talk about state running back history, they aren't just listing players who were born in each state. The meaningful data involves tracking three things: the total number of high-major or NFL-caliber running backs produced by each state over a rolling decade, the positional trends within those states, and the production metrics that separate actual talent from statistically inflated college careers. A state like Texas produces more running backs simply because it plays more football at higher levels. But Kansas consistently overproduces relative to its population. Georgia breaks the curve entirely. You need context data just to read the raw numbers.

The Core Data Sources

You have to pull from at least four independent sources to build something reliable. Pro Football Reference for NFL history. Sports Reference college sections for the collegiate pipeline. High school recruiting sites like 247 Sports or ESPN for the early identification layer. And then the NFL Combine and school-specificCombine results for the measurable data that actually separates prospects. I learned this the hard way. Early on I relied primarily on Pro Football Reference and my Texas-based QB running back dataset showed Oklahoma producing zero NFL running backs between 2010 and 2015. That was obviously wrong. A few later names had slipped through because their NFL careers were short or injury-cut. Cross-referencing with thecombine.com archives and school alumni pages caught the gap. The fix was adding a supplemental sheet for players with fewer than 50 career rushes who still deserve visibility.

Building the Tracking Spreadsheet

Start with a master list. Column headers should include player name, birth state, high school state, college attended, draft year, NFL games played, career rushing yards, yards per carry, and current status. Keep it in Google Sheets because the filter functions are faster than Excel for large datasets and multiple people can access it without licensing headaches. I use conditional formatting to flag states that have produced fewer than three qualifying running backs in a given five-year window. This helps you spot under-the-radar talent pipelines before other analysts catch on. Alabama might not look impressive on total volume but the per-capita talent rate is absurd. Same deal with Mississippi and Alabama actually producing proportionally more NFL running backs than states with larger populations like Illinois or Ohio.

Get the Full Details

Top-12 Ohio State running backs in program history
Top-12 Ohio State running backs in program history

Common Mistakes That Ruin These Projects

The biggest one is counting birth state instead of development state. A lot of running backs are born in one state and grow up playing in another. Deuce Vaughn was born in California but developed his game in Kansas. If you only count birth state you completely miss that pattern. I switched to tracking high school state as the primary category and keep birth state as secondary data. Another issue is timeline creep. NFL career lengths vary enormously. A running back from 1995 has had thirty more years to accumulate stats than one from 2020. Always use rolling windows. I run mine on five-year and ten-year cycles. Any snapshot that tries to compare 2005 to 2025 directly is garbage data.

Practical Workflow for Weekly Updates

Keep it simple. Every Sunday night during the season I pull the week one NFL game logs and highlight any running back who appears for the first time. I check their school, state, and previous production history. The off-season is when the real work happens. I spend August through February building out the high school and early college recruitment layers for incoming classes. I also maintain a separate tracking sheet for injured reserve and practice squad players. These are the guys who get dropped from public datasets but are still part of the state production pipeline. A running back placed on IR after two solid seasons still counts toward your state totals. He just needs a different classification column.

Tools and Automation Shortcuts

There is no good paid API for this specific data. The closest thing is the ESPN Stats API but it requires a partnership and costs thousands. I use a combination of Python scripts running against the sports reference web pages and manual entry for edge cases. The Python portion handles the bulk pulling and the manual work handles players who transferred multiple times or have incomplete records. If you do not know Python, you can still manage this with Google Sheets alone. The filter and pivot table functions are sufficient for a smaller dataset. But once you push past about five hundred entries the performance degrades noticeably and you will wish you had automation in place.

The Top 5 Running Backs in Florida State Seminoles History
The Top 5 Running Backs in Florida State Seminoles History

What This Method Cannot Do Well

State running back history is not predictive. Just because a state produced five good running backs in the last three years does not mean the next one from that state will succeed at the next level. Recruitment patterns, coaching changes, and offensive scheme trends matter far more than geographic clustering. The data describes trends, it does not forecast outcomes. Also, certain states have systemic reporting gaps. Smaller states like Wyoming and Delaware have very sparse records because fewer recruits get tracked publicly. If you are doing deep analysis on these states you need to dig into local high school newspaper archives and state football association records. It is tedious work and the results are often incomplete. The data stays useful as long as you keep the definitions consistent and update it regularly. Once you let a dataset sit for six months or longer without review it accumulates errors that are nearly impossible to fully correct later. Build the habit of weekly maintenance and the whole thing stays clean.