The Numbers Game Nobody Asked For
Football Math is the practice of applying quantitative methods to evaluate player performance, team strategy, and game outcomes. It goes beyond counting yards and touchdowns. The real work lives in probability models, efficiency metrics, and situational decision frameworks. Most people who hear about Football Math think it is either magic or nonsense. It is neither. It is just math applied to a chaotic sport, which means it produces useful patterns with predictable blind spots. At its core, Football Math tracks expected outcomes rather than observed outcomes. The simplest example is Expected Points (EP). Every field position and game situation in the NFL has a historical average point value attached to it. Fourth-and-1 from the 40 with ten minutes left is worth somewhere around 1.2 expected points. A punt from that spot is worth about 0.3. That difference tells you whether going for it makes sense before you even watch the play unfold. Expected Points Added, or EPA per play, measures how much each individual play shifts your point expectation. A running back who gains three yards on third-and-one but fumbles on the next carry has a mixed EPA profile. The third-down conversion added roughly 1.5 expected points. The fumble subtracted maybe 2.8. The raw stat sheet would say he gained nine yards and had one big run. Football Math says he cost his team points. That is the gap between the two approaches.
Win Probability Added works similarly but tracks the likelihood of winning instead of points. After a successful onside kick recovery in the fourth quarter, win probability might jump from 12 percent to 34 percent in a single play. That 22-point swing is the WPA. Aggregated over a season, WPA identifies clutch performers and punts narratives that box score stats completely miss.
Building Your Own Football Math Model
You do not need a PhD to do this. You need a dataset, a spreadsheet, and some willingness to debug your own assumptions. Here is the practical workflow I use. First, get clean play-by-play data. The NFL public API at https://github.com/rbdulduld/nflfastR-data is the standard reference point, though there are commercial sources too. Download a full season and verify the rows match what you expect. I once imported a dataset where overtime plays were labeled with home team advantages that did not exist because the data source folded the extra period into the fourth quarter. That error inflated win probability models for teams that won in OT by roughly 8 percent. I caught it by comparing drive sequences against official box scores. Second, decide what you are predicting. EPA is the most common target. You can model it as a regression problem where features include down, distance, field position, score differential, time remaining, and player-level variables like quarterback dropback time or receiver separation. Keep the initial model simple. Overfitting is the easiest way to produce impressive-looking but useless numbers.
Get the Full Details

Third, calculate baseline probabilities from historical data. Use at least three seasons of play to establish reference points. A two-minute drill from your own 20 with a one-score deficit has a very different success rate than early first down in the same situation. Context matters more than raw ability in isolation. Fourth, validate against out-of-sample data. Train on 2022 and 2023, test on 2024. If your EPA predictions are no better than guessing the league average every time, go back to the feature selection step. My models usually land somewhere between 62 and 68 percent accuracy on directional EPA predictions, which sounds low but is actually decent. Predicting exact point values is nearly impossible. Predicting whether a play will be above or below average is the real goal. If you want ready-made tools, nflscrapR and the fastR packages in R handle most of the heavy lifting for EPA and WPA calculations. The underlying logic is open source and you can inspect exactly how each number is derived. That transparency matters because Football Math becomes unreliable when the calculation method is opaque.
The Counter-Intuitive Stuff People Miss
Here is something most beginners get wrong: high EPA does not always mean a player is good. Sometimes it means the player got targeted in high-leverage situations where big plays are expected. A wide receiver who catches eight passes for 120 yards and one touchdown in the red zone might have an incredible EPA but a mediocre yards-after-catch number. He was in prime situations. Evaluate the process, not just the result. Another counter-intuitive finding: punting on fourth down is correct more often than coaches admit. The conventional wisdom among casual fans is that teams go for it too much now. The data actually shows the opposite in many cases. The break-even point for fourth-and-short is lower than most people think because the alternative, a punt, rarely flips field position enough to justify the risk. I ran a personal analysis on fourth-and-2 situations between our own 30 and the 40 during the 2023 season. Going for it had positive expected value in roughly 74 percent of those cases. Coaches who punted were, on average, making a suboptimal choice.
Where Football Math Breaks Down
It breaks down when you treat it as prediction instead of evaluation. Expected Goals models in soccer work well for season-long player assessment. In football, the sample size per game is small and variance is enormous. A single coin flip on a 50-50 ball can swing EPA by three points. That makes game-to-game Football Math noisy and often misleading if used for forecasting. It also fails when you ignore situational context. A quarterback throwing behind a collapsing pocket has a lower completion percentage and a lower EPA per attempt than his true ability suggests. The context is not noise. It is the signal. Most automated Football Math tools strip that nuance away and you end up ranking players by stats that punish them for playing hard down the middle of the field against blitzes. The biggest limitation is that Football Math cannot measure effort, leadership, or fundamentals that do not show up in play-by-play data. A safety who takes poor pursuit angles might not commit a mistake visible in the box score. His EPA contribution might look fine. His actual impact on the defense is negative. You catch that by watching film, not by looking at spreadsheets.

Another blunt truth: Football Math struggles with small samples by design. Three games of data will produce EPA numbers that look dramatic but mean almost nothing. You need at least forty plays per player position to start seeing stable trends. Quarterbacks need hundreds. Treat any single-week Football Math projection as noise unless the sample is large enough to smooth out variance. If your goal is simply to predict game outcomes with reasonable accuracy, you are better off combining Football Math with traditional scouting and injury reports. The model gives you a baseline. Human judgment fills in the gaps that numbers cannot capture. Using one without the other produces worse results than using both together.