Working With Sample Averages in Practice

The Law Of Big Numbers is one of those concepts that sounds like basic textbook material until you try to apply it to real data and realize half the people using it are misunderstanding what it actually guarantees. I spent about three years dealing with this in insurance actuarial work and later in A/B testing for a payment processing platform, so I have some practical experience with where it holds up and where it quietly breaks. At its core, the theorem states that as your sample size grows, the sample mean converges to the expected population mean. That convergence gets tighter. The standard error of the mean shrinks at a rate proportional to the square root of n, which means to cut your margin of error in half you need four times the sample size, not two. This is the part most people miss when they start running experiments without thinking about it. What it does not say is that individual observations will look like the average. It also does not tell you anything about the speed of convergence for distributions with heavy tails or infinite variance. If your data comes from a Pareto distribution or a Cauchy distribution, increasing your sample size might not help much at all, and the sample mean can behave erratically no matter how large n gets.

I learned this the hard way in 2019 when I was building a pricing model for a small commercial line of business. We had about 4,200 policyholder observations and the aggregate loss ratio looked deceptively stable month over month. The law of large numbers should have been comforting us there, right. Instead we were sitting on a portfolio where about 8 percent of the policies accounted for roughly 63 percent of the claims, and those large claims were arriving in clumps rather than independently. The effective sample size was nowhere near 4,200. I ran a bootstrap resample to check the variability of the aggregate loss ratio under different clustering assumptions and found our confidence intervals were off by a factor of about 2.4 compared to what the textbook formula gave us. We ended up adding a coverage load instead of trusting the raw average, which cost us some competitive price on a renewal but saved us from a underpricing spiral that would have taken two years to unwind.

How To Apply It Without Making Mistakes

When you are actually using this in a workflow, the first step is checking whether your data satisfies the conditions that make the theorem useful. The classical version requires independent and identically distributed observations with a finite expected value and finite variance. Real data rarely meets all three cleanly, so you need to verify each one before proceeding. Independence check. If your observations are correlated in time or space, the effective sample size drops. In my payment processing work, we had transaction data where fraud patterns clustered around certain merchant categories and weekends. A naive application of the law would treat 50,000 transactions as 50,000 independent data points. They were not. I calculated the intraclass correlation coefficient across transaction batches and found an average ICC of about 0.07, which reduced the effective sample size to roughly 28,000. That mattered when we were estimating fraud rate thresholds for a new scoring model. Variance check. Look at your data before you trust any average. Plot the histogram, check the tail behavior, compute the kurtosis. If your kurtosis is above 3, you have heavier tails than a normal distribution and your convergence will be slower. If you suspect infinite variance, use the median or a trimmed mean instead of the arithmetic mean. The law still applies to the median under mild conditions, but the convergence rate changes and you need different confidence interval formulas.

Get the Full Details

Gavel for court of law icon | Free stock photo - 402117
Gavel for court of law icon | Free stock photo - 402117

Convergence diagnostics. Do not just compute one big average and move on. Split your data into sequential blocks and track how the running mean behaves as each block is added. If it is still drifting after you have collected what you think is a large sample, you may be dealing with a non-stationary process. In those cases the law of big numbers is not going to save you because the population mean itself is moving. I ran into this with a SaaS churn model where the baseline churn rate shifted significantly after a product redesign, and my first three months of post-launch data gave a mean that was completely irrelevant to the fourth month onward. I had to segment by cohort and re-estimate after each structural change rather than pooling everything together.

Practical Guidelines For Sample Size Estimation

If you need to determine how much data to collect before an estimate is useful, start with the formula for the confidence interval width around a mean: the half-width is approximately 1.96 times the standard deviation divided by the square root of n for a 95 percent interval. Rearrange that to solve for n and you get n equals 1.96 squared times sigma squared divided by the desired half-width squared. The problem is you usually do not know sigma beforehand, so you run a pilot sample of maybe 50 to 100 observations, compute the sample standard deviation from that, and plug it in as an estimate. This gives you a rough target, not a guarantee, because the pilot estimate itself has sampling error. In practice, for business decisions I usually aim for an effective sample size that puts the half-width of the 95 percent interval below 5 percent of the mean for ratio estimates and below 0.02 for proportions that are not too close to zero or one. That usually means between 400 and 2,500 independent observations depending on the variance in your data. If your variance is high relative to the mean, you need more. If you are dealing with rare events, the variance of a proportion p is p times 1 minus p, so for p equal to 0.01 the variance is 0.0099 and you need roughly 3,800 observations to get a half-width around 0.0098, which is barely acceptable. For rarer events at p equal to 0.001, you are looking at well over 38,000 observations just to get a decent interval, and even then the normal approximation starts breaking down and you should switch to an exact binomial or Poisson-based method. One thing that trips people up is the difference between the law of big numbers and the central limit theorem. The law of big numbers says the average converges to the true mean. The central limit theorem says the distribution of that average becomes approximately normal as n grows. The law does not require normality of the underlying data. The CLT does, in its standard form, and it requires finite variance. You can have a perfectly valid convergence of the mean under the law while the distribution of the standardized sum remains non-normal for a very long time if the underlying distribution is skewed or heavy-tailed. I have seen teams call a result "validated by the law of large numbers" when they really meant the CLT applied, and then they built around a normal approximation that was completely wrong for their n. The distinction matters when you are working with n in the low hundreds and a highly skewed distribution like claim sizes or ad click conversion rates.

Where The Law Of Big Numbers Fails Completely

The theorem assumes a stationary population. If your data generating process changes over time, the sample mean is estimating a moving target and convergence becomes meaningless in the way you expect. This happens constantly in marketing analytics where campaign effects decay, seasonality shifts, and external events alter baseline behavior. I once saw a team average 14 months of website conversion data and then complain that their holdout test results did not match the historical average. The historical average was a fiction because the site had launched a new checkout flow in month 6 and a major brand campaign in month 10. Pooled together, the mean estimated nothing useful. The theorem also fails for processes with memory or long-range dependence. Financial returns, network traffic, and some ecological datasets exhibit patterns where correlations decay so slowly that the effective sample size grows much more slowly than the actual sample size. In those cases you need techniques like block bootstrapping or spectral methods to estimate the true variability, and the naive standard error formula underestimates uncertainty by a large margin. I dealt with this in a time series forecasting project for retail demand where the autocorrelation function decayed hyperbolically rather than exponentially. Using the naive formula for the standard error of the mean gave us intervals that were too narrow by roughly a factor of 3.5 compared to a block bootstrap with block size chosen by the spectral diversity method. Another failure mode is selection bias. If your sample is not representative of the population you care about, increasing n just makes your biased estimate more precisely wrong. This is probably the most common practical problem. I have seen this in survey-based research where response rates dropped below 15 percent and the analysis proceeded as if the respondents were a random sample. With n equal to 8,000 and a nonresponse bias that shifted the mean by 0.3 standard deviations, the confidence interval was extremely narrow and completely missed the true population parameter. The law of large numbers gave a false sense of security because the math was correct for the data that was collected, but the data was not the data that mattered.

Free of Charge Creative Commons criminal law Image - Legal 17
Free of Charge Creative Commons criminal law Image - Legal 17

Alternatives When The Law Of Big Numbers Is Not Enough

If your data is small, non-stationary, or heavily correlated, Bayesian methods often give you more usable results than pure frequentist averaging. You can incorporate prior information, model the uncertainty in the variance directly, and get full posterior distributions rather than point estimates with symmetric intervals. In my insurance work, we switched to hierarchical Bayesian models for small geographic segments where n was under 200. The shrinkage estimates pulled extreme segment averages toward the national mean in proportion to the segment size, which produced more stable and accurate predictions than raw averaging or regression adjustment alone. The computational cost was higher, but with modern MCMC samplers or variational inference, a model that used to take hours now runs in minutes on a standard workstation. For high-dimensional or sparse data where the classical law does not apply cleanly, concentration inequalities like Hoeffding's bound or Chernoff bounds can give you finite-sample guarantees that do not require asymptotic approximations. These are looser than asymptotic confidence intervals but they are valid for any n, which matters when you are making decisions with 50 or 100 observations and cannot afford to wait for asymptotics to kick in. I use these when I need to set upper bounds on risk exposure with limited data, such as estimating the maximum likely loss from a new product line with only a handful of early claims. There is no single tool that fixes every problem. The law of big numbers is a foundational concept that tells you something important about what happens when you collect enough independent, stationary data. It is not a magic wand that makes bad data good or guarantees precision when the underlying assumptions are violated. Knowing when to use it, when to adjust it, and when to walk away and use something else is what separates people who understand it from people who just repeat the definition.