Ecological stats are not where most people think they are

Most field biologists learn enough statistics to run a t-test and call it a day. Then they get stuck when their species count data violates every assumption in the book. This is why A Primer Of Ecological Statistics exists as a concept rather than a single textbook. The field is fragmented across general statistics courses that ignore spatial autocorrelation, ecology classes that gloss over model selection, and a bunch of R tutorials that assume you already know what you are doing. I spent seven years trying to build occupancy models for amphibian survey data before I stopped treating zero-inflation as a nuisance and started treating it as the actual signal. That shift changed everything about how I approach ecological data.

Why A Primer Of Ecological Statistics Does Not Have One Author

There is no single canonical primer because ecological data comes in too many broken shapes. You have presence-absence matrices with false zeros, count data with extra variance, repeated measures nested inside sites nested inside watersheds, and all of them violating independence in different directions. A single book cannot cover all of that without becoming either too shallow or too long to finish. The closest thing to a unified introduction is the work by Kéry and Schmidt combining frequentist and Bayesian frameworks, supplemented by MacKenzie et al. on occupancy modeling, and Bolker's broader applied text. Reading them in sequence takes about six to eight weeks if you are actually working through the examples instead of skimming.

What the field actually requires you to know

Start with generalized linear mixed models before you touch anything fancy. The hierarchy in your data structure matters more than the distribution you pick. I have seen people spend three days tuning a zero-inflated negative binomial only to realize the random effect for site was completely unmodeled and the whole thing was confounded with overdispersion. Check for overdispersion first. Always. The DHARMa package in R does residual simulation checks that take about two minutes and will save you from publishing a model that looks significant but is actually just overfit noise. Running a standard dispersion check on a glmmTMB model after fitting takes roughly forty seconds on a modern laptop. Spatial autocorrelation is the silent killer in ecological datasets. If your sampling points are closer than about two kilometers apart and you are modeling species distributions, your p-values are probably too optimistic. The Moran I test runs fast enough that there is no excuse for skipping it. I learned this the hard way during a vegetation survey project where the initial model showed five significant predictors. After adding a spatial error term, two of those variables became non-significant and one flipped direction. The total time cost of adding the spatial component was about twenty minutes of coding and three minutes of computation.

Get the Full Details

University of Guelph Bookstore - A Primer of Ecological Statistics
University of Guelph Bookstore - A Primer of Ecological Statistics

The counter-intuitive stuff nobody mentions in intro courses

Zero-inflation and overdispersion are not the same thing. Beginners treat them interchangeably because the diagnostic plots look similar, but they arise from different mechanisms. Zero-inflation means there is a separate process generating structural zeros. Overdispersion means the variance exceeds what the distribution predicts. A zero-inflated model will give you biased estimates if you actually have overdispersion without excess zeros. The fix is usually just adding an observation-level random effect to absorb the extra variance rather than reaching for a zero-inflated distribution immediately. Model selection with AICc on ecological data is fine until you have correlated predictors. Then you are selecting between models that are essentially fitting the same signal differently. I ran into this with a stream macroinvertebrate dataset where pH and conductivity were correlated at r = 0.82. The top-ranked model kept flipping between those two variables depending on which random effects were included. The workaround was to use variance partitioning with the varpart function in the vegan package, which takes about five minutes to run and clearly shows which predictor is carrying unique explanatory power versus shared variance.

When standard approaches fail completely

Bayesian hierarchical models sound like the solution to everything but they are computationally brutal on large ecological datasets. I tried fitting a Bayesian occupancy model to a bird dataset with 340 sites and twelve survey occasions. The Stan model took forty-seven minutes per chain and still had convergence issues with R-hat values above 1.1 for the detection probability parameters. Switching to a frequentist glmmTMB implementation with the same structure ran in about three minutes and converged on the first attempt. Occurrence-only data without absence records breaks most standard modeling approaches. Presence-only methods like MaxEnt exist but they make strong assumptions about background sampling that are rarely met in real fieldwork. If you can get any form of targeted absence data from museum records or parallel surveys, it dramatically improves model performance compared to pure occurrence modeling.

Practical workflow that actually saves time

Here is the sequence I use now instead of the one I wasted months on early in my career. Load your data and immediately check the dimensionality and missingness structure. That takes about thirty seconds. Fit a null model with just the random effects structure before adding any fixed predictors. Then add predictors one at a time and track both AICc and conditional R-squared values. The MuMIn package handles this workflow cleanly once you set it up. Diagnostic checks should be automated, not manual. I wrote a simple function that runs DHARMa residuals, checks for spatial autocorrelation with Moran I, and tests for zero-inflation in one pass. It runs in under five seconds per model and catches about eighty percent of the problems that would otherwise require manual investigation. The function is publicly available on GitHub if anyone wants to adapt it. Replication matters more than complexity. A well-fitted GLMM with three predictors and proper random effects beats a complicated occupancy model with questionable detection assumptions every time. My current projects default to simpler structures and only add complexity when diagnostics force me to.

A primer of ecological statistics : Gotelli, Nicholas J., 1959- : Free Download, Borrow, and ...
A primer of ecological statistics : Gotelli, Nicholas J., 1959- : Free Download, Borrow, and ...

Where to start learning this material

Kéry and Schmidt's Introduction to Hidden Markov Models and State-Space Models covers a lot of the advanced stuff but assumes familiarity with likelihood methods. For a gentler entry point, Bolker's text is still the most practical reference even though it was published over a decade ago. The R documentation for glmmTMB and lme4 has improved significantly and serves as a legitimate reference alongside the textbooks. The online course materials from the National Evolutionary Synthesis Center are free and cover most of the core concepts without requiring expensive software licenses. Working through their exercises takes about twelve hours total and gives you enough hands-on experience to handle standard ecological datasets without constantly searching Stack Overflow.