Gene Appel Leaves Willow Creek — What It Actually Is and How to Use It
This is a computational ecology framework built around agent-based species migration modeling. The original implementation traces back to Gene Appel's work on ecological system simulation, and Willow Creek is the particular test environment most people use to validate their models before running larger landscape-scale projections. If you are looking at this from a programming angle, it is IPL-derived logic repackaged for modern ecology workflows. The core idea is straightforward: define organism agents, give them movement rules, and let them traverse a rasterized landscape over simulated time steps. Most people trying this for the first time get stuck on the habitat map configuration. That is where the real friction lives. The framework handles the simulation engine fine, but feeding it clean geospatial data without reprojecting everything first causes subtle bugs that do not throw errors. You just get wrong results and waste two days chasing them.
Gene Appel Leaves Willow Creek Setup and Configuration
Start by cloning the base repository from the canonical source. The package is lightweight. You do not need Docker unless you are planning batch runs across multiple climate scenarios, and even then a local Python environment works fine for models under 100 square kilometers. Here is the installation path that actually works without wasting time: git clone https://github.com/gene-appel/willow-creek.git
cd willow-creek
pip install -e ".[eco]"
After installation, download the Willow Creek reference landscape. It comes as GeoTIFF terrain and land-cover layers. Do not skip the reprojection step. I have seen too many people load the raw WGS84 data directly into the simulator and wonder why their agents move north-south twice as fast as east-west. Reproject to UTM zone 12N using GDAL before feeding anything into the model. The configuration file sits at willow_creek/config/default.yaml. The most important setting is the time step resolution. The default of 1-day steps works for most temperate species, but conifer migration studies need hourly steps during the spring dispersion window. Set that early. Changing it mid-run corrupts the output logs.
Get the Full Details
Running a Basic Migration Simulation
Define your agent species in the YAML config. A single species entry looks like this: species:
name: Pseudotsuga menziesii
movement_model: correlated_random_walk
max_step_km: 0.5
habitat_preference: conifer_forest
dispersal_season: [3, 4, 5]
population_init: 1000 Run the simulation with willow_creek run --config my_run.yaml --output results/basic_migration. A 100-year simulation with 1,000 agents on the standard Willow Creek map takes roughly 4 minutes on a modern laptop. Roughly 200,000 agent-steps per second, give or take depending on your habitat layer complexity.
Check the output immediately after each run. The framework writes three things: agent trajectories as GeoJSON, occupancy rasters per time step, and a summary CSV with population counts and spread rate metrics. If the summary CSV is empty, your habitat preference filter is rejecting all movement. That happens when the land-cover classification codes in your map do not match the ones the config expects. Cross-reference the codebook in the repository docs.
A Real Problem I Ran Into and How I Fixed It
Last year I was running a scenario where the simulated population should have colonized a valley floor but instead vanished entirely by step 12. No error message. No warnings. The agents just disappeared from the trajectory file. I spent three days debugging the movement logic, rewrote the habitat suitability function twice, and nearly gave up on the whole project. The actual cause was a nodata value mismatch between the terrain DEM and the land-cover layer. One layer used -9999 for no data, the other used 0. When an agent stepped onto a pixel flagged as nodata by one layer but valid in the other, the simulator dropped it silently. I wrote a quick preprocessing script that standardized both rasters to the same nodata value and resampled them to identical extents before running the model. The population survived past step 12 and behaved exactly as expected. If you are doing multi-layer habitat models, always validate nodata consistency across every input raster before starting a simulation. Run gdalinfo on each file and check the NoData values. It takes 30 seconds and saves you days of troubleshooting later.

Counter-Intuitive Things Beginners Miss
The first thing most people get wrong is assuming finer resolution always produces better results. With this framework, increasing spatial resolution from 30 meters to 10 meters does not improve accuracy for migration modeling. It slows the simulation by roughly 9x because the agent step count scales with area, and the movement logic does not benefit from the extra detail at that scale. Stick with 30m unless you have a specific reason tied to your landscape features. For most broad-scale species distribution work, 30m is the practical ceiling. The second thing is ignoring the correlation parameter in the movement model. The default correlated random walk uses a turning angle distribution that keeps agents moving in roughly the same direction between steps. If you set the correlation coefficient too low, agents perform a near-diffusive search and your simulated spread rates look unrealistically slow. The literature on animal movement suggests values between 0.7 and 0.9 for most terrestrial mammals and birds. I usually start at 0.85 and adjust based on observed field data if I have it. Also worth noting: the framework does not natively support temporal habitat change within a single simulation run. If your landscape is going to shift due to fire, logging, or climate-driven succession, you need to either precompute static habitat layers for each time period and run disconnected simulations, or write a wrapper script that swaps the habitat raster between simulation blocks. I use a Python wrapper that loads the base config, runs 50 years, replaces the land-cover raster with the next time slice, and continues. It is not elegant. It works.
Limitations and When to Look Elsewhere
This framework is not a general-purpose agent-based modeler. It is narrowly built for species migration and dispersal across raster landscapes. If you need to simulate intraspecific competition, predator-prey dynamics, or landscape-level nutrient flows alongside movement, you are better off using something like NetLogo or Mesa. Gene Appel Leaves Willow Creek handles the core use case well but does not attempt to be a multipurpose platform. Another limitation is memory usage. Running a simulation with more than 50,000 agents on the full Willow Creek extent will consume several gigabytes of RAM. The trajectory storage is the main culprit. If you need large populations, disable per-agent trajectory logging and keep only periodic snapshots. You lose individual movement history, but you keep the occupancy and summary statistics that matter for most papers. Documentation is thin. The repository has a README and a config schema reference, but there is no tutorial library. The examples folder contains three working scenarios, which is enough to get started but not enough to cover edge cases. I rely on reading the source code directly for questions about the movement kernel and habitat filtering logic. The code is readable. It is not obfuscated. Spending an hour in the relevant modules usually answers whatever question you have.
If you need cloud-scale parallel simulations across hundreds of parameter combinations, build a wrapper using the CLI interface. The framework runs headless and accepts config paths, so a simple shell loop or a Python multiprocessing script handles batch execution without modification. I have run parameter sweeps of 200 configurations across 48 cores without issues. The bottleneck becomes disk I/O at that scale, not the simulation itself.
