What Viral Physics Actually Is

Viral Physics is a Python package for simulating how viruses spread through populations using agent-based modeling and spatial dynamics. It sits somewhere between standard epidemiology tools like SIR models and full-scale individual-level simulations. The core idea is simple: you define a population, give individuals traits (location, contact rate, immunity status), and let the simulation run. The package handles the stochastic transitions between states. Installation is straightforward. You can pull it directly from PyPI with pip, and it depends on numpy, scipy, and matplotlib. I usually keep it in a dedicated virtual environment because the scipy dependency can clash with other packages if you've got a messy site-packages directory. Once installed, the typical workflow looks like this. You create a population object, assign spatial coordinates and contact matrices, define transmission parameters like R0 and incubation periods, then run the simulation. The output gives you time-series data on infections, recoveries, and deaths, plus optional spatial heatmaps.

How It Works Under the Hood

Most people treat Viral Physics like a black box where they plug in numbers and get a graph. That works for basic use cases, but you hit problems fast if you actually need your simulation to match real-world data. The key thing most beginners miss is how the contact matrix interacts with the spatial grid. The default configuration assumes homogeneous mixing within spatial cells, which means if your population has clustered subcommunities, the simulation will overestimate transmission speed by a factor of two or three unless you explicitly define the contact structure. I spent about three days debugging a simulation where the outbreak peaked way too early. The issue turned out to be that I was using the default random spatial distribution instead of assigning realistic residential clustering. Once I mapped the agents to actual census block-level density data and adjusted the contact probability based on distance decay functions, the curve shape aligned much closer to what the real epidemic looked like. That single change took the runtime from about forty minutes down to roughly twelve on the same hardware because the optimized spatial indexing kicked in once the grid wasn't completely uniform.

Common Pitfalls

One thing the documentation doesn't emphasize enough is memory usage. The agent-based approach means every individual is a tracked object. A simulation with even a modest population of one hundred thousand agents across multiple time steps can consume several gigabytes of RAM. If you're working on a machine with limited memory, you'll need to either reduce the agent count, increase the spatial cell size, or switch to the approximate mean-field mode the package offers, which trades individual-level accuracy for dramatically lower resource requirements. Another issue is parameter sensitivity. Small changes in the contact rate or recovery period can produce wildly different outcomes because the model is stochastic. Running a single simulation run is basically meaningless. I usually run at least fifty replicates per parameter set and look at the distribution of outcomes rather than any single trajectory. The package has a built-in batch runner for this, but you have to call it explicitly — it doesn't happen automatically.

Get the Full Details

Physics of Viruses – STATISTICAL PHYSICS UB
Physics of Viruses – STATISTICAL PHYSICS UB

When Viral Physics Falls Apart

There are scenarios where this tool simply won't work well. If you need to model airborne transmission in complex indoor environments with ventilation dynamics, Viral Physics isn't the right choice. It handles person-to-person contact and spatial movement, but it doesn't simulate air flow or aerosol physics. For that, you'd need a computational fluid dynamics tool combined with an epidemiological model, which is a completely different engineering problem. Similarly, if your study population is very small — say under ten thousand agents — the computational overhead of the agent-based framework isn't justified. A standard compartmental model would run in seconds instead of minutes and give you sufficiently accurate aggregate results. The agent-based approach really starts paying off around the hundred-thousand-agent mark where individual heterogeneity matters.

Practical Tips for Getting Useful Results

Calibration is where most people struggle. You need real data points to tune your parameters, and without them the simulation is just generating plausible-looking noise. Even crude estimates help — case counts from a few weeks into an outbreak, hospitalization rates, or seroprevalence survey results can all constrain the parameter space enough to make the output actionable. I typically use a simple grid search over the most uncertain parameters, running each combination through the batch runner, then comparing the simulated curves against the observed data by mean squared error. Exporting your results is also worth noting. The default output format is a nested dictionary structure that isn't immediately useful for visualization or further analysis. Converting it to a pandas DataFrame or saving to a parquet file early in your pipeline saves a lot of time later. The package includes a utility function for this, but it's buried in the examples section rather than the main documentation.