Working with Arthur Getis's Approach to Spatial Analysis
If you have ever tried to map crime hotspots or disease clusters and ended up with something that looked nice but meant nothing, you are not alone. Arthur Getis built his reputation on the observation that most people doing spatial analysis skip the fundamentals and jump straight to coloring a map. His Introduction Geography Arthur Getis materials are widely used because they force you to deal with the actual mechanics of spatial relationships before you try to draw conclusions. The core idea is straightforward but easily overlooked. Getis emphasized that geographic phenomena are not randomly distributed. Clusters exist, and identifying them requires specific statistical tools rather than eyeballing a choropleth map. His Gi* statistic, developed with Deborah Ord, became the standard for measuring spatial clustering of high or low values. The formula itself is elegant enough to write on a napkin, which is exactly what most students do before realizing how much depends on your weight matrix.
Understanding the Getis-Ord Gi* Statistic
The Gi* statistic tests whether the local cluster of high or low values around each feature is statistically significant. You calculate it by summing neighboring values, weighting them by distance or adjacency, and comparing that sum to what you would expect under a random distribution. A high Gi* with a significant p-value means you have a hotspot. A low Gi* with significance indicates a coldspot. Values near zero simply mean no clustering pattern exists at that location. Here is where people routinely mess up. The weight matrix choice changes everything. If you use a binary contiguity matrix on irregularly shaped census tracts, your results will be skewed because larger tracts get artificially fewer neighbors while smaller ones dominate the calculation. I spent an entire afternoon debugging a heatmap that showed suspiciously regular clustering patterns until I realized the tracts in the suburban ring were roughly the same size, creating a false equilibrium in the neighbor counts. Switching to a distance-band weight with a sufficiently large threshold resolved it immediately.
Setting Up Your Analysis Correctly
Most tutorials skip the data preparation step, which is unfortunate because it is where 80 percent of failures happen. You need a shapefile or feature class with consistent geometry, attribute data with no nulls in the variable you are testing, and a properly defined coordinate system. Using unprojected lat/long coordinates with a distance-based weight matrix will give you incorrect distance calculations unless your study area is very small. Run the Spatial Statistics Tools in ArcGIS Pro or the pysal package in Python. Both handle the Gi* calculation well. The pysal approach gives you more flexibility with weight matrix construction, which matters significantly if you are working with point data or mixed-scale regions. ArcGIS is faster to set up but harder to customize when your data does not fit the default assumptions. Check your data distribution before running anything. Getis's method assumes you are working with continuous or interval-level data. Throwing ordinal survey responses into a Gi* calculation produces numbers that look precise but are statistically meaningless. I once saw a graduate student publish a paper using Gi* on Likert-scale county voting data. The maps were colorful and the p-values were significant. None of it held up under basic scrutiny because the underlying variable did not meet the assumptions required for the spatial lag term.
Get the Full Details

Interpreting Results Without Overstating Them
A significant Gi* value does not prove causation. It proves spatial clustering. The difference matters enormously when you are presenting to city planners or journal reviewers. The common mistake is treating a hotspot as an explanation rather than a starting point for investigation. Getis himself was careful about this distinction throughout his career. The statistic identifies where clustering exists, not why it exists. You should also check for spatial nonstationarity. The assumption that the process generating your data is the same across your entire study area is often wrong. Rural-urban gradients, for example, can produce misleading Gi* values because the mechanisms driving values in one zone differ fundamentally from another. Running a separate Gi* analysis within subdivided zones or using a geographically weighted version gives you more reliable results when that is the case. The main limitation of Getis's framework is that it works best with complete spatial data. Gaps or missing areas distort the neighbor relationships and can create artificial coldspots at the edges of your study region. If your data has substantial missing coverage, consider using a kernel-based approach or explicitly modeling the absence mechanism rather than forcing a standard weight matrix on incomplete data. It adds complexity but saves you from publishing results that look like artifacts.
For a solid foundational text, search for Introduction Geography Arthur Getis materials through your university library or academic publishers. Getis has published extensively on spatial regression and geographic information systems, and his teaching materials reflect the same emphasis on statistical rigor over visual appeal. His work remains relevant because the mistakes people make with spatial analysis have not changed in any fundamental way since he started writing about them.