Getting Started With Positional Data
Most people come into football analytics thinking positional data is just tracking coordinates. It isn't. It's a mess of noise, gaps, and assumptions that you have to wrestle into something usable before you even think about modelling anything. I work with this stuff daily. What follows is what actually happens when you try to get meaningful outputs from tracking data. Not the textbook version.
Data Analytics In Football Positional Data Collection Modelling And Analysis
The first thing to understand is that positional data comes in different flavours, and they are not interchangeable. Opta provides event-level data with x and y coordinates at around 10Hz. SportLight and TRACAB give you continuous tracking at 25Hz. StatsBomb's new release is somewhere in between. If you mix these without normalising them, your models will produce garbage results. I learned that the hard way during a scouting project where I combined two sources and spent three days debugging why the spatial clusters looked like random noise. You need a few things sorted first. Your equipment source matters more than most people realise. If you are using camera-based systems, expect occlusion gaps. Players getting blocked by others will create holes in the track. If you are using GPS or inertial sensors, you get clean individual data but you lose the team shape context that cameras provide. Most professional clubs use camera systems. If you are an analyst working independently, you are probably stuck with event data unless you can access tracking feeds. Your second requirement is a consistent coordinate system. Different providers use different scales. Opta uses a 105 by 68 grid normalised to 100 by 100. Some systems use metres. Some use pixel coordinates from the broadcast feed. Convert everything to a standard frame immediately. I use metres on a 105 by 68 pitch as my default. It makes everything downstream easier.
Collecting The Data
Collection is the part everyone underestimates. Here is what the pipeline looks like in practice. First, you need the raw files. If you have access to Wyscout or Hudl Sportscode, they export tracking data in JSON or CSV format depending on the provider. If you are working with Opta events, you get the XML or JSON feed directly. SportLight data comes through their API with per-frame coordinates for every player and the ball. Second, you need to parse and store it. A simple pandas DataFrame works if you are small scale. Each row should contain the match ID, frame timestamp, player ID, x coordinate, y coordinate, and team affiliation. Add a player role column if it is not included, because positional labels matter for filtering.
Get the Full Details

Third, handle the pitch alignment. Cameras are positioned at different angles. Broadcast feeds are angled from the sideline, which distorts distances near the edges. You need a homography transformation if you want to convert pixel coordinates to real pitch dimensions. This is where most people give up. The fix is straightforward if you know the camera parameters. Use the pitch markings visible in the frame to calculate the transformation matrix. OpenCV has functions for this. I wrote a small script that auto-detects the goal lines and touchlines from the first frame and applies the correction. It takes about two minutes per match instead of twenty.
Smoothing And Gap Filling
Raw tracking data is jagged. Player positions jump between frames because of sampling limitations and occlusion. You need smoothing before anything else. A simple moving average introduces lag. A Kalman filter is better but overkill for most use cases. I use a Savitzky-Golay filter with a window of five frames and polynomial order two. It preserves peak velocities and accelerations while removing the jitter. The output is smoother trajectories without the time delay that breaks velocity calculations. Gap filling is another separate problem. When a player is occluded for three or more frames, you need to interpolate. Linear interpolation between the last known position and the next visible position works fine for short gaps up to about half a second. Beyond that, the error compounds. I use cubic spline interpolation for larger gaps but flag any interpolation longer than two seconds as low confidence. Models trained on interpolated data will be biased toward smoother movement than reality. Your spatial metrics will look better than they actually are.
Building Basic Features
Once your data is cleaned, you can start extracting features. The core ones every analyst needs are velocity, acceleration, distance covered, and directional changes. Derive these from the positional coordinates using finite differences. Velocity is the change in position divided by the change in time. Acceleration is the change in velocity over time. Be careful with the edge frames. The first and last frames do not have previous or future points, so pad them or drop them. I drop them. It is cleaner. Distance covered is straightforward. Sum the Euclidean distances between consecutive frames for each player across the match. Convert frames to seconds using your sampling rate. A 25Hz system means each frame represents 0.04 seconds. Multiply velocity by that interval to get distance per frame. Directional changes matter more than people realise. Calculate the angle between consecutive velocity vectors. A change above a certain threshold indicates a turn or stop-start movement. This is useful for identifying high-intensity actions that distance alone misses. I threshold turns at 45 degrees. Anything sharper counts as a directional change.

Space And Team Shape Metrics
This is where positional data becomes interesting. Space metrics require you to define what space means in your model. The most common approach is Voronoi tessellation. Each player owns the area of the pitch closest to them. As players move, the cells shift. The aggregate cell area gives you a proxy for spatial control. Voronoi cells are computationally cheap but they assume players move to occupy nearest space, which is not always true. Players hold positions. They cover zones. A better alternative for team shape analysis is convex hulls or alpha shapes. The convex hull connects the outermost players to form a polygon. The area inside that polygon represents the team's defensive or attacking coverage. Alpha shapes refine this by allowing concave boundaries based on a radius parameter. I use alpha shapes with a radius of 10 metres. It produces tighter bounds than convex hulls without being as sensitive as raw Voronoi cells. I encountered a specific problem with alpha shapes during a defensive analysis project. The opposing team played a very compact low block with all eleven players clustered in a small area. The alpha shape collapsed to nearly zero area because the radius was too large relative to the cluster density. I adjusted the radius dynamically based on the mean pairwise distance between players. If the average distance was below eight metres, I halved the radius. This prevented the metric from breaking on compact formations while keeping it stable on spread-out teams.
Modelling Approaches
There are several ways to model positional data. The simplest is clustering. K-means or DBSCAN can identify recurring positional patterns across matches. I use DBSCAN because it does not require you to specify the number of clusters in advance. It finds dense regions of player positions and treats outliers as noise. The output tells you where players naturally operate. Set eps to 15 metres and min_samples to ten for semi-professional data. For top-level tracking at 25Hz, reduce eps to 10 metres because positions are more precisely defined. For temporal modelling, hidden Markov models work surprisingly well for identifying states like attacking build-up, counter-attack, or defensive set shape. Each state has a probability distribution over pitch zones. The model infers the hidden state at each frame based on the observed positions. Training requires labelled data. You can manually annotate a subset of matches and use those labels to train the emission probabilities. I typically label 30 to 50 matches for a stable model. More than that and the returns diminish. Less than that and the state definitions become unreliable. Deep learning approaches exist but they require significant data and computational resources. Convolutional neural networks applied to positional heatmaps can predict outcomes like shot probability or chance creation. The problem is that these models are black boxes and they generalise poorly across leagues. A model trained on Premier League positional data does not transfer well to the Championship or Serie B. The movement patterns and spatial structures differ enough to cause performance drops of 15 to 20 percent. I recommend starting with simpler models unless you have a specific reason to go deeper.
Common Pitfalls
Here are the mistakes I see repeatedly. Overfitting to a single match is the biggest one. One game does not represent a player's tendencies. Aggregate across at least ten matches before drawing conclusions about positional behaviour. Another pitfall is ignoring the off-the-ball context. Positional data alone cannot tell you why a player moved where they did. They might have been drawing a marker, creating space for a teammate, or reacting to a tactical instruction. Combine positional data with event data to add that layer. Match each positional frame with the nearest event within a two-second window. This gives you intent context without requiring video review. A third issue is normalization errors. If you normalise coordinates differently between training and inference, your model breaks. Define your normalisation once and apply it consistently. I store the min and max values from the training set and use them for all subsequent matches. Never re-normalise per match.

Tools And Implementation
The ecosystem for this work is mostly Python-based. The key libraries are numpy for array operations, pandas for data handling, scikit-learn for clustering and basic ML, scipy for signal processing and interpolation, and matplotlib or plotly for visualization. For Voronoi and spatial calculations, scipy.spatial provides Delaunay triangulation and voronoi_finite_regions. For alpha shapes, there is a package called alphashape on PyPI. If you want a complete working example, the most practical starting point is the pybaseball tracking repository adapted for football, or the statsbomb-open-data notebooks on GitHub. They contain parsing scripts and feature extraction pipelines that you can modify. I built my own pipeline by combining elements from both and adding the dynamic alpha shape adjustment I described. The full code is not publicly hosted, but the structure is straightforward enough to replicate.
What This Approach Cannot Do
Positional data modelling has real limitations. It cannot capture technical skill. A player's first touch or passing quality is invisible in coordinates. It cannot replace video analysis for decision-making assessment. It also struggles in set-piece situations where player positioning is highly structured and repetitive. The models will identify the pattern but not the strategic purpose behind it. If your analysis goal is understanding why something happened, positional data alone will not answer that. It tells you what happened spatially and temporally. The why requires contextual information outside the dataset. The honest takeaway is that positional analytics is a powerful tool for describing spatial behaviour and identifying patterns, but it is one input among many. The analysts who get the best results combine it with event data, video, and domain knowledge rather than treating it as a standalone solution.