Eye Tracking Is Getting Practical Again
The hardware finally caught up to what the algorithms promised, and now every major UX lab, marketing firm, and cognitive psychology group is running eye tracking on a weekly basis instead of treating it as a specialized research tool. The market has split into two camps that don't talk to each other much: the high-end lab-grade systems like SMI and Tobii Pro that cost ten thousand dollars or more and require a dedicated room with controlled lighting, and the consumer-grade units built into laptops or attached as standalone cameras that have improved enough to be genuinely useful for lighter work but still introduce noise at the edge cases. I've been setting up and maintaining eye tracking rigs since about 2013, and the shift from "this is barely reliable" to "this works most of the time if you know what you're doing" happened somewhere around 2019 when IR LED placement and camera calibration routines stopped being a black box. Before that, I'd lose half my participants to calibration failure before they even started a task. Now the failure rate is more like five to eight percent when you do it right, which is still annoying but manageable.
Current Trends In Eye Tracking Research
The biggest trend right now is not AI-based gaze prediction the way most people imagine it. It's multimodal integration — combining eye tracking data with EEG, physiological sensors, and sometimes even speech or gesture capture — because a single gaze stream tells you about attention, not about comprehension, frustration, or decision-making. That distinction matters more than the raw accuracy numbers that get thrown around in marketing materials. The second trend is mobile eye tracking becoming actually viable. Head-mounted eye trackers like the Pupil Labs core or the Meta glasses-based systems have been around for years, but the processing pipeline is now fast enough that you can stream and analyze data in real time during field studies instead of collecting a hard drive's worth of data and processing it weeks later. This is what enabled the recent wave of research on naturalistic visual behavior in crowds, stores, and outdoor environments instead of the same sterile lab setups we'd been stuck with. There's also a quiet trend toward open-source toolchains. The old model was proprietary software that cost as much as the hardware. Now there are active projects like OpenSeeFace, GazePoint.js, and various Python libraries that pull raw IR camera data and produce gaze estimates without requiring a vendor license. This has dramatically lowered the barrier for small teams and academic labs that operate on grant budgets that haven't kept pace with hardware costs.
When I set up a new study, the first thing I decide is whether I'm using an absolute or relative calibration scheme. Most people default to absolute calibration because the software makes it easy, but for many experimental designs relative calibration — anchoring to the participant's own gaze points rather than screen coordinates — produces measurably better data. The tradeoff is that your analysis pipeline needs to handle that distinction, and if you're comparing across participants you have to account for individual differences in where they naturally look.
Get the Full Details

What Actually Works In Practice
Here's the part nobody puts in the methods section. Pupil diameter is not a reliable measure of cognitive load unless you control for luminance with extreme precision. I spent about three months trying to use pupil dilation as an indicator of mental effort in a 2021 study, and the signal was completely drowned out by ambient light fluctuations from HVAC vents cycling on and off in the lab. The fix was to put a calibrated lux sensor in the participant's field of view and include it as a covariate in the analysis, which reduced usable data by about forty percent but gave me results I could actually defend. Gaze prediction models that use deep learning to estimate where someone is looking from webcam footage alone are impressive at scale but fail in specific ways that matter. They tend to overpredict central fixation bias and underpredict peripheral scanning patterns, which means your heatmap will look cleaner than it should. If you're building a product recommendation system around this kind of technology, you'll want to validate against a ground-truth dataset that includes a broad range of eye movement patterns, not just controlled lab tasks. The industry-standard approach for data cleaning is still a combination of velocity-based and dispersion-based fixation detection, usually implemented through the Psychtoolbox or custom Python scripts using the EyeLink processing pipeline. The catch is that parameter tuning matters enormously — a default velocity threshold that works for one participant demographic may miss or fragment fixations entirely for another. I calibrated those parameters individually per subject using the artifact rejection tools in the manufacturer's software before switching to automated pipelines, and it took me roughly twenty minutes per participant compared to about five minutes with defaults.
Setup Considerations That Actually Matter
Lighting conditions make or break remote eye tracking more often than camera quality does. An infrared-based system will work fine under normal office fluorescents, but if your test environment has windows with daylight entering from the side, you'll get specular reflections on the cornea that the algorithm interprets as pupil center shifts. The workaround is either polarizing filters on the IR LEDs or positioning the cameras so that the dominant ambient light source is roughly aligned with the optical axis, which is easier said than done in a real office space. Head stabilization is still relevant even on systems that advertise "wide eye-box" or "hands-free tracking." The Tobii Pro Spectrum claims reliable tracking within a fifty-by-fifty-millimeter window, and it's mostly accurate there, but the error rate climbs nonlinearly once the head moves beyond that zone during natural viewing. I keep participants loosely restrained with a chin rest even when using these systems because the data quality improvement is worth the slight increase in participant discomfort. Most people adapt within the first minute anyway. Sample rate selection is another place where people overspend. Ninety hertz is sufficient for nearly all UX and marketing applications. Twelve hundred hertz is needed only if you're measuring saccadic micro-movements or studying specific oculomotor disorders. I've seen research groups pay for high-frequency sampling and then aggregate the data to ninety hertz in post-processing anyway because their experimental questions didn't require the resolution, which is effectively paying twice for nothing.
Common Pitfalls
The most common mistake I see is treating eye tracking data as ground truth for visual attention. A fixation means the eyes paused near a location, not that the person understood, liked, or even consciously registered the stimulus. People frequently conflate gaze with interest and then build entire intervention strategies around that assumption. The literature on this gap is extensive — I'd recommend reading the work by Raymond and Holmqvist if you haven't already. Another issue is the so-called monitor effect, where participants behave differently when they know they're being tracked. This isn't just about reactivity in the social sense. The presence of the equipment changes where people naturally direct their gaze because the hardware subtly reshapes their posture and head position. I've run control conditions with and without the eye tracker mounted and seen meaningful differences in scanpath patterns even when the visual display was identical. Dark-haired participants present a well-known but inconsistently addressed problem for webcam-based and some remote IR systems. The contrast between the iris and surrounding sclera is lower, and the pupil boundary is harder to segment reliably. If your participant pool includes a significant number of individuals with dark irises, you should test your specific setup with those participants before committing to a full study. A few months ago I ran a pilot where the average calibration accuracy for dark-eyed participants was nearly double the error rate compared to light-eyed participants on the same device, and the published papers using that system didn't mention it once.

Where This Field Is Heading
Expect further integration with machine learning pipelines that go beyond describing where people look to predicting what they'll do next. There are already working prototypes that combine gaze sequences with page layout features to forecast click-through rates in e-commerce settings with moderate accuracy. The limitations are real — these models degrade quickly when deployed outside the domain they were trained on, and they encode whatever biases exist in the training data — but the direction of travel is clear. Another area gaining traction is ecologically valid eye tracking, meaning studies conducted outside controlled labs in environments that actually resemble the target context. Shopping malls, classrooms, hospitals, and construction sites have all become feasible with current head-mounted and wearable systems. The data is messier, the calibration is harder, and the sample sizes tend to be smaller, but the external validity improvement is substantial compared to lab-based equivalents. The regulatory landscape is also shifting. GDPR and similar frameworks now explicitly treat biometric data from eye tracking as sensitive personal data in many jurisdictions, which changes how you can collect, store, and share the data. This is less of a technical problem and more of an operational one, but it's the kind of thing that catches research groups off guard when they're already deep into a study.
If you're starting a project and need a pragmatic entry point, begin with a remote tracker at ninety hertz, calibrate per participant, validate your cleaning pipeline on a small pilot before committing to the full sample, and budget at least twice as much time for data processing as you think you'll need. The hardware problem is mostly solved. The data interpretation problem still isn't.