The Hardware Layer
Eye gaze technology sits behind an infrared camera, usually mounted below a monitor or built into a tablet. The camera fires near-infrared light at your eye, and that light reflects off the cornea in a predictable way. A pupil detection algorithm finds the dark pupil in the image. A corneal reflection point tracks where the IR light hits the curved surface of the eye. The vector between those two points tells you where the eye is looking, relative to the camera. It sounds straightforward until you actually try to make it work in a production environment. Most commercial systems like Tobii, SMI, or even the newer smartphone-based solutions use something called Purkinje image tracking or pupil-center corneal reflection (PCCR) models. The PCCR approach measures the angle between the center of the pupil and the first Purkinje image, which is the bright reflection from the corneal surface. That angle maps to a direction vector in 3D space. Software then raycasts that vector against a screen plane to figure out where you're looking. The math is trigonometry and linear algebra, not anything mystical. I spent about eight months debugging a deployment where the raw gaze accuracy was drifting by 2 to 3 degrees of arc after twenty minutes of use. The root cause turned out to be thermal drift in the IR LED driver circuit. The LEDs warmed up, their output intensity shifted, and the pupil segmentation algorithm started misclassifying the pupil boundary. I ended up adding a ten-second warm-up period before the system went into calibration mode and switching to a constant-current LED driver instead of the cheap PWM one the board had come with. Accuracy stabilized within half a degree after that change. That kind of problem doesn't show up in any whitepaper.
Calibration and Its Actual Cost
Every eye tracker requires calibration before it's usable, and this is where most people underestimate the time investment. A standard nine-point calibration takes about two minutes. But that assumes the user keeps their head reasonably still and the lighting conditions stay stable. In real-world deployments, I've seen re-calibration triggered every forty to sixty minutes because of subtle changes in ambient light or because the subject's glasses caused specular reflections that confused the IR camera. There are two main calibration approaches. The explicit kind is where you stare at dots that appear on screen and the system builds a mapping function from gaze angle to screen coordinates. The implicit or continuous calibration approach uses a reference point, sometimes a webcam feed or a known fixed object, to correct drift without interrupting the user. Both have tradeoffs. Explicit calibration is more accurate initially but degrades over time. Implicit calibration stays more stable but introduces latency and can get confused if the reference object moves.
Accuracy Numbers and What They Actually Mean
The marketing says sub-degree accuracy. The reality depends entirely on distance, head freedom, and eye contrast. At sixty centimeters from the camera with the head fixed in a chin rest, you're looking at roughly 0.5 to 1 degree of visual angle error on decent hardware. That translates to about half a centimeter on a screen at that distance. Move the head freely and error jumps to two or three degrees. For interaction tasks that require precision like selecting small UI elements, that difference is the gap between usable and frustrating. Dark eyes present a specific problem. The IR camera needs a strong contrast between the pupil and the iris to segment the pupil correctly. Dark brown or black irises can have very low contrast against the pupil under IR illumination, making the algorithm guess at the boundary. I solved this on one project by increasing the IR illumination intensity and switching to a camera with a larger sensor and better low-light performance. The combination reduced segmentation errors significantly. Cheap webcams with small sensors just can't handle this scenario regardless of the algorithm.
Get the Full Details

Common Pitfalls That Break Deployments
Ambient IR light is the silent killer of gaze systems. Sunlight contains infrared energy. Fluorescent lights flicker in the IR spectrum. If your deployment environment has windows or certain types of lighting, the camera picks up that noise and the pupil detection becomes unreliable. I had a client who installed a gaze-based kiosk in a lobby with large south-facing windows. The system worked fine at nine in the morning and completely failed by noon. We added a narrow-band IR filter on the camera lens and shielded the sensor from direct sunlight with a physical hood. That fixed it, but it wasn't something the integration guide mentioned. Another issue is blink suppression. Most systems discard data during blinks because the pupil becomes invisible. That creates gaps in the data stream. For reading studies or natural scene viewing, those gaps matter less. For interaction systems where you need continuous input, blink patterns can cause missed commands or unintended selections. Some newer systems use predictive interpolation to fill blink gaps, but the interpolation isn't perfect and can introduce artifacts.
Head Tracking as a Force Multiplier
Pure eye-based gaze has a limited field of view. Once your eyes move past about thirty degrees from the camera axis, accuracy drops sharply. That's why most modern systems combine eye tracking with head tracking. A secondary camera or an IMU tracks head position and orientation. The system then compensates for head movement by adjusting the gaze vector in real time. This gives you a much larger effective field of view, though not without cost. Head tracking adds another source of error, and the fusion of the two data streams requires careful temporal alignment. I once worked on a system where the head tracker had a sampling rate of sixty hertz and the eye tracker was running at two hundred and fifty hertz. Without proper timestamp alignment and interpolation, the fused output was visibly jittery. Synchronizing the clocks between the two devices brought the jitter down to an imperceptible level. Most off-the-shelf systems handle this internally, but if you're building your own pipeline, the synchronization detail is where everything falls apart.
When Eye Gaze Technology Fails Completely
Ney bars and certain corneal conditions can make IR-based tracking impossible. Contact lenses with tint or pattern can scatter IR light in unpredictable ways. Some users simply cannot maintain the fixation stability needed for reliable tracking. I recommend having a backup input method always available. Gaze technology should augment interaction, not replace alternatives entirely. The best deployments I've seen treat gaze as one input modality among several and let the user switch seamlessly. The hardware is only half the problem. Software integration determines whether your system is actually usable. APIs vary significantly between vendors. Tobii has its own SDK with good documentation but locks you into their ecosystem. Open-source options like iViewX or PyGaze exist but require more development effort. For research applications, Python libraries like eye-tracking-toolkit or the gaze submodule in PsychoPy work reasonably well. For production, you want something with stable drivers and vendor support, even if it costs money. Data output is another area that gets glossed over. Raw gaze data includes samples at a certain frequency with position, confidence scores, and pupil diameter. Pupil diameter is interesting because it carries cognitive load information, but it's also the most noisy signal in the dataset. Blink rate, saccade amplitude, and fixation duration are more stable metrics. If your application doesn't need pupil data, filtering it out early reduces processing load and simplifies the pipeline.
