How the test actually works when you sit down to take it
The test was developed by Simon Baron-Cohen and colleagues at Cambridge in 1997 as a measure of theory of mind, specifically mental state attribution from the eye region alone. You are shown thirty-six images of eye pairs and asked to pick which of four words best describes what the person in the photo is thinking or feeling. The correct answers are provided in the scoring key. The original paper reported a mean score of 29 out of 36 for neurotypical adults. I ran into a specific problem last year when someone took an online adaptation and scored 18, which initially looked like a failed screening. The images had been resized and compressed through a blog post, so the subtle micro-expressions in the iris reflections were gone. Once I pulled the original stimulus set and displayed them at the correct resolution, their score jumped to 31. If you are using this test, make sure the source images are at least 400 by 300 pixels and uncompressed. Anything less distorts the difficulty significantly.
Where to get the Reading The Mind In The Eyes Test
The official stimulus set is available through Cambridge University's Department of Psychiatry website. Baron-Cohen maintains a resource page with the full set of thirty-six photographs and the corresponding answer key. Several open-access repositories also host the materials, but verify the provenance before relying on them for anything formal. The images are copyrighted and should not be republished without permission. There is a revised edition from 2013 with updated stimuli. The original thirty-six items included some ambiguous cases where inter-rater reliability dropped, so Baron-Cohen and colleagues replaced several photos and adjusted the word lists. The revised version is what most clinicians and researchers use now. The old version still circulates widely on casual websites, which is why scores you find online are not always comparable to each other.
The mechanics of scoring and what the numbers mean
Each correct answer is worth one point. The maximum score is 36. There is no partial credit. A score below 20 in adults is considered clinically significant and typically warrants further assessment with a broader battery. Between 20 and 28 falls in a range that can appear in various populations, including autism spectrum conditions and social anxiety. Above 29 is within the typical range for adult neurotypical participants. One thing most people miss is that the test is not simply a facial recognition task. The stimuli are deliberately selected to contain only the eye region, often with neutral lip areas cropped out or partially obscured. This removes contextual cues from the mouth and cheeks. The difficulty comes from the fact that many of the target mental states look nearly identical in the eyes. "Contemplative" and "concentrating" share the same brow position. "Scheming" and "devious" differ by subtle tension in the upper lid. That is why the four-option format matters, not just the image itself. I keep a reference sheet with the thirty-six items and their correct answers pinned to my desk. Administering this test blind is a mistake if you care about accuracy. Without the key in front of you, you will second-guess ambiguous items and introduce scoring drift. The whole process takes about eight minutes for a trained administrator and two minutes to score once the responses are recorded.
Get the Full Details

Common pitfalls and why people misuse this test
The most frequent error is using an unofficial version with modified options or shuffled answer choices. Some websites present the images with randomized four-word sets that do not match the original stimulus list. The test is standardized, which means changing the distractors invalidates the norm data. A score of 24 on a modified version tells you nothing about the person. Another issue is cultural bias in the stimuli. The original photographs were taken of Caucasian actors in Western settings. Eye reading performance varies across cultural groups in ways the test does not account for. I have seen this play out when colleagues in Japan administered the test to students and interpreted lower scores as clinical deficits rather than recognizing the cross-cultural artifact. The test is not culturally neutral. You should not use this test as a standalone diagnostic tool. It measures one narrow slice of social cognition. Autism diagnosis requires a comprehensive evaluation covering communication, behavior, and developmental history. The Eyes test is a screening aid at best, not a confirmation instrument. Baron-Cohen himself has stated this repeatedly in the literature.
Personality traits also influence performance independent of clinical status. People high in agreeableness tend to select warmer emotional labels, while those higher in narcissism skew toward more dominant or calculating interpretations. If you are comparing scores across individuals without controlling for these variables, your conclusions will be noisy.
Practical administration guidelines
Present the images in a consistent order. Randomizing the order is acceptable, but do not let the participant go back and change answers after seeing later items. The test is designed as a one-pass exercise. Time pressure is not part of the original protocol, but prolonged exposure to individual items can shift responses toward over-analysis. Eight to ten minutes total is typical. Record responses immediately. Write down each chosen word or letter option next to the item number. When scoring, match responses against the official key. Count correct items. Compare against the age-normed ranges if available. The 2013 revision includes updated norms for different age groups, so use the norms that match the version you administered. If someone consistently picks the same wrong answer across multiple items that share a thematic cluster, note that pattern. For example, a participant who confuses "worried" with "concerned" across several items may have a genuine deficit in distinguishing closely related negative states, which is different from random guessing. Document these patterns separately from the total score.

What the test does not tell you
It does not measure empathy as a whole construct. Empathy involves emotional resonance, perspective taking, and behavioral response. This test measures only one cognitive component. A high scorer can still lack compassion. A low scorer can still be emotionally attuned in other contexts. The label "mind reading test" is misleading shorthand for mental state attribution from eyes alone. It also does not predict real-world social success. Correlations between Eyes test scores and everyday social functioning exist but are modest, typically in the 0.25 to 0.40 range depending on the population studied. Real social interaction involves hearing, body language, context, memory, and prior knowledge. Four pictures of eyes cannot capture that complexity. For research purposes, consider pairing this test with the Reading the Mind in the Voice task from the same research group. The combination gives you a clearer picture of whether someone's difficulty is visual, auditory, or domain-general. Using the Eyes test alone leaves you guessing about the source of any deficit you observe.