Working Through the ADI-R in Practice
The Autism Diagnostic Interview-Revised is one of those instruments everyone in developmental diagnostics has to know, but few people actually enjoy administering. It's a semi-structured, caregiver-based interview that takes about two to three hours to complete properly. You're asking parents or guardians detailed questions about early development and current behavior across three domains: social interaction, communication, and restricted/repetitive behaviors. The scoring is algorithmic—there's a cutoff for each domain and a combined total that you compare against established thresholds to determine whether the profile is consistent with an autism spectrum diagnosis. The manual and interview guide are published by Western Psychological Services (now part of CAPS). You can order them directly from their website or through major academic suppliers. The scoring forms come with it, but you'll want to invest in the training workshop if your agency doesn't already provide one. Just buying the book and winging it is not going to work well—the scoring decisions require judgment calls that aren't obvious from reading the manual alone. There are also third-party platforms like Pearson ClinGen or Q-global that offer digital versions, though some clinics avoid those for privacy reasons. I've seen people try to find free PDFs circulating online. Don't bother. The instrument is copyrighted, and more importantly, using an unlicensed copy means you're working from potentially outdated or incomplete materials. The reliability of your results depends on having the current version.
What the Interview Actually Looks Like in the Room
You sit down with the primary caregiver—usually a parent—and work through around 100 items. Each item asks about the child's typical behavior at a specific age window, mostly focusing on the 4-to-6-year period, which matters because early memories are more reliable for onset criteria. You code responses as 0, 1, 2, or currently atypical, and then there are special codes like "unknown" or "not applicable." The algorithm then tallies scores into the three domains and gives you a total. The domains break down like this: social interaction gets roughly 40 items, communication gets around 35, and repetitive/stereotyped behaviors get about 25. The standard cutoff scores are 10 for social, 8 for communication, 3 for repetitive behaviors, and 20 total. Those are the numbers you're aiming at, but the real work happens in the gray areas between them. I spent years administering this and the one thing that never gets mentioned in the manual is how much the quality of your rapport with the caregiver determines whether you get accurate information. A guarded parent who minimizes difficulties will give you scores that look almost normal. A parent who is highly anxious may over-report. I had a case where the mother described a child who clearly met criteria, but every item scored zero because she genuinely believed her son's lack of eye contact and single-word speech was just shyness. It took about twenty minutes of gentle, specific questioning—asking her to describe a typical morning routine, what happened when we walked in the door, how he reacted when I offered him a toy—to get her to recount behaviors she'd been filtering through a different lens the whole time.
Common Scoring Pitfalls
The first thing people get wrong is treating "development at any age" questions as current behavior questions. If a child lost a skill, that's important, and you need to capture the duration. The manual tells you this, but raters skip it under time pressure. Another trap is the "currently atypical" designation—use it sparingly. It's meant for behaviors that are present now but weren't clearly present in early development. If you mark something as currently atypical for everything, you inflate the algorithm scores without actually meeting the developmental onset requirement that DSM-5 insists on. Item 31 about pointing is another place where people go too easy. A declarative point—showing something to share interest—is qualitatively different from an instrumental point, which is just requesting. The scoring hinges on that distinction. I once had a rater score a child as intact on this item because the kid pointed at objects all day to ask for them. He never pointed to share. That child's algorithm score dropped by three points once I re-scored that item correctly, which was the difference between borderline and clinically significant in the social domain. There's also a known problem with item wording in certain cultural contexts. Some questions assume a particular style of play or social expectation that isn't universal. A child raised in a multilingual household might have a different communication pattern that looks like a deficit on paper but reflects adaptation rather than impairment. I ran into this with a family where the 4-year-old was actively learning English as a second language while the home language was Vietnamese. The communication domain scored borderline on the surface, but when I dug into the items about functional communication and alternative methods of expressing needs, the picture changed. The tool has limitations here, and no amount of training eliminates that problem entirely.
Get the Full Details

What It Does and Doesn't Tell You
The ADI-R is not a standalone diagnostic tool. It's part of a comprehensive evaluation. The DSM-5 requires that symptoms be present in early development and cause clinically significant impairment, and this interview is one of the best instruments we have for establishing that developmental history. But it doesn't measure current adaptive functioning, it doesn't assess cognitive ability, and it doesn't rule out other conditions. You need the ADOS-2 on the child directly, along with hearing screening, cognitive testing, and a thorough medical history to build a complete picture. People also over-rely on the algorithm cutoff as if it's a pass-or-fail test. It isn't. A score just below the cutoff doesn't mean the child doesn't have autism. The ADI-R is sensitive to the severity of symptoms in the broadest sense, but individual variation means some children with clear autism profiles score in the subclinical range. Conversely, some children with ADHD, language disorders, or reactive attachment patterns can score above cutoff. The algorithm is a heuristic, not a verdict. I stopped trying to use it as a sorting machine around year five of doing this work. It's better used as a structured way to make sure you're asking the right questions about developmental history rather than relying on parental recall alone. The caregiver fills in gaps in their own memory by answering systematically. That's its actual value—more than the final number it produces.
Practical Tips That Aren't in the Manual
Send the preliminary questions to parents before the session. Not all of them, just a preview so they have time to think about their child's earliest behaviors. Parents consistently report that items about age of first words or first gestures trigger memories they wouldn't have accessed on the spot. I also schedule the interview for the second half of a testing day, after the ADOS-2 observation, because knowing what you saw in the room changes how you ask the follow-up questions. The interview and the observation inform each other, and doing them in the wrong order biases your questioning. If a child is nonverbal or minimally verbal, don't skip the language items and try to compensate elsewhere. Score them accurately based on what the caregiver reports about the child's expressive and receptive abilities at each age. The algorithm handles nonverbal profiles fine—it's designed for that. The problem is raters who quietly change the scoring standard for nonverbal children, which introduces inconsistency that invalidates the results. Documentation matters more than people think. Write down the caregiver's exact words for ambiguous items. Two weeks later, when you're scoring and you can't remember whether a description referred to a single gesture or a consistent pattern, your notes are the only thing that will save you from reconstructing the answer to fit a preferred score. I've lost count of the number of times this has prevented me from inflating a score retrospectively.
When the ADI-R Falls Short
It's a long interview. For families dealing with multiple children, work schedules, transportation issues, or language barriers, completing a full two-to-three-hour session is often impractical. There's a shorter screening tool called the M-CHAT that you can use first, but it only catches some cases and it's not a replacement. If you're working in an underserved area where follow-up is unreliable, you might consider whether the investment in a full ADI-R is going to produce a usable result or just an incomplete form sitting in a file. For adolescents and adults, the ADI-R was never designed for that population. The developmental history it captures relies on parental recall, and parental recall of a 10-year-old's behavior from twenty years earlier is unreliable. Some clinicians adapt it loosely for older individuals, but the validity drops significantly. The ADI-R-R (the revision for older individuals) is still in development and isn't widely available yet. Until then, if you're evaluating an adolescent or adult, you're working with a tool that has known limitations for that age group regardless of how experienced you are. The scoring itself takes time. A competent rater who knows the instrument well finishes the interview in two to three hours, then needs another hour or so to code and compute the algorithm. A novice will take considerably longer and make more errors. If your clinic has a backlog of evaluations and the only person trained in ADI-R scoring is one staff member, that bottleneck will delay reports for everyone. Budget accordingly.
