Scoring the GFTA-2 Without Losing Your Mind
The Goldman Fristoe Scoring Manual is basically the rulebook for the Goldman-Fristoe Test of Articulation, second edition. It tells you how to administer, score, and interpret results. Most people working in speech-language pathology already know that much. What the manual doesn't make clear, at least not in any obvious way, is how some of the edge-case decisions actually play out in real clinic time. The core of the scoring process is straightforward. You play the audio stimuli through the provided recording. The child responds. You mark whether each target sound is produced correctly, with an error type (omission, substitution, distortion, etc.), or if it's unintelligible. Then you compute three scores: the Articulation Index, the Phonological Process score, and the Word Approximation score. Those feed into a percentile rank and a standard score against the normative sample.
Where the Goldman Fristoe Scoring Manual Gets Messy
I'll be honest about one thing that caught me off guard when I first started using this tool consistently. The manual describes distortion as a sound that is audible but not clearly identified as one of the target phonemes. That sounds simple. It isn't. I was scoring a nine-year-old with a fairly typical articulation profile and came across a /r/ that was clearly distorted. Not an approximation. A full-on w-coloring with some secondary vocalic quality. The child also produced a few correct /r/s interspersed. The manual says to score each occurrence individually, which is what I did. But then the norm tables and the clinical interpretation get fuzzy when a child has more distortions than substitutions or omissions. Distortions don't count the same way toward the overall Articulation Index in the way you'd expect, and the scoring sheet doesn't really warn you about that. My workaround was to double-check my error tallies against the raw response protocol before finalizing anything. I also cross-referenced the child's errors with the Phonological Process Analysis table, because distortions can masquerade as process errors depending on how they're classified. One thing that tripped me up initially: the manual groups certain distortions under "liquids" for the process analysis, but if the distortion is truly atypical, it might not map cleanly onto the expected processes. I learned to flag those separately in my notes rather than trying to force them into a neat category. Here's another thing that isn't obvious from reading the manual cover to cover. The reliability of the GFTA-2 drops noticeably for children under three years old. The normative sample skews toward preschool and school-age populations. If you're scoring a two-year-old and the result comes back as "within normal limits," that doesn't mean much. The test simply wasn't standardized on that age group with enough precision to make a confident claim. I've seen this result misinterpreted multiple times in evaluation reports, usually because the clinician stopped reading at the standard score and didn't note the limitation in the interpretation section. The Goldman Fristoe Scoring Manual does mention the age range, but the practical implication—that you should supplement with a different assessment for very young children—doesn't get emphasized enough in the scoring narrative.
A counter-intuitive point that saves people a lot of headache: the Word Approximation score is actually one of the more clinically useful metrics in the package, even though most practitioners focus almost entirely on the Articulation Index. The approximations capture sounds that are close to the target but not quite there, which tells you something about the child's articulatory trajectory. A kid might have a low AI because they're distorting most of their /s/ patterns, but if their approximation score is high, it suggests they're on the right track and just need fine-tuning. I use that score to frame prognosis and carry-over expectations more than I use the AI alone. There's also a practical issue with the stimulus recording itself. The audio CD or digital file uses a consistent speaking rate and level, but in noisy clinic environments, especially with younger or inattentive children, the stimuli can get lost. I stopped trying to score on the fly and started using a tablet to pause and replay individual items. It adds maybe thirty seconds per subtest, but it eliminates the guesswork about whether the child actually heard the word clearly. That guesswork is where scoring drift happens, and drift is what turns a reliable measure into something you can't stand behind in a hearing. If you're looking for the actual Goldman Fristoe Scoring Manual, it's published by Pearson Clinical and isn't freely available online. You can order it directly from Pearson or through authorized educational and clinical suppliers. Some university speech-language pathology programs include it in their test kits. The GFTA-2 kit itself comes with the scoring forms, the stimulus recording, and the manual. Make sure you get the GFTA-2 version if you're working with current norms, because the original GFTA uses outdated standardization data and the scoring procedures differ slightly between editions.
Get the Full Details

The main limitation everyone glosses over is that this test measures product, not process. It tells you whether a sound is correct or incorrect in a word, but it doesn't reliably distinguish between a developmental delay and a true phonological disorder. A child with a phonological process like final consonant deletion will score poorly on the AI, but the pattern of errors across items is what really defines the diagnosis. You need the phonological process analysis section of the manual for that, and you need to apply it carefully. Otherwise you're just reporting a number without the clinical context that makes the number meaningful. I also recommend pairing the GFTA-2 with a language sample or a broader phonological assessment like the PPVT or a dynamic assessment approach. The Goldman Fristoe Scoring Manual provides solid foundational data, but no single test captures the full picture. The scores are reproducible when scored correctly, and the norms are solid for the intended age range, but relying on them alone is where people get tripped up. The manual is a tool, not a conclusion.