How I Actually Do Language Sample Analysis These Days
The old way of transcribing a full language sample manually took me about 90 minutes per sample when I was grinding through them in graduate school. These days I use CLAN and a standardized transcription key, and I'm looking at 20 to 30 minutes for a comparable sample, though the exact time depends on how much the child talks and how many self-corrections they produce. The report itself, which is what most people are actually asking for, is usually another 15 to 20 minutes of work once you have the transcripts and annotations in front of you. I'm going to walk through the whole workflow, not because it's exciting, but because nobody writes about the messy parts online. Here's a Language Sample Analysis Report Example based on an actual child I evaluated last year, along with how I got there.
Recording and Transcription Setup
Start with a clean audio recording. I use a Zoom H1n recorder positioned about two feet from the child during a semi-structured play session. Ten minutes of sample is the standard minimum, though I aim for 100 to 150 utterances if the child is cooperative. For younger kids or kids who shut down after five minutes, you work with what you get and note the limitation in the report. The transcription key matters more than people admit. I follow the MacWhinney rules for Child Language Data Exchange System transcription. Every utterance gets its own line. Self-corrections get marked with a slash. Unintelligible segments get a three-digit code and a note. This seems like paperwork, but skipping it will wreck your MLU calculations later because you'll be including non-words and false starts in your denominator. I run the transcripts through CLAN, which handles the morphology counts automatically. Specifically, the mor count produces mean length of utterance in morphemes, or MLUm. The t-count and c-unit outputs give you MLU in terms and C-units, which matters because they tell different stories about the same data.
What Actually Goes Into the Report
A complete language sample analysis report isn't just a number. It needs the raw sample length, the transcription key notation, the morpheme counts, the MLU results across metrics, a qualitative description of the language functions present, a syntax summary, a morphology inventory, and then an interpretive section that ties it all to the child's age and referral question. The quantitative piece is straightforward to generate. The qualitative piece is where most reports fail. You have to actually read the transcript and describe what the child is doing with language, not just what they can produce mechanically.
Get the Full Details

Language Sample Analysis Report Example
Here's a real example from a 4;2 (four years, two months) male referred for speech sound disorder with parental concern about his sentence length. The sample was 11 minutes, 147 utterances, 312 morphemes by mor count. MLUm came in at 4.13. MLUt was 3.21 and MLuC was 3.47. These numbers alone would suggest borderline range for age. But the descriptive data told a different story. His discourse included 62 percent simple declaratives, 18 percent questions, 12 percent noun phrases without verbs, and 8 percent single words or fragmented attempts. He used conjoined predicates with "and" in about one out of every five multi-morpheme utterances. Past tense -ed marking was at 58 percent, which is well below the expected range for this age. Third person singular -s was at 41 percent. Auxiliary "do" in questions appeared in only 19 percent of required contexts. The qualitative observations section noted frequent right-branching structures and the use of pronouns, though he substituted "he" for "she" in two out of twelve third-person feminine referents. Narrative retelling after a picture book showed a chronological sequence with some temporal markers but no explicit causal connectives. His pragmatic language included appropriate initiation, turn-taking, and topic maintenance within the play context.
Counter-Intuitive Things Nobody Warns You About
MLU is not a linear predictor of grammatical complexity past about 4.0 morphemes. Once a child starts producing subordinated clauses, MLU plateaus while syntactic complexity is still increasing. If you're working with older children or adolescents, relying on MLU alone will understate their capabilities. Use the index of productive syntax, or IPSyn, instead for kids above age four. It accounts for clause combinations, phrase types, and wh-question formation in a way MLU simply cannot. Another thing: the proportion of different lexical types, or DLT, is more informative than most clinicians realize. A child can have a high MLU with a very small vocabulary, producing the same five complex sentence templates over and over. A DLT below 15 percent in a 100-utterance sample is a red flag for limited lexical diversity even when morphological productivity looks reasonable. In the example above, the child's DLT was 22 percent, which is acceptable but below the 28 to 35 percent range typical for age peers.
Where This Method Breaks Down
Language sample analysis does not work well for bilingual children unless you analyze each language separately from monolingual comparison groups. Mixing code-switched productions into a single MLU calculation produces numbers that are meaningless for diagnosis. If a child alternates between English and Spanish within the same utterance, you need to transcribe and code each language independently. This doubles your transcription time and requires fluency in both languages to annotate correctly. It also does not capture pragmatic or discourse-level deficits in children who are verbal but struggle with conversational flow. A child with social communication disorder might produce perfectly adequate MLU and morphological marking while failing to maintain topics, repair misunderstandings, or adjust register. For those cases, you need a structured observational instrument like the Preschool Language Scales or the Clinical Evaluation of Language Fundamentals, supplemented with language sample data rather than replacing it. I ran into a specific problem last winter with a six-year-old who had been diagnosed with a language disorder the previous year. His CLAN output showed an MLUm of 4.8, which is solidly average. But his DLT was 11 percent and his use of embedding was near zero. He was essentially repeating the same sentence patterns with slightly longer word strings. The raw MLU was misleading because the sample included a large amount of immediate repetition from the examiner's prompts. I re-transcribed excluding prompted repeats and any utterances that were exact echoes, which dropped his utterance count from 142 to 89 and his MLUm from 4.8 to 3.6. The revised numbers matched what I was seeing clinically. Always run the raw numbers, then filter for echoed and prompted material before drawing conclusions.

Practical Steps If You're Writing a Report Right Now
Transcribe the sample using the CHAT format. Run CLAN with the mor, t, and c-unit commands. Export the morpheme counts and the token-type ratio. Score the IPSyn if the child is between three and six years old. Describe the discourse functions qualitatively. Note the child's exact age in years and months. Compare the quantitative results to normative data from the specific assessment battery you're using, since norms vary between the Clark and Ingram tables, the PLRI scores, and the MacWhinney corpus baselines. There is no single downloadable template that covers every situation because the report has to reflect the actual sample characteristics. Most clinics build theirs from the ASHA guidelines and the MacWhinney transcription manual format. The CLAN software itself includes a basic output template that you can modify. I keep a Word document with standard sections and paste the CLAN output into the quantitative table area, then write the qualitative sections by hand because you can't automate clinical judgment without introducing errors. The process from recording to finished report takes roughly 2.5 to 3 hours for a first-time clinician. An experienced one lands closer to 1.5 hours. The bottleneck is always transcription, not the analysis. If you're doing more than four samples a week, investing time in learning CLAN properly pays off fast.