The Problem With Bilingual Assessment Tools
Most speech-language pathologists who work with bilingual clients hit a wall when it comes to standardized testing. You pull up a normative assessment tool, look at the demographic tables, and realize the norms were built on monolingual English speakers. Using those scores for a Spanish-English bilingual child or a Mandarin-English adult isn't just inaccurate, it's technically unethical. The field has spent the last fifteen years grappling with this, and the tools that have emerged are mixed in quality. I've been doing bilingual assessments for about a decade. The tools available now are better than what we had in 2015, but they're far from a clean solution. Here's what actually works in practice, what doesn't, and where the real gaps are.
What Bilingual Speech And Language Assessment Tools Actually Are
These aren't single products. They're a category of instruments, frameworks, and databases designed to evaluate speech and language in people who use two or more languages. The category includes standardized tests with bilingual norms, dynamic assessment protocols, language sampling software, and decision-making frameworks that help you interpret results across languages. The field has moved away from the old approach of simply translating a monolingual test and administering it. That didn't work because translation doesn't equal equivalence. A test translated into Spanish still measures English-language-influenced constructs if the items, norms, and administration procedures aren't built for bilingual speakers from the ground up. What actually exists now falls into a few buckets. There are tools built specifically for bilingual populations with separate or combined norms. There are research-grade databases like the Bilingual Atlas of Language Disorders and the Speech Bilingual Assessment Framework that don't give you a score but give you the information to make a clinical judgment. And there are informal measures, observation protocols, and dynamic assessment procedures that bypass the norm-referenced problem entirely.
How to Choose the Right Tool for Your Caseload
This is where most people get it wrong. They look for a single test that covers everything, and they end up using something that doesn't fit their client's profile. The right approach starts with the client's linguistic background, not the catalog of available tests. First, map the client's language exposure. Age of acquisition matters. A child who started English at age four has a different disorder-to-difference profile than a child who has been exposed to both languages since birth. Proficiency in each language matters. A client who is dominant in their heritage language will score differently on an English-only measure regardless of whether they have a disorder. I had a client recently, a ten-year-old girl raised in a Korean-English household, who scored in the clinically significant range on an English expressive vocabulary test. Her Korean was never assessed. When I ran a targeted language sample in Korean, her MLU and morphological complexity were age-appropriate. The English deficit was real but narrower than the initial score suggested. We ended up with a much more accurate picture and a better intervention plan. Second, check the norming population. This sounds obvious but people skip it constantly. Look at the standardization sample. Was it truly bilingual? How many languages were represented? What were the socioeconomic demographics? Many so-called bilingual tools were normed on small samples that don't reflect the diversity of the populations you actually see in practice. A tool normed on Mexican-American children in Texas may not generalize to a Colombian-American child in New Jersey, even though both are Spanish-English bilinguals.
Get the Full Details

Third, consider the construct being measured. Not all assessments measure the same thing. Some assess phonology. Some assess morphology. Some assess narrative discourse. A child might have a phonological process error that is language-specific, meaning it appears in one language but not the other. If you're only testing English phonology, you miss the full picture. The same goes for syntax and semantics. I recommend at minimum two data sources across both languages before drawing any conclusions about the presence or absence of a language disorder.
Tools That Are Actually Worth Using
I'm going to be direct about what has held up in my practice and what hasn't. Some of this is subjective because your caseload determines what works, but here's the reality of the current landscape. The Clinical Evaluation of Language Fundamentals--Fourth Edition (CELF-4) and CELF-5 have a Spanish version. It's not a translation, it was developed alongside the English version. The norms include Spanish-speaking monolinguals but still don't have robust bilingual norms. It's useful as a screening tool but insufficient as a standalone assessment for bilingual clients. The Preschool Language Scales--Fifth Edition (PLS-5) and Expressive One-Word Vocabulary Test--Fourth Edition (EOWPVT-4) also have Spanish versions with similar caveats. Use them as part of a battery, not as the battery.
The Bilingual English-Spanish Assessment (BESA) is one of the few tools built specifically for bilingual children ages three to eight. It uses obviating procedures to reduce cultural and linguistic bias. It's not perfect. The norms are limited to certain age bands and the sample wasn't enormous. But it's one of the few options that was designed from scratch for this population rather than adapted after the fact. For school-age and adolescent clients, the Test of Early Language Development--Fourth Edition (TELD-4) has a Spanish form and covers receptive and expressive language across a wider age range. Still not bilingual-normed, but broader in coverage. Language sampling remains one of the most underutilized tools in bilingual assessment. I know that sounds controversial to people who prefer standardized scores, but I've found it to be more informative than most norm-referenced tests for bilingual speakers. A 100-utterance language sample in each language gives you real data on morphological complexity, syntactic variety, and pragmatic competence. You can analyze it using MLU, type-token ratio, and clause complexity measures. There's no norm to misinterpret. There's just data from the client's actual language use.

Dynamic assessment is another approach that sidesteps the norming problem entirely. You test, then teach, then retest. The amount of learning and the rate of transfer tell you whether a low score reflects a language difference or a language disorder. It takes longer than a standard test, roughly 45 minutes to an hour per language instead of 20, but the diagnostic accuracy is substantially better for bilingual populations.
The Practical Workflow I Use
Here's how I actually structure a bilingual assessment, not the ideal version but the version that fits into a real schedule. I start with a case history interview that covers age of onset, language exposure patterns, family history, educational history, and any prior assessments. This takes 20 to 30 minutes and is essential. A lot of the interpretation depends on these details. Then I administer a bilingual screening measure if one is appropriate for the client's age. This identifies areas of concern across both languages. I typically use the BESA for young children or a combination of the PLS-5 and EOWPVT in both languages for older clients.
Next I collect language samples in both languages. This is the step people skip because it takes time, but it's the step that changes the diagnosis most often. I record a conversational sample and a narrative sample if the client is old enough. Each language sample is about 100 utterances, which takes roughly 15 to 20 minutes per language depending on the client's rate of production. After that, I do targeted standardized subtests based on the screening results. If the screening flagged phonology, I add a phonological process analysis. If it flagged vocabulary, I add a semantic analysis. I don't administer whole tests. I pick the subtests that address the identified concerns in each language. Finally, I integrate the data using a decision-making framework. The American Speech-Language-Hearing Association (ASHA) has guidelines for assessing bilingual individuals that I reference constantly. The core principle is that no single score should determine a diagnosis. You need convergence across multiple data sources, languages, and assessment types.

What Doesn't Work
I want to be blunt about a few things that sound good but fall apart in practice. Translation-only tools are the biggest problem. Any test that was created in English and then translated into another language without going through a full psychometric validation process in the target language is not reliable for diagnosis. Item difficulty changes. Cultural references don't map. Norms are invalid. Don't use them. Single-language assessment followed by interpretation is common but insufficient. A report that says "the child scored below average in English" without addressing the other language is not a bilingual assessment. It's an English assessment with a bilingual label attached. That's how misdiagnosis happens, and it happens at alarming rates in schools.
Overreliance on standardized scores from monolingual norms is the third major pitfall. Yes, some clinicians use English-only norms with a note about the limitation. The note doesn't fix the problem. A score that's two standard deviations below a monolingual mean is not the same thing as a clinically significant deficit in a bilingual context. The gap between the two interpretations is where delays and inappropriate placements come from.
Where the Field Is Going
The quality of bilingual assessment tools is improving, but slowly. The Core Language Scales and other newer instruments are beginning to include more diverse samples, and research on bilingual language development has advanced significantly, which informs tool development. There are also emerging digital tools and apps that support language sampling and analysis, which reduces the time burden that has historically made thorough bilingual assessment impractical. But the biggest gap remains in the middle age groups. There are decent tools for preschoolers, reasonable tools for school-age children, and very few options for adolescents and adults. Bilingual adults who present with language difficulties often fall through the cracks because the available instruments were designed for developmental populations. This is a real problem in clinical and educational settings. My recommendation for anyone working with bilingual clients is to build a toolkit rather than rely on a single product. Combine a purpose-built bilingual measure with language sampling in both languages and dynamic assessment procedures. It takes more time, maybe 90 minutes to two hours total instead of a single 40-minute test session, but the diagnostic accuracy is noticeably better and the reports you write are defensible.

The tools exist. They're not perfect. But they're closer to what we needed five years ago, and the ones coming out now are addressing the gaps that used to be unfillable. The main limitation is access and training, not the tools themselves. Most of these instruments require certification or purchase, and proper interpretation requires understanding of bilingual language development that isn't covered in standard graduate programs. If your organization hasn't invested in professional development around bilingual assessment, that's probably the first place to look.