Why most people get the Standards For Psychological And Educational Testing wrong
The AEA/APA/NCME joint standards aren't a certification you pass and forget. They're a living framework that determines whether your testing program survives a scrutiny audit or gets torn apart because some assumption was never documented. I learned that the hard way during a program review where our validity evidence file couldn't hold up under cross-examination from an outside evaluator. The gap wasn't technical error. It was a documentation habit we didn't realize we lacked. What the standards actually are — The current edition covers 15 chapters across three major domains: tests as measurement instruments, fairness and equity in test use, and the responsibilities of test developers and users. The chapters address validity, reliability, fairness, score interpretation, test construction, administration procedures, and the ethical obligations of anyone who designs, publishes, or uses a test. They don't mandate a single approach. They describe what competent practice looks like and provide the benchmarks against which programs are evaluated. The first thing most people miss is that the standards apply differently depending on whether you are the test developer or the test user. If you build a measure, your obligations center on documentation, validation evidence, and fairness analysis. If you administer an existing test, your obligations shift toward appropriate use, score interpretation, and the administrative conditions that preserve the validity evidence the developer built in.
A practical workflow — Start by mapping every standard to a concrete artifact in your program. I keep a matrix that lists each standard with a column for evidence type, a column for ownership, and a column for the version date of the supporting document. When a new study is run or a revision is made, I update the matrix immediately rather than waiting for a review cycle. This prevents the situation I hit once when we needed to produce evidence of concurrent validity for a client audit and realized three of our source studies were missing their original protocols. We recovered two from the publisher's archives and reconstructed the third from our own lab notes, but it cost nearly two weeks of work that could have been avoided in an hour of routine updates. Validity is not a checkbox — The standards treat validity as the central concern, but validity here means the degree to which evidence supports the intended score interpretations and the decisions made from those scores. It is not a single statistic. It is a cumulative argument built from content relevance, internal structure, relations to other variables, response processes, and consequences of testing. Most programs collapse this into a Cronbach alpha and a factor analysis and call it done. That is insufficient under the standards, especially when fairness and consequential evidence are involved. One counter-intuitive point: high reliability does not guarantee validity. I once saw a well-received aptitude instrument with a reliability coefficient above .90 that produced systematically biased predictions across demographic groups because the scoring model was never examined for differential item functioning. The reliability masked the bias. The standards require both evidence types to be addressed separately.
Fairness and equity in practice — The fairness chapter is where programs tend to be weakest. Fairness is not simply about providing accommodations. It covers accessibility of materials, sampling representativeness, bias in item content, differential item functioning, and the downstream consequences of score-based decisions. A common pitfall is treating fairness as a static property of the test rather than a property of the testing program. A test can be designed with fairness in mind and still produce inequitable outcomes if the administration conditions or scoring models introduce drift. When I run a fairness review, I start with DIF analysis using both categorical and continuous methods, then move to impact analysis across relevant subgroups, and finally examine whether the intended decision thresholds produce disparate outcomes. If any of those steps flag issues, I adjust the scoring model, revise items, or change the decision rules. This is iterative, not one-off. Documentation that holds up — Under the standards, documentation is evidence. It needs to be traceable, versioned, and sufficient for an independent reviewer to reproduce the argument. I recommend maintaining a validation dossier that includes: the test blueprint, item development logs, pilot data, field test results, reliability estimates with confidence intervals, validity evidence organized by the five sources, fairness analyses, score reporting standards, and any decision rules tied to score interpretation. Store it in a system with access controls and version history. PDFs in shared folders are not enough for a serious program.
Get the Full Details

Common failures I see repeatedly First, treating norming as a one-time event. Populations shift. Norms age. If your reference group was collected more than five years ago, you should reassess whether it still represents the current population you serve. Second, ignoring the standards for computerized adaptive testing even when the test is delivered on paper. The standards cover delivery mode, and CAT introduces unique fairness and validity concerns that are not automatically covered by print validations. Third, conflating standard errors of measurement with standard errors of prediction. They are related but not interchangeable, and the standards require the correct metric for the decision context. Where the standards fall short — The standards are descriptive, not prescriptive. They tell you what to consider but rarely specify exact thresholds or mandatory analyses. This leaves room for professional judgment but also creates inconsistency across programs. In practice, reviewers often expect more than the minimum language states, particularly around differential item functioning and consequential evidence. If you only meet the letter of the standards without addressing the spirit, you will be vulnerable during external evaluation.
Another limitation is the slow pace of revision. The field moves faster than the publication cycle for major updates, so new methods like automated scoring, AI-mediated accommodations, and online proctoring often sit in a gray area until guidance catches up. My workaround has been to treat emerging practices as provisional standards and document the rationale explicitly, citing the relevant existing standard and explaining how the new practice satisfies or extends it. Implementation timeline — For an established test, a full standards alignment review typically takes six to twelve months depending on the scope of existing evidence. For a new instrument, plan for eighteen to thirty-six months to build a defensible validity argument that meets the standards across all required sources. Budget time for fairness analysis and impact studies early. These are expensive and difficult to retrofit. If you are starting from scratch, begin with the intended score interpretations and work backward through the standards rather than forward from data collection. The standards are organized to support this deductive approach, and it reduces the risk of gathering evidence that looks good on paper but does not actually support the decisions you intend to make.