Understanding the Owls Ii Scoring Manual
The Owls Ii Scoring Manual is the companion document for the OWLs II (Observing Worker–Learner Interactions) instrument, which is used primarily in health professions education to rate the quality of teaching encounters between clinicians and trainees. It isn't a standalone test. It's a rubric and a set of behavioral anchors that define what each score level actually looks like in practice. Without it, you're just watching a conversation and guessing whether it was good or bad. I've used this for clinical teaching evaluations across internal medicine and surgery rotations. The manual itself is published by the Association for Hospital Medical Education and affiliated groups. You can find it through the AME website or the original OWLs project repository. It's free to download once you've completed the brief registration process. The PDF runs roughly 60 pages including the scoring anchors, the behavioral descriptors, and the training modules.
Owls Ii Scoring Manual: Where to Get It
The primary source is the OWLs project page at awlnetwork.org or the AME publications portal. The file is labeled "OWLs II Scoring Manual" and typically comes as a single PDF. Some institutions also maintain a mirrored copy on their medical education intranet, but those are often outdated. Always check the revision date on the document. The current version includes updates from the 2023 revision cycle, and older versions have slightly different anchor wording that can throw your inter-rater reliability numbers off if you mix them. Here's the thing most people miss when they first pick up the manual. OWLs II doesn't score the learner. It scores the teaching episode. The instrument focuses on the interaction between the clinical preceptor and the trainee, not on the trainee's performance as a clinician. That distinction matters because if you walk into an evaluation thinking you're assessing how well the resident managed the patient, you'll apply the rubric wrong and your scores will be meaningless. The scoring covers six domains. Each domain has specific behavioral descriptors tied to a Likert scale. The domains are:
Activating prior knowledge — whether the teacher checks what the learner already knows before building on it. Demystifying performance — how the teacher models clinical reasoning out loud instead of expecting the learner to figure it out by osmosis. Structuring feedback — whether feedback is given in a way the learner can actually use, with clear action steps.
Get the Full Details

Demonstrating professionalism — modeling appropriate professional behavior during the encounter. Facilitating autonomy — giving the learner an appropriate amount of independence for their level of training. Quality of patient care — ensuring the teaching moment doesn't come at the expense of the patient.
Each domain is scored independently. The manual provides anchor examples for scores ranging from 1 to 5 within each domain. A score of 3 is defined as "adequate" but not exemplary. Most faculty members naturally cluster around a 3 across all domains. That's normal. It's also why valid inter-rater calibration is essential before you start collecting data for any formal evaluation program.
Training and Calibration
The scoring manual includes a set of training videos with demonstrated encounters. You should watch them. I've seen programs skip this step and go straight to live scoring, which produces reliability coefficients that make no sense. After watching the training modules, you score the same video independently, then compare your scores with another rater. If your scores differ by more than one point on any domain, you review the anchors together until you converge. This calibration step typically takes about 90 minutes and needs to be repeated at least once per academic year if you're running a longitudinal program. One practical note: the video examples in the manual use simulated encounters with standardized patients and actors. Real clinical settings are messier. A teaching moment might be interrupted by a code blue, or the attending might be clearly rushed and doing their best in difficult conditions. The anchors assume an ideal environment. When you're scoring real encounters, factor in contextual constraints. The manual acknowledges this briefly in the commentary sections, but it doesn't give you a concrete adjustment formula. You have to apply judgment here.

A Specific Problem I Ran Into
Last year I was scoring a surgery rotation evaluation where the attending was technically an excellent clinician but had zero training in educational methodology. The learner was a senior resident who was highly competent. On the "facilitating autonomy" domain, I initially scored a 2 because the attending was making almost all the decisions and barely involving the resident in procedural planning. But when I reviewed the encounter notes, I realized the resident had explicitly asked to observe only and had requested minimal involvement due to being prepped for an boards simulation that week. The attending was respecting that request. My initial score was wrong because I hadn't accounted for the learner's stated preference. I ended up adjusting to a 4 with a detailed note in the evaluation file explaining the context. This is the kind of edge case that doesn't appear in the manual and only shows up after you've done enough real-world scoring to recognize the pattern. The biggest mistake people make is averaging the domain scores and treating the result as a single "teaching quality" number. The manual explicitly warns against this. Each domain measures something distinct, and collapsing them obscures important information. A teacher might score a 5 on professionalism and structuring feedback but a 2 on activating prior knowledge. That's a meaningful profile, not a problem to be smoothed over. Another common issue is conflating the quality of the clinical content with the quality of the teaching. If an attending gives excellent clinical advice but does so without checking the learner's prior knowledge or providing structured feedback, the teaching score should reflect that gap. The domain of quality of patient care is where people get tripped up most often. A high patient care score does not automatically elevate other domains. Keep them separate.
Limitations of the Instrument
OWLs II was designed for clinical teaching encounters lasting 15 to 45 minutes. Shorter encounters under 10 minutes are difficult to score reliably because there simply isn't enough interaction to observe across all six domains. Similarly, purely virtual or telehealth encounters present challenges that the current manual doesn't fully address. The behavioral anchors were written with in-person encounters in mind, and while the domains still apply, raters need to be aware that certain behaviors manifest differently on screen. Inter-rater reliability is another documented limitation. Even with proper calibration, typical reliability coefficients for OWLs II fall in the 0.65 to 0.75 range, which is acceptable for formative feedback but insufficient for high-stakes faculty promotion decisions. If your institution is using this for tenure or promotion, you should supplement it with other assessment methods. The manual itself acknowledges this limitation in the scoring validity section. For programs looking for a more quantitative alternative for large-scale studies, the Mini-CEX or the Clinical Evaluation Tool (CET) might be more appropriate depending on your goals. OWLs II excels at capturing the qualitative dimension of clinical teaching interactions, but it wasn't designed as a replacement for comprehensive clinical performance assessments.
Practical Tips for Implementation
If you're setting up an OWLs II evaluation program at your institution, budget roughly three hours for initial rater training including the calibration exercise. Plan for biannual re-calibration sessions. Keep the scoring forms digitized rather than paper-based to reduce data entry errors. The scoring manual recommends using a mobile-friendly tablet interface for real-time scoring during live encounters. Paper forms tend to get lost or filled out incorrectly in busy clinical environments. The manual includes a scoring worksheet template in the appendix. Use it. Don't improvise your own form. Deviations from the standard worksheet make it difficult to compare scores across evaluators and over time. I've seen programs create custom spreadsheet-based scoring systems that looked clever but introduced scoring inconsistencies that took months to trace back and fix. One final note on timing. Live scoring during an actual clinical encounter is the gold standard but not always feasible. Record-and-review scoring is a reasonable alternative and often more practical for busy departments. Just be aware that the manual's training data is built around live observation, and some raters report that reviewing recordings feels less intuitive because you can't see the full context of nonverbal cues that matter in the moment.
