Working with Upper Respiratory Tract Labeling in Medical Datasets
I run a small bioinformatics lab and we spend more time than I'd like admit wrestling with anatomy labels in NLP pipelines. A couple years ago I built a system that extracted mentions from laryngoscopy procedure notes and tagged them against the upper respiratory tract — nasal cavity, pharynx, larynx, trachea. That's what I mean by "upper respiratory tract labeled." It sounds obvious until you hit the edge cases and realize most people don't have a clean ontology to point at. When I say the upper respiratory tract is labeled, I'm talking about discrete anatomical regions getting consistent terms across a corpus of clinical text or radiology reports. The tract runs from the nares down through the nasal cavity, paranasal sinuses, pharynx (nasopharynx, oropharynx, laryngopharynx), and then the larynx. After that you're in the lower tract. Splitting that line cleanly is harder than it reads on paper. A labeled dataset for this region usually has entity tags like NASAL_CAVITY, NASOPHARYNX, OROPHARYNX, LARYNX, and TRACHEA. Some annotators lump trachea into upper respiratory; others don't. That inconsistency alone will wreck an evaluation if you're not paying attention.
My Process for Building an Upper Respiratory Tract Labeled Set
I usually start with a small seed annotation using a standard like RadLex or SNOMED CT. Here's how the work actually goes: Step 1 — Pick a gold standard for the ontology. RadLex maps "Upper Respiratory Tract" to a specific concept ID. SNOMED CT has separate concept IDs for nasopharynx, oropharynx, and larynx. UMLS CUIs work too. Don't try to invent your own labels unless you enjoy explaining them to reviewers later. Step 2 — Define the boundary rules up front. This is where people slip. Is "larynx" labeled when a report says "vocal cord" or only when it says "larynx"? Is "epiglottis" upper or lower? My rule: epiglottis and vocal cords are larynx. Pyriform sinus is hypopharynx and counts as upper. Anything past the cricoid cartilage is lower. Write these down before annotators see the data.
Step 3 — Annotate with a pragmatic tool. I use BRAT or Doccano. For clinical text at scale, BRAT's ASCII format is actually easier to parse programmatically than the JSON Doccano output. One person annotates, a second person spots-checks 20 percent. Disagreements go to a third reader. Kappa around 0.75 is acceptable for this kind of domain. Step 4 — Link each mention to a ontology term. Raw character spans aren't enough. Every labeled span needs a concept ID so downstream models can generalize. A mention of "windpipe" should map to TRACHEA, not sit as a free-text synonym.
Get the Full Details

What I Wish I'd Known Before Starting
Here's the thing nobody warns you about: clinical shorthand obliterates anatomy boundaries. A sentence like "posterior rhinoscopy showed turbinate hypertrophy and mild pharyngeal erythema" has three upper respiratory mentions in eight words, but the annotator has to decide whether "turbinate" is nasal cavity or if it needs its own sub-region label. I started a sub-label for turbinates because the ENT docs kept requesting it separately. You make that call during the protocol, not after ten hours of annotation. Another trap: abbreviations. "LPR" means laryngopharyngeal reflux to a gastroenterologist and laryngoscope to an ENT. The label isn't wrong — the surrounding text is ambiguous. I added a rule that any abbreviation without an explicit anatomical noun nearby gets flagged as UNCERTAIN_UPPER and sent to a senior reader. It slowed annotation by maybe 12 percent but cut my inter-annotator disagreement in half.
Common Pitfalls and Where the Method Breaks
Label quality degrades fast when the source text is dictated. Spoken dictation contains false starts, self-corrections, and filler that doesn't appear in typed reports. "The, uh, the arytenoid area, I mean the supraglottis, looked — no, both looked normal." Your pipeline will label both "arytenoid area" and "supraglottis" unless you pre-process for disfluencies. I wrote a lightweight regex pass that removes filled pauses and marks self-corrections before annotation starts. It's not perfect but it stops duplicate labels on the same structure. The bigger failure mode is domain drift. A dataset labeled on outpatient clinic notes won't transfer well to ICU ventilator records, even though both reference the upper airway. The vocabulary is different enough that model performance drops roughly 18 to 25 percent on out-of-domain text without fine-tuning. If you're building a general-purpose extractor, include both types in the training set from the start. And don't try to build a labeled set larger than about 2,000 documents without funding a second annotator and a formal review cycle. My first 600-document set had a systematic error where "tonsil" was inconsistently tagged — sometimes ORPHARYNX, sometimes a separate label, sometimes unlabeled. I caught it three months in during a model evaluation. Fixing it required re-labeling roughly a third of the set. Lesson learned: do a midpoint audit at 500 documents minimum.
Practical Resources
If you need an existing reference, the RadLex browser at radlex.org is the quickest way to verify concept IDs for upper respiratory anatomy. For NLP-specific work, the i2b2/UMN challenge sets from 2008 through 2014 contain de-identified clinical text with entity annotations, though they focus broadly on locations rather than isolating the upper tract. The CLEF eHealth lab has also released some usable corpora. There isn't a single download link for a clean "upper respiratory tract labeled" dataset because the boundary decisions depend on your downstream task. A model that extracts diagnoses needs different granularity than one that parses surgical reports. Define what you're building for before you start labeling.

Bottom Line
Upper respiratory tract labeling is straightforward in theory — pick an ontology, draw boundaries, annotate. It's messy in practice because clinical text is messy. The things that matter most are writing the boundary rules before you annotate, auditing mid-project, and accepting that no labeled set generalizes perfectly across documentation styles. If you're just starting out, build a 500-document pilot, get kappa above 0.7, and only then scale up.