How I Label Respiratory Imaging Data For Training Models

Last year my team needed a labeled dataset for upper respiratory anatomy segmentation. We were training a model to identify structures in CT scans and nasal endoscopy footage. The whole process took about three weeks, and honestly most of the time was spent fixing bad labels from the junior annotators. I want to walk through what we actually did, the mistakes I saw repeatedly, and where the whole approach breaks down. This isn't just drawing boxes around things. You're creating pixel-level or region-level ground truth for structures like the nasal cavity, paranasal sinuses, pharynx, larynx, trachea, and upper bronchi. The anatomical boundaries are messy by nature. The nasal turbinates blur into surrounding mucosa on lower-resolution scans. The pharynx changes shape with each breath. These aren't static objects you can neatly contour. Most people entering this space treat it like a straightforward segmentation task. It isn't. The biggest factor you need to account for is inter-observer variability. Two trained radiologists will draw different boundaries on the same slice. Your label protocol has to capture that reality rather than pretending a single ground truth exists.

When I started, I made the mistake of using standard polygon tools for everything. Nasal septum labels came out jagged and inconsistent. Switching to cubic spline interpolation with control points locked to anatomical landmarks cut our labeling time roughly in half and improved inter-rater agreement from about 0.61 to 0.78 on Dice scores. That's a real number from our validation set, not a theoretical improvement.

The Labeling Pipeline I Actually Use

Here's the workflow we settled on after burning through two failed attempts. Step one: modality selection and preprocessing. CT data needs windowing before any labeling starts. Soft tissue window (350 HU width, 50 HU level) works for pharyngeal and laryngeal structures. Bone window isn't necessary unless you're also labeling bony anatomy like the turbinates or sinus walls. MRI data is a different beast entirely. Most respiratory labeling datasets use CT because the air-tissue contrast is cleaner. If your source data is MRI, budget double the labeling time. Step two: anatomical schema definition. Before opening any annotation tool, write down exactly which structures you're labeling and how you're defining each boundary. This sounds obvious and nobody does it. Our first dataset had no boundary definition for the cricopharyngeus muscle. The annotators basically guessed, and those labels were unusable. We ended up removing 40 percent of our training cases because of it.

Get the Full Details

Label the anatomy of the upper respiratory system. Tongue Larynx Inferior nasal meatus ...
Label the anatomy of the upper respiratory system. Tongue Larynx Inferior nasal meatus ...

Define your schemas using standard terminology. SNOMED CT or RadLex codes help if you plan to share or sell the dataset. Otherwise just be internally consistent and document everything. Step three: tool selection. We use Brainsuite, ITK-SNAP, and Labelme depending on the data type. For volumetric CT, ITK-SNAP with the semi-automatic region growing feature is fast but you need to validate every third slice manually. Labelme gives you more control on 2D endoscopic frames but the workflow is slower. PolyMedical is worth looking at if you're building a clinical product and need DICOM compatibility out of the box. Step four: annotation execution. Label slice by slice for CT volumes. Start from the nasopharynx and work down through the trachea bifurcation. For each structure, draw the outline on every slice where it's visible, then interpolate between slices where the structure is small or partially visible. Never skip more than three consecutive slices without a manual label. The interpolation artifacts from skipping slices create training noise that's hard to clean up later.

Step five: quality control. This is where most projects fail. Have a second annotator review a random 20 percent sample. Then have a radiologist review any cases where the two annotators disagreed on boundary placement. Our QC pass typically catches labeling errors in about 12 percent of cases. That sounds high until you see what happens when bad labels train a model.

Edge Cases That Wasted My Time

Post-surgical anatomy is the hardest thing to label. I spent a week trying to annotate a dataset of post-turbinatectomy patients. The natural boundaries were gone. The surgical scarring created ambiguous tissue planes. What I ended up doing was labeling based on the preoperative anatomical template rather than the postoperative appearance. It's not perfect but it's at least consistent across subjects. The alternative is spending hours arguing with yourself about whether remnant turbinate tissue counts as the superior or middle turbinate. Another problem: pediatric airways. The proportions are different, the cartilage is less calcified, and adult reference atlases don't map well. We built a separate labeling protocol for pediatric cases instead of forcing adult definitions onto kids. It saved us from garbage labels.

Anatomy of the upper respiratory system | Human respiratory system, Respiratory system ...
Anatomy of the upper respiratory system | Human respiratory system, Respiratory system ...

Where This Approach Falls Apart

Pixel-level segmentation of the upper respiratory tract doesn't work well with low-dose CT or compressed video frames from endoscopy. The spatial resolution just isn't there. If your source data is below 1mm slice thickness for CT or 720p for endoscopic video, you're better off using coarse region labels instead of fine segmentation masks. Don't waste time creating false precision. Also, automatic segmentation models trained on one scanner manufacturer's data don't generalize well to another. We trained on Siemens data and tried applying it to GE scans. The boundary predictions drifted by several millimeters on the laryngeal structures. Domain adaptation or per-manufacturer retraining is usually necessary. If you need a downloadable tool to get started, the open-source packages I listed above all have free community editions. The full licensing details are on their respective GitHub pages. Just be aware that the free versions often lack batch processing and automated QC features that become essential once your dataset exceeds a few hundred cases.