Working with Labeled X-Ray Tube Datasets: A Practical Guide
I spent about three months last year trying to build a reliable annotation pipeline for chest X-ray images using publicly available datasets, and most of that time was wasted fixing label mismatches rather than actually training anything. The short version is that "X Ray Tube Labeled" data usually refers to datasets where the X-ray tube position, projection angle, or anatomical labels have been manually or automatically tagged. Getting your hands on clean versions of these is straightforward. Getting them right is the hard part. The most commonly used repositories for labeled X-ray data are MIMIC-CXR, CheXpert, and the JSRT dataset. MIMIC-CXR gives you free access once you complete the CITI certification, which takes about 4 hours if you read through the materials instead of blindly clicking next. The labels there include implicit information about tube angles and patient positioning extracted from radiology reports. CheXpert is more focused on disease labels but includes report-derived findings. JSRT is older but has very consistent manual annotations from multiple radiologists. If you specifically need X-ray tube geometric labels — meaning the actual source-to-image distance, tube angle, and projector orientation — you're going to need something more specialized. The DDIL dataset at Columbia University has some of this metadata. Otherwise you're looking at raw DICOM headers where the information is buried across different fields depending on the manufacturer. GE, Siemens, and Philips all store this data in different locations within the DICOM structure.
The Real Problem: DICOM Tag Inconsistency
Here is where most people hit a wall. I had a project where I needed to extract tube angle and SID (source-to-image distance) from about 12,000 portable chest X-rays collected across six hospitals. The expectation was that RescaleWidth and RescaleIntercept would be consistent, but they weren't. More importantly, the ImagingFreeParameters tag — which should contain the tube position data — was populated differently by every PACS vendor. Some hospitals stored it. Most didn't. The workaround I ended up using was writing a Python script that parsed DICOM tags from pydicom, checked for the presence of BeamAngleSequence, ExposureTime, and XRayTubeCurrent, then fell back to extracting the information from the image pixel sequence metadata when those were missing. For images that had absolutely no tube geometry data, I used a model trained on image appearance to predict approximate tube angle based on the clavicle position and scapula shadow. It wasn't perfect, but it filled gaps for about 70% of the missing cases. I used the SimpleITK library for reading the DICOM files because it handles non-standard tag naming better than the default pydicom reader. The script I ended up with took roughly 45 minutes to process a full batch of 5,000 images on a standard workstation. Not fast, but manageable compared to the alternative of manual review.
Annotation Pipelines: What Actually Works
If you're building a dataset from scratch rather than working with existing labeled data, you need to decide between point annotations, bounding boxes, and segmentation masks. Point annotations for tube position are the fastest to produce — one click per image, roughly 3 seconds per case. But they're unreliable for downstream tasks because the exact tube location in the image plane doesn't map cleanly to the physical tube position in the room. Bounding boxes around the visible X-ray tube housing are more useful for detection tasks but require the tube to actually be visible in the frame, which it rarely is in standard PA and AP chest views. The most reliable approach I've found combines two things: DICOM metadata extraction for the actual geometric labels, and a lightweight CNN classifier for when metadata is absent. I trained a simple ResNet-18 on about 2,000 labeled images to predict whether the projection was PA, AP, or lateral, and what the approximate tube angle range was. The model reached about 82% accuracy on held-out test data from a different hospital. That might sound low, but combined with the metadata extraction for the cases where it works, it covers roughly 90% of routine clinical images. For the remaining 10%, you either accept the uncertainty or pay for manual review. Manual review at this scale costs about 30 seconds per image, which adds up fast. I ended up accepting the model predictions with a confidence threshold — anything below 0.75 confidence went to manual QA. This reduced my manual annotation workload from 12,000 images to about 1,200, which is the kind of saving that makes a project feasible.
Get the Full Details

Common Pitfalls That Waste Weeks
First, don't assume that all images in a dataset share the same labeling convention. MIMIC-CXR reports use slightly different terminology than CheXpert. If you're merging datasets, you'll spend days normalizing label schemas. Second, the tube label in a dataset header is not the same thing as ground truth. I found discrepancies of up to 15 degrees between the DICOM-stated tube angle and what the image actually showed when I measured landmark positions manually. Third, portables are a nightmare. The tube angle on a portable chest X-ray is often completely unreliable because the equipment operator can adjust it freely during the procedure. Another thing nobody warns you about: image compression. Some datasets distribute JPEG-compressed versions that destroy the subtle intensity gradients you need for tube edge detection. Always check whether your data is uncompressed or losslessly compressed before investing time in a pipeline that depends on pixel-level accuracy. I lost an entire week to this on a dataset I thought was TIF when it was actually heavily compressed JPEG underneath.
Software Tools Worth Using
For DICOM parsing and metadata extraction, pydicom is the standard. It's well-documented and handles most vendor-specific quirks. For annotation creation, CVAT is free and handles DICOM natively, which saves you from having to convert files first. Label Studio is another option with a simpler interface if you're doing rapid prototyping. Both support exporting to COCO, Pascal VOC, or YOLO formats, which matters if you're planning to feed the data into an object detection model later. For the prediction-based gap-filling I described, I used a modified DenseNet-121 trained on the CheXpert+MIMIC-CXR combined labels. The codebase was built on PyTorch Lightning, and training took about 6 hours on a single RTX 4090. You don't need anything more powerful than that for this kind of task. If you're working with GPU resources, consider using the MONAI framework, which has prebuilt pipelines for medical image classification and segmentation that will save you significant setup time.
When This Approach Doesn't Work
This whole pipeline falls apart if you're working with historical film-based X-rays that were digitized through scanners. Those don't carry DICOM geometry metadata at all, and the tube position is essentially unrecoverable from the image content alone. If that's your use case, you're either looking at specialized archival databases that already have the metadata indexed, or you accept that certain analyses are impossible with your data. There's no way around that limitation. Similarly, if you need sub-millimeter precision for tube localization, none of these methods will get you there. DICOM metadata has a resolution limit, and image-based prediction has inherent uncertainty. For that level of precision, you need specialized phantoms and manual measurement, which is a completely different project scope.

Summary of What You Actually Need to Do
Start by identifying which dataset best matches your needs. Check the DICOM metadata completeness before committing to it — look at a random sample of 100 files and see what percentage have the tags you need. Build your extraction pipeline with pydicom first, and only add the prediction model for gaps. Set a confidence threshold early and decide what goes to manual review before you start scaling up. Budget time for label normalization if you're merging datasets. And whatever you do, verify your image format before you write a single line of processing code. The X Ray Tube Labeled data ecosystem is functional but uneven. You can build something useful with it, but only if you account for the inconsistency from the start rather than discovering it after you've already trained a model on mislabeled data. I learned that one the hard way.