Getting Computer Vision Training Data Right
Most people approaching CV training don't realize how much time actually gets wasted on bad label files before anything meaningful ever runs. I've sat through too many project post-mortems where the model performance tanked because someone treated the annotation phase as administrative overhead instead of the actual foundation. The difference between a working pipeline and a failed one almost always comes down to the answers you collect during training. This is the phase where you define ground truth for your dataset — bounding boxes, segmentation masks, keypoint placements, or whatever labeling schema your model needs. The answers aren't just "yes this object is a car." They're structured JSON or XML files that every downstream process depends on. Get these wrong and you're not debugging code later, you're debugging data. I've found that the most common mistake people make is creating overly broad category labels. Early in one project, our team labeled a category as "person" across thousands of images, only to discover that the model couldn't distinguish between standing figures, partially occluded humans, and reflective mannequins in shop windows. The training accuracy looked fine at first, but once we deployed to real-world conditions, recall on actual people dropped to about thirty-four percent. We ended up splitting "person" into four subcategories and re-annotated roughly two thousand images. That decision alone shifted our F1 score from 0.61 to 0.89 on the validation set.
The Annotation Workflow
Start with a clear labeling guideline document before you open any annotation tool. I'm not talking about a single page of bullet points. I mean a living document that specifies exactly what constitutes a positive label, what the edge cases are, and how to handle ambiguous situations. When our team skipped this step on a traffic sign detection project, annotators disagreed on whether partially faded signs should still count. About fifteen percent of the dataset ended up being inconsistent just from that ambiguity alone. Pick your tool carefully. LabelImg is fine for simple bounding box work. For segmentation masks, cvat.ai gives you more control, though the interface has a learning curve. If you're doing keypoints or complex polygon work, consider Supervisely or even Polygonist depending on your budget. Free tools work for small datasets, but they don't scale well past five thousand images without significant manual effort. Before you label a single image, create a small test batch — maybe fifty images — and have at least two people annotate them independently. Then measure inter-annotator agreement using IoU for bounding boxes or Dice coefficient for segmentation. If your agreement score is below eighty percent, your guidelines need revision before you scale up. This step usually takes one to two days but prevents weeks of rework later.
Evaluating Your Training Answers
Once your annotations are complete, the real work starts. Most teams jump straight into training their model, but that's premature. You need to validate the answers first. Run a random sample of about ten percent of your labeled data through a visual QA pass. Look for misaligned boxes, missed objects, incorrect categories, and overlapping annotations that shouldn't be there. In one project I managed, our automated quality check flagged about four percent of annotations as problematic, but the visual audit caught another six percent that the automated tools missed — mostly subtle issues like boxes that were technically within tolerance but semantically wrong. After the QA pass, split your dataset into training, validation, and test sets. A common error I see is using overlapping data across these splits. If your images come in clusters — say, multiple shots of the same intersection from the same camera angle — make sure all images from a given cluster go into one split only. Data leakage between sets will inflate your metrics during development and give you a false sense of confidence when you ship.
Get the Full Details

Common Pitfalls and How I Deal With Them
Class imbalance is the thing that catches everyone out eventually. If you're training a model to detect defects on a production line and only three percent of your images contain actual defects, your model will learn to predict "no defect" for everything and still achieve ninety-seven percent accuracy. That accuracy number is useless. Use techniques like class-weighted loss functions, oversampling the minority class, or generating synthetic samples with augmentation. I usually recommend combining at least two of these approaches rather than relying on just one. Another issue that doesn't get enough attention is boundary cases in your labeling schema. What happens when an object is partially outside the frame? What if two objects of the same class overlap significantly? What about objects that are blurry or heavily occluded? These decisions need to be explicit in your guidelines, not left to individual annotator judgment. I ran into a particularly stubborn problem last year where the lighting conditions in our training images were dramatically different from the deployment environment. The training set was shot under bright daylight with high contrast, but the model was being deployed in a warehouse with dim, yellow-tinted fluorescent lighting. The model performed adequately on our test set but failed in production. The fix wasn't just adding more training data — it was specifically acquiring or simulating data from the target deployment environment. We ended up spending about six weeks collecting in-situ images and rebalancing the dataset to reflect the actual operating conditions. It cost more time upfront but saved us from a major rollout failure.
File Formats and Conversion
Make sure your annotation files are in a format that your training framework can actually consume. TensorFlow expects either TFRecord or COCO JSON format. PyTorch projects often use YOLO format with individual text files per image. If you're working across teams, standardize on COCO JSON as the interchange format — it's the most universally supported and handles multi-class segmentation without extra configuration. I've seen teams waste entire days trying to debug training failures that traced back to a simple format mismatch. One person exported annotations as Pascal VOC XML while the training script expected YOLO format. The model trained without errors but learned nothing useful because the coordinate normalization was completely wrong. Always verify your format conversion with a small sample before committing to full-scale training.
When to Bring in External Help
If your dataset exceeds about ten thousand images with complex labeling requirements, consider using a professional annotation service. Internal teams can manage smaller projects, but quality tends to degrade as volume increases because your annotators get fatigued and your review capacity becomes a bottleneck. Professional services charge anywhere from fifteen cents to two dollars per annotation depending on complexity, but they also provide quality guarantees and typically deliver faster turnarounds for large batches. The key is providing them with the same detailed guidelines you would give an internal team, plus a calibrated reference set where you've already marked the correct answers. This lets them calibrate to your standards before they touch your actual dataset. Without that calibration step, you'll get results that look structurally correct but are semantically misaligned with what your model actually needs to learn.
