The Loss Functions You Should Actually Be Using

I spent about three years debugging models that performed well in training and completely fell apart in production. The common thread was almost always the loss function. Everyone defaults to categorical cross-entropy because it's the first thing in the documentation, but that doesn't mean it's the right call for your specific problem. I'm going to walk through the strategies I've actually seen move the needle, not the textbook examples. Focal Loss is probably the most useful modification you'll encounter if you're working with imbalanced data. The standard cross-entropy loss treats every example equally, which means your model spends most of its gradient signal on examples it already handles well. Focal loss reweights the loss so that hard examples get more attention. The formula introduces two parameters: gamma, which controls how much easy examples are downweighted, and alpha, which balances the positive and negative classes. In practice, gamma around 2 and alpha around 0.25 works as a starting point for most object detection tasks. I used this on a fraud detection pipeline where the positive class was roughly 0.3% of the data. Switching from binary cross-entropy to focal loss cut our false positive rate by about 40% without touching the architecture. The tradeoff is that focal loss is more sensitive to learning rate choices. I ended up needing to drop the initial learning rate by half compared to what worked with standard cross-entropy, and I had to run a wider grid search to find a stable configuration. Here's something people don't talk about enough: label smoothing isn't just a regularisation trick. It fundamentally changes how your model calibrates its confidence scores. Standard softmax cross-entropy pushes the model toward predicting probabilities of exactly 0 or 1 for the training classes, which creates overconfident models that fail on distribution shifts. Label smoothing replaces the hard one-hot target with a distribution where each class gets a small probability mass, typically epsilon values between 0.1 and 0.2. This produces better-calibrated outputs and tends to improve generalisation on out-of-distribution samples. I ran into an edge case with label smoothing that I didn't expect. On a multi-label classification task where samples can have multiple active classes simultaneously, applying label smoothing naively caused the model to under-predict the number of active labels. The workaround was to only apply smoothing to the negative labels, leaving the positive labels untouched. This preserved the model's ability to predict the correct number of classes while still benefiting from the regularisation on the negative ones.

Contrastive loss and its variants are worth understanding even if you're not doing self-supervised learning. The basic idea is to pull similar examples together and push dissimilar ones apart in embedding space. triplet loss is a simpler version that uses an anchor, a positive, and a negative. The margin parameter is critical here - too small and the model doesn't learn a meaningful separation, too large and training becomes unstable because examples can't satisfy the constraint. I found that a margin of 0.3 to 0.5 works for most image embedding tasks, but text embeddings often need something closer to 0.1. There's also InfoNCE loss, which is what contrastive language models use. It treats the problem as a classification task over negatives, and it scales better with large batch sizes. The downside is that it requires a lot of negative samples to be effective, which means you need decent GPU memory or a careful batching strategy. I should mention huber loss for regression tasks. Most people default to mean squared error for regression, which makes sense mathematically under Gaussian assumptions, but MSE is extremely sensitive to outliers. A single bad label can dominate your gradient. Huber loss switches to L1 behaviour for large errors and L2 for small ones, controlled by a delta parameter. Delta values between 1.0 and 2.0 are reasonable starting points. This matters more than you'd think in production systems where label quality degrades over time. I worked on a price prediction model where 5% of the training labels were consistently wrong due to a data pipeline bug. With MSE, those outliers pushed the model's predictions off by 12-15% on average. Switching to huber loss with delta=1.5 brought the error down to about 4% without any label cleaning. Mixture density loss is another strategy that gets overlooked. Standard loss functions assume a single point prediction, but many real-world problems have multimodal output distributions. If you're predicting delivery times, there might be a fast mode and a slow mode depending on weather and traffic. A single Gaussian prediction will land somewhere in between and be wrong in both modes. Mixture density networks output the parameters of a mixture distribution instead of a point estimate. You're basically training the network to predict a weighted combination of Gaussians. This is computationally more expensive and harder to tune, but it's genuinely useful when your downstream decisions depend on understanding the full shape of the prediction distribution, not just the mean.

One thing I wish I'd understood earlier: loss function selection should match your evaluation metric, not just your training objective. If you're optimising for AUC but your loss function doesn't correlate well with ranking quality, you'll see a gap between training loss and test performance. In those cases, consider using a ranking-based loss like pairwise hinge loss or listwise losses such as ListNet. These are more expensive to compute but they optimise directly for what you actually care about. I used a pairwise approach on a recommendation system where the business metric was click-through rate. The model trained with standard cross-entropy had good log-loss but mediocre CTR. Switching to a sampled softmax with a temperature of 0.1 improved the ranking quality and CTR by about 8 percentage points. There are scenarios where none of these strategies help and the problem is purely data quality or feature representation. I've seen teams spend weeks tuning loss functions on models that were fundamentally limited by noisy labels or missing features. Before investing effort into sophisticated loss modifications, verify that your baseline model with standard cross-entropy and a clean dataset is already performing close to the bayes optimal. If it's not, the issue is elsewhere. The diminishing returns from loss function tweaks are real and usually happen after you've exhausted data and feature engineering options.

Get the Full Details

Visualizing the Loss Landscape of a Neural Network
Visualizing the Loss Landscape of a Neural Network