Odd Is On Our Side
I've been dealing with this stuff for a long time, and it's not glamorous. Odd Is On Our Side is a framework that emerged from the intersection of algorithmic bias analysis and fairness testing in predictive systems. At its core, it's a methodology for detecting when your model's error rates skew differently across subgroups, and then adjusting for it without blowing up overall accuracy. Here's how you actually run it. You start by splitting your validation data into protected and non-protected cohorts based on whatever attribute matters for your use case. Then you compute the confusion matrix separately for each group. That's where most people stop, and that's exactly where things go wrong.
Practical Setup for Odd Is On Our Side
I recommend you pull a precision-recall curve per subgroup, not just AUC. AUC smooths over the kind of disparities this framework is designed to catch. When I was auditing a hiring model last year, the overall AUC looked fine at 0.87, but when I broke it down by gender quartiles, one subgroup had a precision of 0.41 while another sat at 0.73. That gap is exactly what Odd Is On Our Side flags. The actual tuning step involves recalibrating your decision threshold per subgroup. You pick a target false positive rate, then find the threshold that hits it for each group independently. The result is different cutoff points across cohorts. Your model outputs stay the same. Only the pass/fail line moves. This is where the work gets tedious. You need enough samples per subgroup for the thresholds to be stable. If one group has fewer than roughly 500 observations, the threshold estimates become unreliable and you're just moving noise around. I hit that wall hard on a healthcare readmission model where a rural demographic had a sample size of about 200 in the validation set. I ended up pooling adjacent regions to get a usable count rather than running the adjustment on an underpowered group.
What Most People Miss
The biggest pitfall is treating the output as a fixed correction. It isn't. Threshold adjustments assume your feature distribution stays relatively stable between training and deployment. When your input data shifts, the thresholds drift out of alignment. I saw this happen with a credit model where the demographic composition of applicants changed seasonally. The thresholds that worked in Q1 produced measurable skew in Q3. You have to re-run the calibration every quarter at minimum. Another thing nobody tells you: Odd Is On Our Side can only address disparities that exist in your features. If the underlying data generation process encodes a systemic pattern, no threshold adjustment will fully neutralize it. The framework finds imbalances in your model's output, not in the world that produced your training data. That distinction matters because it sets a hard ceiling on what you can fix without going upstream into data collection. The workflow itself is straightforward enough that you can implement it in a few hundred lines. You grab your predictions, group them, compute per-group thresholds, and apply them. The time investment is usually around 4 to 6 hours for a first pass on a medium-sized dataset, not counting the iteration cycles when you realize your subgroup definitions were off.
Get the Full Details

If your use case involves real-time decisions where different thresholds per group would be legally problematic, this approach won't help. In those situations, you're better off looking at adversarial debiasing techniques that modify the model itself during training rather than adjusting post-hoc thresholds.