Fairness & thresholds
Two groups, one model. Tune thresholds and see what equal opportunity costs.
Two groups, one model. Tune thresholds and see what equal opportunity costs.
A scholarship screen gives each of 200 applicants a score from 0 to 100. Applicants at or above the threshold are selected. Later we learn who was qualified (would have done well with the scholarship). Synthetic data.
■ Group A
■ Group B
Each criterion captures a reasonable idea of fairness. Demographic parity asks that both groups are selected at the same rate. Equal opportunity asks that qualified applicants have the same chance (TPR) in both groups. Equalised odds also asks for the same false-positive rate. Predictive parity asks that a selection means the same thing: equal precision.
When the groups’ base rates p (share qualified) differ, these criteria pull in different directions. Look at the formulas: if TPR and FPR are equal in both groups, the selection rate TPR·p + FPR·(1 − p) still depends on p, so the selection rates come apart (unless TPR = FPR, a model that is no better than a coin). Even a perfect model fails demographic parity: it selects exactly the qualified applicants, so each group is selected at its own base rate. Equal precision on top of equal error rates only happens in extreme cases, such as selecting almost nobody, which is why the criteria are not counted below 10% or above 90% selection. Switch to “Same base rates” and the conflict mostly disappears.
Which criterion matters most is a judgement about values and context, not something the maths can decide. A shared threshold treats equal scores equally; separate thresholds can equalise outcomes or error rates. It also matters why scores or base rates differ (unequal access to coaching, a biased measurement or real differences): the data alone cannot tell you.
| Metric | ■ A | ■ B | Gap |
|---|---|---|---|
| Selection rate (TP+FP)/n | 45% | 16% | 28.7 pp |
| TPR TP/(TP+FN) | 82% | 50% | 31.7 pp |
| FPR FP/(FP+TN) | 8% | 5% | 3.3 pp |
| Precision TP/(TP+FP) | 91% | 77% | 13.8 pp |
| Accuracy (TP+TN)/n | 87% | 84% | 2.9 pp |
Fairness criteria“equal” = gap ≤ 5 pp · 0/4 met
Equal selection rates: both groups are selected at the same rate.
Equal TPR: qualified applicants have the same chance of selection in both groups.
Equal TPR and equal FPR: the same hit rate and the same false-alarm rate in both groups.
Equal precision: a selected applicant is equally likely to be qualified in both groups.
What does each rule cost? most accurate thresholds that satisfy it
| Rule | t A / B | Accuracy | Apply |
|---|---|---|---|
| No fairness rule | 54 / 53 | 87.0% | |
| One shared threshold | 53 / 53 | 86.0%−1.0 pp | |
| Demographic parity | 56 / 47 | 82.0%−5.0 pp | |
| Equal opportunity | 59 / 50 | 86.0%−1.0 pp | |
| Equalised odds | 59 / 50 | 86.0%−1.0 pp | |
| Predictive parity | 55 / 60 | 86.5%−0.5 pp |
Your thresholds now: 85.5% accurate. The cost of a rule depends on the data; here it is a few points, elsewhere it can be much larger or close to zero.
Challenges 0/2 · different base rates
Then check what happened to precision in each case.
Each criterion captures a reasonable idea of fairness. Demographic parity asks that both groups are selected at the same rate. Equal opportunity asks that qualified applicants have the same chance (TPR) in both groups. Equalised odds also asks for the same false-positive rate. Predictive parity asks that a selection means the same thing: equal precision.
When the groups’ base rates p (share qualified) differ, these criteria pull in different directions. Look at the formulas: if TPR and FPR are equal in both groups, the selection rate TPR·p + FPR·(1 − p) still depends on p, so the selection rates come apart (unless TPR = FPR, a model that is no better than a coin). Even a perfect model fails demographic parity: it selects exactly the qualified applicants, so each group is selected at its own base rate. Equal precision on top of equal error rates only happens in extreme cases, such as selecting almost nobody, which is why the criteria are not counted below 10% or above 90% selection. Switch to “Same base rates” and the conflict mostly disappears.
Which criterion matters most is a judgement about values and context, not something the maths can decide. A shared threshold treats equal scores equally; separate thresholds can equalise outcomes or error rates. It also matters why scores or base rates differ (unequal access to coaching, a biased measurement or real differences): the data alone cannot tell you.
Things to try