Quest 103 - AI Bias Auditor
Quest 103: AI Bias Auditor
medium 25-30 minutes🎯 Learning Objectives
- Why measuring fairness across groups is essential for responsible AI
- How to compute demographic parity and equalized odds
- The difference between group-level fairness metrics and global averages
- How bias can hide behind seemingly 'fair' aggregate numbers
📖 Concept: AI Fairness Metrics
AI systems make decisions that affect real people — loan approvals, hiring recommendations, medical diagnoses. When these systems treat different demographic groups unequally, it’s called algorithmic bias. The problem is, bias can hide in plain sight. A system might have 90% overall accuracy, but if it’s 95% accurate for one group and only 75% for another, that’s a serious fairness issue.
Fairness metrics quantify these disparities. Two of the most important are:
- Demographic Parity: Do different groups get positive predictions at the same rate? If a hiring AI recommends 40% of Group A but only 15% of Group B, demographic parity is violated.
- Equalized Odds: Among people who actually qualify, does the AI correctly identify them at the same rate across groups? If the AI catches 90% of qualified candidates from Group A but only 60% from Group B, equalized odds is violated.
The key insight is that fairness must be measured between groups, not by averaging everything together. Averages mask disparity.
⚙️ How It Works
The Fairness Computation Pipeline
1. Collect predictions with group labels ↓2. Group predictions by demographic group ↓3. Compute per-group statistics: - Positive prediction rate - True positive rate ↓4. Compare metrics across groups ↓5. Report disparities (demographic parity gap, equalized odds gap)Why Group-Level Metrics Matter
Consider a simple example with 10 predictions per group:
| Group | Total | Predicted Positive | Actually Positive | True Positives |
|---|---|---|---|---|
| Group A | 10 | 8 | 5 | 5 |
| Group B | 10 | 3 | 5 | 2 |
- Group A positive rate: 80% | True positive rate: 100%
- Group B positive rate: 30% | True positive rate: 40%
- Demographic parity gap: 80% - 30% = 0.50 (large disparity)
- Equalized odds gap: 100% - 40% = 0.60 (large disparity)
If you averaged everything together, you’d see “70% positive rate” and miss the fact that Group B is being systematically underserved.
💡 Example: Computing Fairness Metrics
Here’s how to compute fairness metrics in JavaScript:
Step 1: Group predictions by demographic
const groups = {};for (const p of predictions) { if (!groups[p.group]) groups[p.group] = []; groups[p.group].push(p);}Step 2: Compute per-group statistics
for (const name of Object.keys(groups)) { const preds = groups[name]; const total = preds.length; const positivePredicted = preds.filter(p => p.predicted).length; const actualPositive = preds.filter(p => p.actual); const truePositives = actualPositive.filter(p => p.predicted).length;
groupStats[name] = { positiveRate: positivePredicted / total, truePositiveRate: actualPositive.length > 0 ? truePositives / actualPositive.length : 0, };}Step 3: Compute the gap between groups
const positiveRates = groupNames.map(g => groupStats[g].positiveRate);const demographicParity = Math.max(...positiveRates) - Math.min(...positiveRates);
const tprRates = groupNames.map(g => groupStats[g].truePositiveRate);const equalizedOdds = Math.max(...tprRates) - Math.min(...tprRates);A gap of 0 means perfect fairness. Larger gaps indicate greater disparity.
⚠️ Common Mistakes
Mistake 1: Averaging across all groups
“Overall accuracy is 85%, so the model is fair” → Averages mask group-level disparities. Always compute metrics per group first, then compare.
Mistake 2: Ignoring groups with small sample sizes
“Group C only has 2 predictions, so skip it” → Small groups matter too. Report the sample size alongside metrics so others can judge reliability.
Mistake 3: Confusing demographic parity with equalized odds
“They both measure fairness, so they’re the same” → They measure different things. Demographic parity checks prediction rates; equalized odds checks accuracy among qualified individuals. A model can satisfy one but violate the other.
Mistake 4: Rounding away the problem
“The gap is only 0.003, so it’s negligible” → Even small gaps can have large real-world impact when applied to thousands of decisions. Report exact values and let stakeholders decide the threshold.
📝 Knowledge Check
📝 Knowledge Check
Q1:What is demographic parity?
Q2:Why is computing a single global fairness metric misleading?
Q3:How do you compute the demographic parity gap?
🏋️ Quest: AI Bias Auditor
Now it’s time to practice! Implement a fairness metrics calculator for AI predictions.
-
Download the starter files:
Terminal window npx bluebeltdojo download quest-103-bias-auditorcd quest-103-bias-auditor -
Open
problem.jsin your editor with your AI tool -
Implement
computeFairnessMetrics(predictions)that:- Groups predictions by demographic group
- Computes positive prediction rate and true positive rate per group
- Calculates demographic parity gap and equalized odds gap
-
Run
node test.jsand read the failures carefully -
Fix any edge cases the AI missed
-
Verify all tests pass:
Terminal window node test.js -
When all tests pass, submit your solution:
Terminal window npx bluebeltdojo submit
💡 Tip: The key insight is that fairness must be computed BETWEEN groups, not by averaging. If your code only computes a single global metric, you’re missing the point.
คำใบ้
- อ่าน instructions ใน
problem.jsอย่างละเอียด - Fairness metrics ต้องคำนวณระหว่างกลุ่ม (between groups) ไม่ใช่ค่าเฉลี่ยรวม
- Demographic parity = ความแตกต่างของ positive prediction rates
- Equalized odds = ความแตกต่างของ true positive rates
- ถ้าติดขัด ลองอ่าน “Common Mistakes” อีกครั้ง — อย่าดู solution โดยตรง