Discrimination and Representative Signal Distortion
Discrimination and Representative Signal Distortion
Discrimination
Sevgi Yuksel
Very difficult!
How do define discrimination?
Very difficult!
1. Regression analysis
− Goldberger (1984); Neal and Johnson (1996)
2. Audit studies
− Bertrand and Mullainathan (2004)
3. Quasi-experiments
− Goldin and Rouse (2000); Anwar, Bayer, and Hjalmarsson (2012)
4. Testing Models
− Charles and Guryan (2008) Chandra and Staiger (2010)
Regression analysis
• Old literature (> four decades old) has tested for evidence of
discrimination in labor, housing, and product markets by conducting
‘audit’ field experiments
• Useful overview: Riach and Rich (2002)
• Conclusion of Riach and Rich: “...demonstrated pervasive and
enduring discrimination against non-whites and women”
Bertrand and Mullainathan (2004)
Confirmation bias.
• Distortion of the signal in RSD is not driven directly the prior but by
the contrast between the prior and other distributions.
If t ∼ N µg , σ 2 and s = t + ϵ with ϵ ∼ N 0, ξ 2 ,
ξ2
t̂ Bay = ωµg + (1 − ω)s, where ω Bay = .
ξ2 + σ2
• κ is a normalization factor,
• γ ≥ 0 measures distortion due to representativeness,
R(s, g , −g ) = yy(s(s| |−g
g) R
•
)
with y (s | g ) := h(s | t, g )f (t |, g )dt captures
how representative signal s is of group g given reference group −g .
Implications of contrast-biased evaluation in guassian setting
Assume t ∼ N µg , σ 2 and s ∼ N t, ξ 2 ,
ξ2
∆g = γ (µg − µ−g ).
ξ2 + σ2
Implications of contrast-biased evaluation in guassian setting
Assume t ∼ N µg , σ 2 and s ∼ N t, ξ 2 ,
ξ2
∆g = γ (µg − µ−g ).
ξ2 + σ2
Predictions:
Assume t ∼ N µg , σ 2 and s ∼ N t, ξ 2 ,
ξ2
∆g = γ (µg − µ−g ).
ξ2 + σ2
Predictions:
Assume t ∼ N µg , σ 2 and s ∼ N t, ξ 2 ,
ξ2
∆g = γ (µg − µ−g ).
ξ2 + σ2
Predictions:
• Abstract design.
− Removes confounds (taste-based discrimination).
− Controls prior beliefs and objective.
• Incentives:
− Base payment of $7.5
− Chance of winning bonus $20 is (100 - MSE) percent.
Baseline
75
65
Mean Assesment
55
45
35
25
30 40 50 60 70
Type
A first look at results
Baseline
75
65
Mean Assesment
55
45
35
25
30 40 50 60 70
Type
A first look at results
Baseline
75
65
Mean Assesment
55
45
35
25
30 40 50 60 70
Type
A first look at results
Baseline
75
65
Mean Assesment
55
45
35
25
30 40 50 60 70
Type
A first look at results
Baseline
75
B = 1.8, ω = 0.16
65
Mean Assesment
55
45
35
B = -1.7, ω = 0.20
25
30 40 50 60 70
Type
A first look at results
Baseline NoGroup
75
75
65
65
Mean Assesment
Mean Assesment
55
55
45
45
35
35
25
30 40 50 60 70 30 40 50 60 70
Type Type
A first look at results
OneGroup SignalFist
75
75
B = -0.2, ω = 0.16 B = 0.2, ω = 0.15
65
65
Mean Assesment
Mean Assesment
55
55
45
45
35
35
B = 0.4, ω = 0.09
25 B = -0.3, ω = 0.18
25
30 40 50 60 70 30 40 50 60 70
Type Type
Estimates of representative signal distortion, ∆g
Baseline NoGroup
1
.8
.8
.6
.6
Cdf
Cdf
.4
.4
.2
.2
0
0
-10 -8 -6 -4 -2 0 2 4 6 8 10 -6 -4 -2 0 2 4 6
Δg Δg
OneGroup SignalFirst
1
1
.8
.8
.6
.6
Cdf
Cdf
.4
.4
.2
.2
0
-10 -8 -6 -4 -2 0 2 4 6 8 10 -10 -8 -6 -4 -2 0 2 4 6 8 10
Δg Δg
Individual-level estimates of base-rate neglect, ω Bay − ω
Baseline NoGroup
1
.8
.8
.6
.6
Cdf
Cdf
.4
.4
.2
.2
0
0
-.2 0 .2 .4 .6 -.2 0 .2 .4 .6
ωBay-ω ωBay - ω
OneGroup SignalFirst
1
1
.8
.8
.6
.6
Cdf
Cdf
.4
.4
.2
.2
0
-.2 0 .2 .4 .6 -.2 0 .2 .4 .6
ωBay- ω ωBay- ω
Measures of (in)accuracy and discrimination
R
Group difference in assessments, GD := E t̂h − t̂l | t dF (t).
• Linked to “seperation” criteria in ML fairness literature.
(Barocas Hardt Narayanan 2019, Narayanan 2018; Hutchinson Mitchell 2019)
µl +µh
Focus on linear strategies, i.e., t̂ = ω1 µg + ω2 2
+ (1 − ω1 − ω2 )s.
Accuracy-discrimination frontier
Accuracy-discrimination frontier
Implications for Statistical Discrimination
85
75
Inaccuracy (MSE)
65
Baseline
55
45
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
45
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
OptNoDiscrimination
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
OptNoDiscrimination
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
OptNoDiscrimination NoBias
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
OptNoDiscrimination NoBias
SignalFirst
45
Bayesian
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Implications for Statistical Discrimination
85
pBRN
75
NoGroup
Inaccuracy (MSE)
65
Baseline
55
OptNoDiscrimination NoBias
SignalFirst
45
Bayesian
OneGroup
35
-1 0 1 2 3 4 5 6 7 8 9 10
Discrimination (GD)
Using the model to improve outcomes