Mitigating Algorithmic Bias in Healthcare
Mitigating Algorithmic Bias in Healthcare
Abstract
Background AI and predictive analytics have increased the speed of innovation in medicine. If left unchecked,
however, algorithmic bias can exacerbate health disparities across race, class, or gender. Early bias mitigation literature
has focused on addressing bias in the preparation and development phases of the algorithm life cycle (pre- and in-
processing). Post-processing methods, applied at the point of implementation, are less computationally intensive
and do not require re-building or training the model, allowing lower-resourced health systems to improve bias
in off-the-shelf binary classification models, which are increasingly common within electronic medical records. This
umbrella review sought to identify post-processing bias mitigation methods and tools applicable to binary healthcare
classification models in healthcare and summarize bias reduction effectiveness and accuracy loss.
Methods This review was registered with PROSPERO and reported according to PRISMA 2020. PubMed and Sco-
pus were searched in December 2023 for English-language reviews published post-2013 using an expanded search
string from previous work on machine learning bias. Eligibility criteria followed the PICOT framework. Reviews
were screened independently by two authors. Data were extracted from reviews using the Joanna Briggs Institute
Extraction Form for Review of Reviews, as well as from cited studies (hence, an “extended” umbrella review). Quality
was assessed using the Critical Appraisal Checklist for Systematic Reviews. Evidence was synthesized by mitigation
method and effectiveness.
Results Searches yielded 184 records. After duplicate removal, title/abstract, and full text screening, 11 reviews
were included, citing 16 eligible studies. Post-processing methods tested included threshold adjustment (9 studies,
cited by 8 reviews), reject option classification (6 studies, cited by 4 reviews), and calibration (5 studies, cited by 4
reviews). Threshold adjustment reduced bias across 8/9 trials; reject option classification and calibration reduced bias
in approximately half of trials (5/8 and 4/8). Results were reported with heterogeneous fairness and accuracy metrics,
making comparison difficult. A lack of effectiveness evaluation was noted across reviews. Four reviews identified
16 software libraries for addressing bias. Quality of the majority of reviews was weak due to inadequate reporting
on methods.
*Correspondence:
Shaina Mackin
mackins@[Link]
Full list of author information is available at the end of the article
© The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long
as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if
you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or
parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated
otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not
permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To
view a copy of this licence, visit [Link]
Mackin et al. BMC Digital Health (2025) 3:26 Page 2 of 13
Conclusions Threshold adjustment showed significant promise in post-processing bias mitigation for healthcare
algorithms, followed by reject option classification and calibration. Future research should empirically compare
post-processing methods on binary classification models using real-world healthcare data. As commercial algorithms
proliferate, health systems require proven, achievable strategies to maximize fairness.
Keywords Algorithmic bias, Algorithmic fairness, Bias mitigation, Health equity
Reasons for exclusion upon full text review included lack narrative. Notably, only one review provided a compre-
of post-processing methods or measurement (n = 12) and hensive list of all its included studies and their charac-
not being a review (n = 2). Out of the 11 retained reviews, teristics [5]. The 11 included reviews collectively cited
one was systematic, three were scoping, and seven were 45 studies when discussing post-processing methods.
Mackin et al. BMC Digital Health (2025) 3:26 Page 5 of 13
Upon consulting the texts of those 45 cited studies, only Cited studies consistently failed to provide standard a
16 (36%) actually tested and measured bias mitigation for priori thresholds of success for bias reduction across or
post-processing methods applicable to binary classifica- within fairness metrics. Reported rates of bias reduction
tion models in healthcare. These 16 studies, reporting 25 varied and were difficult to compare across studies due to
trials, were retained in the extended umbrella review, as the heterogeneity of fairness metrics.
seen in Fig. 1. Defining uniform fairness metric improvement of
any magnitude as successful bias reduction, threshold
Available methods: evidence from reviews and cited adjustment showed promise across cited studies. Thresh-
studies old adjustment showed uniform bias reduction in eight
Table 1 charts retained reviews according to their char- out of nine trials (88.9%). Reject option classification
acteristics, including field, review type, post-processing showed uniform bias reduction in five out of eight trials
methods discussed, and open source tools reported. (62.5%), and mixed bias reduction in two (25%). Calibra-
The most common post-processing methods tested tion showed uniform bias reduction in four out of eight
were variants of threshold adjustment (9 studies, cited by trials (50%) and mixed reduction in two (25%). Table 2
8 reviews), reject option classification (6 studies, cited by summarizes the effectiveness of each study by mitigation
4 reviews), and calibration (5 studies, cited by 4 reviews). method using the bias reduction categories defined in
While “equalized odds” (EO) was frequently cited as Methods. Figure 2 summarizes the effectiveness of bias
a mitigation method, EO is a fairness metric popular- reduction by method, across studies.
ized by Hardt et al. in a paper that also introduced one
method of classifier adjustment to reduce bias [31]. Stud- Fairness‑accuracy tradeoffs: evidence from cited studies
ies that cited “equalized odds” as a method were catego- Several cited studies measured success in terms of
rized within a family of threshold adjustment methods fairness-accuracy trade-off curves [36, 40]. Figure 3
[31–35]. Nine studies introduced novel post-processing shows a summary of the loss of accuracy across stud-
methods, which were included when meeting our other ies by method. Eleven of the sixteen cited studies meas-
screening criteria. These included four threshold adjust- ured accuracy loss during their post-processing bias
ment variants [31, 33, 34, 36], two reject option classifi- mitigation trials, although studies used heterogeneous
cations [32, 37], and three calibrations [18, 38, 39]. The sub-group and overall metrics. Five trials saw no accu-
remaining studies evaluated already established methods. racy loss as a result of mitigation [31, 32, 34, 37, 41,
Table 2 classifies studies by novel or empiric method and 42]. In fourteen trials the mitigation method resulted in
provides the definitions and frequencies of all method low losses to accuracy [18, 22, 32, 35, 36, 38]. No trials
categories cited: threshold adjustment, rejection option reported high losses in accuracy. Six trials did not report
classification, and calibration. on accuracy measurement, namely variants of calibration
where the fairness-accuracy tradeoff is less central [39,
Method effectiveness: evidence from cited studies 43, 44]. Threshold adjustment had four trials with no loss
Only one of the 11 included reviews, Huang et al., 2022, to accuracy, and three trials with low loss to accuracy.
reported measures of bias reduction from the primary Reject option classification had two trials with no loss to
studies it included [5]. Across the 16 retained studies accuracy, and five trials with low loss to accuracy. Cali-
cited by reviews, post-processing bias mitigation meth- bration had zero trials with no loss to accuracy and five
ods were evaluated across 25 trials. Bias reduction was trials with low loss to accuracy.
quantified for threshold adjustment, reject option clas- Threshold adjustment had four trials with uniform bias
sification, and calibration. Seventeen of these 25 evalu- reduction and no loss to accuracy. Reject option classifi-
ations (68%) showed uniform bias reduction of a given cation had two trials with uniform bias reduction and no
method. These results are classified and evaluated in Sup- loss to accuracy. Calibration had zero trials with uniform
plementary Table 2 (Table S2). bias reduction and no loss to accuracy. No standards
The 25 trials reported in 16 cited studies used hetero- regarding optimal fairness-accuracy tradeoff levels were
geneous metrics to measure fairness and bias reduction. reported.
Metrics included equal opportunity (false negative rate
or true positive rate parity), equalized odds (true positive Tools for implementation: evidence from reviews
rate and false positive rate parity), and demographic par- Software libraries for bias mitigation were a focus of
ity (equal probability of positive prediction). For a robust three relatively high-quality reviews [5, 6, 16], as well
review of fairness metrics and the algorithmic bias field at as one lower-quality review [4]. These four reviews
large, see Xu et al. and Huang et al., 2022 [5, 16]. Table S2 identified 16 unique tools for addressing bias that were
enumerates metric-specific bias reduction by cited study. advertised as open-source. The most frequently cited
Mackin et al. BMC Digital Health
Banerjee et al 2023 United States Healthcare: radiology Narrative equalized odds, calibrated equalized odds, reject option Aequitas, AIF360, Fairkit-Learn, Fairlearn, Fairness Com-
(2025) 3:26
classification, discrimination-aware ensemble parison, Fairness Measures, FairTest, Google What-If, IBM
Watson OpenScale, Themis-ML, VerifyML
Chaunzwa et al 2022 United States Healthcare: cancer Narrative subgroup-calibration NA
Du et al 2020 United States Data analytics Narrative calibrated distribution, calibrated equalized odds, trou- NA
bling neurons turn off
Huang et al 2022 United States Healthcare: general Scoping calibration, reject option classification, varying cutoffs Aequitas, AIF360, Fairlearn, Tensorflow
Liu et al 2023 Singapore Healthcare: general Narrative equalized odds, adjust sub-group thresholds, fine-tune NA
model according to subgroups
P Chen et al 2023 China Data analytics Systematic equalized odds, calibration, reject option classification, AIF360, Facebook Fairness Flow, Fairness Indicators
priority sampling, threshold optimization, regularization
Paulus & Kent 2020 United States Healthcare: general Narrative fairness constraints, differential threshold setting NA
R Chen et al 2023 United States Healthcare: general Narrative equalized odds, differential threshold setting, predictive NA
parity, calibration
Timmons et al 2023 United States Healthcare: mental health Narrative label flipping, individual and group debiasing, Laplacian NA
regularization, Gaussian randomized smoothing
Wang et al 2023 China Healthcare: general Scoping flipping decisions for equalized odds or equalized NA
opportunity, differentiating thresholds
Xu et al 2022 United States Healthcare: general Scoping calibrated/equalized odds, risk score adjustment AIF360, Fairlearn, FairML, FairMLHealth, Fairness-compar-
with parameterized monotonic function, rank adjust- ison, Fairness Indicators, MEASURES, ML-fairness-gym,
ment with dynamic programming procedure, causal themis-ml
pathway
Page 6 of 13
Mackin et al. BMC Digital Health (2025) 3:26 Page 7 of 13
of these tools were AIF360, cited by four reviews [4– sixteen platforms as having both bias identification and
6, 16], Google’s suite of Fairness Indicators / What-If mitigation functionality [5]. According to P. Chen et al.,
Tool, cited by four reviews [4–6, 16], Fairlearn, cited because of their nascent nature, there is “no mature
by three reviews [4, 5, 16], and Aequitas [4, 5], Fair- standard for how to quantify the risk of AI fairness”
ness Comparison [4, 16], and Themis ML [4, 16], each using these tools [6]. None of the four reviews reported
cited by two reviews. Huang et al. identified six of the
Mackin et al. BMC Digital Health (2025) 3:26 Page 8 of 13
whether the tools supported post-processing methods this space, all reviews were retained, although the major-
(versus pre- or in-processing methods) specifically. ity met < 50% of the JBI quality assessment criteria.
reducing bias. Third, we noted a lack of reporting on The studies in this review provided little justification for
changes to model accuracy resulting from implement- metrics chosen, and reviews provided no synthesized
ing post-processing methods. Fourth, we discovered a guidance on how to choose the right metric. While the
recurring theme of governance itself serving as a post- choice of appropriate group fairness metric is highly
processing bias mitigation tool. Fifth, we identified a lack context-dependent, minimal viable guidelines for health-
of guidance on how to use common software libraries care use-cases are needed in order to build an evidence
for post-processing bias mitigation. Sixth, we found a base of mitigation comparisons. Subgroup-specific per-
significant research-to-practice gap. These findings are formance reporting is now required for studies following
expounded upon below and accompanied by relevant the TRIPOD + AI framework, but fairness metrics and
areas for future research. standard thresholds of concern are not specified [50].
Some open-access resources, like the Aequitas fairness
Top post‑processing methods tree [51], have posited logic and visual aids for choos-
Threshold adjustment, reject option classification, ing the appropriate fairness metric for a given use case,
and calibration emerged as the most frequent meth- but research is needed to evaluate the use of such guid-
ods studied across the literature. These method families ance on real-world healthcare data and scenarios. Studies
are defined in Table 2 and provide a short list of tested which test multiple mitigation methods using a uniform
post-processing methods which can lower the barrier to set of fairness metrics would allow systems to make more
entry for implementing healthcare institutions. They also informed choices on mitigation strategy.
show real-world promise. Within these reviews, post-
processing methods improved equitable treatment by Insufficient accuracy reporting
race/ethnicity across identification of dementia (thresh- While bias mitigation is the focus of this review, reduc-
old adjustment, as seen in Gianatassio et al., 2020 [41]); tion in bias may require accuracy tradeoffs [35]. Eleven
care management resource allocation (reject option clas- out of sixteen cited studies reported some measure of
sification, as seen in Briggs et al., 2020 [22]), and cardio- accuracy loss due to bias mitigation, but heterogeneous
vascular disease prediction (calibration, as seen in Barda metrics for measuring accuracy at sub-group and global
et al., 2021 [39]). In Gianatassio et al. 2021, the authors levels made comparing the magnitudes of these losses
used threshold adjustment on a model predicting demen- difficult. Threshold adjustment was the most frequently
tia status to minimize bias, reducing differences in sen- reported method that uniformly reduced bias while
sitivity and specificity between race/ethnicity subgroups maintaining accuracy. The fairness-accuracy-tradeoff is
to less than three and five percentage points, respec- a focus of a growing portion of the healthcare debiasing
tively [41]. In Briggs et al., 2020, the authors employed literature, and more primary studies are needed to suf-
ROC on a risk prediction model used to aid treatment ficiently evaluate standardized measures of accuracy loss
plan decision-making in health systems across the U.S. in bias reduction [47].
Bias notably improved across a suite of fairness metrics,
with statistical parity difference changing from −0.13 to Governance as a post‑processing approach to mitigating
−0.06; disparate impact from 0.34 to 0.19; average odds bias
difference from −0.09 to −0.02; equal opportunity differ- This extended umbrella review sought to identify effec-
ence from 0.04 to 0.03; and mean fairness measure from tive post-processing statistical methods of bias reduction
0.15 to 0.09 [22]. In Barda et al., 2021, the authors used in healthcare algorithms. An unexpected theme that sur-
a calibration algorithm on a cardiovascular risk model to faced a posteriori in seven of eleven reviews was the role
reduce subgroup calibration slope variance from 1.179 to of governance itself as a post-processing approach to con-
0.019 [39]. All three authors discussed but did not quan- trolling bias. The need for an established regulatory sys-
tify the potential real-world impact of their work. tem specific to healthcare for governing fairness across
model development, deployment, and reporting was uni-
Fairness metric heterogeneity versally mentioned [5, 6, 9, 17, 30, 45, 47]. Reviews rec-
In investigating method effectiveness, we discovered ommended that clinicians and diverse stakeholders be
that the fairness metrics used to measure bias reduction involved in model building [17, 30] and that providers
varied greatly across studies. Choice of metric is cru- be included in developing clinically contextualized bias
cial, because all metrics cannot be optimized simulta- evaluation checklists [30, 47] like the Prediction Model
neously [33, 36, 49]. Choosing a less appropriate metric Risk of Bias Assessment Tool [5]. Reviews also suggested
for particular model use case can inflate perceived suc- that model documentation be published and the Fairness,
cess of bias mitigation, as in Foryciarz et al., 2022 [43]. Accountability, and Transparency in Machine Learning
Mackin et al. BMC Digital Health (2025) 3:26 Page 10 of 13
principles maintained throughout prediction projects a robust evidence base for their implementation, how-
[17, 45]. Centering human judgement at key junctures in ever, these methods – the majority of which emerged
the development and implementation processes (human- in the past four years – remain impractical for health
in-the-loop) was also discussed [9]. Collectively, these systems. Our review did identify a short-list of more
strategies were described as critical to ensuring that bias robustly tested, empiric post-processing methods. It did
was not introduced or exacerbated in any phase of the not, however, surface any findings on which of these
algorithmic lifecycle. Frameworks like Duke’s Algorithm- methods are best fit-for-purpose for implementing insti-
Based Clinical Decision Support (ABCDS) are emerg- tutions at different stages of algorithmic readiness and
ing to meet these calls, but more work remains to meet deployment [58]. Table S2 reports how bias was reduced
review recommendations [52]. in each cited study, but no cited study reported impacts
of this bias reduction after implementation on treatment
Barrier to entry for software library use or downstream care. Implementation research is needed
Four of eleven reviews identified a suite of software to understand how these methods reduce disparities in
libraries advertising bias mitigation functionalities, the practice. Post-processing methods may be more com-
most common being AIF360, Fairlearn, and Aequitas plex to implement than methods which address bias
[51, 53, 54]. These tools are Python repositories adver- during development and produce a single algorithm. For
tising functions that run a variety of statistical mitiga- example, threshold adjustment may require the devel-
tion approaches, including pre-, in-, and post-processing opment of multiple best practice alerts, tailored to each
techniques. However, no reviews compared these plat- risk threshold, if an EMR system does not allow users
forms for post-processing functionality specifically, to adjust alert firing levels by patient sociodemographic
nor did they evaluate implementation costs, resource factor. Post-processing mitigation may be cumbersome
requirements, or level of expertise needed to use each to address intersectional bias (requiring many thresh-
tool. Given the pace of innovation in machine learning old-specific adjustments) and is likely ill-equipped to
and predictive modeling, healthcare systems looking for handle complex classification problems. Implementa-
scalable solutions to accuracy and bias problems that do tion research must keep pace with innovation research
not require upskilling large healthcare workforces would to bridge the gap between theory and practice and safely
benefit from guidance for navigating the tools on the bring algorithmic bias mitigation out of academic data-
market. bases and into delivery of care. Robust standards, paired
Building technical workforce competency has been with specific, accessible pathways for bias monitoring
identified as a major barrier to AI implementation in the and mitigation, are crucial to “do no harm” in the age of
healthcare sector at large [55]. The healthcare workforce, predictive clinical analytics [59].
including both clinical and informatics staff, remains
behind the leading edge on the specialized programming Significance and practical implications
and computing competencies many machine learning Implementing predictive models in healthcare holds
applications on the market require [55–57]. Given that greater risks and ethical liabilities than in many other
health systems are already understaffed for training and fields because of the grave downstream consequences to
localizing predictive models, it is unsurprising that they patient safety and health of inaccurate, unfair decision-
are even further behind in their capacities to safeguard making. Biased predictions disproportionately impact
these models against biases [8]. Mitigating bias during historically underserved patients, who frequently bear
post-processing offers methods to improve bias without the burden of associated biased treatment [5, 9, 12]. This
massive skill investment. A comparative analysis of the risk of algorithmic bias threatens to further entrench sys-
post-processing methods offered by the software librar- temic health and healthcare disparities for low income
ies on the market is needed in order to make them more patients, uninsured patients, and patients outside of the
navigable for the end-user. Future research should con- dominant socioeconomic group [60]. Healthcare sys-
duct head-to-head comparisons of the post-processing tems cannot realistically tackle these biases until scal-
mitigation functions available within all open-source able methods are available that can be implemented with
packages and software libraries reported by reviews using existing system resources.
real-world data and comparable metrics to make them While bias reduction literature is rapidly expanding,
more accessible across health systems. there has been a lack of attention to methods which can
be feasibly implemented without developer expertise.
Research‑to‑practice gap This umbrella review is the first to our knowledge to
There were nine novel methods for post-processing miti- compare real-world effectiveness and accuracy of post-
gation found in our review [18, 31–34, 36–39]. Without processing methods across published reviews. Despite
Mackin et al. BMC Digital Health (2025) 3:26 Page 11 of 13
Author details
Conclusion Office of Population Health, New York City Health + Hospitals, New York, NY,
1
6. Chen P, Wu L, Wang L. AI fairness in data management and analytics: A 29. Riva JJMK, Burnie SJ, Endicott AR, Busse JW. What is your research ques-
review on challenges, methodologies and applications. Applied Sciences tion? An introduction to the PICOT format for clinicians. J Can Chiropr
(Switzerland). 2023;13(18). [Link] Assoc. 2012;56(3):167–71.
7. Thomasian NM, Eickhoff C, Adashi EY. Advancing health equity with 30. Liu M, Ning Y, Teixayavong S, Mertens M, Xu J, Ting DSW, et al. A
artificial intelligence. J Public Health Policy. 2021;42(4):602–11. [Link] translational perspective towards clinical AI fairness. NPJ Digit Med.
org/10.1057/s41271-021-00319-5. 2023;6(1):172. [Link]
8. Panch T, Mattie H, Atun R. Artificial intelligence and algorithmic bias: 31. Hardt M, Price E, Srebro N. Equality of opportunity in supervised learning.
implications for health systems. J Glob Health. 2019;9(2): 010318. [Link] Advances in neural information processing systems. arXiv:1610.02413.
doi.org/10.7189/jogh.09.020318. 2016;29. [Link]
9. Wang Y, Song Y, Ma Z, Han X. Multidisciplinary considerations of fairness 32. Lohia PK, Ramamurthy KN, Bhide M, Saha D, Varshney KR, Puri R, Bias miti-
in medical AI: A scoping review. Int J Med Inform. 2023;178. [Link] gation post-processing for individual and group fairness. In,. IEEE inter-
org/10.1016/j.ijmedinf.2023.105175. national conference on acoustics, speech and signal processing (icassp);
10. Epic. Epic’s cognitive computing models. Epic Userweb Galaxy. 2024. 2019 May 12–17; Brighton, UK. New York: IEEE. 2019;2019:2847–51.
[Link] 33. Mishler A, Kennedy EH, Chouldechova A, editors. Fairness in risk assess-
Galaxy-Redirect. Accessed 23 Dec 2024. ment instruments: Post-processing to achieve counterfactual equalized
11. Efthimiou O, Seo M, Chalkou K, Debray T, Egger M, Salanti G, et al. Devel- odds. In: Proceedings of the 2021 ACM Conference on Fairness, Account-
oping clinical prediction models: a step-by-step guide. BMJ. 2024;386: ability, and Transparency; (FAccT ’21); 2021 Mar; Virtual. New York (NY):
e078276. [Link] ACM; 2021. p. 386–400. [Link]
12. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias 34. Small E, Sokol K, Manning D, Salim FD, Chan J, editors. Equalised odds is
in an algorithm used to manage the health of populations. Science. not equal individual odds: Post-processing for group and individual fair-
2019;366(6464):447–53. [Link] [Link]: Proceedings of the 2024 ACM Conference on Fairness, Account-
13. Vorisek CN, Stellmach C, Mayer PJ, Klopfenstein SAI, Bures DM, Diehl A, ability, and Transparency (FAccT ’24); 2024 Apr 8–11; Rio de Janeiro,
et al. Artificial intelligence bias in health care: Web-based survey. J Med Brazil. New York (NY): Association for Computing Machinery; 2024. p.
Internet Res. 2023;25: e41089. [Link] 1559–1578. [Link]
14. Adams R, Henry K, Soleimani H, Rawat N, Saheed M, Chen E, et al. Assess- 35. Tahir A, Cheng L, Liu H, editors. Fairness through aleatoric uncertainty. In:
ing clinical use and performance of a machine learning sepsis alert for Proceedings of the 32nd ACM International Conference on Information
sex and racial bias. Crit Care Med. 2022;50(1):705. [Link] and Knowledge Management CIKM ’23); 2023 Oct 23–27; Birmingham,
01.ccm.0000811944.77042.17. United Kingdom. New York (NY): Association for Computing Machinery;
15. Edwards SE, Class QA, Ford CE, Alexander TA, Fleisher JD. Racial bias in 2023:2372–2381. [Link]
cesarean decision-making. Am J Obstet Gynecol MFM. 2023;5(5): 100927. 36. Jang T, Shi P, Wang X, editors. Group-aware threshold adaptation for fair
[Link] classification. In: Proceedings of the AAAI Conference on Artificial Intel-
16. Xu J, Xiao Y, Wang WH, Ning Y, Shenkman EA, Bian J, et al. Algorithmic fair- ligence; 2022;36(6):6988–6995. [Link]
ness in computational medicine EBioMedicine. 2022;84: 104250. [Link] 37. Kamiran F, Karim A, Zhang X, editors. Decision theory for discrimination-
doi.org/10.1016/j.ebiom.2022.104250. aware classification. In: Proceedings of the 2012 IEEE 12th International
17. Timmons AC, Duong JB, Simo Fiallo N, Lee T, Vo HPQ, Ahle MW, et al. A Conference on Data Mining (ICDM); 2012 Dec 10–13; Brussels, Belgium.
call to action on assessing and mitigating bias in artificial intelligence Piscataway (NJ): IEEE; 2012:924–929. [Link]
applications for mental health. Perspect Psychol Sci. 2023;18(5):1062–96. 45.
[Link] 38. Hsu B, Chen X, Han Y, Namkoong H, Basu K. An Operational Perspective
18. Nguyen D, Gupta S, Rana S, Shilton A, Venkatesh S. Fairness improvement to Fairness Interventions: Where and How to Intervene. arXiv preprint
for black-box classifiers with Gaussian process. Inf sci. 2021;576:542–56. arXiv:230201574. 2023.
[Link] 39. Barda N, Yona G, Rothblum GN, Greenland P, Leibowitz M, Balicer R, et al.
19. Putzel P, Lee S. Blackbox post-processing for multiclass fairness. arXiv Addressing bias in prediction models by improving subpopulation
preprint arXiv:220104461. 2022. calibration. J Am Med Inform Assoc. 2021;28(3):549–58. [Link]
20. Petersen F, Mukherjee D, Sun Y, Yurochkin M. Post-processing for indi- 10.1093/jamia/ocaa283.
vidual fairness. arXiv. [Link] Published Oct 40. Iosifidis V, Fetahu B, Ntoutsi E, editors. Fae: A fairness-aware ensemble
2021. Accessed 12 Feb 2025. framework. In: IEEE international conference on big data; 2019: IEEE.
21. Hort MCZ, Zhang JM, Harman M, Sarro F. Bias mitigation for machine [Link]
learning classifiers: a comprehensive survey. ACM J Responsible Comput- 41. Gianattasio KZ, Ciarleglio A, Power MC. Development of algorithmic
ing. 2023. [Link] dementia ascertainment for racial/ethnic disparities research in the US
22. Briggs E, Hollmén J, editors. Mitigating discrimination in clinical machine health and retirement study. Epidemiology. 2020;31(1):126–33. [Link]
learning decision support using algorithmic processing techniques. In: doi.org/10.1097/EDE.0000000000001101.
Discovery Science: 23rd International Conference, DS 2020, Thessaloniki, 42. Rodolfa KT, Lamba H, Ghani R. Empirical observation of negligible fair-
Greece, October 19–21, 2020, Proceedings. Cham (CH): Springer-Verlag; ness–accuracy trade-offs in machine learning for public policy. Nat Mach
2020:19–33. [Link] Intell. 2021;3(10):896–904. [Link]
23. Belbasis L, Bellou V, Ioannidis JPA. Conducting umbrella reviews BMJ Med. 43. Foryciarz A, Pfohl SR, Patel B, Shah N. Evaluating algorithmic fairness in
2022;1(1): e000071. [Link] the presence of clinical guidelines: the case of atherosclerotic cardiovas-
24. Aromataris E FR, Godfrey C, Holly C, Khalil H, Tungpunkom P. Chapter 10: cular disease risk estimation. BMJ Health Care Inform. 2022;29(1). [Link]
Umbrella Reviews. In: JBI Manual for Evidence Synthesis. JBI; 2020; [Link] doi.org/10.1136/bmjhci-2021-100460.
doi.org/10.46658/JBIMES-20-11. 44. Thompson HM, Sharma B, Bhalla S, Boley R, McCluskey C, Dligach D, et al.
25. Shaina Mackin, Remle Newton-Dame. Post-processing methods for miti- Bias and fairness assessment of a natural language processing opioid
gating algorithmic bias in healthcare applications: an umbrella review. misuse classifier: detection and mitigation of electronic health record
PROSPERO 2024 CRD42024490667 Available from: [Link] data disadvantages across racial subgroups. J Am Med Inform Assoc.
ac.uk/prospero/display_record.php?ID=CRD42024490667. Accessed 23 2021;28(11):2393–403. [Link]
Dec 2024. 45. Paulus JK, Kent DM. Predictably unequal: understanding and address-
26. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, ing concerns that algorithmic clinical prediction may increase
et al. The PRISMA 2020 statement: an updated guideline for reporting sys- health disparities. NPJ Digit Med. 2020;3:99. [Link]
tematic reviews. Syst Rev. 2021;10(1):89. Published 2021 Mar 29. [Link] s41746-020-0304-9.
doi.org/10.1186/s13643-021-01626-4. 46. Du M, Yang F, Zou N, Hu X. Fairness in deep learning: A computational
27. EndNote Team. EndNote. EndNote X9 ed. Philadelphia, PA: Clarivate; 2013. perspective. IEEE Intell Syst. 2021;36(4):25–34. [Link]
28. Veritas Health Innovation. Covidence systematic review software. Mel- MIS.2020.3000681.
bourne: Veritas Health Innovation; 2023. Available at www.covidence.org.
Mackin et al. BMC Digital Health (2025) 3:26 Page 13 of 13
47. Chen RJ, Wang JJ, Williamson DFK, Chen TY, Lipkova J, Lu MY, et al. Rumi Chunara Dr. Rumi Chunara is an Associate Professor of Bio-
Algorithmic fairness in artificial intelligence for medicine and health- statistics at the New York University School of Global Public Health
care. Nat Biomed Eng. 2023;7(6):719–42. [Link] and Associate Professor of Computer Science at the Tandon School
s41551-023-01056-8. of Engineering.
48. Chaunzwa TL, Del Rey MQ, Bitterman DS. Clinical informatics approaches
to understand and address cancer disparities. Yearb Med Inform.
2022;31(1):121–30. [Link] Remle Newton‑Dame Remle Newton-Dame is the Assistant Vice
49. Pleiss GRM, Wu F, Kleinberg J, Weinberger KQ. On fairness and calibration President of Healthcare Analytics in the Office of Population Health at
arXiv preprint. 2017. [Link] NYC Health + Hospitals.
50. Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of
a multivariable prediction model for individual prognosis or diagnosis
(TRIPOD): the TRIPOD Statement. BMJ. 2015;350: g7594. [Link]
10.1136/bmj.g7594.
51. Saleiro P KB, Hinkson L, London J, Stevens A, Anisfeld A, Rodolfa KT, et al.
Aequitas: A bias and fairness audit toolkit. arXiv preprint. 2019. Available
from: [Link] Accessed 23 Dec 2024.
52. Bedoya AD, Economou-Zavlanos NJ, Goldstein BA, et al. A framework for
the oversight and local deployment of safe and high-quality prediction
models. J Am Med Inform Assoc. 2022;29(9):1631–6. [Link]
1093/jamia/ocac078.
53. Trusted-AI. AIF360. GitHub. 2024. Published 2020. Updated December
2024. [Link] Accessed 23 Dec 2024.
54. Werts H, Dudík M, Edgar R, Jalali A, Lutz R, Madaio M. Fairlearn: Assessing
and improving fairness of AI systems. arXiv. 2023. Available from: [Link]
arxiv.org/abs/2303.16626.
55. Ahmed MI, Spooner B, Isherwood J, Lane M, Orrock E, Dennison A. A
systematic review of the barriers to the implementation of artificial intel-
ligence in healthcare. Cureus. 2023;15(10): e46454. [Link]
7759/cureus.46454.
56. Faes L, Wagner SK, Fu DJ, Liu X, Korot E, Ledsam JR, et al. Automated
deep learning design for medical image classification by health-care
professionals with no coding experience: a feasibility study. Lancet Digit
Health. 2019;1(5):e232–42. [Link]
6. Accessed 23 Dec 2024.
57. Mirin N, Mattie H, Jackson L, Samad Z, Chunara R. Data Science in Public
Health: Building Next Generation Capacity. Harvard Data Science Review.
2022;4(4). Available from: [Link]
58. Warty RR, Smith V, Salih M, Fox D, McArthur SL, Mol BW. Barriers to the
diffusion of medical technologies within healthcare: A systematic review.
IEEE Access. 2021;9:139043–58. [Link]
3118554.
59. van de Sande D, Chung EFF, Oosterhoff J, van Bommel J, Gommers D,
van Genderen ME. To warrant clinical adoption AI models require a multi-
faceted implementation evaluation. NPJ Digit Med. 2024;7(1):58. [Link]
doi.org/10.1038/s41746-024-01064-1.
60. Colon-Rodriguez CJ. Shedding light on healthcare algorithmic and
artificial intelligence bias: U.S. Department of Health and Human Services;
2023 Available from: [Link]
light-healthcare-algorithmic-and-artifi cial-intelligence-bias. Accessed 23
Dec 2024.
61. Papatheodorou SI, Evangelos E. Umbrella reviews: what they are and
why we need them. In: Veroniki A, Evangelou E, editors. Meta-Research:
Methods and Protocols. New York: Humana Press; 2022. p. 135–45.
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in pub-
lished maps and institutional affiliations.