0% found this document useful (0 votes)
12 views13 pages

Mitigating Algorithmic Bias in Healthcare

This umbrella review examines post-processing methods for mitigating algorithmic bias in binary healthcare classification models, highlighting the effectiveness of threshold adjustment, reject option classification, and calibration. The review identifies 11 relevant studies and notes that threshold adjustment showed significant promise in reducing bias, while the overall quality of the reviews was weak due to inadequate reporting. Future research is needed to empirically compare these methods using real-world healthcare data to enhance fairness in clinical decision-making.

Uploaded by

prince1110verma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views13 pages

Mitigating Algorithmic Bias in Healthcare

This umbrella review examines post-processing methods for mitigating algorithmic bias in binary healthcare classification models, highlighting the effectiveness of threshold adjustment, reject option classification, and calibration. The review identifies 11 relevant studies and notes that threshold adjustment showed significant promise in reducing bias, while the overall quality of the reviews was weak due to inadequate reporting. Future research is needed to empirically compare these methods using real-world healthcare data to enhance fairness in clinical decision-making.

Uploaded by

prince1110verma
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Mackin et al.

BMC Digital Health (2025) 3:26 BMC Digital Health


[Link]

RESEARCH Open Access

Post‑processing methods for mitigating


algorithmic bias in healthcare classification
models: An extended umbrella review
Shaina Mackin1*, Vincent J. Major2, Rumi Chunara3 and Remle Newton‑Dame1

Abstract
Background AI and predictive analytics have increased the speed of innovation in medicine. If left unchecked,
however, algorithmic bias can exacerbate health disparities across race, class, or gender. Early bias mitigation literature
has focused on addressing bias in the preparation and development phases of the algorithm life cycle (pre- and in-
processing). Post-processing methods, applied at the point of implementation, are less computationally intensive
and do not require re-building or training the model, allowing lower-resourced health systems to improve bias
in off-the-shelf binary classification models, which are increasingly common within electronic medical records. This
umbrella review sought to identify post-processing bias mitigation methods and tools applicable to binary healthcare
classification models in healthcare and summarize bias reduction effectiveness and accuracy loss.
Methods This review was registered with PROSPERO and reported according to PRISMA 2020. PubMed and Sco-
pus were searched in December 2023 for English-language reviews published post-2013 using an expanded search
string from previous work on machine learning bias. Eligibility criteria followed the PICOT framework. Reviews
were screened independently by two authors. Data were extracted from reviews using the Joanna Briggs Institute
Extraction Form for Review of Reviews, as well as from cited studies (hence, an “extended” umbrella review). Quality
was assessed using the Critical Appraisal Checklist for Systematic Reviews. Evidence was synthesized by mitigation
method and effectiveness.
Results Searches yielded 184 records. After duplicate removal, title/abstract, and full text screening, 11 reviews
were included, citing 16 eligible studies. Post-processing methods tested included threshold adjustment (9 studies,
cited by 8 reviews), reject option classification (6 studies, cited by 4 reviews), and calibration (5 studies, cited by 4
reviews). Threshold adjustment reduced bias across 8/9 trials; reject option classification and calibration reduced bias
in approximately half of trials (5/8 and 4/8). Results were reported with heterogeneous fairness and accuracy metrics,
making comparison difficult. A lack of effectiveness evaluation was noted across reviews. Four reviews identified
16 software libraries for addressing bias. Quality of the majority of reviews was weak due to inadequate reporting
on methods.

*Correspondence:
Shaina Mackin
mackins@[Link]
Full list of author information is available at the end of the article

© The Author(s) 2025. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0
International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long
as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if
you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or
parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated
otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not
permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To
view a copy of this licence, visit [Link]
Mackin et al. BMC Digital Health (2025) 3:26 Page 2 of 13

Conclusions Threshold adjustment showed significant promise in post-processing bias mitigation for healthcare
algorithms, followed by reject option classification and calibration. Future research should empirically compare
post-processing methods on binary classification models using real-world healthcare data. As commercial algorithms
proliferate, health systems require proven, achievable strategies to maximize fairness.
Keywords Algorithmic bias, Algorithmic fairness, Bias mitigation, Health equity

Background Statistical approaches to mitigation that can blunt or


In the next decade, artificial intelligence (AI) and predic- remove algorithmic bias do exist. Bias can be mitigated
tive analytics promise to improve cancer detection, tailor within three stages of the algorithmic development life-
drugs to patients, and supercharge the search for new cycle: pre-, in-, and post-processing. Pre-processing
cures [1–3]. However, these technologies leverage real- methods adjust the data before model development, and
world datasets which often reflect the biases of render- typically include resampling, reweighting, and relabeling
ing health systems and the societies in which they deliver [5, 16]. In-processing methods adjust bias during model
care. To speed innovation while advancing health equity, training and may include prejudice removers, regular-
clinical medicine must address this bias. Systems can izers, and adversarial debiasing [5, 16]. Post-processing
start by confronting algorithmic bias, in which a predic- methods adjust model outputs, attempting to improve
tive model developed on real-world data learns to make fairness after model training is complete [5, 16]. While
recommendations that create unfair differences in access pre- and in-processing methods are practical for model
to treatment or resources based on the protected charac- developers looking to reduce bias, post-processing meth-
teristics of an individual or group [4–7]. This threat risks ods do not require access to any underlying data used
exacerbating systemic inequities by race, class, or gender to train (or retrain), reducing environmental impact
[8]. If left unaddressed, algorithmic bias drives biased as well as cost and expertise barriers to implementa-
treatment in healthcare [9]. tion [17–20]. Because of their computational efficiency,
Risk prediction algorithms – binary classification mod- post-processing methods can be applied to a wider set of
els estimating the likelihood of a certain outcome – are use-cases – namely, the mitigation of bias from ‘off-the-
increasingly deployed for clinical decision support within shelf ’ algorithms, where the inner workings of the model
electronic medical records (EMRs) and third-party tools. are hidden or opaque, including ‘black-box’ models [17].
For example, Epic offers a library of over 40 “cognitive Post-processing methods also allow health systems who
computing models” developed on Epic data that indi- rely on commercial models or lack large data science
vidual systems can implement [10]. Clinical models such teams to reduce bias within classification algorithms in
as these, however, are increasingly flagged for algorith- their EMRs with relatively few resources.
mic bias [11]. A seminal 2019 study by Obermeyer et al. Despite the inherent ease and accessibility of post-
found that at a given risk score, “Black” patients were far processing methods, healthcare literature primarily
sicker than “White” patients, introducing disparities in describes pre- and in-processing methods crafted for
receipt of high-risk care management [12]. An acute kid- model developers [21]. There is a paucity of pragmatic
ney injury model, trained on a dataset of less than 10% literature evaluating the effectiveness of post-processing
females, performed better on male than female patients methods tailored to model implementers. While a num-
[13]. A Johns Hopkins’ sepsis model, in which bias could ber of post-processing methods exist, there is a scant
lead to inequitable prevention of septic shock, was found evidence base to guide selection and implementation
to have lower confirmation rates for “Black” patients than of them in clinical settings. Because of this, biomedical
“white” patients, with higher rates for “Asian” patients researchers often rely on methods developed outside
[14]. The Maternal–Fetal Medicine Unit’s predictive the healthcare sector, such as computer science [22].
model for successful vaginal delivery post-cesarean was With the objective of collating an accessible resource for
found to predict lower success rates for “Black” women, healthcare institutions interested in mitigating bias, this
biasing decisions made by providers and patients and umbrella review sought to identify (1) post-processing
contributing to the disproportionately high C-section bias mitigation methods used to date that are applica-
rates among Black patients [15]. Algorithmic bias is nei- ble to binary classification models in healthcare; (2) the
ther rare nor coincidental, and disparities in performance effectiveness of these methods at reducing bias; (3) any
of diagnostic and mortality risk prediction have been reported impacts of these methods on model accuracy;
replicated across additional care delivery settings and and (4) any open-source tools available for implementa-
conditions with grave consequences [5]. tion of these methods.
Mackin et al. BMC Digital Health (2025) 3:26 Page 3 of 13

Methods Reviews [24]. Variables included: title, author, year, coun-


Extended umbrella review try, field, review type, databases searched, number and
The decision to conduct an umbrella review was made date range of included studies, biases, fairness metrics,
because of the availability of existing, recent reviews post-processing methods, and mitigation tools. Quality
on algorithmic bias mitigation methods in healthcare, of reviews was assessed using the JBI Critical Appraisal
per Belbasis et al.’s guidelines (2022) [23]. The umbrella Checklist for systematic reviews and research syntheses
review was initiated according to standards from the [24]. The variables extracted from cited studies included
Joanna Briggs Institute [24]. During the review process, mitigation method, fairness metric(s), mitigation success
we found that reviews did not consistently report the (y/n), and any mitigation quantification.
effectiveness of methods from their included studies, fail-
ing to accomplish objectives 2 and 3. In order to extract Data synthesis
comparisons of effectiveness by mitigation method, this Extracted data were synthesized into summary tables
umbrella review was extended to include the full texts of based on quality of reviews, characteristics of reviews,
post-processing studies cited by included reviews, hence- characteristics of cited studies, and post-processing
forth referred to as “cited studies”. Objectives 1 and 4 methods tested. Frequencies of methods, fairness met-
were met by the umbrella review; objectives 2 and 3, by rics, and software libraries were also summarized.
the review of cited studies. The protocol for the review Cited studies that quantified the effectiveness of two or
was registered with the international prospective register more different post-processing bias mitigation meth-
of systematic reviews (PROSPERO CRD42024490667) ods were included multiple times. We categorized the
and reported according to the 2020 Preferred Reporting success of mitigating bias by method. Due to the het-
Items for Systematic Reviews and Meta-analyses [25, 26]. erogeneity in fairness metrics used across studies and
the absence of standard bias cutoffs within fairness
Search strategy and eligibility criteria metrics, however, no numeric gradations for compar-
PubMed and Scopus were searched in December 2023 for ing levels of bias reduction were identified. As such, we
English-language reviews of any type published between categorized reported changes in bias as follows: a) uni-
2013–2023. Search strings from Huang et al.’s 2022 semi- form bias reduction (improvement of bias by any mag-
nal review on bias mitigation methods in clinical models nitude across all reported fairness metrics; b) mixed bias
were expanded to include the compound terms “algorith- reduction (improvement of bias across some reported
mic bias”, “AI bias”, “algorithmic fairness”, and “AI fair- fairness metrics but not others); and c) no bias reduc-
ness” [5]. Complete search strings for both databases are tion (no improvement of bias across any fairness metric
presented in Supplementary Table 1 (Table S1). Results reported). No studies reported uniform worsening of bias
were imported to and deduplicated by EndNote (version by any method.
X9) and managed with Covidence software [27, 28]. Title/ It is established in the literature that optimizing fair-
abstracts and full texts were screened independently by ness by reducing bias can result in trade-offs with model
two reviewers (SM and RND). Conflicts were resolved accuracy [30]. Due to the heterogeneity in accuracy
through consensus. measures used across studies, however, no numeric gra-
Eligibility criteria were developed using the PICOT dations for comparing losses to accuracy were identified.
(population, intervention, comparison, outcome, time) Some studies reported numeric accuracy, some numeric
framework [29]. To maximize applicability to the pre- balanced accuracy, and others only qualitatively com-
dictive classification algorithms increasingly driving mented on the direction of accuracy change. As such,
decision support within mainstream electronic medical we categorized reported accuracy changes as follows: a)
records, this study excluded records which focused solely no loss (accuracy was not impacted according to study
on medical imaging, pharmacology, or genetics. Reviews authors); or b) low loss (accuracy decreased a slight, mar-
from any country with any underlying study populations ginal, or negligible amount according to study authors).
of any age, nationality, gender, or demographic were eligi- No high levels of accuracy loss were reported.
ble for inclusion. Reviews not addressing post-processing
methods were excluded. For cited studies, only methods Results
applicable to binary classification and measured by group Search results
fairness metrics were included. Searches yielded 184 eligible records. Figure 1 illustrates
the PRISMA flow diagram, including exclusion criteria
Data extraction and quality assessment for reviews and studies. After duplicate removal (n = 49),
Data were extracted by SM using the Joanna Briggs screening was completed for title/abstracts (n = 135) and
Institute (JBI) Extraction Form for Review of Systematic full texts (n = 25). Eleven full text reviews were retained.
Mackin et al. BMC Digital Health (2025) 3:26 Page 4 of 13

Fig. 1 PRISMA Flow Diagram for Extended Umbrella Review

Reasons for exclusion upon full text review included lack narrative. Notably, only one review provided a compre-
of post-processing methods or measurement (n = 12) and hensive list of all its included studies and their charac-
not being a review (n = 2). Out of the 11 retained reviews, teristics [5]. The 11 included reviews collectively cited
one was systematic, three were scoping, and seven were 45 studies when discussing post-processing methods.
Mackin et al. BMC Digital Health (2025) 3:26 Page 5 of 13

Upon consulting the texts of those 45 cited studies, only Cited studies consistently failed to provide standard a
16 (36%) actually tested and measured bias mitigation for priori thresholds of success for bias reduction across or
post-processing methods applicable to binary classifica- within fairness metrics. Reported rates of bias reduction
tion models in healthcare. These 16 studies, reporting 25 varied and were difficult to compare across studies due to
trials, were retained in the extended umbrella review, as the heterogeneity of fairness metrics.
seen in Fig. 1. Defining uniform fairness metric improvement of
any magnitude as successful bias reduction, threshold
Available methods: evidence from reviews and cited adjustment showed promise across cited studies. Thresh-
studies old adjustment showed uniform bias reduction in eight
Table 1 charts retained reviews according to their char- out of nine trials (88.9%). Reject option classification
acteristics, including field, review type, post-processing showed uniform bias reduction in five out of eight trials
methods discussed, and open source tools reported. (62.5%), and mixed bias reduction in two (25%). Calibra-
The most common post-processing methods tested tion showed uniform bias reduction in four out of eight
were variants of threshold adjustment (9 studies, cited by trials (50%) and mixed reduction in two (25%). Table 2
8 reviews), reject option classification (6 studies, cited by summarizes the effectiveness of each study by mitigation
4 reviews), and calibration (5 studies, cited by 4 reviews). method using the bias reduction categories defined in
While “equalized odds” (EO) was frequently cited as Methods. Figure 2 summarizes the effectiveness of bias
a mitigation method, EO is a fairness metric popular- reduction by method, across studies.
ized by Hardt et al. in a paper that also introduced one
method of classifier adjustment to reduce bias [31]. Stud- Fairness‑accuracy tradeoffs: evidence from cited studies
ies that cited “equalized odds” as a method were catego- Several cited studies measured success in terms of
rized within a family of threshold adjustment methods fairness-accuracy trade-off curves [36, 40]. Figure 3
[31–35]. Nine studies introduced novel post-processing shows a summary of the loss of accuracy across stud-
methods, which were included when meeting our other ies by method. Eleven of the sixteen cited studies meas-
screening criteria. These included four threshold adjust- ured accuracy loss during their post-processing bias
ment variants [31, 33, 34, 36], two reject option classifi- mitigation trials, although studies used heterogeneous
cations [32, 37], and three calibrations [18, 38, 39]. The sub-group and overall metrics. Five trials saw no accu-
remaining studies evaluated already established methods. racy loss as a result of mitigation [31, 32, 34, 37, 41,
Table 2 classifies studies by novel or empiric method and 42]. In fourteen trials the mitigation method resulted in
provides the definitions and frequencies of all method low losses to accuracy [18, 22, 32, 35, 36, 38]. No trials
categories cited: threshold adjustment, rejection option reported high losses in accuracy. Six trials did not report
classification, and calibration. on accuracy measurement, namely variants of calibration
where the fairness-accuracy tradeoff is less central [39,
Method effectiveness: evidence from cited studies 43, 44]. Threshold adjustment had four trials with no loss
Only one of the 11 included reviews, Huang et al., 2022, to accuracy, and three trials with low loss to accuracy.
reported measures of bias reduction from the primary Reject option classification had two trials with no loss to
studies it included [5]. Across the 16 retained studies accuracy, and five trials with low loss to accuracy. Cali-
cited by reviews, post-processing bias mitigation meth- bration had zero trials with no loss to accuracy and five
ods were evaluated across 25 trials. Bias reduction was trials with low loss to accuracy.
quantified for threshold adjustment, reject option clas- Threshold adjustment had four trials with uniform bias
sification, and calibration. Seventeen of these 25 evalu- reduction and no loss to accuracy. Reject option classifi-
ations (68%) showed uniform bias reduction of a given cation had two trials with uniform bias reduction and no
method. These results are classified and evaluated in Sup- loss to accuracy. Calibration had zero trials with uniform
plementary Table 2 (Table S2). bias reduction and no loss to accuracy. No standards
The 25 trials reported in 16 cited studies used hetero- regarding optimal fairness-accuracy tradeoff levels were
geneous metrics to measure fairness and bias reduction. reported.
Metrics included equal opportunity (false negative rate
or true positive rate parity), equalized odds (true positive Tools for implementation: evidence from reviews
rate and false positive rate parity), and demographic par- Software libraries for bias mitigation were a focus of
ity (equal probability of positive prediction). For a robust three relatively high-quality reviews [5, 6, 16], as well
review of fairness metrics and the algorithmic bias field at as one lower-quality review [4]. These four reviews
large, see Xu et al. and Huang et al., 2022 [5, 16]. Table S2 identified 16 unique tools for addressing bias that were
enumerates metric-specific bias reduction by cited study. advertised as open-source. The most frequently cited
Mackin et al. BMC Digital Health

Table 1 Characteristics of included reviews


Author Year Country Field Review type Post-processing methods discussed Open source tools reported

Banerjee et al 2023 United States Healthcare: radiology Narrative equalized odds, calibrated equalized odds, reject option Aequitas, AIF360, Fairkit-Learn, Fairlearn, Fairness Com-
(2025) 3:26

classification, discrimination-aware ensemble parison, Fairness Measures, FairTest, Google What-If, IBM
Watson OpenScale, Themis-ML, VerifyML
Chaunzwa et al 2022 United States Healthcare: cancer Narrative subgroup-calibration NA
Du et al 2020 United States Data analytics Narrative calibrated distribution, calibrated equalized odds, trou- NA
bling neurons turn off
Huang et al 2022 United States Healthcare: general Scoping calibration, reject option classification, varying cutoffs Aequitas, AIF360, Fairlearn, Tensorflow
Liu et al 2023 Singapore Healthcare: general Narrative equalized odds, adjust sub-group thresholds, fine-tune NA
model according to subgroups
P Chen et al 2023 China Data analytics Systematic equalized odds, calibration, reject option classification, AIF360, Facebook Fairness Flow, Fairness Indicators
priority sampling, threshold optimization, regularization
Paulus & Kent 2020 United States Healthcare: general Narrative fairness constraints, differential threshold setting NA
R Chen et al 2023 United States Healthcare: general Narrative equalized odds, differential threshold setting, predictive NA
parity, calibration
Timmons et al 2023 United States Healthcare: mental health Narrative label flipping, individual and group debiasing, Laplacian NA
regularization, Gaussian randomized smoothing
Wang et al 2023 China Healthcare: general Scoping flipping decisions for equalized odds or equalized NA
opportunity, differentiating thresholds
Xu et al 2022 United States Healthcare: general Scoping calibrated/equalized odds, risk score adjustment AIF360, Fairlearn, FairML, FairMLHealth, Fairness-compar-
with parameterized monotonic function, rank adjust- ison, Fairness Indicators, MEASURES, ML-fairness-gym,
ment with dynamic programming procedure, causal themis-ml
pathway
Page 6 of 13
Mackin et al. BMC Digital Health (2025) 3:26 Page 7 of 13

Table 2 Frequency of post-processing mitigation methods across cited studies


Method category Definition Bias reduction Cited studies (n)
Novel Empiric

Threshold adjustment Moving decision thresholds 8/9 uniform reduction 4 5


for model predictions by sub- 1/9 no reduction Hardt et al., Small et al., Jang Gianattasio et al., Rodolfa et al.,
groups to optimize fairness et al., Mishler et al. Thompson et al.a, Lohia et al.a,
and equalize error rates Tahir et al.a
Reject option classification Allowing the model 5/8 uniform reduction 2 4
to not classify cases nearest 2/8 mixed reduction Kamiran et al., Lohia et al.a Briggs et al., Tahir et al.a, Tubella
the decision boundary; then, 1/8 no reduction et al., Lohia et al.a, Nguyen et al.a
classifying those protected
cases favorably and those
privileged cases unfavorably
Calibration Recalibrating predictions 4/8 uniform reduction 3 3
for different subgroups 2/8 mixed reduction Hsu et al.a, Nguyen et al.a, Foryciarz et al., Thompson et al.a,
to more accurately represent 2/8 no reduction Barda et al. Hsu et al.a
the true probabilities of out-
comes
a
Study evaluated more than one post-processing bias mitigation method, as follow:
Hsu et al. evaluated four calibration methods
Lohia et al. evaluated one threshold adjustment and two reject option classification methods
Nguyen et al. evaluated one calibration and two reject option classification methods
Tahir et al. evaluated one reject option classification and one threshold adjustment method
Thompson et al. evaluated one calibration and one threshold adjustment method

Fig. 2 Effectiveness of Bias Reduction Across Studies, by Method

of these tools were AIF360, cited by four reviews [4– sixteen platforms as having both bias identification and
6, 16], Google’s suite of Fairness Indicators / What-If mitigation functionality [5]. According to P. Chen et al.,
Tool, cited by four reviews [4–6, 16], Fairlearn, cited because of their nascent nature, there is “no mature
by three reviews [4, 5, 16], and Aequitas [4, 5], Fair- standard for how to quantify the risk of AI fairness”
ness Comparison [4, 16], and Themis ML [4, 16], each using these tools [6]. None of the four reviews reported
cited by two reviews. Huang et al. identified six of the
Mackin et al. BMC Digital Health (2025) 3:26 Page 8 of 13

Fig. 3 Loss of Accuracy due to Mitigation Across Studies, by Method

whether the tools supported post-processing methods this space, all reviews were retained, although the major-
(versus pre- or in-processing methods) specifically. ity met < 50% of the JBI quality assessment criteria.

Strength of evidence Discussion


The JBI Critical Appraisal Checklist for Systematic Overview
Reviews was used to assess the quality of evidence [24]. To our knowledge, this is the first synthesis of reviews
Supplementary Table 3 (Table S3) includes full JBI evalu- across the algorithmic bias mitigation literature that
ation across all included reviews. Ten of the 11 JBI criteria specifically examined the performance of different post-
were relevant to our review. Item 9, likelihood of publica- processing methods in healthcare classification settings.
tion bias, is not recommended for use outside of quan- Our search reinforced that post-processing bias mitiga-
titative reviews. Six reviews clearly stated their review tion methods are less commonly covered in the literature
questions and provided appropriate recommendations than pre- and in-processing methods. Of 135 reviews
for policy and practice as well as directions for further screened, 11 discussed post-processing methods. These
research, meeting three of ten relevant JBI criteria (items 11 reviews cited 45 primary studies in reference to post-
1, 10, 11). These included Banerjee et al., 2023; Timmons processing techniques, of which only 16 actually tested
et al., 2023; Paulus & Kent, 2020; Liu et al., 2023; Du the effectiveness of post-processing methods aiming to
et al., 2020; and R. Chen et al., 2023 [4, 17, 30, 45–47]. In improve group fairness for binary classification models in
addition to meeting these criteria, Chaunzwa et al., 2022 healthcare. The quality of included reviews was hetero-
provided their eligibility and search strategies, meeting geneous, with six reviews scoring 3/10 on the JBI Criti-
five criteria (items 1,2,3,10,11) [48]. Wang et al., 2023 and cal Appraisal checklist. Despite these weaknesses, some
P. Chen et al., 2023 also reported on searching adequate promising findings emerged with important implications
databases, meeting six criteria (items 1,2,3,4,10,11) [6, for algorithmic equity in research and practice.
9]. Xu et al., 2022 met seven criteria by adding meth- Six key insights surfaced from the results of our
ods used to minimize errors in data extraction (items extended umbrella review. First, we identified the most
1,2,3,4,7,10,11) [16], and Huang et al., 2022 met eight cri- frequently reported methods used for post-processing
teria by providing all of the above plus appropriate meth- algorithmic bias mitigation across healthcare literature.
ods to combine studies (items 1,2,3,4,7,8,10,11) [5]. To Second, we found heterogeneity in the fairness met-
document the underdeveloped state of the literature in rics used for measuring the effectiveness of methods for
Mackin et al. BMC Digital Health (2025) 3:26 Page 9 of 13

reducing bias. Third, we noted a lack of reporting on The studies in this review provided little justification for
changes to model accuracy resulting from implement- metrics chosen, and reviews provided no synthesized
ing post-processing methods. Fourth, we discovered a guidance on how to choose the right metric. While the
recurring theme of governance itself serving as a post- choice of appropriate group fairness metric is highly
processing bias mitigation tool. Fifth, we identified a lack context-dependent, minimal viable guidelines for health-
of guidance on how to use common software libraries care use-cases are needed in order to build an evidence
for post-processing bias mitigation. Sixth, we found a base of mitigation comparisons. Subgroup-specific per-
significant research-to-practice gap. These findings are formance reporting is now required for studies following
expounded upon below and accompanied by relevant the TRIPOD + AI framework, but fairness metrics and
areas for future research. standard thresholds of concern are not specified [50].
Some open-access resources, like the Aequitas fairness
Top post‑processing methods tree [51], have posited logic and visual aids for choos-
Threshold adjustment, reject option classification, ing the appropriate fairness metric for a given use case,
and calibration emerged as the most frequent meth- but research is needed to evaluate the use of such guid-
ods studied across the literature. These method families ance on real-world healthcare data and scenarios. Studies
are defined in Table 2 and provide a short list of tested which test multiple mitigation methods using a uniform
post-processing methods which can lower the barrier to set of fairness metrics would allow systems to make more
entry for implementing healthcare institutions. They also informed choices on mitigation strategy.
show real-world promise. Within these reviews, post-
processing methods improved equitable treatment by Insufficient accuracy reporting
race/ethnicity across identification of dementia (thresh- While bias mitigation is the focus of this review, reduc-
old adjustment, as seen in Gianatassio et al., 2020 [41]); tion in bias may require accuracy tradeoffs [35]. Eleven
care management resource allocation (reject option clas- out of sixteen cited studies reported some measure of
sification, as seen in Briggs et al., 2020 [22]), and cardio- accuracy loss due to bias mitigation, but heterogeneous
vascular disease prediction (calibration, as seen in Barda metrics for measuring accuracy at sub-group and global
et al., 2021 [39]). In Gianatassio et al. 2021, the authors levels made comparing the magnitudes of these losses
used threshold adjustment on a model predicting demen- difficult. Threshold adjustment was the most frequently
tia status to minimize bias, reducing differences in sen- reported method that uniformly reduced bias while
sitivity and specificity between race/ethnicity subgroups maintaining accuracy. The fairness-accuracy-tradeoff is
to less than three and five percentage points, respec- a focus of a growing portion of the healthcare debiasing
tively [41]. In Briggs et al., 2020, the authors employed literature, and more primary studies are needed to suf-
ROC on a risk prediction model used to aid treatment ficiently evaluate standardized measures of accuracy loss
plan decision-making in health systems across the U.S. in bias reduction [47].
Bias notably improved across a suite of fairness metrics,
with statistical parity difference changing from −0.13 to Governance as a post‑processing approach to mitigating
−0.06; disparate impact from 0.34 to 0.19; average odds bias
difference from −0.09 to −0.02; equal opportunity differ- This extended umbrella review sought to identify effec-
ence from 0.04 to 0.03; and mean fairness measure from tive post-processing statistical methods of bias reduction
0.15 to 0.09 [22]. In Barda et al., 2021, the authors used in healthcare algorithms. An unexpected theme that sur-
a calibration algorithm on a cardiovascular risk model to faced a posteriori in seven of eleven reviews was the role
reduce subgroup calibration slope variance from 1.179 to of governance itself as a post-processing approach to con-
0.019 [39]. All three authors discussed but did not quan- trolling bias. The need for an established regulatory sys-
tify the potential real-world impact of their work. tem specific to healthcare for governing fairness across
model development, deployment, and reporting was uni-
Fairness metric heterogeneity versally mentioned [5, 6, 9, 17, 30, 45, 47]. Reviews rec-
In investigating method effectiveness, we discovered ommended that clinicians and diverse stakeholders be
that the fairness metrics used to measure bias reduction involved in model building [17, 30] and that providers
varied greatly across studies. Choice of metric is cru- be included in developing clinically contextualized bias
cial, because all metrics cannot be optimized simulta- evaluation checklists [30, 47] like the Prediction Model
neously [33, 36, 49]. Choosing a less appropriate metric Risk of Bias Assessment Tool [5]. Reviews also suggested
for particular model use case can inflate perceived suc- that model documentation be published and the Fairness,
cess of bias mitigation, as in Foryciarz et al., 2022 [43]. Accountability, and Transparency in Machine Learning
Mackin et al. BMC Digital Health (2025) 3:26 Page 10 of 13

principles maintained throughout prediction projects a robust evidence base for their implementation, how-
[17, 45]. Centering human judgement at key junctures in ever, these methods – the majority of which emerged
the development and implementation processes (human- in the past four years – remain impractical for health
in-the-loop) was also discussed [9]. Collectively, these systems. Our review did identify a short-list of more
strategies were described as critical to ensuring that bias robustly tested, empiric post-processing methods. It did
was not introduced or exacerbated in any phase of the not, however, surface any findings on which of these
algorithmic lifecycle. Frameworks like Duke’s Algorithm- methods are best fit-for-purpose for implementing insti-
Based Clinical Decision Support (ABCDS) are emerg- tutions at different stages of algorithmic readiness and
ing to meet these calls, but more work remains to meet deployment [58]. Table S2 reports how bias was reduced
review recommendations [52]. in each cited study, but no cited study reported impacts
of this bias reduction after implementation on treatment
Barrier to entry for software library use or downstream care. Implementation research is needed
Four of eleven reviews identified a suite of software to understand how these methods reduce disparities in
libraries advertising bias mitigation functionalities, the practice. Post-processing methods may be more com-
most common being AIF360, Fairlearn, and Aequitas plex to implement than methods which address bias
[51, 53, 54]. These tools are Python repositories adver- during development and produce a single algorithm. For
tising functions that run a variety of statistical mitiga- example, threshold adjustment may require the devel-
tion approaches, including pre-, in-, and post-processing opment of multiple best practice alerts, tailored to each
techniques. However, no reviews compared these plat- risk threshold, if an EMR system does not allow users
forms for post-processing functionality specifically, to adjust alert firing levels by patient sociodemographic
nor did they evaluate implementation costs, resource factor. Post-processing mitigation may be cumbersome
requirements, or level of expertise needed to use each to address intersectional bias (requiring many thresh-
tool. Given the pace of innovation in machine learning old-specific adjustments) and is likely ill-equipped to
and predictive modeling, healthcare systems looking for handle complex classification problems. Implementa-
scalable solutions to accuracy and bias problems that do tion research must keep pace with innovation research
not require upskilling large healthcare workforces would to bridge the gap between theory and practice and safely
benefit from guidance for navigating the tools on the bring algorithmic bias mitigation out of academic data-
market. bases and into delivery of care. Robust standards, paired
Building technical workforce competency has been with specific, accessible pathways for bias monitoring
identified as a major barrier to AI implementation in the and mitigation, are crucial to “do no harm” in the age of
healthcare sector at large [55]. The healthcare workforce, predictive clinical analytics [59].
including both clinical and informatics staff, remains
behind the leading edge on the specialized programming Significance and practical implications
and computing competencies many machine learning Implementing predictive models in healthcare holds
applications on the market require [55–57]. Given that greater risks and ethical liabilities than in many other
health systems are already understaffed for training and fields because of the grave downstream consequences to
localizing predictive models, it is unsurprising that they patient safety and health of inaccurate, unfair decision-
are even further behind in their capacities to safeguard making. Biased predictions disproportionately impact
these models against biases [8]. Mitigating bias during historically underserved patients, who frequently bear
post-processing offers methods to improve bias without the burden of associated biased treatment [5, 9, 12]. This
massive skill investment. A comparative analysis of the risk of algorithmic bias threatens to further entrench sys-
post-processing methods offered by the software librar- temic health and healthcare disparities for low income
ies on the market is needed in order to make them more patients, uninsured patients, and patients outside of the
navigable for the end-user. Future research should con- dominant socioeconomic group [60]. Healthcare sys-
duct head-to-head comparisons of the post-processing tems cannot realistically tackle these biases until scal-
mitigation functions available within all open-source able methods are available that can be implemented with
packages and software libraries reported by reviews using existing system resources.
real-world data and comparable metrics to make them While bias reduction literature is rapidly expanding,
more accessible across health systems. there has been a lack of attention to methods which can
be feasibly implemented without developer expertise.
Research‑to‑practice gap This umbrella review is the first to our knowledge to
There were nine novel methods for post-processing miti- compare real-world effectiveness and accuracy of post-
gation found in our review [18, 31–34, 36–39]. Without processing methods across published reviews. Despite
Mackin et al. BMC Digital Health (2025) 3:26 Page 11 of 13

the observed heterogeneity of fairness metrics, threshold Supplementary Information


adjustment and reject option classification emerged as The online version contains supplementary material available at [Link]
viable routes to bias reduction for healthcare classifica- org/​10.​1186/​s44247-​025-​00166-4.
tion algorithms without drastic accuracy costs. We hope
Supplementary Material 1.
that our study serves as a resource for low-income, safety
net hospitals and all systems looking to tackle algorith-
mic bias and safeguard patients from algorithmic harm. Acknowledgements
The authors wish to acknowledge New York City Health + Hospitals and New
York University for their support throughout this project.
Limitations
Authors’ contributions
This study had important limitations. First, while we SM and RND wrote the protocol, conducted title and abstract review and full
conducted an “extended” umbrella review that included text review, analyzed the data, and wrote the main manuscript. Additionally,
SM searched the literature, extracted the data, assessed the quality of reviews,
both reviews and their cited studies to capture post-pro- synthesized the data, and prepared all tables and figures. VJM analyzed the
cessing methods, we only assessed the quality of reviews. data, collaborated on Tables 2 and S2, and contributed to the main manu-
Examination of the quality of each of the cited studies in script. RC contributed to the interpretation of results and main manuscript. All
authors reviewed the manuscript.
the reviews fell outside the scope of an umbrella review
and was not reported [61]. Second, we limited our review Funding
search strategy to algorithms deployed in healthcare This research was funded by Schmidt Futures with Rockefeller Philanthropy
Advisors as fiscal sponsor, grant ID G-23–2135530. The funder played no role in
settings and thus automatically excluded reviews from study design, data collection, analysis and interpretation of data, or the writing
other computer science disciplines that may have more of this manuscript.
robustly developed post-processing methods. This nar-
Data availability
rowing of our scope was intentional, as healthcare models No primary data was generated. All data analyzed during this study are
and patient data have unique needs. Third, our extended included in this published article and its supplementary information files.
umbrella review was limited to include only the methods
that have been cited within already published reviews. Declarations
Any methods not yet covered in published reviews were
Ethics approval and consent to participate
not included, which may have overlooked some emerg- Not applicable.
ing and novel methods. Relatedly, the small sample size
and mixed quality of included studies is a fundamental Consent for publication
Not applicable.
limitation of this review, but this limitation is itself an
important finding, and one that underscores our recom- Competing interests
mendation for more robust research in this space. The authors declare no competing interests.

Author details
Conclusion Office of Population Health, New York City Health + Hospitals, New York, NY,
1

USA. 2 Department of Population Health, NYU Grossman School of Medicine,


Our extended umbrella review provides evidence that New York, NY, USA. 3 Center for Health Data Science, New York University, New
post-processing methods can successfully mitigate algo- York, NY, USA.
rithmic bias in healthcare classification models. Hetero-
Received: 23 December 2024 Accepted: 16 April 2025
geneity in fairness metrics made quantification of bias
reduction difficult. When looking for significant bias
reduction across any fairness metric, threshold adjust-
ment emerged as the most effective post-processing References
method followed by reject option classification and cali- 1. Nierengarten MB, New AI model shows promise for cancer diagnosis.
bration. Further research is needed to investigate the Cancer. 2025;131(3). [Link]
2. Taherdoost H, Ghofrani A. AI’s role in revolutionizing personalized
scalability of these methods and to quantify accuracy medicine by reshaping pharmacogenomics and drug therapy. Intelligent
losses by approach. Implementation-focused research Pharmacy. 2024;2(5):643–50. [Link]
comparing open source tools like AIF360 and Aequitas 3. Chopra H, Annu, Shin DK, Munjal K, Priyanka, Dhama K, et al. Revolutioniz-
ing clinical trials: the role of AI in accelerating medical breakthroughs. Int
could improve translation and speed adoption. As com- J Surg. 2023;109(12):4211–4220. [Link]
mercial algorithms proliferate, lower-resourced health 000705.
systems who rely on off-the-shelf risk models in their 4. Banerjee I, Bhattacharjee K, Burns JL, Trivedi H, Purkayastha S, Seyyed-
Kalantari L, et al. “Shortcuts” causing bias in radiology artificial ntelligence:
EMRs can benefit from the proven, accessible bias miti- Causes, evaluation, and mitigation. J Am Coll of Radiol. 2023;20(9):842–51.
gation strategies identified in this review to achieve algo- [Link]
rithmic equity in clinical practice. 5. Huang J, Galal G, Etemadi M, Vaidyanathan M. Evaluation and mitigation
of racial bias in clinical machine learning models: Scoping review. JMIR
Medical Informatics. 2022;10(5). [Link]
Mackin et al. BMC Digital Health (2025) 3:26 Page 12 of 13

6. Chen P, Wu L, Wang L. AI fairness in data management and analytics: A 29. Riva JJMK, Burnie SJ, Endicott AR, Busse JW. What is your research ques-
review on challenges, methodologies and applications. Applied Sciences tion? An introduction to the PICOT format for clinicians. J Can Chiropr
(Switzerland). 2023;13(18). [Link] Assoc. 2012;56(3):167–71.
7. Thomasian NM, Eickhoff C, Adashi EY. Advancing health equity with 30. Liu M, Ning Y, Teixayavong S, Mertens M, Xu J, Ting DSW, et al. A
artificial intelligence. J Public Health Policy. 2021;42(4):602–11. [Link] translational perspective towards clinical AI fairness. NPJ Digit Med.
org/​10.​1057/​s41271-​021-​00319-5. 2023;6(1):172. [Link]
8. Panch T, Mattie H, Atun R. Artificial intelligence and algorithmic bias: 31. Hardt M, Price E, Srebro N. Equality of opportunity in supervised learning.
implications for health systems. J Glob Health. 2019;9(2): 010318. [Link] Advances in neural information processing systems. arXiv:​1610.​02413.
doi.​org/​10.​7189/​jogh.​09.​020318. 2016;29. [Link]
9. Wang Y, Song Y, Ma Z, Han X. Multidisciplinary considerations of fairness 32. Lohia PK, Ramamurthy KN, Bhide M, Saha D, Varshney KR, Puri R, Bias miti-
in medical AI: A scoping review. Int J Med Inform. 2023;178. [Link] gation post-processing for individual and group fairness. In,. IEEE inter-
org/​10.​1016/j.​ijmed​inf.​2023.​105175. national conference on acoustics, speech and signal processing (icassp);
10. Epic. Epic’s cognitive computing models. Epic Userweb Galaxy. 2024. 2019 May 12–17; Brighton, UK. New York: IEEE. 2019;2019:2847–51.
[Link] 33. Mishler A, Kennedy EH, Chouldechova A, editors. Fairness in risk assess-
Galaxy-​Redir​ect. Accessed 23 Dec 2024. ment instruments: Post-processing to achieve counterfactual equalized
11. Efthimiou O, Seo M, Chalkou K, Debray T, Egger M, Salanti G, et al. Devel- odds. In: Proceedings of the 2021 ACM Conference on Fairness, Account-
oping clinical prediction models: a step-by-step guide. BMJ. 2024;386: ability, and Transparency; (FAccT ’21); 2021 Mar; Virtual. New York (NY):
e078276. [Link] ACM; 2021. p. 386–400. [Link]
12. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias 34. Small E, Sokol K, Manning D, Salim FD, Chan J, editors. Equalised odds is
in an algorithm used to manage the health of populations. Science. not equal individual odds: Post-processing for group and individual fair-
2019;366(6464):447–53. [Link] [Link]: Proceedings of the 2024 ACM Conference on Fairness, Account-
13. Vorisek CN, Stellmach C, Mayer PJ, Klopfenstein SAI, Bures DM, Diehl A, ability, and Transparency (FAccT ’24); 2024 Apr 8–11; Rio de Janeiro,
et al. Artificial intelligence bias in health care: Web-based survey. J Med Brazil. New York (NY): Association for Computing Machinery; 2024. p.
Internet Res. 2023;25: e41089. [Link] 1559–1578. [Link]
14. Adams R, Henry K, Soleimani H, Rawat N, Saheed M, Chen E, et al. Assess- 35. Tahir A, Cheng L, Liu H, editors. Fairness through aleatoric uncertainty. In:
ing clinical use and performance of a machine learning sepsis alert for Proceedings of the 32nd ACM International Conference on Information
sex and racial bias. Crit Care Med. 2022;50(1):705. [Link] and Knowledge Management CIKM ’23); 2023 Oct 23–27; Birmingham,
01.​ccm.​00008​11944.​77042.​17. United Kingdom. New York (NY): Association for Computing Machinery;
15. Edwards SE, Class QA, Ford CE, Alexander TA, Fleisher JD. Racial bias in 2023:2372–2381. [Link]
cesarean decision-making. Am J Obstet Gynecol MFM. 2023;5(5): 100927. 36. Jang T, Shi P, Wang X, editors. Group-aware threshold adaptation for fair
[Link] classification. In: Proceedings of the AAAI Conference on Artificial Intel-
16. Xu J, Xiao Y, Wang WH, Ning Y, Shenkman EA, Bian J, et al. Algorithmic fair- ligence; 2022;36(6):6988–6995. [Link]
ness in computational medicine EBioMedicine. 2022;84: 104250. [Link] 37. Kamiran F, Karim A, Zhang X, editors. Decision theory for discrimination-
doi.​org/​10.​1016/j.​ebiom.​2022.​104250. aware classification. In: Proceedings of the 2012 IEEE 12th International
17. Timmons AC, Duong JB, Simo Fiallo N, Lee T, Vo HPQ, Ahle MW, et al. A Conference on Data Mining (ICDM); 2012 Dec 10–13; Brussels, Belgium.
call to action on assessing and mitigating bias in artificial intelligence Piscataway (NJ): IEEE; 2012:924–929. [Link]
applications for mental health. Perspect Psychol Sci. 2023;18(5):1062–96. 45.
[Link] 38. Hsu B, Chen X, Han Y, Namkoong H, Basu K. An Operational Perspective
18. Nguyen D, Gupta S, Rana S, Shilton A, Venkatesh S. Fairness improvement to Fairness Interventions: Where and How to Intervene. arXiv preprint
for black-box classifiers with Gaussian process. Inf sci. 2021;576:542–56. arXiv:230201574. 2023.
[Link] 39. Barda N, Yona G, Rothblum GN, Greenland P, Leibowitz M, Balicer R, et al.
19. Putzel P, Lee S. Blackbox post-processing for multiclass fairness. arXiv Addressing bias in prediction models by improving subpopulation
preprint arXiv:220104461. 2022. calibration. J Am Med Inform Assoc. 2021;28(3):549–58. [Link]
20. Petersen F, Mukherjee D, Sun Y, Yurochkin M. Post-processing for indi- 10.​1093/​jamia/​ocaa2​83.
vidual fairness. arXiv. [Link] Published Oct 40. Iosifidis V, Fetahu B, Ntoutsi E, editors. Fae: A fairness-aware ensemble
2021. Accessed 12 Feb 2025. framework. In: IEEE international conference on big data; 2019: IEEE.
21. Hort MCZ, Zhang JM, Harman M, Sarro F. Bias mitigation for machine [Link]
learning classifiers: a comprehensive survey. ACM J Responsible Comput- 41. Gianattasio KZ, Ciarleglio A, Power MC. Development of algorithmic
ing. 2023. [Link] dementia ascertainment for racial/ethnic disparities research in the US
22. Briggs E, Hollmén J, editors. Mitigating discrimination in clinical machine health and retirement study. Epidemiology. 2020;31(1):126–33. [Link]
learning decision support using algorithmic processing techniques. In: doi.​org/​10.​1097/​EDE.​00000​00000​001101.
Discovery Science: 23rd International Conference, DS 2020, Thessaloniki, 42. Rodolfa KT, Lamba H, Ghani R. Empirical observation of negligible fair-
Greece, October 19–21, 2020, Proceedings. Cham (CH): Springer-Verlag; ness–accuracy trade-offs in machine learning for public policy. Nat Mach
2020:19–33. [Link] Intell. 2021;3(10):896–904. [Link]
23. Belbasis L, Bellou V, Ioannidis JPA. Conducting umbrella reviews BMJ Med. 43. Foryciarz A, Pfohl SR, Patel B, Shah N. Evaluating algorithmic fairness in
2022;1(1): e000071. [Link] the presence of clinical guidelines: the case of atherosclerotic cardiovas-
24. Aromataris E FR, Godfrey C, Holly C, Khalil H, Tungpunkom P. Chapter 10: cular disease risk estimation. BMJ Health Care Inform. 2022;29(1). [Link]
Umbrella Reviews. In: JBI Manual for Evidence Synthesis. JBI; 2020; [Link] doi.​org/​10.​1136/​bmjhci-​2021-​100460.
doi.​org/​10.​46658/​JBIMES-​20-​11. 44. Thompson HM, Sharma B, Bhalla S, Boley R, McCluskey C, Dligach D, et al.
25. Shaina Mackin, Remle Newton-Dame. Post-processing methods for miti- Bias and fairness assessment of a natural language processing opioid
gating algorithmic bias in healthcare applications: an umbrella review. misuse classifier: detection and mitigation of electronic health record
PROSPERO 2024 CRD42024490667 Available from: [Link] data disadvantages across racial subgroups. J Am Med Inform Assoc.
ac.​uk/​prosp​ero/​displ​ay_​record.​php?​ID=​CRD42​02449​0667. Accessed 23 2021;28(11):2393–403. [Link]
Dec 2024. 45. Paulus JK, Kent DM. Predictably unequal: understanding and address-
26. Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, ing concerns that algorithmic clinical prediction may increase
et al. The PRISMA 2020 statement: an updated guideline for reporting sys- health disparities. NPJ Digit Med. 2020;3:99. [Link]
tematic reviews. Syst Rev. 2021;10(1):89. Published 2021 Mar 29. [Link] s41746-​020-​0304-9.
doi.​org/​10.​1186/​s13643-​021-​01626-4. 46. Du M, Yang F, Zou N, Hu X. Fairness in deep learning: A computational
27. EndNote Team. EndNote. EndNote X9 ed. Philadelphia, PA: Clarivate; 2013. perspective. IEEE Intell Syst. 2021;36(4):25–34. [Link]
28. Veritas Health Innovation. Covidence systematic review software. Mel- MIS.​2020.​30006​81.
bourne: Veritas Health Innovation; 2023. Available at www.​covid​ence.​org.
Mackin et al. BMC Digital Health (2025) 3:26 Page 13 of 13

47. Chen RJ, Wang JJ, Williamson DFK, Chen TY, Lipkova J, Lu MY, et al. Rumi Chunara Dr. Rumi Chunara is an Associate Professor of Bio-
Algorithmic fairness in artificial intelligence for medicine and health- statistics at the New York University School of Global Public Health
care. Nat Biomed Eng. 2023;7(6):719–42. [Link] and Associate Professor of Computer Science at the Tandon School
s41551-​023-​01056-8. of Engineering.
48. Chaunzwa TL, Del Rey MQ, Bitterman DS. Clinical informatics approaches
to understand and address cancer disparities. Yearb Med Inform.
2022;31(1):121–30. [Link] Remle Newton‑Dame Remle Newton-Dame is the Assistant Vice
49. Pleiss GRM, Wu F, Kleinberg J, Weinberger KQ. On fairness and calibration President of Healthcare Analytics in the Office of Population Health at
arXiv preprint. 2017. [Link] NYC Health + Hospitals.
50. Collins GS, Reitsma JB, Altman DG, Moons KG. Transparent reporting of
a multivariable prediction model for individual prognosis or diagnosis
(TRIPOD): the TRIPOD Statement. BMJ. 2015;350: g7594. [Link]
10.​1136/​bmj.​g7594.
51. Saleiro P KB, Hinkson L, London J, Stevens A, Anisfeld A, Rodolfa KT, et al.
Aequitas: A bias and fairness audit toolkit. arXiv preprint. 2019. Available
from: [Link] Accessed 23 Dec 2024.
52. Bedoya AD, Economou-Zavlanos NJ, Goldstein BA, et al. A framework for
the oversight and local deployment of safe and high-quality prediction
models. J Am Med Inform Assoc. 2022;29(9):1631–6. [Link]
1093/​jamia/​ocac0​78.
53. Trusted-AI. AIF360. GitHub. 2024. Published 2020. Updated December
2024. [Link] Accessed 23 Dec 2024.
54. Werts H, Dudík M, Edgar R, Jalali A, Lutz R, Madaio M. Fairlearn: Assessing
and improving fairness of AI systems. arXiv. 2023. Available from: [Link]
arxiv.​org/​abs/​2303.​16626.
55. Ahmed MI, Spooner B, Isherwood J, Lane M, Orrock E, Dennison A. A
systematic review of the barriers to the implementation of artificial intel-
ligence in healthcare. Cureus. 2023;15(10): e46454. [Link]
7759/​cureus.​46454.
56. Faes L, Wagner SK, Fu DJ, Liu X, Korot E, Ledsam JR, et al. Automated
deep learning design for medical image classification by health-care
professionals with no coding experience: a feasibility study. Lancet Digit
Health. 2019;1(5):e232–42. [Link]
6. Accessed 23 Dec 2024.
57. Mirin N, Mattie H, Jackson L, Samad Z, Chunara R. Data Science in Public
Health: Building Next Generation Capacity. Harvard Data Science Review.
2022;4(4). Available from: [Link]
58. Warty RR, Smith V, Salih M, Fox D, McArthur SL, Mol BW. Barriers to the
diffusion of medical technologies within healthcare: A systematic review.
IEEE Access. 2021;9:139043–58. [Link]
31185​54.
59. van de Sande D, Chung EFF, Oosterhoff J, van Bommel J, Gommers D,
van Genderen ME. To warrant clinical adoption AI models require a multi-
faceted implementation evaluation. NPJ Digit Med. 2024;7(1):58. [Link]
doi.​org/​10.​1038/​s41746-​024-​01064-1.
60. Colon-Rodriguez CJ. Shedding light on healthcare algorithmic and
artificial intelligence bias: U.S. Department of Health and Human Services;
2023 Available from: [Link]
light-​healt​hcare-​algor​ithmic-​and-​artif​i cial-​intel​ligen​ce-​bias. Accessed 23
Dec 2024.
61. Papatheodorou SI, Evangelos E. Umbrella reviews: what they are and
why we need them. In: Veroniki A, Evangelou E, editors. Meta-Research:
Methods and Protocols. New York: Humana Press; 2022. p. 135–45.

Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in pub-
lished maps and institutional affiliations.

Shaina Mackin Shaina Mackin is a Research Manager in Healthcare


Analytics in the Office of Population Health at NYC Health + Hospi-
tals, the largest municipal safety net healthcare system in the United
States.

Vincent J. Major Dr. Vincent Major is an Assistant Professor in the


Department of Population Health at NYU Grossman School of Medi-
cine and the Associate Director of the Division of Applied AI and
Technologies in the Department of Health Informatics.

You might also like