Algorithm
Algorithm
Digital health technologies can generate data that can be used to train artificial intelligence (AI) algorithms, which Lancet Digit Health 2024;
have been particularly transformative in cardiovascular health-care delivery. However, digital and health-care data 6: e749–54
repositories that are used to train AI algorithms can introduce bias when data are homogeneous and health-care Published Online
August 29, 2024
processes are inequitable. AI bias can also be introduced during algorithm development, testing, implementation,
[Link]
and post-implementation processes. The consequences of AI algorithmic bias can be considerable, including missed S2589-7500(24)00155-9
diagnoses, misclassification of disease, incorrect risk prediction, and inappropriate treatment recommendations. This is the second in a Series of
This bias can disproportionately affect marginalised demographic groups. In this Series paper, we provide a brief four papers about artificial
overview of AI applications in cardiovascular health care, discuss stages of algorithm development and associated intelligence and digital
innovations in cardiovascular
sources of bias, and provide examples of harm from biased algorithms. We propose strategies that can be applied
care. All papers in the Series are
during the training, testing, and implementation of AI algorithms to mitigate bias so that all those at risk for or living available at [Link]/
with cardiovascular disease might benefit equally from AI. series/AI-and-digital-tools-in-
cardiovascular-care.
Introduction The risk of AI bias in health care has come to attention Department of Medicine,
Cardiovascular disease is a leading cause of death and in recent years, and there have been several publications— Faculty of Health Sciences,
McMaster University,
disability globally. More than most other health conditions, largely in the grey literature—describing select sources Hamilton, ON, Canada
cardiovascular disease relies heavily on multi-modality of biases, possible downstream effects, and strategies to (A Mihan MPH,
diagnostics and digital health technology for diagnosis, mitigate them.8–13 This Viewpoint consolidates this litera- H G C Van Spall MD); Cardiology
Division, University of Texas
risk prediction, and treatment.1 Big data serves as a reposi- ture and focuses on cardiovascular care, where AI bias
Southwestern Medical Center,
tory to train artificial intelligence (AI) algorithms to could have grave consequences; several cardiovascular Dallas, TX, USA (A Pandey MD);
improve care. The health-care applications of AI in cardi- diagnoses are high risk and missing or undertreating Baim Institute for Clinical
ology include identifying risk factors, predicting disease them by virtue of biased algorithms could result in death Research, Boston, MA, USA
(H G C Van Spall)
outcomes, interpreting diagnostic imaging, generating or disability. Thus, we focus on the causes and conse-
treatment recommendations, and guiding health-care quences of AI bias in cardiovascular disease, and propose Correspondence to:
Harriette GC Van Spall,
resource allocation.2,3 AI-integrated clinical decision strategies to mitigate them with the goal of achieving Department of Medicine, Faculty
support tools can also enhance care in underserved cardiovascular health equity. of Health Sciences, McMaster
regions.4 Thus, AI can bridge gaps in care and generate University, Hamilton,
data that can further improve care.4 Improving cardiovascular care with digital ON L8L 0A3, Canada
[Link]@[Link]
Notwithstanding the potential of AI, there are disparities health technology and AI
in access to health-care services, diagnostics, and digital Digital health technologies, including online education
health technologies that disproportionately affect socioeco- tools, electronic health or pharmacy records, external wear-
nomically deprived, rural, and ethnic minority able and implantable cardiac devices, digitised diagnostic
populations,1,5 resulting in their under-representation or data, telemedicine visits, mobile applications, and AI, have
misrepresentation in digital datasets—ie, electronic reposi- transformed cardiovascular care.2–5,14–16 Using existing
tories of data in a format that machines can read, interpret, digital datasets, AI algorithms can be trained to analyse
or analyse.6 Furthermore, there are long-standing inequi- patterns of disease progression, classify diseases, predict
ties in the diagnosis, risk stratification, and treatment clinical events, generate differential diagnoses, and
of marginalised groups.1 Training AI algorithms on data- generate treatment recommendations.2 Machine learning
sets that under-represent specific demographic groups or algorithms can use single-lead electrocardiograms (ECGs)
reflect biased clinical decision making could contribute to in wearable devices to detect cardiomyopathies or heart
AI bias.6 In addition, there can be biases in the algorithm failure.15 Using changes in voice and activity, machine-
development, implementation, and post-implementation learning algorithms within some digital applications can
practices that inform iterative AI learning.7–9 Such biases also detect decompensation of conditions, such as heart
can result in inaccurate diagnoses, risk prediction, disease failure.14 Automated analysis of diagnostics, such as ECGs,
classification, or treatment recommendations, particularly chest x-rays, echocardiograms, or CT scans can increase
in those who already face barriers and disparities in health the speed of cardiovascular diagnostic interpretation by
care.10 clinicians.3 AI-based applications, such as RadTranslate,
Identification of the problem that the algorithm will solve Panel: Types of artificial intelligence bias7,8,19,20
• Algorithmic bias: results from the algorithm itself and
Selection of datasets to train the algorithm includes systematic, intrinsic, and repetitive errors that
cause unfair outputs
• Annotation bias: results from those who are annotating
Management of the data used for algorithm training
data such that different labels are applied to demographic
groups
Development and training of the algorithm • Content production bias: caused from structural, lexical,
semantic, and syntactic differences in contents generated
by users based on different demographic factors
Testing or validation of the algorithm
• Evaluation bias: caused by using inappropriate evaluation
criteria when assessing applications
Implementation of the algorithm • Historical bias: a type of societal bias that currently exists
and can be exemplified in current data collection
• Latent bias: can develop over time despite an initially fair
Selection of outcomes to evaluate performance artificial intelligence algorithm
• Measurement bias: results from how a factor is selected,
Ongoing evaluation of algorithm performance analysed, and measured
• Outcome labelling bias: occurs when an outcome is not
Figure 1: Bias-prone steps in algorithm development and implementation objectively defined or detected across groups
• Representation bias: occurs when the sample in the
training or testing dataset has different characteristics
can allow patients to communicate in their language than the population in which the algorithm will be used
of preference, reducing the risk of medical complications • Sampling bias: limits the generalisability of the data and is
that arise from language barriers.17 Algorithms in natural caused by non-random sampling of subgroups
language processing can analyse and interpret large • Simpson’s paradox: bias that can occur in the analysis of
amounts of text from different data sources, generating heterogeneous data that consists of subgroups;
cardiovascular diagnoses, decision support, or treatment associations observed in one subgroup might be different
recommendations in response to queries from users; this from what is observed in another subgroup
creation of new content—generative AI—has transformed • Variable selection bias: results when important variables
knowledge acquisition and exchange in health care.16 are not included in the training dataset, resulting in
Digital datasets that train AI algorithms should include incorrect associations, or associations attributed to the
a broad range of people to whom applications will be rele- wrong variable
vant, but inclusion can vary depending on geography,
regional economics, and individual demographics. Health-
care services, including digital health care, can be restricted conditions can have serious implications on health
in rural and remote regions, and for women, minority outcomes.
ethnic people, and people who are socioeconomically
deprived.1,5 In addition to being under-represented in Biased training data
digital datasets, these groups might also be misrepresented Data selection and sampling can be a source of bias
in online text and language models due to biases in health- (panel). Representation bias can result from how
care descriptors, decision making, diagnoses, and resource the study population is defined; if the groups included
allocation; these biases might be related to nationality, age, in the training dataset do not adequately represent
sex or gender, race or ethnicity, presumed religion, and those with the condition of interest, bias can occur.19
socioeconomic status of patients.18 The multiple sources Some groups might be under-represented in cardiac
of data gaps and biases are amplified when used to train AI diagnostic imaging datasets,21 for example, if they
algorithms, culminating in bias that can increase dispari- receive less care and are under-referred for imaging in
ties in historically marginalised groups. clinical settings. Similarly, female patients are under-
referred for intracardiac devices and cardiac
Sources and types of AI bias synchronisation therapy22 and would be under-repre-
There are several steps in developing, testing, and sented in related digital datasets. Another source of bias
implementing AI algorithms (figure 1), and bias in any is sampling bias or non-random sampling,9 which can
of these steps could result in AI bias.11 Cardiovascular result in biased training data; however, even random
diagnoses can be particularly high risk or life limiting sampling is inadequate for mitigating bias if the training
relative to other health-care conditions,2,3 and biased AI data are not representative of those for whom the algo-
algorithms with poor performance in cardiovascular rithm will used.
Some biases affect the quality and reliability of training detect moderate or severe valvular heart disease in
data (panel). Measurement bias can result in a training a cohort of 7163 patients from three centres and was then
dataset with erroneous data or measurements (eg, using externally validated. The model demonstrated decreased
readings from pulse oximeters or other wearable technolo- accuracy across the lifespan, and there was a numeric
gies that perform sub-optimally or lack validation in people trend towards worse performance in Black patients such
with darker skin).8,21 Another example is the use of single- that valvular disease was detected in 4·4% of Black
photon emission CT for chest pain evaluation algorithms
in women, despite the greater accuracy of cardiac MRI.21 AI application Outcome
Variable selection bias occurs when important variables
Elias et al AI tool to detect moderate or severe The model showed decreased accuracy in
are not included in the training dataset. Incorrect associa- (2022)24 valvular heart disease older versus younger adults; there was a
tions are made or attributed to the wrong variable when numeric trend towards worse performance
the relevant variables are not included or when inappro- in detecting valvular heart disease in Black
patients versus White patients
priate proxy variables are used.23 Developers can also be
biased in how they weight variables for algorithm decision Cheema et al An ensemble machine learning model to In the test set, the model showed
(2022)25 detect patients with stage C or stage D numerically poorer accuracy in female versus
making, resulting in biased programming. Annotation heart failure male patients, and in patients who were
bias can be introduced by annotating text, voice, imaging, Black, or other races, versus patients who
or video data such that different labels are applied for were White
different demographic groups.20 A facial expression anno- Hong et al Machine learning models for stroke All algorithms showed poorer risk
(2023)26 prediction compared with existing stroke- discrimination in Black patients compared
tated as in pain if the face belongs to a White man, for prediction algorithms and with the with White patients
example, might be labelled as upset in a Black woman atherosclerotic cardiovascular disease pooled
despite similarities in the expression. Biases can be intro- cohort equation
duced with incorrect labels applied to training data, Jabbour et al Randomised clinical vignette survey Compared with clinician judgement alone,
perpetuating stereotypes and inequalities in the algorithm. (2023)27 comparing non-biased and systematically unbiased AI models improved clinician
biased AI model predictions on clinician diagnostic accuracy while biased AI models
Outcome labelling bias can result when an outcome is not diagnostic accuracy substantially decreased clinician accuracy
objectively defined or detected across different groups;8 for Kaur et al A convolutional neural network model Risk discrimination was poorer in older than
example, myocardial ischaemia is under-diagnosed as (2024)28 trained on 12-lead electrocardiograms to younger patients and, among younger
a cause of chest pain in women, and relying on clinician predict incident heart failure within 5 years patients, in Black versus other patients
diagnosis for establishing the outcomes can impair Li et al Machine learning models to predict the risk The model underestimated risk in racial or
(2022)29 of prolonged length of hospital stay or in- ethnic minorities, females, and
the performance of an AI algorithm in female patients.
hospital mortality for heart failure socioeconomically deprived individuals
subpopulations
Biased algorithms Li et al Machine learning-based predictive models True positives and positive prediction values
AI algorithm bias describes systematic, intrinsic, and (2023)30 for 10-year cardiovascular disease risk were lower in female than male patients
repetitive errors in algorithms that produce unfair outputs assessment
and compound existing inequities in health-care systems,11 Obermeyer et al Commercial risk-prediction tool to target The model, when implemented, led to racial
(2019)23 patients for high-risk care management bias by systematically underestimating risk
which can arise from biases in how algorithms are used in programmes in Black patients and under-referring them
clinical settings and how they learn over time.7 Blind trust for care management
in an AI algorithm, selective uptake in practices with AI=artificial intelligence.
homogeneous patients, biased or disparate health-care
patterns, and application of the algorithms to improve sub- Table: Examples of cardiovascular AI applications with observed biases in performance
optimally selected outcomes could introduce biases in
algorithm learning over time.7 Latent bias describes biases
that are waiting to happen and can develop even when Inaccurate risk
the algorithm was initially fair.7 Evaluation bias can result prediction
from using inappropriate evaluation criteria to assess
the performance of algorithms and can contribute to
the temporal propagation of bias.19
Health-care
Missed
disparities and Artificial
Consequences of AI bias on cardiovascular care discrimination intelligence bias
diagnoses
patients, versus 10·0% in White patients. The results stroke) and compared it with the American Heart
highlighted the importance of model development and Association’s Pooled Cohort Risk Equations (PCE). All
validation across diverse cohorts.24 In a 2024 study, a deep machine learning models that were trained with features
learning model trained on ECGs to predict 5-year inci- of the PCE model had significant bias (measured by equal
dent heart failure demonstrated poorer risk opportunity difference and disparate impact) across race
discrimination in older than younger adults, and in male and sex. True positive rates and positive prediction rates
than female patients. Among younger patients, risk were lower in female than male patients. A deep learning
discrimination was poorer in Black patients versus other model that was assessed in the same cohort showed
patients.28 A 2019 study assessed a commercial risk- significant bias across sex and race.30 In a 2022 retrospec-
prediction algorithm and found that sicker Black patients tive study, machine learning models trained to predict
were given similar risk scores as healthier White patients. the risk of prolonged length of hospital stay or in-hospital
This biased risk score classification resulted in the under- mortality in a cohort of 210 368 patients with heart failure
referral of Black patients for complex care programmes.23 underestimated risk in female patients; Asian, Black, or
Biased AI models reduce performance and decrease Hispanic patients; and patients of low socioeconomic
clinician diagnostic accuracy. For example, in a 2023 status.29 Last, a 2023 retrospective study of 62 482 patients
survey study of 457 clinicians randomly assigned to AI assessed the performance of machine learning and
model predictions with and without explanations, clini- existing stroke prediction algorithms in 10-year stroke
cians’ baseline diagnostic accuracy improved as long as prediction and found that all algorithms showed poorer
the models were not biased. Systematically biased AI risk discrimination (measured by C indexes) in Black
model predictions decreased clinician accuracy compared patients than in White patients.26 Overall, an algorithm
with baseline clinician judgement. Model explanations that performs poorly could lead to inappropriate testing,
did not mitigate the detrimental effect of biased AI inaccurate diagnoses, or incorrect treatment.10
models.27 In a 2022 retrospective study, a machine
learning model used to detect patients with stage C or D Strategies to mitigate AI bias and promote AI
heart failure had numerically poorer reported accuracy in health equity
female patients (0·81) than in male patients (0·85); and We propose a framework for AI health equity that inter-
in patients of racial groups (0·77) other than Black (0·82) sects AI, digital technology, and health equity in cardiology
or White (0·84) patients.25 A 2023 retrospective cohort care and research. This framework would consider health
study of 109 490 patients evaluated machine learning equity in each step of algorithm development, from
models used to predict the risk of cardiovascular disease the purpose to data collection, variable and outcome selec-
(coronary heart disease, myocardial infarction, and tion, and algorithm testing. An AI equity framework has
the potential to bridge rather than amplify existing dispari-
ties and overcome biased human decision making that
Sources of artificial intelligence bias Mitigation strategies
disadvantages marginalised populations. The dual prin-
! Training data homogeneity Representative training data and random
sampling
cipal goals of AI health equity should be to prevent AI
from maintaining or worsening known health disparities
and to use AI to enhance health equity.
! Selective variables Inclusion of variables relevant to all As a first step, a research team should be composed
demographic groups of diverse voices, experiences and backgrounds, with
a health equity lens applied throughout the algorithm
development process.8,11,18 Bias can be mitigated with
! Inaccurate or missing data Reliable source measurements and the
data cleaning processes guidance from diverse and equity-trained data scientists
who actively assess for biases in study questions and are
aware of clinical biases and structural inequities in
! Biased algorithms Base models designed to avoid bias
(ie, bias checks, bias effect assessment, and the health-care system; “keeping the human in the loop”
de-biasing algorithms) can mitigate bias as trained experts are typically cognisant
of biases that some datasets are prone to.10
Mitigation strategies should address sources of bias
! Lack of external validation Externally validated algorithms accounting
for changes throughout time arising from various stages of the AI algorithm develop-
ment, training, and testing process (figure 3).8,9 The
algorithm training data should represent people living
! Annotation and outcome labelling Clear definitions and adjudicated outcomes
with cardiovascular disease,9 with transparent reporting
on the source of training data, variable and outcome
selection, algorithm development, and algorithm testing
! Research team homogeneity Diverse research team and health equity lens
or validation.9,11 Selected variables should be relevant to
and represent patients with cardiovascular disease
Figure 3: Mitigation strategies for artificial intelligence bias across age, sex and gender, race and ethnicity,
diagnoses and recommendations for incorrect 15 Attia ZI, Harmon DM, Dugan J, et al. Prospective evaluation of
treatments. An AI health equity framework and bias smartwatch-enabled detection of left ventricular dysfunction.
Nat Med 2022; 28: 2497–503.
mitigation strategies at each development step could 16 Gala D, Makaryus AN. The utility of language models in
help ensure that all people living with cardiovascular cardiology: a narrative review of the benefits and concerns of
disease benefit from AI without being subjected to ChatGPT-4. Int J Environ Res Public Health 2023; 20: 6438.
17 Chonde DB, Pourvaziri A, Williams J, et al. RadTranslate:
avoidable risk. an artificial intelligence-powered intervention for urgent imaging
Contributors to enhance care equity for patients with limited English
HGCV conceptualised the study. AM and HGCV developed the figures proficiency during the COVID-19 pandemic. J Am Coll Radiol
and wrote the manuscript. All authors conducted the literature search, 2021; 18: 1000–08.
developed the tables, and edited revisions. All authors had full access to 18 Arora A, Alderman JE, Palmer J, et al. The value of standards for
the data and had final responsibility for the decision to submit for health datasets in artificial intelligence-based applications.
publication. Nat Med 2023; 29: 2929–38.
19 Belenguer L. AI bias: exploring discriminatory algorithmic
Declaration of interests decision-making models and the application of possible machine-
AP has received research support from the National Institutes of Health; centric solutions adapted from the pharmaceutical industry.
received grant funding from Applied Therapeutics and Gilead Sciences; AI Ethics 2022; 2: 771–87.
received consulting fees for Tricog, Novo Nordisk, Bayer, Medtronic, 20 Tat E, Bhatt DL, Rabbat MG. Addressing bias: artificial
Edward Lifesciences, Cytokinetics, Roche, Sarfez Pharma, Science37, intelligence in cardiovascular medicine. Lancet Digit Health 2020;
Rivus, Axon Therapies, Alleviant, and Lilly; received non-financial 2: e635–36.
support from Pfizer and Merck; participated on data and safety 21 van Assen M, Beecy A, Gershon G, Newsome J, Trivedi H,
monitoring boards for Bayer, Cytokinetics, Novo Nordisk, and Gichoya J. Implications of bias in artificial intelligence:
Medtronic; and is a consultant for Palomarin with stocks compensation. considerations for cardiovascular imaging. Curr Atheroscler Rep
HCGV and AM declare no competing interests. 2024; 26: 91–102.
22 Sullivan K, Doumouras BS, Santema BT, et al. Sex-specific
Acknowledgments differences in heart failure: pathophysiology, risk factors,
No funding was received for the preparation of this manuscript. management, and outcomes. Can J Cardiol 2021; 37: 560–71.
References 23 Obermeyer Z, Powers B, Vogeli C, et al. Dissecting racial bias in
1 Mihan A, Van Spall HGC. Interventions to enhance digital health an algorithm used to manage the health of populations. Science
equity in cardiovascular care. Nat Med 2024; 30: 628–30. 2019; 366: 447–53.
2 Khan MS, Arshad MS, Greene SJ, et al. Artificial intelligence and 24 Elias P, Poterucha TJ, Rajaram V, et al. Deep learning
heart failure: a state-of-the-art review. Eur J Heart Fail 2023; electrocardiographic analysis for detection of left-sided valvular
25: 1507–25. heart disease. J Am Coll Cardiol 2022; 80: 613–26.
3 Averbuch T, Sullivan K, Sauer A, et al. Applications of artificial 25 Cheema B, Mutharasan RK, Sharma A, et al. Augmented
intelligence and machine learning in heart failure. intelligence to identify patients with advanced heart failure in an
Eur Heart J Digit Health 2022; 3: 311–22. integrated health system. JACC Adv 2022; 1: 100123.
4 Myrick N, Gilbert S. How digital technologies can address 5 sources 26 Hong C, Pencina MJ, Wojdyla DM, et al. Predictive accuracy of
of health inequity. Sept 8, 2021. [Link] stroke risk prediction models across black and white race, sex,
agenda/2021/09/how-digital-technologies-can-address-5-sources-of- and age groups. JAMA 2023; 329: 306–17.
health-inequity/ (accessed Oct 26, 2023). 27 Jabbour S, Fouhey D, Shepard S, et al. Measuring the impact of
5 Reddy H, Joshi S, Joshi A, Wagh V. A critical review of global digital AI in the diagnosis of hospitalized patients: a randomized clinical
divide and the role of technology in healthcare. Cureus 2022; vignette survey study. JAMA 2023; 330: 2275–84.
14: e29739. 28 Kaur D, Hughes JW, Rogers AJ, et al. Race, sex, and age
6 Ibrahim H, Liu X, Zariffa N, Morris AD, Denniston AK. Health disparities in the performance of ECG deep learning models
data poverty: an assailable barrier to equitable digital health care. predicting heart failure. Circ Heart Fail 2024; 17: e010879.
Lancet Digit Health 2021; 3: e260–65. 29 Li Y, Wang H, Luo Y. Improving fairness in the prediction of heart
7 DeCamp M, Lindvall C. Latent bias and the implementation of failure length of stay and mortality by integrating social
artificial intelligence in medicine. J Am Med Inform Assoc 2020; determinants of health. Circ Heart Fail 2022; 15: e009473.
27: 2020–23. 30 Li F, Wu P, Ong HH, Peterson JF, Wei WQ, Zhao J. Evaluating
8 Nazer LH, Zatarah R, Waldrip S, et al. Bias in artificial intelligence and mitigating bias in machine learning models for
algorithms and recommendations for mitigation. PLoS Digit Health cardiovascular disease prediction. J Biomed Inform 2023;
2023; 2: e0000278. 138: 104294.
9 Vokinger KN, Feuerriegel S, Kesselheim AS. Mitigating bias in 31 Segar MW, Jaeger BC, Patel KV, et al. Development and validation
machine learning for medicine. Commun Med 2021; 1: 25. of machine learning-based race-specific models to predict 10-year
risk of heart failure: a multicohort analysis. Circulation 2021;
10 Mittermaier M, Raza MM, Kvedar JC. Bias in AI-based models for
143: 2370–83.
medical applications: challenges and mitigation strategies.
NPJ Digit Med 2023; 6: 113. 32 US Food and Drug Administration. Software as a medical device
(SaMD) action plan. January, 2021. [Link]
11 Panch T, Mattie H, Atun R. Artificial intelligence and algorithmic
media/145022/download (accessed Jan 14, 2024).
bias: implications for health systems. J Glob Health 2019; 9: 010318.
33 US Food and Drug Administration. Marketing submission
12 Accuray. Overcoming AI bias: understanding, identifying and
recommendations for a predetermined change control plan for
mitigating algorithmic bias in healthcare. 2023. [Link]
artificial intelligence/machine learning (AI/ML)-enabled device
com/blog/overcoming-ai-bias-understanding-identifying-and-
software functions: draft guidance for industry and food and drug
mitigating-algorithmic-bias-in-healthcare/ (accessed June 14, 2024).
administration staff. April, 2023. [Link]
13 SuperAnnotate. Bias in machine learning: types and examples. information/search-fda-guidance-documents/marketing-
March 17, 2022. [Link] submission-recommendations-predetermined-change-control-plan-
machine-learning#mitigating-bias-in-machine-learning (accessed artificial (accessed Jan 16, 2024).
June 14, 2024).
14 American Heart Association. AI-phone app detected worsening Copyright © 2024 The Author(s). Published by Elsevier Ltd. This is an
heart failure based on changes in patients’ voices. Nov 13, 2023. Open Access article under the CC BY-NC 4.0 license.
[Link]
worsening-heart-failure-based-on-changes-in-patients-voices
(accessed Dec 25, 2023).