0% found this document useful (0 votes)
13 views11 pages

Predictive

This document discusses the use of electronic health records for predictive analytics and data science in healthcare. It outlines some key advantages of analyzing EHR data compared to other research methods, but also identifies several potential pitfalls and biases that can arise from variations in data collection practices and limitations of retrospective analysis. The goals of the document are to provide guidance on avoiding common methodological errors when conducting research using EHR data.

Uploaded by

sans
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views11 pages

Predictive

This document discusses the use of electronic health records for predictive analytics and data science in healthcare. It outlines some key advantages of analyzing EHR data compared to other research methods, but also identifies several potential pitfalls and biases that can arise from variations in data collection practices and limitations of retrospective analysis. The goals of the document are to provide guidance on avoiding common methodological errors when conducting research using EHR data.

Uploaded by

sans
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

Data Science in Healthcare: Leveraging Electronic


Health Records for Predictive Analytics
DR. RUPESH SHUKLA1, AARYESH SHUKLA2
1
Principal, Department of Computer Science, ILVA Commerce and Science college, Indore / DAVV
Indore, India.
2
Department of CSE, RGPV, Bhopal, India.

Abstract- Analysis of electronic health records, I. INTRODUCTION


often known as EHR analysis, is a technique that
is gaining popularity and is being used Electronic health records, more often referred to as
increasingly frequently to do research on patient EHRs, are rapidly being utilised to conduct out
data from the real world. When compared to other population health surveys, construct categorization
approaches to study, the use of data that is and prediction models for decision support,
routinely obtained offers a variety of advantages, determine which treatment techniques are the most
such as fewer "administrative costs, the successful, and even duplicate randomised clinical
opportunity to update studies when new patterns of trials. Other common abbreviations for EHRs are
behavior emerge, and larger sample sizes. EHR EMRs and EHRs. The use of data that has already
analysis comes with its own distinct set of been gathered is the primary benefit of EHR analysis
methodological challenges as a result of the fact when contrasted with the use of other "data sources
that the data in question were not collected with the and research methods," such as cohort studies or
goal of doing research. In this Viewpoint, we randomised controlled trials. This is the most
elaborate on the necessity of having an in-depth significant advantage of using EHR analysis. Due to
grasp of clinical procedures and outline six the presence of this benefit, it is preferable to the
potential pitfalls that should be avoided while utilisation of several other data sources and research
working with EHR data. Both of these topics are methodologies.[1,2] This helps to reduce the
covered in the context of dealing with electronic quantity of administrative work that is required, as
health record data. In order to do this, we rely on well as the costs and the possibility of bias in the
examples from the research that has already been process of selecting samples. Electronic health
conducted in addition to our own personal records, often known as EHRs, are less likely to be
experiences. We provide solutions that may be used affected by inclusion bias than randomised
to avoid or lessen the impact of each of these six controlled trials are. This is due to the fact that EHRs
concerns, which are as follows: subjective are more representative of the overall population
treatment allocation, sample selection bias, that is being targeted. This is as a result of the fact
imprecise variable definitions, restrictions to that data are collected from any and all persons who
deployment, variable measurement frequency, and interact with health care in any capacity. As time
model over fitting. In conclusion, we have great goes on, electronic health records, often known as
expectations that this Viewpoint will serve as a EHRs, will be able to combine more specific data
roadmap for researchers to follow in order to from patients and will provide access to datasets
further increase the methodological rigour of EHR with increasing quantities. This is especially true in
analysis. This optimism is based on the fact that we situations when the EHR study in issue comprises
have high hopes that this Viewpoint will serve as a large integrated health-care systems or a network of
roadmap. health-care providers that make use of information
systems that are interoperable. Despite the fact that
Indexed Terms- Healthcare, Electronic, Data there are still significant epidemiological
Science consequences, such as those that are discussed in
this Viewpoint, electronic health records (EHRs)
provide increasingly complete data from
patients.[3,4] At the end of the day, having a sample
size that is sufficient enough might give enhanced

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 383


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

"statistical power to carry out subgroup analyses and the evaluation of the information that is included in
reduce the chance of generating type II errors."[5] EHRs should never be carried out on one's own
without the support of investigators who have a
• Importance of methodological degree of knowledge that is comparable to that of an
The complexity of the data included in EHRs can authority on the subject matter that is being studied.
frequently result in methodological issues, which This degree of competence need to include
can restrict both the application and validity of the everything, from the procedures that go into giving
findings from research conducted using EHRs. Even therapy to the processing of data, among other
while the examination of EHR data may offer things.[16,17]
fruitful research opportunities, the practical
implementation and scientific credibility of the • Goals of this Viewpoint"
resulting conclusions are occasionally impeded by a In the past twenty years, with the ever-increasing use
number of factors.[6] To begin, the potential for of electronic health record (EHR) data, previous
variance in the method of data gathering utilised for investigations and our research have observed that
EHRs is significantly higher compared to that which issues arising during clinical care (for example,
is utilised for clinical archives and clinical trials.[7] people from minority ethnic groups seeking health
This is due to the fact that the process of evaluating care less frequently) are also reflected in the EHR
numerous variables makes use of a diverse range of data, thus introducing bias in EHR studies. This bias
technologies and sample frequencies, which in turn can be traced back to the fact that people from
increases the possibility that variation will occur. minority ethnic groups seek health care less
Second, because the vast majority of EHR data frequently.[18] This discrimination may be traced
studies are carried out using a retrospective back to the fact that persons who belong to ethnic
methodology, the cohorts, exposures, and outcomes groups that are underrepresented in the majority are
are all characterised using language that focuses on more likely to seek medical attention than those who
the past.[8,9] This is due to the fact that the vast belong to ethnic groups that are underrepresented in
majority of EHR data is archived in a retroactive the minority.[19] In the corpus of research that
fashion. This retrospective technique makes it already exists, there does not seem to be an
possible to modify such criteria in light of the results integrated and comprehensive evaluation of the
of the analysis, which might lead to findings that can ways in which clinical processes and information
be interpreted in a way that is unsuitable as a system design effect EHR data analysis, in our view.
consequence of the testing of a large number of This is something that we believe needs to be
hypotheses without the necessary statistical addressed. In this Viewpoint, we elaborate on these
adjustment. [10] In other words, the findings might ideas by combining important clinical concerns with
be misinterpreted as a result of the testing of the machine learning, epidemiological, and statistical
hypotheses.[11] Not to mention the fact that clinical factors.[20] We do this by conducting an analysis of
practise patterns not only have a considerable the relevant literature and presenting a number of
influence on the sort of data that is gathered and the case stories.[21] There are still others that are
quality of that data, but also on the patients who are necessary for any sort of EHR study, despite the fact
selected to participate in the research.[12,13] These that the bulk of the risks that we provide are largely
patterns add biases into the process of selecting important to models that are based on machine
samples, which leads to the formation of misleading learning. In this Viewpoint, we adopt a mixed
linkages between variables. It is vital to keep in approach, based both on the assessment of experts
mind that the data gathered through electronic health and on a study of the relevant literature, to identify
records are not collected with the primary goal of six common clinical and methodological
undertaking research.[14] This is something that blunders.[22] These mistakes include both clinical
must be kept in mind at all times. Instead, they serve and methodological errors.[23,24] Our literature
the aim of maintaining all of the information about review as well as the references of the publications
patients that is gathered during clinical treatment, or that were discovered were compared to one another
in other cases, they play a role in administration such in order to check that we had not overlooked any
as invoicing. In addition, they are responsible for the relevant studies. "The authors (CMS, SLH, and
safekeeping of patient records.[15] These two roles LAC) were the ones who made the choice of which
are equally essential to the whole. Because of this, papers to include in their review. We believe that by

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 384


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

explaining these six pitfalls of critical importance, unit research be recognised in EHR
we will be able to point researchers in the direction research.[32,33]Protocols that are adhered to in
of avoiding making the same errors in future studies, hospitals may decide which data are collected, when
and we hope that the solutions that we propose will they are collected, and how they are collected.
enhance the scientific robustness of descriptive and Additionally, these protocols may also influence
predictive models that incorporate EHR data.[25] whether or not certain data are obtained. For
Even though some of the points that we present here instance, in order for the electronic health record
are elaborations or expansions of essential ideas that (EHR) to provide an odd blood test result, the patient
have already been published in reporting must have previously had a blood test.[34]
recommendations such as the Strengthening the Therefore, the identification of people with aberrant
Reporting of Observational studies in Epidemiology results can be related with both local processes and
checklist, we also discuss additional dangers and different testing frequencies for distinct patient
concepts" that it is essential to keep in mind. In groups. This is because of the fact that the testing
addition, the provision of solutions that are capable frequencies might vary. [35] However, despite the
of being implemented is an additional part of this fact that this issue of missing data is extremely
Viewpoint that adds to the total value of the common in research, it is often ignored. This results
proposition.[26,27] in bias in the selection of samples, which is very
seldom adjusted for. The criteria that must be met in
• Pitfalls and solutions order for a patient to be admitted to an intensive care
It is possible to make the case that the vast majority unit (ICU) will have an effect on the findings of any
of mistakes are not made on purpose or as a result of research that involves the analysis of EHR data
laziness, but rather because the individual who obtained from patients treated in an ICU. These
makes them is unaware of the knowledge gaps in criteria are different from one intensive care unit
their own understanding. When trying to develop (ICU) to the next, and they may even shift based on
more sophisticated statistical models, it is vital to the circumstances within a single institution.[36]
have an in-depth understanding of how data are During the COVID-19 pandemic, these changing
gathered and choices are made. [28] This is circumstances were easily evident when the number
particularly true when trying to construct more of patients who required critical care unit beds
complex statistical models.[29] For instance, it is of exceeded the available beds in the hospitals. This
the utmost importance to be aware of the patients caused a shortage of beds. When a patient is ready
who are admitted to a hospital or an intensive care for departure from the hospital (selective censoring),
unit (ICU), the process by which choices about for instance, "or when a patient should be readmitted
treatment are made, and the appropriate time for from the ward to the critical care unit, are both
patients to be released from an ICU or from instances of clinical decisions that are sensitive to
hospital.[30] In spite of the fact that non-medical the subjectivity of the health care practitioners who
researchers may be able to gain some insights into are making them.[37,38] In the majority of these
these processes by reading about them in the media instances, the decisions about triage and other topics
or by looking them up on the internet, we strongly could be impacted in some way by elements of the
advise including doctors, nurses, or any other healthcare practitioners themselves, such as their
relevant health-care practitioners in the research clinical experience and cultural differences, as well
project at every stage of its development. In addition as by current events in general, such as the
to this, it is important to note that non-medical pandemic. We strongly advise doing regular reviews
researchers may be able to gain some insights into of the design choices and research assumptions with
these processes by reading about them in the media physicians or other health-care professionals who
or by looking them up on the internet.[31] Because are familiar with the local practises in the area.
the influence of hospital-specific practises was not Because of this, you will be able to find a solution to
taken into account in a number of the earlier study, the problems that were brought up before. In
this oversight ultimately led to issues with the addition, studies that attempt to explore or predict
external validity of the studies, and in some outcomes that occur from treatment decisions (for
instances, with the internal validity of the studies as example, which patients should be administered
well. in addition to the need that the influence of renal replacement therapy) should have causal
domain knowledge on the policy of local health-care inference frameworks[39]. These frameworks help

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 385


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

researchers determine which patients should be between 0.6 and 0.8.38–40 In addition to this, the
treated with renal replacement therapy and how nature of the sickness itself creates the likelihood of
often. It is possible that these frameworks will assist inaccurate classification being applied to the
decide which individuals need to have renal patient.[45] When adopting the gold-standard
replacement therapy delivered to them. We have criterion for sepsis, which is 3.41, this is one
provided an overview of the many solutions that are illustration of this phenomenon. However, the use of
conceivable, as well as a list of the probable "more inclusive definitions of what constitutes a
difficulties that may occur in the future (figure). The suspected infection rather than restrictive ones (for
dangers are structured in a way that corresponds to example, whether a positive nitrite urine test is
the stage of the analysis in which they are most sufficient or whether other signs of tissue invasion
likely to occur. This ensures that the information is also need to be present) could significantly affect
easy to find and understand".[40] both the size of the cohort and the performance of
the model". A positive nitrite urine test is one of the
• Sample Selection Bias most important components of the Sepsis-3
The step of data analysis known as "building definition.[46] When carrying out research, it is
cohorts" is a challenging one since it requires essential to take into consideration essential criteria
converting a large number of case descriptions into such as cohort definitions and external validation.
specific data criteria. As a direct consequence of The poor performance of the Epic Sepsis Model
this, this stage is very challenging.[41] The capacity serves as an illustration of why this is the case. It is
of the researchers to precisely interpret and translate important to note that even when algorithms are
these definitions, as well as "the quality of the used to identify illnesses like sepsis, retrospective
definition that was used in the literature to correctly examinations of people who are presumed to have
identify instances, are both variables that have an sepsis may indicate significant
impact on the composition of the cohort. misclassification.[47] This is something that has to
Importantly, sample selection bias may occur if be taken into account, so keep that in mind. A further
these case criteria are either too strict, which may example of the use of criteria that are overly general
result in the exclusion of a subgroup, or overly wide, may be found in the field of mortality prediction. It
which may result in an increase in the number of has been shown that the produced cohorts that are
patients who have been mistakenly identified as used in the different research projects are
being part of the cohort.[42] Both of these scenarios comparable to one another in terms of their
may lead to the same end result: an increase in the characteristics.
number of patients who" have been misclassified as
being a member of the cohort. Studies that used
different methods to identify patients with sepsis are
a prominent example of definitions that are too
general.[43] These studies include "the quick
Sequential Organ Failure Assessment (qSOFA)
score, codes proposed by Martin and colleagues32
(also known as the Martin methodology)", codes
proposed by Angus and van der Poll33 (also known
as the Angus methodology), and the systemic
inflammatory response syndrome score. In spite of
the fact that it has been shown that these strategies
may be useful as proxies for sepsis, the specificity Figure: 1 In studies that use electronic health
of such approaches is missing. In the meanwhile, it records, there are often occurring clinical and
has been shown on several times that the capacity to methodological difficulties, as well as possible
recognise individuals with sepsis in a range of remedies.[48]
contexts is varied and somewhat discriminatory.[44]
This is the case even though sepsis is a rather variable, despite the fact that all of the studies used
common condition. This is shown by areas under the the exact "same dataset (Medical Information Mart
receiver operating characteristic curve (AUROC) for Intensive Care III) and investigated the
values for qSOFA scores of 2 or above that lie performance of the exact same model (i.e., mortality

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 386


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

prediction). This was the case even though all of the feasible that the procedure codes may not cover all
publications were published in the same of the available procedures; rather, they may just
year.[49]Age limits (for instance, excluding people include those that can be invoiced. Before settling
younger than 18 years of age), the exclusion of on a choice about the outcome measure, it is
individuals who had numerous stays in the essential to take into account both the
intensive" care unit, or the necessity of certain epidemiological ramifications and the clinical
measures, such as those of infection indicators, were context of the study. An example that comes up
some of the factors that contributed to the diversity rather often is the conundrum of whether or not to
of the inclusion and exclusion criteria. Other factors use in-hospital events (like hospital mortality), as
that contributed to this diversity included the opposed to a fixed time point (like 28-day
necessity of certain measures, such as those of mortality), as a measure of patient outcomes. When
infection indicators. We strongly suggest that you investigating the consequences of direct hospital
always make use of the definitions that have been treatments, it is customary to focus on events that
thoroughly reviewed; nevertheless, you should bear take place within the hospital.[54] An effect that
in mind that even these definitions are prone to should be examined using in hospital events is the
mistake and are not cast in stone.[50] The examples influence of using prophylactic heparin on the
that were shown before provided the foundation for incidence of venous thromboembolism in patients
this advice. In addition, if the performance of the being treated in hospitals. This is only one example
model is supplied, it need to be compared with the of an effect that should be studied using inhospital
algorithms that are now in use, regardless of whether events. On the other hand, if you are interested in the
such algorithms are based on the opinions of incidence of postsurgical thrombosis, you should
specialists or on machine learning. In the event that choose a predetermined time period as your starting
comparisons are made using algorithms that are point for the study. This is due to the fact that the
derived from machine learning, it is essential that length of stay a patient has in the hospital is related
differences in terminology be researched. We to the likelihood that they may develop thrombosis
recommend addressing any possible limits, during their stay. In general, we recommend
emphasising any potential implications on the consulting the expertise of a group of experts from a
robustness of the findings, and giving a clear variety of fields in order to get advice on definitions.
explanation of the criteria that were utilised for the Patients should be included on this team in addition
selection of cohorts. We also advocate addressing to specialists who are knowledgeable in relevant
any potential implications on the robustness of the topic areas. Examples of such experts are medical
results. In the event that there are no criteria that are professionals, epidemiologists, statisticians, and
universally acknowledged as the benchmark, we social scientists.
suggest doing a sensitivity analysis in order to
ascertain which approach is most effective in • Limits To Deployment
locating the population that is the subject of the The fact that the data that is available in EHRs is
investigation.[51,52] unable to be simply turned into clinical practise
offers a substantial impediment for the
• Imprecise Variable Definitions implementation and deployment in the actual world
After the building of the cohort, the next key pitfall of machine learning or other prediction models that
to look out for is the definition of the variables were generated from EHR research. This issue arises
incorrectly. The degree to which the definitions are whenever the data format of EHRs deviates from the
sensitive and particular has a significant bearing on practise that takes place in the real world, as well as
whether or not the analyses correctly depict how if there is a lack of time-stamped results. Time-
things are done in the actual world. Because of the stamped data, for instance, may be judged to be
impact of this factor, definitions need to be subjected available at the moment of measurement and then
to a close and careful analysis. This is also true when reviewed as if it had been gathered at that point in
forming cohorts.[53] In addition, it is essential to time during a retrospective analysis. This evaluation
have a clear understanding of the difference between may be performed as if the data had been obtained
the data that is being collected for the purposes of at that point in time. On the other hand, it's likely
billing and the data that is being recorded for that doctors and other people who make decisions
therapeutic treatment purposes. For instance, it is won't really have access to this data at the time that

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 387


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

was first expected. This sort of disparity may appear, leakage of patient data from the training dataset into
for example, when the findings "of blood tests or the test dataset".
blood cultures are used.[55] This is because the
results may be included in the analysis with the • variable measuring frequency
timestamp of the registration of the blood draw The natural link that exists between the frequency of
rather than the time at which they were made measurements and the severity of the illness is
available to the doctors. A further illustration of this another issue that is often neglected, despite the fact
would be the exclusion of patients who have very that it is unavoidably there. When a patient's health
long hospital stays since this information is not is considered to be unstable, practitioners may
available early on in the course of the patient's typically issue orders for more laboratory tests or
hospital stay (which is when the bulk of algorithms analyse records of vital sign readings on a more
are used). In addition, International Classification of regular basis. It is essential that, regardless of
Diseases codes 55 are not often allocated to patients whether one is creating predictive or descriptive
until after they have been released from the hospital models, the link that exists between the frequency of
or have gone away, and the time stamps that are measurements and the severity of the disease be
connected with the occurrence of these codes might taken into consideration. For example, if patients
vary. Because of this variation in timing, the use of who had a substantial proportion of missing (that is,
these codes is not something that would be suitable not executed) data were deleted from the study, this
for models that utilise hospital admission as their may have an effect on the validity of the model.
baseline. The performance of the model may be Patients who had a significant percentage of missing
overstated in many of the aforementioned situations, (that is, not executed) data were omitted from the
and it may not be possible to reproduce its results in research. This exclusion often results in a biassed
applications that take place in the real world. The model, which, in the same way as sample selection
appearance of training data in the test datasets is a bias (trap 1) has the ability to overstate the
separate but related issue, which is connected to the seriousness of the condition, it also has the potential
leaking of data, which is a distinct but related one. to be misleading. It is vital to keep in mind that
This is a different issue, yet there is a connection imputation, in which the values of the cohort's mean
between both. This leakage may often take place if or median are used, is not a technique that may be
repeated admissions from the same patients are used to compensate for missing data. This is because
used, or if the timeseries data from a single patient the frequency of taking measurements is not a
are not restricted to either the training set or the test completely arbitrary occurrence. This example
set, but instead appear in both. Alternatively, this highlights how data that are regularly obtained
leakage can take place if repeated admissions from implicitly reflect the judgements of the doctors and
the same patients are used. There is a serious flaw in how there may be a significant degree of diversity
the methodology that is referred to as leakage of data amongst the various providers of medical care. If
and it has the potential to result in an inaccurately there is a substantial amount of variation in clinical
high estimate of the performance of the model. As a practises and if the behaviour of doctors were to
direct result of this, the performance of the model as change at some time in the future, a model that was
well as its utility in the application of clinical trained on these kinds of data may be more likely to
research are both diminished. The challenges that have poor performance. "As a consequence of this,
are associated with deployment may be avoided by we recommend involving statisticians in order to
first acquiring an understanding of when data will discuss appropriate epidemiological (for example,
become accessible in real time, and then selecting weighting) or statistical strategies (for example,
features based on the data that will be made multiple imputation, or removal of highly unreliable
available to physicians. This will allow for the variables entirely), and we also recommend
deployment process to go more smoothly. In involving clinicians in order to identify
addition, data from time series as well as data from circumstances in which variable measurement
a large number of admissions to an institution such frequency could result in a biassed analysis. This is
as a hospital or intensive care unit should be due to the fact that we advocate for incorporating
randomly assigned to either the training dataset or doctors in the process of determining the conditions
the test dataset. Because of this, there will be no under which varied measurement frequency may
lead to an incorrect interpretation of the data".

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 388


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

There is a large amount of subjectivity involved in highly recommend giving serious thought to the
the therapy assigning process question of whether or not it is even conceivable to
carry out a study involving causal inference. "When
In most instances, the goal of research using causal in doubt, sensitivity studies should be undertaken to
inference is to make an informed prediction as to the investigate the potential effect of unmeasured
unobserved, counterfactual treatment outcome, confounders based on data from earlier research.
which then makes it possible to evaluate the impacts These analyses should be performed to analyse the
of the treatment. In order to quantify these possible influence of unmeasured confounders".
consequences, it is required to first identify and then
take into consideration all of the variables that are • Model Over fitting And Decreased Capacity To
associated with the treatment allocation. Only then Generalise Results
can these impacts be quantified. It has been shown There is a possibility that the results cannot be
that this approach can function in the same way as generalized beyond the data source from which they
randomised controlled clinical trials (RCTs), but were produced because of the existence of
despite the fact that it may provide certain significant disparities across the various institutions
advantages, it is not without limitations. Studies that and regions. It is essential to keep in mind that not
utilise causal inference begin with the supposition all research and models need to be generalizable
that every factor that has a role in treatment since there is a possibility of model over fitting
allocation is really observed. This is the starting happening. Therefore, it is necessary to inquire of
point for these types of investigations. However, researchers the answer to the issue of whether or
previous research has shown that the allocation of whether the treatments or results of interest are
treatment can be affected by "both differences in impacted by local practicing patterns. For instance,
interphysician decision making (i.e., different if researchers working inside of a hospital wanted to
physicians prescribing different treatments in the determine how the handoff of a patient from one
same clinical setting) and differences in service to another influences clinical outcomes, the
intraphysician decision making (i.e., bias on best performing model could be constructed without
patient's socioeconomic factors, such as race and attaching an excessive amount of emphasis to issues
ethnicity). In other words, different physicians can resulting from inadequate generalisability. This
prescribe different treatments in the same clinical would allow for the model to have the most
setting". When conducting studies that are based on predictive power. If, on the other hand, the objective
EHRs, researchers should employ causal is to solve this problem outside of the limits of a
frameworks whenever it is practical to do so in order particular institution, then it is necessary to conduct
to minimise introducing bias due to confounding external validation using a separate cohort. In
variables. This is because causal frameworks allow addition, changes in the frequency of important
for more accurate interpretation of results. Before factors or the absence of these variables should be
initiating exploratory studies, it is strongly advised highlighted across clinical settings and correctly
that causal diagrams be developed to illustrate the accounted for whenever it is possible to do so.[56]
team's knowledge of the process by which the data The issue of external validity is not a yes-or-no
were produced (for instance, by making use of statement; rather, it is concerned with establishing
directed acyclic networks). This may be done before which specific clinical settings the analysis is
moving on to the actual exploratory analyses. These appropriate to. This is not a simple yes-or-no
diagrams might be helpful in identifying any argument. When trying to derive the generalizability
missing confounding factors or other crucial of models, a causal diagram may be of aid since it
components that need to be taken into consideration makes it evident which correlations in the data are
over the course of the investigation. Even though most likely to differ depending on the organisation
controlling for confounders is a crucial part of or the region. Additionally, for the purpose of
replicating randomization, electronic health records evaluating the performance of the model, acceptable
(EHRs) often fail to capture important clinical performance measures that are not affected by class
indicators or socioeconomic determinants of health. imbalance should be applied. When there is an
This is despite the fact that mimicking "imbalance in the frequency of a feature or outcome
randomization is a core component of research. In of interest, such as when there is a very low hospital
light of the circumstances presented above, we death rate for certain medical conditions, it is

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 389


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

recommended to avoid using metrics that are prone standing on, or looking over, the shoulders of
to class imbalance as performance measures for the clinicians? NPJ Digit Med 2021; 4: 62.
model". One example of this would be when there is [3] Bonomi S. The electronic health record: a
a very low hospital mortality rate for certain medical comparison of some European countries. In:
diseases. Accuracy is a good illustration of this Ricciardi F, Harfouche A, eds. Information and
concept. In addition, reporting just aggregate communication technologies in organizations
measures of discrimination (such as AUROC) might and society. Lecture notes in information
mask a lower therapeutic usefulness since, for systems and organisation, vol 15. Cham:
instance, the sensitivity or specificity may not be Springer, 2016: 33–50.
high enough. It's likely that utilising measurements
[4] Tambone V, Boudreau D, Ciccozzi M, et al.
that are based on a single operating point, like the F1
Ethical criteria for the admission and
score, which is aimed to strike a balance between
management of patients in the ICU under
accuracy and recall, would be more informative than
conditions of limited medical resources: a
using measures that are based on many operating
shared international proposal in view of the
points. In the end, the assessment of the model has
COVID-19 pandemic. Front Public Health
to be modified in accordance with the use scenario
2020; 8: 284.
for which the system was developed (for instance,
screening as opposed to therapeutic guidance). [5] American Thoracic Society. Fair allocation of
intensive care unit resources. Am J Respir Crit
CONCLUSION Care Med 1997; 156: 1282–301.
[6] Curtis JR, Vincent J-L. Ethics and end-of-life
We discovered six prevalent methodological and care for adults in the intensive care unit. Lancet
clinical difficulties that affect the robustness, 2010; 376: 1347–53.
validity, and reproducibility of research that makes [7] Piers RD, Azoulay E, Ricou B, et al.
use of electronic health record data (EHR data). The Perceptions of appropriateness of care among
most significant challenges consist of a biassed European and Israeli intensive care unit nurses
sample selection, imprecise variable definitions, and physicians. JAMA 2011; 306: 2694–703.
deployment limits, a lack of adjustment for the [8] Usman OA, Usman AA, Ward MA.
relationship between frequency of measurements Comparison of SIRS, qSOFA, and NEWS for
and severity of sickness, subjective treatment the early identification of sepsis in the
allocation, and limited generalisability of findings. emergency department. Am J Emerg Med
These challenges were encountered because the 2019; 37: 1490–97.
sample selection process was biassed. Because of all
[9] Singer M, Deutschman CS, Seymour CW, et
of these different factors, it is challenging to
al. The third international consensus
generalise the results. Although this list is not
definitions for sepsis and septic shock (Sepsis-
exhaustive and the possible solutions to these
3). JAMA 2016; 315: 801–10.
worries do not apply in every circumstance, we have
penned this Viewpoint in the hopes that it will bring [10] Johnson AEW, Aboab J, Raffa JD, et al. A
attention to a number of critical risks associated with comparative analysis of sepsis identification
EHR data and encourage researchers to take them methods in an electronic database. Crit Care
into consideration when working with EHR-based Med 2018; 46: 494–99.
research". [11] Wong A, Otles E, Donnelly JP, et al. External
validation of a widely implemented proprietary
REFERENCES sepsis prediction model in hospitalized
patients. JAMA Intern Med 2021; 181: 1065–
[1] Harutyunyan H, Khachatrian H, Kale DC, Ver 70.
Steeg G, Galstyan A. Multitask learning and [12] Lapsley I, Melia K. Clinical actions and
benchmarking with clinical time series data. financial constraints: the limits to rationing
Sci Data 2019; 6: 96. intensive care. Sociol Health Illn 2001; 23:
[2] Beaulieu-Jones BK, Yuan W, Brat GA, et al. 729–46.
Machine learning for patient risk stratification: [13] Trentini F, Marziano V, Guzzetta G, et al. The
pressure on healthcare system and intensive

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 390


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

care utilization during the COVID-19 outbreak [20] Rajesh Kumar Kaushal, Rajat Bhardwaj,
in the Lombardy region of Italy: a retrospective Naveen Kumar, Abeer A. Aljohani, Shashi
observational study in 43 538 hospitalized Kant Gupta, Prabhdeep Singh, Nitin Purohit,
patients. Am J Epidemiol 2022; 191: 137–46. "Using Mobile Computing to Provide a Smart
[14] Jacoba CMP, Celi LA, Silva PS. Biomarkers and Secure Internet of Things (IoT)
for progression in diabetic retinopathy: Framework for Medical Applications",
expanding personalized medicine through Wireless Communications and Mobile
integration of AI with electronic health Computing, vol. 2022, Article ID 8741357, 13
records. Semin Ophthalmol 2021; 36: 250–57. pages, 2022.
[Link]
[15] Robles Arévalo A, Maley JH, Baker L, et al.
Data-driven curation process for describing the [21] Bramah Hazela et al 2022 ECS Trans. 107
blood glucose management in the intensive 2651 [Link]
care unit. Sci Data 2021; 8: 80. 3 Sauer CM, [22] Ashish Kumar Pandey et al 2022 ECS Trans.
Gómez J, Botella MR, et al. Understanding 107 2681
critically ill sepsis patients with normal serum [Link]
lactate levels: results from US and European [23] G. S. Jayesh et al 2022 ECS Trans. 107 2715
ICU cohorts. Sci Rep 2021; 11: 20076. [Link]
[16] Komorowski M, Celi LA, Badawi O, Gordon [24] Shashi Kant Gupta et al 2022 ECS Trans. 107
AC, Faisal AA. The Artificial Intelligence 2927 [Link]
Clinician learns optimal treatment strategies
[25] S. Saxena, D. Yagyasen, C. N. Saranya, R. S.
for sepsis in intensive care. Nat Med 2018; 24:
K. Boddu, A. K. Sharma and S. K. Gupta,
1716–20.
"Hybrid Cloud Computing for Data Security
[17] 1. Navaneetha Krishnan Rajagopal, System," 2021 International Conference on
Mankeshva Saini, Rosario Huerta-Soto, Rosa Advancements in Electrical, Electronics,
Vílchez-Vásquez, J. N. V. R. Swarup Kumar, Communication, Computing and Automation
Shashi Kant Gupta, Sasikumar Perumal, (ICAECA), 2021, pp. 1-8, doi:
"Human Resource Demand Prediction and 10.1109/ICAECA52838.2021.9675493.
Configuration Model Based on Grey Wolf
[26] S. K. Gupta, B. Pattnaik, V. Agrawal, R. S. K.
Optimization and Recurrent Neural Network",
Boddu, A. Srivastava and B. Hazela, "Malware
Computational Intelligence and Neuroscience,
Detection Using Genetic Cascaded Support
vol. 2022, Article ID 5613407, 11 pages, 2022.
Vector Machine Classifier in Internet of
[Link]
Things," 2022 Second International
[18] 2. Navaneetha Krishnan Rajagopal, Naila Conference on Computer Science, Engineering
Iqbal Qureshi, S. Durga, Edwin Hernan and Applications (ICCSEA), 2022, pp. 1-6,
Ramirez Asis, Rosario Mercedes Huerta Soto, doi: 10.1109/ICCSEA54677.2022.9936404.
Shashi Kant Gupta, S. Deepak, "Future of
[27] Natarajan, R.; Lokesh, G.H.; Flammini, F.;
Business Culture: An Artificial Intelligence-
Premkumar, A.; Venkatesan, V.K.; Gupta,
Driven Digital Framework for Organization
S.K. A Novel Framework on Security and
Decision-Making Process", Complexity, vol.
Energy Enhancement Based on Internet of
2022, Article ID 7796507, 14 pages, 2022.
Medical Things for Healthcare 5.0.
[Link]
Infrastructures2023, 8, 22.
[19] Eshrag Refaee, Shabana Parveen, Khan [Link]
Mohamed Jarina Begum, Fatima Parveen, M. 2
Chithik Raja, Shashi Kant Gupta, Santhosh
[28] V. S. Kumar, A. Alemran, D. A. Karras, S.
Krishnan, "Secure and Scalable Healthcare
Kant Gupta, C. Kumar Dixit and B. Haralayya,
Data Transmission in IoT Based on Optimized
"Natural Language Processing using Graph
Routing Protocols for Mobile Computing
Neural Network for Text Classification," 2022
Applications", Wireless Communications and
International Conference on Knowledge
Mobile Computing, vol. 2022, Article ID
Engineering and Communication Systems
5665408, 12 pages, 2022.
(ICKES), Chickballapur, India, 2022, pp. 1-5,
[Link]
doi: 10.1109/ICKECS56523.2022.10060655.

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 391


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

[29] M. Sakthivel, S. Kant Gupta, D. A. Karras, A. Gupta. " AI-Based Smart Education System for
Khang, C. Kumar Dixit and B. Haralayya, a Smart City Using an Improved Self-Adaptive
"Solving Vehicle Routing Problem for Leap-Frogging Algorithm." CRC Press, 2022.
Intelligent Systems using Delaunay [Link]
Triangulation," 2022 International Conference [36] Rosak-Szyrocka, J., Żywiołek, J., & Shahbaz,
on Knowledge Engineering and M. (Eds.). (2023). Quality Management, Value
Communication Systems (ICKES), Creation and the Digital Economy (1st ed.).
Chickballapur, India, 2022, pp. 1-5, doi: Routledge.
10.1109/ICKECS56523.2022.10060807. [Link]
[30] S. Tahilyani, S. Saxena, D. A. Karras, S. Kant [37] Dr. Shashi Kant Gupta, Hayath T M., Lack of
Gupta, C. Kumar Dixit and B. Haralayya, it Infrastructure for ICT Based Education as an
"Deployment of Autonomous Vehicles in Emerging Issue in Online Education,
Agricultural and using Voronoi Partitioning," TTAICTE. 2022 July; 1(3): 19-24. Published
2022 International Conference on Knowledge online 2022 July,
Engineering and Communication Systems [Link]/10.36647/TTAICTE/01.03.A004
(ICKES), Chickballapur, India, 2022, pp. 1-5,
[38] Hayath T M., Dr. Shashi Kant Gupta,
doi: 10.1109/ICKECS56523.2022.10060773.
Pedagogical Principles in Learning and Its
[31] V. S. Kumar, A. Alemran, S. K. Gupta, B. Impact on Enhancing Motivation of Students,
Hazela, C. K. Dixit and B. Haralayya, TTAICTE. 2022 October; 1(2): 19-24.
"Extraction of SIFT Features for Identifying Published online 2022 July,
Disaster Hit areas using Machine Learning [Link]/10.36647/TTAICTE/01.04.A004
Techniques," 2022 International Conference
[39] Shaily Malik, Dr. Shashi Kant Gupta, “The
on Knowledge Engineering and
Importance of Text Mining for Services
Communication Systems (ICKES),
Management”, TTIDMKD. 2022 November;
Chickballapur, India, 2022, pp. 1-5, doi:
2(4): 28-33. Published online 2022 November
10.1109/ICKECS56523.2022.10060037.
[Link]/10.36647/TTIDMKD/02.04.A006
[32] V. S. Kumar, M. Sakthivel, D. A. Karras, S.
[40] Dr. Shashi Kant Gupta, Shaily Malik,
Kant Gupta, S. M. Parambil Gangadharan and
“Application of Predictive Analytics in
B. Haralayya, "Drone Surveillance in Flood
Agriculture”, TTIDMKD. 2022 November;
Affected Areas using Firefly Algorithm," 2022
2(4): 1-5. Published online 2022 November
International Conference on Knowledge
[Link]/10.36647/TTIDMKD/02.04.A001
Engineering and Communication Systems
(ICKES), Chickballapur, India, 2022, pp. 1-5, [41] Dr. Shashi Kant Gupta, Budi Artono,
doi: 10.1109/ICKECS56523.2022.10060857. “Bioengineering in the Development of
Artificial Hips, Knees, and other joints.
[33] Parin Somani, Sunil Kumar Vohra, Subrata
Ultrasound, MRI, and other Medical Imaging
Chowdhury, Shashi Kant Gupta.
Techniques”, TTIRAS. 2022 June; 2(2): 10–
"Implementation of a Blockchain-based Smart
15. Published online 2022 June
Shopping System for Automated Bill
[Link]/10.36647/TTIRAS/02.02.A002
Generation Using Smart Carts with
Cryptographic Algorithms." CRC Press, 2022. [42] Dr. Shashi Kant Gupta, Dr. A. S. A. Ferdous
[Link] Alam, “Concept of E Business Standardization
and its Overall Process” TJAEE 2022 August;
[34] Shivlal Mewada, Dhruva Sreenivasa
1(3): 1–8. Published online 2022 August
Chakravarthi, S. J. Sultanuddin, Shashi Kant
Gupta. "Design and Implementation of a Smart [43] A. Kishore Kumar, A. Alemran, D. A. Karras,
Healthcare System Using Blockchain S. Kant Gupta, C. Kumar Dixit and B.
Technology with A Dragonfly Optimization- Haralayya, "An Enhanced Genetic Algorithm
based Blowfish Encryption Algorithm." CRC for Solving Trajectory Planning of
Press, 2022. Autonomous Robots," 2023 IEEE
[Link] International Conference on Integrated
Circuits and Communication Systems
[35] Ahmed Muayad Younus, Mohanad S.S.
(ICICACS), Raichur, India, 2023, pp. 1-6, doi:
Abumandil, Veer P. Gangwar, Shashi Kant
10.1109/ICICACS57338.2023.10099994

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 392


© AUG 2023 | IRE Journals | Volume 7 Issue 2 | ISSN: 2456-8880

[44] S. K. Gupta, V. S. Kumar, A. Khang, B. [51] Shashi Kant Gupta, Alex Khang, Parin
Hazela, N. T and B. Haralayya, "Detection of Somani, Chandra Kumar Dixit, Anchal Pathak
Lung Tumor using an efficient Quadratic (2023). Data Mining Processes and Decision-
Discriminant Analysis Model," 2023 Making Models in Personnel Management
International Conference on Recent Trends in System (1st Ed.), CRC Press.
Electronics and Communication (ICRTEC), [Link]
Mysore, India, 2023, pp. 1-6, doi: [52] Alex Khang, Shashi Kant Gupta, Chandra
10.1109/ICRTEC56977.2023.10111903. Kumar Dixit, Parin Somani (2023). Data-
[45] S. K. Gupta, A. Alemran, P. Singh, A. Khang, driven Application of Human Capital
C. K. Dixit and B. Haralayya, "Image Management Databases, Big Data, and Data
Segmentation on Gabor Filtered images using Mining (1st Ed.), CRC Press.
Projective Transformation," 2023 International [Link]
Conference on Recent Trends in Electronics [53] Chandra Kumar Dixit, Parin Somani, Shashi
and Communication (ICRTEC), Mysore, Kant Gupta, Anchal Pathak (2023). Data-
India, 2023, pp. 1-6, doi: centric Predictive Modelling of Turnover Rate
10.1109/ICRTEC56977.2023.10111885. and New Hire in Workforce Management
[46] S. K. Gupta, S. Saxena, A. Khang, B. Hazela, System (1st Ed.), CRC Press.
C. K. Dixit and B. Haralayya, "Detection of [Link]
Number Plate in Vehicles using Deep Learning [54] Anchal Pathak, Chandra Kumar Dixit, Parin
based Image Labeler Model," 2023 Somani, Shashi Kant Gupta (2023). Prediction
International Conference on Recent Trends in of Employee’s Performance Using Machine
Electronics and Communication (ICRTEC), Learning (ML) Techniques (1st Ed.), CRC
Mysore, India, 2023, pp. 1-6, doi: Press. [Link]
10.1109/ICRTEC56977.2023.10111862.
[55] Worakamol Wisetsri, Varinder Kumar, Shashi
[47] S. K. Gupta, W. Ahmad, D. A. Karras, A. Kant Gupta, “Managerial Autonomy and
Khang, C. K. Dixit and B. Haralayya, "Solving Relationship Influence on Service Quality and
Roulette Wheel Selection Method using Human Resource Performance”, Turkish
Swarm Intelligence for Trajectory Planning of Journal of Physiotherapy and Rehabilitation,
Intelligent Systems," 2023 International Vol. 32, pp2, 2021.
Conference on Recent Trends in Electronics
and Communication (ICRTEC), Mysore,
India, 2023, pp. 1-5, doi:
10.1109/ICRTEC56977.2023.10111861.
[48] Shashi Kant Gupta, Olena Hrybiuk, NL
Sowjanya Cherukupalli, Arvind Kumar Shukla
(2023). Big Data Analytics Tools, Challenges
and Its Applications (1st Ed.), CRC Press.
ISBN 9781032451114
[49] Shobhna Jeet, Shashi Kant Gupta, Olena
Hrybiuk, Nupur Soni (2023). Detection of
Cyber Attacks in IoT-based Smart Cities using
Integrated Chain Based Multi-Class Support
Vector Machine (1st Ed.), CRC Press. ISBN
9781032451114
[50] Parin Somani, Shashi Kant Gupta, Chandra
Kumar Dixit, Anchal Pathak (2023). AI-based
Competency Model and Design in the
Workforce Development System (1st Ed.),
CRC Press.
[Link]

IRE 1704994 ICONIC RESEARCH AND ENGINEERING JOURNALS 393

You might also like