Missing Data Imputation Techniques Review
Missing Data Imputation Techniques Review
[Link]
ISSN Online: 2327-5227
ISSN Print: 2327-5219
Majed Alwateer1, El-Sayed Atlam1,2, Mahmoud Mohammed Abd El-Raouf3, Osama A. Ghoneim4,
Ibrahim Gad2
1
Department of Computer Science, College of Computer Science and Engineering, Taibah University, Yanbu, Saudi Arabia
2
Computer Science Department, Faculty of Science, Tanta University, Tanta, Egypt
3
Basic and Applied Science Institute, College of Engineering and Technology, Arab Academy for Science and Technology (AAST),
Alexandria, Egypt
4
Department of Computer Science, Faculty of Computers and informatics, Tanta University, Tanta, Egypt
Keywords
Missing Data, Machine Learning, Prediction, Deep Learning, Imputation
1. Introduction
Missing data can occur due to various reasons, such as participant non-response
in surveys, equipment malfunction in experimental settings, or data entry errors
[1] [2]. The presence of missing data can significantly impact statistical analyses
and lead to incorrect conclusions if not properly addressed [3]. Missing data is a
pervasive issue in empirical research across various disciplines, including social
sciences, medical research, and data science [4]. It occurs when no value is avail-
able for a variable in an observation, potentially leading to incomplete datasets
that can compromise the validity and reliability of statistical analyses [5] [6].
Missing values can significantly impact your analysis: They can introduce bias
if not handled properly. Many machine learning algorithms can’t handle missing
values that are out of the box. They can lead to the loss of important information
if instances with missing values are simply discarded. Improperly handled missing
values can lead to incorrect conclusions or predictions.
Missing values can sneak into your data for a variety of reasons. Here are some
common reasons: Data Entry Errors: Sometimes, it’s just human error. Someone
might forget to input a value or accidentally delete one. Sensor Malfunctions: In
IoT or scientific experiments, a faulty sensor might fail to record data at certain
times [7]. Survey Non-Response: In surveys, respondents might skip questions
they are uncomfortable answering or don’t understand. Merged Datasets: When
combining data from multiple sources, some entries might not have correspond-
ing values in all datasets. Data Corruption: During data transfer or storage, some
values might get corrupted and become unreadable. Intentional Omissions: Some
data might be intentionally left out due to privacy concerns or irrelevance. Sam-
pling Issues: The data collection method might systematically miss certain types
of data. Time-Sensitive Data: In time series data, values might be missing for pe-
riods when data wasn’t collected (e.g., weekends, holidays) [8].
Researchers typically encounter three main types of missing data mechanisms, as
defined by Rubin [4]: 1) Missing Completely at Random (MCAR): The probability
of missing data is unrelated to both observed and unobserved variables. 2) Missing
at Random (MAR): The probability of missing data depends on observed variables
but not on unobserved variables. 3) Missing Not at Random (MNAR): The probabil-
ity of missing data depends on unobserved variables, including the missing data itself.
Addressing missing data is of paramount importance in research for several
reasons: 1) Bias reduction: Ignoring missing data or using simplistic methods like
complete case analysis can lead to biased estimates and incorrect inferences [9].
Proper handling of missing data helps minimize this bias and improve the accu-
racy of research findings. 2) Statistical power: Missing data reduces the effective
sample size, leading to decreased statistical power. Appropriate imputation tech-
niques can help maintain or even increase power by utilizing all available infor-
mation [10]. 3) Generalizability: Incomplete datasets may not be representative of
the population of interest, potentially limiting the generalizability of research
findings. Addressing missing data can help improve the external validity of results
[11]. 4) Ethical considerations: In clinical trials and other human subjects re-
search, ignoring missing data may lead to the waste of valuable data that partici-
pants have provided, raising ethical concerns. 5) Regulatory compliance: In some
fields, such as clinical trials, regulatory bodies require proper handling and reporting
values [4]. MNAR is the most challenging type of missing data to handle. Standard
imputation methods can yield biased results, necessitating specialized techniques
or sensitivity analyses to address the issue effectively [32].
MNAR occurs when the probability of missing data depends on unobserved
variables or the missing values themselves [4]. This is the most challenging type
of missingness to address and often requires specialized techniques. MNAR oc-
curs when: P ( R | Y ) ≠ P ( R | Yobs ) .
An example of MNAR is in a mental health survey, participants with severe
depression may be less likely to complete questions about their symptoms, with
this likelihood directly related to the severity of their undisclosed condition. Sim-
ilarly, individuals with high incomes might be less inclined to report their income
in a survey.
In practice, it is often impossible to definitively determine whether data are
MAR or MNAR based solely on observed data [33]. Therefore, researchers often
must rely on subject-matter knowledge and conduct sensitivity analyses to assess
the robustness of their findings under different missing data assumptions [17].
The problem of missing data imputation arises when a dataset contains unob-
=
served values for certain features. Let X ( X 1 , X 2 , , X n ) ∈ n represent a ran-
dom vector, where each X i corresponds to a feature and follows a distribution
P ( X ) . A binary mask vector M = ( M 1 , M 2 , , M n ) indicates the presence or
4. Methods
This study evaluated several imputation methods for handling missing data, rang-
ing from simple statistical techniques to more sophisticated deep learning ap-
proaches. Specific methods included.
4.1.2. MissForest
This iterative approach uses mean/mode imputation to initialize the dataset [16].
Then, a random forest is trained to predict the missing values in each feature,
iteratively refining the imputed values until convergence (defined by a lack of im-
provement in the imputed matrix). Convergence criteria included a maximum of
20 iterations and 100 trees. The difference between successive imputed matrices
( M imp _ new and M imp _ old ) for numerical ( N ) and categorical ( F ) features was
measured as:
∑ ( M new )
imp imp 2
− M old
δ N = j∈N (1)
∑ j∈N ( M new )
imp 2
i =n
∑ j∈F ∑ i=I
1 M imp
≠M imp
δF = new old
(2)
FNA
(e.g., age, gender, income) based on the variables with available data. 2) For each
missing value, find a donor record that matches the criteria. 3) Replace the miss-
ing value with the corresponding value from the donor record. Advantages: Sim-
ple, can be effective for handling missing values in categorical variables.
1 n
V = ∑Vi
n i =1
(
f enc ( X ) = s X ⋅ W T + b )
(
f dec ( Y ) = s Y ⋅ W T + b )
where Y = f enc ( X ) and Z = f dec ( Y ) are the hidden and output vectors, re-
spectively; W and b are the encoder weights and bias; and W and b are
the decoder weights and bias. Training minimizes the reconstruction error be-
tween X and Z .
1 N +C
BCE =− ∑ yi ⋅ log ( p ( yi ) ) + (1 − yi ) ⋅ log (1 − p ( yi ) )
N =i N +1
Loss RMSE + BCE
=
Accounts for
Handles complex
uncertainty in
Multiple missing data patterns, Can be computationally “IterativeImputer”
imputation, can provide Regression/
Imputation accounts for expensive, requires (scikit-learn),
more accurate estimates Classification
(MI) uncertainty in specialized software. “fancyimpute”
than single imputation
imputation.
methods.
Primarily for
Simple, can be Can introduce bias if the
categorical data, “KNNImputer”
Hot-Deck effective for handling donor records are not truly
replaces missing values (scikit-learn) can Classification
Imputation missing values in similar to the record with
with values from a be adapted
categorical variables. the missing value.
similar donor record.
neighbors.
Missing values were imputed using the average (numerical features) or mode
(categorical features) of the k nearest neighbors in feature space, using Euclid-
ean distance. For example, k = 4 . Formally, for a sample S ( X , Y ,0 ) with four
nearest neighbors
= N 4 {=
( X i , Yi ,1) | i 1, 2,3, 4} , the imputed value Y is calcu-
lated as:
arg max
{ }
z ∑ ( X i ,Yi ,1)∈N 4 1( Yi = z ) if Y is categorical
Y =
1
∑ i4=1Yi if Y is numerical
4
where z ∈ {0,1} and 1(Yi = z ) is an indicator function. The Euclidean distance
between points x and y is defined as:
d ( x, y ) ∑ i=1 ( xi − yi )
n 2
=
easy to interpret, as you can visualize how decisions are made. The disadvantages
of Decision Trees Imputation: 1) Overfitting: Decision trees can easily overfit to
the training data, especially if not properly pruned. 2) Sensitivity to Noise: Outli-
ers can affect the structure of the tree, impacting the imputation results.
The advantages of Random Forests Imputation: 1) Improved Accuracy: Ran-
dom forests generally provide better accuracy than single decision trees due to
their ensemble nature, reducing overfitting and variance. 2) Robust to Outliers:
The averaging mechanism makes random forests less sensitive to outliers com-
pared to individual trees. 3) Handles Large Datasets: Random forests can effec-
tively manage large datasets with high dimensionality. The disadvantages of Ran-
dom Forests Imputation: 1) Complexity: The model is less interpretable than a
single decision tree, as it’s harder to visualize how predictions are made. 2) Com-
putationally Intensive: Training multiple trees can be resource-intensive, espe-
cially for large datasets.
tions but may be challenging to train and tune. The Generative Adversarial Im-
putation Network (GAIN) [50] employs a generative adversarial network (GAN)
architecture. Unlike standard GANs, the discriminator in GAIN does not classify
the entire generated output as real or fake; instead, it classifies each individual
variable as either imputed or observed. Convergence is achieved when the gener-
ator produces imputations indistinguishable from the true data distribution. A
“hint” mechanism augments the discriminator’s input with partial information
about the missing values (M), represented by the hint vector H. This hint is typi-
cally a proportion of M (e.g., 90% identical). The generator then learns to impute
the remaining values. The original work demonstrates that insufficient hints lead
to multiple optimal generator outputs.
Formally, given a random vector Z, the generator produces an imputed dataset
X i , from which X i is derived (Equation (1)). The discriminator loss function,
LD , is defined as:
,H )
LD ( M , M= ∑ M i ⋅ log ( M i ) + (1 − M i ) ⋅ log (1 − M i )
i:H i =0
where M is the true mask vector, M is the generated mask vector, and H is
the hint vector. The summation is restricted to indices where H i = 0 to prevent
overfitting to the hint. The discriminator is trained to minimize this loss:
batchsize
min D − ∑ LD ( M i , M i , H i )
i =1
The generator loss function comprises two terms: LG , which measures the
generator’s ability to deceive the discriminator, and LM , which quantifies the ac-
curacy of the imputation for observed values:
LG ( M , M , H ) =
− ∑ (1 − M i ) ⋅ log ( M i )
i:H i =0
d
(
LM X , X = )
−∑ M i ⋅ Diff X i , X i ( )
i =1
) ( )
X − X 2 if X i is numerical
(
Diff X , X = i i
can be employed to impute missing values in time series data by leveraging tem-
poral dependencies [51] [52]. They are particularly useful when the data exhibits
trends and seasonality. An ARIMA model is denoted as ARIMA (p, d, q), where
“p” represents the order of the autoregressive (AR) component, “d” represents the
degree of difference required to make the time series stationary, and “q” repre-
sents the order of the moving average (MA) component. The AR component
models the relationship between the current observation and previous observa-
tions; the MA component models the relationship between the current observa-
tion and past forecast errors, and differencing (d) removes trends and makes the
series stationary. A general ARIMA (p, d, q) model can be represented by the fol-
lowing equation:
φ ( B )(1 − B ) yt =
θ ( B ) t
d
crucial to ensure that the imputed data maintains the integrity and reliability of
the original data. Several metrics are commonly used to assess the effectiveness of
imputation techniques. The choice of evaluation metric depends on the specific
objective of the imputation and the nature of the data. These evaluation metrics
can be broadly categorized into two groups, as shown in Table 4.
1 n TP + TN
∑ ( yi − yˆi )
2
Mean Squared Error (MSE) MSE
= Accuracy Accuracy =
n i =1 TP + TN + FP + FN
1 n TP
Mean Absolute Error (MAE) MAE
= ∑ yi − yˆi
n i =1
Recall (Sensitivity) Recall =
TP + FN
c) Informed Consent: Participants should be informed about how their data, in-
cluding any imputed values, will be used in research. This includes potential im-
plications for privacy and the integrity of their responses. d) Appropriateness of
Methods: Different imputation methods (mean, median, predictive modeling,
etc.) have different assumptions. Choosing the right method is crucial to avoid
distorting the data and the conclusions drawn from it. e) Impact on Decision-
Making: The results derived from imputed data can influence policy or clinical
decisions. Researchers should ensure that their imputation practices do not lead
to harmful outcomes. f) Equity: Consider whether the imputation methods used
could disproportionately affect certain groups. Ensuring that imputation methods
do not reinforce existing inequalities is vital. g) Ethical Oversight: It’s beneficial
to have an ethical review process in place to assess the imputation strategies and
their potential implications for participants and broader societal contexts. h) Data
Integrity: Strive to maintain the integrity of the original dataset. Imputation
should not compromise the authenticity of the data, and researchers should be
mindful of the limitations that come with imputed values. i) Training and Exper-
tise: Ensure that those involved in the imputation process have the necessary
training and understanding of the ethical implications of their work.
9. Conclusion
As data collection and analysis continue to grow in importance across various do-
mains, the field of missing data imputation is likely to see further advancements
and innovations. This review has provided a comprehensive overview of missing
data imputation techniques, from traditional methods, and statistical methods to
advanced machine learning approaches. Key observations highlight: 1) The criti-
cal role of understanding missing data mechanisms. 2) The trade-offs between
simple and complex imputation methods. 3) There is a need for careful evaluation
and selection of imputation techniques. and 4) The potential of machine learning
and deep learning approaches for handling complex missing data patterns. Future
research should focus on developing more robust, efficient, and adaptable impu-
tation methods that can handle the increasing complexity and scale of modern
datasets while addressing ethical concerns and preserving data integrity.
Conflicts of Interest
The authors declare no conflicts of interest regarding the publication of this paper.
References
[1] Enders, C.K. (2022) Applied Missing Data Analysis. Guilford Publications.
[2] Mitchel, J.T., Kim, Y.J., Choi, J., Park, G., Cappi, S., Horn, D., et al. (2011) Evaluation
of Data Entry Errors and Data Changes to an Electronic Data Capture Clinical Trial
Database. Drug Information Journal, 45, 421-430.
[Link]
[3] Schafer, J.L. and Graham, J.W. (2002) Missing Data: Our View of the State of the Art.
Psychological Methods, 7, 147-177. [Link]
[4] Little, R. and Rubin, D. (2019) Statistical Analysis with Missing Data. Third Edition,
Wiley. [Link]
[5] Lazar, N.A. (2003) Statistical Analysis with Missing Data. Technometrics, 45, 364-
365. [Link]
[6] Hajjar, S. (2018) Statistical Analysis: Internal-Consistency Reliability and Construct
Validity. International Journal of Quantitative and Qualitative Research Methods, 6,
27-38.
[7] Noor, T.H., Atlam, E., Almars, A.M., Noor, A. and Malki, A.S. (2023) An IoT-Based
Energy Conservation Smart Classroom System. Intelligent Automation & Soft Com-
puting, 35, 3785-3799. [Link]
[8] Zhan, Y., Xia, Y., Liu, Y., Li, F. and Wang, Y. (2017) Incentive-Aware Time-Sensitive
Data Collection in Mobile Opportunistic Crowdsensing. IEEE Transactions on Ve-
hicular Technology, 66, 7849-7861. [Link]
[9] Molenberghs, G. and Verbeke, G. (2013) Missing Data. In: Scott, M.A., Simonoff, J.S.
and Marx, B.D., Eds., The SAGE Handbook of Multilevel Modeling, SAGE Publica-
tions Ltd, 403-424. [Link]
[10] Graham, J.W. (2009) Missing Data Analysis: Making It Work in the Real World. An-
nual Review of Psychology, 60, 549-576.
[Link]
[11] Sterne, J.A.C., White, I.R., Carlin, J.B., Spratt, M., Royston, P., Kenward, M.G., et al.
(2009) Multiple Imputation for Missing Data in Epidemiological and Clinical Re-
search: Potential and Pitfalls. BMJ, 338, b2393-b2393.
[Link]
[12] Masconi, K.L., Matsha, T.E., Echouffo-Tcheugui, J.B., Erasmus, R.T. and Kengne,
A.P. (2015) Reporting and Handling of Missing Data in Predictive Research for Prev-
alent Undiagnosed Type 2 Diabetes Mellitus: A Systematic Review. EPMA Journal, 6,
Article No. 7. [Link]
[13] Ferris, J.A. (2009) Missing Data: A Gentle Introduction. Drug and Alcohol Review,
28, 90-91. [Link]
[14] Enders, C.K. (2017) Multiple Imputation as a Flexible Tool for Missing Data Han-
dling in Clinical Research. Behaviour Research and Therapy, 98, 4-18.
[Link]
[15] Bertsimas, D., Pawlowski, C. and Zhuo, Y.D. (2018) From Predictive Methods to
Missing Data Imputation: An Optimization Approach. Journal of Machine Learning
Research, 18, 1-39.
[16] Stekhoven, D.J. and Bühlmann, P. (2011) Missforest—Non-Parametric Missing
Value Imputation for Mixed-Type Data. Bioinformatics, 28, 112-118.
[Link]
[17] Van Buuren, S. (2018) Flexible Imputation of Missing Data. CRC Press.
[18] Rubin, D.B. (1976) Inference and Missing Data. Biometrika, 63, 581-592.
[Link]
[19] Doreswamy, Gad, I. and Manjunatha, B.R. (2017) Performance Evaluation of Predic-
tive Models for Missing Data Imputation in Weather Data. 2017 International Con-
ference on Advances in Computing, Communications and Informatics (ICACCI),
Udupi, 13-16 September 2017, 1327-1334.
[Link]
[20] Hosahalli, D. and Gad, I. (2018) A Generic Approach of Filling Missing Values in
NCDC Weather Stations Data. 2018 International Conference on Advances in
[35] Andridge, R.R. and Little, R.J.A. (2010) A Review of Hot Deck Imputation for Survey
Non-Response. International Statistical Review, 78, 40-64.
[Link]
[36] van Buuren, S. (2007) Multiple Imputation of Discrete and Continuous Data by Fully
Conditional Specification. Statistical Methods in Medical Research, 16, 219-242.
[Link]
[37] Aidos, H. and Tomas, P. (2021) Neighborhood-Aware Autoencoder for Missing
Value Imputation. 2020 28th European Signal Processing Conference (EUSIPCO),
Amsterdam, 18-21 January 2021, 1542-1546.
[Link]
[38] Rubin, D.B. (1987) Multiple Imputation for Nonresponse in Surveys. Wiley.
[Link]
[39] Rombach, I., Gray, A.M., Jenkinson, C., Murray, D.W. and Rivero-Arias, O. (2018)
Multiple Imputation for Patient Reported Outcome Measures in Randomised Con-
trolled Trials: Advantages and Disadvantages of Imputing at the Item, Subscale or
Composite Score Level. BMC Medical Research Methodology, 18, Article No. 87.
[Link]
[40] Murti, D.M.P., Pujianto, U., Wibawa, A.P. and Akbar, M.I. (2019) K-Nearest Neigh-
bor (K-NN) Based Missing Data Imputation. 2019 5th International Conference on
Science in Information Technology (ICSITech), Yogyakarta, 23-24 October 2019, 83-
88. [Link]
[41] Troyanskaya, O., Cantor, M., Sherlock, G., Brown, P., Hastie, T., Tibshirani, R., et al.
(2001) Missing Value Estimation Methods for DNA Microarrays. Bioinformatics, 17,
520-525. [Link]
[42] Gad, I., Elmezain, M., Alwateer, M.M., Almaliki, M., Elmarhomy, G. and Atlam, E.
(2023) Breast Cancer Diagnosis Using a Machine Learning Model and Swarm Intel-
ligence Approach. 2023 1st International Conference on Advanced Innovations in
Smart Cities (ICAISC), Jeddah, 23-25 January 2023, 1-5.
[Link]
[43] Mallinson, H. and Gammerman, A. (2003) Imputation Using Support Vector Ma-
chines. University of London Egham, UK: Department of Computer Science Royal
Holloway.
[44] Noor, T.H., Almars, A., Gad, I., Atlam, E. and Elmezain, M. (2022) Spatial Impres-
sions Monitoring during COVID-19 Pandemic Using Machine Learning Techniques.
Computers, 11, Article 52. [Link]
[45] Vincent, P., Larochelle, H., Bengio, Y. and Manzagol, P. (2008) Extracting and Com-
posing Robust Features with Denoising Autoencoders. Proceedings of the 25th inter-
national conference on Machine learning—ICML’08, Helsinki, 5-9 July 2008, 1096-
1103. [Link]
[46] Noor, T.H., Almars, A.M., Atlam, E. and Noor, A. (2022) Deep Learning Model for
Predicting Consumers’ Interests of IoT Recommendation System. International Jour-
nal of Advanced Computer Science and Applications, 13, 161-170.
[Link]
[47] Elmezain, M., Malki, A., Gad, I. and Atlam, E. (2022) Hybrid Deep Learning Model-
Based Prediction of Images Related to Cyberbullying. International Journal of Ap-
plied Mathematics and Computer Science, 32, 323-334.
[Link]
[48] Shahbazian, R. and Trubitsyna, I. (2022) DEGAIN: Generative-Adversarial-Net-
work-Based Missing Data Imputation. Information, 13, Article 575.
[Link]
[49] Malki, A., Atlam, E. and Gad, I. (2022) Machine Learning Approach of Detecting
Anomalies and Forecasting Time-Series of IoT Devices. Alexandria Engineering
Journal, 61, 8973-8986. [Link]
[50] Yoon, J., Jordon, J. and van der Schaar, M. (2018) GAIN: Missing Data Imputation
Using Generative Adversarial Nets, Proceedings of the 35th International Conference
on Machine Learning, Stockholm, 10-15 July 2018, 5689-5698.
[Link]
[51] Chhabra, G. (2023) Comparison of Imputation Methods for Univariate Time Series.
International Journal on Recent and Innovation Trends in Computing and Commu-
nication, 11, 286-292. [Link]
[52] Malki, A., Atlam, E., Hassanien, A.E., Ewis, A., Dagnew, G. and Gad, I. (2022)
SARIMA Model-Based Forecasting Required Number of COVID-19 Vaccines Glob-
ally and Empirical Analysis of Peoples’ View towards the Vaccines. Alexandria Engi-
neering Journal, 61, 12091-12110. [Link]
[53] Collins, L.M., Schafer, J.L. and Kam, C. (2001) A Comparison of Inclusive and Re-
strictive Strategies in Modern Missing Data Procedures. Psychological Methods, 6,
330-351. [Link]
[54] Emmanuel, T., Maupong, T., Mpoeleng, D., Semong, T., Mphago, B. and Tabona, O.
(2021) A Survey on Missing Data in Machine Learning. Journal of Big Data, 8, Article
No. 140. [Link]
[55] Gustems-Carnicer, J. and Calderón, C. (2012) Coping Strategies and Psychological
Well-Being among Teacher Education Students. European Journal of Psychology of
Education, 28, 1127-1140. [Link]
[56] Atlam, E., Masud, M., Rokaya, M., Meshref, H., Gad, I. and Almars, A.M. (2024)
EASDM: Explainable Autism Spectrum Disorder Model Based on Deep Learning.
Journal of Disability Research, 3, 1-15. [Link]
[57] Masud, M., Almars, A.M., Rokaya, M.B., Meshref, H., Gad, I. and Atlam, E. (2024) A
Novel Light-Weight Convolutional Neural Network Model to Predict Alzheimer’s
Disease Applying Weighted Loss Function. Journal of Disability Research, 3, 1-10.
[Link]
[58] McMahon, P., Zhang, T. and Dwight, R.A. (2020) Approaches to Dealing with Miss-
ing Data in Railway Asset Management. IEEE Access, 8, 48177-48194.
[Link]
[59] Bennin, K.E., Tahir, A., MacDonell, S.G. and Börstler, J. (2021) An Empirical Study
on the Effectiveness of Data Resampling Approaches for Cross-Project Software De-
fect Prediction. IET Software, 16, 185-199. [Link]
[60] Hsu, C.C. and Sandford, B.A. (2019) Minimizing Non-Response in the Delphi Pro-
cess: How to Respond to Non-Response. Practical Assessment, Research, and Evalu-
ation, 12, 17.
[61] Carpenter, J.R., Roger, J.H. and Kenward, M.G. (2013) Analysis of Longitudinal Tri-
als with Protocol Deviation: A Framework for Relevant, Accessible Assumptions, and
Inference via Multiple Imputation. Journal of Biopharmaceutical Statistics, 23, 1352-
1371. [Link]
[62] Adnan, F.A., Jamaludin, K.R., Wan Muhamad, W.Z.A. and Miskon, S. (2022) A Re-
view of the Current Publication Trends on Missing Data Imputation over Three Dec-
ades: Direction and Future Research. Neural Computing and Applications, 34,
18325-18340. [Link]
[63] Pedersen, A., Mikkelsen, E., Cronin-Fenton, D., Kristensen, N., Pham, T.M.,
Pedersen, L., et al. (2017) Missing Data and Multiple Imputation in Clinical Epide-
miological Research. Clinical Epidemiology, 9, 157-166.
[Link]
[64] Li, T., Sahu, A.K., Talwalkar, A. and Smith, V. (2020) Federated Learning: Challenges,
Methods, and Future Directions. IEEE Signal Processing Magazine, 37, 50-60.
[Link]
[65] Singh, P., Singh, M.K., Singh, R. and Singh, N. (2022) Federated Learning: Chal-
lenges, Methods, and Future Directions. In: Yadav, S.P., Bhati, B.S., Mahato, D.P. and
Kumar, S., Eds., Federated Learning for IoT Applications, Springer International
Publishing, 199-214. [Link]
[66] Ma, C., Tschiatschek, S., Palla, K., Hernández-Lobato, J.M., Nowozin, S. and Zhang,
C. (2018) Eddi: Efficient Dynamic Discovery of High-Value Information with Partial
Vae. arXiv: 1809.11142.
[67] Sultana, Z., Akter, S. and Yeasmin, A. (2022) Transfer Learning Approach Applied to
Data Imputation. University of Liberal Arts Bangladesh.
[Link]
[68] Lyu, L., Hu, Y., Wang, N., Zhou, X. and Fang, M. (2022) Application of Deep Learn-
ing and Transfer Learning in Continuous Missing Value Imputation of Water Quality
Data. 2022 IEEE 8th International Conference on Computer and Communications
(ICCC), Chengdu, 9-12 December 2022, 1211-1216.
[Link]
[69] Anagnostopoulos, I. (2018) Fintech and Regtech: Impact on Regulators and Banks.
Journal of Economics and Business, 100, 7-25.
[Link]
[70] Gupta, M. and Gupta, B. (2020) A New Scalable Approach for Missing Value Impu-
tation in High-Throughput Microarray Data on Apache Spark. International Journal
of Data Mining and Bioinformatics, 23, 79-100.
[Link]
[71] Petrozziello, A., Jordanov, I. and Sommeregger, C. (2018) Distributed Neural Net-
works for Missing Big Data Imputation. 2018 International Joint Conference on Neu-
ral Networks (IJCNN), Rio de Janeiro, 8-13 July 2018, 1-8.
[Link]
[72] Rajkomar, A., Hardt, M., Howell, M.D., Corrado, G. and Chin, M.H. (2018) Ensuring
Fairness in Machine Learning to Advance Health Equity. Annals of Internal Medi-
cine, 169, 866-872. [Link]
[73] Chandler, R.K., Fletcher, B.W. and Volkow, N.D. (2009) Treating Drug Abuse and
Addiction in the Criminal Justice System. JAMA, 301, 183-190.
[Link]
[74] De Pau, M., Vruggink, R., Vandevelde, S. and Vander Laenen, F. (2023) Culturally
Sensitive Forensic Mental Healthcare for Racialized People Labeled as Not Criminally
Responsible: A Scoping Review. International Journal of Forensic Mental Health, 22,
276-288. [Link]
KNN is more effective than mean/median imputation in preserving data relationships, as it considers the proximity of data points and maintains the data structure. Mean/median imputation, while simpler, can introduce bias and fails to account for inter-variable relationships, especially in skewed datasets .
In healthcare, imputation techniques address incomplete electronic health records by maintaining data integrity, enabling accurate analyses, which in turn improve diagnostic accuracy and patient care outcomes. Methods like Multiple Imputation and advanced algorithms tailored to healthcare datasets are vital to achieving these goals .
Computational complexity significantly impacts MI and KNN implementations, making them resource-intensive, especially for large datasets and complex models. This complexity can limit their usability in practice, as they require specialized software and extended time for processing, hindering scalability and efficiency .
Bias occurs in imputation when data is NMAR because the missingness is directly related to unobserved data values or other variables, leading to systematic errors if these variables are not accounted for, skewing the imputed results .
Assuming a specific data distribution, such as normality, helps streamline the imputation process but can lead to biased imputed values if the actual data distribution differs significantly. This discrepancy can misrepresent data characteristics, affecting the validity of subsequent analyses .
The NOCB method preserves trends and can reduce bias by using more representative future data, potentially offering more accurate results compared to LOCF. However, it assumes continuity which might not be valid, introduces temporal distortion bias, and is generally more complex and harder to justify than LOCF .
MI generates multiple plausible estimates to account for uncertainty in missing data, leading to more realistic confidence intervals and p-values. It reduces bias as multiple datasets are imputed, especially when data is not missing at random, offering more accurate parameter estimates and statistical tests than single imputation .
LOCF is suboptimal when data lacks a strong temporal trend or when consecutive data points are missing, as it can introduce bias and propagate errors by not reflecting actual variability in the dataset .
Hot Deck Imputation is advantageous for handling categorical missing data effectively by preserving data distribution through similar donor records. However, it faces challenges in scalability and complexity when applied to large datasets, requiring careful definition of matching criteria and donor selection .
In medical data, domain knowledge is crucial for identifying relationships between variables and understanding the causes of missing data, guiding the choice of imputation method and improving the accuracy and relevance of the imputed values .