0% found this document useful (0 votes)
7 views6 pages

Multiple Imputation in Respiratory Studies

Uploaded by

Chua Su Jen
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views6 pages

Multiple Imputation in Respiratory Studies

Uploaded by

Chua Su Jen
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

bs_bs_banner

INVITED REVIEW SERIES:


MODERN STATISTICAL METHODS IN RESPIRATORY MEDICINE
SERIES EDITORS: RORY WOLFE AND MICHAEL ABRAMSON

Introduction to multiple imputation for dealing with missing data

KATHERINE J. LEE1,2 AND JULIE A. SIMPSON3

1
Clinical Epidemiology and Biostatistics Unit, Murdoch Childrens Research Institute, 2Department of Paediatrics, The
University of Melbourne, 3Centre for Molecular, Environmental, Genetic & Analytic Epidemiology, Melbourne School of
Population and Global Health, The University of Melbourne, Melbourne, Victoria, Australia

ABSTRACT INTRODUCTION
Missing data are common in both observational and
experimental studies. Multiple imputation (MI) is a Missing data are common in both experimental and
two-stage approach where missing values are imputed observational studies of respiratory health. The
a number of times using a statistical model based on problem of missing data can arise in a variety of
the available data and then inference is combined forms, from item non-response, where an individual
across the completed datasets. This approach is has missing data on a particular measure (e.g. when
becoming increasingly popular for handling missing a participant refuses to provide a blood sample for
data. In this paper, we introduce the method of MI, as cytokine measurement), to unit non-response,
well as a discussion surrounding when MI can be a where a participant does not have data for a range of
useful method for handling missing data and the draw- measures (e.g. a participant missing a wave of data
backs of this [Link] illustrate MI when exploring collection in a longitudinal study). It can also occur
the association between current asthma status and in any number of variables, including the outcome,
forced expiratory volume in 1 s after adjustment for the exposure of interest and in confounding vari-
potential confounders using data from a population- ables required for adjustment, all of which can lead
based longitudinal cohort study. to issues for the analysis of interest (see Wolfe and
Abramson1, and Kasza and Wolfe2 which provide a
Key words: experimental study, missing data, multiple impu-
tation, observational study.
prelude to this paper). For example, the study by
Vermeulen et al. considers the scenario where there
Abbreviations: FEV1, forced expiratory volume in 1 s; MI, are missing data on health-related quality of life at
multiple imputation; SD, standard deviation; TAHS, Tasmanian various follow-up times within a longitudinal study
Longitudinal Health Study. of lung transplantation.3 This would mean missing
outcome data if the research question was around
quality of life following lung transplantation, or
missing covariate data if the research question
was around the association between quality of life
following lung transplantation and longer term
outcomes.
Correspondence: Katherine Lee, Clinical Epidemiology and The most common method of dealing with missing
Biostatistics Unit, Murdoch Childrens Research Institute, data is to exclude participants who have one or more
Flemington Road, Parkville, Melbourne, Vic. 3052, Australia. missing values, a so-called complete case analysis.
Email: [Link]@[Link]
The Authors: Dr Katherine Lee is a biostatistician with 11 years
Excluding participants in this way can introduce
of experience in clinical and statistical research and over 65 selection bias as participants who have complete data
peer-reviewed publications. Her main areas of expertise are clini- may be different to those with missing data. This
cal trials and the method of multiple imputation for dealing with means that the sample used for analysis may no
missing data. Associate Professor Julie Simpson is a biostatis- longer be representative of the population of interest.
tician with 20 years of experience in clinical and population health Complete case analysis can also be inefficient as it can
research and over 130 peer-reviewed papers. Her main area of
expertise is the application of nonlinear mixed effects models to
mean excluding a large number of participants. A
pharmacokinetic-pharmacodynamic data and statistical methods number of methods have been suggested in the litera-
for handling missing data in longitudinal cohort studies. ture which enable participants with missing data to
Received 6 October 2013; accepted 13 October 2013 be included in the analysis, such as last observation
© 2013 The Authors Respirology (2014) 19, 162–167
Respirology © 2013 Asian Pacific Society of Respirology doi: 10.1111/resp.12226
14401843, 2014, 2, Downloaded from [Link] by National Health And Medical Research Council, Wiley Online Library on [13/11/2025]. See the Terms and Conditions ([Link] on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Multiple imputation for missing data 163

carried forward, the missing indicator method, mean Incomplete Impute missing data Analyse each imputed Combine
dataset mulple mes dataset and esmate esmates
substitution, and regression imputation. However, parameter of interest
these unprincipled, simple approaches have also
been shown to lead to biased results.4,5 Multiple impu-
tation (MI) is an alternative approach whereby 1
missing values are imputed based on the observed
data, repeating this process a number of times to
account for the uncertainty in the imputed values.
This approach is becoming increasingly popular 2
to deal with missing data as it has the potential to
correct the bias in the complete case and alternative

...
analyses.6,7

...
Medical journals now often request that research-
ers indicate the amount of missing data in their m
study, assess differences between those individuals
with complete data and those with missing data,
and explain how the missing data were handled Figure 1 Illustration of the method of multiple imputation. Each
in the statistical analyses.8 Furthermore, leading box represents a data value where the columns are variables and
journals are starting to request justification for the the rows are individuals. Blank spaces represent the missing
statistical approach chosen by the authors to deal values. In this figure, β̂i is the estimate of interest from the
with the missing data, and state that additional sen- completed dataset number i, in our case the regression coeffi-
sitivity analyses may be requested when there are cient for current asthma status in a multivariable linear regres-
extensive missing data.9 This increased awareness sion model for forced expiratory volume in 1 s, βMI is the
regarding the impact of missing data means that estimate obtained from multiple imputation, and m is the
number of imputed datasets.
researchers need to consider the best way to handle
missing data in their analyses, and justify their deci-
sion regarding the approach chosen to deal with
missing data. WHAT IS MI?
In this paper, we provide an introduction to MI as a
statistical method for dealing with missing data, illus- MI is a two-stage process. In stage 1, the missing
trating the approach using data from a population- values are imputed by sampling from an imputation
based longitudinal cohort study of respiratory health. model. This imputation model should include all vari-
We also discuss scenarios where MI can be a useful ables that are in the analysis model (outcome, expo-
method for handling missing data and present some sure, confounders), as well as additional (at least
of the drawbacks of this approach. partially) observed variables that are not included in
the analysis model, but are associated with the vari-
able(s) with missing data. These additional variables
are known as auxiliary variables. The imputation
MOTIVATING EXAMPLE process is repeated a number of times, that is, a
number of completed datasets are created (Fig. 1), to
As an example, we consider the question of whether capture the uncertainty in the missing values.
current asthma status is associated with forced In the second stage, the epidemiological analysis of
expiratory volume in 1 s (FEV1), after adjustment for interest is performed on each of the ‘completed’
age, gender, socio-economic status, smoking status, datasets (observed plus imputed values) (Fig. 1).11 The
height and waist circumference using multivariable final MI estimate is simply the average of the esti-
linear regression (see Kasza and Wolfe2 for a descrip- mates (see Kasza and Wolfe2 for a description of
tion of multivariable linear regression). We use data (a regression parameter estimates) derived from each
random sample) from the fifth decade of follow-up of the completed datasets. The standard error of the
(commencing in 2004) from the Tasmanian Longitu- MI estimate incorporates both the uncertainty in
dinal Health Study (TAHS). TAHS is a population- the estimate within the completed datasets and the
based longitudinal cohort study of 8683 children born uncertainty across the completed datasets due to the
in 1961 and attending school in Tasmania in 1968.10 In missingness (see Appendix for details). This ensures
this study, waist circumference was not available for that MI produces a valid 95% confidence interval and
approximately one quarter of the participants, P-value for the MI estimate of the regression param-
meaning that when we adjust for this covariate using eter of interest.11
the standard complete case analysis, those partici- In the simplest case, when there are missing data in
pants are excluded. We present MI as an alternative a single (continuous) variable, the imputation model
approach for dealing with these missing data. consists of a linear regression model for the variable
For illustrative purposes, we restrict our analysis to that has missing values (e.g. waist circumference in
316 TAHS participants who have complete data on all our example), regressed on the other variables to be
variables in the analysis aside from waist circumfer- used for imputation (e.g. the other variables to be
ence. We note that the estimated associations are for used in the analysis plus the auxiliary variables).
illustrative purposes only, and should not be inter- When there are missing data in a number of variables,
preted as a definitive analysis of this study. there are currently two approaches for imputing the
© 2013 The Authors Respirology (2014) 19, 162–167
Respirology © 2013 Asian Pacific Society of Respirology
14401843, 2014, 2, Downloaded from [Link] by National Health And Medical Research Council, Wiley Online Library on [13/11/2025]. See the Terms and Conditions ([Link] on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
164 KJ Lee and JA Simpson

missing values. One approach is to impute the simply because they have missing data in that one (or
missing values using a series of conditional regression more) variable(s), hence can be inefficient (i.e. can
models. Under this approach, a regression model is mean wide confidence intervals around parameter
set up for each variable with missing data as estimates). MI can improve efficiency (i.e. results in
described above, cycling through the regression narrower confidence intervals) as it allows all partici-
models sequentially to impute missing values for pants to be included in the analysis.
each variable conditional on the imputed values for The largest potential for efficiency gains from MI
the other variables with missing data.12,13 The second over the complete case analysis is when the vari-
approach is to impute all of the variables with missing able(s) of interest are fully observed, such as in our
data simultaneously using a joint normal distribu- motivating example, where the exposure of interest
tion.14 Both of these MI approaches are available in (asthma) and the outcome (FEV1) are fully observed,
standard computerized statistical packages (e.g. but there are missing values in important confound-
Stata15 and SAS16). ers (waist circumference). In this scenario, excluding
incomplete cases means that we are losing informa-
tion about the exposure–outcome relationship in
WHEN MIGHT MI OFFER BENEFITS cases where the covariate is missing, information that
OVER A COMPLETE CASE ANALYSIS? can be recovered using MI. In contrast, if there are
missing exposure or outcome data, then we are less
MI can offer gains over a complete case analysis in likely to gain information about the association
terms of reducing bias and/or improving precision. between exposure and outcome from MI, unless there
are auxiliary variables that are very highly correlated
with the incomplete variable, for example a similar
Reducing bias outcome measured at a previous wave of data
In some scenarios, there may be differences between collection.17,18
participants with and without complete data, for The gain in efficiency from MI is due to the inclu-
example, those with asthma and/or allergies may be sion of auxiliary variables in the imputation model
more motivated to attend follow-up visits for a res- (see Section 2). The stronger the association between
piratory study. In this context, carrying out a com- the incomplete and auxiliary variables, the more
plete case analysis can lead to biased results as it will accurate the imputed values will be, and the larger the
be based on a non-representative sample from the potential gains in precision from MI. In practice, there
population. needs to be a reasonably strong correlation between
If the reasons for missingness are known, that is, the the incomplete and auxiliary variable for MI to have
missingness depends on observed data (e.g. age in notable gains over a complete case analysis,19 and
our data example), but not on data that are unob- sometimes there are not many (if any variables) with
served, we refer to the data as missing at random. If such strong correlation.20
data are missing at random, MI can use the observed
data to fill in the missing values. Filling in the missing
values enables all participants to be included in the MI IS NOT A MIRACLE CURE
analysis, hence correcting the bias in the complete
case analysis. This bias correction is, however, only Although MI is intuitively appealing, it is not a miracle
possible if there are auxiliary variables that can be cure. Importantly, MI can introduce bias over a com-
included in the model used to impute the missing plete case analysis if not carried out appropriately.17
values (otherwise the imputation and analysis models Specifically, there are a number of decisions which
are analogous). Data may also be missing not missing must be made when setting up the imputation model
at random, that is the missingness may depend on (stage 1), which can affect the validity of the resulting
unobserved data (e.g. unmeasured genetic factors). In inference:
this context, MI can still reduce bias compared with 1 How much missing data are there?—If there is a lot
complete case analysis if there are auxiliary variables of missing data, any bias introduced by the decisions
that are strong predictors of missingness, although made in setting up the imputation model will be
some bias may remain as the imputed values are esti- inflated as a large amount of data is being imputed
mated from the observed data only. from a potentially mis-specified model.21
2 Which variables to include in the imputation
model?—It is important to include all variables that
Improving Precision are in the analysis model in the imputation model,
If the missingness occurs completely at random, for including any interaction terms and terms that
example, if the spirometer was not available to describe a nonlinear association, for example quad-
measure lung function for a random set of appoint- ratic or logarithm transformations.19 Leaving these
ments, then a complete case analysis will be unbiased variables out of the imputation model can lead to
since it includes a random sample of the original biased results. As discussed in the previous section, it
study participants and hence a random sample from is also important to include auxiliary variables that
the population (assuming that the original sample aid information recovery.
was a random sample from the population). However, 3 How to include non-normally distributed continu-
performing a complete case analysis still means ous variables?—Both approaches to MI (described
throwing away information from study participants briefly under “What is MI?”) assume normality for
Respirology (2014) 19, 162–167 © 2013 The Authors
Respirology © 2013 Asian Pacific Society of Respirology
14401843, 2014, 2, Downloaded from [Link] by National Health And Medical Research Council, Wiley Online Library on [13/11/2025]. See the Terms and Conditions ([Link] on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Multiple imputation for missing data 165

continuous variables (at least conditionally on the Step 1: Invesgaon of the missing data
other variables in the imputation model). This means - Which variables?
that including the raw (i.e. original) scale values of any - How much missing data?
non-normally distributed variables (e.g. cytokine - Are there predictors of missingness?
levels pmol/l, which often follow a highly skewed dis-
tribution) in the imputation model can produce
imputed values that are quite different to the Step 2: Set up the imputaon model
observed values. This, in turn, can have flow on effects - Which method of imputaon?
to the inference obtained.22 In contrast, it has been - p
Which variables to include in the imputaon model?
- What form of variables to use?
suggested that transforming data to improve normal-
ity prior to imputation can also lead to biased results
in some scenarios.23
4 How to impute categorical variables when using a Step 3: Carry out mulple imputaon
joint normal distribution?—Since this approach - Generate m sets off imputed
d values
l to produce
d m completed
l dd datasets
- Perform epidemiological analysis on each completed dataset
assumes normality for all variables in the imputation - Combine the results using Rubin’s Rules
model, it is unclear how best to impute missing values
for a categorical variable under this framework.24,25
5 How to impute and analyse variables that have a
restricted range of values?—Some variables are Step 4: Invesgate the sensivity of the results to the decisions made
when seng up the imputaon model
restricted in the range of values that they can take, for - Visualise the distribuon of the imputed and observed values
example FEV1 must be greater than 0 by definition - Carry out diagnosc checks for your imputaon model
and is also usually < 8 L. Again there are various
approaches that have been suggested for the imputa- Figure 2 Process for carrying out multiple imputation.
tion and analysis of restricted range variables.23,26,27
All of the above issues (see Graham19 for a detailed
discussion) should be considered prior to imputation,
and with respect to the dataset under investigation.
The fact that this approach is so flexible means that it is using the Stata command, ‘mi impute regress’. Note
desirable to have some expertise in the methodology that there was no need to decide whether to use
of MI prior to using this approach and making these sequential imputation or a joint normal model in this
decisions. Furthermore, it is important to explore the example as there was only one variable with missing
sensitivity of the results of the epidemiological analy- data, hence imputation was carried out using a
sis of interest to the decisions made in the imputation univariate linear regression model. Since waist cir-
process, which is reassuring if all imputation models cumference was approximately normally distributed,
lead to the same overall conclusion. it was imputed using the original scale (i.e. no trans-
formation was required). All variables in the analysis
model were included in the imputation model, along
with indicators for high blood pressure and high
RESULTS FROM OUR CASE STUDY cholesterol. Other categorical variables were also
included in the imputation model as a series of indi-
Figure 2 shows a flow chart of the steps that should be cators (single indicators for current asthma, female
taken when carrying out MI. and ever smoker, and four indicators for the socio-
In our case study, waist circumference was missing economic status—a five-level ordinal variable).
in 81 of the 316 participants in the analysis dataset Twenty imputed datasets were generated, with the
(26%). Participants with available data on their waist resulting estimates of the model parameters (from
circumference were slightly younger than those separate analyses of each completed dataset) com-
missing waist circumference (mean age 42.5 (stand- bined using the Stata command, ‘mi estimate’.15
ard deviation (SD) 0.5) years compared with 42.7 (SD Table 1 shows the results from our case study. In
0.2) years, note at the fifth decade of follow-up par- the complete case analysis, asthma was associated
ticipants were aged between 41 and 44 years), sug- with a lower FEV1 (regression coefficient −266 mL,
gesting that a complete case analysis may lead to standard error 84 mL). The estimate from MI shows a
biased results. slightly stronger relationship, regression coefficient
In the observed data, mean waist circumference −294 mL, but importantly has decreased uncertainty
was greater for participants who had high blood pres- (standard error = 74 mL), and hence increased preci-
sure (99.3 (SD 14.9) cm) compared with those without sion, compared with the complete case analysis due
high blood pressure (92.3 (SD 13.0) cm), and was to the recovery of the 81 cases with missing waist cir-
greater for participants with high cholesterol (95.7 cumference measurements. This means a narrower
(SD 11.4)) compared with those without high choles- 95% confidence interval and smaller P-value from
terol (93.2 (SD 14.2)). These findings suggest that high the MI analysis. In contrast, the estimate of the asso-
blood pressure and high cholesterol may be useful ciation between waist circumference and FEV1 (and
auxiliary variables for imputing waist circumference. the corresponding standard error) are similar in the
We present the results from our case study using a two analyses, since there is less to be gained from
complete case analysis and MI. Analyses were con- MI (in terms of bias or precision) regarding this
ducted using Stata Release 12.15 MI was carried out relationship.
© 2013 The Authors Respirology (2014) 19, 162–167
Respirology © 2013 Asian Pacific Society of Respirology
14401843, 2014, 2, Downloaded from [Link] by National Health And Medical Research Council, Wiley Online Library on [13/11/2025]. See the Terms and Conditions ([Link] on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
166 KJ Lee and JA Simpson

Table 1 Results from the analysis of the Tasmanian Longitudinal Health Study

Asthma Waist circumference

Estimate (SE) 95% CI P Estimate (SE) 95% CI P

Complete case analysis −266 (84) −431, −102 0.002 −6.0 (2.9) −11.6, −0.3 0.04
Multiple imputation −294 (74) −439, −149 <0.001 −5.8 (2.9) −11.6, 0.0 0.05

Estimates are regression coefficients from a multivariable linear regression model. The outcome variable is forced expiratory volume
in 1 s (mL), and the covariates are asthma (exposure of interest), age (years), gender, socio-economic status, smoking status, height
(cm) and waist circumference (cm; variable with 26% missing data).
CI, confidence interval; SE, standard error.

CONCLUDING REMARKS (ViCBiostat), and project grant 607400. Research at the Murdoch
Childrens Research Institute is supported by the Victorian Gover-
nment’s Operational Infrastructure Support Program.
We have presented MI as a useful approach for
dealing with missing data. In particular, it can reduce
bias and increase precision compared with a com-
plete case analysis when there are additional REFERENCES
observed variables associated with the variable with
missing data. However, we note that this may not 1 Wolfe R, Abramson M. Modern statistical methods in respiratory
always be the case, particularly when the variables of medicine. Respirology 2014; 19: 9–13.
interest (e.g. the outcome or the exposure of interest) 2 Kasza J, Wolfe R. Statistical regression models: interpretation of
have missing data.17,28 commonly-used models. Respirology 2014; 19: 14–21.
Although MI can be useful, we do not recommend 3 Vermeulen KM, Post WJ, Span MM, van der Bij W, Koëter GH,
TenVergert EM. Incomplete quality of life data in lung transplant
it be used as a blanket approach for dealing with all
research: comparing cross sectional, repeated measures ANOVA,
missing data. In particular, MI can introduce bias not and multi-level analysis. Respir. Res. 2005; 6: 101–10.
present in a complete case analysis if not carried out 4 Molenberghs G, Kenward MG. Missing Data in Clinical Studies.
appropriately. Instead, we highlight the importance John Wiley and Sons Ltd, Chichester, 2007.
of MI as a sensitivity analysis surrounding the 5 Van Buuren S. Flexible Imputation of Missing Data. CRC Press,
robustness of the results to the missing data. As dem- Hoboken, 2012.
onstrated in our case study, it is reassuring when the 6 Sterne JAC, White IR, Carlin JB, Royston P, Kenward MG, Wood
complete case analysis and MI, both of which make AM, Carpenter JR. Multiple imputation for missing data in epi-
different assumptions about the missingness, result demiological and clinical research: potential and pitfalls. BMJ
in the same overall conclusion. We also recommend 2009; 338: b2393.
7 Mackinnon A. The use and reporting of multiple imputation in
that the researcher performing MI be suitably
medical research—a review. J. Intern. Med. 2010; 268: 586–93.
familiar with the methodology of this approach 8 von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche P,
before using MI. Vandenbroucke JP, STROBE Initiative. Strengthening the report-
When considering MI, it is important to first assess ing of observational studies in epidemiology (STROBE) state-
whether MI is likely to offer gains, in terms of either ment: guidelines for reporting observational studies in
reducing bias or increasing precision, over a complete epidemiology. BMJ 2007; 335: 806–8.
case analysis. Once it has been decided to use MI, the 9 Ware JH, Harrington D, Hunter DJ, D’Agostino RB. Missing data.
analyst needs to consider carefully the most appropri- N. Engl. J. Med. 2012; 367: 1353–4.
ate imputation model, and the sensitivity of the 10 Wharton C, Dharmage S, Jenkins M, Dite G, Hopper J, Giles G,
Abramson M, Walters EH. Tracing 8,600 participants 26 years
approach to the decisions made during imputation.
after recruitment at age seven for the Tasmanian Asthma Study.
There is an increasing focus on diagnostic checks of Aust. N. Z. J. Public Health 2006; 30: 105–10.
the imputation model to help guide the researcher in 11 Rubin DB. Multiple Imputation for Nonresponse in Surveys.
building an appropriate model.5,29 Finally, MI assumes Wiley, New York, 1987.
that data are missing dependent on observed vari- 12 Raghunathan TE, Lepkowski JM, Van Hoewyk J, Solenberger P. A
ables. In practice, missingness may depend on unob- multivariate technique for multiply imputing missing values
served data. Given it is not possible to assess whether using a sequence of regression models. Surv. Methodol. 2001; 27:
missingness depends on unobserved data, it may be 85–95.
important to assess the sensitivity of the analysis to 13 VanBuuren S, Boshuizen HC, Knook DL. Multiple imputation of
the missing at random assumption when using MI.30,31 missing blood pressure covariates in survival analysis. Stat. Med.
1999; 18: 681–94.
14 Schafer JL. Analysis of Incomplete Multivariate Data. Chapman &
Acknowledgements Hall, London, 1997.
We thank the TAHS Steering Committee for providing us with a 15 StataCorp. Stata Statistical Software: Release 12. StataCorp LP,
random subset of the data from the fifth decade of follow up of College Station, TX, 2011.
the TAHS cohort which was funded by the NHMRC (ID 299901). 16 SAS Institute Inc. PROC MI. SAS Procedures Guide,Version 92. SAS
This work was supported by funding from the National Health Institute Inc, Cary, NC, 2008.
and Medical Research Council: Career Development Fellowship 17 Lee KJ, Carlin JB. Recovery of information from multiple
ID 1053609 (K.J.L.), a Centre of Research Excellence grant, ID imputation: a simulation study. Emerg. Themes Epidemiol. 2012;
1035261, awarded to the Victorian Centre for Biostatistics 9: 3.

Respirology (2014) 19, 162–167 © 2013 The Authors


Respirology © 2013 Asian Pacific Society of Respirology
14401843, 2014, 2, Downloaded from [Link] by National Health And Medical Research Council, Wiley Online Library on [13/11/2025]. See the Terms and Conditions ([Link] on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Multiple imputation for missing data 167
18 Marshall A, Altman D, Royston P, Roger L. Comparison of tech- regression model) from the first completed
niques for handling missing covariate data within prognostic dataset = −298, and then β̂ 2 = −267 , and so on up
modelling studies: a simulation study. BMC Med. Res. Methodol.
2010; 10: 7. to β̂ 20 = −287 , so that:
19 Graham J. Missing Data: Analysis and Design. Springer, New York,
NY, 2012. 1
20 Karahalios A, Baglietto L, Lee KJ, English DR, Carlin JB, Simpson β MI = × (( −298) + ( −267 ) +  + ( −287 )) = −294.
JA. The impact of missing data on analyses of a time-dependent
20
exposure in a longitudinal cohort: a simulation study. Emerg.
Themes Epidemiol. 2013; 10: 6.
21 Rubin DB. Multiple imputation after 18+ years. J. Am. Stat. Assoc.
The standard error of the MI estimate, used to calcu-
1996; 91: 473–89. late the 95% confidence interval (CI) and P-value
22 Lee KJ, Carlin JB. Multiple imputation for missing data: fully for β MI , is derived using the formula:
conditional specification versus multivariate normal imputa-

( m1 )B
tion. Am. J. Epidemiol. 2010; 171: 624–32.
23 von Hippel PT. Should a normal imputation model be modified SE (βMI ) = V + 1 +
to impute skewed variables? Sociol. Methods Res. 2013; 42: 105–
38.
24 Galati JC, Seaton KA, Lee KJ, Simpson JA, Carlin JB. Rounding where V is a measure of the uncertainty of the esti-
non-binary categorical variables following multivariate normal
mate within each of the imputed datasets, calculated
imputation: evaluation of simple methods and implications for
practice. J. Stat. Comput. Simul. 2012; iFirst: 1–14. as the average of the square of the standard errors of
25 Lee KJ, Galati JC, Simpson JA, Carlin JB. Comparison of methods the β̂ i ’s derived from each of the i imputed datasets:
for imputing ordinal data using multivariate normal imputation:
a case study of non-linear effects in a large cohort study. Stat.
1
( ( )) .
2
∑ SE βˆi
m
Med. 2012; 31: 4164–74. V=
26 Enders C. Applied Missing Data Analysis. The Guilford Press, New m i=1
York, 2010.
27 Royston P, Carlin JB, White IR. Multiple imputation of missing
In our example, ( )
SE β̂1 = 72 , ( )
SE β̂1 = 73 , ...
values: new features for ‘mim’. Stata J. 2009; 9: 252–64.
28 White IR, Carlin JB. Bias and efficiency of multiple imputation ( )
SE β̂1 = 73 , so that
compared with complete-case analysis for missing covariate
values. Stat. Med. 2010; 29: 2920–31. 1
29 Abayomi K, Gelman A, Levy M. Diagnostics for multivariate V= (722 + 732 +  + 732 ) = 5361.
imputations. Appl. Stat. 2008; 57: 273–91.
20
30 Carpenter JR, Kenward MG, White IR. Sensitivity analysis after
multiple imputation under missing at random: a weighting And B is a measure of how far the estimates of β from
approach. Stat. Methods Med. Res. 2007; 16: 259–75. each of the imputed datasets are from the overall MI
31 Carpenter J, Kenward MG. Multiple Imputation and Its Applica- estimate ( β MI ), representing the between imputation
tion. WIley, Chichester, 2013.
variability. This is calculated from the sum of the
squared differences between β̂i and β MI :

1
( )
2

m
APPENDIX—RUBIN’S RULES B= βˆi − β MI .
m −1 i =1

The MI estimate of a parameter which we will


denote β, in our example the regression coefficient for In our example, β̂1 = −298 , β̂ 2 = −267 , . . . , β̂ 20 −287,
current asthma status in a multivariable linear regres- and β MI = −294 so that:
sion model for FEV1, is simply the average of the esti-
mated β from the analysis of each of the m completed B=
1
( −298 − ( −294))2 + ( −267 − ( −294))2 +  +
datasets (observed plus imputed data, Fig. 1): 20 − 1
( −287 − ( −294))2 = 85.
1 m ˆ
β MI = ∑ βi
m i =1 Combing these gives the standard error of the MI
estimate:
where β̂i is the estimate of the parameter in the ith

( )
(i = 1, . . . ,m) completed dataset. In our example with
1
20 completed datasets, the estimate of β (the regres- SE (β MI ) = 5361 + 1 + × 85 = 74.
20
sion coefficient obtained from the multivariable

© 2013 The Authors Respirology (2014) 19, 162–167


Respirology © 2013 Asian Pacific Society of Respirology

You might also like