0% found this document useful (0 votes)
3 views6 pages

Guidelines for Reporting Research Analyses

The document provides guidelines for authors on how to report multivariate analyses in scientific articles, focusing on regression analysis and ANOVA. It outlines various types of regression, including simple and multiple linear regression, logistic regression, and the importance of reporting assumptions, outliers, and missing data treatment. Additionally, it emphasizes the need for clarity in presenting statistical results, including coefficients, confidence intervals, and validation of models.

Uploaded by

Harsyah Ahmad
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views6 pages

Guidelines for Reporting Research Analyses

The document provides guidelines for authors on how to report multivariate analyses in scientific articles, focusing on regression analysis and ANOVA. It outlines various types of regression, including simple and multiple linear regression, logistic regression, and the importance of reporting assumptions, outliers, and missing data treatment. Additionally, it emphasizes the need for clarity in presenting statistical results, including coefficients, confidence intervals, and validation of models.

Uploaded by

Harsyah Ahmad
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

See discussions, stats, and author profiles for this publication at: [Link]

net/publication/6508026

Documenting Research in Scientific Articles: Guidelines for Authors *

Article in Chest · March 2007


DOI: 10.1378/chest.06-2088 · Source: PubMed

CITATIONS READS

70 9,707

1 author:

Tom Lang
Tom Lang Communications and Training International
45 PUBLICATIONS 49,644 CITATIONS

SEE PROFILE

All content following this page was uploaded by Tom Lang on 12 July 2021.

The user has requested enhancement of the downloaded file.


balt2/zcb-chest/zcb-chest/zcb-orig/zcb0543-07a mortonk2 S!6 11/8/06 11:49 Art: 06-2088 Input-gt

Medical Writing Tips

Documenting Research in Scientific


Articles: Guidelines for Authors*
3. Reporting Multivariate Analyses
Tom Lang, MA

(CHEST 2007; 131:1–??) • Simple linear regression is used to assess the

M ultivariate analyses include two broad statistical relationship between a single continuous explana-
techniques, regression analysis and analysis of tory variable and a single continuous response
variance (ANOVA). The reporting guidelines for variable that varies linearly over a range of values.
each are similar and here have been condensed from • Multiple linear regression is used to assess the
the book How To Report Statistics in Medicine.1 linear relationship between two or more continu-
ous or categorical explanatory variables and a
single continuous response variable.
Reporting Regression Analysis • Simple logistic regression is used to assess the
relationship between a single continuous or cate-
Regression analysis attempts to predict or estimate gorical explanatory variable and a single categori-
the value of a response variable or outcome from the cal response variable, usually a binary variable,
known values of one or more explanatory variables or such as whether or not a heart attack has occurred.
predictors. The type of regression analysis is deter- • Multiple logistic regression is used to assess the
mined by the number of explanatory (or indepen- relationship between two or more continuous or
dent) variables and of the response (or dependent) categorical explanatory variables and a single cat-
variables, as well as by the “level of measurement” of egorical response variable.
these variables. • Nonlinear regression is used to assess variables
The phrase level of measurement refers to the kind that are not linearly related and that cannot be
of information collected about a variable. Nominal transformed into a linear relationship. These equa-
data are categorical data with no inherent ranking, tions model more complex relationships than the
such as blood type (eg, A, B, AB, and O); ordinal data other forms of regression analysis.
are categorical data that do have an inherent ranking, • Polynomial regression can be used for any of the
such as severity categories (eg, mild, moderate, and above combinations of explanatory and response
severe); and continuous data are measurements variables when the relationship among the vari-
made on a continuous scale of equal intervals. The ables is curvilinear, which requires, say, squaring
level of measurement can also be set by the re- or cubing one or more explanatory variables in the
searcher. For example, data on BP can be collected model.
as a nominal variable (hypertensive or not hyperten- • Cox proportional hazards regression, an aspect of
sive), an ordinal variable (hypotensive, normotensive, time-to-event (survival) analysis, is used to assess
or hypertensive) or a continuous variable (systolic BP the relationship between two or more continuous
measured in millimeters of mercury.) or categorical explanatory variables and a single
The most common types of regression analyses are continuous response variable (the time to the
as follows: event). Typically, the event (usually death) has not
yet occurred for all participants in the sample,
*From Tom Lang Communications and Training, Davis, CA. which creates censored observations.
AQ: E Manuscript received August 21, 2006; revision accepted August
24, 2006.
Reproduction of this article is prohibited without written permission Guideline: Describe the Relationship of
from the American College of Chest Physicians ([Link].
org/misc/[Link]).
Interest or the Purpose of the Analysis
Correspondence to: Tom Lang, MA, 1925 Donner Ave, No. 3,
Davis, CA 95618; e-mail: tomlangcom@[Link] In addition to predicting one value from or more
DOI: 10.1378/[Link] others, regression analysis can be used to “control AQ: A

[Link] CHEST / 131 / 0 / 000, 2007 1


balt2/zcb-chest/zcb-chest/zcb-orig/zcb0543-07a mortonk2 S!6 11/8/06 11:50 Art: 06-2088 Input-gt

for” the potential confounding effects of explanatory most recent observed value for the person (called the
variables that are associated with the response vari- last-observation-carried-forward method, which is
able. Regression analysis can separate the effects of, commonly used in pharmaceutical research). Other
say, age and sex on survival after surgery, for exam- methods of imputing data are possible, but they
ple. should be based on sound judgment.
Regression analysis can also be used to create risk
scores. Here, the variables of the risk score are those
of the regression equation, and the score itself is the Guideline: Report How Any Outlying
value predicted by the regression model. Values Were Treated in the Analysis
Outliers are extreme values that appear to be
Guideline: Identify the Variables Used in anomalies. Outliers cannot be ignored: even a single
the Analysis and Summarize Each With outlier can have a profound effect on the relationship
Descriptive Statistics derived from the regression line.2,3 All outliers must
be reported, but it is permissible to report the results
Continuous variables should be summarized with with and without the outliers to indicate their effect
medians and ranges or interquartile ranges (or on the results.
means and SDs if the data are normally distributed),
and categorical data can be summarized with counts
or percentages. Guideline: Report the Regression Model
A simple linear regression equation can be re-
Guideline: Confirm That the Assumptions ported in the text or in a scatter plot of the data.
of the Analysis Were Met and State How Multiple linear regression models can be reported as
Each Was Checked equations (Fig 1) or in tables (Table 1); logistic F1 T1
regression models are typically reported in tables
A statement that the assumptions were verified because the equations are so complex (Table 2). T2
and by which methods is all that need be included.
There are both formal checks (eg, hypothesis tests)
and informal checks (eg, inspection of graphs of Guideline: Report the actual p Value and
residuals) for these assumptions. Sometimes, data the 95% Confidence Interval for the
that violate the assumptions can be adjusted (eg, with Regression Coefficient(s) of the
data transformations) to meet the assumptions. If Explanatory Variable(s), and in Logistic
such adjustments were made, they should be iden- Regression, Report the Odds Ratio and the
tified. Associated 95% Confidence Interval
In regression analysis, the regression coefficient
Guideline: Report How Any Missing Data for an explanatory variable indicates how much the
Were Treated in the Analyses average value of the response variable, Y, varies with
each unit change in the explanatory variable, X. The
Missing data can be a problem in multivariate coefficient, or !-weight, is an estimate and so should
analysis because it reduces the sample size unless
corrective measures are taken. To create a model for
predicting weight from age and height, for example,
values for each of these variables must be collected
for each patient. If age is missing from one patient,
Figure 1. A multiple linear regression equation. In this example,
the patient is excluded from the analysis, and the the model predicts overall function score, Y, for patients with
sample size is reduced by one. In regression models multiple sclerosis based on: disease severity, X1 (level 1 being
with several variables, losses to missing data can be least severe and level 15 being most severe); ambulatory ability
(measured as the rate of walking in laps per minute), X2; and
common. number of lesions, X3. Here, X1, X2, and X3 are explanatory
However, missing data can be replaced in a pro- variables (sometimes called risk factors); the numbers in front of
cess called imputation. Simple imputation methods the X values are called regression coefficients or !-weights.
Coefficients are interpreted as follows: if X1 and X3 are held
include using the mean of all observed values for all constant (or “controlling for” disease severity and number of
people in place of the missing value; using the mean lesions), then mean functional score increases by about 1.25
observed value for the same person in other time times (1.22, the coefficient for X2) for each additional lap per
minute. The final model had a coefficient of multiple determi-
periods; using the mean of the previous and follow- nation, R2, of 0.58, indicating that the three variables in the
ing values for the person, if they exist; or using the model explain 58% of the variation in the response variable.

2 Medical Writing Tips


balt2/zcb-chest/zcb-chest/zcb-orig/zcb0543-07a mortonk2 S!6 11/8/06 11:50 Art: 06-2088 Input-gt

Table 1—A Table for Reporting a Multiple Linear Regression Model With Three Explanatory Variables*

Variables Coefficient (!) SE 95% CI Wald $2 p Value†

Intercept 40.79 2.55


X1 3.98 2.37 % 0.67 to 8.63 1.68 0.10
X2 1.23 0.29 0.66 to 1.80 4.20 # 0.001
X3 % 2.09 0.28 % 2.64 to % 1.54 % 7.34 # 0.001
*Intercept & a mathematical constant (no clinical interpretation); X1 to X3 & the explanatory variables; Coefficient & the mathematical weightings
of the explanatory variables in the equation (the regression coefficient or ! -weight); SE & estimated precision of the coefficients; 95% CI & 95%
confidence intervals for the coefficients; Wald $2 & the Wald test statistic calculated from the data to be compared with the $2 distribution with
1 degree of freedom.
†Variables X2 and X3 are statistically significant predictors of the response variable.

be accompanied by a confidence interval that indi- ables to include in the model. In simultaneous
cates its precision. regression, all of the explanatory variables are in-
Odds ratios are widely used in logistic regression cluded in the model and are tested as a group. In
analysis. For a binary explanatory variable, the odds hierarchical regression, the investigator defines the
ratio is the ratio of the odds that an event will occur number and order in which the explanatory variables
in one group to the odds that the event will occur in are entered into the model. Common procedures are
the other group. An odds ratio of 1 means that both forward, backward, stepwise, and best-subset tech-
groups have a similar likelihood of having a heart niques.
attack. The larger the odds ratio, the more likely the
event is expected to occur in the group used in the Guideline: In Multiple Regression Models,
numerator. Specify Whether All Potential
Explanatory Variables Were Assessed for
Collinearity (Nonindependence)
Guideline: Specify How the Explanatory
Variables That Appear in the Final The explanatory variables in a multiple linear
Regression Model Were Chosen regression equation should be independent of one
another.4 If two or more explanatory variables are AQ: C
One of the first steps in building a multiple correlated, that is, if their regression lines are parallel
regression model is to identify the explanatory vari- or “collinear,” then they are not independent. Col-
ables that are significantly related to the response linear variables add much the same information to
AQ: B variable.4 Several dozens of variables may be consid- the model, so only one is needed. The variable with
ered one at a time in this process, called univariate the strongest relationship with the response variable
analysis. Often, a less-restrictive "-level, such as 0.1, should be considered for inclusion in the final model.
is used in the univariate analysis to identify a broad
range of explanatory variables that might be associ- Guideline: In Multiple Regression Models,
ated with the response variable. That is, variables Specify Whether the Explanatory
with p values # 0.1 on univariate analysis are con- Variables Were Tested for Interaction
sidered for inclusion in the model.
The second step in building a regression model is Two explanatory variables are said to interact if the
to identify the best combination of explanatory vari- effect of one explanatory variable on the response

Table 2—A Table for Reporting a Multiple Logistic Regression Model With Four Explanatory Variables*
2
Variable Coefficient (! ) SE Wald $ P Odds Ratio 95% CI

Intercept % 1.88 0.48


X1 1.435 0.589 5.93 0.02 4.2 1.32–13.33
X2 % 0.847 0.690 1.51 0.22 0.43 0.11–1.66
X3 3.045 1.260 5.84 0.02 21.01 1.78–248.29
X4 2.200 0.990 4.94 0.03 9.03 1.30–62.83
*Odds Ratio & controlling for other variables in the model, for every unit increase in, say, variable 1, the odds of having the event of interest
increase by 4.2 (likewise, controlling for other variables in the model, for every unit increase in, say variable 2, the odds of having the event
decrease by 0.43); 95% CI & the 95% confidence interval for the estimated odds ratio. See Table 1 for other abbreviations or explanations not
used in the text.

[Link] CHEST / 131 / 0 / 000, 2007 3


balt2/zcb-chest/zcb-chest/zcb-orig/zcb0543-07a mortonk2 S!6 11/8/06 11:50 Art: 06-2088 Input-gt

variable depends on the level of the second explan- Formal goodness-of-fit tests calculate a p value. If
atory variable. Interaction implies that the variables the p value is statistically significant, the model does
should be considered together, not separately. So, not appropriately fit the data.
for example, if alcohol interacts with antibiotics in
the blood, the model should have a variable for blood
alcohol level, one for blood antibiotic level, and an Guideline: Specify Whether the Model
interaction term that expresses the relationship be- Was Validated
tween serum alcohol and antibiotic level.
Regression models can be validated or tested
against a similar set of data to show that they explain
what they seek to explain. One method used when
Guideline: Provide a Measure of the the sample is large is to develop the model on, say,
“Goodness of Fit” of the Model to the 75% of the data, then to create another model on the
Data remaining 25% of the data, and determine whether
The predictive value of a regression model is the models are similar. Another method involves
AQ: D affected by how well it “fits” the data.5,6 Thus, a removing the data from one subject at a time and
measure of goodness of fit is useful because it reveals recalculating the model. The coefficients and the
how well the model reflects the data on which it was predictive validity of all the models can then be
created. assessed. Such methods are called jack-knife proce-
Simple linear regression analysis can be thought of dures. A third method involves developing another
as an extension of correlation analysis, except that model on a separate set of similar data and deter-
now one variable is being used to predict the other mining whether the models differ.
with the addition of a regression line. As in correla-
tion analysis, scatter plots can be useful for showing
this relationship. The correlation coefficient itself Guideline: Name the Statistical Package
can indicate indirectly how well the model can or Program Used in the Analysis
predict. Correlations have to be high, say, ' 0.7, as
well as statistically significant, if a simple linear Although commercial statistical programs gener-
regression model is to predict with any degree of ally are validated and updated, and have met the test
accuracy. of time, the performance characteristics of privately
In simple linear regression analysis, the correlation developed programs are often unknown.
coefficient associated with the scatter plot is also
useful in the form of the coefficient of determination
(r2). This coefficient indicates how much of the Reporting ANOVA
variability in the response variable is explained by the
ANOVA is a form of hypothesis testing for studies
explanatory variable. For example, if the correlation
involving two or more variables. It is closely related
between skin-fold thickness and body fat is 0.8, then
to regression analysis and should be reported accord-
r2 & 0.64, or 64%. That is, 64% of the variability in
ing to the same general guidelines. Usually, ANOVA
body fat can be accounted for by skin-fold thickness.
is used to assess categorical explanatory variables,
In multiple linear regression analysis, the coefficient
whereas regression analysis is used to assess contin-
of multiple determination (R2) has the same func-
uous explanatory variables. When a study includes
tion.
both continuous and categorical explanatory vari-
A residual is the difference between the value
ables, the analysis may be called multiple regression
predicted by the model and the actual value of the
or analysis of covariance.
data point as collected. The smaller the residual, the
ANOVA is a “group comparison” that determines
better the prediction. Residuals can also be graphed
whether a statistically significant difference exists
to determine how well the assumption of linearity
somewhere among the groups studied. If a signifi-
was met. Thus, a graph of residuals (one kind of
cant difference is indicated, ANOVA is usually fol-
“model diagnostic plot”) in which the values are
lowed by a multiple comparison procedure that
small for all values of X, meaning that they stay close
compares combinations of groups to examine further
to an average difference of zero, indicates that the
any differences among them.
assumption of linearity was met and that the model
The most common ANOVA procedures used in
predicts reasonably well. Outlier assessments work
biomedical research are as follows:
the same way as residual assessments, in that they
and their associated residuals are apparent on the • One-way ANOVA assesses the effect of a single
graph as data points to investigate. (hence the “one-way” designation) categorical ex-

4 Medical Writing Tips


balt2/zcb-chest/zcb-chest/zcb-orig/zcb0543-07a mortonk2 S!6 11/8/06 11:50 Art: 06-2088 Input-gt

Table 3—A Table for Presenting the Results of a Two-Way ANOVA for Analyzing the Two Factors Group and Age*

Source of Sums of Mean


Variation df† Squares Square F Statistic p Value

Group 1 0.64 0.64 2.24 0.16


Age 3 3.92 1.31 4.57 0.02
Group ( age 3 4.91 1.64 5.72 0.01
Error 12 3.43 0.29 ... ...
*ANOVA & includes the two factors: group (two levels or categories) and age (four categories or levels), and the levels of each category should
be stated in the description of the study (group and age significantly interact and so must be considered together); Source of
variation & identification of the sources of variability in the response variable as the factors in the model (group, age, and the interaction between
group and age) and as random error (the variability not explained by the factors); df & the degrees of freedom, a mathematical concept; Sums
of squares & unlike one-way ANOVA, the sums of squares in multiway ANOVA are not easily explained and are best regarded as simply steps
in the calculation of the mean squares; Mean square & the sums of squares divided by the degrees of freedom (essentially, estimates of the
variation in the data); F statistic & the test statistic for the F distribution, for testing for interaction effects and main effects, equals the mean
square for each factor divided by the mean square of the error; p Value & the probability values indicating the statistical significance of the effect
of each factor on the response variable (eg, age and group interact )p & 0.01* in affecting the response variable and should be further investigated
together; ie, the main effect of group or the main effect of age should not be investigated alone).
†For two groups, the df is 2 % 1, or 1. For four age categories, the df is 4 % 1, or 3. For the interaction effect between group and age (ie,
group ( age), the df values for each factor are multiplied (3 ( 1 & 3).

planatory variable (sometimes called a factor) on a can also be expanded to include additional explana-
single continuous response variable. Note, too, tory variables and can assess their simultaneous
that the factor (category) has three or more alter- effects on the response variable. Whereas the pur-
natives (or “levels” or “values”; eg, blood type is A, pose of regression analyses is usually to predict the
B, AB, or O). When there are only two alternatives value of the response variable, the purpose of
(two groups), this analysis reduces to Student t ANOVA is usually to compare groups for differences
test. in the means of the response variable. ANOVA
• Two-way ANOVA assesses the effect of two cate- models are also usually reported in tables (Table 3). T3
gorical explanatory variables (again, sometimes
ACKNOWLEDGMENT: This article draws heavily from How
called factors) on a single continuous response To Report Statistics in Medicine, by Tom Lang.1
variable.
• Multiway ANOVA assesses the effect of three or
more categorical explanatory variables (still called References
factors) on a single continuous response variable. 1 Lang T, Secic M. How to report statistics in medicine. 2nd ed.
• Analysis of covariance assesses the effect of one or Philadelphia, PA: American College of Physicians, 2006
more categorical explanatory variables while con- 2 Godfrey K. Simple linear regression in medical research. In:
trolling for the effects of some other (possibly Bailar JC, Mosteller F, eds. Medical uses of statistics. 2nd ed.
Boston, MA: NEJM Books, 1992; 201–232
continuous) explanatory variables (now called co-
3 Altman DG, Gore SM, Gardner MJ, et al. Statistical guide-
variates) on a single continuous response variable. lines for contributors to medical journals. BMJ 1983; 286:
• Repeated-measures ANOVA is used to assess sev- 1489 –1493
eral, or repeated, measurements of the same 4 Shutty M. Guidelines for presenting multivariate statistical
participants under different conditions (such as analyses in rehabilitation psychology. Rehabil Psych 1994;
39:141–144
BP measurements taken while the patient is su-
5 Bagley SC, White H, Golomb BA. Logistic regression in the
pine, sitting, or standing) or at different points medical literature: standards for use and reporting, with
over time (such as muscle strength measured 1, 5, particular attention to one medical domain. J Clin Epidemiol
10, and 20 days after surgery). 2001; 54:979 –985
6 Hosmer DW, Taber S, Lemeshow S. The importance of
ANOVA is typically used to compare three or assessing the fit of logistic regression models: a case study.
more group means on a certain response variable. It Am J Public Health 1991; 81:1630 –1635

[Link] CHEST / 131 / 0 / 000, 2007 5

View publication stats

Common questions

Powered by AI

Cox proportional hazards regression accommodates censored observations, which occur when the event of interest (usually death) has not yet occurred for all sample participants, by using survival analysis techniques that model the time until the event occurs .

'Goodness of fit' measures indicate how well a regression model reflects the underlying data by comparing observed and predicted values. A good fit suggests that the model accurately captures the relationships among variables, thereby enhancing the model's predictive power .

Testing for interaction effects is necessary to account for situations where the effect of one explanatory variable on the response variable depends on another variable. Interaction implies that variables should be considered together because their combined effect differs significantly from their individual effects .

Assessing potential explanatory variables for collinearity is crucial because collinear variables are not independent and add similar information to the model, potentially distorting results. Typically, only one of the variables should be included, preferably the one with the strongest relationship with the response variable .

Reporting multiple logistic regression models involves specifying the choice and testing of explanatory variables for significance, independence, and interaction. Additionally, presenting coefficients, standard errors, odds ratios with confidence intervals, and the model's fit helps ensure clarity and replicability .

ANOVA can include both categorical and continuous variables through analysis of covariance (ANCOVA), which assesses the impact of categorical variables while controlling for continuous covariates. This extension allows the study of mean differences while accounting for variability caused by continuous factors .

Logistic regression models use odds ratios to convey the likelihood of an event occurring in one group compared to another. An odds ratio of 1 indicates equal likelihood between groups, while values greater or less than 1 suggest increased or decreased likelihood, respectively .

Univariate analysis in multiple regression modeling identifies variables significantly related to the response variable, using a less-restrictive alpha level to cast a wide net for possible explanatory variables. Hierarchical regression further refines this by determining the order and number of explanatory variables for inclusion based on their statistical significance and theoretical relevance .

ANOVA is primarily used for comparing one or more categorical explanatory variables and their effects on a continuous response variable to see if there are significant differences in means among groups. In contrast, regression analysis assesses the relationship between continuous explanatory and response variables, often to predict outcomes and control for confounding effects .

Polynomial regression primarily deals with situations where the relationship among variables is curvilinear, requiring transformations such as squaring or cubing one or more explanatory variables. This allows it to model more complex, non-linear relationships that cannot be transformed into a linear relationship .

You might also like