56-59_YMAM538_Thomp_CP1 2/24/09 5:34 PM Page 56
Basics of Research Part 13 Cheryl Bagley Thompson, PhD, RN
Descriptive Data Analysis
This 13th article of the Basics of Research series is first in a short series on statistical analysis. These articles
will discuss creating your statistical analysis plan, levels of measurement, descriptive statistics, probability theo-
ry, inferential statistics, and general considerations for interpretation of the results of a statistical analysis.
Statistical Analysis Plan If the investigators have no idea of what statistics to use,
The most important part of any research project is the the first question should be what statistics are appropriate
planning process. This statement is as true for data analysis for the research questions being addressed.
as for any of the other steps in the research process. The Investigators should take advantage of the meeting with
development of your statistical analysis plan should not be the statistician to find out why the analysis is appropriate
delayed until after you have your data in hand. Rather, the and to increase their knowledge of statistics. They need to
investigator should select the statistics to describe the sam- be able to defend their choice of statistics at presentations
ple and to analyze the data for each research question or and within publications.
hypothesis before initiating the study. Most grant applica- Several advantages result from having a plan for data
tions will require this information, but these decisions analysis before starting the study. The most obvious is that
should be made regardless of application for funding. the investigators are not left wondering what to do with all
Investigators should plan to first describe their sample. of the data they now have in their computer. A plan speeds
They should identify the important demographic character- the process of data analysis. If a computer program will be
istics of the sample, such as sex, age, and race. These vari- used, the commands for the analysis can even be written
ables will be the same for most studies. Other sample before data collection is complete. In this case, as soon as all
characteristics, such as diagnosis, weight, height, Glasgow of the data are entered, the investigators run the predeter-
Coma Scale, and so forth, also may be important to pro- mined programs, and the analysis is ready for interpretation.
vide. Descriptive statistics will describe these variables. The second advantage of planning the statistical analysis
Next, investigators should plan the analyses for each before the study is an increase in scientific integrity. The
research question/hypothesis. A table may be useful for this investigators who have a plan ahead of time are less likely
activity. In the first column should be the research ques- to bend the analysis to suit their purpose. A plan also pre-
tion/hypothesis; in the second, all relevant variables (and vents the process of repeating analyses until something is
timing information if needed); and in the third, the statisti- found that is statistically significant. A post hoc (after the
cal test to be used. This process helps investigators ensure fact) approach to statistical analysis is inappropriate and
that they are collecting all needed data, at the right time, to increases the chance of making a type I statistical error (see
answer their question. After all study data have been col- a future issue of this series on hypothesis testing for a dis-
lected is not the time to discover that an important piece of cussion of type I errors).2 If post hoc analyses are used, a
data has been missed. technique such as Bonferroni adjustment is needed to
Investigators uncomfortable with statistical analysis should decrease the chance of a type 1 error.2
consult a statistician early in the planning phase. A statistician Before beginning a study, the investigators should iden-
will help them determine what statistical analyses are most tify the computer, the data entry method,3 and the data
appropriate for answering the research questions/hypotheses, analysis software they will use for the study. They also
taking into consideration the types of data to be collected. should spend time during the early phases of the project
Although statisticians may seem intimidating, investigators becoming familiar with the software to be used. Data analy-
should consider them an important member of the research sis will proceed more smoothly if the investigators do not
team and avail themselves of their expertise. need to stop and ask for technical assistance.
To help diminish the stress of a statistical consultation,
investigators should prepare a list of questions before the Statistical Analysis
meeting. In creating the list of questions, they need to start
Overview
with the research question.1 If the investigators have an
idea of what statistics to use, then the questions for the sta- Statistical analysis can be a complex process. However, the
tistician are related to whether the proposed analyses are statistics required for studies most commonly done in critical
appropriate and what other statistics should be considered. care transport research are fairly straightforward. Statistics are
56 Air Medical Journal 28:2
56-59_YMAM538_Thomp_CP1 2/24/09 5:34 PM Page 57
generally descriptive (describing what is) or inferential (deter- the absence of all energy serves as the reference point.
mining the likelihood of a real difference being present in the Unlike with Fahrenheit or Celsius, at 20 degrees Kelvin,
population). To select the most appropriate statistics, investi- molecules are moving twice as fast as at 10 degrees Kelvin.
gators need to know which type of question they are asking Data elements that have the characteristic of a true zero
and the level of measurement being used for the variables. are measured at the ratio (continuous) level of measure-
This article presents information on levels of measurement ment. Common examples are height, weight, and heart
and descriptive statistics, leaving probability theory and infer- rate. Ratio level data are measured at the highest level of
ential statistics for a future issue. measurement and contain the greatest amount of informa-
tion. Ratio level data can be transformed by addition, sub-
Level of Measurement traction, multiplication, or division without altering their
Level of measurement refers to the amount of informa- relative values. Ratio level data can be analyzed with the
tion contained within the data element and to some extent widest range of statistical methods. Ratio level data are
the degree of detail present. Data elements are measured at often required for use with the most powerful statistics.2
the nominal (categorical), ordinal, interval, or ratio (con- Data elements can always be reduced in their level of mea-
tinuous) level. The term nominal level (or categorical) data surement but can never be increased. For example, if income
refers to data that can only be put into groups. For exam- is measured at the ratio level (in exact dollar amounts), the
ple, the demographic data elements of race and religion are data elements can be reduced to ordinal level data by the cre-
measured at the nominal level. This means that the values ation of categories (Table 1). However, data measured at the
consist of categories such as white, African American, ordinal level of measurement cannot be increased to the ratio
Native American, Asian, and other. With nominal level level of measurement. If we ask a subject what category her
data, no category is better than another, and the difference income fits into, we can never determine from the raw data
between categories cannot be determined. For example, her exact income level. Consequently, data should always be
neither Catholic nor Protestant is better than the other, and collected at the highest level of measurement possible and
whether Catholic is closer to Protestant or closer to Jewish converted at the time of data analysis if a lower level of mea-
is not known. The only interpretation possible is that two surement is desired. This recommendation does not hold
subjects are or are not the same on this variable. when there is strong reason to believe that a subject will not
A specific subset of nominal level data is dichotomous be truthful if she is required to provide exact data or when
data. Dichotomous data are nominal but have only two pos- the data cannot be expected to be accurately measured at the
sible categories. A common example is mortality. The values higher level of measurement.
for mortality are either live or die. Other dichotomous vari-
ables are things that can be measured as yes or no, on or off, Descriptive Statistics
or present or absent. Dichotomous variables possess charac- Descriptive statistics are numbers that summarize the
teristics beyond those of other nominal level data, but such data with the purpose of describing what occurred in the
a discussion is beyond the scope of this article. sample. In contrast, inferential statistics are numbers that
Ordinal level data are one step up from nominal data. As allow the investigator to determine whether there are dif-
the name implies, ordinal data have an inherent order. Data ferences between two or more samples and whether these
values such as never, sometimes, often, and always have differences are likely to be present in the population of
order. An individual would interpret sometimes as being interest. Descriptive statistics also can be used to compare
more frequently than never and always as more frequently samples from one study with another. Descriptive statistics
than often. However, the difference in magnitude is not also help researchers detect sample characteristics that may
known with ordinal level data. It cannot be said of ordinal influence their conclusions. For example, if a sample of air
level data that always is twice as frequent as often or that the medical personnel included 400 women and only 20 men,
distance between never and sometimes is the same as the dis- the investigator would need to be careful about generaliz-
tance between sometimes and often. The only interpretation ing the findings to male air transport personnel.
available is that of which is greater or which is lesser. Frequency distributions are often the first analyses to be
At the interval level of measurement, distances between done on a data set. Frequency distributions are a valuable
data elements can be determined. Temperature is the most
common variable measured at the interval level. With tem-
perature, the difference between 40 and 50 degrees is the Table 1. Yearly Income
same as the difference between 50 and 60 degrees. Ratio Level Data Ordinal Level Data
However, interval level data have no true zero. 4,590 < 10,000
Consequently, multiplication is not allowed with interval 11,230 10,000-20,000
level data. This means that you cannot say that 10 degrees 25,600 20,001-30,000
is twice as hot as 5 degrees. Because there is no true zero, a 33,775 30,001-40,000
reference point does not exist. Zero degrees Fahrenheit or
Data from the ratio level column can be reclassified as a value from the
Celsius are arbitrary numbers that do not relate to the
ordinal level column but not vice versa.
amount of temperature present. In contrast, temperature
measured in Kelvin has a true zero4 (absolute zero), where
March-April 2009 57
56-59_YMAM538_Thomp_CP1 2/24/09 5:34 PM Page 58
sents. For example, a frequency distribution for gender would
Table 2. Frequency Distribution describe how many men and how many women were in the
sample. A frequency distribution can be shown using
Gender Number (%)
numeric values or using graphical techniques (Table 2 and
Male 154 (43.4)
Figure 1). Frequency distributions are often univariate (one
Female 201 (56.6)
variable only) but may be bivariate (including two variables).
A bivariate frequency distribution is often presented as a table
with the name and values of one variable across the top and
Figure 1. Frequency Distribution (Histogram) the name and values of the second variable down the left side.
Table 3 is an example of a bivariate frequency distribution.
Multivariate frequency distributions describing more than
two variables at one time are possible but become more com-
plex and are beyond the scope of this article.
Measures of central tendency are statistics that describe
where the middle of the sample lies. The lowest level measure
of central tendency is the mode. The mode is the value most
frequently occurring within the dataset. A mode can be used
with all levels of measurement and is the primary measure of
central tendency available for nominal level data. When
examining the number of patients with trauma, cardiac, or
other medical problems, the category having the most
patients represents the mode. If evaluating the ages of a sam-
ple, the age represented by the most subjects is the mode.
Although the mode provides helpful information for nomi-
nal or ordinal level data with only a few categories, the mode
Table 3. Bivariate Frequency Distribution may be of little value with interval or ratio level data. In
Professional Role Table 4, age 18 is the most common age; however, the sam-
Gender Paramedic Nurse Physician ple is generally much older than that. Overall the sample
Male 74 62 12 consists mostly of individuals in their 40s, although not
Female 53 142 5 many subjects have the exact same age in that small range.
The next measure of central tendency is the median, the
value that is in the exact middle of the sample. The median
is the point at which half of the subjects lie above this value
and half of the subjects lie below it. For example, the ages
Table 4. Mode for Interval or Ratio Level Data of the nine subjects in sample 1 of Table 5 are arranged in
Age Frequency numeric order. The fifth age, 45, is in the exact middle of
18 5 the sample; consequently, age 45 is the median. The
23 1 median is a better measure of central tendency than mode
31 1 because it is not influenced by an accidental grouping of
40 3 values away from the true center of the data. However, the
42 4 median cannot be determined for nominal level data
43 2 because no order is present within the data.
45 1 The mean (or average) is the most common measure of
46 2 central tendency. The mean is calculated by adding up the
54 1 value for all subjects and dividing by the total number of
65 1 subjects (n). For sample 1 in Table 5, the mean is 38.3.
Mode = 18 The mean is more sensitive to outliers and more influenced
by the distribution of the values than is the median. In Table
5 two data sets are demonstrated. Both have the same median
but have very different means. Consequently, different pieces
method for describing nominal or ordinal level data (dis- of information are available with the two measures, and one
creet data). Because discreet data only characterize the or both may be relevant to the research at hand.
quantity within categories, a frequency distribution ade- The use of means to describe a dataset should be limited to
quately describes nominal or ordinal level data. Frequency interval and ratio level data. Nominal level data do not have a
distributions also can help detect data entry errors. true numeric value, so it is not possible to compute a mean.
A frequency distribution consists of a description of the Although ordinal data might be represented using numeric
number of subjects selecting each possible option and may values, the conceptual intervals between the values may not
include the percentage of the sample that this number repre- be the same; therefore, the mean would be difficult to inter-
58 Air Medical Journal 28:2
56-59_YMAM538_Thomp_CP1 2/24/09 5:34 PM Page 59
being bell shaped, with a middle that is exactly in the center of
Table 5. Median and Mean the distribution. In addition, the tails (sides) of the distribu-
tion are symmetric, having the exact same shape (Figure 2).
Sample 1 Sample 2
The presence of a normal distribution of data in the popula-
18 34
tion is a common assumption for inferential statistics.
20 36
One measure of variability is the range. Range is the dif-
21 38
ference between the greatest value and the smallest value.
36 42
For the example in Table 6, the range in sample 1 is 57 and
45 45
in sample 2 is 29. The range informs the reader that one set
46 55
of data is more spread out (more variance) than the other.
52 60
The range, like the mean of a sample, is very sensitive to
53 58
outliers or measurements that are greatly different than the
54 62
rest of the sample. In Table 6, the range is wide, mostly
Median 45 45
because of the influence of only two values, 19 and 76. The
Mean 38.3 47
rest of the values are clustered between 36 and 66, so a
range of 56 could be misleading.
A measure of variability that minimizes the effects of out-
Figure 2. Normal Distribution (Bell Curve) liers is standard deviation. A standard deviation is a mathe-
matical calculation of the variance of all the measurements in
a sample. The standard deviation can be viewed as the average
distance from the mean that each of the values lies. The math-
ematical equations for calculating the standard deviation can
be found in Burns and Grove2 and are easily performed by
basic statistical software programs or standard spreadsheets.
Pearson’s R is another common descriptive statistic. This
statistic describes the relationship between two variables.
Although Pearson’s R can be used descriptively, it is more
commonly used as an inferential statistic. This topic will be
covered in a later issue of the series.
Table 6. Range and Standard Deviation Example
Sample 1 Sample 2
Conclusion
19 33 In conclusion, I would like to stress that the best research
31 42 studies are initiated with a statistical plan already created.
42 48 This plan may or may not have been developed with the
44 51 assistance of a statistician. The first step of data analysis is
55 53 usually to describe the sample and then subgroups within
61 54 the sample. Frequency distribution, mean, median, mode,
62 56 range, and standard deviation are the most commonly used
69 60 statistics for accomplishing this task.
76 62 In the next issue in the series, the basics of probability
Range 57 29 theory will be discussed. This information will be used as a
Mean 51 51 background to the discussion of inferential statistics.
Standard Deviation 17.4 9.0
References
1. Thompson CB, Panacek EA. Clinical research and critical care transport: How to get
started. Air Med J 2006;25:107-11.
2. Burns N, Grove SK. The practice of nursing research: Conduct, critique, and utiliza-
pret. For example, for never (0), occasionally (1), and always tion. 5th ed. St. Louis: Elsevier Saunders; 2005.
(2), the conceptual difference between never and occasion- 3. Thompson CB, Panacek EA. Data management. Air Med J 2008;27:156-8.
4. U.S. Metric Association. Metric system temperature (Kelvin and degree Celsius).
ally may not be the same as the difference between occasion-
Available at [Link]/~hillger/[Link]. Accessed November 30,
ally and always. Thus, the numbers 0, 1, and 2 do not 2008.
accurately represent the conceptual distance between values,
and the mean would be skewed accordingly. Cheryl Bagley Thompson, PhD, RN, is an associate professor and
assistant dean of informatics and learning technologies at the
Measures of central tendency provide information on where University of Nebraska Medical Center College of Nursing in
the majority of data lie. However, these measures do not Omaha. She can be reached at cbthompson@[Link].
inform the reader regarding the distribution of data across
possible values or their variability from one subject to the 1067-991X/$36.00
Copyright 2009 by Air Medical Journal Associates
next. One method for describing a collection of values is called doi:10.1016/[Link].2008.12.001
distribution. A normal distribution is typically described as
March-April 2009 59