Stata Guide to Factor Analysis Techniques
Stata Guide to Factor Analysis Techniques
Keywords
Akaike Information Criterion (AIC) • Anti-image • Bartlett method • Bayes
Information Criterion (BIC) • Communality • Components • Confirmatory factor
analysis • Correlation residuals • Covariance-based structural equation
modeling • Cronbach’s alpha • Eigenvalue • Eigenvectors • Exploratory factor
analysis • Factor analysis • Factor loading • Factor rotation • Factor scores •
Factor weights • Factors • Heywood cases • Internal consistency reliability •
Kaiser criterion • Kaiser–Meyer–Olkin criterion • Latent root criterion •
Measure of sampling adequacy • Oblimin rotation • Orthogonal rotation •
Oblique rotation • Parallel analysis • Partial least squares structural equation
modeling • Path diagram • Principal axis factoring • Principal components •
Principal component analysis • Principal factor analysis • Promax rotation •
Regression method • Reliability analysis • Scree plot • Split-half reliability •
Structural equation modeling • Test-retest reliability • Uniqueness • Varimax
rotation
Learning Objectives
After reading this chapter, you should understand:
8.1 Introduction
Principal component analysis (PCA) and factor analysis (also called principal
factor analysis or principal axis factoring) are two methods for identifying
structure within a set of variables. Many analyses involve large numbers of
variables that are difficult to interpret. Using PCA or factor analysis helps find
interrelationships between variables (usually called items) to find a smaller number
of unifying variables called factors. Consider the example of a soccer club whose
management wants to measure the satisfaction of the fans. The management could,
for instance, measure fan satisfaction by asking how satisfied the fans are with the
(1) assortment of merchandise, (2) quality of merchandise, and (3) prices of
merchandise. It is likely that these three items together measure satisfaction with
the merchandise. Through the application of PCA or factor analysis, we can
determine whether a single factor represents the three satisfaction items well.
Practically, PCA and factor analysis are applied to understand much larger sets of
variables, tens or even hundreds, when just reading the variables’ descriptions does
not determine an obvious or immediate number of factors.
PCA and factor analysis both explain patterns of correlations within a set of
observed variables. That is, they identify sets of highly correlated variables and
infer an underlying factor structure. While PCA and factor analysis are very similar
in the way they arrive at a solution, they differ fundamentally in their assumptions
of the variables’ nature and their treatment in the analysis. Due to these differences,
the methods follow different research objectives, which dictate their areas of
application. While the PCA’s objective is to reproduce a data structure, as well
as possible only using a few factors, factor analysis aims to explain the variables’
correlations by means of factors (e.g., Hair et al. 2013; Matsunaga 2010; Mulaik
2009).1 We will discuss these differences and their implications in this chapter.
Both PCA and factor analysis can be used for exploratory or confirmatory
purposes. What are exploratory and confirmatory factor analyses? Comparing the
left and right panels of Fig. 8.1 shows us the difference. Exploratory factor
analysis, often simply referred to as EFA, does not rely on previous ideas on the
factor structure we may find. That is, there may be relationships (indicated by the
arrows) between each factor and each item. While some of these relationships may
be weak (indicated by the dotted arrows), others are more pronounced, suggesting
that these items represent an underlying factor well. The left panel of Fig. 8.1
illustrates this point. Thus, an exploratory factor analysis reveals the number of
factors and the items belonging to a specific factor. In a confirmatory factor
1
Other methods for carrying out factor analyses include, for example, unweighted least squares,
generalized least squares, or maximum likelihood. However, these are statistically complex and
inexperienced users should not consider them.
8.2 Understanding Principal Component and Factor Analysis 267
Fig. 8.1 Exploratory factor analysis (left) and confirmatory factor analysis (right)
Researchers often face the problem of large questionnaires comprising many items.
For example, in a survey of a major German soccer club, the management was
particularly interested in identifying and evaluating performance features that relate
to soccer fans’ satisfaction (Sarstedt et al. 2014). Examples of relevant features
include the stadium, the team composition and their success, the trainer, and the
268 8 Principal Component and Factor Analysis
as manifestations of the factor capturing the “joint meaning” of the items related to
it. The arrows pointing from the factor to the items in Fig. 8.1 indicate this point. In
our example, the “joint meaning” of the three items could be described as satisfac-
tion with the stadium, since the items represent somewhat different, yet related,
aspects of the stadium. Likewise, there is a second factor that relates to the two
items x4 and x5, which, like the first factor, shares a common meaning, namely
satisfaction with the merchandise.
PCA and factor analysis are two statistical procedures that draw on item
correlations in order to find a small number of factors. Having conducted the
analysis, we can make use of few (uncorrelated) factors instead of many variables,
thus significantly reducing the analysis’s complexity. For example, if we find six
factors, we only need to consider six correlations between the factors and overall
satisfaction, which means that the recommendations will rely on six factors.
Like any multivariate analysis method, PCA and factor analysis are subject to
certain requirements, which need to be met for the analysis to be meaningful. A
crucial requirement is that the variables need to exhibit a certain degree of correla-
tion. In our example in Fig. 8.1, this is probably the case, as we expect increased
correlations between x1, x2, and x3, on the one hand, and between x4 and x5 on the
other. Other items, such as x1 and x4, are probably somewhat correlated, but to a
lesser degree than the group of items x1, x2, and x3 and the pair x4 and x5. Several
methods allow for testing whether the item correlations are sufficiently high.
Both PCA and factor analysis strive to reduce the overall item set to a smaller set
of factors. More precisely, PCA extracts factors such that they account for
variables’ variance, whereas factor analysis attempts to explain the correlations
between the variables. Whichever approach you apply, using only a few factors
instead of many items reduces its precision, because the factors cannot represent all
the information included in the items. Consequently, there is a trade-off between
simplicity and accuracy. In order to make the analysis as simple as possible, we
want to extract only a few factors. At the same time, we do not want to lose too
much information by having too few factors. This trade-off has to be addressed in
any PCA and factor analysis when deciding how many factors to extract from
the data.
Once the number of factors to retain from the data has been identified, we can
proceed with the interpretation of the factor solution. This step requires us to
produce a label for each factor that best characterizes the joint meaning of all the
variables associated with it. This step is often challenging, but there are ways of
facilitating the interpretation of the factor solution. Finally, we have to assess how
well the factors reproduce the data. This is done by examining the solution’s
goodness-of-fit, which completes the standard analysis. However, if we wish to
continue using the results in further analyses, we need to calculate the factor scores.
270 8 Principal Component and Factor Analysis
Factor scores are linear combinations of the items and can be used as variables in
follow-up analyses.
Figure 8.2 illustrates the steps involved in the analysis; we will discuss these in
more detail in the following sections. In doing so, our theoretical descriptions will
focus on the PCA, as this method is easier to grasp. However, most of our
descriptions also apply to factor analysis. Our illustration at the end of the chapter
also follows a PCA approach but uses a Stata command (factor, pcf), which
blends the PCA and factor analysis. This blending has several advantages, which
we will discuss later in this chapter.
Before carrying out a PCA, we have to consider several requirements, which we can
test by answering the following questions:
– the scale points are equidistant, which means that the difference in the wording
between scale steps is the same (see Chap. 3), and
– there are five or more response categories.
– When all communalities (we will discuss this term in Sect. [Link]) are above
0.60, small sample sizes of below 100 are adequate.
– With communalities around 0.50, sample sizes between 100 and 200 are
sufficient.
– When communalities are consistently low, with many or all under 0.50, a sample
size between 100 and 200 is adequate if the number of factors is small and each
of these is measured with six or more indicators.
– When communalities are consistently low and the factors numbers are high or
are measured with only few indicators (i.e., 3 or less), 300 observations are
recommended.
expect lower correlations between, for example, x1 and x4 and between x3 and x5.
Thus, not all of the correlation matrix’s elements need to have high values. The
PCA depends on the relative size of the correlations. Therefore, if single
correlations are very low, this is not necessarily problematic! Only when all the
correlations are around zero is PCA no longer useful. In addition, the statistical
significance of each correlation coefficient helps decide whether it differs signifi-
cantly from zero.
There are additional measures to determine whether the items correlate suffi-
ciently. One is the anti-image. The anti-image describes the portion of an item’s
variance that is independent of another item in the analysis. Obviously, we want all
items to be highly correlated, so that the anti-images of an item set are as small as
possible. Initially, we do not interpret the anti-image values directly, but use a
measure based on the anti-image concept: The Kaiser–Meyer–Olkin (KMO)
statistic. The KMO statistic, also called the measure of sampling adequacy
(MSA), indicates whether the other variables in the dataset can explain the
correlations between variables. Kaiser (1974), who introduced the statistic,
recommends a set of distinctively labeled threshold values for KMO and MSA,
which Table 8.2 presents.
To summarize, the correlation matrix with the associated significance levels
provides a first insight into the correlation structures. However, the final decision of
whether the data are appropriate for PCA should be primarily based on the KMO
statistic. If this measure indicates sufficiently correlated variables, we can continue
the analysis of the results. If not, we should try to identify items that correlate only
weakly with the remaining items and remove them. In Box 8.1, we discuss how to
do this.
(continued)
8.3 Principal Component Analysis 273
Which umbrella term can we use to summarize a set of variables that loads highly on a
specific factor?
What is the common reason for the strong correlations between a set of variables?
From a theoretical perspective, the assumption that there is a unique variance for
which the factors cannot fully account, is generally more realistic, but simulta-
neously more restrictive. Although theoretically sound, this restriction can some-
times lead to complications in the analysis, which have contributed to the
widespread use of PCA, especially in market research practice.
Researchers usually suggest using PCA when data reduction is the primary
concern; that is, when the focus is to extract a minimum number of factors that
account for a maximum proportion of the variables’ total variance. In contrast, if the
primary concern is to identify latent dimensions represented in the variables, factor
analysis should be applied. However, prior research has shown that both approaches
arrive at essentially the same result when:
2
Related discussions have been raised in structural equation modeling, where researchers have
heatedly discussed the strengths and limitations of factor-based and component-based approaches
(e.g., Sarstedt et al. 2016, Hair et al. 2017a, b).
276 8 Principal Component and Factor Analysis
indicating that the factor and the variable correlate perfectly. On the other hand, if
the factor and the variable are uncorrelated, the angle between these two is 90 . This
correlation between a (unit-scaled) factor and a variable is called the factor
loading. Note that factor weights and factor loadings essentially express the same
thing—the relationships between variables and factors—but they are based on
different scales.
After extracting F1, a second factor (F2) is extracted, which maximizes the
remaining variance accounted for. The second factor is fitted at a 90 angle into
the vector space (Fig. 8.4) and is therefore uncorrelated with the first factor.3 If we
extract a third factor, it will explain the maximum amount of variance for which
factors 1 and 2 have hitherto not accounted. This factor will also be fitted at a 90
angle to the first two factors, making it independent from the first two factors
(we don’t illustrate this third factor in Fig. 8.4, as this is a three-dimensional space).
The fact that the factors are uncorrelated is an important feature, as we can use them
to replace many highly correlated variables in follow-up analyses. For example,
using uncorrelated factors as independent variables in a regression analysis helps
solve potential collinearity issues (Chap. 7).
3
Note that this changes when oblique rotation is used. We will discuss factor rotation later in this
chapter.
8.3 Principal Component Analysis 277
2.10, it covers the information of 2.10 variables or, put differently, accounts for
2.10/5.00 ¼ 42% of the overall variance (Fig. 8.5).
Extracting a second factor will allow us to explain another part of the remaining
variance (i.e., 5.00 – 2.10 ¼ 2.90 units, Fig. 8.5). However, the eigenvalue of the
second factor will always be smaller than that of the first factor. Assume that the
second factor has an eigenvalue of 1.30 units. The second factor then accounts for
1.30/5.00 ¼ 26% of the overall variance. Together, these two factors explain
(2.10 + 1.30)/5.00 ¼ 68% of the overall variance. Every additional factor extracted
increases the variance accounted for until we have extracted as many factors as
there are variables. In this case, the factors account for 100% of the overall
variance, which means that the factors reproduce the complete variance.
Following the PCA approach, we assume that factor extraction can reproduce
each variable’s entire variance. In other words, we assume that each variable’s
variance is common; that is, the variance is shared with other variables. This differs
in factor analysis, in which each variable can also have a unique variance.
uniqueness values should be below 0.50. Every additional factor extracted will
increase the explained variance, and if we extract as many factors as there are items
(in our example five), each variable’s communality would be 1.00 and its unique-
ness equal to 0. The factors extracted would then fully explain each variable; that is,
the first factor will explain a certain amount of each variable’s variance, the second
factor another part, and so on.
However, since our overall objective is to reduce the number of variables
through factor extraction, we should extract only a few factors that account for a
high degree of overall variance. This raises the question of how to decide on the
number of factors to extract from the data, which we discuss in the following
section.
Determining the number of factors to extract from the data is a crucial and
challenging step in any PCA. Several approaches offer guidance in this respect,
but most researchers do not pick just one method, but determine the number of
factors resulting from the application of multiple methods. If multiple methods
suggest the same number of factors, this leads to greater confidence in the results.
[Link] Expectations
When, for example, replicating a previous market research study, we might have a
priori information on the number of factors we want to find. For example, if a
previous study suggests that a certain item set comprises five factors, we should
extract the same number of factors, even if statistical criteria, such as the scree plot,
suggest a different number. Similarly, theory might suggest that a certain number of
factors should be extracted from the data.
Strictly speaking, these are confirmatory approaches to factor analysis, which
blur the distinction between these two factor analysis types. Ultimately however,
we should not fully rely on the data, but keep in mind that the research results
should be interpretable and actionable for market research practice.
280 8 Principal Component and Factor Analysis
When using factor analysis, Stata allows for estimating two further criteria
called the Akaike Information Criterion (AIC) and the Bayes Information
Criterion (BIC). These criteria are relative measures of goodness-of-fit and
are used to compare the adequacy of solutions with different numbers of
factors. “Relative” means that these criteria are not scaled on a range of, for
example, 0 to 1, but can generally take any value. Compared to an alternative
solution with a different number of factors, smaller AIC or BIC values
indicate a better fit. Stata computes solutions for different numbers of factors
(up to the maximum number of factors specified before). We therefore need
to choose the appropriate solution by looking for the smallest value in each
criterion. When using these criteria, you should note that AIC is well known
for overestimating the “correct” number of factors, while BIC has a slight
tendency to underestimate this number.
4
Note that factor rotation primarily applies to factor analysis rather than PCA—see Preacher and
MacCallum (2003) for details. However, our illustration draws on the factor, pcf command,
which uses the factor analysis algorithm to compute PCA results for which rotation applies.
8.3 Principal Component Analysis 281
variables in the set. However, the first factor appears to generally correlate more
strongly with the variables, whereas the second factor only correlates weakly with
the variables (to clarify, we look for small angles between the factors and
variables). This implies that we “assign” all variables to the first factor without
taking the second into consideration. This does not appear to be very meaningful, as
we want both factors to represent certain facets of the variable set. Factor rotation
can resolve this problem. By rotating the factor axes, we can create a situation in
which a set of variables loads highly on only one specific factor, whereas another
set loads highly on another. Figure 8.6 illustrates the factor rotation graphically.
On the left side of the figure, we see that both factors are orthogonally rotated
49 , meaning that a 90 angle is maintained between the factors during the rotation
procedure. Consequently, the factors remain uncorrelated, which is in line with the
PCA’s initial objective. By rotating the first factor from F1 to F10 , it is now strongly
related to variables x1, x2, and x3, but weakly related to x4 and x5. Conversely, by
rotating the second factor from F2 to F20 it is now strongly related to x4 and x5, but
weakly related to the remaining variables. The assignment of the variables is now
much clearer, which facilitates the interpretation of the factors significantly.
Various orthogonal rotation methods exist, all of which differ with regard to
their treatment of the loading structure. The varimax rotation (the default option
for orthogonal rotation in Stata) is the best-known one; this procedure aims at
maximizing the dispersion of loadings within factors, which means a few variables
will have high loadings, while the remaining variables’ loadings will be consider-
ably smaller (Kaiser 1958).
Alternatively, we can choose between several oblique rotation techniques. In
oblique rotation, the 90 angle between the factors is not maintained during
rotation, and the resulting factors are therefore correlated. Figure 8.6 (right side)
illustrates an example of an oblique factor rotation. Promax rotation (the default
option for oblique rotation in Stata) is a commonly used oblique rotation technique.
The Promax rotation allows for setting an exponent (referred to as Promax power in
Stata) that needs to be greater than 1. Higher values make the loadings even more
extreme (i.e., high loadings are amplified and weak loadings are reduced even
further), which is at the cost of stronger correlations between the factors and less
total variance explained (Hamilton 2013). The default value of 3 works well for
most applications. Oblimin rotation is a popular alternative oblique rotation type.
282 8 Principal Component and Factor Analysis
5
When the gamma is set to 1, this is a special case, because the value of 1 represents orthogonality.
The result of setting gamma to 1 is effectively a varimax rotation.
8.3 Principal Component Analysis 283
solution (i.e., the goodness-of-fit) (Graffelman 2013). More precisely, to assess the
solution’s goodness-of-fit, we can make use of the differences between the
correlations in the data and those that the factors imply. These differences are
also called correlation residuals and should be as small as possible.
In practice, we check the proportion of correlation residuals with an absolute
value higher than 0.05. Even though there is no strict rule of thumb regarding the
maximum proportion, a proportion of more than 50% should raise concern. How-
ever, high residuals usually go hand in hand with an unsatisfactory KMO measure;
consequently, this problem already surfaces when testing the assumptions.
After the rotation and interpretation of the factors, we can compute the factor
scores, another element of the analysis. Factor scores are linear combinations of the
items and can be used as separate variables in subsequent analyses. For example,
instead of using many highly correlated independent variables in a regression
analysis, we can use few uncorrelated factors to overcome collinearity problems.
The simplest ways to compute factor scores for each observation is to sum all the
scores of the items assigned to a factor. While easy to compute, this approach
neglects the potential differences in each item’s contribution to each factor
(Sarstedt et al. 2016).
Drawing on the eigenvectors that the PCA produces, which include the factor
weights, is a more elaborate way of computing factor scores (Hershberger 2005).
These weights indicate each item’s relative contribution to forming the factor; we
simply multiply the standardized variables’ values with the weights to get the factor
scores. Factor scores computed on the basis of eigenvectors have a zero mean. This
means that if a respondent has a value greater than zero for a certain factor, he/she
scores above the above average in terms of the characteristic that this factor
describes. Conversely, if a factor score is below zero, then this respondent exhibits
the characteristic below average.
284 8 Principal Component and Factor Analysis
Different from the PCA, a factor analysis does not produce determinate factor
scores. In other words, the factor is indeterminate, which means that part of it
remains an arbitrary quantity, capable of taking on an infinite range of values (e.g.,
Grice 2001; Steiger 1979). Thus, we have to rely on other approaches to computing
factor scores such as the regression method, which features prominently among
factor analysis users. This method takes into account (1) the correlation between the
factors and variables (via the item loadings), (2) the correlation between the
variables, and (3) the correlation between the factors if oblique rotation has been
used (DiStefano et al. 2009). The regression method z-standardizes each factor to
zero mean and unit standard deviation.6 We can therefore interpret an observation’s
score in relation to the mean and in terms of the units of standard deviation from this
mean. For example, an observation’s factor score of 0.79 implies that this observa-
tion is 0.79 standard deviations above the average with regard to the corresponding
factor.
Another popular approach is the Bartlett method, which is similar to the
regression method. The method produces factor scores with zero mean and standard
deviations larger than one. Owing to the way they are estimated, the factor scores
that the Bartlett method produces are considered are considered more accurate
(Hershberger 2005). However, in practical applications, both methods yield highly
similar results. Because of the z-standardization of the scores, which facilitates the
comparison of scores across factors, we recommend using the regression method.
In Table 8.3 we summarize the main steps that need to be taken when conducting
a PCA in Stata. Our descriptions draw on Stata’s factor, pcf command, which
carries out a factor analysis but rescales the resulting factors such that the results
conform to a standard PCA. This approach has the advantage that it follows the
fundamentals of PCA, while allowing for analyses that are restricted to factor
analysis (e.g., factor rotation, use of AIC and BIC).
Many researchers and practitioners acknowledge the prominent role that explor-
atory factor analysis plays in exploring data structures. Data can be analyzed
without preconceived ideas of the number of factors or how these relate to the
variables under consideration. Whereas this approach is, as its name implies,
exploratory in nature, the confirmatory factor analysis allows for testing
hypothesized structures underlying a set of variables.
In a confirmatory factor analysis, the researcher needs to first specify the
constructs and their associations with variables, which should be based on previous
measurements or theoretical considerations.
6
Note that this is not the case when using factor analysis if the standard deviations are different
from one (DiStefano et al. 2009).
8.4 Confirmatory Factor Analysis and Reliability Analysis 285
interest, as they indicate whether the construct has been reliably and validly
measured.
Reliability analysis is an important element of a confirmatory factor analysis and
essential when working with measurement scales. The preferred way to evaluate
reliability is by taking two independent measurements (using the same subjects)
and comparing these by means of correlations. This is also called test-retest
reliability (see Chap. 3). However, practicalities often prevent researchers from
surveying their subjects a second time.
An alternative is to estimate the split-half reliability. In the split-half reliability,
scale items are divided into halves and the scores of the halves are correlated to
obtain an estimate of reliability. Since all items should be consistent regarding what
they indicate about the construct, the halves can be considered approximations of
alternative forms of the same scale. Consequently, instead of looking at the scale’s
test-retest reliability, researchers consider the scale’s equivalence, thus showing the
extent to which two measures of the same general trait agree. We call this type of
reliability the internal consistency reliability.
In the example of satisfaction with the stadium, we compute this scale’s split-
half reliability manually by, for example, splitting up the scale into x1 on the one
side and x2 and x3 on the other. We then compute the sum of x2 and x3 (or calculate
the items’ average) to form a total score and correlate this score with x1. A high
correlation indicates that the two subsets of items measure related aspects of the
same underlying construct and, thus, suggests a high degree of internal consistency.
8.5 Structural Equation Modeling 289
However, with many indicators, there are many different ways to split the variables
into two groups.
Cronbach (1951) proposed calculating the average of all possible split-half
coefficients resulting from different ways of splitting the scale items. The
Cronbach’s Alpha coefficient has become by far the most popular measure of
internal consistency. In the example above, this would comprise calculating the
average of the correlations between (1) x1 and x2 + x3, (2) x2 and x1 + x3, as well as
(3) x3 and x1 + x2. The Cronbach’s Alpha coefficient generally varies from 0 to
1, whereas a generally agreed lower limit for the coefficient is 0.70. However, in
exploratory studies, a value of 0.60 is acceptable, while values of 0.80 or higher are
regarded as satisfactory in the more advanced stages of research (Hair et al. 2011).
In Box 8.3, we provide more advice on the use of Cronbach’s Alpha. We illustrate a
reliability analysis using the standard Stata module in the example at the end of this
chapter.
Whereas a confirmatory factor analysis involves testing if and how items relate to
specific constructs, structural equation modeling involves the estimation of
relations between these constructs. It has become one of the most important
methods in social sciences, including marketing research.
There are broadly two approaches to structural equation modeling: Covariance-
based structural equation modeling (e.g., J€oreskog 1971) and partial least
290 8 Principal Component and Factor Analysis
7
Note that we omitted the error terms for clarity’s sake.
8.6 Example 291
constructs, Y1 or Y2, exerts the greater influence on Y3. The result would guide us
when developing marketing plans in order to increase overall satisfaction and,
ultimately, loyalty by answering the research question whether we should rather
concentrate on increasing the fans’ satisfaction with the stadium or with the
merchandise.
The evaluation of a path model analysis requires several steps that include the
assessment of both measurement models and the structural model. Diamantopoulos
and Siguaw (2000) and Hair et al. (2013) provide a thorough description of the
covariance-based structural equation modeling approach and its application. Acock
(2013) provides a detailed explanation of how to conduct covariance-based struc-
tural equation modeling analyses in Stata. Hair et al. (2017a, b, 2018) provide a
step-by-step introduction on how to set up and test path models using partial least
squares structural equation modeling.
8.6 Example
In this example, we take a closer look at some of the items from the Oddjob
Airways dataset ( Web Appendix ! Downloads). This dataset contains eight
items that relate to the customers’ experience when flying with Oddjob Airways.
For each of the following items, the respondents had to rate their degree of
agreement from 1 (“completely disagree”) to 100 (“completely agree”). The vari-
able names are included below:
Our aim is to reduce the complexity of this item set by extracting several factors.
Hence, we use these items to run a PCA using the factor, pcf procedure in
Stata.
With 1,065 independent observations, the sample size requirements are clearly met,
even if the analysis yields very low communality values.
Determining if the variables are sufficiently correlated is easy if we go to ►
Statistics ► Summaries, tables, and tests ► Summary and descriptive statistics ►
Pairwise correlations. In the dialog box shown in Fig. 8.9, either enter each variable
separately (i.e., s1 s2 s3 etc.) or simply write s1-s8 as the variables appear in this
order in the dataset. Then also tick Print number of observations for each entry,
Print significance levels for each entry, as well as Use Bonferroni-adjusted
significance level. The latter option corrects for the many tests we execute at the
same time and is similar to what we discussed in Chap. 6.
Table 8.4 shows the resulting output. The values in the diagonal are all 1.000,
which is logical, as this is the correlation between a variable and itself! The
off-diagonal cells correspond to the pairwise correlations. For example, the pairwise
correlation between s1 and s2 is 0.7392. The value under it denotes the p-value
(0.000), indicating that the correlation is significant. To determine an absolute
minimum standard, check if at least one correlation in all the off-diagonal cells is
significant. The last value of 1037 indicates the sample size for the correlation
between s1 and s2.
The correlation matrix in Table 8.4 indicates that there are several pairs of highly
correlated variables. For example, not only s1 is highly correlated with s2 (correla-
tion ¼ 0.7392), but also s3 is highly correlated with s1 (correlation ¼ 0.6189), just
8.6 Example 293
| s1 s2 s3 s4 s5 s6 s7
-------------+---------------------------------------------------------------
s1 | 1.0000
|
| 1038
|
s2 | 0.7392 1.0000
| 0.0000
| 1037 1040
|
s3 | 0.6189 0.6945 1.0000
| 0.0000 0.0000
| 952 952 954
|
s4 | 0.7171 0.7655 0.6447 1.0000
| 0.0000 0.0000 0.0000
| 1033 1034 951 1035
|
s5 | 0.5111 0.5394 0.5593 0.4901 1.0000
| 0.0000 0.0000 0.0000 0.0000
| 1026 1027 945 1022 1041
|
s6 | 0.4898 0.4984 0.4972 0.4321 0.8212 1.0000
| 0.0000 0.0000 0.0000 0.0000 0.0000
| 1025 1027 943 1022 1032 1041
|
s7 | 0.4530 0.4555 0.4598 0.3728 0.7873 0.8331 1.0000
| 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000
| 1030 1032 947 1027 1038 1037 1048
|
s8 | 0.5326 0.5329 0.5544 0.4822 0.8072 0.8401 0.7773
| 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000 0.0000
| 1018 1019 937 1014 1029 1025 1032
|
| s8
-------------+---------
s8 | 1.0000
|
| 1034
|
values in the table are all above the threshold value of 0.50. For example, s1 has an
MSA value of 0.9166.
The output in Table 8.5 shows three blocks. On the top right of the first block,
Stata shows the number of observations used in the analysis (Number of obs ¼ 921)
and indicates that the analysis yielded two factors (Retained factors ¼ 2). In the
second block, Stata indicates the eigenvalues for each factor. With an eigenvalue of
5.24886, the first factor extracts a large amount of variance, which accounts for
5.24886/8 ¼ 65.61% of the total variance (see column: Proportion). With an
eigenvalue of 1.32834, factor two extracts less variance (16.60%). Using the Kaiser
criterion (i.e., eigenvalue >1), we settle on two factors, because the third factor’s
eigenvalue is clearly lower than 1 (0.397267). The Cumulative column indicates
the cumulative variance extracted. The two factors extract 0.8222 or 82.22% of the
variance, which is highly satisfactory. The next block labeled Factor loadings
(pattern matrix) and unique variances shows the factor loadings along with the
Uniqueness, which indicates the amount of each variable’s variance that the factors
cannot reproduce (i.e., 1-communality) and is therefore lost in the process. All
uniqueness values are very low, indicating that the factors reproduce the variables’
variance well. Specifically, with a value of 0.3086, s3 exhibits the highest unique-
ness value, which suggests a communality of 1–0.3086 ¼ 0.6914 and is clearly
above the 0.50 threshold.
Beside the Kaiser criterion, the scree plot helps determine the number of factors.
To create a scree plot, go to ► Statistics ► Postestimation ► Principal component
analysis reports and graphs ► Scree plot of eigenvalues. Then click on Launch and
296 8 Principal Component and Factor Analysis
(obs=921)
--------------------------------------------------------------------------
Factor | Eigenvalue Difference Proportion Cumulative
-------------+------------------------------------------------------------
Factor1 | 5.24886 3.92053 0.6561 0.6561
Factor2 | 1.32834 0.93107 0.1660 0.8222
Factor3 | 0.39727 0.13406 0.0497 0.8718
Factor4 | 0.26321 0.03202 0.0329 0.9047
Factor5 | 0.23119 0.03484 0.0289 0.9336
Factor6 | 0.19634 0.00360 0.0245 0.9582
Factor7 | 0.19274 0.05067 0.0241 0.9822
Factor8 | 0.14206 . 0.0178 1.0000
--------------------------------------------------------------------------
LR test: independent vs. saturated: chi2(28) = 6428.36 Prob>chi2 = 0.0000
-------------------------------------------------
Variable | Factor1 Factor2 | Uniqueness
-------------+--------------------+--------------
s1 | 0.7855 0.3995 | 0.2235
s2 | 0.8017 0.4420 | 0.1619
s3 | 0.7672 0.3206 | 0.3086
s4 | 0.7505 0.5090 | 0.1777
s5 | 0.8587 -0.3307 | 0.1533
s6 | 0.8444 -0.4203 | 0.1104
s7 | 0.7993 -0.4662 | 0.1439
s8 | 0.8649 -0.3291 | 0.1436
-------------------------------------------------
OK. Stata will produce a graph as shown in Fig. 8.13. There is an “elbow” in the
line at three factors. As the number of factors that the scree plot suggests is one
factor less than the elbow indicates, we conclude that two factors are appropriate.
This finding supports the conclusion based on the Kaiser criterion.
Stata also allows for plotting each eigenvalue’s confidence interval. However,
to display such a scree plot requires running a PCA with a different command
in combination with a postestimation command. The following syntax
produces a scree plot for our example with a 95% confidence interval
(heteroskedastic), as well as a horizontal reference line at the 1 threshold.
estat kmo
-----------------------
Variable | kmo
-------------+---------
s1 | 0.9166
s2 | 0.8839
s3 | 0.9427
s4 | 0.8834
s5 | 0.9308
s6 | 0.8831
s7 | 0.9036
s8 | 0.9180
-------------+---------
Overall | 0.9073
298 8 Principal Component and Factor Analysis
While the Kaiser criterion and the scree plot are helpful for determining the
number of factors to extract, parallel analysis is a more robust criterion. Parallel
analysis can only be accessed through the free add-on package paran. To install the
package, type in help paran in the command window and follow the instructions
to install the package. Having installed the package, type in paran s1 s2 s3 s4
s5 s6 s7 s8, centile(95) q all graph in the command window and Stata
will produce output similar to Table 8.7 and Fig. 8.14.
Table 8.7 contains two rows of eigenvalues, with the first column (Adjusted
Eigenvalue) indicating the sampling error-adjusted eigenvalues obtained by paral-
lel analysis. Note that your results will look slightly different, as parallel analysis
draws on randomly generated datasets. The second column (Unadjusted Eigen-
value) contains the eigenvalues as reported in the PCA output (Table 8.5). Analo-
gous to the original analysis, the two factors exhibit adjusted eigenvalues larger
than 1, indicating a two-factor solution. The scree plot in Fig. 8.14 also supports this
result, as the first two factors exhibit adjusted eigenvalues larger than the randomly
generated eigenvalues. Conversely, the random eigenvalue of the third factor is
clearly larger than the adjusted one.
Finally, we can also request the model selection statistics AIC and BIC for
different numbers of factors (see Fig. 8.15). To do so, go to ► Statistics ►
Postestimation ► AIC and BIC for different numbers of factors. As our analysis
draws on eight variables, we restrict the number of factors to consider to 4 (Specify
the maximum number of factors to include in summary table: 4). Table 8.8
shows the results of our analysis. As can be seen, AIC has the smallest value
(53.69378) for a four-factor solution, whereas BIC’s minimum value occurs for a
8.6 Example 299
Computing: 10% 20% 30% 40% 50% 60% 70% 80% 90% 100%
--------------------------------------------------
Component Adjusted Unadjusted Estimated
or Factor Eigenvalue Eigenvalue Bias
--------------------------------------------------
1 5.1571075 5.2488644 .09175694
2 1.2542191 1.3283371 .07411802
3 .3409911 .39726683 .05627573
4 .23833934 .26320842 .02486908
5 .22882924 .23118517 .00235593
6 .23964682 .19634077 -.04330605
7 .2775334 .19273589 -.0847975
8 .26333341 .14206139 -.12127203
--------------------------------------------------
Criterion: retain adjusted components > 1
----------------------------------------------------------
#factors | loglik df_m df_r AIC BIC
---------+------------------------------------------------
1 | -771.1381 8 20 1558.276 1596.88
2 | -19.72869 15 13 69.45738 141.8393
3 | -10.09471 21 7 62.18943 163.5241
4 | -.8468887 26 2 53.69378 179.1557
----------------------------------------------------------
no Heywood cases encountered
rotate, kaiser
--------------------------------------------------------------------------
Factor | Variance Difference Proportion Cumulative
-------------+------------------------------------------------------------
Factor1 | 3.41063 0.24405 0.4263 0.4263
Factor2 | 3.16657 . 0.3958 0.8222
--------------------------------------------------------------------------
LR test: independent vs. saturated: chi2(28) = 6428.36 Prob>chi2 = 0.0000
-------------------------------------------------
Variable | Factor1 Factor2 | Uniqueness
-------------+--------------------+--------------
s1 | 0.2989 0.8290 | 0.2235
s2 | 0.2817 0.8711 | 0.1619
s3 | 0.3396 0.7590 | 0.3086
s4 | 0.1984 0.8848 | 0.1777
s5 | 0.8522 0.3470 | 0.1533
s6 | 0.9031 0.2719 | 0.1104
s7 | 0.9017 0.2076 | 0.1439
s8 | 0.8557 0.3524 | 0.1436
-------------------------------------------------
--------------------------------
| Factor1 Factor2
-------------+------------------
Factor1 | 0.7288 0.6847
Factor2 | -0.6847 0.7288
--------------------------------
The upper part of Table 8.9 is the same as the standard PCA output (Table 8.5),
showing that the analysis draws on 921 observations and extracts two factors, which
jointly capture 82.22% of the variance. As its name implies, the Rotated factor
loadings block shows the factor loadings after rotation. Recall that rotation is
carried out to facilitate the interpretation of the factor solution. To interpret the
factors, we first “assign” each variable to a certain factor based on its maximum
absolute factor loading. That is, if the highest absolute loading is negative, higher
values of a particular variable relate negatively to the assigned factor. After that, we
should find an umbrella term for each factor that best describes the set of variables
associated with that factor. Looking at Table 8.9, we see that s1–s4 load highly on
the second factor, whereas s5-s8 load on the first factor. For example, s1 has a
0.2989 loading on the first factor, while its loading is much stronger on the second
factor (0.8290). Finally, note that the uniqueness and, hence, the communality
values are unaffected by the rotation (see Table 8.5).
302 8 Principal Component and Factor Analysis
To facilitate the identification of labels, we can plot each item’s loading against
each factor. To request a factor loadings plot, go to ► Statistics ► Postestimation ►
Factor analysis reports and graphs ► Plot of factor loadings and click on Launch.
In the dialog box that follows, retain the default settings and click on OK. The
resulting plot (Fig. 8.16) shows two cluster of variables, which strongly load on one
factor while having low loadings on the other factor. This result supports our
previous conclusion in terms of the variable assignments.
Having identified which variables load highly on which factor in the rotated
solution, we now need to identify labels for each factor. Looking at the variable
labels, we learn that the first set of variables (s1–s4) relate to reliability aspects of
the journey and related processes, such as the booking. We could therefore label
this factor (i.e., factor 2) reliability. The second set of variables (s5–s8) relate to
different aspects of the onboard facilities and the travel experience. Hence, we
could label this factor (i.e., factor 1) onboard experience. The labeling of factors is
of course subjective and you could provide different descriptions.
--------------------------------------------------------------------------------------
Variable | s1 s2 s3 s4 s5 s6 s7 s8
-------------+------------------------------------------------------------------------
s1 | 0.0000
s2 | -0.0525 0.0000
s3 | -0.1089 -0.0602 0.0000
s4 | -0.0598 -0.0557 -0.0907 0.0000
s5 | -0.0174 -0.0063 -0.0029 0.0076 0.0000
s6 | 0.0046 0.0023 -0.0216 0.0159 -0.0464 0.0000
s7 | 0.0136 0.0182 -0.0118 0.0042 -0.0529 -0.0393 0.0000
s8 | -0.0056 -0.0135 -0.0056 0.0052 -0.0404 -0.0265 -0.0645 0.0000
--------------------------------------------------------------------------------------
When examining the lower part of Table 8.10, we see that there are several
residuals with absolute values larger than 0.05. A quick count reveals that 8 out of
29 (i.e., 27.59%) residuals are larger than 0.05. As the percentage of increased
residuals is well below 50%, we can presume a good model fit.
Similarly, our previous analysis showed that the two factors reproduce a suffi-
cient amount of each variable’s variance. The uniqueness values are clearly below
0.50 (i.e., the communalities are larger than 0.50), indicating that the factors
account for more than 50% of the variables’ variance (Table 8.5).
To illustrate its usage, let’s carry out a reliability analysis of the first factor onboard
experience by calculating Cronbach’s Alpha as a function of variables s5 to s8. To
run the reliability analysis, click on ► Statistics ► Multivariate analysis ►
Cronbach’s Alpha. A window similar to Fig. 8.19 will appear. Next, enter variables
s5–s8 into the Variables box.
8.6 Example 305
The Options tab (Fig. 8.20) provides options for dealing with missing values and
requesting descriptive statistics for each item and the entire scale or item
correlations. Check Display item-test and item-rest correlations and click on
OK.
The results in Table 8.11 show that the scale exhibits a high degree of internal
consistency reliability. With a value of 0.9439 (see row Test scale), the Cronbach’s
Alpha coefficient lies well above the commonly suggested threshold of 0.70. This
result is not surprising, since we are simply testing a scale previously established by
means of item correlations. Keep in mind that we usually carry out a reliability
analysis to test a scale using a different sample—this example is only for illustration
purposes! The rightmost column of Table 8.11 indicates what the Cronbach’s Alpha
would be if we deleted the item indicated in that row. When we compare each of the
values with the overall Cronbach’s Alpha value, we can see that any change in the
scale’s set-up would reduce the Cronbach’s Alpha value. For example, by removing
s5 from the scale, the Cronbach’s Alpha of the new scale comprising only s6, s7,
and s8 would be reduced to 0.9284. Therefore, deleting this item (or any others)
makes little sense. In the leftmost column of Table 8.11, Stata indicates the number
of observations (Obs), as well as whether that particular item correlates positively
or negatively with the sum of the other items (Sign). This information is useful for
determining whether reverse-coded items were also identified as such. Reverse-
coded items should have a minus sign. The columns item-test, item-rest, and
average interitem covariance are not needed for a basic interpretation.
306 8 Principal Component and Factor Analysis
average
item-test item-rest interitem
Item | Obs Sign correlation correlation covariance alpha
-------------+-----------------------------------------------------------------
s5 | 1041 + 0.9211 0.8578 421.5723 0.9284
s6 | 1041 + 0.9422 0.8964 412.3225 0.9168
s7 | 1048 + 0.9222 0.8506 399.2565 0.9330
s8 | 1034 + 0.9203 0.8620 434.3355 0.9282
-------------+-----------------------------------------------------------------
Test scale | 416.8996 0.9439
-------------------------------------------------------------------------------
The company’s relationships with its customers are usually long-term oriented
and complex. Since the company’s philosophy is to help customers and business
partners solve their challenges or problems, they often customize their products and
services to meet the buyers’ needs. Therefore, the customer is no longer a passive
buyer, but an active partner. Given this background, the customers’ satisfaction
plays an important role in establishing, developing, and maintaining successful
customer relationships.
Very early on, the company’s management realized the importance of customer
satisfaction and decided to commission a market research project in order to
identify marketing activities that can positively contribute to the business’s overall
success. Based on a thorough literature review, as well as interviews with experts,
the company developed a short survey to explore their customers’ satisfaction with
specific performance features and their overall satisfaction. All the items were
measured on 7-point scales, with higher scores denoting higher levels of satisfac-
tion. A standardized survey was mailed to customers in 12 countries worldwide,
which yielded 281 fully completed questionnaires. The following items (names in
parentheses) were listed in the survey:
Your task is to analyze the dataset to provide the management of Haver and
Boecker with advice for effective customer satisfaction management. The dataset is
labeled haver_and_boecker.dta ( Web Appendix ! Downloads).
For further information on the dataset and the study, see Festge and Schwaiger
(2007), as well as Sarstedt et al. (2009).
1. What is factor analysis? Try to explain what factor analysis is in your own
words.
2. What is the difference between exploratory factor analysis and confirmatory
factor analysis?
3. What is the difference between PCA and factor analysis?
4. Describe the terms communality, eigenvalue, factor loading, and uniqueness.
How do these concepts relate to one another?
5. Describe three approaches used to determine the number of factors.
6. What are the purpose and the characteristic of a varimax rotation? Does a
rotation alter eigenvalues or factor loadings?
References 309
7. Re-run the Oddjob Airways case study by carrying out a factor analysis and
compare the results with the example carried out using PCA. Are there any
significant differences?
8. What is reliability analysis and why is it important?
9. Explain the basic principle of structural equation modeling.
Nunnally, J. C., & Bernstein, I. H. (1993). Psychometric theory (3rd ed.). New York:
McGraw-Hill.
Psychometric theory is a classic text and the most comprehensive introduction to
the fundamental principles of measurement. Chapter 7 provides an in-depth
discussion of the nature of reliability and its assessment.
Sarstedt, M., Hair, J. F., Ringle, C. M., Thiele, K. O., & Gudergan, S. P. (2016).
Estimation issues with PLS and CBSEM: where the bias lies! Journal of
Business Research, 69(10), 3998–4010.
This paper discusses the differences between covariance-based and partial least
squares structural equation modeling from a measurement perspective. This
discussion relates to the differentiation between factor analysis and PCA and
the assumptions underlying each approach to measure unobservable
phenomena.
Stewart, D. W., (1981). The application and misapplication of factor analysis in
marketing research. Journal of Marketing Research, 18(1), 51–62.
David Stewart discusses procedures for determining when data are appropriate for
factor analysis, as well as guidelines for determining the number of factors to
extract, and for rotation.
References
Acock, A. C. (2013). Discovering structural equation modeling using Stata (Revised ed.). College
Station: Stata Press.
Brown, J. D. (2009). Choosing the right type of rotation in PCA and EFA. JALT Testing &
Evaluation SIG Newsletter, 13(3), 20–25.
Carbonell, L., Izquierdo, L., Carbonell, I., & Costell, E. (2008). Segmentation of food consumers
according to their correlations with sensory attributes projected on preference spaces. Food
Quality and Preference, 19(1), 71–78.
Cattell, R. B. (1966). The scree test for the number of factors. Multivariate Behavioral Research,
1(2), 245–276.
Cliff, N. (1987). Analyzing multivariate data. New York: Harcourt Brace Jovanovich.
Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3),
297–334.
Diamantopoulos, A., & Siguaw, J. A. (2000). Introducing LISREL: A guide for the uninitiated.
London: Sage.
Dinno, A. (2009). Exploring the sensitivity of Horn’s parallel analysis to the distributional form of
random data. Multivariate Behavioral Research, 44(3), 362–388.
310 8 Principal Component and Factor Analysis
DiStefano, C., Zhu, M., & Mı̂ndriă, D. (2009). Understanding and using factor scores:
Considerations fort he applied researcher. Practical Assessment, Research & Evaluation,
14(20), 1–11.
Festge, F., & Schwaiger, M. (2007). The drivers of customer satisfaction with industrial goods: An
international study. Advances in International Marketing, 18, 179–207.
Gorsuch, R. L. (1983). Factor analysis (2nd ed.). Hillsdale: Lawrence Erlbaum Associates.
Graffelman, J. (2013). Linear-angle correlation plots: New graphs for revealing correlation
structure. Journal of Computational and Graphical Statistics, 22(1), 92–106.
Grice, J. W. (2001). Computing and evaluating factor scores. Psychological Methods, 6(4),
430–450.
Hair, J. F., Black, W. C., Babin, B. J., & Anderson, R. E. (2013). Multivariate data analysis. A
global perspective (7th ed.). Upper Saddle River: Pearson Prentice Hall.
Hair, J. F., Ringle, C. M., & Sarstedt, M. (2011). PLS-SEM: Indeed a silver bullet. Journal of
Marketing Theory and Practice, 19(2), 139–151.
Hair, J. F., Hult, G. T. M., Ringle, C. M., & Sarstedt, M. (2017a). A primer on partial least squares
structural equation modeling (PLS-SEM) (2nd ed.). Thousand Oaks: Sage.
Hair, J. F., Hult, G. T. M., Ringle, C. M., Sarstedt, M., & Thiele, K. O. (2017b). Mirror, mirror on
the wall. A comparative evaluation of composite-based structural equation modeling methods.
Journal of the Academy of Marketing Science, 45(5), 616–632.
Hair, J. F., Sarstedt, M., Ringle, C. M., & Gudergan, S. P. (2018). Advanced issues in partial least
squares structural equation modeling (PLS-SEM). Thousand Oaks: Sage.
Hamilton, L. C. (2013), Statistics with Stata: Version 12: Cengage Learning.
Hayton, J. C., Allen, D. G., & Scarpello, V. (2004). Factor retention decisions in exploratory factor
analysis: A tutorial on parallel analysis. Organizational Research Methods, 7(2), 191–205.
Henson, R. K., & Roberts, J. K. (2006). Use of exploratory factor analysis in published research:
Common errors and some comment on improved practice. Educational and Psychological
Measurement, 66(3), 393–416.
Hershberger, S. L. (2005). Factor scores. In B. S. Everitt & D. C. Howell (Eds.), Encyclopedia of
statistics in behavioral science (pp. 636–644). New York: John Wiley.
Horn, J. L. (1965). A rationale and test for the number of factors in factor analysis. Psychometrika,
30(2), 179–185.
oreskog, K. G. (1971). Simultaneous factor analysis in several populations. Psychometrika, 36(4),
J€
409–426.
Kaiser, H. F. (1958). The varimax criterion for factor analytic rotation in factor analysis. Educa-
tional and Psychological Measurement, 23(3), 770–773.
Kaiser, H. F. (1974). An index of factorial simplicity. Psychometrika, 39(1), 31–36.
Kim, J. O., & Mueller, C. W. (1978). Introduction to factor analysis: What it is and how to do it.
Thousand Oaks: Sage.
Longman, R. S., Cota, A. A., Holden, R. R., & Fekken, G. C. (1989). A regression equation for the
parallel analysis criterion in principal components analysis: Mean and 95th percentile
eigenvalues. Multivariate Behavioral Research, 24(1), 59–69.
MacCallum, R. C., Widaman, K. F., Zhang, S., & Hong, S. (1999). Sample size in factor analysis.
Psychological Methods, 4(1), 84–99.
Matsunga, M. (2010). How to factor-analyze your data right: Do’s and don’ts and how to’s.
International Journal of Psychological Research, 3(1), 97–110.
Mulaik, S. A. (2009). Foundations of factor analysis (2nd ed.). London: Chapman & Hall.
Preacher, K. J., & MacCallum, R. C. (2003). Repairing Tom Swift’s electric factor analysis
machine. Understanding Statistics, 2(1), 13–43.
Russell, D. W. (2002). In search of underlying dimensions: The use (and abuse) of factor analysis
in Personality and Social Psychology Bulletin. Personality and Social Psychology Bulletin,
28(12), 1629–1646.
Sarstedt, M., Schwaiger, M., & Ringle, C. M. (2009). Do we fully understand the critical success
factors of customer satisfaction with industrial goods? Extending Festge and Schwaiger’s
References 311