0% found this document useful (0 votes)
92 views18 pages

Elbow Method

The document discusses the use of cluster analysis in strategic management research. It analyzes 45 published studies using cluster analysis to investigate important strategic issues. The implementation of cluster analysis in these studies was often less than ideal, potentially limiting their ability to generate knowledge. Suggestions are provided for improving the use of cluster analysis in future research.

Uploaded by

Osmar Salvador
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
92 views18 pages

Elbow Method

The document discusses the use of cluster analysis in strategic management research. It analyzes 45 published studies using cluster analysis to investigate important strategic issues. The implementation of cluster analysis in these studies was often less than ideal, potentially limiting their ability to generate knowledge. Suggestions are provided for improving the use of cluster analysis in future research.

Uploaded by

Osmar Salvador
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Strategic Management Journal, Vol.

17, 441 -458 ( 1996)

THE APPLICATION OF CLUSTER ANALYSIS IN


STRATEGIC MANAGEMENT RESEARCH: AN

c ANALYSIS AND CRITIQUE


DAVID J. KETCHEN, JR.
College of Business Administration, Louisiana State University, Baton Rouge,
Louisiana, U S A .
CHRISTOPHER L. SHOOK
College of Business, Northern Illinois University, DeKalb, Illinois, U.S.A.

Cluster analysis is a statistical technique that sorts observations into similar sets or groups. The
use of cluster analysis presents a complex challenge because it requires several methodological
choices that determine the quality of a cluster solution. This paper chronicles the application of
cluster analysis in strategic management research, where the technique has been used since the
late 1970s to investigate issues of central importance. Analysis of 45 published strategy studies
reveals that the implementation of cluster analysis has been ofen less than ideal, perhaps
detracting from the ability of studies to generate knowledge. Given these findings, suggestions are
offered for improving the application of cluster analysis in future inquiry.

Strategic management research focuses on the firms that share a common profile along concep-
relationships among strategy, environment, tually distinct variables (Meyer, Tsui, and Hinings,
leadership/organization, and performance 1993; Miller and Mintzberg, 1983). For example,
(Summer et al., 1990). Each of these four con- Mintzberg ( 1989) identified configurations whose
structs is multidimensional. Strategy, for example, members were similar in size, age, formality, and
can be viewed as composed of process and content centralization. One is an ‘entrepreneurial’ con-
concerns (Ansoff, 1965), scope and resource figuration that consists of firms that are simul-
deployments (Hofer and Schendel, 1978), or cor- taneously small, young, informal, and have cen-
porate, business, and functional-level issues tralized decision making. In contrast, ‘professional
(Andrews, 1971). Similarly, environment may be bureaucracies’ are large, old, formalized, and rely
divided into task and general elements (Thompson, on decentralized decision making. Examination of
1967). Leadership/organization encapsulates a va- organizational configurations have been conducted
riety of firm characteristics, including structure andunder many labels, including strategic groups (e.g.,
culture (Summer et al., 1990). Performance con- Hatten and Schendel, 1977), organizational
sists of at least three categories: financial, oper- typologies (e.g., Miles and Snow, 1978), taxon-
ational, and overall effectiveness (Venkatraman omies (e.g., Galbraith and Schendel, 1983) and
and Ramanujam, 1986). The multidimensionality archetypes (e.g., Miller and Friesen, 1978). Thus,
of these constructs creates a conceptual challenge in essence, organizational configurations (or, more
in that a vast array of specific combinations could simply, configurations) as used here and elsewhere
be developed along these dimensions to describe (e.g., Dess, Newport, and Rasheed, 1993; Ketchen,
organizations. Thomas, and Snow, 1993) is a cover term that
One popular response to this challenge has been encapsulates a variety of research streams. Regard-
to identify ‘organizational configurations’: sets of less of the specific label, the underlying assumption
is that configurations represent a way to meaning-
fully capture the complexity of organizational
Key words: classification; cluster analysis; configu- reality.
rational research; strategic groups Due to strategic management’s emphasis on

CCC 0143-2095/96/060441-18 Received 21 November 1994


0 1996 by John Wiley & Sons, Ltd. Final revision received 24 October 1995
442 D. J. Ketchen, Jr. and C. L. Shook
identifying groups of similar organizations, cluster real conditions but instead may simply be stat-
analysis has been a popular methodology following istical artifacts that capitalize on random numeri-
its introduction into the field in a stream of cal variation across organizations (Thomas and
research often referred to as the Purdue brewing Venkatraman, 1988). Further, cluster analysis’
studies (i.e., Hatten, 1974; Hatten and Schendel, sorting ability is powerful enough that it will pro-
1977; Hatten, Schendel, and Cooper, 1978). Prior vide clusters even if no meaningful groups are
to these efforts, the search for configurations had embedded in a sample (Barney and Hoskisson,
been centered in the industrial/organization eco- 1990). Thus, cluster analysis has the potential
nomics literature, where groups were defined not only to offer inaccurate depictions of the
across narrow sets of variables, often only one or groupings in a sample but also to impose group-
two (e.g., Hunt, 1972; Porter, 1973). Such analyses ings where none exist.
allowed mapping of the structure of samples, but These concerns and others have led some
were too coarse-grained to capture the multidimen- observers to suggest that the frequent use of clus-
sionality of the constructs of interest in strategy ter analysis is an embarrassment to strategic
research (Hatten and Hatten, 1987). Cluster analy- management. For example, when Meyer ( 1991)
sis, which takes a sample of elements (e.g., asked prominent researchers to assess the field,
organizations) and groups them such that the stat- cluster analysis was singled out as a source of
istical variance among elements grouped together ‘methodological stigma’. Further, one claimed
is minimized while between-group variance is that ‘the empiricism inherent in this method has
maximized, addresses this limitation. Specifically, been forever branded upon our collective back-
cluster analysis permits the inclusion of multiple side’ (Meyer, 1991: 826). Fueling this contempt
variables as sources of configuration definition. For are the equivocal results that cluster analysis
example, Hatten et al. (1978) drew on 16 variables often provides. The best illustration is strategic
representative of manufacturing, financial, and groups research, which has been unable to con-
market strategies as well as environmental charac- sistently find a group membership-performance
teristics. Thus, cluster analysis can provide very link (Barney and Hoskisson, 1990). Thus, at the
rich descriptions of configurations without over- moment, the value of cluster analysis to strategy
specifying the model. research is in doubt.
However, the use of cluster analysis in stra- Despite the controversy surrounding cluster
tegic management research has come under fre- analysis, to date there has been no comprehen-
quent attack (e.g., Barney and Hoskisson, 1990; sive assessment of the efficacy of its use. Thus,
Meyer, 1991; Thomas and Venkatraman, 1988). the purpose of this paper is to examine the appli-
One cause for concern is the extensive reliance cation of cluster analysis in the field of strategic
on researcher judgment that is inherent in cluster management. Our aim is not only to evaluate the
analysis. This is an issue throughout the process past use of cluster analysis but also to lay out an
of using the technique, but perhaps most troub- agenda to guide future research toward making
ling is the fact that, unlike techniques such as the best possible use of the technique. We begin
regression and analysis of variance, cluster by describing the issues identified in methodol-
analysis does not offer a test statistic (such as an ogical research as critical when using cluster
F-statistic) that provides a clear answer regard- analysis as well as the appropriate ways for strat-
ing the support or lack of support of a set of egy research to address them. Next, we document
results for a hypothesis of interest. Instead, to a all of the strategy articles published in the Acad-
large extent, it is the researchers who are the emy of Management Journal, Journal of Inter-
arbiters of the meaning of results acquired national Business Studies, Journal of Manage-
through cluster analysis. ment, Management Science, and Strategic
A second major issue for critics is their per- Management Journal from 1977 to 1993 that use
ception that most applications of cluster analysis the technique. Based on a comparison of existing
in strategy have lacked an underlying theoretical practices with the guidelines developed through
rationale. Often clustering dimensions seem to be our review of methodological research, we then
selected haphazardly (Reger and Huff, 1993). offer recommendations designed to maximize the
Without a theoretical foundation, however, the value of cluster analysis to strategic manage-
clusters that are identified may not reflect any ment.
Application of Cluster Analysis 443

CRITICAL ISSUES IN THE USE OF follow McKelvey’s (1975, 1978) suggestion to


CLUSTER ANALYSIS consider as many variables as possible because one
cannot know in advance which variables differen-
This section describes the key issues involved tiate among observations. Thus, the use of many
when using cluster analysis. It is important to note clustering variables is expected to maximize the
that there is considerable variation in how unified likelihood of discovering meaningful differences.
methodological experts are regarding these issues. One example of an inductive study is Hambrick
Specifically, experts tend to agree on what the (1983), where the author used 10 environmental
important issues are, but often disagree about how variables to develop a taxonomy of mature indus-
to address them. As described below, there are tries with no a priori expectations about the likely
some limited areas of agreement regarding appro- name of the resultant types.
priate ‘remedies’ for clustering issues. There is a When following a deductive approach, the num-
larger body of issues, however, for which no con- ber and suitability of clustering variables, as well
sensus has emerged. Here we describe the dis- as the expected number and nature of groups in a
agreements and discuss how the issue should be cluster solution, are strongly tied to theory
dealt with in strategic management research to (Ketchen et al., 1993). Methodological research
satisfy its own, often unique, requirements. Thus, suggests that using deductive theory to guide vari-
in sum, one aspect of this section is the collection able choice is often wise. Cluster analysis derives
and synthesis of the wisdom offered by methodol- the most internally consistent groups across all
ogical experts. A perhaps more important element variables, thus irrelevant variables can cause a
is the identification of the strengths, weaknesses, deterioration of a solution’s validity (Punj and Ste-
and trade-offs involved with each issue as it relates wart, 1983). An implication is that, when possible,
to the needs of the strategic management literature. studies should focus on variables with solid theor-
To avoid overlap with material that is available etical foundations. One example of a deductive
elsewhere, we do not provide detailed explanations study is Lawless and Finch (1989), where theory-
of the mechanics of specific aspects of cluster based predictions of the relative performance of
analysis. Such information is offered by a number different configurations are tested within each type
of prior efforts (e.g., Aldenderfer and Blashfield, of environment identified by Hrebiniak and
1984; Everitt, 1980; Hair et al., 1992; LOK,1983; Joyce (1985).
Punj and Stewart, 1983); the interested reader is The cognitive approach can be viewed as a ‘con-
referred to these sources. ceptual cousin’ of the inductive approach because
both avoid making theory-based predictions. While
inductive configurations are defined along dimen-
Clustering variables
sions that researchers view as important, the cogni-
Choosing the variables along which to group tive approach relies on the perceptions of expert
observations is the most fundamental step in the informants such as industry executives (e.g., Mas-
application of cluster analysis, and thus, perhaps carenhas and Aaker, 1989a, 1989b; Reger and
the most important. This process involves three Huff, 1993) to define clustering variables. This lat-
critical issues: (1) how to select variables; (2) ter approach has its roots in research on interpret-
whether or not to standardize variables; and (3) ation in organizations, which posits that it is the
how to address multicollinearity among variables. meaning that top managers attach to phenomena,
not ‘objective’ characteristics, which directs sub-
sequent organizational action (Dutton, Fahey, and
Selection of variables
Narayanan, 1983) and performance (Thomas,
There are three basic approaches to identifying Clark, and Gioia, 1993). One implication is that
appropriate clustering variables: inductive; deduc- configurations based on the perceptions of top
tive; and cognitive (Ketchen et al., 1993). The managers may be crucial to understanding any
inductive approach focuses on exploratory classi- given setting (Porac and Thomas, 1990).
fication of observations. In other words, neither the We suggest that the approach to selecting vari-
clustering variables nor the number and nature of ables should match a study’s purpose. When
the resultant groups are tightly linked to deductive attempting to explain or predict relationships, a
theory. Instead, the inductive approach seems to theoretical foundation is advisable, if not required
444 D.J. Ketchen, Jr. and C. L. Shook
(Bacharach, 1989). Hence, studies designed to dis- inconsistent across the two solutions, the validity
cern the nature and extent of links between key of each should be assessed; the solution exhibiting
constructs (e.g., organizational configurations and the highest validity might then be adopted.
performance) should rely on a deductive approach
(Ketchen et al,, 1993). However, strategy research
Multicollinearity among variables
is often exploratory, with a focus on theory build-
ing rather than testing. Here variables should be High correlation among clustering variables can be
chosen in a way that fosters rich description of a problematic because it may overweight one or
sample’s characteristics (Meyer et al., 1993). Both more underlying constructs. Thus, researchers may
the inductive and cognitive approaches fit this want to correct multicollinearity, especially if it is
requirement. The latter may often be preferred, desirable that constructs be equally weighted. Hair
however, because its use of experts (often top et al. (1992) suggest using the Mahalanobis dis-
managers) enhances confidence that the variables tance measure, which both standardizes variables
are important in a particular data set. and adjusts for high correlations. As noted above,
however, standardization is controversial. Another
problem is that statistical programs such as SAS
Standardization of variables
and SPSS do not offer this measure.
Because cluster analysis groups elements (e.g., Multicollinearity can also be addressed through
firms) such that the distance between groups along subjecting variables to factor analysis (specifically,
all clustering variables is maximized, variables principal components analysis with orthogonal
with large ranges (i.e., where elements are separ- rotation) and using the resultant uncorrelated factor
ated by large distances) are given more weight in scores for each observation as the basis for clus-
defining a cluster solution than those with small tering ( h n j and Stewart, 1983). However, this
ranges (Hair et aL, 1992). As a result, a subset of technique is controversial because researchers
variables can dominate the definition of clusters. often drop all factors with low eigenvalues (a stat-
The ‘remedy’ is standardization, which transforms istic representing the amount of variance explained
the distribution of elements along variables so that by a factor). The excluded factors may represent
each has a mean of zero and a standard deviation unique, important information (Dillon, Mulani, and
of one. This process allows variables to contribute Frederick, 1989), meaning that a less-than-optimal
equally to the definition of clusters but may also set of clusters may result. Thus, as with standardiz-
eliminate meaningful differences among elements ation, any remedy for multicollinearity has a cost.
(Edelbrock, 1979). This illustrates a dilemma of Because both methods of correcting multicolli-
using cluster analysis: for any remedy, there is nearity have potential pitfalls, researchers should
almost always an associated cost.’ attempt to assess the impact of their chosen tech-
Given this trade-off, whether to standardize clus- nique. We suggest that the ideal approach is to per-
tering variables is an equivocal issue. Some experts form a cluster analysis multiple times changing
(Hair er al., 1992; Hanigan, 1985) recommend only the method of addressing multicollinearity.
standardization, perceiving a need to eliminate the Consistent group assignments despite different
potential effects of scale differences among vari- methods would be evidence of stability whereas
ables. Others offer evidence that standardization inconsistent assignments would suggest a tenuous
has no significant effects (Edelbrock, 1979; Milli- cluster solution.
gan, 1980). Aldenderfer and Blashfield (1984)
advise that because standardization may have
Clustering algorithms
adverse effects, it should be addressed on a case-
by-case basis. However, they do not offer specific The selection of appropriate clustering algorithms
guidance about how to approach a particular case. (i.e., the rules or procedures followed to sort
Because results may differ solely based on stan- observations) is critical to the effective use of clus-
dardization, we suggest that analyses be done both ter analysis ( h n j and Stewart, 1983). There are
using and not using standardization. If clusters are two basic types of algorithms: hierarchical and
nonhierarchical.
Hierarchical algorithms progress through a ser-
’ We wish to thank an anonymous reviewer for this insight. ies of steps that build a tree-like structure by either
Application of Cluster Analysis 445
adding individual elements to (i.e., agglomerative) into one group. It is the task of the researcher to
or deleting them from (i.e., divisive) clusters. The decide at what point the number of groups is
five most popular agglomerative algorithms are appropriate. Polythetic divisive methods follow the
single linkage, complete linkage, average linkage, opposite approach: all observations are in one
centroid method, and Ward’s method (Hair et al., group initially, then observations are divided into
1992). The differences among them lie in the smaller groups until eventually each observation
mathematical procedures used to calculate the dis- becomes a separate cluster. Again, the researcher
tance between clusters. Each has different system- must decide what level of division is appropriate.
atic tendencies (or biases) in the way it groups Although the methods start at opposite ends of the
observations. For example, the centroid method clustering process, the number of groups identified
has a bias toward producing irregularly shaped should be the same regardless of which one is used.
clusters. Further, it can only be used with interval Thus, the distinction between the two methods is
or ratio data (Hair et al., 1992). Ward’s method of little practical consequence. Should one wish to
tends to produce clusters with roughly the same use divisive methods, however, a procedure using
number of observations (SAS Institute, 1990) and matrix algebra is described in Everitt ( 1980).
the solutions it provides tend to be heavily dis- All hierarchical algorithms suffer from several
torted by outliers (i.e., observations with extreme problems. First, researchers often do not know the
values-Milligan, 1980). Given such tendencies, underlying structure of a sample in advance, mak-
there should be a match between the algorithm ing it difficult to select the ‘correct’ algorithm.
selected and the underlying structure of focal data Second, these algorithms make only one pass
(i.e., sample size, distribution of observations, and through a data set, thus poor cluster assignments
what types of variables are included-nominal, cannot be modified.* Finally, solutions are often
ordinal, ratio, or interval). Thus, for example, the unstable when cases are dropped, especially when
centroid method should only be used when (a) data a sample is small (Jardine and Sibson, 1971). This
are measured with interval or ratio scales and (b) is troublesome for strategy research, where sample
clusters are expected to be very dissimilar from sizes are often small (e.g., 19 in Dess and Davis,
each other. Likewise, Ward’s method is best suited 1984; 16 in Lewis and Thomas, 1990). Because
for studies where (a) the number of observations of these problems, confidence in the validity of a
in each cluster are expected to be approximately solution obtained using only hierarchical methods
equal and (b) there are no outliers. is limited.
The use of divisive methods in the social sci- Nonhierarchical algorithms (also referred to as
ences has been limited to the field of archaeology K-means or iterative methods) partition a data set
(e.g., Whallon, 1972). As a result, divisive into a prespecified number of clusters. Specific
methods are not well known in strategic manage- nonhierarchical methods vary slightly, but function
ment. There are two types of divisive techniques: in essentially the same manner (Hair et al., 1992).
monothetic and polythetic (Everitt, 1980). The
monothetic techniques are used with binary (i.e.,
dichotomous) variables, A sample is divided into Only one pass is made through the data because of the hier-
archical methods treatment of the [Link] on the
groups based on each observation’s possession (or agglomerativealgorithms for a moment, each observation starts
lack) of an attribute. Groups are then broken into as its own individual cluster. On each successive step, the two
smaller groups based on the presence or absence closest clusters are joined into a new aggregate cluster, thus
reducing the number of clusters by one in each step. Eventually,
of individual attributes. Because this procedure all clusters are joined into one cluster. Multiple passes can not
groups observations through successive rather than be made because the algorithm would treat the observations as
simultaneous application of variables, it would not individual clusters in the first step of the next pass. Thus, the
same clusters would emerge from clustering the observations
be useful for configurational research as it has tra- again. Given that the divisive methods follow the same process
ditionally been conducted in strategic management. in reverse, the same arguments apply: multiple passes would
Polythetic divisive methods, in essence, are the produce the same results.
However, multiple passes with hierarchical algorithms might
logical opposite or ‘mirror image’ of agglomer- be helpful if the initial pass is used as a tool to identify outliers.
ative methods. Agglomerative methods initially One could then remove the outliers from the data set (assuming
view each observation as a separate cluster and that this would be consistent with the purpose of the study) and
run the analysis again. This might be particularly valuable if
then compile them into successively smaller num- one has selected an algorithm that is particularly sensitive to
bers of groups, eventually putting all observations outliers (e.g., Ward’s method).
446 D. J. Ketchen, Jr. and C. L. Shook

After initial cluster centroids (the ‘center points’ looks for natural clusters of the data that are indi-
of clusters along input variables) are selected, each cated by relatively dense ‘branches’. This method’s
observation is assigned to the group with the near- reliance on interpretation requires that it be used
est centroid. As each new observation is allocated, cautiously (Aldenderfer and Blashfield, 1984).
the cluster centroids are recomputed. Multiple The agglomeration coefficient (i.e., a numerical
passes are made through a data set to allow obser- value at which various cases merge to form a
vations to change cluster membership based on cluster) is the basis for two related techniques. The
their distance from the recomputed centroids. To first method involves graphing the coefficient on a
arrive at an optimal solution, passes through a data y-axis and the number of clusters on an x-axis. A
set continue until no observations change clusters marked flattening of the graph suggests that the
(Anderberg, 1973). clusters being combined are very dissimilar, thus
Nonhierarchical methods have two potential the appropriate number of clusters is found at the
advantages over hierarchical methods. First, by ‘elbow’ of the graph. Interpreting a graph, how-
allowing observations to switch cluster member- ever, may be difficult; for example, the elbow may
ship, nonhierarchical methods are less impacted by not be pronounced, indicating that there may not
outlier elements. Although outliers can initially be any natural groups in the data (Hambrick and
distort clusters, this is often corrected in sub- Schecter, 1983). Alternatively, the graph may have
sequent passes as the observations switch cluster more than one elbow, indicating that more than one
membership (Aldenderfer and Blashfield, 1984; natural set of clusters fit the data (Aldenderfer and
Hair et al., 1992). Second, by making the multiple Blashfield, 1984). The second procedure involves
passes through the data, the final solution optim- examining the incremental changes in the coef-
izes within-cluster homogeneity and between-clus- ficient. A large increase implies that dissimilar
ter heterogeneity. Obtaining this improvement, clusters have been merged; thus, the number of
however, requires that the number of clusters be clusters prior to the merger is most appropriate. A
specified a priori (Milligan, 1980). In many fields major limitation with this approach is that there
(including strategic management), this is problem- may be no large jumps in the coefficient, indicating
atic because cluster analyses are often exploratory. that there may not be any natural groups in the
A solution advocated by many experts is to use data. In some cases, there may be several large
a two-stage procedure where a hierarchical algor- jumps; this would be evidence for more than one
ithm is used to define the number of clusters and natural set of clusters.
cluster centroids; these results then serve as the The cubic clustering criterion (CCC) is a meas-
starting points for subsequent nonhierarchical clus- ure of within-cluster homogeneity relative to
tering (Hair et aL, 1992; Milligan, 1980; Punj and between-cluster heterogeneity. The ‘appropriate’
Stewart, 1983). Research has shown that this pro- number of clusters is indicated by the peaking of
cedure increases validity of solutions (Milligan, the CCC; however, Milligan and Cooper (1985)
1980; Punj and Stewart, 1983). The only cost is found that this test may suggest too many clusters.
the extra time and effort required on the Strategy researchers need to be aware of this
researchers’ part; a cost we contend is worth bear- potential because several use the CCC (e.g., Fom-
ing. Thus, in summary, the best solutions may be brun and Zajac, 1987; Mascarenhas, 1989), but
those obtained by using hierarchical and nonhier- many oft-cited methodological works predate
archical methods in tandem. Milligan and Cooper’s finding (e.g., Aldenderfer
and Blashfield, 1984) or ignore the CCC (e.g., Hair
et al., 1992).
Determining the number of clusters Finally, a priori theory can serve as a nonstat-
A variety of techniques are available to determine istical tool for determining the number of clusters
the number of clusters in a data set. When using (Hair et aL, 1992). Although a priori theory is, by
hierarchical methods, the most basic procedure is definition, not central to exploratory research, it
to visually inspect a dendogram, a graph of the does provide a benchmark for assessing the results
order that observations join clusters and the simi- of theory-testing inquiry. For example, comparison
larity of observations joined. Dendograms of emergent clusters with a theory-based typology
resemble decision trees with short ‘limbs’ rep- can provide evidence regarding the typology’s
resenting the joining of observations. A researcher descriptive validity (e.g., Ketchen et aL, 1993).
Application of Cluster Analysis 447
In summary, using a single method to determine can turn to external validity. This may be done by
the number of clusters is questionable because each cluster analyzing both the sample of interest and a
method has limitations (Everitt, 1980). Thus, we second, similar sample and then assessing the simi-
advocate the use of multiple techniques that can larity of the results (Hairet al., 1992; Hambrick,
overcome each others’ shortcomings. For example, 1983). In many strategy studies, however, a ‘hold-
as noted above, the CCC may indicate too many out’ sample is not available. In other instances, the
clusters. If the CCC is part of an array of tech- use of a second sample may not even be appropri-
niques and other techniques point to fewer clusters ate. For example, strategic groups are often viewed
than does the CCC, it might be wise to discount the as industry specific (Thomas and Venkatraman,
CCC. Thus, confidence in the number of clusters 1988), and thereby cannot be generalized to
identified may be greater when determined through another setting. Thus, validation using multiple
the convergence of multiple methods. samples should be used only if consistent with the
assumptions underlying a study.
Criterion-related validity can be assessed through
Validating clusters
significance tests (often multivariate analysis of
The goals of validation are to ensure that a cluster variance) with external variables (Aldenderfer and
solution has external validity (i.e., is representative Blashfield, 1984, Anderberg, 1973). Such variables
of the general population of interest-Cook and should be theoretically related to the clusters, but not
Campbell, 1979) and criterion-related validity (ie., used in defining clusters. Given the field’s emphasis
is useful for the prediction of important on defining the strategy-performance relationship
outcomes-Kerlinger, 1986). Extreme care in vali- (Summer et af., 1990), the external variables in strat-
dation is warranted because, despite the rigor used egy research are often performance measures (e.g.,
in previous steps, without validation one is not Miller, 1988; Robinson and Pearce, 1988). Signifi-
assured of having arrived at a meaningful and use- cance tests with external variables offer a powerful
ful set of clusters (Punj and Stewart, 1983). tool to establish validity of a cluster solution because
Reliability (i.e., consistency) is a necessary but the technique uses a test static (often an F-statistic),
not sufficient condition of validity (Kerlinger, thereby avoiding having the researcher provide the
1986). Therefore, the reliability of a cluster sol- meaning of results. External variables are expensive
ution must be established before validity is tested. to obtain in many fields (Bailey, 1994),but, in strat-
There are two primary ways to evaluate reliability. egy research, the availability of archival data often
First, as advocated above, researchers may perform solves this problem. Thus, we strongly advocate the
a cluster analysis multiple times, changing algor- use of this technique whenever possible.
ithms and methods for addressing multicollinear- In sum, techniques are available to help establish
ity. The degree of consistency in solutions indi- reliability and external validity, but the value of
cates reliability (Hair et af., 1992). Second, these techniques is limited because they use cluster
researchers may split a sample and analyze the two analysis and thus are subject to its inherent prob-
halves independently (Hambrick, 1983). A modi- lems, most notably the reliance on researcher judg-
fied version of this latter procedure is to obtain ment. More promising is the use of significance
cluster centroids from half of a sample and use tests with external variables to establish criterion-
them to define clusters in the other half. In either related validity. Overall, we suggest that reliability
case, consistency across sample halves indicates and validity will be questionable whenever a
reliability (Hair et af., 1992). However, there is no research design uses clustering techniques in iso-
standard for assessing a satisfactory level of con- lation. Only when cluster analysis is augmented
sistency, leaving this determination largely to with additional techniques-especially ones that
researcher judgment. Also, in some strategy stud- are less subject to researchers’ biases-can confi-
ies, sample sizes may be too small for meaningful dence in the results obtained be strong.
clusters to be derived from sample halves. For
example, splitting their sample of 18 companies Summary
probably would not have been helpful to Reger and
Huffs (1993) efforts to establish the reliability of As described above, strategy researchers using
the three strategic groups they found. cluster analysis face an array of challenges. None-
If reliability has been demonstrated, attention theless, the technique has been widely used over
448 D.J. Ketchen, Jr. and C. L. Shook
the last 15 years. Much of this research (e.g., the SMJ publishes only strategy-related inquiry, while
strategic groups literature) has seemingly produced AMJ and JOM contain studies representative of the
equivocal results, perhaps in part because of the entire management discipline and JIBS and MS
role played by researcher judgment throughout the include research from a wider spectrum of business
clustering process. Given the controversy sur- disciplines. It is interesting to note, however, that
rounding cluster analysis, we examined the extent the other six empirical journals contained no stud-
to which the technique has been used appropri- ies using cluster analysis.
ately. A description of this effort and the related
results are presented below.
Coding procedure
The critical issues described above served as the
METHOD basis for our coding scheme. Specifically,each study
was coded according to issues involving clustering
To identify important strategy studies that have variables, clustering algorithms, determining the
used cluster analysis, we examined the 16 journals number of clusters, and validating clusters.
comprising the ‘forum for strategy research’
(MacMillan, 1991). Of these journals, five with a Clustering variables
non-empirical orientation were eliminated? leav-
ing 11 journals as sources of empirical strategy Several issues related to the selection of clustering
research, including the Academy of Management variables were coded. The justification of variables
Journal (AMJ),Administrative Science Quarterly, was coded as inductive, deductive, or cognitive.
Decision Science, Journal of General Manage- Studies focused on exploratory classification of
ment, Jounzal of International Business Studies observations (i-e., the clustering variables were not
(JIBS), Journal of Management (JOM), Journal of tightly linked to deductive theory) were coded as
Management Studies, Management Science (MS), inductive. Studies that tightly linked clustering
Omega, Rand Journal of Economics, and Strategic variables to theory were coded as deductive. Stud-
Management Journal (SMJ). These remaining ies where clustering variables were based on
journals were searched for strategy studies using experts’ opinions were coded as cognitive. We also
cluster analysis for the period from 1977 (the year coded whether or not variables were standardized
of the first published strategy research using cluster and if the Mahalanobis distance measure or factor
analysis-Hatten and Schendel, 1977) through analysis were used to address multicollinearity.
1993. Strategy studies were defined as empirical
research efforts examining relationships among
Clustering algorithms
strategy, environment, leadership/organization,
and performance (Summer et al., 1990). Studies’ clustering algorithms were coded as hier-
A total of 45 studies (see Appendix) were archical, nonhierarchical, hierarchical and non-
located? The use of cluster analysis has increased hierarchical used in tandem, or not specified.
over time: seven studies were found in the period’s Where appropriate, the specific hierarchical clus-
first half (1977-85) whereas 38 were found in the tering algorithms were coded as well.
second half (1986-93). Most studies (n=28,
62%) were found in SMJ. Eight studies (18%)
Determining the number of clusters
were found in AMJ, five ( 11%) in JIBS, three (7%)
in JOM, and one (2%) in MS. The distribution of The methods used to determine the appropriate num-
studies across these journals is not surprising, as ber of clusters were coded. The methods included:
dendogram observation, change in agglomeration
coefficient, cubic clustering criterion, a priori theory,
These non-empirical journals include Academy of Manage- and other methods specified by authors.
ment Executive, Academy of Management Review, California
Management Review, Harvard Business Review, and Sloan
Management Review.
4 0 n e study (Lawless and Finch, 1989) contained two inde- Validating clusters
pendent applications of cluster analysis (one of environment
types and the other of strategy types); each application was Similarly, the methods used for validating sol-
counted as one ‘study’. utions were coded, including: multiple clustering
Application of Cluster Analysis 449
algorithms, split sample, hold-out sample, statisti- algorithms alone were used in eight studies (18%).
cal tests on nonclustering variables, and other Finally, in seven studies (16%), there was no men-
methods specified by authors. tion of the algorithm(s) used.

Coding reliability Determining the number of clusters


All studies were coded independently by the two Confidence in the number of clusters is greater
authors. To ensure consistency, a random sample when multiple methods converge; however, only
of nine studies was coded. The resultant interrater 18 studies (40%) used this type of approach. In 21
reliability, as measured by percentage of agree- studies (47%), one method was used; techniques
ment, was 87 percent. A meeting was held to dis- used were not identified in six studies (13%).
cuss the discrepancies. The remaining 36 studies Individual methods were also tallied; change in
were then coded with a 93 percent interrater the agglomeration coefficient (n = 19, 42%) and
reliability, which compares favorably to the 83 per- observing breaks in the dendogram ( n = 16, 36%)
cent interrater reliability obtained in a similar study were the most widely cited. The CCC was used in
by Ford, MacCallum, and Tait (1986). All discrep- five studies (ll%), while a priori theory guided
ancies were resolved by the authors reviewing the cluster definition in four studies (9%). In 13 stud-
study and coming to a joint decision; relevant items ies (29%), other methods not grounded in the
were recoded accordingly. methodological literature were indicated, includ-
ing, for example, ensuring equal-size groups and
‘interpretability’.
RESULTS
Clustering variables Validating clusters
Most studies (n=35, 78% of the total studies Validation may be the most neglected issue in clus-
located) used an inductive approach to select clus- ter analysis. Ideally, reliability and both external
tering variables. Four studies (9%) took a deduct- and criterion-related validity should be assessed,
ive approach; in six studies (13%), a cognitive but none of the 45 studies examined all three.
approach was evident. In 12 of the 45 studies Further, no validation techniques were cited in 17
(27%), variables were standardized; 14 studies studies (38%).
(32%) used factor scores as the basis for clustering Reliability was addressed in 13 studies (29%);
while none used the Mahalanobis distance meas- all of these efforts used multiple clustering algo-
ure. The results for these and other coded items are rithms and four (9%) also analyzed split samples.
presented in Table 1. External validity was tested by analyzing hold-out
samples in two studies (4%). While such efforts
can be valuable in establishing validity, this value
Clustering algorithms is sharply limited because each technique shares
Only six studies (13%) used the preferred the fundamental concern inherent in cluster analy-
approach of hierarchical and nonhierarchical algor- sis: the heavy reliance on researcher judgment. In
ithms in tandem. In combination with nonhier- 10 studies (22%), statistical tests on nonclustering
archical clustering, Ward’s method was cited three variables were done to assess criterion-related val-
times, and the centroid method, complete linkage, idity.6 These latter studies are noteworthy because
and average linkage were each used in one study. they validated results using a technique (i.e.,
Hierarchical algorithms alone were cited in 24 ANOVA or MANOVA) where researcher judg-
studies (35%). Ward’s method was used in 16 of ment is limited to the choice of criterion variables.
these studies, the centroid method in four, com-
plete linkage in and average linkage and We did not include in this count several studies that examined
CONCOR’ each in one study. differences between clusters along non-clustering variables as
part of their hypothesis testing but did not mention that such
efforts also provide evidence regarding validity. Our reasoning
CONCOR is an algorithm introduced by Breiger, Booman, was that many readers would not be aware of this relation to
and Arabie (1975) for clustering relational data in social net- validity unless alerted to it and thus could not use this infor-
work analysis. mation to make informed judgments about the results.
450 D. J. Ketchen, Jr. and C. L. Shook
Table 1. Summary of decisions made by strategy researchers when using cluster analysis

Number of studies Percent of total

Clustering variables
Justification of input variables
Inductive 35 78
Deductive 4 9
Cognitive 6 13
Factor analysis used 14 32
Variables standardized 12 27

Clustering algorithms
Hierarchical
Ward’s method 16 36
Centroid method 4 9
Complete linkage 2 4
Average linkage 1 2
CONCOR 1 2
Nonhierarchical 8 18
Combination
Ward’s method and K-means 7
Centroid method and K-means 2
Complete linkage and K-means 2
Average linkage and K-means 2
Not specified 16

Determining the numbers of clusters


Multiple methods 18 40
Single method 21 47
Not specifiedhone 6 13
Specific methods
Change in agglomeration coefficient 19 42
Dendogram observation 16 36
Cubic clustering criterion 5 11
A-prion theory 4 9
Other techniques 13 29

Validating clusters
Reliability
Multiple algorithms 13 29
Split-half samples 4 9
External validity
Hold-out samples 2 4
Criterion-related validity
Statistical tests on non-clustering variables 10 22
Other
Statistical tests on clustering variables 11 24
Expert opinion 3 7
Not specifiedhone 17 38

Two techniques not advocated in the methodo- Kim and Lim, 1988). Such tests, however, should
logical literature were found. Eleven studies (24%) always be significant because cluster analysis, by
used MANOVA to demonstrate that the means of design, creates groups that have minimal overlap
the clustering variables differ across clusters and along the clustering variables (Aldenderfer and
took this as evidence that the cluster solution Blashfield, 1984). Hence, this technique does little
reflects the underlying structure of a data set (e.g., to establish the validity of a solution. A final
Application of Cluster Analysis 45 1

technique-expert opinions of clusters’ validity- relevant methodological research. Methodological


is not widely discussed by methodologists but was texts generally agree that the critical issues when
used in three studies (7%). using cluster analysis are: selecting clustering vari-
In total, our findings indicate that the ‘norm’ for ables; whether or not to standardize variables;
strategy research using cluster analysis is to (1) addressing multicollinearity among variables; se-
select variables inductively, (2) use a hierarchical lecting appropriate clustering algorithms;
or nonhierarchical algorithm by itself, and (3) pay determining the number of clusters; and validating
limited attention to determining the number of and clusters. These texts often disagree, however, about
validating clusters. how to address the issues. This may help to explain
why many studies suffer from shortcomings:
because the available guidance is often unclear or
DISCUSSION even contradictory, rigor is an elusive achieve-
ment.
The search for organizational configurations has In analyzing extant research, we found problems
long been a focus of the field of strategic manage- at each step in the clustering process. First, far
ment. Cluster analysis has played a key role in this more authors have used an inductive approach to
research because it allows for the inclusion of mul- the selection of clustering variables than have con-
tiple variables as sources of configuration defi- ducted deductive or cognitive inquiry. In 1993, for
nition, thus enabling the creation of potentially rich example, two-thirds of the studies derived con-
descriptions. Despite this strength, the use of clus- figurations based on an inductive approach.
ter analysis has been widely criticized (e.g., Barney Although such exploratory efforts were rec-
and Hoskisson, 1990; Meyer, 1991). As with any ommended in earlier guidance (Hambrick, 1984;
technique, however, the results provided by cluster Harrigan, 1985), they are less desirable now given
analysis are only as good as its implementation the advancing state of theory development about
and, more generally, the overall design of a study. strategy and the emergence of the cognitive
With this in mind, below we assess the current approach. Indeed, the relatively slow rate of
state of cluster analysis research. Our efforts here knowledge accumulation in studies using cluster
are both general and specific. We discuss overall analysis (most notably in the strategic groups
trends but also examine specific methodological literature-Ketchen et al., 1993) may be attribu-
decisions that authors made (Le., what they did table in part to the imbalance of the past. Thus, we
‘right’ and ‘wrong’), assess the likely empirical believe that the divergent contributions to knowl-
consequences of poor choices, and note how such edge promised by each approach require that schol-
consequences might have been prevented or at ars pursue a more balanced research agenda (cf.
least minimized. We then look to the future and Montgomery, Wernerfelt and Balakrishnan, 1989).
offer suggestions to maximize the value of cluster Specifically, there is a need to ‘reconfigure’ con-
analysis to strategy inquiry. Underlying these figurational research toward a substantial increase
suggestions is our assertion that the ill-effects of in the attention paid to deduction and cognition.
the weaknesses of cluster analysis can be con- The standardization of clustering variables has
trolled through triangulation; bringing to bear mul- advocates and detractors among methodologists.
tiple, disparate methods (Denzin, 1978; Jick, When no clear guidance is available, as is the case
1979). Specifically, for cluster analysis to be a here, researchers need to make choices guided by
helpful tool in the effort to create knowledge about the goals of their study and the nature of their data
organizations, the technique needs to be embedded (McGrath, Martin, and Kulka, 1982). We suggest
in research designs that include other methods that that researchers take a very conservative path:
are far less subject to researcher judgment. specifically, that they perform analyses both using
and not using standardization and compare the
results. No strategy studies to date have taken this
The state of cluster analysis research
approach; most did not standardize. Careful atten-
Our results suggest that the ability of research tion to this issue is needed, however, because strat-
using cluster analysis to generate knowledge has egy studies often include both large- and small-
been hindered by the technique’s implementation. scale variables. For example, in Ketchen et al.
This is perhaps not surprising given the state of the ( 1993), clustering variables included both number
452 D. J. Ketchen, Jr. and C.L. Shook
of hospital beds (a relatively large-scale measure, ated when using a single method, we consider the
with a mean of approximately 250) and hospitals’ use of Ward’s method, the most often used hier-
current ratio (a relatively small-scale measure, archical algorithm in strategy research. As noted
with a mean of approximately 2.3). This disparity earlier, Ward’s method is greatly impacted by out-
led the authors to (appropriately) standardize their liers. Although we can not be conclusive (due to
variables, but confidence in the results would have the lack of detail presented in many studies), at
been heightened if (a) they had done their analysis least some researchers may not have paid sufficient
with nonstandardized variables as well and (b) the attention to the role of outliers. In most cases, all
two sets of results were consistent. observations in the sample were represented in the
Addressing multicollinearity presents a similar cluster solution (e.g., Birley and Westhead, 1990;
dilemma; clear guidance is lacking, so researchers Lewis and Thomas, 1990; Manu, 1992; Zahra and
must use their best judgment, given the specifics Covin, 1993), indicating either there were no out-
of each study. Some pieces have seemingly fallen liers or the outliers were ignored. In each of these
short here. We found some instances, for example, studies, using an iterative method to refine the sol-
where high correlations among clustering variables ution derived using Ward’s method would have
(i.e., 0.50 or higher) were evident but no corrective enhanced validity because the iterative methods do
action was reported nor was justification offered not share the biases of Ward’s. Thus, while the
for not acting (e.g., Hitt and Tyler, 1991; Ketchen varied, often contradictory guidance offered by
et al., 1993; Roth, 1992). Many other authors did methodologists can be frustrating, strategy
not report the correlations among clustering vari- researchers need to pay careful attention to those
ables (e.g., Dess and Davis, 1984; Hambrick, 1983; instances where methodologists agree on the cor-
Manu, 1992; Miles, Snow, and Sharfman, 1993). rect tack.
Authors may feel that reporting such correlations Turning to the issue of determining the number
is unnecessary because the linear, continuous of clusters in a solution, each of the available tech-
relationships among variables are not central to niques has biases, leading us to echo Everitt’s
configurational research. Without correlations, (1980) call for the use of multiple techniques.
however, the reader is unable to make an informed Unfortunately, the majority of strategy studies did
judgment about whether or not multicollinearity not employ multiple methods; in these studies,
was a concern. This is illustrative of a more general clusters may have been shaped by the biases of a
problem: articles often do not describe how crucial solitary method. Given this potential, it is discon-
issues were addressed. In any study, the rationale certing that two influential articles (Le., Dess and
underlying methodological decisions should be Davis, 1984; Hambrick, 1983) are in this group.
presented in sufficient detail to allow readers to Also troublesome are the studies that did not spec-
make informed judgments about the findings (Daft, ify how clusters were identified (e.g., Dominguez
1985). This is vital for studies using cluster analy- and Sequeira, 1993; Hitt and Tyler, 1991; Miller,
sis, however, because of the role played by 1988). At this and other steps of the clustering pro-
researcher judgment. cess, the burden of proof lies with the authors.
The selection of clustering algorithms is an issue Absent relevant ‘evidence’, the reader is unable to
for which there is consensus among methodolog- render any verdict regarding the clusters identified.
ists. Unfortunately, however, strategy research has A final concern raised by our findings is the
rarely followed the guidance that is offered. The scant attention often paid to validating cluster sol-
two-stage procedure we advocate (hierarchical and utions. Reliability was ignored in over two-thirds
non-hierarchical methods used in tandem) has been of prior studies. A particularly effective method for
recommended by many (Aldenderfer and Blash- assessing reliability was noted, however: two
field, 1984; Hartigan, 1975; Hair et al., 1992; authors independently derived clusters using dif-
Milligan, 1980; Punj and Stewart, 1983) because ferent algorithms and then assessed the conver-
both methods have weaknesses when used alone, gence between their solutions (Reger and Huff,
hence the ‘true’ structure of data sets may not be 1993). Researchers should be encouraged to devise
reflected in the solutions that either provides. Most other innovative methodologies. External validity
strategy researchers have ignored or are unaware (generalizability) was tested only twice (i.e., Ham-
of this guidance, however, only six studies used brick, 1983; Mascarenhas, 1989). Unless general-
the dual technique. To illustrate the problems cre- izability is assessed, however, there can be no con-
Application of Cluster Analysis 453

fidence that results are relevant beyond the sample multiple clustering techniques in an effort to estab-
at hand. Likewise, criterion-related validity often lish validity (e.g., Hawes and Crittenden, 1984;
was not assessed; as a result, the predictive utility Smith and Grimm, 1987). While preferable to the
of solutions frequently remains unknown. Ironi- use of a single clustering approach, the results are
cally, the number of studies that tested criterion- still vulnerable to methodological critique because
related validity through significance tests on exter- all clustering techniques share the reliance on
nal (i.e., non-clustering) variables was very close researcher judgment. Between-methods triangu-
to the number that inappropriately cited signifi- lation can occur when multiple techniques that do
cance tests on clustering variables as evidence of not share a common weakness are applied. If these
validity. Thus, in sum, researchers need to be wary methods arrive at the same set of results or support
of errors of omission and commission. the same conclusions, confidence can be high that
results are valid and not driven by one’s method
(Denzin, 1978). Given that cluster analysis’ ‘Ach-
The future of cluster analysis research
illes heel’ is its reliance on researcher judgment,
Despite the problems associated with its past use, we the logical implication is that studies using cluster
believe that cluster analysis can be valuable to future analysis should include other techniques that do
strategy research because of the technique’s unparal- not rely heavily on such interpretations. When such
leled ability to classify a large number of obser- techniques support the conclusions of cluster
vations along multiple variables. At the same time, analysis, there is assurance that findings are not
the long-running interest in identifying groups of driven by researcher judgment.
similar organizations continues to grow, as shown by Some methodological research on cluster analy-
the explosion of studies using cluster analysis since sis advocates triangulation in efforts to establish
the late 1980s and the publication of an issue on validity (though usually without using the term).
‘configurational approaches to organizational analy- Authors’ suggestions are generally limited to the
sis’ in the Academy of Management J o u m l in 1993. pursuit of within-method triangulation. For exam-
For cluster analysis to be of maximum value, how- ple, recommendations offered include the use of
ever, researchers must take steps to overcome its multiple clustering techniques (Aldenderfer and
weaknesses. The main problem is that cluster analy- Blashfield, 1984), clustering of hold-out samples
sis’ reliance on researcher judgment makes the val- (Hair et al., 1992), and supplementing cluster sol-
idity of results subject to serious doubts. We believe utions with other judgment-laden techniques, such
that the key to surmounting this problem is the vigor- as multidimensional scaling (Arabie, Carroll and
ous pursuit of triangulation; i.e., the application of DeSarbo, 1987). Perhaps not surprisingly, most tri-
multiple techniques to a single research problem. All angulation in strategy research using cluster analy-
research methods have both strengths and weak- sis has been of the within-method type.
nesses but, importantly, different methods have dif- While the guidance offered by methodologists is
ferent strengths and weaknesses. Thus, in many helpful, strategy researchers should be vigilant for
instances, the strengths of one method may comp opportunities to develop within-method triangu-
lement the strengths of another while neutralizing the lation in order to help meet the specific needs of
latter’s weaknesses (Denzin, 1978; Jick, 1979). strategy research. We believe that the inclusion of
Applying multiple techniques helps to ensure that monothetic clustering techniques in studies offers
results are not methodological artifacts; specifically, one such opportunity. Specifically, this might help
agreement between two or more methods provides to address a key limitation of most cluster analysis
evidence that results are not a product of the biases techniques (agglomerative hierarchical as well as
of one method. nonhierarchical). Regardless of the algorithm
In considering triangulation, it is important to selected, cluster analysis treats all variables as equ-
note that there are two types. Within-method tri- ally important to defining clusters. As a result,
angulation occurs when a research issue has been none of the strategy studies to date have purpose-
examined using multiple approaches, but these fully weighted variables.’ This has not been prob-
approaches share a common weakness which pre-
vents high levels of confidence in the results
’Unintentional weighting of clustering variables may occur if,
obtained (Denzin, 1978). For example, within- as noted above, there are scale differences between variables
method triangulation is apparent in studies that use or if multicollinearity is present.
454 D.J. Ketchen, Jr. and C. L. Shook
lematic because the literature has been dominated (1993) offer evidence that while managers per-
by exploratory (i.e., inductive) inquiry, where no ceive groups of firms, these groups are often based
a priori rationale for differential weights is avail- on different variables than those used by
able. As strategy research moves toward an empha- researchers. Expert opinions also can establish the
sis on theory testing (i.e., deductive studies), how- practical value of a study. This is a central concern
ever, the relative weights of strategic concepts in in strategic management because of its aim to be an
our theories must be represented empirically. Simi- applied field (e.g., Rumelt, Schendel, and Teece,
larly, when using top managers’ cognitions as the 1994). When the experts are relevant practitioners
source of clustering variables (i-e., cognitive like executives in a focal industry, their views can
studies), managers may suggest that variables are establish the ‘real world‘ value of a set of results.
not equally important; such differences must be In sum, experts’ opinions can help maximize val-
reflected in empirical efforts. By splitting samples idity and establish practical value; thus, they
on the most important variable first, then along the should be sought whenever possible.
second, and so on, monothetic techniques offer a Another tack with multiple benefits is to use
means for weighting variables such that their rela- cluster analysis in conjunction with other statistical
tive importance in clustering is better aligned with tools. This not only helps establish validity (to the
their importance in the perspective underlying the extent that the tools are objectively based) but also
derivation of clusters. Thus, the inclusion of mono- can enable the testing of powerful and sophisti-
thetic techniques as part of an array of clustering cated theoretical models. The intent behind past
methods might often be helpful in efforts to applications of cluster analysis has been to capture
develop valid cluster solutions. It is important to the complexity of organizational reality. Yet the
caution, however, that monothetic techniques’ need models examined have been relatively simple;
for dichotomous variables would usually require research generally depicts clusters as stand-alone
the manipulation of data because many variables entities or as a means of explaining outcome vari-
of interest are continuous. ables. These approaches have provided insights,
The stiffer challenge to researchers is the need but are limited in the ability to capture complexity.
to develop between-methods triangulation, where Thus, we believe that the advancement of knowl-
methodological guidance has been sparse. One edge will be hastened by viewing cluster analysis
exception is the advocation of significance testing as one possible piece of any methodological
on external variables through ANOVA or puzzle. Some research has begun to follow this
MANOVA (Anderberg, 1973; Aldenderfer and path. One example is the use of strategic groups as
Blashfield, 1984). In the remainder of this section, a context within which to examine the link between
we expand upon this initial foray into building firm-specificcharacteristics and performance (Cool
between-methods triangulation by suggesting how and Schendel, 1988; Lawless, Bergh, and Wilsted,
researchers can use techniques that do not share 1989). To take a more comprehensive view, stud-
cluster analysis’ strong reliance on researcher judg- ies might include strategic group membership as
ment to develop confidence in empirical results, one of several factors in a structural equation or
and more broadly, well-founded knowledge about path analysis model designed to predict perform-
organizations. ance. Interestingly, the earliest research using clus-
We look first to extant strategy research. A ter analysis emphasized the use of the technique
promising validation tool for inductive and deduc- alongside others (e.g., Hatten and Schendel, 1977;
tive studies that is not discussed in methodological Hatten e? al., 1978) as did early methodological
works but has been used by a few strategy authors guidance (Anderberg, 1973). We assert that such
is experts’ opinions of clusters’ validity. Like clus- designs are now a necessity.
ter analysis itself, the use of expert opinions relies To examine specifically how such designs might
heavily on perceptions. If the experts used are take shape, we consider the possible value of time
other researchers, this technique may only help series analysis to studies that include clustering.
build within-method triangulation. If the experts The dominant thmst of research using cluster
are practitioners, however, between-methods tri- analysis has been the attempt to explain the per-
angulation is possible because the perspectives and formance of organizations through their member-
assumptions of researchers and managers are gen- ship in a configuration. Studies of the link between
erally quite different. Indeed, Reger and Huff organizational configurations and performance
Application of Cluster Analysis 455

across multiple time periods generally use repeated Further, to maximize the development of knowl-
cross-sectional tests (e.g., Fiegenbaum and edge, cluster analysis should be used in combi-
Thomas, 1990). Such a design can uncover pat- nation with other methods to test sophisticated
terns within years (e.g., finding differentially per- theoretical models.
forming groups in a particular year) and across
years (e.g., the number of years that configurations
are and are not linked to performance) but is lim- ACKNOWLEDGEMENTS
ited in its ability to address other interesting ques-
tions. Specifically, time series analysis could be We would like to thank Arthur Bedeian, Kevin
used to address the cumulative effect of configur- Mossholder, Timothy Palmer, Craig Russell,
ation membership on performance over time and Charles Snow, James Thomas, and two anonymous
the extent to which the passage of time impacts the reviewers for their insights on this article.
configurations-performance relationship (Bergh,
1993). Time series analysis also can be used to
explore different lag times between the derivation
REFERENCES
of configurations and the measurement of perform- Aldenderfer, M. S. and R. K. Blashfield (1984). Cluster
ance; an important but unexamined issue (cf. Analysis. Sage, Newbury Park, CA.
Ketchen et al., 1993). In terms of methodological Anderberg, M. R. (1973). Cluster Analysis for Appli-
concerns, time series analysis offers a heightened cations. Academic Press, New York.
Andrews, K. R. (1971). The Concept of Corporate Strat-
level of protection from potential statistical prob- egy. Irwin, Homewood, IL.
lems, most notably the chance of committing Type Ansoff, H. I. (1965). Colporate Strategy. McGraw-Hill,
I errors (Bergh, 1993). Given that the use of cluster New York.
analysis introduces significant methodological dif- Arabie, P., J. D. Carroll and W. S. DeSarbo (1987).
ficulty into a study’s design, such protection is Three-way Scaling and Clustering. Sage, Newbury
Park, CA.
highly desirable. Thus, we advocate the incorpor- Bacharach, S. B. (1989). ‘Organizational theories: Some
ation of time series analysis into research on the criticism for evaluation’, Academy of Management
organizational configurations-performance Review, 14, pp. 496-515.
relationship and, more generally, to any study of Bailey, K. D. (1994). Typologies and Taxonomies: An
multiple time periods that uses cluster analysis. Introduction to Classification Techniques. Sage,
Newbury Park,CA.
Barney, J. B. and R. E. Hoskisson (1990). ‘Strategic
groups: Untested assertions and research proposals’,
CONCLUSION Managerial and Decision Economics, 11, pp. 187-
198.
Cluster analysis has been an important tool for Bergh, D. D. (1993). ‘Watch the time carefully: The use
and misuse of time effects in management research’,
examining the relationships among strategy, Journal of Management, 19, pp. 683-705.
environment, leadership/organization, and per- Birley, S. and P. Westhead (1990). ‘Growth and per-
formance; links which define the field of strategic formance contrasts between “types” of small firms’,
management (Summer et aL, 1990). In the field’s Strategic Management Journal, 11(7), pp. 535-557.
infancy, knowledge about these relationships was Breiger, R. L., S. A. Boorman and P. Arabie (1975). ‘An
algorithm for clustering relational data with appli-
obtained primarily through in-depth case studies cations to social network analysis and comparison
focusing on small numbers of firms (e.g., Chand- with multidimensional scaling’, Journal of Mathemat-
ler, 1962). However, strategic . management’s ical Psychology, 12, pp. 328-383.
dominant epistemology changed as the field Chandler, A. D. (1962). Strategy and Structure. MIT
matured. Indeed, cluster analysis’ introduction in Press, Cambridge, MA.
Cook, T. D. and D. T. Campbell (1979). Quasi-exper-
the late 1970s and frequent use in the 1980s paral- imentation: Design and Analysis Issues for Field Set-
leled an increased emphasis on developing knowl- tings. Houghton-Mifflin, Boston, MA.
edge through sophisticated statistical tests of large Cool, K. and I. Dierickx (1993). ‘Rivalry, strategic
data bases. We believe that cluster analysis can groups and firm profitability’, Strategic Management
have a place in strategic management’s methodol- Journal, 14(1), pp. 47-59.
Cool, K. and D. Schendel (1987). ‘Strategic group for-
ogical tool box in the 1990s and beyond, but the mation and performance: The case of the [Link]-
technique must be applied prudently in order to aceutical industry, 1963-1982’, Management Sci-
ensure the validity of insights that it provides. ence, 33, pp. 1102-1 124.
456 D.J. Ketchen, Jr. and C. L. Shook
Cool, K. and D. Schendel (1988). ‘Performance differ- ound strategies for mature industrial-product business
ences among strategic group members’, Strategic units’, Academy of Management Journal, 26,
Management Journal, 9(3), pp. 207-223. pp. 23 1-248.
Daft, R. L. (1985). ‘Why I recommend that your manu- Harrigan, K. R. ( 1985). ‘An application of clustering for
script be rejected and what you can do about it’. In strategic group analysis’, Strategic Management
L. L. Cummings and P. J. Frost (eds.), Publishing in Journal, 6( l), pp. 55-73.
the Organizational Sciences. Irwin, Homewood, IL, Hartigan, J. A. (1975). Clustering Algorithms. Wiley,
pp. 193-209. New York.
Denzin, N. K. (1978). The Research Act. McGraw-Hill, Hatten, K. J. (1974). ‘Strategic models in the brewing
New York. industry’, unpublished doctoral dissertation, Purdue
Dess, G. and P. Davis (1984). ‘Porter’s (1980) generic University.
strategies as determinants of strategic group member- Hatten, K. J. and M. L. Hatten (1987). ‘Strategic groups,
ship and organizational performance’, Academy of asymmetrical mobility barriers and contestability’,
Management Journal, 27, pp. 467-488. Strategic Management Journal, 8(4), pp. 329-342.
Dess, G., S . Newport and A. M. A. Rasheed (1993). Hatten, K. J. and D. E. Schendel(l977). ‘Heterogeneity
‘Configurationresearch in strategic management: Key within an industry: Firm conduct in the U.S. brewing
issues and suggestions’, Journal of Management, 19, industry’, Journal of Industrial Economics, 26,
pp. 775-795. pp. 97-1 13.
Dillon, W. R., N. Mulani and D. G. Frederick (1989). Hatten, K. J., D. E. Schendel and A. C. Cooper (1978).
‘On the use of component scores in the presence of ‘A strategic model of the U.S. brewing industry:
group structure’, Journal of Consumer Research, 16, 1952-1971’, Academy of Management Journal, 21,
pp. 106-112. pp. 592-610.
Dominguez, L. V. and C. G. Sequeira (1993). ‘Determi- Hawes, J. and W. Crittenden (1984). ‘A taxonomy of
nants of LDC exporters’ performance: A cross- competitive retailing strategies’, Strategic Manage-
national study’, Journal of lnternational Business ment Journal, 5(3), pp. 275-287.
Studies, 24, pp. 19-40. Hitt, M. A. and B. B. Tyler (1991). ‘Strategic decision
Douglas, S. P. and D. K. Rhee ( 1989). ‘Examining gen- making models: Integrating different perspectives’,
eric competitive strategy types in U.S. and European Strategic Management Journal, 12(5), pp. 327-35 1.
markets’, Journal of International Business Studies, Hofer, C. W. and D. Schendel (1978). Strategy Formu-
20, pp. 437-463. lation: Analytical Concepts. West Publishing, St.
Dutton, J. E., L. Fahey and V. K. Narayanan (1983). Paul, MN.
‘Toward understanding strategic issue diagnosis’, Hoskisson, R., M. Hitt, R. Johnson and D. Moesel
Strategic Management Journal, 4(4), pp. 307-323. (1993). ‘Construct validity of an objective (entropy)
Edelbrock, C. (1979). ‘Comparing the accuracy of hier- measure of diversification strategy’, Strategic Man-
archical clustering algorithms: The problem of classi- agement Journal, 14(3), pp. 215-235.
fying everybody’, Multivariate Behavioral Research, Hrebiniak, L. G. and W. F. Joyce (1985). ‘Organiza-
14, pp. 367-384. tional adaptation: Strategic choice and environmental
Everitt, B. (1980). Cluster Analysis (2nd ed.). Heineman determinism’, Administrative Science Quarterly, 30,
Educational Books, London. pp. 336-349.
Fiegenbaum, A. and H. Thomas (1990). ‘Strategic Hunt, M. S . (1972). ‘Competition in the major home
groups and performance: The [Link] industry, appliance industry 1960- 1970’, unpublished doctoral
1970-1984’, Strategic Management Journal, 11(3), dissertation, Harvard University.
pp. 197-215. Jardine, N. and R. Sibson (1971). Mathematical Tax-
Fombrun, C. J. and E. J. Zajac (1987). ‘Structural and onomy, Wiley, New York.
perceptual influences on intraindustry stratification’, Jick, T. D. (1979). ‘Mixing qualitative and quantitative
Academy of Management Journal, 30, pp, 33-50. methods: Triangulation in action’, Administrative Sci-
Ford, J. K., R. C. MacCallum and M. Tait (1986). ‘The ence Quarterly, 24, pp. 602-61 1.
application of exploratory factor analysis in applied Kerlinger, F. N. (1986). Foundations of Behavioral
psychology: A critical review and analysis’, Person- Research. Holt, Rinehart & Winston, Fort Worth, TX.
nel Psychology, 39, pp. 291-314. Ketchen, D. J., J. B. Thomas and C. C. Snow (1993).
Galbraith, C. and D. Schendel (1983). ‘An empirical ‘Organizational configurations and performance: A
analysis of strategy types’, Strategic Management comparison of theoretical approaches’, Academy of
Journal, 4(2), pp, 153-173. Management Journal, 36, pp. 1278- 1313.
Hair, J. F., R. E. Anderson, R. L. Tatham and W. C. Kim, L. and Y. Lim (1988). ‘Environment, generic stra-
Black (1992). Multivariate Data Analysis (3rd ed.). tegies, and performance in a rapidly developing coun-
Macmillan, New York. try: A taxonomic approach’, Academy of Management
Hambrick, D. C. (1983). ‘An empirical typology of Journal, 31, pp. 802-827.
mature industrial-product environments’, Academy of Kim, W., P. Hwang and W. Burgers (1993). ‘Multina-
Management Journal, 26, pp. 213-230. tionals’ diversification and the risk-retum trade-off’,
Hambrick, D. C. (1984). ‘Taxonomic approaches to Strategic Management Journal, 14(4), pp. 275-286.
studying strategy: Some conceptual and methodolog- Lafuente, A., and V. Salas (1989). ‘Types of entrepreneurs
ical issues’, Journal of Management, 10, pp. 27-41. and firms: The case of new Spanish firms’, Strategic
Hambrick, D. C. and S . M. Schecter (1983). ‘Turnar- Manaeement
” Journal. lo(.1). DD. 17-30.
I
Application of Cluster Analysis 457
Lawless, M. and L. Finch (1989). ‘Choice and determin- figuration’. In G. Morgan (ed.), Beyond Method.
ism: A test of Hrebiniak and Joyce’s framework on Sage, Beverly Hills, CA, pp. 57-73.
strategy-environment fit’, Strategic Management Milligan, G. W. (1980). ‘An examination of the effect
Journal, 10(4), pp. 35 1-365. of six types of error perturbation on fifteen clustering
Lawless, M., D. D. Bergh and W. Wilsted (1989). ‘Per- algorithms’, Psychometrika, 45, pp. 325-342.
formance variations among strategic group members: Milligan, G. W. and M. C. Cooper (1985). ‘An examin-
An examination of individual firm capability’, Jour- ation of procedures for determining the number of
nal of Management, 15, pp. 649-661. clusters in a data set’, Psychometrika, 50, pp. 159-
Lewis, P. and H. Thomas (1990). The linkage between 179.
strategy, strategic groups and performance in the U.K. Mintzberg, H. (1989). Mintzberg on Management. Free
retail grocery industry’, Strategic Management Jour- Press, New York.
nal, 11(5), pp. 385-397. Montgomery, C. A., B. Wernerfelt and S. Balakrishnan
Lorr, M. (1983). Cluster Analysis for the Social Sci- (1989). ‘Strategy content and the research process:
ences. 3ossey-Bass, San Francisco, CA. A critique and commentary’, Strategic Management
Macmillan, I. C. (1991). ‘The emerging forum for busi- Journal, 10(2), pp. 189-197.
ness policy scholars’, Strategic Management Journal, Morrison, A. and K. Roth (1992). ‘A taxonomy of busi-
12(2), pp. 161-165. ness-level strategies in global industries’, Strategic
Manu, F. A. (1992). ‘Innovation orientation, environ- Management Journal, 13(6), pp. 399-41 8.
ment and performance: A comparison of U.S. and Momson, A. and K. Roth (1993). ‘Relating Porter’s
European markets’, Journal of International Business configuration/coordinationframework to competitive
Studies, 23, pp. 333-359. strategy and structural mechanisms: Analysis and
Mascarenhas, B. (1986). ‘International strategies of non- implications’, Journal of Management, 19, pp. 797-
dominant firms’, Journal of International Business 818.
Studies, 17, pp. 1-25. Nohria, N. and C. Garcia-Pont (1991 ). ‘Global strategic
Mascarenhas, B. (1989). ‘Strategic group dynamics’, linkages and industry structure’, Strategic Manage-
Academy of Management Journal, 32, pp. 333-352. ment Journal, Summer Special Issue, 12, pp. 105-
Mascarenhas, B. and D. Aaker (1989a). ‘Strategy over 124.
the business cycle’, Strategic Management Journal, Porac, J. F. and H. Thomas (1990). ‘Taxonomic mental
10(3), pp. 199-210. models in competitor definition’, Academy of Man-
Mascarenhas, B. and D. Aaker (1989b). ‘Mobility bar- agement Review, 15, pp. 224-240.
riers and strategic groups’, Strategic Management Porter, M. E. (1973). ‘Consumer behavior, retailer
Journal, 10(5), pp. 475-485. power, and manufacturer strategy in consumer goods
McDougall, P. and R. Robinson (1990). ‘New venture industries’, unpublished doctoral dissertation, Har-
strategies: An empirical identification of eight “arche- vard University.
types” of competitive strategies for entry’, Strategic Punj, G. and D. W. Stewart (1983). ‘Cluster analysis in
Management Journal, 11(6), pp. 447-467. marketing research: Review and suggestions for
McGrath, J. E., J. Martin and R. A. Kulka (1982). Judg- application’, Journal of Marketing Research, 20,
ment Calls in Research. Sage, Beverly Hills, CA. pp. 134-148.
McKelvey, B. (1975). ‘Guidelines for empirical classi- Rajagopalan, N. and S. Finkelstein (1992). ‘Effects of
fication of organizations’, Administrative Science strategic orientation and environmental change on
Quarterly, 20, pp. 509-525. senior management reward systems’, Strategic Mun-
McKelvey, B. ( 1978). ‘Organizational systematics: agement Journal, Summer Special Issue, 13,
Taxonomic lessons from biology’, Management Sci- pp. 127-142.
ence, 24, pp. 1428-1440. Reger, R. K. and A. S. Huff (1993). ‘Strategic groups:
Meyer, A. D. (1991). ‘What is strategy’s distinctive A cognitive perspective’, strategic Management
competence?’, Journal of Management, 17, Journal, 14(2), pp. 103-124.
pp. 821-833. Robinson, R. and J. Pearce (1988). ‘Planned patterns of
Meyer, A. D., A. S. Tsui and C. R. Hinings (1993). strategic behavior and their relationship to business-
‘Configurational approaches to organizational analy- unit performance’, Strategic Management Journal,
sis, Academy of Management Journal, 36, 9( l ) , pp. 43-60.
pp. 1175-1195. Roth, K. (1992). ‘International configuration and coordi-
Miles, G., C. C. Snow and M. P. Sharfman (1993). nation archetypes for medium-sized firms in global
‘Industry variety and performance’, Strategic Man- industries’, Journal of International Business Studies,
agement Journal, 14(3), pp. 163-177. 23, pp. 533-549.
Miles, R. E. and C. C. Snow (1978). Organizational Rumelt, R. P., D. E. Schendel and D. J. Teece (1994).
Strategy, Structure, and Process. McGraw-Hill, Fundamental Issues in Strategy: A Research Agenda.
New York. Harvard Business School Press, Boston, MA.
Miller, A. (1988). ‘A taxonomy of technological set- SAS Institute, Inc. (1990). SAS Users Guide: Statistics,
tings, with related strategies and performance levels’, Version 6. SAS Institute, Cary, NC.
Strategic Management Journal, 9(3), pp. 239-254. Smith, K. and C. Grimm (1987). ‘Environmental vari-
Miller, D. and P. Friesen (1978). ‘Archetypes of strategy ation, strategic change and firm performance: A study
formulation’, Management Science, 24, pp. 92 1-933. of railroad deregulation’, Strategic Management
Miller, D. and H. Mintzberg (1983). ‘The case for con- Journal, 8(4), pp. 363-376.
458 D. J. Ketchen, Jr. and C. L. Shook
Summer, C. E.,R. A. Bettis, I. H. Duhaime, J. H. Grant, Journal of International Business Studies
D. C. Hambrick, C. C. Snow and C. P. Zeithaml
(1990). ‘Doctoral education in the field of business
policy and strategy’, Journal of Management, 16, Mascarenhas ( 1986)
pp. 361-398. Douglas and Rhee (1989)
Thomas,H. and N. Venkatraman (1988). ‘Research on Manu (1992)
strategic groups: Progress and prognosis’, Journal of Roth (1992)
Management Studies, 25, pp. 537-555. Dominguez and Sequeira (1993)
Thomas, J. B., S. M. Clark and D. A. Gioia (1993).
‘Strategic sensemaking and organizational perform-
ance: Linkages among scanning, interpretation, Journal of Management
action, and outcomes’, Academy of Management
Journal, 36, pp. 239-270. Lawless, Bergh, and Wilsted (1989)
Thompson, J. D. (1967). Organizations in Action. Waddock and Isabella (1989)
McGraw-Hill, New York.
Venkatraman, N. and J. E. Prescott (1990). Momson and Roth (1993)
‘Environment-strategy coalignment: An empirical
test of its performance implications’, Strategic Man- Management Science
agement Journal, 11( l), pp. 1-23.
Venkatraman, N. and V. Ramanujam (1986). ‘Measure- Cool and Schendel (1987)
ment of business performance in strategy research’,
Academy of Management Review, 11, pp. 801-814.
Waddock, S. and L. Isabella (1989). ‘Strategy, beliefs Strategic Management Journal
about the environment and performance in a banking
simulation’, Journal of Management, 15, pp. 617- Woo and Cooper (1981)
632. Galbraith and SchendeI (1983)
Walter, G. and J. Barney (1990). ‘Management objec-
tives in mergers and acquisitions’, Strategic Manage- Hawes and Crittenden ( 1984)
ment J o u m l , 11(l), pp. 79-86. Smith and Grimm (1987)
Whallon, R. W. (1972). ‘A new approach to pottery Cool and Schendel (1988)
typlogy’, American Antiquity, 37,pp. 13-33. Miller ( 1988)
Woo, C. Y. Y. and A. C. Cooper ( 1981). ‘Strategies of Robinson and Pearce ( 1988)
effective low share businesses’, Strategic Manage-
ment Journal,2(3), pp. 301-318. Lafuente and Salas (1989)
Zahra, S. and J. Covin (1993). ‘Business strategy, tech- Lawless and Finch (1989)
nology policy, and firm performance’, Strategic Man- Mascarenhas and Aaker (1989a)
agement Journal, 14(6), pp. 451-478. Mascarenhas and Aaker (1989b)
Birley and Westhead ( 1990)
Fiegenbaum and Thomas (1990)
APPENDIX: Strategy Articles That Use Lewis and Thomas (1990)
Cluster Analysis McDougall and Robinson ( 1990)
Venkatraman and Prescott (1990)
Academy of Management Journal Walter and Barney (1990)
Hitt and Tyler ( 1991)
Hatten, Schendel, and Cooper (1978) Nohria and Garcia-Pont ( 1991)
Hambrick (1983) Momson and Roth (1992)
Hambrick and Schecter (1983) Rajagopalan and Finkelstein (1992)
Dess and Davis (1984) Cool and Dierickx (1993)
Fombrun and Zajac ( 1987) Hoskisson, Hitt, Johnson and Moesel ( 1993)
Kim and Lim (1988) Kim, Hwang, and Burgers (1993)
Mascarenhas ( 1989) Miles, Snow, and Sharfman (1993)
Ketchen, Thomas and Snow (1993) Reger and Huff ( 1993)
Zahra and Covin (1993)

You might also like