Conjoint Analysis: A Practical Guide
Conjoint Analysis: A Practical Guide
Introduction
Conjoint analysis is a marketing research technique designed to help managers
determine the preferences of customers and potential customers. In particular, it
seeks to determine how consumers value the different attributes that make up a prod-
uct and the trade-offs they are willing to make among the different attributes or fea-
tures that compose the product. As such, conjoint analysis is best suited for products
that have very tangible attributes that can be easily described or quantified.
Although the history of conjoint analysis can be traced to early work in math-
ematical psychology,1 its popularity has grown tremendously over the last few years
as access to easy-to-use software has allowed its widespread implementation. There
have been probably hundreds of applications of conjoint analysis in industrial set-
tings.2 Some of the more important issues for which modern conjoint analysis is used
are the following:
• Predicting the market share of a proposed new product, given the current offer-
ings of competitors
• Predicting the impact of a new competitive product on the market share of any
given product in the marketplace
• Determining consumers’ willingness to pay for a proposed new product
• Quantifying the trade-offs customers or potential customers are willing to make
among the various attributes or features that are under consideration in the new
product design
55
56 CUTTING-EDGE MARKETING ANALYTICS
Continuing with the car example, an experimental design might look like the
information presented in Table 5-1.
This is a simple design that contains a total of 15 attribute levels. Real designs
often contain more attributes and levels than are presented here.
When constructing an experimental design, it is important to keep the following
points in mind:
• The more tangible and understandable the levels of each attribute are to the
respondents, the more valid the results of the research will be. For example,
attribute levels such as really roomy are vague, meaning different things to dif-
ferent people, and should be avoided.
• The greater the number of attribute levels to be tested, the more data that will
be needed to achieve the same degree of output accuracy.
• For quantitative variables (price and horsepower, in this example), the greater
the distance between any two consecutive levels, the harder it will be to get a
good idea of how a consumer might evaluate something in between the two (for
example, $24,000).
Data Collection
Collecting data for a conjoint analysis has been made relatively simple by the
advent of dedicated off-the-shelf software. The type of conjoint analysis used dic-
tates the exact nature of the data collected. An exhaustive discussion of the benefits
and drawbacks of each of the many different types of conjoint analysis now in use is
beyond the scope of this book. But those interested are encouraged to read Orme’s
technical paper for a good discussion of this topic.3
58 CUTTING-EDGE MARKETING ANALYTICS
The state of the art in conjoint data collection involves using personal comput-
ers or a web-based version of the software to guide respondents through an interac-
tive conjoint survey. The software creates the hypothetical product profiles using the
experimental design provided by the researcher and estimates the attribute-level utili-
ties from participant ratings or choices.
utilities are generally scaled in such a way that they add up to zero. So, a negative
number does not mean that a given level has negative utility; it just means that this
level is on average less preferred than a level with an estimated utility that is positive.
Conjoint analysis output is also often accompanied by t-values, a standard metric
for evaluating statistical significance. Because of the way conjoint utilities are scaled,
the standard interpretation of t-values can yield misleading results. For example, the
level Saturn of the attribute Brand has a t-value of 0.87. In general, a t-value of this
magnitude would fail a test of statistical significance; however, this t-value is gener-
ated because within the attribute Brand, the level Saturn has neither a very high nor
very low relative preference. It is basically in the middle in terms of overall prefer-
ence. Because of the scaling, levels that have more moderate levels of preference
within a given attribute are likely to have estimated utilities close to zero, which tends
to produce very low t-values (recall that the t-test is measuring the probability that the
true value of a parameter is not different from zero).
A better way to think about statistical significance in this context is to examine the
t-values of the levels with the highest and lowest preference within a given attribute.
An applicable common practice would be if the sum of the absolute value of these
two statistics is greater than three, then that given attribute is significant in the overall
choice process of consumers. At a practical level, it is rare that an attribute will not
be significant, and, if you find one that is, it means it probably should not have been
included in the experimental design in the first place because respondents are not
considering that attribute’s information when they make choices.
Trade-Off Analysis
The utility of any given product that you might consider can be easily computed
by simply summing the utilities of its attribute levels. For example, a Toyota with 280
horsepower, leather interior, no sunroof, and a price of $23,000 has a utility of 0.75 +
1.18 + 1.60 – 0.68 + 2.10 = 4.95. If the car with the same basic specifications were a
Volkswagen, the overall utility would drop to 0.65 + 1.18 + 1.60 – 0.68 + 2.10 = 4.85,
60 CUTTING-EDGE MARKETING ANALYTICS
a drop of 0.10. This drop can be seen directly by noticing that the difference between
the utility for the brand Toyota (0.75) and Volkswagen (0.65) is 0.10. In addition,
because nothing else in the profile of the car has changed, this will be the exact utility
difference between two cars that are the same except for this brand difference.
A natural consequence of this observation is that you use the utilities to analyze
what average consumers would be willing to give up on one particular attribute to
gain improvements in another. For example, how much money would they be will-
ing to give up (price) if a sunroof was added to the vehicle? Let’s look directly at this
issue of the hypothetical car detailed in the previous paragraph. Adding a sunroof to
the (Toyota) car would yield an overall utility of 0.75 + 1.18 + 1.60 + 0.68 + 2.10 =
6.31. This represents an increase in utility of 6.31 – 4.95 = 1.36 over the identical car
without a sunroof.
This information directly implies that you can reduce the utility of price by 1.36,
and average consumers would be just as happy as before the sunroof was installed. To
find out how much the price can be raised, you must convert the change in utility with
a change in price. You do this by first noting how much the original car costs ($23,000)
and the utility associated with that figure, 2.10. You know that you can reduce the
price utility by 1.36. This is equivalent to saying that you can reduce the price utility to
2.10 – 1.36 = 0.74. By referring to Table 5-2, you can immediately see that this implies
a price between $25,000 and $27,000 because –1.56 < 0.74 < 1.15. In fact, if you
assume a linear relationship between price and utility in the range between $25,000
and $27,000, you can solve for the exact price by performing a linear interpolation
within this range.4 Specifically, the interpolation yields:
1.15⫺0.74
$25,000 ⫹ ⫻ $2,000 ⫽$25,302.58
1.15⫺(⫺1.56)
This implies that, if the sunroof is added, the price of the vehicle could be raised
from $23,000 to about $25,300, and the average consumer’s attitude would be one of
indifference between the two vehicles. Qualitatively, it shows that the value of a sun-
roof to consumers is very substantial.
This same kind of analysis can be performed for other attributes. You could ask
how much additional horsepower you would need to add if the interior was changed
CHAPTER 5 • A PRACTICAL GUIDE TO CONJOINT ANALYSIS 61
from leather to cloth. This particular question does present a problem, however.
Because the current vehicle under consideration has 280 horsepower, and that is the
maximum amount of horsepower tested by the conjoint analysis, it will be impossible
to determine how much consumers will value additional horsepower. This leads to
an important consideration when constructing the experimental design. That is, if
the output is to be used for trade-off analysis, it is important that the range of the
levels tested within each attribute span the entire range of that attribute before man-
agement would ever consider it as a realistic design alternative. If the experimental
design takes this into account, you can perform trade-off analysis between any two
attributes in the design.
• The company must know the other products, besides its own offering, that a
consumer is likely to consider when making a selection in the category.
• Each of these competitive products’ important features must be included in the
experimental design. In other words, you must be able to calculate the utility of
not only your own product offering, but also that of the competitive products.
Market share prediction relies on the use of a multinomial logit model.5 The basic
form of the logit model is
eUi
Sharei =
∑
n Uj
j =1
e
where
Ui is the estimated utility of product i,
Uj is the estimated utility of product j, and
n is the total number of products in the competitive set, including product i.
To make things clear, consider the following example. Suppose you are interested
in predicting the market share of a car with the following profile: Saturn; $23,000;
220HP; cloth interior; no sunroof. You believe that when consumers consider your
car, they will also consider purchasing cars that are currently on the market with the
following profiles:
62 CUTTING-EDGE MARKETING ANALYTICS
For the Saturn and its associated product profile, the estimated utility is 2.10 –
0.13 – 2.24 – 1.60 – 0.68 = –2.55. Similarly, the utilities of the three competing prod-
ucts can be calculated:
With these utilities in hand, you can now directly apply the logit model to forecast
market share for the Saturn. This is given by the following:
e −2.55
ShareSaturn = = 0.025 or 2.5%
e −2.55 + e −2.03 + e1.06 + e −3.69
This implies that this particular Saturn vehicle will achieve a 2.5% market share
within the specified competitive set. The market share of any vehicle that can be
described by the experimental design and a set of competitive vehicles, also described
by the experimental design, can be found in a similar manner.
Ui − U i
Ii =
∑
n
j =1
Uj −U j
where:
Ii is the importance of any given attribute i,
U is the highest utility level within a given attribute (subscripts indicate which
attribute), and
U is the lowest utility level within a given attribute.
This equation is really quite intuitive. To calculate the importance of any given
attribute, you just take the difference between the highest and lowest utility level of
that attribute and divide this by the sum of the differences between the highest and
lowest utility level for all attributes (including the one in question). The resulting
number will always lie between zero and one and is generally interpreted as the per-
cent decision weight of an attribute in the overall choice process.
It also should be clear at this point that this estimated attribute importance
depends critically on your experimental design. In particular, if you increase the dis-
tance between the most extreme levels of any given attribute, you will almost certainly
increase the overall attribute importance. For example, if the tested price range was
$21,000–$31,000 instead of $23,000–$29,000 (Table 5-1), this is very likely to increase
the estimate attribute importance of price.
Let’s now consider a concrete example using the attribute Horsepower. The
importance of this attribute is calculated as follows:
1.18 + 2.24
I Horsepower = = 0.25
( (2.10 + 1.69) + (0.75 + 1.27) + (1.18 + 2.24 ) + (1.60 + 1.60) + (0.68 + 0.68) )
In the example, 25% of the overall decision weight is assigned to Horsepower.
You may verify through analogous calculations that the decision weight for Price is
about 27%; Brand, about 15%; Sunroof, about 10%; and Upholstery, about 23%. The
numbers provide a very intuitive metric for thinking about the importance of each
attribute in the decision process.
64 CUTTING-EDGE MARKETING ANALYTICS
Conclusion
Conjoint analysis has a broad array of possible applications. Many of these appli-
cations are variants of the three very common applications presented here. The
increasingly widespread availability of conjoint analysis software—both PC and web-
enabled—points to its continued growth as a marketing decision aid.
This chapter has presented what is generally known as aggregate-level conjoint
analysis. That is, all of the respondents are pooled into one group, and a single set of
attribute-level utilities are estimated from the ratings or choices provided by the peo-
ple in this group. Recent advancements in conjoint analysis have enabled researchers
to estimate different utilities for different groups of respondents and even, in some
cases, for individual respondents. Although the mathematics necessary for this pro-
cedure is sometimes quite complex, it is now possible to estimate the attribute-level
utilities and to compute trade-off analyses for each individual respondent. This has
some significant advantages over aggregate-level analysis, particularly when consider-
ing marketing segmentation issues. Either way, the data collection and the basic inter-
pretation of the output remain the same. Although there is currently no textbook that
can provide answers to all the questions that might arise when applying this technique
in a business setting, there is, as of this writing, a very good and surprisingly compre-
hensive collection of technical papers located on the site of a company that markets
conjoint analysis software ([Link] These
provide answers to many of the practical implementation questions a user may face.
Endnotes
1. R. Duncan Luce and John W. Tukey, “Simultaneous Conjoint Measurement: A New Type of Fun-
damental Measurement,” Journal of Mathematical Psychology 1 (February 1964): 1–27.
2. Paul E. Green, Abba M. Krieger, and Yoram Wind, “Thirty Years of Conjoint Analysis: Reflections
and Prospects,” Interfaces 31 (May–June 2001): 56–73.
3. Bryan K. Orme, “Which Conjoint Method Should I Use?,” Sawtooth Software Technical Paper
(2003). A copy of this paper is available at [Link]
[Link] (accessed April 3, 2012).
4. This is a common way to approximate the relationship between the value of the attribute and its
utility for attribute values that were not directly tested by the conjoint analysis. The closer the tested
levels are to each other, the more accurate this approximation. Also notice that this interpolation
can only be performed for quantitative attributes such as price. Interpolating between qualitative
attributes, such as brand, is nonsensical.
5. Please refer to Chapter 13, “Logistic Regression,” for the basics of the logit model. Also, most econo-
metrics textbooks will have information on logit models.
7
Multiple Regression in Marketing-Mix Models
Introduction
The movie Moneyball has a lot to teach us about optimizing a company’s market-
ing mix. In the movie, the management of the Oakland Athletics discovers that the
baseball team can get ahead by changing its perspective and looking at data differently
than its competitors.
The A’s know most major-league teams use batting average (hits over real oppor-
tunities) as the prevailing metric for determining the worth of a hitter. Traditional
wisdom says, “You hit more, you win more.” So the players who have more hits per at
bat1 are generally the most sought after and are paid the most money. But by examin-
ing the outcome of decades of baseball games, the A’s find a variable they believe to
be more predictive of success. It is not only hits that help a baseball team win; walks
count too. Getting on base and not making outs is more closely correlated with win-
ning games than hits alone.
The team’s management takes the analysis of the data and uses it to buy underval-
ued players—players who don’t necessarily have the highest batting averages but who
do have high on-base percentages. For a small-market team such as the A’s, which has
less money to spend on players than other franchises, this strategy changes the game.
Moneyball is about baseball, but the idea also works in the context of business
marketing. Although management often makes assumptions, by actually analyzing the
data, a business can better understand how to succeed. And if a business can find an
important variable before others begin using it, management can build its strategy
around that variable to gain an advantage.
78
CHAPTER 7 • MULTIPLE REGRESSION IN MARKETING-MIX MODELS 79
25
y = 1.42x + 9.90
20
15
$ Spent by
Customer
10
0
0 1 2 3 4 5 6 7 8 9
Number of Promotions
In this example, No More Germs has data covering a time period of 29 weeks,
promotions ranging between 0 and 9, and corresponding sales from 10 to 23. A linear,
single-variable regression analysis can be run on this data with the aid of computer
software.2 The results will help No More Germs examine the relationship between
the number of promotions and the number of sales to customers by producing a func-
tion that describes the relationship. The objective is to draw a line that at each point
80 CUTTING-EDGE MARKETING ANALYTICS
represents the number of sales that are likely for any given number of promotions.
In this case, the x variable—or independent variable—is the number of promotions.
The y variable—or dependent variable (known as such because it depends on x)—is
units sold.
The function produced by the regression is intended to cover as many of the
known data points as possible and/or reduce the distance between the line and the
points as much as possible. This allows the data analyst to accurately predict sales that
are likely, given the number of promotions in other sample sets of data (in this case,
if data from other weeks is used). The equation from the regression analysis for the
best-fit straight line for No More Germs is y = 1.42x + 9.9.
The most critical outputs of the regression for the marketing manager are two
coefficients: the intercept (9.9) and the slope (1.42) of the line. The intercept rep-
resents the number of sales that are likely when promotions are 0, which is equal to
9.9 in this example. This is the point where the line crosses the y-axis. The slope of
the line describes the relationship between sales (y, or the dependent variable) and
promotions (x, or the independent variable) by stating the ratio of the change in y to
a unit change in x. In the example, the number of sales changes 1.42 per one-unit
increase in promotions (Figure 7-2). The slope (often referred to as “rise over run”)
is, therefore, 1.42 ÷ 1, or 1.42.
25
y = 1.42x + 9.90
20
1.42
15
$ Spent by
Customer
10
0
0 1 2 3 4 5 6 7 8 9
Number of Promotions
Three things can be determined immediately by looking at the slope of the line:
(1) If the number is positive, the relationship between the two variables is positive,
meaning as the independent variable increases, so does the dependent variable; (2)
if the slope of the line is 0, no changes are observed in the dependent variable as the
independent variable changes (in other words, the variables are not correlated); and
(3) if the slope of the line is negative, a change in the independent variable will pro-
duce the opposite effect in the dependent variable (in this case, No More Germs’ sales
would decrease if promotions increased).
Remember that although in this case the relationship between promotions and
sales is obvious, in most cases a regression analysis is used to show a relationship
between variables that are not as clearly related. For example, what if No More Germs
wanted to know what kind of effect web advertising had on sales of its products? The
company’s marketing manager might not know how effective web ads are compared
with print ads, for example, and the regression would assist him or her in deciding
where to put the company’s advertising dollars.
The output of No More Germs’ sample regression (which is typical of these
reports) is shown in Table 7-1. Although the analysis yields multiple statistics, the
most critical for marketing analysts (in addition to the coefficients of the equation) are
r squared and p-value. In this example, r squared is 60%, meaning the line described
by this function is appropriate for explaining 60% of the data points. This indicates
how accurate the function is within the current sample of data. (Note: A typical
marketing-focused regression would have an r squared of about 20% to 30%, as there
are numerous factors that affect sales—such as competition, weather, and so on—that
would be unknown before running the analysis.)
ANOVA
df SS MS F Sig F
Regression 1 267.28 267.28 40.60 0.00
Residual 27 177.75 6.58
Total 28 445.03
82 CUTTING-EDGE MARKETING ANALYTICS
Illustration of R Squared = 0%
1
0.9
0.8
Predicted Units Purchased
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
0 1 1 2 2 3 3 4 4 5
A line cannot be produced that will explain any of the data. Now imagine r squared
is 100%. In this case, all of the data points (dots) will be on a line (Figure 7-4). The line
accounts for all of the points in the data set. All regression analyses will result in lines
with accuracy somewhere in between these extremes.
CHAPTER 7 • MULTIPLE REGRESSION IN MARKETING-MIX MODELS 83
3.5
Predicted Units Purchased
2.5
1.5
0.5
0
0 1 1 2 2 3 3 4 4 5
P-value describes the significance of the findings given the sample size. But what
does significant mean? In this population sample, 29 observations are made. Because
this is a regression analysis of a small sample, you want to know whether you will still
see the resulting coefficients if you include another 29 observations or another 29,000
observations. Will the slope of the line be 1.42, or will it be 0 or negative? Here, the
p-value indicates there is a 0% chance the coefficients will change beyond the stan-
dard error given the addition of more data points or different samples. Most impor-
tant, it indicates a 0% chance the slope will become negative, indicating the opposite
relationship between the variables than what is indicated by the regression. In other
words, regardless of how many times the data is sampled, the relationship will hold.
In addition to these critical outputs of a regression analysis, it might be beneficial
to be familiar with one other value. In this example, t-stat is a reflection of p-value;
however, depending on the regression or model used, the name of this value may
change (for example, chi-square). P-value, on the other hand, will always be referred
to in the same way. Particularly for marketing managers, who in most cases will need
to be smart consumers of regression outputs but will not have to run the analyses
themselves, p-value will provide adequate information about the significance of the
findings.
84 CUTTING-EDGE MARKETING ANALYTICS
oversimplified and, at least in terms of marketing, fail to fully explain most real-world
situations).
Price Bias
Direction of Bias in Feature = [Sign of Units –ve (negative)
Correlation Between Price and Units]
Feature –ve (negative) +ve (positive)
× [Sign of Correlation Between Price
and Feature] Display –ve (negative) +ve (positive)
CHAPTER 7 • MULTIPLE REGRESSION IN MARKETING-MIX MODELS 87
Another way to think about omitted variables is shown in Figure 7-5. Here, x and
y are shown with respect to some omitted variable, z. By examining this relationship,
the marketing manager can determine the direction of the bias created by omitting
that variable.
Z
(Price)
⫺ ⫺
Y X
(Units Sold) (Feature / Display)
⫹
Note that an omitted variable is only a problem when it affects both whatever is
included in the model and the dependent variable. If it is not correlated with other
independent variables in the model, removing it will reduce r squared, but it will not
affect the coefficient of the variables included in the model. In the current example,
the variation in units is being assigned to feature and display, when in fact it should
be assigned to price. If the changes in one variable did not affect another, whatever
variation in the dependent variable was being captured would still reflect reality. For
example, weather can have a profound effect on sales (for example, a hurricane keep-
ing buyers in Florida from making it to stores for an extended period), but hurricanes
need not necessarily affect the feature or display plans for a brand. If weather and
feature and display plans are not correlated, then inclusion of weather is not necessary
to obtain accurate estimates of feature or display.
In this example, you have what is known as an optimistic model, which can be
a concern for marketing managers. When presenting such results to decision mak-
ers, the findings will be overstated because a significant variable (price) was omitted.
Although you cannot include everything in your model, knowing whether the results
are conservative or optimistic is beneficial. Typically, a conservative model (one that
has a negative bias) is best. Investing in a marketing channel shown to be effective by
a conservative model may still represent lost opportunity if the amount of the invest-
ment is low, but it will not represent an outright mistake in resource allocation.
When do you know if you have the true model? You never know, but examining
the four Ps (product, price, place, and promotion) is a good place to start. The results
of a regression analysis are only hypotheses, and they should be tested in field experi-
ments to ensure their validity.
88 CUTTING-EDGE MARKETING ANALYTICS
In this example, the company will make $6.60 per promotion. But if the cost of
the promotion increases or the company makes less gross profit per unit, the eco-
nomic significance of the promotion could quickly be lost. In other words, even if
your regression findings are significant, you must first use a profit/loss function before
taking action.
Conclusion
A regression analysis is intended to help marketing managers understand the rela-
tionship between two or more variables or concepts. Typically, a company will use
historic sales data or data generated through experiments to identify factors that most
affect a brand’s sales.
The value of a regression model is only as good as the variables selected to be in
the model. Strong managerial intuition is required to identify variables (such as price,
feature, and display, among others) that are most closely related to sales. For the
CHAPTER 7 • MULTIPLE REGRESSION IN MARKETING-MIX MODELS 89
best results, managers should also have some insight into how these variables actu-
ally relate in the real world to determine whether the results of a regression might be
conservative or overly optimistic. This intuition is the artistic or creative side of analyt-
ics and is necessary to move a regression beyond a statistical exercise and turn it into
something valuable for a business.
Endnotes
1. Webster’s Third New International Dictionary, Unabridged, defines “at bat” as “an official turn at
batting charged to a baseball player except when the player walks, sacrifices, is hit by a pitched ball,
or is interfered with by the catcher.”
2. For more information on how to perform a regression using computer software, please visit Darden
Marketing Analytics at [Link]