Advanced Marketing Research Methods
Advanced Marketing Research Methods
Marketing research is the function that links the consumer, customer, and public to
the marketer through information – information used to
identify and define marketing opportunities and problems,
generate, refine, and evaluate marketing actions,
monitor marketing performance, and
improve understanding of marketing as a process
Marketing research
specifies the information required to address these issues,
designs the method for collecting information,
manages and implements the data collection process,
analyzes, and
communicates the findings and their implications
In short: marketing research is a set of formal practices for doing what people do
all the time – gathering information and using it to make better decisions
How might the following organizations effectively employ marketing research? Be specific!
A small sporting goods store
Städel Museum
International School of Management
Rhein-Main-Verkehrsverbund
MyZeil shopping mall
A public campaign against cigarette smoking
Eintracht Frankfurt soccer team
Member of the Deutsche Bundestag representing the electoral district Frankfurt am Main II
Design sample
Collect data
RarelyDetail
does research objectives
an initial request for and
helpinformation
adequatelyneedsestablish the need for research
information. Managers often react to hunches and symptoms rather than to clearly
Set decision
identified researchsituations.
design andYetdata
thesources
need for research information must be precisely
defined if the research project is to provide information pertinent to the necessary
Design
decision. Too data
often,collection procedure
the importance of this initial step is overlooked in the excitement
of undertaking a research project, resulting in research findings that do not
Design
adequately sample
inform decisions.
Collect data
Symptoms are performance measures, metrics, and diagnostics that signal the
presence of a problem / an opportunity
− E.g., market share below expectation, increase in sales of snack products with
less fat content
Symptoms themselves rarely contain information regarding their causes
Therefore, the decision-maker‘s task is to respond to symptoms by analyzing the
underlying problems / opportunities to determine whether the situation calls for a
decision
The process of identifying problems / opportunities involves analyzing past,
present, and possible future situations facing an organization to uncover those
variables that either
− cause poor performance or
− represent opportunities for further growth
It is essential that the researcher understands the problem situation from the
perspective of the decision maker
Consequently, the first task is to identify the decision maker(s)
− Often, the person who first requests assistance from marketing research is not
the decision maker and may or may not know how the decision maker views the
specifics of the decision situation
In a second step, the researcher needs to translate goals likely lacking operational
significance (e.g., “we need to hold off the competition”) into focused objectives
Decision objectives include
− organizational objectives (e.g., increasing earnings per share by 10 percent next
year) and
− personal objectives of the decision maker(s) and those influencing those
individuals (e.g., getting promoted, acquiring more prestige)
Setgeneral
After the research design
need and data sources
for information is clearly established, researchers must list,
specifically and in detail, the objectives and information needs of the proposed
Design
research. data collection
Research objectivesprocedure
answer the question, “why is this project being
conducted?” and are typically put in writing before the project is undertaken.
Design
Information sample
needs answer the question, “what specific information is required to
attain the objectives?”.
Collect data
Research objectives answer the question, „what is the purpose of the research
project?“
Research objectives serve to guide the research project by
− giving direction to the specific information to be gathered (information need),
which in turn
− guides the specific questions developed for the questionnaire
Each question on the questionnaire should have a direct correspondence to an
information need, and each information need should have a direct
correspondence to at least one research objective
− Otherwise, unneeded data will be collected
When developing the information needs, both the manager and the researcher
should ask, for each item: realistically, can this information be obtained?
Research objectives can be stated so broadly that they fail to communicate the specifics of why
the study is being conducted. Evaluate the following statement: To study consumer reactions to
cartoon characters in advertising. Propose a more precise and useful formulation.
Do we cover all information Do we know what we are going Does the research project
needs properly? to do given a specific outcome? economically make sense?
In each of the following situations, identify the fundamental source of the marketing problem or
opportunity, a decision objective arising from the marketing problem or opportunity, and a possible
research objective .
Cool Pool Supply is a manufacturer of swimming pool maintenance chemicals. Recently, a
malfunction of the equipment that mixes anti-algae compound resulted in a batch of the product
that not only inhibits algae growth but also causes the pool water to turn a beautiful shade of
light blue.
The MBA director of a local college recently extended the offers to 20 promising students. Only
five offers were accepted. In the past, acceptance rates have averaged 90%. A survey of non-
acceptors conducted by the director revealed that the primary reason students declined the offer
was their perception that the college’s course requirements are too “restrictive”.
Chocoholic Candy Company has enjoyed great success in its small regional market.
Management attributes much of this success to Chocoholic’s unique distribution system, which
ensures twice-weekly delivery of fresh product to retail outlets. The directors of the company
have instructed management to expand Chocoholic’s geographical market if it can be done
without altering the twice-weekly delivery policy.
Research objectives and hypothesis typical for the three categories of research
There are three types of marketing research: exploratory, descriptive, and causal reseach. Indicate
which type each item in the list below illustrates. Explain your answers.
Airways Luggage is a producer of lightweight, cloth-covered luggage. The company distributes its
luggage through major department stores, mail-order houses, clothing retailers, and other retail
outlets, such as stationary stores and leather-goods stores. The company advertises rather
heavily, and supplements this promotional effort with a large field staff of sales reps, numbering
around 400. There is a large turnover in sales reps (10%-20% per year). Because the cost of
training a new person is estimated at $15.000 to $20.000, not including the lost sales that might
result because of a personal switch, Ms. Books had been conducting exit interviews with each
departing sales rep. On the basis of these interviews, she has formed the opinion that the major
reason is dissatisfaction with promotional opportunities and pay. But top management has not
been sympathetic to Ms. Books’ pleas regarding the changes needed in these areas of corporate
policy, saying her suggestions are based on intuition and little hard data. Ms. Books had decided to
call on their in-house Marketing Research Department.
Identify the general hypothesis that would guide your research efforts.
What type of research design would you recommend to Ms. Books? Why?
Category of research
Data collection method Exploratory Descriptive Causal
Secondary sources
Information systems
Databanks of other organizations
Syndicated services
Primary sources
Qualitative research
Surveys
Experiments
Secondary data
Commercial Published
There are a lot of useful sources of internal secondary data. For each of the following documents,
indicate what kind of information it provides.
Document Information provided
Sales invoice
Salesperson’s call
report
Salesperson’s
expense account
Individual customer
record
Financial record
Credit memo
Warranty card
A large chain of building supply yards was aiming to grow at a rate of three yards per year. From
past experience, this meant carefully reviewing as many as 20 to 30 possible locations. You have
been assigned the task of making this process more systematic. The first step is to specify the
types of secondary information that should be available for the market area of each location. The
second step is to identify the possible sources of this information and appraise their usefulness.
From studies of the patrons of the present yard, you know that 60 percent of the dollar volume is
accounted for by building contractors and tradesmen. The rest of he volume is sold to farmers,
householders, and hobbyists. However, the sales to do-it-yourselfers have been noticeably
increasing. About 75 percent of the sales were lumber and building materials, although appliances,
garden supplies, and home entertainment systems are expected to grow in importance.
Small convenience or quota samples are often used, rather than rigorous,
statistically meaningful samples
The information sought relates to respondents’ motivations, beliefs, feelings, and
attitudes, not to facts about their lives and behavior
An intuitive, subjective approach is used to gathering the data
The data collection format is open ended, allowing respondents to express
themselves in their own words at a length they deem appropriate
The approach is not intended to provide statistically or scientifically accurate data,
but often to guide further investigation that will
Main categories
of primary data
Communication Observation
Quick-Stop Inc. recently opened a new convenience store in Northglenn, Colorado. The store is
open every day from 07:00 a.m. to 11:00 p.m. In order to better plan the location of other units in
the Denver metro area, management is interested in determining the trading area from which this
store draws its customers. How would you determine this information by questionnaire? By
observation method? Which method would be preferred? Be sure to specify how you would define
“trading area”.
The Metal Products Division of Geni Ltd. devised a special metal container to store plastic garbage
bags. Plastic bags posed household problems, as they gave off unpleasant odors, looked
disorderly, and provided a breeding place for insects. The container overcame these problems, as
it had a bag-support apparatus that held the bag open for filing and sealed the bag when the lid
was closed. In addition, the storage area held at least four full bags. The product was priced at $
59,99 and was sold through hardware stores. The company has done little advertising and has
relied on in-store promotion and display. The divisional manager was wondering about the
effectiveness of these displays and has called on you to do the necessary research. Should the
communication or observational method be used? Justify your choice.
Designing forms for primary data collection is a key component of most studies
It is crucial to control measurement error, as it is often the largest source of
preventable error in marketing research, far exceeding errors in sampling plans or
administration
− The function of a questionnaire is measurement in a formalized way
− If a question does not truly measure what it is supposed to, measurement error
is present
The only way to develop the skill of questionnaire design is to write a
questionnaire, use it in a series of interviews, analyze its weaknesses, revise it,
and then repeat this process for new surveys
− No steps, principles, or guidelines can absolutely guarantee an effective and
efficient questionnaire
− Questionnaire design is a skill that the researcher learns through experience,
after some basic principles are assimilated
Design must
Researchers sampledevelop a data-collection procedure that establishes an effective
link between the information needs and the question to be asked or the observations
Collect data
to be recorded. The success of the study is dependent on the researchers’ skill and
creativity in establishing this link.
Process and code data
Brands in a category
Unique definition of
Store types (e.g., supermarket, drug store)
Nominal numerical labels
Sales territories
(0, 1, 2, …)
Demographics (e.g., gender, occupation)
Preferences
Order of numerals Subjective frequencies
Ordinal
(0 < 1 < 2 …) Purchase likelihood
Demographics (e.g., age group, education)
Equality of Attitudes
Interval differences Opinions
(2-1 = 7-6) Index numbers
Objective frequencies
Equality of ratios Costs
Ratio
(2/4 = 4/8) Number of customers
Sales (units, value)
Main characteristics
Main characteristics
1) Statistics appropriate for nominal scale also apply to the ordinal scale
Main characteristics
1) Statistics appropriate for nominal and ordinal scales also apply to the interval scale
Assume that we have scaled brands A, B, and C on an interval scale regarding buyers’ degree of
liking of the brands. (Liking is typically assumed to be measured on an ordinal scale. For the sake
of analyzing what is and is not meant by an interval scale, assume that the liking-based scale
example is reasonable.) Brand A receives a 6, the highest liking score, B receives a 3, and C
receives a 2. What can be said about these interval-scaled data?
The liking for brand B is less favorable than that for brand A
Brand A is liked twice as much as brand B
The degree of liking between A and B is three times greater than the liking between B and C
Main characteristics
1) Statistics appropriate for nominal, ordinal and interval scales also apply to the ratio scale
Identify the type of scale being used in each of the following questions. Justify your answer.
(1) During which season of the year were you born? (winter / spring / summer / fall)
(2) What is your total household income?
(3) Which are your three most preferred candy bars? Rank them from 1 to 3 according to your
preference, with 1 as most preferred! (M&Ms plain / M&Ms peanut / Reese’s / Almond Joy /
Good & Plenty)
(4) How much time do you spend traveling to school every day? (under 5 minutes / 5-10 minutes /
11-15 minutes / 16-20 minutes / 21 minutes or more)
(5) How satisfied are you with Newsweek magazine? (very satisfied / satisfied / neither satisfied
nor dissatisfied / dissatisfied / very dissatisfied)
(6) On an average, how many beers do you drink in a day? (more than 3 / 2 to 3 / less than 2)
(7) Which one of the following courses have you taken? (marketing research / sales management /
advertising management / consumer behavior)
(8) What is the level of education for the head of the household? (some high school / high school
graduate / some college / college graduate and/or graduate work)
The analysis for each of the questions on the previous page is given below. Is the analysis
appropriate for the scale used? Is the conclusion appropriate?
(1) About 50% of the sample were born in the fall, 25% of the sample were born in spring, and the
remaining 25% were born in the winter. It can be concluded that the fall is twice as popular as
the spring and the summer seasons.
(2) The average income is $25.000. There are twice as many individuals with an income of less
than $9.999 than individuals with an income of $40.000 and over.
(3) M&Ms plain is the most preferred brand. The mean preference is 3,52.
(4) The median time spent traveling to school is 8,5 minutes. Three times as many respondents
travel fewer than 5 minutes than respondents traveling 16-20 minutes.
(5) The average satisfaction score is 4,5, which seems to indicate a high level of satisfaction with
Newsweek magazine.
(6) 10% of the respondents drink less than two bottles of beer a day, whereas three times as many
respondents drink over three bottles a day.
(7) Sales management is the most frequently taken course because the median is 3,2
(8) The responses indicate that 40% of the sample has some high school education, 25% of the
sample are high school graduates, 20% have some college education, and 10% are college
graduates. The mean education level is 2,6.
Respondents characteristics
Personal factors such as mood, fatigue, willingness to participate, and even health
at the time of administration
Situational factors
Variations in the environment in which the measurements are reached, such as
time of the day, temperature, presence of family members, and point-of-purchase
marketing activities
Data collection factors
Variations in how the questions are administered, including the influence of the
interviewing method (e.g., phone, personal contact, web-based, or mail)
Measuring instrument factors
The degree of ambiguity and difficulty of the question and the ability of the
respondents to answer them
Data analysis factors
Errors made in the coding and tabulations process
Error that causes a constant bias in Error implied by influences that bias
the measurements: it affects the measurement but are not systematic:
measurement in a predictable way it produces inconsistency in repeated
− Measuring time with a stopwatch measures
that runs 10% faster than it should − Measuring time with several
− Measuring height with a poorly stopwatches that run without
calibrated wooden yardstick systematic error
− Measuring height with an elastic
ruler
Reliability refers to the extent to which the measurement is free from random error,
validity to the extent it is free from both systematic and random error
Two different measurements of the same Predictive validity involves the ability of a
marketing phenomenon administered at the measured marketing phenomenon at one
same point in time are correlated point in time to predict another marketing
Concurrent validity is primarily used to phenomenon at a future point
determine the validity of new measuring If the correlation between the two measures
techniques by correlating them with is high, the initial measure is said to have
established ones, which have presumably predictive validity
performed well in the past
Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 177-179, case
1.7 Hepworth Golden Auto.
(1) What developments in HGA performance led Bill Douglass to ask Rightway Research’s help?
What is the decision-making process that Mr. Douglass expects to be facilitated with the
insights provided by the research findings?
(2) Evaluate the stated research objective. Does it thoroughly address HGA’s core problems?
(3) Evaluate the information needs stated in the proposal. Have these been correctly identified?
Should any be altered, deleted, or added?
(4) Consider the data sources listed in the proposal. Are these sources appropriate for the stated
information needs and research objective? Why or why not? What other data sources could be
used? Consider accessibility, accuracy and cost in making source suggestions.
(5) Evaluate the data collection framework.
(6) What other possible research design and data sources could be used? How would these fit the
information needs relative to the given design?
(7) Considering the research proposal, do you think Rightway Research truly understood why the
information was needed, and for what eventual purpose? Explain your answer.
(8) Building on your previous answers, prepare a new research proposal for the HGA research
project.
Design sample
Collect need
Researchers data to clearly define who or what is to be included in the sample, the
population from which the sample is to be drawn, the methods used to select the
Process and code data
sample, and the sample size.
Analyze and interpret data
Benefits of sampling
If speed, cost, and efficiency of effort are not vitally important in a specific project,
neither is sampling; but such situations are exceedingly rare
Design sample
Sampling practice
30,0
10 20 50 100 1.000 10.000
Sample size
A fundamental error lies in the belief that large samples imply less error and therefore better re-
sults; although this is true for sampling errors, it is most certainly not true for non-sampling errors
Pete Thames, the general manager of the Winona Wildcats, a minor league baseball team, is
concerned about the declining level of attendance the team’s games in the past two seasons. He is
unsure whether the decline is due to a national decrease in the popularity of baseball or the factors
that are specific to the Wildcats. Having worked in the marketing research department of the
team’s major league affiliate, the New Jersey Lights, Thames is prepared to conduct a study on the
subject. However, because it is a small organization, limited financial resources are available for
the project. Fortunately, the Wildcats have a large group of volunteers who can be used to
implement the survey. The study’s primary objective is to discover the reasons why Winona
residents are not attending games. A list of 1.200 names, which includes all attenders for the past
two seasons, is available as a mailing list.
What is the target population for this study?
What is the appropriate sampling frame for this study of the attitudes of both attenders and
nonattenders?
Which kind of sample would provide the most efficient sampling?
Why would this method of sampling be the most efficient in this situation?
Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 377-379, case
2.7 Delta Dairy local market taste test survey.
(1) Define the population and the sampling frame for this survey project.
(2) What sampling procedure would you recommend for this taste test survey? Why? What potential
deficiencies do you foresee with your suggested procedure?
(3) One alternative for how to conduct the survey was to place taste tables on the most crowded
streets of the “capital” of northern Greece (Thessaloniki) and select passersby who would be
willing to participate in the survey. Evaluate this alternative and discuss its advantages and
disadvantages. What might be a better method, from the point of representativeness, bias, and
other principles of sampling?
(4) Should the taste test be blind or nonblind? Why? Should explicit brand names be tested? Would
it be a good idea for Delta to present information about the various milks being tasted, or even
mock-up of containers in which it will be sold? Should pricing information be tested
concurrently?
(5) Considering the survey’s objective and the time and budget constraints, what milk samples
would you include in the taste test? Would you vary them across different respondents? How?
Might you allow consumers to simply try whatever milk products they like best?
(6) Is it important to study different regions of the country, or different regions within northern
Greece? Should certain segments be deliberately overrepresented in the sample? Should any
type of stratification be imposed? Finally, what sort of biases might arise from simply using a
random sample – assuming one can be accessed – of the northern Greek population?
Design sample
Collect data
Process
Collecting dataand code involves
typically data a large proportion of the research budget and a
sizeable proportion of the total error in the research results. Consequently, selection,
Analyze
training, and interpret
and control data are essential to effective marketing research
of interviewers
studies.
Present results and conclusions
Design sample
Collect data
EditingAnalyze
involvesand interpretthe
reviewing data
data forms as to legibility, consistency, and
completeness. Coding involves establishing categories for responses or groups of
Present
responses results
so that and conclusions
numerals can be used to represent the categories.
Design sample
Collect data
Present must
Data analysis results
beand conclusions
consistent with the requirements of the detailed information
needs identified at the beginning of the project. Analysis is usually performed by
using appropriate statistical software packages.
For each of the following research questions indicate (1) the number of variables to analyze, (2) if
it’s descriptive or inferential statistics, as well as (3) the scale level.
What is the level of association between age (in years) and weight (in kg) in the sample?
What is the impact of height (in cm), age (in years) and gender (male / female) on weight (in kg)
in the sample?
Is the average age (in years) of the customer base 25?
What is the average age (in years) in the sample?
Is the level of association between age (0-18, 19-45, 46-65, >65) and weight (0-50, 51-70, 71-
90, >90) in the population greater than zero?
Univariate data analysis procedures allow to get a “feel” for the data and are a good
opportunity to weed out coding errors, outliers and other potential problems
The mode is the category The median is defined as The mean is the sum of the
of a nominal variable that the middle value when the values divided by the
occurs most often data are arranged in order sample size
The mode should not be of magnitude; that is, half The mean is – in contrast
applied to ordinal or the values fall above the to the median – not robust
interval data unless these median and half fall below1)
data have been grouped The median is an
first especially useful measure
when the data contain
outliers (= values well
outside the range of most
of the data), as it resists
dramatic change in the
presence of outliers; i.e.,
the median is robust
1) When there are an even number of data points, the median is calculated by taking the midpoint of the two middle values
The data in the file MA2 AdMR Data file 1 [Link] stems from a comparison of 17 different beer
brands.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)!
Calculate the appropriate measures for the central tendency. Interpret the results!
The product manager for Schmidts decided to reposition his beers and to enter the premium
segment. While keeping the recipe unchanged, the price was changed to 1,30.
Change the data file accordingly and re-calculate the appropriate measures for the central
tendency. Interpret the findings!
Absolute frequencies are the number of The variance is the sum of the squared
items in the sample in each category; deviations of the data (observations) from
relative frequencies present the data as the sample mean divided by the degrees of
proportions freedom
Frequencies are typically presented in a − Squaring accomplishes the desirable
histogram (bar graph of the relevant cell feature to have both positive and
counts) negative deviations count equally
− The degree of freedom equals the
number of independent observations (on
the variable of interest) minus the number
of statistics calculated (from the same
data) used in any particular formula
The standard deviation is the square root of
the variance and measures the dispersion of
the data around the (sample) mean
Due to a significant decline in profits, the price of Schmidts was changed back to 0,30.
As far as appropriate, calculate for the variables the variance and the standard deviation and
interpret the results. What is the main challenge in interpreting the figures?
Copy the data of costs into an Excel file (column A) and calculate the mean (column B). Subtract
the mean from each data point (column C) and sum these differences up (column D). Interpret
the sum! Square each difference from column C (column E) and sum these squared differences
up. Interpret the sum! Divide the sum by the degrees of freedom and compare the result with the
corresponding value in the SPSS output.
Calculate absolute and relative frequencies including a histogram for the variable calories and
interpret the results.
Re-do the calculation superimposing the histogram with a normal distribution (with the same
mean and standard deviation as the variable calories). Interpret the results
The null hypothesis H0 states that a population parameter takes on a particular value
By contrast, there is an alternative hypothesis H1, and this is what the researcher is attempting
to verify
− There are three types of alternative hypotheses: the population parameter (e.g., µ)
(1) does not take a specific value (e.g., 25): H0: µ = 25 and H1: µ ≠ 25
(2) is greater than that value: H0: µ ≤ 25 and H1: µ > 25
(3) is less than that value: H0: µ ≥ 25 and H1: µ < 25
− Whenever the alternative hypothesis is directional (no 2 and 3), a one-tailed test applies,
because only large deviations in one direction (large positive deviations in case of no 2, large
negative deviations in case of no 3) would count against H0 in favor of H1
− Otherwise (no 1), a two-tailed test applies, i.e. either large or small values of the sample
statistic would cause the researcher to reject H0
Depending on the empirical evidence, we will be able to say “we reject H0 in favor of H1”
− We will never accept H0, but merely fail to reject it
− In a situation where we fail to reject H0, we do not conclude that H0 is valid; all we can say is
that we do not have suitable evidence to reject it
True condition
H0 is true H0 is false
Convicting the
innocent
Suppose you have to choose the best promotional strategy for a new product introduction from
among three possible strategies. The null hypothesis is that all the strategies are equally effective.
In business terms, what does the type I error mean in this specific case? Evaluate the level or
risk associated with a type I error.
In business terms, what does the type II error mean in this specific case? Evaluate the level of
risk associated with a type II error.
Yes No
Yes
Population
z-test or t-test
standard
deviation σ
is known
No t-test
Two-tailed test
α/2 α/2
µ0
H0: µ = 25
H1: µ ≠ 25
α α
µ0 µ0
H0: µ ≤ 25 H0: µ ≥ 25
H1: µ > 25 H1: µ < 25
The data in the file MA2 AdMR Data file 2 Customer [Link] stems from 372 customers of
three different brands which were asked about their satisfaction with the respective brand.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)! Variables with “none” in column “values” were
measured on a scale from 1 to 10. Afterwards, the values were recalculated according to the
formula ((x – 1) * (100 / 9)), meaning they now range from 0 to 100.
Select brand A and calculate the appropriate measures for the central tendency as well as the
dispersion (incl. histogram) for the variables Sat_Ov, Fulfill and Closeness. Interpret the results!
From last years study, a value of 78 for the overall satisfaction is still in the mind of the CEO. He
asks you whether this years results will allow the conclusion that the population overall satisfaction
is greater than 78.
Formulate the corresponding null and alternative hypothesis!
Perform a t-test (significance level of 5%). Interpret the results.
Compare the empirical t-value (SPSS output) to the theoretical t-value and interpret the results.
How does the interrelation between the theoretical t-value and the significance level α look like?
The companies target for the running year was to achieve a rating of 81 for Closeness. By
achieving the target, the employees would be eligible to a bonus. The CFO claims, that the target
was missed.
Formulate the corresponding null and alternative hypothesis!
Perform a t-test (significance level of 5%). Interpret the results.
Compare the empirical t-value (SPSS output) to the theoretical t-value and interpret the results.
Researchers often need to make inferences about how respondents are distributed across the
possible categories of a nominal variable
− E.g., has the consumer base changed in terms of geographic location (countries, regions,
states, ZIP codes, …), gender split, educational entertainment, or income categories?
The chi-square test is a procedure for comparing a hypothesized population distribution (across
nominal chi-square test categories) against an observed distribution
In that way, it is just like all hypothesis tests we have seen thus far: it compares what we
hypothesize in a population with what we observe in a sample
Customers of brand A participating in the customer satisfaction study were asked about their trade
(customer segment) and the number of their employees.
Calculate frequency tables for each of the two variables (Trade, No_employ) and interpret the
results!
Based on the frequency tables you assume that the categories of Trade and No_employ
respectively do not occur with equal probabilities.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.
Before conducting the study, the Sales Director requested that the sample represents a population
with the following characteristic: the trades 1, 2, 3, and 8 contribute 20 % / 12 % / 45 % / 12 % to
the population, the other trades contribute equally to the remaining share.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.
On top, he requested that the sample represents a population with the following characteristic: the
size classes 8, 7, and 6 contribute 30 % / 35 % / 25 % to the population, the other size classes
contribute equally to the remaining share.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.
In examining the relationship between two interval variables, a useful beginning is to plot the
data on a scatter diagram
− Including the mean values for each of the variables divides the scatter diagram into four
quadrants
− If the data points are concentrated mostly in diagonal quadrants, this would be evidence of a
relationship between the two variables
The covariance measures the degree to which X and Y tend to co-vary, that is, vary in the same
direction from their respective means
− Covariance is positive if the values of Xi and Yi tend to deviate from the respective means in
the same direction, and negative if they tend to deviate in the opposite direction
− If X and Y are statistically independent, cov(X,Y) will be near zero in reasonably large samples
− The reverse is not true: the covariance can be exactly zero in any sample without the variables
being statistically independent
The linear correlation coefficient rxy is just a standardized measure of covariation (the linear
relationship between X and Y)
− It does not measure all possible relationships between the two variables
− It does not make claims about X and Y causing one another
No matter what units we choose to measure two variables, their correlation will not change (the
covariance will!)
The correlation coefficient may take on any value between -1 and +1
When rxy = 1, this indicates a perfect positive correlation; plotting the two variables in question
will show all points exactly on a straight line with positive slope (negative slope in case of rxy = -1)
If rxy = 0, there is no linear relationship between the variables; the best line through the points will
be flat (i.e., a slope of zero)
The exact percentage of variation shared by two variables is calculated by squaring rxy.
r2xy is called the coefficient of determination
The data in the file MA2 AdMR Data file 3 [Link] contains technical data of 406 different car
models.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)!
For the variables mpg and cu_inc, calculate the mean and produce a scatter plot (cu_inc on the
x- and mpg on the y-axis). Interpret the results! Which kind of relationship between the two
variables does the scatter plot indicate?
Calculate the covariance as well as the correlation between the two variables and interpret the
results.
Calculate the bivariate covariance as well as the bivariate correlation between each pair of the
variables mpg, cu_inc, hp, weight, and accel. Interpret the results! Compare the absolute values
of these two measures.
Broadly speaking, regression helps to understand how one or more independent variables are
related to a dependent variable of interest, and to make predictions based on this understanding
Specifically, simple regression is appropriate for one interval-scaled dependent variable (Y) and
one interval-scaled independent variable (X)
− In the absence of other information (i.e., values of X), our best guess at the value of the
¯
variable Y is its mean (Y)
− Our task in using a simple linear regression will be to do better than this, to use another
variable, X, to help explain Y
The principle of least squares suggests that while trying to “do it better than this” we should
minimize the sum of squared deviations between observed values (Yi) and those predicted by
^)
the model (Y i
Parameters used to explain and to predict a dependent variable in a simple regression model
The F-test tells you about all the The t-test tells you about each of the
independent variables in your model taken independent variables in your model
together considered separately
The F-value is calculated by dividing the The t-value is calculated by dividing the
mean square regression by the mean estimate for the unstandardized coefficient
square error by the standard error
− The mean square error measures how Large values indicate that the estimates are
much variance in each data point can be many standard errors away from zero, and
attributed to error therefore are unlikely to have come about
− The mean square regression measures purely by chance
how much variance in each data point The p-values indicate the probability
can be attributed to the regression associated with a t-statistic (two-tailed test)
Note that the F-test is always a one-tailed
test
In simple regressions the F-value will always be the square of the t-value for the slope;
but this will not be the case when there are multiple independent variables
Guess at a certain car models mpg and cu_inc. Justify your choice.
You intend to make a more accurate prediction by running a simple linear regression model
using mpg and cu_inc. Which variable should function as the dependent variable?
Conduct a regression analysis and interpret the results.
Predict mpg for a car with 304 cubic inches.
For case no 1 (first row in the data file), forecast mpg and calculate the deviation due to error. As
well, calculate the total variation and the share of variation explained by the regression.
Researchers will often wish to test whether a regression coefficient has a specific value
− E.g., is the slope significantly different from 0, or in general is it significantly different from any
(hypothesized) value
As in other statistical tests, a null and an alternative hypothesis needs to be specified1)
− H0: β = 0
− H1: β ≠ 0
As a standard, most statistical programs test whether β = 0 and provide the corresponding p-
value as well as the confidence interval
For all other hypothesized values, H0 can be tested using one of the following rules
− In case the confidence interval does not include the hypothesized value of β, H0 can be
rejected (at the confidence level used to calculate the interval); in case it does contain the
hypothesized value, we are not able to reject H0
− In case the (empirical) p-value is bigger than the defined significance level α, H0 can not be
rejected
• The p-value for any hypothesized value of β is determined by the degrees of freedom and
the t-value
• The t-value equals (b – β) / sb (with s being the standard error)
1) In statistics, generally Greek letters are used for population parameters (which will never be known without collecting
data on the entire population) and Roman letters for statistics calculated from a sample
You intend to predict mpg by hp. Run a simple linear regression model and calculate a 95 %
confidence interval for the regression parameters b0 and b1. Interpret the results. Is β1
statistically significant different from 0? Formulate H0 and H1. Justify your conclusion.
Test whether β0 is statistically different from 41,0 (α = 5 %). Formulate H0 and H1. Justify your
conclusion.
The R&D department claims that according to their calculations β1 equals -0,172 (α = 1 %).
Formulate H0 and H1. Justify your conclusion.
Predict mpg for a car with 150 hp.
For case no 2 (second row in the data file), forecast mpg and calculate the deviation due to
error. As well, calculate the total variation and the share of variation explained by the regression.
The R&D department asks whether there is a difference in mpg between cars from the US and
from Europe. Run a simple linear regression model and calculate a 99 % confidence interval for
the regression parameters b0 and b1. Interpret the results. Is β1 statistically significant different
from 0? Formulate H0 and H1. Justify your conclusion.
Among the most common research questions in marketing practice is whether there is an association
between two nominal variables
− E.g., is there a relationship between consumer ethnicity and media consumption habits?
A relationship between two nominal variables can be illustrated with a cross tabulation, its statistical
significance can be tested with a chi square-test
− The null hypothesis for this chi square-test is that the two variables are independent of each other; the
alternative hypothesis is that they are not independent, that is, that there is a relationship between the
two variables
In case the two variables (A and B) are independent, the expected cell counts are determined by the
multiplication rule
− If A and B are independent, the probability of Ai and Bj occurring is the product of the probability of Ai
times the probability of Bj (multiplication rule)
− The expected number (or cell count) of cell Ai / Bj is the product of the total number of observations
times the probability of Ai and Bj occurring
A chi square-test should only be conducted if all the expected cell counts were 5 or greater
− If this is not the case, it is generally recommended that cells be combined to give an expected
frequency of at least 5, as otherwise the chi square-test can yield misleading results
In case the calculated chi square-value exceeds a critical (theoretical) chi square-value (given a certain
significance level α), the null hypotheses is rejected
− The chi square-value reflects the difference between the observed and the expected number of cell
counts across all cells
The file MA2 AdMR Data file 4 [Link] stems from 180 interviewees which were asked
about fifteen statements reflecting their opinion on foreigners.
What is the scale of the different variables? Correct, when necessary, the scale level indicated in
the column “measure” (variable view)!
Test whether there is an association between the variables SocAct and EcoSit, between SocAct
and Occup as well as between EcoSit and Occup. To do so, run cross tabulations and chi
square-tests (α = 5 %). Interpret the results!
Investigate the nature of the relationship between SocAct and Occup. To do so, calculate the
percentage deviations between the observed value and the expected value. Interpret the results!
Test whether there is an association between the variables Age and EcoSit. Interpret the results!
1) Further multivariate analysis techniques are, among others, multivariate analysis of variance (MANOVA),
multidimensional scaling (MDS), correspondence analysis, and structural equation modeling (SEM)
Regression analysis is …
… a way to put a line through a group of points
This line minimizes the total sum-of-squares; that is, the squared distance to the line, summed
over all the points
… a method for testing the validity of relationships
Regression can help us determine whether marketing actions actually work; Verify
e.g., is there a relationship between ad spending and sales?
… a flexible methodology for measuring how things influence one another
Regression quantifies the nature of a relationship and by that allows to Quantify
determine how well marketing actions work
… a scientific approach to forecasting and prediction
Regression can help to determine how well another marketing action may Predict
work; e.g., add spending levels we did not try
When trying to say anything at all about the empirical world, regression is an
indispensable tool for an exceptionally wide variety of applications
Theoretical
model Y = β0 + β1 X1+ ε Y = β0 + β1 X1 + β2 X2 + β3 X3 + … + βK XK+ ε
(population)
Assumptions concerning the relationship between the dependent and an independent variable
I. There is a linear relationship between the dependent variable and each independent variable1)
To estimate a regression model, the errors do not need to be normally distributed (as claimed in
many textbooks); nevertheless, being normally distributed is a prerequisite to conduct a t-test
1) Note that multiple regression can be used to assess all nonlinear relationships as long as they can be linearized
A key feature of the error is that it should be information-free, lacking meaningful patterns
− A meaningful pattern indicates that the initial model is leaving something out (e.g., seasonality)
or is just plain wrong
− If there are meaningful patterns, the researcher should come up with a model to explain the
patterns, not chalk them up to error
There are many varieties of autocorrelation, and some complex error relationships can be difficult
to detect
− One way to see whether the error is pattern-free is to ask a simple question: will knowing the
value of the error for one data point (ei) tell anything about the error value at another point?
The Durbin-Watson (DW) test is a simple test to detect the most common form of autocorrelation,
the so-called first-order1) autocorrelation
− It tests whether the residuals (errors) from a linear regression are autocorrelated
− The test will result in a value between 0 and 4, values near 2 indicate that autocorrelation is
not a problem
− On top, autocorrelation can be detected by plotting the standardized residuals against the
dependent variable
In situations where autocorrelation is present, it could be corrected for by a transformation
(particularly logarithms, exponents, and powers), using (time) lags, first differences (change in a
variable), or dummy variables (e.g., seasonality)
1) The name refers to the fact that the test looks at relationships between adjacent points
1) Note that multiple regression can be used to assess all nonlinear relationships as long as they can be linearized
2) In case the number of observations is smaller than the number of parameters, the parameters will (= can) not be estimated
To conduct inferential tests, the error ei, that is, the difference between the predicted value for the
dependent variable and the value actually observed, should be normally distributed
Creating a histogram (which is typically called a normal probability plot) for the standardized
residuals and superimposing the best-fitting normal distribution helps to detect outliers and a
significant deviation from normality
− If the histogram and the best-fitting normal distribution appear to diverge substantially, there
may be a problem with non-normality
The degree of deviation from normality is also assessed by specific statistical tests, such as the
Kolmogorov-Smirnov, Anderson-Darling, and Shapiro-Wilk tests
Fortunately, small deviations from normality are no cause for worry
− However, if the normal probability plot indicates an extreme deviation between the histogram
and the superimposed normal distribution, the regression is likely to be misleading
− If non-normality is encountered, and if it is pronounced, it needs to be fixed, e.g., by using a
transformation and by checking whether some critical variable was omitted
Interval Interval
Ordinal
Nominal Nominal
− Binary − Binary
− Multinominal − Multinominal
Rank-ordered
Count
Coefficients
Fit measures (t and F)
p values
Although all regressions can be interpreted similarly, the mechanics of carrying out
regression can vary dramatically based on the type of dependent variable
The F-test tells you if all the variables, taken The t-test determines if parts of the model –
together, help explain the variation in the that is, the different independent variables –
dependent variable, Y help explain the variation in the dependent
variable, Y
“Fit” (or prediction) should be (much) bigger, Coefficients (the bi) should be different from
on average, than “error” zero
If the p-value associated with the F-test is If the p-value associated with a t-test is non-
non-significant, it says that the entire model significant, it says that the independent
is not providing sufficient explanatory power variable Xi does not have a statistical
− In that case, the model must be changed, significant impact on the dependent variable
usually by attempting to remove under- Y
performing independent variables by − In that case, Xi should be removed from
looking at their t-tests the model
You intend to run a multiple linear regression of cu_inc, hp, and weight on mpg (file MA2 AdMR
Data file 3 [Link]). Based on your theoretical knowledge and your experience, there is a linear
relationship between the mpg and the mentioned independent variables.
Check whether the assumptions of a multiple linear regression are met.
Based on the findings you decide to exclude all cases with a standardized residual of less than -3 or
more than +3.
Check whether the assumptions of a multiple linear regression are met.
You decide to not exclude the outliers. For the following calculations, assume that all the
assumptions of a multiple linear regression are met.
Run a simple linear regression model of hp on mpg (α = 5%). Interpret the results!
Run a multiple linear regression as described above (α = 5%). Interpret the results!
Modify the model based on the findings of the previous multiple linear regression. Run a multiple
linear regression for the modified model (α = 5%). Interpret the results!
Do the results allow the conclusion that the impact of hp is < -0,08 (α = 5%)? Justify your answer.
You intend to explain the account status (1 = no debt history; 2 = no current debt; 3 = payments
current; 4 = payments delayed; 5 = critical account) of a customer. To do so, you might run an
ordinal regression (file MA2 AdMR Data file 5 Credit [Link]). Based on preliminary analysis, you
identified three predictors (factors) number of credits at the bank, other installment debts, and
housing type as well as three covariates: age and duration of loan,
Run an ordinal regression (α = 5%). Interpret the results!
Being able to include nominal (or categorical) independent variables is critical in applying
regression models to real-world problems
− E.g., seasonality, country of origin, or educational level
A binary variable can be simply integrated into a multiple linear regression model as an
independent variable
− Binary variables are quite easy to interpret because the coefficient always represents a “one-
unit increase” in the quantity in question
A nominal variable (with c categories) needs to be transformed into c-1 binary dummy variables
first
− E.g., the four seasons (winter, spring, summer, and fall) are transformed into three binary
variables winter (0/1), spring (0/1), summer (0/1), and if we know it is not winter, not spring,
and not summer, it must be fall
Researchers must carefully weight the pros and cons of adding dummy variables just because
they can
− We loose statistical power – represented by degrees of freedom – whenever we add more
independent variables
− Sometimes, a simpler model (one with fewer independent variables) is not only easier to
understand, it may actually be more powerful and offer superior forecasts
You intend to run a multiple linear regression of weight, year, and country on mpg (file MA2 AdMR
Data file 3 [Link]). Based on your theoretical knowledge and your experience, there is a linear
relationship between the mpg and the mentioned independent variables.
First of all, recode country into two binary variables USA and Europe
Run a multiple linear regression as described above (α = 5%). Interpret the results!
Check whether the standardized residuals are normally distributed.
Modify the model based on the findings of the previous multiple linear regression. Run a multiple
linear regression for the modified model (α = 5%). Interpret the results!
Factor analysis takes a large number of variables and searches to see whether they have a small
number of factors in common that account for the correlations among the variables
Factor analysis has a number of possible applications in marketing research
− Data reduction
Reducing a mass of data (e.g., attributes) to a (far) smaller number of factors that underlie the
variables
− Structure identification
Discovering the basic structure underlying a set of measures
− Scaling
Identifying the optimal weights of variables being combined to form a scale (a weighted sum)
− Data transformation
Transforming the data (variables) into independent factors which can be used as an input for
many predictive techniques in statistics (e.g., multiple linear regression)
Among many other application areas, factor analysis is used for the development of
personality scales, market segments based on psychographic data, the identification of
critical product attributes, as well as similarities among products and lines
The factor loadings are just correlations between variables and factors
− If a factor loading is high (near -1 or 1), it means that the factor loads high on that variable;
that is, that variable will be used to interpret the factor later on
In factor analysis, most every quantity we would like to know stems directly from the factor
loadings; if we sum squared loading across each of the following, this is what we get:
− For any factor, across all variables: that factor’s eigenvalue
An eigenvalue represents how much variance a factor explains relative to how much it
would be expected to explain by chance alone; that is, on average
In terms of interpreting a factor analysis, the typical approach is that factors with
eigenvalues less than1 should be discarded; those with eigenvalues not too much greater
than 1 are suspect; and those with large eigenvalues should be retained
− For any variable, across just the factors in our solution: that variable’s communality
A communality represents how well all factors together explain each of the variables
In terms of interpreting a factor analysis, the total communality should be compared with a
best possible value; on top, if the objective is to explain the original variables equally well,
the communalities for the variables should be on a similar level
− For any variable, across all factors: 1
Varimax rotation “reorients” the original factors so that their loadings are as near -1, 0, or 1 as
possible while ensuring that the factors remain uncorrelated with one another (orthogonal
rotation)1)
− Although individual variables can correlate non-trivially with several factors, it must be
stressed that the factors themselves are perfectly uncorrelated (their correlation is zero)
Due to the varimax rotation, the variance explained by each factor will change
− The “best” of the factors will have smaller eigenvalue than before, whereas the “worst” will
have a larger eigenvalue
− The sum of the eigenvalues, however, will be the same
1) In contrast to an orthogonal rotation, an oblique rotation allows the factors to become correlated with one another
You intend to identify factors that account for the correlation between the fifteen statements
reflecting the opinion of British on foreigners (file MA2 AdMR Data file 4 [Link], variables
Stat01 to Stat15).
Run a factor analysis on the fifteen statements. Check the univariate descriptives as well as the
scree plot. Interpret the results!
Rotate the factor solution using varimax rotation. Interpret the results!
Based on the outcome of the analysis you decide to check a solution with four factors. Run a
factor analysis generating four factors and interpret the results!
Finally, you would like to check a solution with two factors. Run a factor analysis generating two
factors and interpret the results! Which alternative solution (2, 3, or 4 factors respectively) do you
evaluate as the most appropriate one?
Cluster analysis allows the researcher to place objects / items / people into groups; these
groups are often called clusters
− Cluster procedures form groups, assign objects (“cases”) to each of them, and help
determine a reasonable overall number of groups
− There are clustering algorithms available that take nominal, ordinal, interval, or ratio
measures as input1)
− Cluster procedures assume that natural clusters exist within the data
Among the most common uses of clustering techniques in marketing are segmenting customers
and segmenting products / brands
Cluster analyses involves a trade-off between two quantities: (1) the distance of each
point in a group to the group center, and (2) the distance between group centers
1) The following slides focus on metric variables, but the underlying principles for non-metric variables are similar
How a set of objects should be clustered The clustering procedure will count all
depends critically on variables included in the procedure as
equally important; it is up to the researcher
− which other objects are being clustered
to assess whether this is desirable or not
along with them and
An implicit weighting might be implied by the
− how much latitude within a cluster versus
across clusters should be allowed for fact that the variables are
− measured on different scales
Option: transform the variables first,
usually with a z-transform1)
− correlated (redundancy among variables)
Option: use (uncorrelated) factors as
variables
Based on theoretical considerations an
explicit weighting might be introduced by the
researcher, e.g. by multiplying the values by
some number (greater than 1 to emphasize,
less than 1 to de-emphasize)
1) This ensures that each variable is on an “equal footing” statistically, transformed to have a 0 mean and standard
deviation of 1
In order to group objects together, some There are two approaches to clustering, a
kind of similarity measure is needed hierarchical and a nonhierarchical approach
Users of clustering need to check carefully The distinguishing feature of hierarchical
that the metric is consistent with the cluster analysis (as opposed to non-
research question, especially whether hierarchical) is that once objects are
distance (volume) or dissimilarity (pattern) is clustered together, they are always together
of interest
The main hierarchical clustering method is
The most common distance measure for the Ward’s method (minimize within-cluster
metric variables is the (squared) Euclidian variation)
distance
− Additional methods are single linkage
− Additional distance measures are the (shortest distance), complete linkage
Minkowski metric, the Chebyshev (longest distance), average linkage
distance, and the city block metric (average distance), and centroid method
Common dissimilarity measures for metric (centroid distance)
variables are the Pearson correlation and The main nonhierarchical approach is the k-
the cosine means method
The merits of both approaches are combined, and hence the results should be better
Unfortunately, unlike in regression or factor analysis, where there are objective measures (such
as r2) of fit, one must take a more “exploratory” (that is, trial-and-error, until the results “look
good”) approach in cluster analysis, and base one’s judgment on more or less pictorial evidence
First, the researcher can specify in advance the number of clusters or a range for the number of
clusters based on theoretical, logical, or practical considerations
Second, the distance between clusters (error variability measure) at successive steps may
serve as a useful guideline
− As a rule of thumb, one should stop when the successive distances between steps make a
sudden jump
− A plot of error sums of squares with the number of clusters may help to identify the jumps
Finally, the total cluster pattern (dendrogram) can provide a feel for an appropriate number of
clusters
− Based on theoretical and practical considerations, it is worthwhile to check whether a
proposed cluster solution does make sense or not
− A lack of statistical significant differences between clusters with regard to variables
considered as relevant indicate that the cluster solution is of limited use
The question as to the “best” number or clusters can only be answered by the researcher,
by balancing the project’s need for accuracy (which will argue for more clusters) against
the universal desire for a simple, robust explanation (which will argue for fewer)
The data in the file MA2 AdMR Data file 6 [Link] stems from a comparison of 20 different beer
brands. You intend to cluster the different brands based on the variables Calories, Sodium,
Alcohol, and Price.
Check whether an implicit weighting might be an issue. Propose measures to handle possible
redundancies among the variables!
You decide to proceed with standardized values (z-scores), but to keep all four variables.
Run a hierarchical cluster analyses using the squared Euclidean distance and the Ward’s
method. Ask for a dendrogram. Interpret the results!
Propose how many clusters to extract. Justify your proposal!1)
Based on the assessment of the different options you favor the solution with four clusters.
Describe the four clusters based on the four cluster variables! (Note: you can save the cluster
membership as a variable)
Check whether there is a statistical significant difference between the four clusters with regard
to the four variables Calories, Sodium, Alcohol, and Price (α = 10%)!
1) To do so, please refer to the file MA2 AdMR Determine the number of clusters
Discriminant analysis is appropriate when one seeks to understand a nominal dependent variable
(e.g., a grouping) in terms of several – interval or binary – independent variables
Discriminant analysis has a number of possible applications in marketing research
− Estimate “discriminant functions” (discrimination)
Which linear combination of the given (independent) variables best distinguishes known
groups (dependent variable)?
− Make group predictions (classification / prediction)
Given a new set of items (e.g., customers, products, firms) whose group membership we do
not know, which of the pre-established groups are they likely to fall into?
− Determine whether the groups really seem different (testing / verification)
Are the various groups significantly different, based on the “profiles” (independent variables) of
the individuals found in them?
− Identify the most useful predictors in discrimination (influence / importance)
Which input variables seem to best predict group differences?
Which variables are most helpful in prediction who will reply to our
Influence / importance
promotions? Do some seem completely useless?
The basic idea of discriminant analysis is to find a linear combination of the independent
variables that makes the predicted mean for each category as different as possible
This linear combination of the n independent variables (known as the discriminant function or
axis) is derived from an equation that takes the form
DF = b1 X1 + b2 X2 + b3 X3 + … + bN XN,
which should look exactly like an ordinary multiple regression
When there are m different categories (groups) to distinguish (with m < n), discriminant analysis
produces (m-1) discriminant functions DF1, DF2, …, DFm-1
Making the mean for each category as different as possible is achieved by maximizing the
between-group variance relative to the within-group variance
Discriminant analysis output typically includes the
− values of the b’s, along with a
− confusion matrix which categorizes correct and incorrect predictions by cross-tabulating the
predicted with the actual category of the dependent variable, and some
− test statistics to asses significance
Discriminant analysis can handle many sorts of independent variables, not only interval-scaled
ones; however, nominal independent variables need to be converted to (binary) dummy variables
You are a loan officer at a bank and want to identify characteristics that are indicative of people
who are likely to default on loans. Information on 850 past and prospective customers is contained
in the file MA2 AdMR Data file 7 Bank [Link]. The first 700 cases are customers who were
previously given loans.
Run a discriminant analysis using the variables age, employ, address, income, debtinc,
creddebt, othdebt and default. Ask for the following statistics: means, Box’s M, univariate
ANOVA, Fisher’s as well as unstandardized function coefficients. Make sure to compute prior
probabilities from group sizes and to display a summary table and a leave-one-out
classification. Interpret the results!
Re-run the discriminant analysis using a separate-groups covariance matrix. Interpret the
results!
You decide to stay with the discriminant analysis using the within-groups covariance matrix.
Re-run the discriminant analyses, but assume that a case is equally likely to be a defaulter or a
nondefaulter. Interpret the results!
Assume that the average volume of a requested bank loan is the same in both groups.
According to internal data, the profitability of nondefaulters is 1% and of defaulters -5%. Draw
conclusions!
In a next step, the 150 prospective customers should be classified, in other words: a decision
needs to be made whether a loan should be made to them or not.
Classify the prospective customers. Make sure to compute prior probabilities from group sizes.
Interpret the results!
Re-run the discriminant analyses, but assume that a case is equally likely to be a defaulter or a
nondefaulter. Interpret the results and compare both classifications!
You decide to classify the prospective customers based on prior probabilities reflecting the group
sizes. However, an applicant will only be eligible to receive a loan in case the probability being a
nondefaulter is at least 55%.
How many prospective customers will receive a loan?
Conjoint allows marketers to find the “sweet spot” where consumers, manufacturers, and
retailers can most mutually agreeably meet
It is important to note that conjoint predicts preference, not sales or market share
Pairing every possible attribute level with every other in a conjoint analysis would typically result
in a large number of possible product profiles, making conjoint a practical impossibility if
respondents had to evaluate them all
− Six attributes and five levels per each attribute result in more than 15.000 possible profiles
An orthogonal design allows researchers to offer a subset of, rather than all, possible
combinations
− Orthogonal design works because it is assumed that what is observed for one variable is
unrelated to what is observed for any of the other
− This means that, when describing a product by its attribute levels, including a particular level of
one attribute gives no information about what levels are present for other attributes
− In the example mentioned above, the 15.000 profiles can be “orthogonalized” into just 25
The conjoint task involves ranking the subset of possible combinations from “most preferred” to
“least preferred”
Please rank the following 18 movie theater configurations according to your preference, 1
being the most and 18 the least preferred one!
Ticket price Line of sight Seat comfort Concessions
$6 Staggered Average seat Gourmet snacks
$6 Not staggered Big seat Hot dogs / popcorn
$8 Staggered Average seat Hot dogs / popcorn
The motivating idea behind adaptive conjoint is: prior choices or responses are utilized to
determine future comparisons
− If the first few conjoint responses strongly indicate that a person is insensitive to some aspects
of a product, he / she should be questioned about trade-offs more relevant to that person
Hence, an adaptive conjoint program will calculate part worths after each new piece of data is
supplied by the respondent, then use special algorithms to better measure those part worths that
are not yet determined with sufficient accuracy
− The researcher can specify exactly how much accuracy is required
− In a nonadaptive conjoint program, even when all the data are collected, there is no guarantee
of sufficiently accurate results
In a two-option adaptive conjoint task, the response is given on an ordinal scale (e.g., a scale
from 1 to 9), with one end of the scale representing “strongly prefer option A” and the other end
representing “strongly prefer option B”
Which of the following laptop computers would you rather purchase?
Which ofFastest
the following laptop computers would you rather
processor Verypurchase?
fast processor
Fastest 6processor
h battery life Very fast4processor
h battery life
6 h battery$life
999 4 h battery$life
799
$ 999
Strongly $ 799 Strongly …
prefer left 1 2 3 4 5 6 7 8 9 prefer right
Strongly Strongly
prefer left 1 2 3 4 5 6 7 8 9 prefer right
The idea behind choice-based conjoint is to have respondents simply make the sort of trade-offs
they normally do: in their heads, reporting only the item that they would actually choose
− For this reason, choice-based conjoint is dramatically more “real world” than other conjoint
tasks
Despite the need to present respondents with a larger number of experimental tasks when using
choice-based conjoint instead of ranking or rating, it can be highly efficient because several
product profiles can be offered at once
− Furthermore, because each task is far simpler and clearer to the respondent, this increase in
the number of tasks rarely inflates the complexity of the study or the time respondents must
put in
In a choice-based conjoint task, the respondent indicates which of the product profiles presented
to him he / she would choose or, importantly, that none of them is acceptable
− The no-choice option accounts for situations when the presented attribute level combinations
are all below some threshold for purchase
Which of the following desktop computers would you purchase?
WhichDell
of the followingHP
desktop computers
Sony would youNone:
purchase?
if these
Fast Fastest Very fast were my only
Dell HP Sony None: if these
processor processor processor choices, I
Fast Fastest Very fast were my only
27‘‘ monitor 24‘‘ monitor 21‘‘ monitor would defer
processor
$ 1.000
27‘‘ monitor
processor
$ 900
24‘‘ monitor
processor
$ 850
21‘‘ monitor
choices, I
my purchase
would defer
…
$ 1.000 $ 900 $ 850 my purchase
Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 531-534,
marketing research focus 11.1: designing your own movie theater using conjoint analysis and an
orthogonal design.
Check the relative importance of the different attributes by referring to the part worths. Explain
your calculation!
Assume that the part worths are calculated for a specific person, let’s call him Mr. Smith. There are
three cinemas Mr. Smith might choose from. Cinema A offers staggered seats, average seats
without cup holder, a small screen with plain sound, hot dogs and popcorn at a price of $ 6.
Cinema B differs from A that the seats are equipped with a cup holder, that the screen is large, the
sound is digital and the ticket is $ 8. In cinema C (in contrast to B) the seats are not staggered, the
sound plain and the ticket – like in A – is $ 6.
Which cinema do you expect Mr. Smith to prefer?
The owner of cinema B would like to gain Mr. Smith as a customer. By how much would he
need to reduce the ticket price if he would like to offer Mr. Smith a total utility that is 5% above
the best alternative? Use a linear interpolation!
Humans are visual creatures: they understand information best when it is integrated into a
picture, not when presented as a disembodied table of numbers
Multidimensional scaling is a way to construct “pictures” of markets that do not rely on someone
deciding beforehand which dimensions are important, or what data to collect in order to draw the
(perceptual) map
MDS summarizes data about associations between a fixed set of objects to reveal relationships
between them (usually brands in a particular product class) by determining the
− minimum dimensionality required to represent the objects’ interrelationship well, and the
− position of each object on each dimension (i.e., its location in the map)
without collecting any attribute-based data
MDS will always provide as faithful a representation of relative similarity or distance as possible,
but we can only attach meaning to the dimensions in the derived graph by tying in other kinds of
data
The MDS routine, unlike those for linear regression, is iterative, determining a good initial guess,
calculating a goodness-of-fit measure, and continuing to derive better and better estimates,
always attempting to better represent the distance data
You are a product manager at Volvo cars, being responsible for the V40. The file MA2 AdMR Data
file 8 [Link] contains a rank order of similarities between pairs of – in total – 11 car models
belonging to the so-called Golf-class. The rank number “1” represents the most similar pair.
Run a multidimensional scaling using the Euclidean distance. Ask for group plots, data matrix,
as well as model and options summary. Interpret the results!
How could you use the scaling results? Which conclusions would you draw from the results of
the analyses?
Design sample
Collect data
Regardless of the rigor, thoroughness, and soundness of the methodology underlying the report,
the research will be useless to executives if the report fails to make an airtight, compelling case