0% found this document useful (0 votes)
39 views143 pages

Advanced Marketing Research Methods

This document outlines the key steps and concepts in marketing research. It discusses establishing the need for research by precisely defining decision situations and opportunities. Next, it addresses detailing research objectives and information needs to guide the project. Finally, it examines how information needs specify the data required to achieve the objectives and make better decisions. The document provides an overview of best practices in defining and planning a marketing research study.

Uploaded by

j
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
39 views143 pages

Advanced Marketing Research Methods

This document outlines the key steps and concepts in marketing research. It discusses establishing the need for research by precisely defining decision situations and opportunities. Next, it addresses detailing research objectives and information needs to guide the project. Finally, it examines how information needs specify the data required to achieve the objectives and make better decisions. The document provides an overview of best practices in defining and planning a marketing research study.

Uploaded by

j
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Advanced Market Research

Prof. Dr. Martin Ohlwein

Frankfurt am Main I March 2015

MA2 AdMR 140205 I Page 1


Prof. Dr. Martin Ohlwein
Lecturer International School of Management I Frankfurt
[Link]@[Link]

Extension of the professional competence in the area of marketing


research, especially as to
the theoretical foundations of the main methods for analyzing
Content of
quantitative data,
lecture
their adequacy (e.g., typical areas of application, underlying
assumptions) as well as
the interpretation of the output

The student is in possession


of comprehensive, detailed, specialist and state-of-the-art knowledge
regarding the different steps in the marketing research process,
Qualification
especially as to data analysis and interpretation
targets of
of specialized skills in the area of marketing research relating to the
lecture
solution of strategic problems
of capabilities to assess the adequacy of different methods for data
analysis in the light of a specific research question

MA2 AdMR 140205 I Page 2


Marketing research is an aid to decision making

Definition of marketing research (American Marketing Association)

Marketing research is the function that links the consumer, customer, and public to
the marketer through information – information used to
identify and define marketing opportunities and problems,
generate, refine, and evaluate marketing actions,
monitor marketing performance, and
improve understanding of marketing as a process

Marketing research
specifies the information required to address these issues,
designs the method for collecting information,
manages and implements the data collection process,
analyzes, and
communicates the findings and their implications

In short: marketing research is a set of formal practices for doing what people do
all the time – gathering information and using it to make better decisions

MA2 AdMR 140205 I Page 3


Key task is to guide captain obvious

MA2 AdMR 140205 I Page 4


Questions for review and critical thinking

How might the following organizations effectively employ marketing research? Be specific!
A small sporting goods store
Städel Museum
International School of Management
Rhein-Main-Verkehrsverbund
MyZeil shopping mall
A public campaign against cigarette smoking
Eintracht Frankfurt soccer team
Member of the Deutsche Bundestag representing the electoral district Frankfurt am Main II

MA2 AdMR 140205 I Page 5


The research process consists of nine steps

Steps of the research process

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect data

Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 6


The need for research information must be
precisely defined

Establish need for information

RarelyDetail
does research objectives
an initial request for and
helpinformation
adequatelyneedsestablish the need for research
information. Managers often react to hunches and symptoms rather than to clearly
Set decision
identified researchsituations.
design andYetdata
thesources
need for research information must be precisely
defined if the research project is to provide information pertinent to the necessary
Design
decision. Too data
often,collection procedure
the importance of this initial step is overlooked in the excitement
of undertaking a research project, resulting in research findings that do not
Design
adequately sample
inform decisions.
Collect data

Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 7


Establishment of a clear statement of problems /
opportunities is a must
Establish need for information

Interdependence between symptoms and underlying problems / opportunities

Symptoms are performance measures, metrics, and diagnostics that signal the
presence of a problem / an opportunity
− E.g., market share below expectation, increase in sales of snack products with
less fat content
Symptoms themselves rarely contain information regarding their causes
Therefore, the decision-maker‘s task is to respond to symptoms by analyzing the
underlying problems / opportunities to determine whether the situation calls for a
decision
The process of identifying problems / opportunities involves analyzing past,
present, and possible future situations facing an organization to uncover those
variables that either
− cause poor performance or
− represent opportunities for further growth

Decisions must aim at solving problems or taking advantage


of opportunities, not at treating symptoms

MA2 AdMR 140205 I Page 8


Once there is a need for a decision, the decision
objectives need to be defined
Establish need for information

Tasks in defining the decision objectives

It is essential that the researcher understands the problem situation from the
perspective of the decision maker
Consequently, the first task is to identify the decision maker(s)
− Often, the person who first requests assistance from marketing research is not
the decision maker and may or may not know how the decision maker views the
specifics of the decision situation
In a second step, the researcher needs to translate goals likely lacking operational
significance (e.g., “we need to hold off the competition”) into focused objectives
Decision objectives include
− organizational objectives (e.g., increasing earnings per share by 10 percent next
year) and
− personal objectives of the decision maker(s) and those influencing those
individuals (e.g., getting promoted, acquiring more prestige)

A clear and concise statement of problems / opportunities and


decision objectives is critical to the formal research process

MA2 AdMR 140205 I Page 9


Finally, alternative courses of action should be
identified and evaluated
Establish need for information

Challenge of identifying alternative courses of action

A course of action specifies how the organization’s resources are to be deployed


in a given time period
The real challenge is to identify the course of action that will result in high
performance and give the organization a competitive edge
Merely identifying new and different courses of action is hard work, involving
“thinking outside the box”
− Creativity and innovative thinking can make a decisive difference
− Exploratory research can be especially helpful in identifying innovative courses
of action
The value of research is typically commensurate with the ability of research
information to reduce uncertainty regarding the selection of a course of action

The key check question is: what information is needed to properly


choose among the alternative courses of action?

MA2 AdMR 140205 I Page 10


Information needs specify the information
required to attain the research objectives

Establish need for information

Detail research objectives and information needs

Setgeneral
After the research design
need and data sources
for information is clearly established, researchers must list,
specifically and in detail, the objectives and information needs of the proposed
Design
research. data collection
Research objectivesprocedure
answer the question, “why is this project being
conducted?” and are typically put in writing before the project is undertaken.
Design
Information sample
needs answer the question, “what specific information is required to
attain the objectives?”.
Collect data

Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 11


Research objectives answer the question: what
is the purpose of the research project?
Detail research objectives and information needs

Connection between research objectives and information needs

Research objectives answer the question, „what is the purpose of the research
project?“
Research objectives serve to guide the research project by
− giving direction to the specific information to be gathered (information need),
which in turn
− guides the specific questions developed for the questionnaire
Each question on the questionnaire should have a direct correspondence to an
information need, and each information need should have a direct
correspondence to at least one research objective
− Otherwise, unneeded data will be collected

When developing the information needs, both the manager and the researcher
should ask, for each item: realistically, can this information be obtained?

MA2 AdMR 140205 I Page 12


Detail research objectives and information needs

Questions for review and critical thinking

Research objectives can be stated so broadly that they fail to communicate the specifics of why
the study is being conducted. Evaluate the following statement: To study consumer reactions to
cartoon characters in advertising. Propose a more precise and useful formulation.

MA2 AdMR 140205 I Page 13


Detail research objectives and information needs

Questions for review and critical thinking

Given the following decision objectives, identify the research objectives.


What pricing strategy to follow for a new product?
Whether to increase the level of advertising expenditures on print or online?
Whether to increase in-store promotion of existing products?
Whether to increase training for frontline service providers?
Whether to change the sales force compensation package?
Whether to change the combination of ticket price, entertainers, and security at the Indiana
State Fair?
Whether to revise a bank’s electronic payment service?

MA2 AdMR 140205 I Page 14


Potential findings should be visualized, decision
criteria developed and cost-benefit ratio evaluated
Detail research objectives and information needs

Visualize research findings Develop decision criteria Evaluate cost-benefit ratio

Mocking up the potential Decision criteria (if-then Most activities in an


research findings is a statements) are rules for organization are evaluated
valuable way to ensure that selecting among courses of on a cost-benefit basis; so
the data to be collected will action, given various should research projects
fit the information needs outcomes Quantifying the costs
specified by the decision
It is crucial that decision associated with a research
maker project is fairly
criteria be developed
In essence, the research before anyone involved straightforward, but it is
process and presentation learns the actual results to difficult to quantify the
are simulated before ensure that organizational benefits, which are often
conducting the project to objectives take priority over subjective in nature
ensure that no conceptual personal objectives The number of units of a
bottleneck prevent the product that need to be
project from being swiftly
sold (on top) to break even
and professionally on the cost of the research
completed project qualifies as a
pragmatic indicator

Do we cover all information Do we know what we are going Does the research project
needs properly? to do given a specific outcome? economically make sense?

MA2 AdMR 140205 I Page 15


A research request form imposes a contractual
degree of commitment
Detail research objectives and information needs

Content of a research request form

Topic Guiding questions


What led to the recognition that research is needed?
Background Were there specific metrics which suggested research be
conducted?
Objectives What, specifically, are the decision objectives?
Problem / opportunity What are the causal factors underlying the decision situation?
Decision alternatives What, specifically, are the alternative courses of action?
What is the overall purpose of the research?
Research objectives
How will one know if that purpose has been satisfied?
What types of information are needed?
Information needs
By what means can each be attained?
Examples of questions What sorts of representative questions might be asked?
What concrete criteria should be used to select the best
Decision criteria
alternative?
Why is the research useful?
Value of research Can the usefullness be measured against hard figures such as
sales, share or profit?

MA2 AdMR 140205 I Page 16


Detail research objectives and information needs

Questions for review and critical thinking

In each of the following situations, identify the fundamental source of the marketing problem or
opportunity, a decision objective arising from the marketing problem or opportunity, and a possible
research objective .
Cool Pool Supply is a manufacturer of swimming pool maintenance chemicals. Recently, a
malfunction of the equipment that mixes anti-algae compound resulted in a batch of the product
that not only inhibits algae growth but also causes the pool water to turn a beautiful shade of
light blue.
The MBA director of a local college recently extended the offers to 20 promising students. Only
five offers were accepted. In the past, acceptance rates have averaged 90%. A survey of non-
acceptors conducted by the director revealed that the primary reason students declined the offer
was their perception that the college’s course requirements are too “restrictive”.
Chocoholic Candy Company has enjoyed great success in its small regional market.
Management attributes much of this success to Chocoholic’s unique distribution system, which
ensures twice-weekly delivery of fresh product to retail outlets. The directors of the company
have instructed management to expand Chocoholic’s geographical market if it can be done
without altering the twice-weekly delivery policy.

MA2 AdMR 140205 I Page 17


The research design guides data collection and
analysis

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Once Design data


the study collection
objectives areprocedure
determined, researchers design the formal research
project, including the appropriate sources of date for the study. A research design is
Design
the general sample
plan guiding the data collection and analysis phases of the research. It is
the framework that specifies the type of information to be collected, the sources of
Collect
data, and the data
data-collection procedures and analysis.
Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 18


There are three categories of research:
exploratory, descriptive, and causal
Set research design and data sources

Exploratory research Descriptive research Causal research

Purpose is to provide Purpose is to provide an Purpose is to provide


insights into the general accurate snapshot of some evidence regarding the
nature of a problem, the aspect of the market cause-and-effect
possible decision environment relationships operating in
alternatives, and relevant the marketing system
Research methods are
variables that need to be
marked by a clear Research methods are
considered marked by a planned and
statement of the decision
Research methods are problem, specific research structured design that will
highly flexible, objectives, detailed minimize systematic error,
unstructured, and information needs, and a maximize reliability and
qualitative, including carefully planned and permit reasonable
secondary data, structured research design unambiguous conclusions
observation, interviews with regarding causality
It aims to portray the
experts, focus groups, and
characteristics of marketing It aims to understand which
case studies variables combine to give
phenomena including
It aims to generate formal, frequency of occurrence, rise to the effect, focusing
testable hypothesis determine the degree to on why the effect happens,
regarding potential which marketing variables and to understand the
problems or opportunities, are associated, and predict functional relationship
or both the occurrence of between the causal factors
marketing phenomena and the effect

MA2 AdMR 140205 I Page 19


The three categories can be characterized by
typical research objectives
Set research design and data sources

Research objectives and hypothesis typical for the three categories of research

Decision objective Research objective Hypothesis


Exploratory research
What new product should What alternative ways are there to Boxed lunches are better than other
be developed? provide lunches for school children? forms
How can our service be What is the nature of any customer Suspect that an image of
improved? dissatisfaction? impersonalization is a problem
Descriptive research
Older people buy our brand, the
What should be the target What kinds of people now buy the
young married are heavy users of
segment? product, and who buys our brand?
competitors’
How should our product be We are regarded as being
What is our current image?
changed? conservative and behind the times
Causal research
An increase of 50 percent or less
Will an increase in the What is the relationship between
will generate marginal revenue in
service staff be profitable? size of service staff and revenue?
excess of marginal costs
Which advertising program
What would get people out of cars Advertising program A generates
for public transit should be
into public transit? more new riders than program B
run?

MA2 AdMR 140205 I Page 20


Questions for review and critical thinking

There are three types of marketing research: exploratory, descriptive, and causal reseach. Indicate
which type each item in the list below illustrates. Explain your answers.

Research task Type of marketing research


Establishing the relationship between
advertising and sales in the beer industry
Identifying target market demographics for a
shopping center located in Omaha, Nebraska
Estimating the 5-year sales potential for Cat-
Scan machines in the United States
Testing the effect of the inside temperature of
a chlothing store on sales of outerwear
Discovering the ways that people who live in
apartments actually use vacuum cleaners,
and identifying cleaning tasks for which they
do not use a vacuum

MA2 AdMR 140205 I Page 21


Set research design and data sources

Questions for review and critical thinking

Airways Luggage is a producer of lightweight, cloth-covered luggage. The company distributes its
luggage through major department stores, mail-order houses, clothing retailers, and other retail
outlets, such as stationary stores and leather-goods stores. The company advertises rather
heavily, and supplements this promotional effort with a large field staff of sales reps, numbering
around 400. There is a large turnover in sales reps (10%-20% per year). Because the cost of
training a new person is estimated at $15.000 to $20.000, not including the lost sales that might
result because of a personal switch, Ms. Books had been conducting exit interviews with each
departing sales rep. On the basis of these interviews, she has formed the opinion that the major
reason is dissatisfaction with promotional opportunities and pay. But top management has not
been sympathetic to Ms. Books’ pleas regarding the changes needed in these areas of corporate
policy, saying her suggestions are based on intuition and little hard data. Ms. Books had decided to
call on their in-house Marketing Research Department.
Identify the general hypothesis that would guide your research efforts.
What type of research design would you recommend to Ms. Books? Why?

MA2 AdMR 140205 I Page 22


Depending on the category of research, specific
data collection methods are suitable
Set research design and data sources

Relationship between data collection method and category of research

Category of research
Data collection method Exploratory Descriptive Causal
Secondary sources
Information systems
Databanks of other organizations
Syndicated services
Primary sources
Qualitative research
Surveys
Experiments

Appropriate method Somewhat appropriate method

In practice, it is rare for a major research project to rely


on just one of these data collection methods

MA2 AdMR 140205 I Page 23


There is a wide variety of secondary data
sources available
Set research design and data sources

Types of secondary sources

Secondary data

Internal sources External sources

Commercial Published

Sales invoice Geo-demographic Directories


Loyalty / CRM data Periodicals
records Diary panel data Statistical sources
Salesperson’s call Store audit data Financial records
reports Scanner data …
Salesperson’s Advertising exposure
expense accounts data
Warranty cards …

MA2 AdMR 140205 I Page 24


Set research design and data sources

Questions for review and critical thinking

There are a lot of useful sources of internal secondary data. For each of the following documents,
indicate what kind of information it provides.
Document Information provided

Sales invoice

Cash register receipt

Salesperson’s call
report
Salesperson’s
expense account
Individual customer
record

Financial record

Credit memo

Warranty card

MA2 AdMR 140205 I Page 25


Set research design and data sources

Questions for review and critical thinking

A large chain of building supply yards was aiming to grow at a rate of three yards per year. From
past experience, this meant carefully reviewing as many as 20 to 30 possible locations. You have
been assigned the task of making this process more systematic. The first step is to specify the
types of secondary information that should be available for the market area of each location. The
second step is to identify the possible sources of this information and appraise their usefulness.
From studies of the patrons of the present yard, you know that 60 percent of the dollar volume is
accounted for by building contractors and tradesmen. The rest of he volume is sold to farmers,
householders, and hobbyists. However, the sales to do-it-yourselfers have been noticeably
increasing. About 75 percent of the sales were lumber and building materials, although appliances,
garden supplies, and home entertainment systems are expected to grow in importance.

MA2 AdMR 140205 I Page 26


Qualitative research techniques don’t follow a
rigid format
Set research design and data sources

Characteristics of qualitative techniques

Small convenience or quota samples are often used, rather than rigorous,
statistically meaningful samples
The information sought relates to respondents’ motivations, beliefs, feelings, and
attitudes, not to facts about their lives and behavior
An intuitive, subjective approach is used to gathering the data
The data collection format is open ended, allowing respondents to express
themselves in their own words at a length they deem appropriate
The approach is not intended to provide statistically or scientifically accurate data,
but often to guide further investigation that will

MA2 AdMR 140205 I Page 27


Depth interview and focus group are important
qualitative techniques
Set research design and data sources

Depth interview Focus group

Unstructured personal interview that uses Loosely structured interactive discussion by


extensive probing to encourage a single a trained moderator among a small group of
respondent to talk freely in order to express respondents based on an interviewer guide
detailed beliefs and feelings on a topic Main purposes are to
Purpose is to transcend the respondent’s − help generate hypotheses,
surface reactions, discovering more − provide overall background information
fundamental reasons underlying the on a product category,
respondent’s attitudes and behavior − generate ideas for new concepts and
It is useful when − interpret previously obtained quantitative
results
− the marketing problem relates to
particularly confidential, sensitive, or The value lies in the potential to discover
potentially embarrassing issues, or when the unexpected, which is far more possible
− group pressure or norms would affect the in a free-flowing group discussion than in a
responses typical survey setting

MA2 AdMR 140205 I Page 28


Quantitative research focuses on past behavior,
attitudes, and respondents characteristics
Set research design and data sources

Main categories
of primary data

Evidence regarding the respondent’s past behavior has wide


usage as a predictor of future behavior
Past behavior Understanding past behavior involves many dimensions;
researchers must be sensitive to the key behavioral
dimensions relevant to predicting future behavior

Attitudes are important in marketing because of the presumed


relationship between attitudes and behavior
Attitudes Simply put, we do not just do things but do them for certain
reasons, reasons that reflect our underlying attitudes and
intentions

Respondents characteristics refer to descriptions of


respondents along particular dimensions of interest, including
Respondents demographic, socioeconomic, and psychological variables
characteristics (e.g., respondents’ activities, interests, and opinions)
For many products, such variables are known to correlate well
with eventual purchase behavior

MA2 AdMR 140205 I Page 29


For quantitative research, two basic methods of
collecting data exist
Set research design and data sources

Communication Observation

The communication method is based on the Observation involves the process of


direct questioning of respondents, typically recognizing and recording the behavior of
based on a questionnaire objects
There are four types of communication In practice, observational techniques are
approaches: face-to-face, telephone, mail, used in conjunction with other data
and internet-based interview collection techniques
Main advantages of the communication Main advantages of the observation method
method are are
− the ability of the method to collect data on − that there is no substantial potential bias
a range of information needs, caused by the interviewing process,
− speed and cost of communication, and − that certain types of data can be collected
− a better control over the data collection only by observation, and
process − that it does not rely on the respondent’s
explicit willingness to engage
Main disadvantages are
− the respondents unwillingness to provide Main disadvantages are
the desired data, − that only things leading to a physical or
− the respondents inability to provide the verbal manifestation can be observed,
data, and − that due to cost and time requirements,
− the influence of the questioning process observation needs to be of short duration,
on the responses and
− substantial ethical issues in the use of
any sort of observation

MA2 AdMR 140205 I Page 30


Set research design and data sources

Questions for review and critical thinking

Quick-Stop Inc. recently opened a new convenience store in Northglenn, Colorado. The store is
open every day from 07:00 a.m. to 11:00 p.m. In order to better plan the location of other units in
the Denver metro area, management is interested in determining the trading area from which this
store draws its customers. How would you determine this information by questionnaire? By
observation method? Which method would be preferred? Be sure to specify how you would define
“trading area”.

MA2 AdMR 140205 I Page 31


Set research design and data sources

Questions for review and critical thinking

The Metal Products Division of Geni Ltd. devised a special metal container to store plastic garbage
bags. Plastic bags posed household problems, as they gave off unpleasant odors, looked
disorderly, and provided a breeding place for insects. The container overcame these problems, as
it had a bag-support apparatus that held the bag open for filing and sealed the bag when the lid
was closed. In addition, the storage area held at least four full bags. The product was priced at $
59,99 and was sold through hardware stores. The company has done little advertising and has
relied on in-store promotion and display. The divisional manager was wondering about the
effectiveness of these displays and has called on you to do the necessary research. Should the
communication or observational method be used? Justify your choice.

MA2 AdMR 140205 I Page 32


Questionnaire design is roughly equal parts art
form, apprenticeship, and scientific undertaking
Set research design and data sources

Challenge of designing data collection instruments

Designing forms for primary data collection is a key component of most studies
It is crucial to control measurement error, as it is often the largest source of
preventable error in marketing research, far exceeding errors in sampling plans or
administration
− The function of a questionnaire is measurement in a formalized way
− If a question does not truly measure what it is supposed to, measurement error
is present
The only way to develop the skill of questionnaire design is to write a
questionnaire, use it in a series of interviews, analyze its weaknesses, revise it,
and then repeat this process for new surveys
− No steps, principles, or guidelines can absolutely guarantee an effective and
efficient questionnaire
− Questionnaire design is a skill that the researcher learns through experience,
after some basic principles are assimilated

Best practice guidelines are useful in avoiding serious errors; fine-tuning


questionnaire design requires the creative inspiration of the skilled researcher

MA2 AdMR 140205 I Page 33


The data collection procedures link information
needs and questions to be asked

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design must
Researchers sampledevelop a data-collection procedure that establishes an effective
link between the information needs and the question to be asked or the observations
Collect data
to be recorded. The success of the study is dependent on the researchers’ skill and
creativity in establishing this link.
Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 34


While developing the data collection procedure,
measurement issues are a constant challenge
Data collection procedure

Basic concept of measurement

The process of measurement is a fundamental component of essentially all


marketing research activities
The measurement process assigns numbers to represent each of the marketing
phenomena under investigation (objects, characteristics, or events) according to a
set of known, specific rules
There are four characteristics of the number system that could (or could not) hold
true for a marketing phenomenon under investigation as well as the assigned
number to represent the phenomenon
− Identity (1 = 1)
− Order (0 < 1 < 2)
− Equality of differences (2-1 = 7-6)
− Equality of ratios (2/4 = 4/8)
A characteristic that holds true for a marketing phenomenon does not necessarily
hold true for the number assigned to represent this phenomenon

MA2 AdMR 140205 I Page 35


There are four types of measurement scales

Data collection procedure

Scale Properties Marketing phenomena (examples)

Brands in a category
Unique definition of
Store types (e.g., supermarket, drug store)
Nominal numerical labels
Sales territories
(0, 1, 2, …)
Demographics (e.g., gender, occupation)

Preferences
Order of numerals Subjective frequencies
Ordinal
(0 < 1 < 2 …) Purchase likelihood
Demographics (e.g., age group, education)

Equality of Attitudes
Interval differences Opinions
(2-1 = 7-6) Index numbers

Objective frequencies
Equality of ratios Costs
Ratio
(2/4 = 4/8) Number of customers
Sales (units, value)

MA2 AdMR 140205 I Page 36


A nominal scale is one on which numbers serve
only as labels to categorize objects
Data collection procedure

Main characteristics

Nominal scales are used for Appropriate statistics


classification and identification
Percentages
− E.g., use of numbers to identify
members of a sports team, ZIP Logistic regression
codes to identify geographic areas Mode
The measurement rule is simple: do Cross tabulation
not assign the same number to
Chi square test
different attributes or different
numbers to the same attribute

A good informal test for whether a scale is nominal is to ask whether


anything important would be affected if two categories simply swapped their
numbers; if not, the scale is almost certainly nominal

MA2 AdMR 140205 I Page 37


An ordinal scale defines the ordered relationship
among objects
Data collection procedure

Main characteristics

Ordinal scales measure whether an


object has more or less of a
characteristic than some other Appropriate statistics1)
object, but they do not provide
information on how much more or Percentiles
less
Median
That is, it offers only information on
the order, not the degree, of Rank order correlation
differences Ordinal regression
With an ordinal scale, all increasing
strictly monotonic transformations
(i.e., X2 > X1 => Y2 > Y1) are
permissible transformations

Although it is fairly common to see ordinal data represented by means, this is


statistically not justified and should only be done in those rare cases when
there is explicit evidence that the ordinal scale is close to an interval one

1) Statistics appropriate for nominal scale also apply to the ordinal scale

MA2 AdMR 140205 I Page 38


An interval scale ranks objects and secures the
equality of difference
Data collection procedure

Main characteristics

Ordinal scales rank objects such


that the distance between the
numerals correspond to the distance Appropriate statistics1)
between the objects on the
characteristic being measured Range
− I.e., it is up to the researcher to Arithmetic mean
choose spacings and the zero Standard deviation
point
Linear regression
With an interval scale, all linear
transformations (i.e., Y = a + bX with
b > 0) are permissible
transformations

In marketing, it is common for attitudinal, opinion, and predisposition judgment to be


treated as interval data; to be technically correct, these judgments are ordinal

1) Statistics appropriate for nominal and ordinal scales also apply to the interval scale

MA2 AdMR 140205 I Page 39


Data collection procedure

Questions for review and critical thinking

Assume that we have scaled brands A, B, and C on an interval scale regarding buyers’ degree of
liking of the brands. (Liking is typically assumed to be measured on an ordinal scale. For the sake
of analyzing what is and is not meant by an interval scale, assume that the liking-based scale
example is reasonable.) Brand A receives a 6, the highest liking score, B receives a 3, and C
receives a 2. What can be said about these interval-scaled data?
The liking for brand B is less favorable than that for brand A
Brand A is liked twice as much as brand B
The degree of liking between A and B is three times greater than the liking between B and C

MA2 AdMR 140205 I Page 40


With a ratio scale, both differences and
quotients of scale points make sense
Data collection procedure

Main characteristics

A ratio scale has all the properties of Appropriate statistics1)


an interval scale plus an absolute
(or natural) zero point Geometric mean
With ratio measurement, only one Harmonic mean
number may be assigned arbitrarily,
namely, the unit of measurement or Coefficient of variation
distance Ratios and products of variables
With a ratio scale, all proportional (interactions)
transformations (i.e., Y = cX with c >
0) are permissible transformations

Many important marketing phenomena possess the properties of a ratio scale,


e.g., sales, market share, costs, number of customers

1) Statistics appropriate for nominal, ordinal and interval scales also apply to the ratio scale

MA2 AdMR 140205 I Page 41


Data collection procedure

Questions for review and critical thinking

Identify the type of scale being used in each of the following questions. Justify your answer.
(1) During which season of the year were you born? (winter / spring / summer / fall)
(2) What is your total household income?
(3) Which are your three most preferred candy bars? Rank them from 1 to 3 according to your
preference, with 1 as most preferred! (M&Ms plain / M&Ms peanut / Reese’s / Almond Joy /
Good & Plenty)
(4) How much time do you spend traveling to school every day? (under 5 minutes / 5-10 minutes /
11-15 minutes / 16-20 minutes / 21 minutes or more)
(5) How satisfied are you with Newsweek magazine? (very satisfied / satisfied / neither satisfied
nor dissatisfied / dissatisfied / very dissatisfied)
(6) On an average, how many beers do you drink in a day? (more than 3 / 2 to 3 / less than 2)
(7) Which one of the following courses have you taken? (marketing research / sales management /
advertising management / consumer behavior)
(8) What is the level of education for the head of the household? (some high school / high school
graduate / some college / college graduate and/or graduate work)

MA2 AdMR 140205 I Page 42


Data collection procedure

Questions for review and critical thinking

The analysis for each of the questions on the previous page is given below. Is the analysis
appropriate for the scale used? Is the conclusion appropriate?
(1) About 50% of the sample were born in the fall, 25% of the sample were born in spring, and the
remaining 25% were born in the winter. It can be concluded that the fall is twice as popular as
the spring and the summer seasons.
(2) The average income is $25.000. There are twice as many individuals with an income of less
than $9.999 than individuals with an income of $40.000 and over.
(3) M&Ms plain is the most preferred brand. The mean preference is 3,52.
(4) The median time spent traveling to school is 8,5 minutes. Three times as many respondents
travel fewer than 5 minutes than respondents traveling 16-20 minutes.
(5) The average satisfaction score is 4,5, which seems to indicate a high level of satisfaction with
Newsweek magazine.
(6) 10% of the respondents drink less than two bottles of beer a day, whereas three times as many
respondents drink over three bottles a day.
(7) Sales management is the most frequently taken course because the median is 3,2
(8) The responses indicate that 40% of the sample has some high school education, 25% of the
sample are high school graduates, 20% have some college education, and 10% are college
graduates. The mean education level is 2,6.

MA2 AdMR 140205 I Page 43


Typically, measurements possess some degree
of error
Data collection procedure

Common potential sources of error

Respondents characteristics
Personal factors such as mood, fatigue, willingness to participate, and even health
at the time of administration
Situational factors
Variations in the environment in which the measurements are reached, such as
time of the day, temperature, presence of family members, and point-of-purchase
marketing activities
Data collection factors
Variations in how the questions are administered, including the influence of the
interviewing method (e.g., phone, personal contact, web-based, or mail)
Measuring instrument factors
The degree of ambiguity and difficulty of the question and the ability of the
respondents to answer them
Data analysis factors
Errors made in the coding and tabulations process

MA2 AdMR 140205 I Page 44


The total error of measurement consists of two
components: systematic error and random error
Data collection procedure

Systematic error Random error

Error that causes a constant bias in Error implied by influences that bias
the measurements: it affects the measurement but are not systematic:
measurement in a predictable way it produces inconsistency in repeated
− Measuring time with a stopwatch measures
that runs 10% faster than it should − Measuring time with several
− Measuring height with a poorly stopwatches that run without
calibrated wooden yardstick systematic error
− Measuring height with an elastic
ruler

Observed score = true score + systematic error + random error

Reliability refers to the extent to which the measurement is free from random error,
validity to the extent it is free from both systematic and random error

MA2 AdMR 140205 I Page 45


Reliability can be assessed based on three
methods
Data collection procedure

Test-retest reliability Alternative form reliability Split-half reliability

The results of repeated The results of two forms A multi-item measurement


measurements of the same given to the subject and device is divided into equi-
person / group using the judged equivalent (by prior valent groups and the item
same scaling device under testing or a panel of experts) responses are correlated to
conditions judged to be but not identical are estimate reliability
similar are compared compared The greater the discrepancy
The greater the discrepancy The greater the discrepancy in scores (i.e., the lower
in scores, the greater the in scores, the greater the correlation coefficients), the
random error present in the random error and thus the greater the random error
measurement and thus the lower the reliability and thus the lower the
lower the reliability reliability
Limitations are
Limitations are that − the expense and delay Limitations are
− it may not be logical / associated with − that possible interdepen-
possible to administer the developing a second dencies between items
measurement twice with measuring instrument, are not assessed, and
the same subject, and − the difficulty of making
− the first measurement my − the difficulty of making the two groups equivalent
change the subject’s the two instruments An even better way to
response to the second equivalent
assess the internal homo-
measurement, and
geneity of a set of items it to
− situational factors may
use coefficient alpha
change

MA2 AdMR 140205 I Page 46


Four major methods allow to infer the validity of
a measurement
Data collection procedure

Construct validity Content validity (face validity)

The construct of interest is related to other Content validity involves subjective


constructs in order to develop an entire judgments by experts as to appropriateness
theoretical framework for the marketing of the measurement
phenomenon being measures This is a common method used in marketing
Construct validity increases as the research to determine the validity of a
correlation between the construct of interest measurement
and the related constructs increases in the
predicted manner

Concurrent validity Predictive validity

Two different measurements of the same Predictive validity involves the ability of a
marketing phenomenon administered at the measured marketing phenomenon at one
same point in time are correlated point in time to predict another marketing
Concurrent validity is primarily used to phenomenon at a future point
determine the validity of new measuring If the correlation between the two measures
techniques by correlating them with is high, the initial measure is said to have
established ones, which have presumably predictive validity
performed well in the past

MA2 AdMR 140205 I Page 47


Data collection procedure

Questions for review and critical thinking

Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 177-179, case
1.7 Hepworth Golden Auto.
(1) What developments in HGA performance led Bill Douglass to ask Rightway Research’s help?
What is the decision-making process that Mr. Douglass expects to be facilitated with the
insights provided by the research findings?
(2) Evaluate the stated research objective. Does it thoroughly address HGA’s core problems?
(3) Evaluate the information needs stated in the proposal. Have these been correctly identified?
Should any be altered, deleted, or added?
(4) Consider the data sources listed in the proposal. Are these sources appropriate for the stated
information needs and research objective? Why or why not? What other data sources could be
used? Consider accessibility, accuracy and cost in making source suggestions.
(5) Evaluate the data collection framework.
(6) What other possible research design and data sources could be used? How would these fit the
information needs relative to the given design?
(7) Considering the research proposal, do you think Rightway Research truly understood why the
information was needed, and for what eventual purpose? Explain your answer.
(8) Building on your previous answers, prepare a new research proposal for the HGA research
project.

MA2 AdMR 140205 I Page 48


Virtually every marketing research study
requires the selection of some kind of sample

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect need
Researchers data to clearly define who or what is to be included in the sample, the
population from which the sample is to be drawn, the methods used to select the
Process and code data
sample, and the sample size.
Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 49


Sampling offers some major benefits over taking
a census
Design sample

Benefits of sampling

A sample saves money


A sample saves time
A sample may be more accurate
A sample is better in situations where the study itself entails contamination or even
destruction of the sampled elements

If speed, cost, and efficiency of effort are not vitally important in a specific project,
neither is sampling; but such situations are exceedingly rare

MA2 AdMR 140205 I Page 50


Selecting a sample involves five steps

Design sample

Steps of the research process

Establish need for information


Steps involved in selecting a sample
Detail research objectives and information needs
Define the population
Set research design and data sources
Identify the sampling frame
Design data collection procedure
Decide on the sample size
Design sample
Select a specific procedure
Collect data
Physically select the sample
Process and code data

Analyze and interpret data

Present results and conclusions

MA2 AdMR 140205 I Page 51


Two categories of sampling procedures exist:
non-probability and probability sampling
Design sample

Probability sampling Non-probability sampling

In probability sampling, each element In non-probability sampling, the


of the population has a known chance selection of a population element to
of being selected for the sample be part of the sample is based in
Probability sampling allows us to some part on the judgment of the
calculate the likely extent to which the researcher or field interviewer
sample value differs from the As there is no known chance of any
population value of interest (sampling particular element in the population
error) being selected, the sampling error
Probability sampling procedures can not be calculated
include Non-probability sampling procedures
− simple random sampling include
− stratified sampling − convenience sampling
− cluster sampling − judgment sampling
o systematic sampling − quota sampling
o area sampling

MA2 AdMR 140205 I Page 52


The stratified sampling procedure tends to entail
smaller variance than unstratified ones
Design sample

Process of setting up and benefits of a stratified sample

There are two steps in setting up a stratified sample


− Divide the defined population into mutually exclusive and collectively exhaustive
strata
− Select an independent simple random sample from each strata
Key is that the designed strata are more homogeneous (less varying) on the
variable(s) for which we are calculating our statistics as the population
− Therefore, first we must answer the question, “what factors contribute to the
variability in the quantity we intended to measure?”
As long as the strata are more homogeneous, stratified sampling results in a
decrease in the standard error of an estimator; thus,
− the confidence interval we calculate will be smaller, we will have
− more statistical power, and / or we can
− accomplish the same degree of confidence with a smaller sample, saving time,
funds, and energy
For a fixed sample size, we can reduce the overall standard error of the estimate
further by sampling more heavily in strata with higher variability (disproportionate
stratified sampling)

MA2 AdMR 140205 I Page 53


Selecting a systematic sample is easy and
inexpensive
Design sample

Process of setting up and benefits of a cluster sample

In cluster sampling, a cluster or group of elements is randomly selected at one


time
− Before we can select a cluster sample, the population must be divided into
mutually exclusive and collectively exhaustive groups
− The elements in one cluster should be as close, in terms of overall
heterogeneity on the variables of interest, as the population as a whole
Out of the selected groups, either all elements might be used (one-stage cluster
sampling) or a random sample of elements from within the selected groups might
be selected (two-stage cluster sampling)
Cluster samples are in most cases less statistically efficient than simple random
samples; for reasons of cost, however, cluster samples are used extensively
− Sample clusters are usually (much) less heterogeneous than the population

In stratified sampling we want homogeneous groups,


whereas in cluster sampling we want heterogeneous groups

MA2 AdMR 140205 I Page 54


A researcher must balance the requirements for
precise measurement and the costs of doing so
Design sample

Sampling practice

Based on a random sample, a company calculates the


average age of its customers as 31,7 years. Depending on
the sample size, a 95% confidence interval looks as
Sampling theory follows:
Average age
Based on statistical theory, 33,0
we can determine the
“optimal” sample size given
− a specified level of
precision, 32,0
− the required level of 31,7
confidence and
− the sample standard 31,0
deviation

30,0
10 20 50 100 1.000 10.000
Sample size

A fundamental error lies in the belief that large samples imply less error and therefore better re-
sults; although this is true for sampling errors, it is most certainly not true for non-sampling errors

MA2 AdMR 140205 I Page 55


Design sample

Questions for review and critical thinking

Pete Thames, the general manager of the Winona Wildcats, a minor league baseball team, is
concerned about the declining level of attendance the team’s games in the past two seasons. He is
unsure whether the decline is due to a national decrease in the popularity of baseball or the factors
that are specific to the Wildcats. Having worked in the marketing research department of the
team’s major league affiliate, the New Jersey Lights, Thames is prepared to conduct a study on the
subject. However, because it is a small organization, limited financial resources are available for
the project. Fortunately, the Wildcats have a large group of volunteers who can be used to
implement the survey. The study’s primary objective is to discover the reasons why Winona
residents are not attending games. A list of 1.200 names, which includes all attenders for the past
two seasons, is available as a mailing list.
What is the target population for this study?
What is the appropriate sampling frame for this study of the attitudes of both attenders and
nonattenders?
Which kind of sample would provide the most efficient sampling?
Why would this method of sampling be the most efficient in this situation?

MA2 AdMR 140205 I Page 56


Design sample

Questions for review and critical thinking

Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 377-379, case
2.7 Delta Dairy local market taste test survey.
(1) Define the population and the sampling frame for this survey project.
(2) What sampling procedure would you recommend for this taste test survey? Why? What potential
deficiencies do you foresee with your suggested procedure?
(3) One alternative for how to conduct the survey was to place taste tables on the most crowded
streets of the “capital” of northern Greece (Thessaloniki) and select passersby who would be
willing to participate in the survey. Evaluate this alternative and discuss its advantages and
disadvantages. What might be a better method, from the point of representativeness, bias, and
other principles of sampling?
(4) Should the taste test be blind or nonblind? Why? Should explicit brand names be tested? Would
it be a good idea for Delta to present information about the various milks being tasted, or even
mock-up of containers in which it will be sold? Should pricing information be tested
concurrently?
(5) Considering the survey’s objective and the time and budget constraints, what milk samples
would you include in the taste test? Would you vary them across different respondents? How?
Might you allow consumers to simply try whatever milk products they like best?
(6) Is it important to study different regions of the country, or different regions within northern
Greece? Should certain segments be deliberately overrepresented in the sample? Should any
type of stratification be imposed? Finally, what sort of biases might arise from simply using a
random sample – assuming one can be accessed – of the northern Greek population?

MA2 AdMR 140205 I Page 57


Collecting data is critical as to the total error in
the research results

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect data

Process
Collecting dataand code involves
typically data a large proportion of the research budget and a
sizeable proportion of the total error in the research results. Consequently, selection,
Analyze
training, and interpret
and control data are essential to effective marketing research
of interviewers
studies.
Present results and conclusions

MA2 AdMR 140205 I Page 58


Data processing includes the functions of editing
and coding

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect data

Process and code data

EditingAnalyze
involvesand interpretthe
reviewing data
data forms as to legibility, consistency, and
completeness. Coding involves establishing categories for responses or groups of
Present
responses results
so that and conclusions
numerals can be used to represent the categories.

MA2 AdMR 140205 I Page 59


Statistical methods help to understand what is
going on under all the noise in the real world

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect data

Process and code data

Analyze and interpret data

Present must
Data analysis results
beand conclusions
consistent with the requirements of the detailed information
needs identified at the beginning of the project. Analysis is usually performed by
using appropriate statistical software packages.

MA2 AdMR 140205 I Page 60


Selection of a suitable statistical method is
influenced by three overall research goals
Analyze and interpret data

Research objective / information need

Number of variables to analyze Description vs. inference Level of measurement

How many variables do we Are we interested in Have the variables been


wish to interrelate: one (uni-), describing the sample or in measured at a nominal,
two (bi-), or multiple making inferences about the ordinal, interval, or ratio scale
(multivariate analysis)? population? level?

Data analysis technique to be used

Matching the scale level(s) is not merely a technical matter:


if one asks the computer to, for example, estimate an ordinary regression on categorical
(nominal) data, it will dutifully perform the analysis, but the results will be meaningless

MA2 AdMR 140205 I Page 61


Analyze and interpret data

Questions for review and critical thinking

For each of the following research questions indicate (1) the number of variables to analyze, (2) if
it’s descriptive or inferential statistics, as well as (3) the scale level.
What is the level of association between age (in years) and weight (in kg) in the sample?
What is the impact of height (in cm), age (in years) and gender (male / female) on weight (in kg)
in the sample?
Is the average age (in years) of the customer base 25?
What is the average age (in years) in the sample?
Is the level of association between age (0-18, 19-45, 46-65, >65) and weight (0-50, 51-70, 71-
90, >90) in the population greater than zero?

MA2 AdMR 140205 I Page 62


Analyze and interpret data

Examples of univariate analyses abound in marketing

The manager might want a description of


the sample’s demographic characteristics,
the usage of the company’s product (sales volume), or
respondents attitudes toward a competitive activity

MA2 AdMR 140205 I Page 63


A first step in almost all marketing research
projects is to perform univariate analyses
Analyze and interpret data

Univariate data analysis procedures

Nominal data Ordinal data Interval data

Descriptive Descriptive Descriptive


− Central tendency: mode − Central tendency: − Central tendency: mean
− Dispersion: relative and median − Dispersion: standard
absolute frequencies by − Dispersion: interquartile deviation
category range
Inferential Inferential Inferential
− Chi-square test − Kolmogorov-Smirnov − z-test
test
− t-test

Univariate data analysis procedures allow to get a “feel” for the data and are a good
opportunity to weed out coding errors, outliers and other potential problems

MA2 AdMR 140205 I Page 64


Mode, median and mean are measures of central
tendency
Analyze and interpret data

Mode Median Mean

The mode is the category The median is defined as The mean is the sum of the
of a nominal variable that the middle value when the values divided by the
occurs most often data are arranged in order sample size
The mode should not be of magnitude; that is, half The mean is – in contrast
applied to ordinal or the values fall above the to the median – not robust
interval data unless these median and half fall below1)
data have been grouped The median is an
first especially useful measure
when the data contain
outliers (= values well
outside the range of most
of the data), as it resists
dramatic change in the
presence of outliers; i.e.,
the median is robust

1) When there are an even number of data points, the median is calculated by taking the midpoint of the two middle values

MA2 AdMR 140205 I Page 65


Analyze and interpret data

Application of the statistical methods

The data in the file MA2 AdMR Data file 1 [Link] stems from a comparison of 17 different beer
brands.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)!
Calculate the appropriate measures for the central tendency. Interpret the results!
The product manager for Schmidts decided to reposition his beers and to enter the premium
segment. While keeping the recipe unchanged, the price was changed to 1,30.
Change the data file accordingly and re-calculate the appropriate measures for the central
tendency. Interpret the findings!

MA2 AdMR 140205 I Page 66


Frequencies and standard deviation help to sense
the distribution of a variable
Analyze and interpret data

Relative and absolute frequencies Standard deviation

Absolute frequencies are the number of The variance is the sum of the squared
items in the sample in each category; deviations of the data (observations) from
relative frequencies present the data as the sample mean divided by the degrees of
proportions freedom
Frequencies are typically presented in a − Squaring accomplishes the desirable
histogram (bar graph of the relevant cell feature to have both positive and
counts) negative deviations count equally
− The degree of freedom equals the
number of independent observations (on
the variable of interest) minus the number
of statistics calculated (from the same
data) used in any particular formula
The standard deviation is the square root of
the variance and measures the dispersion of
the data around the (sample) mean

MA2 AdMR 140205 I Page 67


Analyze and interpret data

Application of the statistical methods

Due to a significant decline in profits, the price of Schmidts was changed back to 0,30.
As far as appropriate, calculate for the variables the variance and the standard deviation and
interpret the results. What is the main challenge in interpreting the figures?
Copy the data of costs into an Excel file (column A) and calculate the mean (column B). Subtract
the mean from each data point (column C) and sum these differences up (column D). Interpret
the sum! Square each difference from column C (column E) and sum these squared differences
up. Interpret the sum! Divide the sum by the degrees of freedom and compare the result with the
corresponding value in the SPSS output.
Calculate absolute and relative frequencies including a histogram for the variable calories and
interpret the results.
Re-do the calculation superimposing the histogram with a normal distribution (with the same
mean and standard deviation as the variable calories). Interpret the results

MA2 AdMR 140205 I Page 68


Hypothesis testing is the basis for making
inferences from the sample to the population
Analyze and interpret data

The concept of a null hypothesis

The null hypothesis H0 states that a population parameter takes on a particular value
By contrast, there is an alternative hypothesis H1, and this is what the researcher is attempting
to verify
− There are three types of alternative hypotheses: the population parameter (e.g., µ)
(1) does not take a specific value (e.g., 25): H0: µ = 25 and H1: µ ≠ 25
(2) is greater than that value: H0: µ ≤ 25 and H1: µ > 25
(3) is less than that value: H0: µ ≥ 25 and H1: µ < 25
− Whenever the alternative hypothesis is directional (no 2 and 3), a one-tailed test applies,
because only large deviations in one direction (large positive deviations in case of no 2, large
negative deviations in case of no 3) would count against H0 in favor of H1
− Otherwise (no 1), a two-tailed test applies, i.e. either large or small values of the sample
statistic would cause the researcher to reject H0
Depending on the empirical evidence, we will be able to say “we reject H0 in favor of H1”
− We will never accept H0, but merely fail to reject it
− In a situation where we fail to reject H0, we do not conclude that H0 is valid; all we can say is
that we do not have suitable evidence to reject it

MA2 AdMR 140205 I Page 69


In hypothesis testing, two types of error might
occur
Analyze and interpret data

True condition

H0 is true H0 is false

Correct decision Type II error Freeing


Do not reject H0 Confidence level the guilty
Probability = 1 - α Probability = β
Sample
conclusion
Type I error Correct decision
Reject H0 Significance level Power of the test
Probability = α Probability = 1 - β

Convicting the
innocent

MA2 AdMR 140205 I Page 70


Analyze and interpret data

Questions for review and critical thinking

Suppose you have to choose the best promotional strategy for a new product introduction from
among three possible strategies. The null hypothesis is that all the strategies are equally effective.
In business terms, what does the type I error mean in this specific case? Evaluate the level or
risk associated with a type I error.
In business terms, what does the type II error mean in this specific case? Evaluate the level of
risk associated with a type II error.

MA2 AdMR 140205 I Page 71


The z- and the t-test are appropriate for tests for
the population mean based on interval data
Analyze and interpret data

Sample size is large (n > 30)

Yes No

Yes
Population
z-test or t-test
standard
deviation σ
is known
No t-test

MA2 AdMR 140205 I Page 72


The α error reflects the critical region of the
distribution
Analyze and interpret data

Critical region for the t distribution

Two-tailed test

α/2 α/2

µ0
H0: µ = 25
H1: µ ≠ 25

One-tailed test One-tailed test

α α

µ0 µ0
H0: µ ≤ 25 H0: µ ≥ 25
H1: µ > 25 H1: µ < 25

MA2 AdMR 140205 I Page 73


The rule when to reject H0 depends on the type
of the hypotheses
Analyze and interpret data

H0: µ = µ0 H0: µ ≥ µ0 H0: µ ≤ µ0


H1: µ ≠ µ0 H1: µ < µ0 H1: µ > µ0

Reject H0 in favor of H1 Reject H0 in favor of H1 Reject H0 in favor of H1


in case of in case of in case of
I T I > t α/2, df T < -t α, df T > t α, df
with with with
T = empirical t-value1) T = empirical t-value1) T = empirical t-value1)
t = theoretical t-value2) t = theoretical t-value2) t = theoretical t-value2)

1) Calculated based on the sample characteristics


2) Taken from a table given a specific significance level and the degrees of freedom

MA2 AdMR 140205 I Page 74


Analyze and interpret data

Application of the statistical methods

The data in the file MA2 AdMR Data file 2 Customer [Link] stems from 372 customers of
three different brands which were asked about their satisfaction with the respective brand.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)! Variables with “none” in column “values” were
measured on a scale from 1 to 10. Afterwards, the values were recalculated according to the
formula ((x – 1) * (100 / 9)), meaning they now range from 0 to 100.
Select brand A and calculate the appropriate measures for the central tendency as well as the
dispersion (incl. histogram) for the variables Sat_Ov, Fulfill and Closeness. Interpret the results!
From last years study, a value of 78 for the overall satisfaction is still in the mind of the CEO. He
asks you whether this years results will allow the conclusion that the population overall satisfaction
is greater than 78.
Formulate the corresponding null and alternative hypothesis!
Perform a t-test (significance level of 5%). Interpret the results.
Compare the empirical t-value (SPSS output) to the theoretical t-value and interpret the results.
How does the interrelation between the theoretical t-value and the significance level α look like?

MA2 AdMR 140205 I Page 75


Analyze and interpret data

Application of the statistical methods

The companies target for the running year was to achieve a rating of 81 for Closeness. By
achieving the target, the employees would be eligible to a bonus. The CFO claims, that the target
was missed.
Formulate the corresponding null and alternative hypothesis!
Perform a t-test (significance level of 5%). Interpret the results.
Compare the empirical t-value (SPSS output) to the theoretical t-value and interpret the results.

MA2 AdMR 140205 I Page 76


The chi-square test compares a hypothesized
against an observed distribution
Analyze and interpret data

The concept of the chi-square test

Researchers often need to make inferences about how respondents are distributed across the
possible categories of a nominal variable
− E.g., has the consumer base changed in terms of geographic location (countries, regions,
states, ZIP codes, …), gender split, educational entertainment, or income categories?
The chi-square test is a procedure for comparing a hypothesized population distribution (across
nominal chi-square test categories) against an observed distribution
In that way, it is just like all hypothesis tests we have seen thus far: it compares what we
hypothesize in a population with what we observe in a sample

MA2 AdMR 140205 I Page 77


Analyze and interpret data

Application of the statistical methods

Customers of brand A participating in the customer satisfaction study were asked about their trade
(customer segment) and the number of their employees.
Calculate frequency tables for each of the two variables (Trade, No_employ) and interpret the
results!
Based on the frequency tables you assume that the categories of Trade and No_employ
respectively do not occur with equal probabilities.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.
Before conducting the study, the Sales Director requested that the sample represents a population
with the following characteristic: the trades 1, 2, 3, and 8 contribute 20 % / 12 % / 45 % / 12 % to
the population, the other trades contribute equally to the remaining share.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.
On top, he requested that the sample represents a population with the following characteristic: the
size classes 8, 7, and 6 contribute 30 % / 35 % / 25 % to the population, the other size classes
contribute equally to the remaining share.
Formulate the corresponding null and alternative hypothesis!
Perform a chi-square test (significance level of 5%). Interpret the results.

MA2 AdMR 140205 I Page 78


Analyze and interpret data

Examples of bivariate analyses abound in marketing

A Manager will typically be ultimately interested in relationships among variables:


What is the relationship between use of our brand and media viewing habits?
Are higher levels of sales force turnover associated with greater sales manager age?
Does the season of the year help us predict attitudes toward our brand?
In each of the cases, we are asking whether the values of one variable offer useful information
about the values of another

MA2 AdMR 140205 I Page 79


As with univariate techniques, the appropriate
bivariate technique depends on the scale level
Analyze and interpret data

Bivariate data analysis procedures

Nominal data Ordinal data Interval data

Descriptive Descriptive Descriptive


− Contingency coefficient − Rank correlation − Linear correlation
− Lambda coefficient coefficient
− Gamma − Simple regression
− Tau
Inferential Inferential Inferential
− Chi-square test − Mann-Whitney U test − t-test on regression
− Kolmogorov-Smirnov coefficient
test − z-test on the difference
between means
− t-test on the difference
between means

MA2 AdMR 140205 I Page 80


The main measure of association is the linear
correlation coefficient rxy
Analyze and interpret data

The concept of the linear correlation coefficient

In examining the relationship between two interval variables, a useful beginning is to plot the
data on a scatter diagram
− Including the mean values for each of the variables divides the scatter diagram into four
quadrants
− If the data points are concentrated mostly in diagonal quadrants, this would be evidence of a
relationship between the two variables
The covariance measures the degree to which X and Y tend to co-vary, that is, vary in the same
direction from their respective means
− Covariance is positive if the values of Xi and Yi tend to deviate from the respective means in
the same direction, and negative if they tend to deviate in the opposite direction
− If X and Y are statistically independent, cov(X,Y) will be near zero in reasonably large samples
− The reverse is not true: the covariance can be exactly zero in any sample without the variables
being statistically independent
The linear correlation coefficient rxy is just a standardized measure of covariation (the linear
relationship between X and Y)
− It does not measure all possible relationships between the two variables
− It does not make claims about X and Y causing one another

MA2 AdMR 140205 I Page 81


The linear correlation coefficient may take on any
value between -1 and +1
Analyze and interpret data

Characteristics of the linear correlation coefficient

No matter what units we choose to measure two variables, their correlation will not change (the
covariance will!)
The correlation coefficient may take on any value between -1 and +1
When rxy = 1, this indicates a perfect positive correlation; plotting the two variables in question
will show all points exactly on a straight line with positive slope (negative slope in case of rxy = -1)
If rxy = 0, there is no linear relationship between the variables; the best line through the points will
be flat (i.e., a slope of zero)

The exact percentage of variation shared by two variables is calculated by squaring rxy.
r2xy is called the coefficient of determination

MA2 AdMR 140205 I Page 82


Analyze and interpret data

Application of the statistical methods

The data in the file MA2 AdMR Data file 3 [Link] contains technical data of 406 different car
models.
What is the scale level of the different variables? Correct, when necessary, the scale level
indicated in the column “measure” (variable view)!
For the variables mpg and cu_inc, calculate the mean and produce a scatter plot (cu_inc on the
x- and mpg on the y-axis). Interpret the results! Which kind of relationship between the two
variables does the scatter plot indicate?
Calculate the covariance as well as the correlation between the two variables and interpret the
results.
Calculate the bivariate covariance as well as the bivariate correlation between each pair of the
variables mpg, cu_inc, hp, weight, and accel. Interpret the results! Compare the absolute values
of these two measures.

MA2 AdMR 140205 I Page 83


Regression analysis is a dominant method of data
analysis throughout the social sciences
Analyze and interpret data

The concept of the simple linear regression analysis

Broadly speaking, regression helps to understand how one or more independent variables are
related to a dependent variable of interest, and to make predictions based on this understanding
Specifically, simple regression is appropriate for one interval-scaled dependent variable (Y) and
one interval-scaled independent variable (X)
− In the absence of other information (i.e., values of X), our best guess at the value of the
¯
variable Y is its mean (Y)
− Our task in using a simple linear regression will be to do better than this, to use another
variable, X, to help explain Y
The principle of least squares suggests that while trying to “do it better than this” we should
minimize the sum of squared deviations between observed values (Yi) and those predicted by
^)
the model (Y i

MA2 AdMR 140205 I Page 84


A regression model can be used to explain and to
predict a dependent variable
Analyze and interpret data

Parameters used to explain and to predict a dependent variable in a simple regression model

Any observation of the dependent Y


variable Y (i.e., Yi) can be
explained by ^=b +b X
Y 0 1
− a model parameter b0 that
represents the mean value of
the dependent variable Y when
the value of the independent Yi = b0 + b1 Xi + ei
variable X is zero (intercept)
Yi
− a model parameter b1 that
b1 units
measures the change in the
value of the independent 1 unit
variable Y associated with a
one-unit increase in the value
of the independent variable X b0
(slope), and
− an error term ε that describes X
0
the effects on Y of all variables Xi
other than X
Any dependent variable Y can be predicted by assuming that e = 0; i.e., ^Y = b0 + b1 X

MA2 AdMR 140205 I Page 85


To do this requires fitting, or estimating, a
regression line
Analyze and interpret data

Partitioning total deviation in the simple regression model

The total deviation of each Y


observation of Y (Yi – Y)¯ can be
partitioned into a deviation due to ^ =b +b X
Y
^ – Y),
Xi (Y ¯ which is explained by
0 1
i
the regression, and a deviation ^)
(Yi – Y
^ ), which is i
due to “error” (Yi – Yi
Yi
unexplained by regression
^ – Y)
(Y ¯ ¯
(Yi – Y)
i
In order to come up with the
“best” regression line, we must ¯
Y ^ – Y)
(Y ¯
specify its slope and intercept so i

that the line comes close to the (Y – Y^) ¯


(Yi – Y)
i i
data (fits the data well)
Yi
So, in choosing the slope and
intercept, we will seek to minimize
the variation unexplained by the
regression, that is, the sum-of-
squared errors Σ (Yi – Y^ )2 0 X
i

In practical situations, we will never


be able to reduce all these unexplained deviations to zero

MA2 AdMR 140205 I Page 86


R square measures the quality of the regression
fit
Analyze and interpret data

Measures of regression fit and quality

R represents the correlation coefficient1)


R square, the coefficient of determination, represents the degree of variation reduction and is
used to judge the quality of the overall fit of a specific regression model
− R square is calculated as variation of Y explained by the regression divided by total variation
of Y
− In other words, it equals the percentage of spread of the dependent variable Y which was
lessened by the independent variable(s)
In case of a multiple regression analysis, r square adjusted should be used to assess the relative
quality of alternative regression models
− R square always gets larger if additional independent variables (X) are introduced, even when
those variables are utterly meaningless
− In contrast, if a statistically meaningless variable is put into a regression, r square adjusted will
get smaller
− Nevertheless, it is possible to add in a statistically nonsignificant variable and have r square
adjusted increase; to assess significance, the values of t and p for the newly added variable
needs to be looked at

1) See slide on the concept of the linear correlation coefficient

MA2 AdMR 140205 I Page 87


The significance of a regression is assessed on
two levels
Analyze and interpret data

Model as a whole Single independent variable

The F-test tells you about all the The t-test tells you about each of the
independent variables in your model taken independent variables in your model
together considered separately
The F-value is calculated by dividing the The t-value is calculated by dividing the
mean square regression by the mean estimate for the unstandardized coefficient
square error by the standard error
− The mean square error measures how Large values indicate that the estimates are
much variance in each data point can be many standard errors away from zero, and
attributed to error therefore are unlikely to have come about
− The mean square regression measures purely by chance
how much variance in each data point The p-values indicate the probability
can be attributed to the regression associated with a t-statistic (two-tailed test)
Note that the F-test is always a one-tailed
test

In simple regressions the F-value will always be the square of the t-value for the slope;
but this will not be the case when there are multiple independent variables

MA2 AdMR 140205 I Page 88


Analyze and interpret data

Application of the statistical methods

Guess at a certain car models mpg and cu_inc. Justify your choice.
You intend to make a more accurate prediction by running a simple linear regression model
using mpg and cu_inc. Which variable should function as the dependent variable?
Conduct a regression analysis and interpret the results.
Predict mpg for a car with 304 cubic inches.
For case no 1 (first row in the data file), forecast mpg and calculate the deviation due to error. As
well, calculate the total variation and the share of variation explained by the regression.

MA2 AdMR 140205 I Page 89


Whether a regression coefficient has a specific
value can be tested with a t-test
Analyze and interpret data

Basic inferential test for regression coefficients

Researchers will often wish to test whether a regression coefficient has a specific value
− E.g., is the slope significantly different from 0, or in general is it significantly different from any
(hypothesized) value
As in other statistical tests, a null and an alternative hypothesis needs to be specified1)
− H0: β = 0
− H1: β ≠ 0
As a standard, most statistical programs test whether β = 0 and provide the corresponding p-
value as well as the confidence interval
For all other hypothesized values, H0 can be tested using one of the following rules
− In case the confidence interval does not include the hypothesized value of β, H0 can be
rejected (at the confidence level used to calculate the interval); in case it does contain the
hypothesized value, we are not able to reject H0
− In case the (empirical) p-value is bigger than the defined significance level α, H0 can not be
rejected
• The p-value for any hypothesized value of β is determined by the degrees of freedom and
the t-value
• The t-value equals (b – β) / sb (with s being the standard error)

1) In statistics, generally Greek letters are used for population parameters (which will never be known without collecting
data on the entire population) and Roman letters for statistics calculated from a sample

MA2 AdMR 140205 I Page 90


Analyze and interpret data

Application of the statistical methods

You intend to predict mpg by hp. Run a simple linear regression model and calculate a 95 %
confidence interval for the regression parameters b0 and b1. Interpret the results. Is β1
statistically significant different from 0? Formulate H0 and H1. Justify your conclusion.
Test whether β0 is statistically different from 41,0 (α = 5 %). Formulate H0 and H1. Justify your
conclusion.
The R&D department claims that according to their calculations β1 equals -0,172 (α = 1 %).
Formulate H0 and H1. Justify your conclusion.
Predict mpg for a car with 150 hp.
For case no 2 (second row in the data file), forecast mpg and calculate the deviation due to
error. As well, calculate the total variation and the share of variation explained by the regression.

MA2 AdMR 140205 I Page 91


Regression analysis can on top be used to test for
difference in population group means
Analyze and interpret data

Inferential test of difference in population group means

It is common to carry out tests for differences in population means


− E.g., one wishes to test whether sales are the same in two regions or whether two brands
reach a similar customer satisfaction level
Such a test can – both technically and conceptually – easily be carried out as part of a simple
regression:
− code one group as 0 and the other as 1, then
− “regress out” the group variable
In such a regression, both the intercept (b0) as well as the slope (b1) have a specific meaning
− b0 is the best guess of the mean for the group coded as 0 in the population from which the
sample was drawn
− b1 is the best guess regarding the difference in the mean between the two groups
The statistical significance of both b0 and b1 can be tested with a t-test for regression coefficients

MA2 AdMR 140205 I Page 92


Analyze and interpret data

Application of the statistical methods

The R&D department asks whether there is a difference in mpg between cars from the US and
from Europe. Run a simple linear regression model and calculate a 99 % confidence interval for
the regression parameters b0 and b1. Interpret the results. Is β1 statistically significant different
from 0? Formulate H0 and H1. Justify your conclusion.

MA2 AdMR 140205 I Page 93


The chi square-test allows to test nominal
associations
Analyze and interpret data

Inferential test for nominal associations

Among the most common research questions in marketing practice is whether there is an association
between two nominal variables
− E.g., is there a relationship between consumer ethnicity and media consumption habits?
A relationship between two nominal variables can be illustrated with a cross tabulation, its statistical
significance can be tested with a chi square-test
− The null hypothesis for this chi square-test is that the two variables are independent of each other; the
alternative hypothesis is that they are not independent, that is, that there is a relationship between the
two variables
In case the two variables (A and B) are independent, the expected cell counts are determined by the
multiplication rule
− If A and B are independent, the probability of Ai and Bj occurring is the product of the probability of Ai
times the probability of Bj (multiplication rule)
− The expected number (or cell count) of cell Ai / Bj is the product of the total number of observations
times the probability of Ai and Bj occurring
A chi square-test should only be conducted if all the expected cell counts were 5 or greater
− If this is not the case, it is generally recommended that cells be combined to give an expected
frequency of at least 5, as otherwise the chi square-test can yield misleading results
In case the calculated chi square-value exceeds a critical (theoretical) chi square-value (given a certain
significance level α), the null hypotheses is rejected
− The chi square-value reflects the difference between the observed and the expected number of cell
counts across all cells

MA2 AdMR 140205 I Page 94


Analyze and interpret data

Application of the statistical methods

The file MA2 AdMR Data file 4 [Link] stems from 180 interviewees which were asked
about fifteen statements reflecting their opinion on foreigners.
What is the scale of the different variables? Correct, when necessary, the scale level indicated in
the column “measure” (variable view)!
Test whether there is an association between the variables SocAct and EcoSit, between SocAct
and Occup as well as between EcoSit and Occup. To do so, run cross tabulations and chi
square-tests (α = 5 %). Interpret the results!
Investigate the nature of the relationship between SocAct and Occup. To do so, calculate the
percentage deviations between the observed value and the expected value. Interpret the results!
Test whether there is an association between the variables Age and EcoSit. Interpret the results!

MA2 AdMR 140205 I Page 95


The term multivariate analysis refers to the
simultaneous analysis of multiple variables
Analyze and interpret data

Brief description of selected multivariate analysis techniques1)

Technique Enables the researcher to …


Multiple regression … predict the level of magnitude of a dependent variable based on
analysis the levels of more than one independent variable
... reduce a set of variables to a smaller set of factors or composite
Factor analysis
variables by identifying underlying dimensions in the data
… identify subgroups of individuals or items that are homogeneous
Cluster analysis
within subgroups and different from other subgroups
Multiple discriminant … predict group membership on the basis of two or more
analysis independent variables
… estimate the utility that different product features or attributes
Conjoint Measurement
provide to consumers

1) Further multivariate analysis techniques are, among others, multivariate analysis of variance (MANOVA),
multidimensional scaling (MDS), correspondence analysis, and structural equation modeling (SEM)

MA2 AdMR 140205 I Page 96


At least in terms of applications, statistics is
almost synonymous with regression
Analyze and interpret data

Overview of regression and its uses

Regression analysis is …
… a way to put a line through a group of points
This line minimizes the total sum-of-squares; that is, the squared distance to the line, summed
over all the points
… a method for testing the validity of relationships
Regression can help us determine whether marketing actions actually work; Verify
e.g., is there a relationship between ad spending and sales?
… a flexible methodology for measuring how things influence one another
Regression quantifies the nature of a relationship and by that allows to Quantify
determine how well marketing actions work
… a scientific approach to forecasting and prediction
Regression can help to determine how well another marketing action may Predict
work; e.g., add spending levels we did not try

When trying to say anything at all about the empirical world, regression is an
indispensable tool for an exceptionally wide variety of applications

MA2 AdMR 140205 I Page 97


The model presumed by the multiple regression is
an extension of the simple regression
Analyze and interpret data

Simple regression Multiple regression

Theoretical
model Y = β0 + β1 X1+ ε Y = β0 + β1 X1 + β2 X2 + β3 X3 + … + βK XK+ ε
(population)

Estimated Yi = b0 + b1 X1+ e Yi = b0 + b1 X1i + b2 X2i + b3 X3i + … + bK XKi + ei


model ^ +e ^ +e
(sample) =Y i i =Y i i

MA2 AdMR 140205 I Page 98


The multiple regression model is based on six
assumptions
Analyze and interpret data

Assumptions of the regression analysis

Assumptions concerning the relationship between the dependent and an independent variable
I. There is a linear relationship between the dependent variable and each independent variable1)

Assumptions concerning the errors, ei


II. The expected value of the errors is zero: E(ei) = 0
III. All errors have the same variance: var(ei) = δ2 (homoscedasticity)
IV. The errors are not correlated: cov(ei) = 0 (no autocorrelation)

Assumptions concerning the relationship between the independent variables


V. The independent variables are linear independent from each other (no multicollinearity)

Assumptions concerning the number of observations


VI. The number of observations need to be at least as big as the number of parameters to be
estimated (K + 1)

To estimate a regression model, the errors do not need to be normally distributed (as claimed in
many textbooks); nevertheless, being normally distributed is a prerequisite to conduct a t-test

1) Note that multiple regression can be used to assess all nonlinear relationships as long as they can be linearized

MA2 AdMR 140205 I Page 99


A way to detect heteroscedasticity is to plot the
residuals against the independent variables
Analyze and interpret data

The challenge of heteroscedasticity

Heteroscedasticity means that the errors do not have a constant variance


− E.g., there is far greater variation in the weight of people who are 1,85 m tall than those who
are 1,65 m tall
− This sort of pattern where the spread of the data varies at different levels of the independent
variable is common in real-world applications
When our data are heteroscedastic (that is, the error variance is not the same everywhere), we
run the risk of claiming far more certainty about our results than we should be
− The best way to see if heteroscedasticity is present is to plot the residuals against each of the
independent variables and simply look at them closely
− On top, heteroscedasticity can be detected with the White test and the Breusch-Pagan test
If heteroscedasticity is strongly present, it should be corrected for
− One way to do this is to use specific models that have been created for that purpose, like
GARCH or WLS models
− Another way to correct for heteroscedasticity is to transform either the independent of
dependent variables; logarithms often work well

In real-world applications, heteroscedasticity is not often considered a very serious problem,


unless it is extreme

MA2 AdMR 140205 I Page 100


Autocorrelation is a very serious problem, and can
be detected by visual inspection and the DW test
Analyze and interpret data

The challenge of autocorrelation

A key feature of the error is that it should be information-free, lacking meaningful patterns
− A meaningful pattern indicates that the initial model is leaving something out (e.g., seasonality)
or is just plain wrong
− If there are meaningful patterns, the researcher should come up with a model to explain the
patterns, not chalk them up to error
There are many varieties of autocorrelation, and some complex error relationships can be difficult
to detect
− One way to see whether the error is pattern-free is to ask a simple question: will knowing the
value of the error for one data point (ei) tell anything about the error value at another point?
The Durbin-Watson (DW) test is a simple test to detect the most common form of autocorrelation,
the so-called first-order1) autocorrelation
− It tests whether the residuals (errors) from a linear regression are autocorrelated
− The test will result in a value between 0 and 4, values near 2 indicate that autocorrelation is
not a problem
− On top, autocorrelation can be detected by plotting the standardized residuals against the
dependent variable
In situations where autocorrelation is present, it could be corrected for by a transformation
(particularly logarithms, exponents, and powers), using (time) lags, first differences (change in a
variable), or dummy variables (e.g., seasonality)
1) The name refers to the fact that the test looks at relationships between adjacent points

MA2 AdMR 140205 I Page 101


There are several options to take care of the
remaining assumptions
Analyze and interpret data

Main options to check assumptions I, II, V, and VI of the regression analysis

Assumption Qualitative inspection Statistical test


I To be judged by the researcher
based on theory and experience1)
II t-test on mean of residuals
(H0: µ = 0)
V To be judged by the researcher t-test on bivariate correlation
based on theory and experience coefficients
Change in b values when one (H0: rij = 0)
independent variable is added /
removed
VI Implicitly done by the statistical
program2)

1) Note that multiple regression can be used to assess all nonlinear relationships as long as they can be linearized
2) In case the number of observations is smaller than the number of parameters, the parameters will (= can) not be estimated

MA2 AdMR 140205 I Page 102


The Kolmogorov-Smirnov test assesses the
degree of deviation from normality
Analyze and interpret data

The challenge of non-normality

To conduct inferential tests, the error ei, that is, the difference between the predicted value for the
dependent variable and the value actually observed, should be normally distributed
Creating a histogram (which is typically called a normal probability plot) for the standardized
residuals and superimposing the best-fitting normal distribution helps to detect outliers and a
significant deviation from normality
− If the histogram and the best-fitting normal distribution appear to diverge substantially, there
may be a problem with non-normality
The degree of deviation from normality is also assessed by specific statistical tests, such as the
Kolmogorov-Smirnov, Anderson-Darling, and Shapiro-Wilk tests
Fortunately, small deviations from normality are no cause for worry
− However, if the normal probability plot indicates an extreme deviation between the histogram
and the superimposed normal distribution, the regression is likely to be misleading
− If non-normality is encountered, and if it is pronounced, it needs to be fixed, e.g., by using a
transformation and by checking whether some critical variable was omitted

MA2 AdMR 140205 I Page 103


The regression model can accommodate a wide
variety of variable types
Analyze and interpret data

Independent variables (X1, X2, …, Xk) Dependent variable (Y)

Interval Interval
Ordinal
Nominal Nominal
− Binary − Binary
− Multinominal − Multinominal
Rank-ordered
Count

Output of regression analysis

Coefficients
Fit measures (t and F)
p values

Although all regressions can be interpreted similarly, the mechanics of carrying out
regression can vary dramatically based on the type of dependent variable

MA2 AdMR 140205 I Page 104


Two criteria help to determine if a model may be
useful
Analyze and interpret data

Model as a whole Single independent variable

The F-test tells you if all the variables, taken The t-test determines if parts of the model –
together, help explain the variation in the that is, the different independent variables –
dependent variable, Y help explain the variation in the dependent
variable, Y

“Fit” (or prediction) should be (much) bigger, Coefficients (the bi) should be different from
on average, than “error” zero

If the p-value associated with the F-test is If the p-value associated with a t-test is non-
non-significant, it says that the entire model significant, it says that the independent
is not providing sufficient explanatory power variable Xi does not have a statistical
− In that case, the model must be changed, significant impact on the dependent variable
usually by attempting to remove under- Y
performing independent variables by − In that case, Xi should be removed from
looking at their t-tests the model

MA2 AdMR 140205 I Page 105


Multiple linear, ordinal, binary, and multinominal
regression are used extensively in marketing
Analyze and interpret data

Dependent variable Type of regression analysis

Interval Multiple linear regression

Ordinal Ordinal regression

Binary Binary regression

Multinominal Multinominal regression

Rand-ordered Rank-ordered or “exploded” regression

Count Poisson or count regression

MA2 AdMR 140205 I Page 106


Analyze and interpret data

Application of the statistical methods

You intend to run a multiple linear regression of cu_inc, hp, and weight on mpg (file MA2 AdMR
Data file 3 [Link]). Based on your theoretical knowledge and your experience, there is a linear
relationship between the mpg and the mentioned independent variables.
Check whether the assumptions of a multiple linear regression are met.
Based on the findings you decide to exclude all cases with a standardized residual of less than -3 or
more than +3.
Check whether the assumptions of a multiple linear regression are met.
You decide to not exclude the outliers. For the following calculations, assume that all the
assumptions of a multiple linear regression are met.
Run a simple linear regression model of hp on mpg (α = 5%). Interpret the results!
Run a multiple linear regression as described above (α = 5%). Interpret the results!
Modify the model based on the findings of the previous multiple linear regression. Run a multiple
linear regression for the modified model (α = 5%). Interpret the results!
Do the results allow the conclusion that the impact of hp is < -0,08 (α = 5%)? Justify your answer.

MA2 AdMR 140205 I Page 107


Analyze and interpret data

Application of the statistical methods

Modify the model by including year into the equation.


Run a multiple linear regression for the modified model (α = 5%). Interpret the results!
In a next step, exclude the independent variable Xi that does not have a statistical significant impact
on the dependent variable Y from the regression model.
Run a multiple linear regression for the modified model (α = 5%). Interpret the results!
In a next step, include no_cylin into the regression model
Run a multiple linear regression for the modified model (α = 5%). Interpret the results!

MA2 AdMR 140205 I Page 108


Whenever the dependent variable is ordinal, an
ordinal regression model should be applied
Analyze and interpret data

Basic concept of ordinal regression

There are many scales in marketing research that are ordinal


− E.g., frequency scales (like “never”, “rarely”, “sometimes”, “often”, and “always”) or adjectival
scales (like “child”, “young adult”, “middle-aged”, and “old”)
− Furthermore, the Likert-based scale (like the 1-to-7 agree-disagree scale) is ordinal,
nevertheless it is often treated as interval (and the multiple linear regression is applied)
Whenever the dependent variable is ordinal, an ordinal regression model should be applied
Although the mechanics of carrying out an ordinal regression varies from that of conducing a
multiple linear model, the results can be interpreted similarly
− The main extra output are the so called thresholds, and they help account for the different
distances between the scale points

MA2 AdMR 140205 I Page 109


Analyze and interpret data

Application of the statistical methods

You intend to explain the account status (1 = no debt history; 2 = no current debt; 3 = payments
current; 4 = payments delayed; 5 = critical account) of a customer. To do so, you might run an
ordinal regression (file MA2 AdMR Data file 5 Credit [Link]). Based on preliminary analysis, you
identified three predictors (factors) number of credits at the bank, other installment debts, and
housing type as well as three covariates: age and duration of loan,
Run an ordinal regression (α = 5%). Interpret the results!

MA2 AdMR 140205 I Page 110


A nominal independent variable can be included in
a regression model as a set of binary variables
Analyze and interpret data

Basic concept of nominal regression

Being able to include nominal (or categorical) independent variables is critical in applying
regression models to real-world problems
− E.g., seasonality, country of origin, or educational level
A binary variable can be simply integrated into a multiple linear regression model as an
independent variable
− Binary variables are quite easy to interpret because the coefficient always represents a “one-
unit increase” in the quantity in question
A nominal variable (with c categories) needs to be transformed into c-1 binary dummy variables
first
− E.g., the four seasons (winter, spring, summer, and fall) are transformed into three binary
variables winter (0/1), spring (0/1), summer (0/1), and if we know it is not winter, not spring,
and not summer, it must be fall
Researchers must carefully weight the pros and cons of adding dummy variables just because
they can
− We loose statistical power – represented by degrees of freedom – whenever we add more
independent variables
− Sometimes, a simpler model (one with fewer independent variables) is not only easier to
understand, it may actually be more powerful and offer superior forecasts

MA2 AdMR 140205 I Page 111


Analyze and interpret data

Application of the statistical methods

You intend to run a multiple linear regression of weight, year, and country on mpg (file MA2 AdMR
Data file 3 [Link]). Based on your theoretical knowledge and your experience, there is a linear
relationship between the mpg and the mentioned independent variables.
First of all, recode country into two binary variables USA and Europe
Run a multiple linear regression as described above (α = 5%). Interpret the results!
Check whether the standardized residuals are normally distributed.
Modify the model based on the findings of the previous multiple linear regression. Run a multiple
linear regression for the modified model (α = 5%). Interpret the results!

MA2 AdMR 140205 I Page 112


Factor analysis is used for data reduction, structu-
re identification, scaling and data transformation
Analyze and interpret data

Overview of factor analysis and its uses

Factor analysis takes a large number of variables and searches to see whether they have a small
number of factors in common that account for the correlations among the variables
Factor analysis has a number of possible applications in marketing research
− Data reduction
Reducing a mass of data (e.g., attributes) to a (far) smaller number of factors that underlie the
variables
− Structure identification
Discovering the basic structure underlying a set of measures
− Scaling
Identifying the optimal weights of variables being combined to form a scale (a weighted sum)
− Data transformation
Transforming the data (variables) into independent factors which can be used as an input for
many predictive techniques in statistics (e.g., multiple linear regression)

Among many other application areas, factor analysis is used for the development of
personality scales, market segments based on psychographic data, the identification of
critical product attributes, as well as similarities among products and lines

MA2 AdMR 140205 I Page 113


There are essentially three steps in a factor
analysis solution
Analyze and interpret data

The three steps in factor analysis

1. Develop a set of correlations between all combinations of the variables of interest


− Thus, it must be reasonable to treat the input variables as interval-scaled
2. Extract a set of initial factors from the correlation matrix developed in the first step
− The objective is to find a set of n factors that are linear combinations of the m variables
(with n ≤ m)
− A linear combination (called a principal component or a principal factor) can be defined as
z = b1X1 + b2X2 + …. + bmXm
− The principal components methodology determines values for b1, b2, …, bm that explain as
much variance in the correlation matrix as possible (first principal factor)
− After calculating the variance not explained by the first factor, a second factor is extracted,
then a third, and so on; we can always extract as many factors as there are variables
− The n factors extracted are uncorrelated with one another (they are said to be orthogonal)
3. Rotate the initial factors in a way that helps the researcher interpret them
− The initial factors face the problem of having each factor correlate modestly with all original
variables, thus making them very difficult to interpret
− The basic idea of rotation is to yield factors that each have some variables that correlate well,
while the rest correlate poorly
− There are two classes of rotation:
orthogonal rotation, which preserves the factors as uncorrelated with one another, and
oblique rotation, which allows the factors to become correlated with one another

MA2 AdMR 140205 I Page 114


Factor loadings are key in interpreting the results

Analyze and interpret data

Details about factor loadings, eigenvalues and communalities

The factor loadings are just correlations between variables and factors
− If a factor loading is high (near -1 or 1), it means that the factor loads high on that variable;
that is, that variable will be used to interpret the factor later on
In factor analysis, most every quantity we would like to know stems directly from the factor
loadings; if we sum squared loading across each of the following, this is what we get:
− For any factor, across all variables: that factor’s eigenvalue
An eigenvalue represents how much variance a factor explains relative to how much it
would be expected to explain by chance alone; that is, on average
In terms of interpreting a factor analysis, the typical approach is that factors with
eigenvalues less than1 should be discarded; those with eigenvalues not too much greater
than 1 are suspect; and those with large eigenvalues should be retained
− For any variable, across just the factors in our solution: that variable’s communality
A communality represents how well all factors together explain each of the variables
In terms of interpreting a factor analysis, the total communality should be compared with a
best possible value; on top, if the objective is to explain the original variables equally well,
the communalities for the variables should be on a similar level
− For any variable, across all factors: 1

MA2 AdMR 140205 I Page 115


Factor analysis requires two decisions based
mainly on the researcher’s judgment
Analyze and interpret data

How many factors do I need? How to label the different factors?

There is never a single correct answer to The interpretation of a factor is somewhat


this question valid for all situations subjective, but it is extremely helpful in
understanding the structure of the full set of
− Researchers need to consider the context
original variables
of their study
Part and parcel of performing a factor
There are several ways to select the
number of factors analysis is looking over the set of variables
that accord with a factor (load high on that
− Consult the so-called scree plot and factor) and using them to determine what
search for “kinks” that factor means
The scree plot visualizes how quickly the
quality of the factors degrades, in terms − Again, this is a subjective enterprise, but
a critical one
of the incremental variance they explain
− The label is used as a sort of shorthand
− Consult the eigenvalues
Select all factors with an eigenvalue to refer to the factor
greater than one
− Consult alternative rotated factor solution
Select the factor solution which allows for
the most practice-oriented (short and
snappy) interpretation of the factors

MA2 AdMR 140205 I Page 116


To ease interpretation, the original factors should
be rotated
Analyze and interpret data

Basic idea of varimax rotation

Varimax rotation “reorients” the original factors so that their loadings are as near -1, 0, or 1 as
possible while ensuring that the factors remain uncorrelated with one another (orthogonal
rotation)1)
− Although individual variables can correlate non-trivially with several factors, it must be
stressed that the factors themselves are perfectly uncorrelated (their correlation is zero)
Due to the varimax rotation, the variance explained by each factor will change
− The “best” of the factors will have smaller eigenvalue than before, whereas the “worst” will
have a larger eigenvalue
− The sum of the eigenvalues, however, will be the same

1) In contrast to an orthogonal rotation, an oblique rotation allows the factors to become correlated with one another

MA2 AdMR 140205 I Page 117


Analyze and interpret data

Application of the statistical methods

You intend to identify factors that account for the correlation between the fifteen statements
reflecting the opinion of British on foreigners (file MA2 AdMR Data file 4 [Link], variables
Stat01 to Stat15).
Run a factor analysis on the fifteen statements. Check the univariate descriptives as well as the
scree plot. Interpret the results!
Rotate the factor solution using varimax rotation. Interpret the results!
Based on the outcome of the analysis you decide to check a solution with four factors. Run a
factor analysis generating four factors and interpret the results!
Finally, you would like to check a solution with two factors. Run a factor analysis generating two
factors and interpret the results! Which alternative solution (2, 3, or 4 factors respectively) do you
evaluate as the most appropriate one?

MA2 AdMR 140205 I Page 118


Segmentation is the most common use of cluster
analysis
Analyze and interpret data

Overview of cluster analysis and its uses

Cluster analysis allows the researcher to place objects / items / people into groups; these
groups are often called clusters
− Cluster procedures form groups, assign objects (“cases”) to each of them, and help
determine a reasonable overall number of groups
− There are clustering algorithms available that take nominal, ordinal, interval, or ratio
measures as input1)
− Cluster procedures assume that natural clusters exist within the data
Among the most common uses of clustering techniques in marketing are segmenting customers
and segmenting products / brands

Cluster analyses involves a trade-off between two quantities: (1) the distance of each
point in a group to the group center, and (2) the distance between group centers

1) The following slides focus on metric variables, but the underlying principles for non-metric variables are similar

MA2 AdMR 140205 I Page 119


Cluster analysis requires five decisions based
mainly on the researcher’s judgment
Analyze and interpret data

Judgments to be made by researcher in the context of a cluster analysis

Which distance metric


Which objects to cluster?
to use? How many clusters to
Which clustering criterion extract?
Which variables to include?
to use?

Input Algorithm Output

MA2 AdMR 140205 I Page 120


Objects and variables need to be selected
consciously
Analyze and interpret data

Which objects to cluster? Which variables to include?

How a set of objects should be clustered The clustering procedure will count all
depends critically on variables included in the procedure as
equally important; it is up to the researcher
− which other objects are being clustered
to assess whether this is desirable or not
along with them and
An implicit weighting might be implied by the
− how much latitude within a cluster versus
across clusters should be allowed for fact that the variables are
− measured on different scales
Option: transform the variables first,
usually with a z-transform1)
− correlated (redundancy among variables)
Option: use (uncorrelated) factors as
variables
Based on theoretical considerations an
explicit weighting might be introduced by the
researcher, e.g. by multiplying the values by
some number (greater than 1 to emphasize,
less than 1 to de-emphasize)

1) This ensures that each variable is on an “equal footing” statistically, transformed to have a 0 mean and standard
deviation of 1

MA2 AdMR 140205 I Page 121


In statistics there are several choices what
distance is
Analyze and interpret data

Which distance metric to use? Which clustering criterion to use?

In order to group objects together, some There are two approaches to clustering, a
kind of similarity measure is needed hierarchical and a nonhierarchical approach
Users of clustering need to check carefully The distinguishing feature of hierarchical
that the metric is consistent with the cluster analysis (as opposed to non-
research question, especially whether hierarchical) is that once objects are
distance (volume) or dissimilarity (pattern) is clustered together, they are always together
of interest
The main hierarchical clustering method is
The most common distance measure for the Ward’s method (minimize within-cluster
metric variables is the (squared) Euclidian variation)
distance
− Additional methods are single linkage
− Additional distance measures are the (shortest distance), complete linkage
Minkowski metric, the Chebyshev (longest distance), average linkage
distance, and the city block metric (average distance), and centroid method
Common dissimilarity measures for metric (centroid distance)
variables are the Pearson correlation and The main nonhierarchical approach is the k-
the cosine means method

Different choices of distance metric as well as clustering criterion can yield


strikingly different final cluster solutions, so care must be taken to select the
one most concordant with the researcher’s needs

MA2 AdMR 140205 I Page 122


Hierarchical and nonhierarchical methods can be
used in sequence
Analyze and interpret data

Steps in combining hierarchical and nonhierarchical clustering models

First, a hierarchical approach can be used to identify the


− number of clusters and any
− outliers1), and to obtain
− cluster centers
Then, the outliers (if any) are removed
Finally, a nonhierarchical approach is used with the input on the
− number of clusters and the
− cluster centers
obtained from the hierarchical approach

The merits of both approaches are combined, and hence the results should be better

1) Single linkage is a favorable clustering method to identify outliers

MA2 AdMR 140205 I Page 123


The question as to the “best” number of clusters
can only be answered by the researcher
Analyze and interpret data

How many clusters to extract?

Unfortunately, unlike in regression or factor analysis, where there are objective measures (such
as r2) of fit, one must take a more “exploratory” (that is, trial-and-error, until the results “look
good”) approach in cluster analysis, and base one’s judgment on more or less pictorial evidence
First, the researcher can specify in advance the number of clusters or a range for the number of
clusters based on theoretical, logical, or practical considerations
Second, the distance between clusters (error variability measure) at successive steps may
serve as a useful guideline
− As a rule of thumb, one should stop when the successive distances between steps make a
sudden jump
− A plot of error sums of squares with the number of clusters may help to identify the jumps
Finally, the total cluster pattern (dendrogram) can provide a feel for an appropriate number of
clusters
− Based on theoretical and practical considerations, it is worthwhile to check whether a
proposed cluster solution does make sense or not
− A lack of statistical significant differences between clusters with regard to variables
considered as relevant indicate that the cluster solution is of limited use

The question as to the “best” number or clusters can only be answered by the researcher,
by balancing the project’s need for accuracy (which will argue for more clusters) against
the universal desire for a simple, robust explanation (which will argue for fewer)

MA2 AdMR 140205 I Page 124


Analyze and interpret data

Application of the statistical methods

The data in the file MA2 AdMR Data file 6 [Link] stems from a comparison of 20 different beer
brands. You intend to cluster the different brands based on the variables Calories, Sodium,
Alcohol, and Price.
Check whether an implicit weighting might be an issue. Propose measures to handle possible
redundancies among the variables!
You decide to proceed with standardized values (z-scores), but to keep all four variables.
Run a hierarchical cluster analyses using the squared Euclidean distance and the Ward’s
method. Ask for a dendrogram. Interpret the results!
Propose how many clusters to extract. Justify your proposal!1)
Based on the assessment of the different options you favor the solution with four clusters.
Describe the four clusters based on the four cluster variables! (Note: you can save the cluster
membership as a variable)
Check whether there is a statistical significant difference between the four clusters with regard
to the four variables Calories, Sodium, Alcohol, and Price (α = 10%)!

1) To do so, please refer to the file MA2 AdMR Determine the number of clusters

MA2 AdMR 140205 I Page 125


Analyze and interpret data

Application of the statistical methods

You still favor a solution with four clusters.


Run a nonhierarchical cluster analyses (k-means cluster). Interpret the results!
Compare both cluster solutions (hierarchical and nonhierarchical cluster analyses)! What do
you recommend?
Check whether there is a statistical significant difference between the four clusters (nonhierar-
chical cluster analysis) with regard to the four variables Calories, Sodium, Alcohol, and Price
(α = 10%)!
Based on the findings, you implement the improvement to the nonhierarchical cluster analysis
which is obvious .
Re-run the nonhierarchical cluster analyses (k-means cluster). Interpret the results!
Compare both cluster solutions (hierarchical and nonhierarchical cluster analyses)! What do
you recommend?

MA2 AdMR 140205 I Page 126


Discriminant analysis is a technique that often
complements cluster analysis
Analyze and interpret data

Overview of discriminant analysis and its uses

Discriminant analysis is appropriate when one seeks to understand a nominal dependent variable
(e.g., a grouping) in terms of several – interval or binary – independent variables
Discriminant analysis has a number of possible applications in marketing research
− Estimate “discriminant functions” (discrimination)
Which linear combination of the given (independent) variables best distinguishes known
groups (dependent variable)?
− Make group predictions (classification / prediction)
Given a new set of items (e.g., customers, products, firms) whose group membership we do
not know, which of the pre-established groups are they likely to fall into?
− Determine whether the groups really seem different (testing / verification)
Are the various groups significantly different, based on the “profiles” (independent variables) of
the individuals found in them?
− Identify the most useful predictors in discrimination (influence / importance)
Which input variables seem to best predict group differences?

MA2 AdMR 140205 I Page 127


Key challenge is to produce better-than-chance
assignment of objects to pre-defined groups
Analyze and interpret data

Typical research questions which might be addressed with a discriminant analyses

Application Typical research question


Which information in our customer database1) best explains which
Discrimination
customers did reply to our promotional offer last month?

Classification / Can we predict which customers will reply to next month’s


prediction promotional offer?

Does the customer data help identify promo-sensitive customers?


Testing / verification
Or could the groupings be arising from chance alone?

Which variables are most helpful in prediction who will reply to our
Influence / importance
promotions? Do some seem completely useless?

1) E.g., prior purchases, age, income, geodemographics

MA2 AdMR 140205 I Page 128


The basic idea of discriminant analysis is one
familiar from regression
Analyze and interpret data

Basic idea and output of discriminant analysis

The basic idea of discriminant analysis is to find a linear combination of the independent
variables that makes the predicted mean for each category as different as possible
This linear combination of the n independent variables (known as the discriminant function or
axis) is derived from an equation that takes the form
DF = b1 X1 + b2 X2 + b3 X3 + … + bN XN,
which should look exactly like an ordinary multiple regression
When there are m different categories (groups) to distinguish (with m < n), discriminant analysis
produces (m-1) discriminant functions DF1, DF2, …, DFm-1
Making the mean for each category as different as possible is achieved by maximizing the
between-group variance relative to the within-group variance
Discriminant analysis output typically includes the
− values of the b’s, along with a
− confusion matrix which categorizes correct and incorrect predictions by cross-tabulating the
predicted with the actual category of the dependent variable, and some
− test statistics to asses significance

Discriminant analysis can handle many sorts of independent variables, not only interval-scaled
ones; however, nominal independent variables need to be converted to (binary) dummy variables

MA2 AdMR 140205 I Page 129


Analyze and interpret data

Application of the statistical methods

You are a loan officer at a bank and want to identify characteristics that are indicative of people
who are likely to default on loans. Information on 850 past and prospective customers is contained
in the file MA2 AdMR Data file 7 Bank [Link]. The first 700 cases are customers who were
previously given loans.
Run a discriminant analysis using the variables age, employ, address, income, debtinc,
creddebt, othdebt and default. Ask for the following statistics: means, Box’s M, univariate
ANOVA, Fisher’s as well as unstandardized function coefficients. Make sure to compute prior
probabilities from group sizes and to display a summary table and a leave-one-out
classification. Interpret the results!
Re-run the discriminant analysis using a separate-groups covariance matrix. Interpret the
results!
You decide to stay with the discriminant analysis using the within-groups covariance matrix.
Re-run the discriminant analyses, but assume that a case is equally likely to be a defaulter or a
nondefaulter. Interpret the results!
Assume that the average volume of a requested bank loan is the same in both groups.
According to internal data, the profitability of nondefaulters is 1% and of defaulters -5%. Draw
conclusions!

MA2 AdMR 140205 I Page 130


Analyze and interpret data

Application of the statistical methods

In a next step, the 150 prospective customers should be classified, in other words: a decision
needs to be made whether a loan should be made to them or not.
Classify the prospective customers. Make sure to compute prior probabilities from group sizes.
Interpret the results!
Re-run the discriminant analyses, but assume that a case is equally likely to be a defaulter or a
nondefaulter. Interpret the results and compare both classifications!
You decide to classify the prospective customers based on prior probabilities reflecting the group
sizes. However, an applicant will only be eligible to receive a loan in case the probability being a
nondefaulter is at least 55%.
How many prospective customers will receive a loan?

MA2 AdMR 140205 I Page 131


Conjoint measurement is a set of methods to
understand and quantify trade-offs
Analyze and interpret data

Overview of conjoint measurement and its uses

Every market is characterized by a key tension between the


− consumer wanting “the best of everything at the lowest price” and the
− manufacturer wanting “the lowest production costs and the highest price”
A first step for the manufacturer, then, is understanding two fundamental issues:
− which attributes (e.g., price, safety, acceleration, softness) and levels (e.g., $ 6, $ 8, $ 10)
consumers value, and
− how much consumers value those attributes and levels
The overall goal of conjoint measurement is to quantify
− the relative importance of the various attributes and
− the utilities of the various levels of any particular attribute
using information supplied by respondents, such as ranks, ratings, or choices

Conjoint allows marketers to find the “sweet spot” where consumers, manufacturers, and
retailers can most mutually agreeably meet

MA2 AdMR 140205 I Page 132


There are essentially four steps in conjoint
measurement
Analyze and interpret data

The four steps in conjoint measurement

1. Decide on the attributes and the levels


− Usually, there is a need to do some preliminary research to determine the most relevant
attributes and a realistic set of levels for each
2. Collect trade-off data from consumers
− There are four alternative methods to collect trade-off data:
ranking-based conjoint,
orthogonal designs,
adaptive methods, and
choice-based conjoint
3. Estimate the value system (relative importance of attributes and utilities of the various levels)
4. Make choice predictions

It is important to note that conjoint predicts preference, not sales or market share

MA2 AdMR 140205 I Page 133


Orthogonal design allows to offer a subset of
possible product profiles
Analyze and interpret data

Main characteristics of the orthogonal method

Pairing every possible attribute level with every other in a conjoint analysis would typically result
in a large number of possible product profiles, making conjoint a practical impossibility if
respondents had to evaluate them all
− Six attributes and five levels per each attribute result in more than 15.000 possible profiles
An orthogonal design allows researchers to offer a subset of, rather than all, possible
combinations
− Orthogonal design works because it is assumed that what is observed for one variable is
unrelated to what is observed for any of the other
− This means that, when describing a product by its attribute levels, including a particular level of
one attribute gives no information about what levels are present for other attributes
− In the example mentioned above, the 15.000 profiles can be “orthogonalized” into just 25
The conjoint task involves ranking the subset of possible combinations from “most preferred” to
“least preferred”
Please rank the following 18 movie theater configurations according to your preference, 1
being the most and 18 the least preferred one!
Ticket price Line of sight Seat comfort Concessions
$6 Staggered Average seat Gourmet snacks
$6 Not staggered Big seat Hot dogs / popcorn
$8 Staggered Average seat Hot dogs / popcorn

MA2 AdMR 140205 I Page 134


Adaptive conjoint is especially suitable if there
are many attribute levels
Analyze and interpret data

Main characteristics of an adaptive conjoint

The motivating idea behind adaptive conjoint is: prior choices or responses are utilized to
determine future comparisons
− If the first few conjoint responses strongly indicate that a person is insensitive to some aspects
of a product, he / she should be questioned about trade-offs more relevant to that person
Hence, an adaptive conjoint program will calculate part worths after each new piece of data is
supplied by the respondent, then use special algorithms to better measure those part worths that
are not yet determined with sufficient accuracy
− The researcher can specify exactly how much accuracy is required
− In a nonadaptive conjoint program, even when all the data are collected, there is no guarantee
of sufficiently accurate results
In a two-option adaptive conjoint task, the response is given on an ordinal scale (e.g., a scale
from 1 to 9), with one end of the scale representing “strongly prefer option A” and the other end
representing “strongly prefer option B”
Which of the following laptop computers would you rather purchase?
Which ofFastest
the following laptop computers would you rather
processor Verypurchase?
fast processor
Fastest 6processor
h battery life Very fast4processor
h battery life
6 h battery$life
999 4 h battery$life
799
$ 999
Strongly $ 799 Strongly …
prefer left 1 2 3 4 5 6 7 8 9 prefer right
Strongly Strongly
prefer left 1 2 3 4 5 6 7 8 9 prefer right

MA2 AdMR 140205 I Page 135


Professional conjoint applications are
overwhelmingly choice-based
Analyze and interpret data

Main characteristics of a choice-based conjoint

The idea behind choice-based conjoint is to have respondents simply make the sort of trade-offs
they normally do: in their heads, reporting only the item that they would actually choose
− For this reason, choice-based conjoint is dramatically more “real world” than other conjoint
tasks
Despite the need to present respondents with a larger number of experimental tasks when using
choice-based conjoint instead of ranking or rating, it can be highly efficient because several
product profiles can be offered at once
− Furthermore, because each task is far simpler and clearer to the respondent, this increase in
the number of tasks rarely inflates the complexity of the study or the time respondents must
put in
In a choice-based conjoint task, the respondent indicates which of the product profiles presented
to him he / she would choose or, importantly, that none of them is acceptable
− The no-choice option accounts for situations when the presented attribute level combinations
are all below some threshold for purchase
Which of the following desktop computers would you purchase?
WhichDell
of the followingHP
desktop computers
Sony would youNone:
purchase?
if these
Fast Fastest Very fast were my only
Dell HP Sony None: if these
processor processor processor choices, I
Fast Fastest Very fast were my only
27‘‘ monitor 24‘‘ monitor 21‘‘ monitor would defer
processor
$ 1.000
27‘‘ monitor
processor
$ 900
24‘‘ monitor
processor
$ 850
21‘‘ monitor
choices, I
my purchase
would defer

$ 1.000 $ 900 $ 850 my purchase

MA2 AdMR 140205 I Page 136


Analyze and interpret data

Application of the statistical methods

Examine thoroughly Feinberg / Kinnear / Taylor (Modern Marketing Research), pp. 531-534,
marketing research focus 11.1: designing your own movie theater using conjoint analysis and an
orthogonal design.
Check the relative importance of the different attributes by referring to the part worths. Explain
your calculation!
Assume that the part worths are calculated for a specific person, let’s call him Mr. Smith. There are
three cinemas Mr. Smith might choose from. Cinema A offers staggered seats, average seats
without cup holder, a small screen with plain sound, hot dogs and popcorn at a price of $ 6.
Cinema B differs from A that the seats are equipped with a cup holder, that the screen is large, the
sound is digital and the ticket is $ 8. In cinema C (in contrast to B) the seats are not staggered, the
sound plain and the ticket – like in A – is $ 6.
Which cinema do you expect Mr. Smith to prefer?
The owner of cinema B would like to gain Mr. Smith as a customer. By how much would he
need to reduce the ticket price if he would like to offer Mr. Smith a total utility that is 5% above
the best alternative? Use a linear interpolation!

MA2 AdMR 140205 I Page 137


Multidimensional scaling summarizes data about
associations to create a perceptual map
Analyze and interpret data

Overview of multidimensional scaling (MDS) and its uses

Humans are visual creatures: they understand information best when it is integrated into a
picture, not when presented as a disembodied table of numbers
Multidimensional scaling is a way to construct “pictures” of markets that do not rely on someone
deciding beforehand which dimensions are important, or what data to collect in order to draw the
(perceptual) map
MDS summarizes data about associations between a fixed set of objects to reveal relationships
between them (usually brands in a particular product class) by determining the
− minimum dimensionality required to represent the objects’ interrelationship well, and the
− position of each object on each dimension (i.e., its location in the map)
without collecting any attribute-based data
MDS will always provide as faithful a representation of relative similarity or distance as possible,
but we can only attach meaning to the dimensions in the derived graph by tying in other kinds of
data
The MDS routine, unlike those for linear regression, is iterative, determining a good initial guess,
calculating a goodness-of-fit measure, and continuing to derive better and better estimates,
always attempting to better represent the distance data

Multidimensional scaling uses a pretty simple data type: item (dis-)similarity

MA2 AdMR 140205 I Page 138


There are three general types of MDS

Analyze and interpret data

Fully metric Nonmetric Fully nonmetric

Interval- or ratio-scaled Rank-ordered input Rank-ordered input


input measures measures required measures required
required
Interval- or ratio-scaled Interval-scaled Rank order of each
relationships among relationship among object on each
objects generated objects generated dimension generated

That is, the distances between


objects in the perceptual space
are meaningful, and can be com-
pared like any ordinary distances

MA2 AdMR 140205 I Page 139


MDS can be applied whenever rigorous, objective,
pictorial presentations of objects prove useful
Analyze and interpret data

Some guidelines for running a MDS

A (perceptual) map should be based on dissimilarities; that is, distances


− It is advisable to define “0” as the point indicating the greatest possible “closeness”
− Note that the program does not need the original individual-level data, but merely their averages, to
create the perceptual map
MDS starts being more reliable (in terms of unambiguous interpretation of the resulting spatial
dimensions) for large number of object comparisons
− For n objects, data on n*(n-1)/2 pair-wise comparisons will be used to estimate x*n coordinates in a x-
dimensional space
− The ratio of number of comparisons to number of coordinates should be “big enough” (e.g., a ratio of 2
is considered as low)
Few dimensions (in practice, two) are always preferred because they allow for easy visualization, but
should not be chosen if they are incapable of representing the data faithfully
− The stress is a goodness-of-fit measure between the input rank order and the output
− A stress level of over 0,20 is considered quite poor, a stress of 0,10 is fair, of 0,05 is good, and of 0,01
or less is excellent
− On top, most programs provide the r2 value
All that matters in a MDS are the relative positions of the objects, and that we can overlay meaningful
data
− The X- and the Y-axes do not have some special, privileged meaning in a map, just because the
computer placed them there

MA2 AdMR 140205 I Page 140


Analyze and interpret data

Application of the statistical methods

You are a product manager at Volvo cars, being responsible for the V40. The file MA2 AdMR Data
file 8 [Link] contains a rank order of similarities between pairs of – in total – 11 car models
belonging to the so-called Golf-class. The rank number “1” represents the most similar pair.
Run a multidimensional scaling using the Euclidean distance. Ask for group plots, data matrix,
as well as model and options summary. Interpret the results!
How could you use the scaling results? Which conclusions would you draw from the results of
the analyses?

MA2 AdMR 140205 I Page 141


Writing the research report is the most crucial
part of the research process

Establish need for information

Detail research objectives and information needs

Set research design and data sources

Design data collection procedure

Design sample

Collect data

Process and code data

Analyze and interpret data

Present results and conclusions


The research results are typically communicated through a written report and an oral
presentation. Findings should be presented in a simple format and addressed to the
information needs of the decision situation. No matter how proficiently all previous
steps have been dispatched, the project will be no more successful than the report.

MA2 AdMR 140205 I Page 142


Above all else, the report is designed to
communicate knowledge to decision makers
Present results and conclusions

Guidelines for written marketing research reports

Consider the audience


− Use only words familiar to a broad readership
− Make wide usage of percentages, rounded-off “ballpark” figures, ranks, or ratios
− Use graphical aids (charts, graphs, pictures, etc.) as frequently as possible
Address promised information needs
− The research report is not meant to showcase the cornucopia of intriguing findings discovered
by the research team or the scorching power of its statistical analyses
Be concise, yet complete
− Knowing what to include and what to leave out is a difficult task, but it is yours, not the client’s
Be objective
− There is a world of difference between presenting conclusions diplomatically and altering their
underlying nature to avoid hurt feelings
Adopt a concise, businesslike writing style
− Have someone unrelated to the project read the final report through; if numerous clarifying
questions arise, rewriting is required

Regardless of the rigor, thoroughness, and soundness of the methodology underlying the report,
the research will be useless to executives if the report fails to make an airtight, compelling case

MA2 AdMR 140205 I Page 143

You might also like