0% found this document useful (0 votes)
2 views54 pages

WMSM Notes

This document is a comprehensive study guide on marketing research, covering theory, methods, analysis, and advanced techniques. It outlines key concepts such as theory, constructs, propositions, and hypotheses, as well as the marketing research process, research designs, measurement, scaling, and questionnaire design. It provides definitions, examples, and steps involved in each area, aimed at equipping readers with essential knowledge for conducting effective marketing research.

Uploaded by

Sumit Joshi cs21
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views54 pages

WMSM Notes

This document is a comprehensive study guide on marketing research, covering theory, methods, analysis, and advanced techniques. It outlines key concepts such as theory, constructs, propositions, and hypotheses, as well as the marketing research process, research designs, measurement, scaling, and questionnaire design. It provides definitions, examples, and steps involved in each area, aimed at equipping readers with essential knowledge for conducting effective marketing research.

Uploaded by

Sumit Joshi cs21
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MARKETING RESEARCH

Comprehensive Study Guide


Theory | Methods | Analysis | Advanced Techniques

Topics Covered
Theory & Constructs | Research Process | Research Design | Measurement & Scaling |
Questionnaire | Sampling | Data Collection & Preparation | Validity & Reliability | Hypothesis
Testing | ANOVA | Correlation & Regression | Discriminant & Logit Analysis | Factor Analysis
| Cluster Analysis | Interdependence & Independence Techniques
UNIT 1: THEORY, CONSTRUCTS, PROPOSITIONS &
HYPOTHESIS

1.1 Theory in Research


Definition
A theory is a systematic set of interrelated statements (concepts, definitions, and propositions)
that explains or predicts events or situations by specifying relations among variables.
It provides a framework for understanding phenomena and guides research design.

Characteristics of a Good Theory


• Conceptually clear and precisely defined
• Operationally testable and falsifiable
• Parsimonious — explains much with few propositions
• Internally consistent — no contradictory propositions
• Explains and predicts real-world phenomena
• Example: Theory of Planned Behavior (TPB) — explains consumer purchase intentions through
attitudes, subjective norms, and perceived behavioral control.

1.2 Concepts and Constructs


Concept
Definition
A concept is a generalized idea about a class of objects, attributes, occurrences, or processes.
It is a mental image or abstraction of reality.
Example: 'Brand Loyalty', 'Customer Satisfaction', 'Price Sensitivity' are all concepts.

Construct
Definition
A construct is a specific type of concept that is deliberately invented (constructed) for a specific
scientific purpose. It is an abstract idea that cannot be directly observed but must be measured
indirectly through indicators.
Example: 'Intelligence' is a construct — it cannot be directly seen, only measured via IQ tests,
academic performance, problem-solving, etc.
Difference: Concept vs. Construct
Aspect Concept Construct
Nature General mental abstraction Specifically built for research
Observability May be directly observable Always latent / unobservable
Measurement Direct or indirect Always indirect (via indicators)
Example Color, Size, Age Brand Equity, Self-Esteem, Loyalty

1.3 Proposition
Definition
A proposition is a statement that specifies the relationship between two or more concepts or
constructs. It is declarative in nature and describes how concepts are related.
Propositions are qualitative statements — they are NOT statistically testable in their general
form.

Examples of Propositions:
• 'Higher advertising expenditure is associated with higher brand awareness.'
• 'Customer satisfaction leads to customer loyalty.'
• 'Price perception influences purchase intention.'

Types of Propositions
• Axioms: Propositions assumed to be true without direct testing (foundational statements).
• Theorems: Propositions derived logically from axioms.
• Hypotheses: Operationalized propositions that are empirically testable.

1.4 Hypothesis
Definition
A hypothesis is a specific, testable statement about the expected relationship between two or
more variables. It is an operationalized proposition — it specifies how to measure the variables
and what the expected result is.
A hypothesis must be: Specific, Measurable, Empirically testable, Stated before data collection.

Types of Hypotheses
1. Null Hypothesis (H0)
States that no relationship, difference, or effect exists. It is the hypothesis of 'no change'.
• Example: H0: There is no significant relationship between advertising spend and sales revenue.

2. Alternative Hypothesis (H1 or Ha)


States that a relationship, difference, or effect does exist. It is what the researcher expects to find.
• Example: H1: Higher advertising spend is significantly associated with higher sales revenue.

3. Directional Hypothesis
Specifies the direction of the expected relationship (positive/negative, greater/lesser).
• Example: H1: Brand awareness positively influences purchase intention.

4. Non-Directional Hypothesis
States that a relationship exists but does not specify direction.
• Example: H1: There is a significant difference in customer satisfaction between urban and rural
consumers.

Theory → Construct → Proposition → Hypothesis (Flow)


Level Element Example Purpose
Abstract Theory Theory of Consumer Behavior Overall framework
Conceptual Construct Brand Loyalty, Satisfaction Unmeasured concept
Relational Proposition 'Satisfaction leads to loyalty' States relationship
Empirical Hypothesis 'H1: Satisfaction score is Testable statement
positively correlated with repeat
purchase rate (r>0, p<0.05)'
UNIT 2: MARKETING RESEARCH PROCESS

2. Steps in the Marketing Research Process


The marketing research process is a systematic series of steps to plan, collect, analyze, and interpret
information for decision-making.

Step Stage Description Example


Step 1 Define the Problem & Identify the management decision Why are sales declining in the
Research Objectives problem and translate it into a eastern region? Objective:
researchable question. Define the Identify factors causing the
research objectives (exploratory, 15% decline.
descriptive, or causal).
Step 2 Develop the Research Determine data sources Online survey of 500
Plan (primary/secondary), research customers aged 18-45 in the
approach, instruments eastern region using stratified
(questionnaire/scale), sampling random sampling.
plan (who, how many, how
chosen), and budget.
Step 3 Collect the Execute the data collection plan. Administer the questionnaire
Information This is the most expensive, error- via email/app; monitor
prone, and critical stage. Involves response rates daily.
fieldwork, supervision, and quality
checks.
Step 4 Analyze the Process, clean, code, and Use SPSS to run frequency
Information statistically analyze the collected tables, cross-tabulations, and
data. Use appropriate tests (t-test, regression analysis.
ANOVA, regression, factor
analysis, etc.).
Step 5 Present the Findings Prepare a research report with key PowerPoint presentation
findings, charts, conclusions, and showing 3 key drivers of sales
actionable recommendations. decline with recommended
Present to management. solutions.
Step 6 Make the Decision Management uses the research Increase customer service
findings to make an informed staff and revise pricing in the
decision and implement solutions. eastern region.
UNIT 3: RESEARCH DESIGNS

3. Research Design
Definition
A research design is the overall plan or blueprint for conducting a marketing research study. It
specifies the methods and procedures for collecting, measuring, and analyzing data.
It answers: What data is needed? From whom? How? When? And at what cost?

A. Exploratory Research Design


Definition
Used when the problem is vague or not clearly defined. It aims to generate insights, ideas, and
hypotheses — not conclusive answers.

Feature Details
Purpose Discover ideas, understand problems, generate hypotheses
Structure Flexible, unstructured, qualitative
Sample Size Small (10-50 typically)
Data Type Qualitative
Output Insights, tentative hypotheses for further research
Methods Focus groups, in-depth interviews, case studies, literature review, projective
techniques
Example A startup conducts 6 focus groups with millennials to understand what
features they want in a new banking app before investing in development.

B. Descriptive Research Design


Definition
Used to describe characteristics of a market, group, or phenomenon. Answers WHO, WHAT,
WHERE, WHEN, and HOW questions. Most common design in marketing research.

Feature Details
Purpose Describe market characteristics, consumer profiles, market size, frequency
Structure Formal, structured
Sample Size Large (hundreds to thousands)
Feature Details
Data Type Quantitative
Sub-types Cross-sectional (one-time snapshot) vs. Longitudinal (tracked over time)
Methods Surveys, questionnaires, observation, secondary data analysis
Example A telecom company surveys 5,000 customers to find: What % use mobile
data daily? Which age group uses most data? What features do they value?
This creates a customer profile.

C. Causal (Experimental) Research Design


Definition
Used to determine cause-and-effect relationships. Tests whether manipulating one variable
(independent) causes a change in another (dependent). Most rigorous design.

Feature Details
Purpose Establish cause-and-effect between variables
Structure Highly structured, controlled
Key Concept Manipulation of independent variable + control of extraneous variables
Internal Validity Control over confounding variables (lab experiments > field experiments)
Types Laboratory experiments vs. Field experiments vs. Test markets
Example (A/B Test) An e-commerce brand tests two checkout page designs (A and B) on equal
groups. Group A (red button) shows 3.2% conversion; Group B (green
button) shows 4.8% conversion. Conclusion: button color causes conversion
difference.
Example (Test Market) FMCG company launches new shampoo in 3 cities with different price
points to determine which price maximizes sales before national launch.

Comparison of All Three Research Designs


Feature Exploratory Descriptive Causal
Objective Discover/explore Describe/measure Establish cause-effect
Flexibility High Low Very Low
Sample Small Large Moderate-Large
Data Qualitative Quantitative Quantitative
Hypothesis Not yet formed Sometimes Clearly stated
Cost Low Moderate High
Used When Problem unclear Problem defined Relationship to test
Feature Exploratory Descriptive Causal
Example Focus group for new Customer satisfaction A/B price testing
product survey
UNIT 4: MEASUREMENT, SCALING & QUESTIONNAIRE

4.1 Measurement
Definition
Measurement is the process of assigning numbers or labels to objects, persons, events, or
states according to specific rules to represent quantities or qualities of attributes.
In marketing research, measurement converts real-world observations into data that can be
statistically analyzed.

4.2 Scaling
Definition
Scaling is the process of creating a continuum upon which measured objects are located. It
generates a range of responses to a measurement question.
Scaling answers not just 'what' but 'how much' or 'how strongly' — it adds intensity to
measurement.

Four Types of Measurement Scales


Scale Order? Equal True Statistics Allowed Marketing Example
Intervals? Zero?
Nominal No No No Mode, Chi-square, Gender (M/F), Brand
Frequency used (A/B/C)
Ordinal Yes No No Median, Percentile, Customer ranking of
Rank correlation brands (1st, 2nd, 3rd)
Interval Yes Yes No Mean, SD, t-test, Likert scale, Semantic
ANOVA, Correlation differential
Ratio Yes Yes Yes All statistics including Sales (Rs), Age, Income,
geometric mean Market share (%)

Comparative vs. Non-Comparative Scales


Category Type Description Example
Comparative Paired Comparison Respondent compares two Do you prefer Pepsi or Coke?
objects at a time
Comparative Rank Order Scale Respondent ranks all Rank these 5 brands from
objects in order best to worst
Category Type Description Example
Comparative Constant Sum Respondent divides 100 Give 100 points: Quality=40,
Scale points among objects Price=30, Service=30
Non-Comparative Likert Scale Agreement-disagreement Strongly Agree to Strongly
on statements Disagree (5/7 points)
Non-Comparative Semantic Bipolar adjectives on a Good----Bad; Fast----Slow (7-
Differential continuum point scale)
Non-Comparative Stapel Scale Rate object on +5 to -5 Rate product quality: -5 to +5
scale without neutral
Non-Comparative Rating Scale Rate object on a single Overall satisfaction: 1=Poor
dimension to 10=Excellent

4.3 Likert Scale


Definition
The Likert Scale (Rensis Likert, 1932) is a psychometric rating scale used to measure attitudes,
opinions, beliefs, and perceptions by having respondents indicate their level of agreement or
disagreement with a series of statements.
It is an interval scale. Each item uses a 5-point or 7-point response format.

Standard 5-Point Likert Scale


Scale Point Label Numeric Code
1 Strongly Disagree 1
2 Disagree 2
3 Neutral / Neither Agree nor Disagree 3
4 Agree 4
5 Strongly Agree 5

Example Likert Scale Items


Statement: 'I am satisfied with the overall quality of this product.'
Strongly Disagree Neutral Agree Strongly Agree
Disagree
[] [] [] [] []

Advantages & Disadvantages of Likert Scale


Advantages Disadvantages
Easy to construct and administer Central tendency bias (choosing middle option)
Respondents find it simple to understand Social desirability bias
Produces interval-level data for analysis Does not explain WHY respondent chose that
option
Can measure intensity of attitudes Acquiescence bias (tendency to agree)
Can be summed to create composite scores Assumes equal intervals between scale points

Variations of Likert Scale


Type Points Feature Best Used For
Standard 5-point Includes neutral point General consumer research
Extended 7-point More sensitivity, finer discrimination Academic/psychological
research
Forced 4-point No neutral — forces opinion When fence-sitters must
Choice commit
Visual Analog Continuous Slider on a line from 0-100 Online surveys, UX research

4.4 Questionnaire
Definition
A questionnaire is a structured set of questions (instrument) designed to collect specific
information from respondents. It is the most widely used primary data collection tool in marketing
research.
A good questionnaire translates research objectives into specific, answerable questions.

Types of Questions
Type Description Example
Open-Ended Free-form answer, no fixed options 'What do you like most about our
product?'
Dichotomous Only 2 options (Yes/No) 'Do you own a smartphone? Yes / No'
Multiple Choice Select from several options 'How often do you shop online?
Daily/Weekly/Monthly/Rarely'
Likert Scale Rate agreement on a statement 'Product meets expectations: 1(SD)
to 5(SA)'
Semantic Differential Rate on bipolar adjectives 'Service quality: Excellent 1-2-3-4-5-
6-7 Terrible'
Type Description Example
Ranking Rank options in order 'Rank these features by importance:
Price, Quality, Brand'
Rating Rate on a numeric scale 'Rate your experience: 1(Very Poor)
to 10(Excellent)'

Steps to Design a Good Questionnaire


Step Action Key Consideration
1 Define the Information What decisions will this data support?
Needed
2 Define the Target Who has the information you need?
Respondents
3 Choose Interview Method Online, phone, face-to-face, mail?
4 Decide Question Content Is each question necessary? Will it yield usable data?
5 Choose Question Wording Simple, clear, unambiguous, unbiased language
6 Determine Question Logical flow: general to specific; easy to difficult
Sequence
7 Choose Physical Form Layout, spacing, fonts, section breaks
8 Pilot Test Test on 10-20 respondents; revise based on feedback
9 Revise & Finalize Incorporate pilot feedback; check for completeness

Characteristics of a Good Questionnaire


• Uses simple, clear, and unambiguous language
• Each question measures one thing only (no double-barreled questions)
• Avoids leading, loaded, or biased questions
• Logical and smooth flow from easy to difficult questions
• Appropriate length — not too long (respondent fatigue)
• Pre-tested and validated before fieldwork
• Covers all research objectives completely
UNIT 5: SAMPLING & SAMPLING TECHNIQUES

5.1 Sampling — Definition & Key Terms


Definition
Sampling is the process of selecting a representative subset (sample) from a larger group
(population) in order to make inferences or generalizations about the whole population.
It saves time, cost, and effort while still producing reliable and valid results.

Term Definition Example


Population The entire group the researcher All smartphone users in India
wants to study
Sample A subset selected from the 2,000 smartphone users selected
population
Sampling Frame Complete list from which sample is Telecom subscriber database
drawn
Sampling Unit Each individual element in the frame One individual subscriber
Sample Size (n) Number of elements in the sample n = 2,000
Parameter Descriptive measure of a population Population mean income
Statistic Descriptive measure of a sample Sample mean income
Sampling Error Difference between sample result ±3% margin of error
and true population result
Non-Sampling Error Errors from sources other than Question misunderstood by
sampling (data entry, bias) respondents

5.2 Probability Sampling Techniques


Definition
In probability sampling, every element of the population has a KNOWN, NON-ZERO chance of
being selected. Results can be generalized to the whole population. Sampling error can be
measured.

1. Simple Random Sampling (SRS)


Every member has an EQUAL and INDEPENDENT chance of selection. Done via random number
tables, lottery, or computer.
Feature Detail
Method Lottery system / Random number generator
Best For Small, homogeneous, easily accessible populations
Advantage Completely unbiased; simple to understand
Disadvantage Impractical for large populations; requires complete sampling frame
Example A researcher randomly selects 200 students from a list of 2,000 enrolled
students using a random number generator.

2. Stratified Random Sampling


Population is divided into mutually exclusive subgroups (strata) based on a shared characteristic, then a
random sample is drawn from EACH stratum.
Feature Detail
Method Proportional stratification (equal %) or disproportional (different %)
Best For Heterogeneous populations with known subgroups
Advantage Ensures representation of all subgroups; reduces sampling error
Disadvantage Requires prior knowledge of population composition
Example Dividing 10,000 customers by income: Low (40%), Medium (40%), High (20%).
Sample proportionally: 400, 400, 200 from each stratum.

3. Cluster Sampling
Population is divided into clusters (often geographic). Some clusters are RANDOMLY SELECTED; all
members within selected clusters are studied.
Feature Detail
Types Single-stage (select all in cluster) vs. Two-stage (random sample within cluster)
Best For Geographically dispersed populations where a frame is unavailable
Advantage Reduces cost and time dramatically
Disadvantage Higher sampling error; clusters must be representative
Example A national market research firm randomly selects 30 districts across India, then
surveys all households within those districts about FMCG usage.

4. Systematic Sampling
Every k-th element is selected from the population list after a random start. k = Population size / Sample
size.
Feature Detail
Formula k = N/n (sampling interval); Start at random number between 1 and k
Feature Detail
Best For Large lists where random selection is difficult
Advantage Simple, fast, spread across entire population
Disadvantage Periodicity bias if list has a cyclical pattern
Example From 10,000 alumni, k=10. Start at #7. Select: 7, 17, 27, 37, ... (n=1,000 alumni).

5. Multi-Stage Sampling
Combines two or more sampling techniques in successive stages. Common in large national studies.
• Stage 1: Cluster by state → Stage 2: Stratify by city size → Stage 3: Randomly select
respondents
• Example: NSSO surveys in India use multi-stage sampling across districts, villages, and
households.

5.3 Non-Probability Sampling Techniques


Definition
In non-probability sampling, selection is based on the researcher's judgment or convenience.
Not all members have an equal chance of selection. Results CANNOT be statistically
generalized to the full population. Used for exploratory research.

Technique Description Pros / Cons Example


1. Convenience Selecting whoever is Quick and cheap; high Surveying shoppers at a
Sampling most easily available or non-response bias mall entrance near the
accessible. researcher's office.
2. Purposive Researcher selects Targeted; high bias Selecting only senior
(Judgmental) sample based on own potential marketing managers for a
Sampling judgment about who is study on B2B brand
most useful for the study. strategy.
3. Quota Sampling Researcher sets quotas Ensures subgroup Interview 50 men and 50
for subgroups (like coverage; not truly women for a total of 100
stratified) but selection representative respondents to ensure
within quotas is NOT gender balance.
random.
4. Snowball Existing respondents Useful for special Researching users of a
Sampling refer new participants. groups; high referral rare medical device —
Used for rare or hard-to- bias initial contacts refer others
reach populations. with the same condition.
5. Self-Selection Respondents volunteer Very easy to collect; 'Click here to share your
Sampling themselves (e.g., online extreme self-selection opinion about our app' —
survey with open link). bias open to anyone who visits
the website.
Probability vs. Non-Probability Sampling
Feature Probability Non-Probability
Selection Random — equal chance Non-random —
judgment/convenience
Generalizability High — can generalize to population Low — cannot generalize
Sampling Error Measurable and controllable Cannot be measured
Time & Cost Higher Lower
Research Purpose Conclusive, descriptive Exploratory, qualitative
Example Stratified survey of 1,000 voters Focus group of 12 target customers
UNIT 6: DATA COLLECTION

6. Data Collection
Definition
Data collection is the systematic process of gathering information from relevant sources to
address the research objectives. It involves deciding what data is needed, from whom, how,
when, and where.
The quality of data collected directly determines the quality of research conclusions.

Primary Data Collection Methods


Method Description Advantage Disadvantage Example
Surveys / Structured Economical, Low response Post-purchase email
Questionnaires questions large samples rate, response survey by Flipkart
distributed to possible bias
respondents
(online, mail,
phone, personal)
Personal Face-to-face verbal High quality Expensive, In-depth homemaker
Interviews interaction responses, time- interviews on detergent
(structured, semi- probing consuming, brands
structured, possible interviewer
unstructured) bias
Telephone Interviews Fast, Declining Political opinion polls
Interviews conducted over moderate cost response before elections
phone rates, no
visual cues
Focus Groups Moderated group Rich Groupthink, Consumers discussing
discussion with 6- qualitative not new cola flavor
12 participants insights, generalizable preferences
group synergy
Observation Researcher Actual Observer bias, Retail store tracking
observes behavior behavior expensive, customer movement via
without interaction captured, no time- sensors
(natural or recall bias consuming
controlled)
Experiment Manipulate Establishes Artificial A/B test on two website
independent causality conditions, landing page designs
variable, measure expensive
effect on dependent
variable
Panel Research Same respondents Trend Panel fatigue, Nielsen tracking 1,000
tracked over time analysis, panel attrition households' TV habits for
behavioral 1 year
Method Description Advantage Disadvantage Example
(consumer panels, change
retail panels) tracking
Projective Indirect methods to Reveals Subjective Brand personality
Techniques uncover hidden unconscious interpretation association: 'If Pepsi were
attitudes (word attitudes a person, who would it
association, TAT) be?'

Secondary Data Collection Methods


Source Examples
Internal Sources Sales records, CRM data, financial reports, customer complaints, past
research
Government Sources Census data, NSSO reports, RBI bulletins, Ministry of Statistics
Industry Reports FICCI, CII, NASSCOM, ASSOCHAM industry surveys
Academic Sources Journals (Journal of Marketing, Harvard Business Review), theses
Online Databases Statista, IBIS World, Bloomberg, Euromonitor, Prowess
Syndicated Data Nielsen, IMRB, Kantar retail and consumer panels
UNIT 7: DATA PREPARATION

7. Data Preparation
Definition
Data preparation is the process of transforming raw, collected data into a clean, organized, and
analysis-ready format. It is a critical bridge between data collection and statistical analysis.
Poor data preparation leads to garbage-in, garbage-out (GIGO) — invalid analysis results.

Step Description Methods/Tools Example


Step 1: Editing Reviewing collected Field editing (on-site Questionnaire with age
questionnaires for review), Office editing '999' or salary blank is
completeness, legibility, (central review) flagged for correction or
consistency, and deletion.
accuracy.
Step 2: Coding Assigning numerical codes Pre-coding (before Male=1, Female=2;
to all responses to enable fieldwork) for closed SD=1, D=2, N=3, A=4,
data entry and computer questions; Post-coding SA=5 on Likert scale
analysis. (after fieldwork) for
open-ended
Step 3: Data Entry Transferring coded data Double-entry Entering 500
into a database or verification; scanning; questionnaire responses
software (SPSS, Excel, R, optical mark recognition into an SPSS data file
SAS, Stata). (OMR)
Step 4: Data Identifying and Missing value treatment; Age=150 is an outlier;
Cleaning correcting/removing errors, Outlier detection; empty 'income' field
inconsistencies, missing Consistency checks replaced with mean
values, and outliers. income
Step 5: Data Converting data into a Recoding (merging Recoding age into
Transformation form suitable for analysis. categories); Computing groups: 18-25=1, 26-
new variables; 35=2, 36+=3
Normalization;
Standardization
Step 6: Data Organizing data into Simple tabulation (one Gender vs. Brand
Tabulation frequency tables and variable); Cross- preference cross-
cross-tabulations for initial tabulation (two+ tabulation reveals female
inspection. variables) preference for Brand A

Handling Missing Data


Method When to Use Risk
Listwise Deletion (Complete When missing data is random and Reduces sample size
Case) small (<5%) significantly
Method When to Use Risk
Mean/Median Substitution When data is missing at random for Reduces variance, distorts
continuous variables distribution
Regression Imputation When missing is predictable from Complex; can over-fit
other variables
Multiple Imputation Best practice for missing data Computationally intensive
(creates multiple complete datasets)
Keep as Missing Category For categorical variables where 'no Adds category but may bias
response' is meaningful analysis
UNIT 8: VALIDITY AND RELIABILITY

8.1 Validity
Definition
Validity is the degree to which a measurement instrument actually measures what it is intended
to measure. A valid measure is ACCURATE — it captures the true concept.
Mnemonic: Validity = Accuracy. Think of a scale that measures your weight — is it measuring
weight or height?

Types of Validity
Type Definition How to Assess Example
Content Validity Instrument covers all Expert panel review; A 'marketing knowledge' test
dimensions of the literature review must cover all 4Ps, not just
construct pricing
Face Validity Instrument appears to Subjective judgment Experts review job
measure what it by experts satisfaction scale and confirm
claims, on the surface it 'looks right'
Criterion Validity Instrument correlates Concurrent (same New sales aptitude test
with an established time) or Predictive predicts actual sales
external criterion (future criterion) performance after 6 months
Construct Validity Instrument captures Factor analysis; Brand loyalty scale correlates
the theoretical Multitrait-multimethod with repurchase behavior
construct it is matrix (convergent) but not browsing
designed to measure (discriminant)
Convergent Validity Sub-type of construct Average Variance Customer satisfaction and
validity: high Extracted (AVE) > 0.5 customer delight scales
correlation with similar correlate highly
constructs
Discriminant Validity Sub-type: low Correlation between Customer satisfaction does
correlation with distinct constructs < NOT correlate with unrelated
unrelated constructs square root of AVE construct like 'store location
preference'
Internal Validity (For causal research) Random assignment A/B test uses matched
The IV truly causes to groups; control of groups to ensure button color
changes in DV, no extraneous variables (not other factors) caused
confounds conversion difference
External Validity Findings can be Representative Survey results from Delhi
generalized to other sampling; replication customers also apply to
settings, populations, Mumbai customers of same
times demographic
8.2 Reliability
Definition
Reliability is the consistency and stability of a measurement instrument. A reliable measure
produces the SAME result under the same conditions every time it is applied.
Mnemonic: Reliability = Consistency. Think of a bathroom scale that always shows 70 kg when
you are actually 70 kg.

Types of Reliability
Type Definition Assessment Acceptable Example
Method Threshold
Test-Retest Reliability Same instrument Correlation r ≥ 0.7 Brand attitude
administered to between Time 1 scale administered
same group at two and Time 2 in Jan and March; r
different times scores = 0.82 = reliable
Alternate Form Two equivalent Correlation r ≥ 0.7 Two versions of a
Reliability versions of between Form A math test given to
instrument and Form B same students
administered to scores
same group
Split-Half Reliability Instrument split into Pearson r ≥ 0.7 20-item scale split
two halves; scores correlation + into odd/even
correlated Spearman- items; both halves
Brown should correlate
correction
Internal Consistency All items in a scale Alpha (α) α ≥ 0.7 A 10-item brand
(Cronbach's Alpha) measure the same coefficient from acceptable; ≥ loyalty scale gives
underlying construct 0 to 1 0.8 good; ≥ α = 0.83 — good
0.9 excellent reliability
Inter-Rater Reliability Degree of Cohen's Kappa Kappa ≥ 0.7 Two researchers
agreement between or % agreement coding focus group
two or more transcripts should
raters/observers agree ≥70% of the
time

Cronbach's Alpha Interpretation Table


Alpha Value Reliability Level Action
α ≥ 0.9 Excellent Use as-is
0.8 ≤ α < 0.9 Good Use as-is; minor improvement optional
0.7 ≤ α < 0.8 Acceptable Use with caution; consider item revision
0.6 ≤ α < 0.7 Questionable Revise weaker items; pilot test again
α < 0.6 Poor / Unacceptable Redesign scale; items are inconsistent
Relationship Between Validity and Reliability
Critical Rule
A measure can be RELIABLE without being VALID (consistently wrong!).
A VALID measure MUST also be RELIABLE. Reliability is necessary but not sufficient for
validity.
Analogy: A broken clock is reliable (consistently shows 3:00) but not valid. A working clock is
both reliable AND valid.
UNIT 9: HYPOTHESIS TESTING

9. Hypothesis Testing
Definition
Hypothesis testing is a formal statistical procedure for making decisions about population
parameters based on sample data. It determines whether there is sufficient evidence to reject
the null hypothesis (H0) in favor of the alternative hypothesis (H1).
It controls the probability of making incorrect decisions (Type I and Type II errors).

Steps in Hypothesis Testing


Step Action Detail
1 State H0 and H1 H0: No effect/difference. H1: Effect/difference exists.
2 Set Level of Significance (α) α = 0.05 (5%) is standard; α = 0.01 for critical decisions
3 Select Appropriate Statistical Based on data type, number of groups, and research
Test question
4 Calculate the Test Statistic Compute t, z, F, χ² value from sample data using formula
5 Determine p-value or Critical p-value from software (SPSS/R) or critical value from table
Value
6 Make Decision If p ≤ α: Reject H0. If p > α: Fail to reject H0.
7 Interpret Results Translate statistical decision into business/managerial
conclusion

Type I and Type II Errors


H0 is TRUE (No real effect) H0 is FALSE (Real effect exists)

Reject H0 TYPE I ERROR (α) — False CORRECT DECISION — True


Positive: Saying there is an effect Positive (Power = 1-β)
when there isn't
Fail to Reject H0 CORRECT DECISION — True TYPE II ERROR (β) — False
Negative Negative: Missing a real effect

Choosing the Right Statistical Test


Research Question Data Type Test to Use
Compare mean of ONE sample to Continuous One-sample t-test or z-test
known value
Research Question Data Type Test to Use
Compare means of TWO Continuous Independent samples t-test
independent groups
Compare means BEFORE and Continuous Paired samples t-test
AFTER (same group)
Compare means of THREE or Continuous One-way ANOVA
MORE groups
Test association between two Categorical Chi-square test (χ²)
categorical variables
Test linear relationship between two Continuous Pearson Correlation
continuous variables
Predict one variable from another Continuous Simple/Multiple Regression
Classify into groups based on Mixed Discriminant Analysis
multiple variables
UNIT 10: ANOVA (ANALYSIS OF VARIANCE)

10. ANOVA — Analysis of Variance


Definition
ANOVA is a statistical technique used to compare the means of THREE or MORE groups
simultaneously to determine if at least one group mean is significantly different from the others.
It partitions total variance into: variance BETWEEN groups (due to group differences) and
variance WITHIN groups (due to random error).

When to Use ANOVA


• Dependent variable (DV): Continuous/interval/ratio (e.g., sales, satisfaction score)
• Independent variable (IV): Categorical with 3+ groups (e.g., region: North/South/East/West)
• Example: Do customer satisfaction scores differ significantly across 4 retail store locations?

Types of ANOVA
Type Description Example
One-Way ANOVA One independent variable (factor) Effect of 3 ad types (TV, Social, Print)
with 3+ levels on brand recall
Two-Way ANOVA Two independent variables Effect of Ad Type AND Gender on
simultaneously; tests main effects brand recall — does gender moderate
and interaction the ad effect?
MANOVA Multiple dependent variables Effect of training program on sales,
(Multivariate) simultaneously customer satisfaction, AND employee
morale together
Repeated Measures Same subjects measured at multiple Customer satisfaction measured at
ANOVA time points Month 1, 3, and 6 after product launch
ANCOVA ANOVA with a covariate (control Compare ad effectiveness across 3
variable) to reduce error variance regions, controlling for age differences

ANOVA Calculation Components


Source of Variation Symbol Meaning Formula (Conceptual)
Between Groups SSB Variation due to group Sum of squared deviations of
(Treatment) differences group means from grand mean
Within Groups (Error) SSW Variation within each group Sum of squared deviations of
individual scores from their group
mean
Source of Variation Symbol Meaning Formula (Conceptual)
Total SST Total variation in data SST = SSB + SSW
Mean Square MSB SSB divided by df between MSB = SSB / (k-1); k = number
Between of groups
Mean Square Within MSW SSW divided by df within MSW = SSW / (N-k); N = total
sample
F-statistic F Ratio of between to within F = MSB / MSW
variance

Decision Rule
If F-calculated > F-critical (from table) → Reject H0 → At least one group mean is significantly
different.
Equivalently: If p-value < α (0.05) → Reject H0.
ANOVA tells you THAT differences exist, NOT which groups differ. Use POST-HOC tests for
that.

Post-Hoc Tests (After ANOVA)


Purpose
When ANOVA rejects H0, post-hoc tests identify WHICH specific group pairs are significantly
different.

Post-Hoc Test When to Use


Tukey's HSD (Honestly Significant Equal sample sizes; most commonly used; controls Type I error
Difference) well
Bonferroni Correction Conservative; best when few comparisons are made
Scheffe Test Most conservative; appropriate for all types of contrasts
LSD (Least Significant Difference) Liberal; more statistical power but higher Type I error risk
Duncan's Multiple Range Test Multiple group comparison; less conservative than Tukey

ANOVA — Worked Example


Research Question: Do average monthly sales differ across 4 regions (North, South, East, West)?
H0: μN = μS = μE = μW (all regional means are equal)
H1: At least one regional mean is significantly different

Region Mean Monthly Sales (Rs Sample Size


Lakhs)
North 45.2 50
South 52.8 50
Region Mean Monthly Sales (Rs Sample Size
Lakhs)
East 41.5 50
West 48.9 50

F-calculated = 5.84, p = 0.001 < α(0.05) → Reject H0


Conclusion: Regional sales differ significantly. Post-hoc (Tukey) shows South > East (p=0.002) and
South > North (p=0.038).

Assumptions of ANOVA
• Normality: DV is normally distributed within each group (test: Shapiro-Wilk)
• Homogeneity of Variance: Equal variances across groups (test: Levene's Test)
• Independence: Observations are independent of each other
• Note: ANOVA is robust to minor violations of normality for large samples (Central Limit Theorem)
UNIT 11: CORRELATION AND REGRESSION

11.1 Correlation Analysis


Definition
Correlation measures the STRENGTH and DIRECTION of the linear relationship between two
continuous variables. It does NOT imply causation.
Key statistic: Pearson's r (ranges from -1 to +1)

Pearson Correlation Coefficient (r)


Value of r Interpretation Business Example
r = +1.0 Perfect positive correlation Every Rs spent on ads → exactly proportional
sales increase
0.7 ≤ r < 1.0 Strong positive correlation Ad spend strongly increases brand awareness
0.3 ≤ r < 0.7 Moderate positive correlation Customer satisfaction moderately increases
loyalty
0 < r < 0.3 Weak positive correlation Store location weakly related to footfall
r=0 No linear correlation Price of tea unrelated to car sales
-0.3 < r < 0 Weak negative correlation Complaint rate weakly decreases with service
quality
-0.7 < r ≤ -0.3 Moderate negative correlation Price increases moderately decrease demand
-1.0 < r ≤ -0.7 Strong negative correlation Heavy promotion strongly reduces price
perception
r = -1.0 Perfect negative correlation Every price increase → exact proportional sales
decrease

Types of Correlation
Type Variables Statistic Use Case
Pearson Correlation Both continuous r Relationship between ad spend
(interval/ratio) (Rs) and sales (Rs)
Spearman Rank Both ordinal or non- rho (ρ) Relationship between brand rank
Correlation normal continuous and customer rank
Point Biserial One continuous, one rpb Relationship between gender
dichotomous (M/F) and income
Phi Coefficient Both dichotomous φ Relationship between gender
(M/F) and product purchase
(Y/N)
Type Variables Statistic Use Case
Partial Correlation Two variables, r12.3 Ad spend-sales relationship,
controlling for a third controlling for seasonality

Hypothesis Testing for Correlation


• H0: ρ = 0 (no linear relationship between the two variables)
• H1: ρ ≠ 0 (a significant linear relationship exists)
• If p < 0.05 → Reject H0 → Correlation is statistically significant
• Important: r2 (coefficient of determination) = % of variance in Y explained by X
• Example: r = 0.75, r2 = 0.56 → Ad spend explains 56% of variation in sales.

11.2 Regression Analysis


Definition
Regression analysis is a statistical technique that establishes a mathematical relationship
between a DEPENDENT variable (Y) and one or more INDEPENDENT variables (X). It is used
for PREDICTION and understanding the nature of relationships.
Unlike correlation, regression specifies direction: X predicts/causes Y.

Simple Linear Regression (SLR)


One dependent variable (Y) predicted by ONE independent variable (X).
Equation: Y = a + bX + e
Symbol Meaning Example
Y Dependent variable (outcome) Monthly Sales (Rs Lakhs)
X Independent variable Advertising Spend (Rs Lakhs)
(predictor)
a (intercept) Value of Y when X = 0 Baseline sales with zero ad spend
b (slope) Change in Y for one unit Sales increase per Rs 1 lakh ad spend
change in X
e (error term) Unexplained variation in Y Other factors affecting sales

Multiple Linear Regression (MLR)


One dependent variable (Y) predicted by TWO or MORE independent variables (X1, X2, X3...).
Equation: Y = a + b1X1 + b2X2 + b3X3 + ... + e
• Example: Sales = 10 + 2.3(Ad Spend) + 1.8(Distribution) - 0.9(Price) + e
• Interpretation: Controlling for distribution and price, each Rs 1 lakh increase in ad spend
increases sales by Rs 2.3 lakhs.
Key Regression Statistics
Statistic Symbol Meaning Good Value
R-squared R² Proportion of variance in Y Higher is better; >0.6 in
explained by all X variables social sciences
Adjusted R² Adj. R² R² adjusted for number of Compare models with
predictors; prevents inflation with different numbers of
more variables predictors
F-statistic F Tests if the overall regression p < 0.05 → model is
model is significant significant
t-statistic (each β) t Tests if each individual predictor is p < 0.05 for each variable
significant
Beta coefficient β Comparable measure of each Largest |β| = most important
(Standardized) predictor's relative importance predictor
Standard Error SE Measures precision of regression Smaller = more precise
coefficients estimates
VIF (Variance Inflation VIF Detects multicollinearity among VIF < 10; ideally < 5
Factor) predictors

Assumptions of Linear Regression


• Linearity: Relationship between X and Y is linear (check: scatter plot)
• Independence: Observations are independent (no autocorrelation — check: Durbin-Watson)
• Homoscedasticity: Constant variance of residuals (check: residual plot)
• Normality of Residuals: Residuals are normally distributed (check: histogram, Q-Q plot)
• No Multicollinearity: Predictors are not highly correlated with each other (check: VIF < 10)

Types of Regression
Type Dependent Use Case
Variable
Simple Linear Regression Continuous One predictor: Ad spend → Sales
Multiple Linear Regression Continuous Multiple predictors: Ad spend + Price + Distribution
→ Sales
Logistic Regression Binary (0/1) Purchase (Yes/No) predicted by age, income, ad
exposure
Ordinal Regression Ordinal Satisfaction rating (1-5) predicted by service quality
Polynomial Regression Continuous Non-linear curves (e.g., diminishing returns of
advertising)
Ridge / Lasso Regression Continuous When multicollinearity is present (regularized
regression)
UNIT 12: DISCRIMINANT ANALYSIS & LOGIT ANALYSIS

12.1 Discriminant Analysis


Definition
Discriminant Analysis (DA) is a multivariate statistical technique used to CLASSIFY
observations into pre-defined groups (categories) and to identify which independent variables
(predictors) best DISCRIMINATE between those groups.
It answers: 'Given a set of predictor variables, which group does this observation most likely
belong to?'

Purpose of Discriminant Analysis


• Classification: Assign new cases to existing groups based on predictor variables
• Description: Identify which variables best distinguish between groups
• Prediction: Predict group membership of new observations

Marketing Examples
1. Will a customer CHURN or STAY? (churn prediction model)
2. Is a loan applicant a GOOD or BAD credit risk?
3. Which of 3 consumer SEGMENTS does a new customer belong to: Budget, Mid-range, or
Premium?
4. Will a product SUCCEED or FAIL in the market based on test market data?

Types of Discriminant Analysis


Type Groups Description Example
Two-Group DA 2 groups Classic linear discriminant Buyer vs. Non-buyer
analysis
Multiple Discriminant 3+ groups Multiple discriminant functions Low / Medium / High brand
Analysis (MDA) extracted loyalty segment
Linear Discriminant 2 or more Assumes equal covariance Customer segment
Analysis (LDA) across groups classification
Quadratic 2 or more Does NOT assume equal When group covariances
Discriminant Analysis covariance differ substantially
(QDA)

Discriminant Function
The discriminant function is a linear combination of predictor variables that maximally separates groups:
D = b0 + b1X1 + b2X2 + b3X3 + ... + bnXn
Symbol Meaning
D Discriminant Score (Z-score assigned to each observation)
b0 Constant (intercept)
b1, b2, ... Discriminant coefficients (unstandardized or standardized)
X1, X2, ... Predictor variables (age, income, brand attitude, etc.)

Number of Discriminant Functions


Maximum number of discriminant functions = Min(number of groups - 1, number of predictors)
• Example: 3 groups and 10 predictors → Max 2 discriminant functions
• First function maximizes group separation; subsequent functions maximize remaining separation
orthogonally

Steps in Discriminant Analysis


Step Action Details
1 Formulate the Problem Define groups (DV) and select predictors (IVs)
2 Check Assumptions Normality, homogeneity of covariance, no multicollinearity
3 Estimate Discriminant Compute coefficients using SPSS/R/SAS
Function
4 Assess Significance Wilks' Lambda (λ): Lower = better group separation; Chi-square
test
5 Interpret the Function Standardized coefficients show relative importance of
predictors; structure matrix shows correlation of each predictor
with function
6 Validate the Model Classification matrix (hit ratio); Holdout sample; Press's Q
statistic
7 Apply for Classification Use centroids (group means on discriminant function) to
classify new cases

Key Statistics in Discriminant Analysis


Statistic Meaning Interpretation
Wilks' Lambda (λ) Overall test of group λ near 0 = excellent separation; λ near 1 =
separation (0 to 1) poor separation. p < 0.05 = significant
Eigenvalue Variance explained by each Higher eigenvalue = stronger function; ratio of
discriminant function between-group to within-group variance
Canonical Correlation Correlation between Squared value = % variance in groups
discriminant scores and explained by the function
group membership
Statistic Meaning Interpretation
Structure Matrix Correlations between Used to identify/name the function (like factor
predictors and discriminant loadings)
function
Standardized Predictor importance Larger |value| = more important predictor
Coefficients controlling for different
scales
Classification Matrix Table of actual vs. predicted % correctly classified; compare to chance-level
(Hit Ratio) group membership accuracy
Press's Q Statistic Tests if classification is Significant Q (> 6.63) = model classifies better
better than chance than chance

Classification Matrix Example


Predicting customer segment: Budget vs. Mid-Range vs. Premium (n=300, 100 per group)
Actual \ Budget Mid-Range Premium % Correct
Predicted
Budget 82 12 6 82%
Mid-Range 10 78 12 78%
Premium 4 8 88 88%
Overall 82.7%

Hit ratio = 82.7%. Chance accuracy for 3 groups = 33.3%. Model classifies far better than chance.

Assumptions of Discriminant Analysis


• Normality: Each predictor is normally distributed within each group
• Homogeneity of Covariance: Covariance matrices are equal across groups (test: Box's M test; p >
0.05 = assumption met)
• No Multicollinearity: Predictors should not be highly correlated (VIF < 10)
• Adequate Sample Size: Minimum 5 observations per predictor per group; ideally 20 per group
• No Outliers: Outliers can significantly distort discriminant functions

12.2 Logistic Regression (Logit Analysis)


Definition
Logistic Regression (Logit Analysis) is a statistical technique used to predict the probability of a
BINARY (dichotomous) outcome (0/1, Yes/No) based on one or more predictor variables.
Unlike linear regression (which predicts a continuous value), logit predicts PROBABILITIES that
are bounded between 0 and 1 using the logistic function (sigmoid curve).
Why Not Use Linear Regression for Binary Outcomes?
• Linear regression can predict values outside [0,1] — probabilities cannot be negative or >1
• Residuals from binary outcomes are not normally distributed
• Logistic regression solves this by transforming the linear equation using the LOGIT (log-odds)
transformation

The Logistic Model


Logit(p) = ln[p / (1-p)] = a + b1X1 + b2X2 + ... + bnXn
Probability of outcome: P(Y=1) = 1 / [1 + e^-(a + b1X1 + b2X2 + ...)]
Symbol Meaning
p Probability of outcome Y = 1 (e.g., probability of purchase)
ln[p/(1-p)] Log-odds or LOGIT transformation (ensures output is unbounded)
Odds Ratio (OR) = For one unit increase in X, the odds of Y=1 multiply by e^b
e^b
a Constant (intercept)
b1, b2, ... Logistic regression coefficients

Types of Logistic Regression


Type Outcome Categories Example
Binary Logistic 2 categories (0 or 1) Will customer churn? Yes/No
Regression
Multinomial Logistic 3+ unordered Which brand will customer choose: A, B, or C?
Regression categories
Ordinal Logistic 3+ ordered categories Rating: Low / Medium / High satisfaction
Regression

Interpreting Logistic Regression Output


Output Meaning Interpretation Example
Coefficient (b) Log-odds change in b(income)=0.42: higher income increases log-
outcome for one unit odds of premium purchase by 0.42
increase in predictor
Odds Ratio (Exp(b)) Multiplicative change in Exp(0.42)=1.52: each unit increase in income
odds for one unit increase multiplies odds of premium purchase by 1.52
in predictor (52% increase)
Wald Statistic Tests significance of each Wald p < 0.05 = predictor is significant
predictor
Nagelkerke R² Pseudo R² (0 to 1) — how Higher = better fit; analogous to R² in linear
well model fits regression
Output Meaning Interpretation Example
-2 Log Likelihood Overall model fit statistic Lower value = better fit
Hosmer-Lemeshow Goodness-of-fit test p > 0.05 = model fits data well (do NOT want to
Test reject)
Classification Table Actual vs. predicted % Sensitivity (correctly predicting Y=1) and
outcomes Specificity (correctly predicting Y=0)

Discriminant Analysis vs. Logistic Regression


Feature Discriminant Analysis Logistic Regression
DV Type Categorical (2 or more groups) Binary (2 groups); Multinomial for 3+
Assumptions Normality, equal covariance (stricter) Fewer assumptions; no normality
required
IV Type Continuous (metric) only Continuous AND categorical
predictors
Output Discriminant score + classification Probability (0 to 1) + classification
Sample Size Min 20 per group recommended Min 10 events per predictor
Preferred When Assumptions met; metric predictors Assumptions violated; mixed
only predictor types
Popular Use Market segmentation classification Churn prediction, credit scoring,
purchase propensity
UNIT 13: FACTOR ANALYSIS (IN DETAIL)

13. Factor Analysis


Definition
Factor Analysis is a multivariate statistical technique used to REDUCE a large number of
observed variables (items/indicators) into a smaller number of underlying latent variables called
FACTORS, based on the patterns of correlations among the original variables.
It identifies the hidden structure in a dataset — which variables 'go together' because they share
a common underlying dimension.

Purposes of Factor Analysis


• DATA REDUCTION: Reduce many variables to a few meaningful factors (e.g., 30 survey items →
5 factors)
• STRUCTURE IDENTIFICATION: Identify the latent dimensions underlying a set of variables
• SCALE DEVELOPMENT: Validate multi-item scales for constructs (e.g., brand loyalty scale)
• MULTICOLLINEARITY REMOVAL: Replace correlated predictors with uncorrelated factors for
regression
• MARKET SEGMENTATION: Identify underlying consumer needs/motivations from a battery of
attitude items

Types of Factor Analysis


Type Purpose Key Difference
Exploratory Factor Discover underlying factor No pre-specified factor structure; let data
Analysis (EFA) structure when no prior theory decide how many factors and which
exists variables load on each
Confirmatory Factor Test a hypothesized factor Pre-specified which variables load on
Analysis (CFA) structure derived from theory which factors; tests model fit (requires
SEM software like AMOS/LISREL)

Mathematical Model of Factor Analysis


Each observed variable is expressed as a linear combination of factors plus a unique error term:
Xi = ai1F1 + ai2F2 + ... + aipFp + eiUi
Symbol Meaning
Xi The ith observed variable (e.g., survey item score)
Fk The kth common factor (latent, unobserved)
aik Factor loading = correlation between variable i and factor k
Symbol Meaning
Ui Unique factor (specific variance unique to variable i)
ei Error/uniqueness coefficient

Steps in Factor Analysis (Detailed)


Step 1: Formulate the Problem
• Select the variables to be included (should be conceptually related and measured at interval/ratio
level)
• Define the objective: Data reduction, scale development, or structure identification
• Example: 25 brand perception items from a consumer survey are to be reduced to key brand
dimensions

Step 2: Construct the Correlation Matrix


• Compute pairwise correlations among all variables
• If most correlations are low (< 0.30), factor analysis may not be appropriate
• Check: Bartlett's Test of Sphericity — tests if correlation matrix is an identity matrix
• H0: No correlations exist (correlation matrix = identity matrix); Rejection (p < 0.05) is REQUIRED
to proceed
• Check: Kaiser-Meyer-Olkin (KMO) Measure of Sampling Adequacy

KMO Value Interpretation Action


≥ 0.90 Marvelous — Excellent Proceed confidently
0.80 – 0.89 Meritorious — Very Good Proceed
0.70 – 0.79 Middling — Good Proceed
0.60 – 0.69 Mediocre — Acceptable Proceed with caution
0.50 – 0.59 Miserable — Barely Consider dropping variables
Acceptable
< 0.50 Unacceptable Do NOT proceed with factor analysis

Step 3: Select the Extraction Method


Method Description When to Use
Principal Component Extracts maximum variance; Data reduction; maximize explained
Analysis (PCA) treats all variance (common + variance; most commonly used
unique) as common
Principal Axis Factoring Uses communalities; extracts When interest is in latent structure; scale
(PAF) only COMMON variance development
(removes unique/error
variance)
Method Description When to Use
Maximum Likelihood (ML) Estimates factors that are most When distributional assumptions met;
likely to produce the observed CFA or when significance tests needed
correlation matrix
Alpha Factoring Maximizes reliability (alpha) of Scale development focused on reliability
factors
Image Factoring Uses image of each variable When interested in shared variance only
(variance explained by all other
variables)

Step 4: Determine the Number of Factors to Retain


Criterion Rule Example
Kaiser's Eigenvalue Retain factors with Eigenvalue If 25 variables analyzed and 5 have
Rule ≥ 1.0 (most widely used) eigenvalue >1, retain 5 factors
Scree Plot Plot eigenvalues against factor Plot shows clear elbow after factor 4 →
number; retain factors before retain 4 factors
the 'elbow' (point of inflection)
Percentage of Retain factors until cumulative 5 factors explain 68% cumulative variance
Variance variance explained ≥ 60% → retain 5
(social sciences)
A Priori Criterion Decide number of factors Theory predicts 3 brand dimensions →
based on theory BEFORE extract 3 factors
analysis
Parallel Analysis Compare eigenvalues to those Most rigorous criterion; recommended for
from random data; retain if accurate factor count
actual > random

Step 5: Rotate the Factor Matrix (Factor Rotation)


Why Rotate?
The initial (unrotated) factor solution is mathematically correct but often difficult to interpret
because variables may load moderately on many factors.
Rotation redistributes the variance to make the factor structure SIMPLER and more interpretable
(each variable loads high on one factor and low on others — called Simple Structure).

Rotation Type Assumption Methods When to Use


Orthogonal Rotation Factors are Varimax (most common), When factors should be
UNCORRELATED Quartimax, Equamax conceptually independent;
(independent) cleaner interpretation
Oblique Rotation Factors ARE Promax, Direct Oblimin, When factors are
correlated (can Oblique theoretically related (e.g.,
correlate with each service quality dimensions);
other) more realistic
Varimax Rotation (Most Common)
• Orthogonal rotation that maximizes the variance of squared loadings within each factor
• Result: Each factor has a few high loadings and many near-zero loadings — clear simple
structure
• Example: After varimax rotation, Factor 1 loads high on 'reliability', 'consistency', 'dependability' →
named 'Brand Trust'

Step 6: Interpret the Factors


Term Definition Threshold/Rule
Factor Loading Correlation between a Significant: |loading| ≥ 0.40 (practical), ≥ 0.30
variable and a factor (range: (minimum)
-1 to +1)
Communality (h²) Proportion of each variable's h² close to 1.0 is better; h² < 0.30 may indicate
variance explained by ALL variable is poorly explained
retained factors
Eigenvalue Amount of total variance Factor 1 eigenvalue = 5.2 means it explains
explained by each factor 5.2 out of 25 total variance units
% Variance Explained Eigenvalue / Number of Factor 1 explaining 20.8% of total variance
variables × 100 (5.2/25 × 100)
Cumulative % Sum of % variance explained 5 factors explaining 65% cumulative variance
Variance across all retained factors
Cross-Loading A variable loads ≥ 0.40 on Problem: Variable meaning ambiguous;
TWO or more factors consider dropping or retain with caution

Factor Naming Example:


Variable Factor 1 Loading Factor 2 Loading Factor 3 Loading
Brand is trustworthy 0.81 0.12 0.09
Brand is consistent 0.78 0.15 0.11
Brand is dependable 0.74 0.08 0.14
Brand is innovative 0.13 0.79 0.10
Brand is modern 0.09 0.76 0.08
Brand is cutting-edge 0.11 0.73 0.12
Brand is expensive 0.08 0.11 0.82
Brand is premium 0.14 0.09 0.78
Factor Name Brand Trust Brand Innovation Brand Prestige

Step 7: Calculate Factor Scores


• Factor scores are computed values for each respondent on each factor
• Used in subsequent analyses (regression, cluster analysis, ANOVA) as composite scores
• Methods: Regression method (most common), Bartlett method, Anderson-Rubin method
• Alternative: Create composite scale scores by averaging items loading on each factor (if
Cronbach's alpha is acceptable)

Step 8: Validate the Factor Solution


• Split-sample validation: Run factor analysis on two halves of the sample; compare factor
structures
• Replication: Compare factor structure across different samples or studies
• Theory alignment: Do derived factors make conceptual sense?
• Cronbach's Alpha: Calculate internal consistency for each factor (α ≥ 0.70 required)
• Average Variance Extracted (AVE) ≥ 0.50 for each factor (construct validity in CFA)

Complete Factor Analysis Example


Research Context: A researcher conducts a study on consumer perceptions of a retail bank using 20
Likert-scale items. Factor analysis is used to identify underlying service quality dimensions.

Phase Result Interpretation


KMO KMO = 0.84 Meritorious — Factor analysis is appropriate
Bartlett's Test χ² = 2847.3, p = 0.000 Correlation matrix is not an identity matrix —
proceed
Factors Retained 5 factors 5 underlying dimensions identified
(Eigenvalue > 1)
Cumulative Variance 5 factors = 67.4% Acceptable — factors explain majority of
Explained variance
Rotation Method Varimax (orthogonal) Dimensions assumed independent for clarity
Factor 1 — Reliability Items: 'Staff deliver promised Named 'Service Reliability'
service', 'Service is consistent',
'Errors are rare'
Factor 2 — Items: 'Staff respond quickly', Named 'Staff Responsiveness'
Responsiveness 'Waiting time is short', 'Help is
available'
Factor 3 — Empathy Items: 'Staff understand my Named 'Customer Empathy'
needs', 'Staff are caring',
'Personal attention given'
Factor 4 — Assurance Items: 'Staff are Named 'Service Assurance'
knowledgeable', 'Staff are
courteous', 'I feel safe'
Factor 5 — Tangibles Items: 'Branch is clean', Named 'Physical Tangibles'
'Equipment is modern', 'Staff
appear professional'
Cronbach's Alpha α = 0.87, 0.83, 0.79, 0.82, 0.76 All 5 factors meet α ≥ 0.70 — reliable scales
respectively
Assumptions of Factor Analysis
• Variables should be measured at interval or ratio level (or treated as such)
• Variables should be roughly normally distributed (especially for ML extraction)
• Sample size: Minimum 5 observations per variable; ideally n ≥ 200; absolutely minimum n ≥ 100
• No outliers: Outliers distort correlations and loadings
• Variables must be correlated (KMO ≥ 0.60; Bartlett's test significant)
• No perfect multicollinearity: Variables should not be perfectly correlated
UNIT 14: CLUSTER ANALYSIS (IN DETAIL)

14. Cluster Analysis


Definition
Cluster Analysis is a multivariate technique used to CLASSIFY objects (respondents, brands,
products, etc.) into mutually exclusive and exhaustive groups (clusters) such that objects within
a cluster are as SIMILAR as possible to each other, while objects from different clusters are as
DIFFERENT as possible.
It is an UNSUPERVISED technique — there are no pre-defined groups. The data itself
determines the clusters.

Purposes of Cluster Analysis in Marketing


• MARKET SEGMENTATION: Group consumers with similar needs, behaviors, or characteristics
• PRODUCT POSITIONING: Identify which products compete most closely based on consumer
perceptions
• BRAND CLASSIFICATION: Group brands by perceived similarities in a product category
• CUSTOMER PROFILING: Identify customer types for CRM and targeting strategies
• RESEARCH TAXONOMY: Group respondents for further analysis (e.g., input to discriminant
analysis)

Marketing Example
A bank analyzes 50,000 customers on 15 variables (account balance, transaction frequency,
product usage, age, income, channel preference). Cluster analysis reveals 5 distinct customer
segments:
1. 'Mass Market Savers' | 2. 'Digital-First Millennials' | 3. 'High-Net-Worth Investors' | 4. 'Small
Business Owners' | 5. 'Passive Account Holders'
Each segment receives a tailored product offer and communication strategy.

Steps in Cluster Analysis


Step 1: Formulate the Problem
• Select clustering variables that are conceptually relevant and distinguish meaningful segments
• Example: Use psychographic (values, lifestyle), behavioral (usage, loyalty), and demographic
variables for consumer segmentation
• Avoid including irrelevant or redundant variables — they dilute cluster distinctiveness

Step 2: Choose a Similarity / Distance Measure


Cluster analysis groups objects by MAXIMIZING similarity (or minimizing distance) within clusters:
Measure Type Formula/Meaning Best For
Euclidean Distance Distance Straight-line distance between Continuous variables; most
two objects in multi-dimensional common
space
Squared Euclidean Distance Square of Euclidean distance; Emphasizes outlying
Distance amplifies large differences differences
City Block Distance Sum of absolute differences Robust to outliers
(Manhattan) across all dimensions
Distance
Chebychev Distance Maximum absolute difference When extreme differences
Distance across dimensions matter most
Pearson Correlation Similarity Correlation between two objects' When variable patterns
profiles matter more than
magnitude
Jaccard Coefficient Similarity For binary variables — proportion Binary/dichotomous data
of shared presence

Important Note
Before computing distances, STANDARDIZE variables (z-scores) if they are measured on
different scales. Otherwise, variables with large ranges will dominate the distance calculation.

Step 3: Choose the Clustering Procedure


Category Method Description Advantages Disadvantages
Hierarchical Agglomerative Starts with each No need to Sensitive to
(Ward's, Complete, object as its own specify number outliers;
Single, Average cluster; progressively of clusters in computationally
Linkage) merges similar advance; good intensive for large
clusters; visualized via for small samples
dendrogram samples
Hierarchical Divisive Starts with all objects Conceptually Computationally
in one cluster; intuitive; useful intensive; rarely
progressively splits for top-down used in practice
into smaller clusters segmentation
Non- K-Means Clustering Researcher specifies Computationally Must specify k in
Hierarchical k clusters; iteratively efficient; advance; sensitive
assigns objects to handles large to initial seed
nearest centroid and samples selection; only
recalculates centroids (n>1000) works with
until stable continuous data
Non- K-Medoids (PAM) Like K-means but Robust to Slower than K-
Hierarchical uses actual data outliers means
points (medoids) as
cluster centers
Model-Based Gaussian Mixture Assumes data comes Soft cluster Requires
Models (GMM) from a mixture of assignments distributional
Gaussian (probability of assumptions
Category Method Description Advantages Disadvantages
distributions; membership);
estimates parameters handles elliptical
probabilistically clusters
Density- DBSCAN Clusters based on No need to Sensitive to density
Based density of data points; specify k; parameter settings
identifies outliers handles non-
naturally spherical
clusters; robust
to outliers

Step 4: Hierarchical Clustering — Linkage Methods


Linkage Method Definition Formula (Distance Characteristic
Between Clusters)
Single Linkage Distance between Min distance between Tends to produce long,
(Nearest Neighbor) clusters = distance any pair 'chaining' clusters; sensitive to
between the two outliers
closest objects
Complete Linkage Distance = distance Max distance between Produces compact, roughly
(Furthest Neighbor) between the two any pair equal-sized clusters; sensitive
FARTHEST objects to outliers
Average Linkage Distance = average Mean of all pairwise Compromise between single
(UPGMA) distance between distances and complete; less distortion
all pairs of objects
across two clusters
Centroid Method Distance = distance Distance between May produce reversals in
between cluster group means dendrogram
centroids (means)
Ward's Method Minimizes the total Increase in total Produces most compact,
within-cluster sum within-cluster SS equal-sized clusters; MOST
of squared WIDELY USED in marketing
deviations research

Step 5: Determine the Number of Clusters


Method Description How to Apply
Dendrogram Visual tree diagram of The level where merging distance suddenly
Inspection hierarchical clustering; increases = cut point for number of clusters
look for large jumps in
distance
Agglomeration Table showing distances Large increase in 'Coefficients' column between
Schedule at each fusion step two steps indicates optimal cluster count
Elbow Method (K- Plot within-cluster sum of Choose k at the 'elbow' where WSS decrease
means) squares (WSS) vs. starts to flatten
number of clusters
Method Description How to Apply
Silhouette Analysis Measures how similar Silhouette width closer to 1 = well-clustered; near
each object is to its own 0 = borderline; negative = misclassified
cluster vs. other clusters
Calinski-Harabasz Ratio of between-cluster Higher value = better-defined clusters; choose k
Index to within-cluster with maximum index
dispersion
Theory / Business Use domain knowledge If theory suggests 4-5 consumer segments,
Logic about expected number validate statistically
of meaningful segments

Step 6: Interpret and Profile the Clusters


• Calculate mean values of each clustering variable for each cluster (cluster centroids)
• Use ANOVA or Chi-square to test if cluster differences are statistically significant
• Examine non-clustering variables (demographics, behaviors) to create rich cluster profiles
• Name each cluster based on its distinguishing characteristics
• Example: Cluster 1 has high income + high brand loyalty + low price sensitivity → 'Affluent Brand
Advocates'

Step 7: Validate and Apply the Clusters


• Split-sample validation: Run cluster analysis on two subsamples; compare cluster structures
• Predictive validity: Do clusters differ on external variables not used in clustering?
• Discriminant analysis: Use clusters as DV to verify separability of segments
• Application: Develop targeted marketing strategies for each segment

K-Means Clustering — Step-by-Step Algorithm


Step Action Detail
1 Choose k Specify number of clusters (e.g., k=3)
2 Initialize Centroids Randomly select k objects as initial centroids (or use K-means++
for smarter initialization)
3 Assign Objects Assign each object to the nearest centroid based on Euclidean
distance
4 Recalculate Centroids Compute new centroid (mean) for each cluster based on current
assignments
5 Reassign Objects Reassign objects to nearest new centroid
6 Check Convergence Repeat Steps 4-5 until cluster assignments do not change
(convergence)
7 Final Solution Stable cluster assignments represent the final K-means solution

Cluster Analysis Example


Research: Segment 500 retail bank customers into groups based on: Account Balance, Transaction
Frequency, Products Used, Digital Usage Score.

Cluster n Avg Balance Trans/Month Products Digital Segment Name


Score
1 120 Rs 2.1L 8.2 1.2 3.1/10 Low-Value
Traditional
2 180 Rs 8.5L 22.5 3.8 8.7/10 Active Digital
Users
3 95 Rs 45.2L 15.3 6.2 6.5/10 High-Value Multi-
Product
4 65 Rs 82.4L 18.8 8.5 7.2/10 Premium Wealth
Clients
5 40 Rs 1.2L 3.1 1.0 2.2/10 Dormant Low-
Activity

Assumptions and Limitations of Cluster Analysis


Assumption/Limitation Detail How to Handle
No Multicollinearity Highly correlated Run factor analysis first; use factor scores as
clustering variables create inputs
redundancy — over-
weight that dimension
Variable Variables on different Always standardize (z-scores) before clustering
Standardization scales distort distance
measures
Outliers Outliers can create Detect and remove outliers before clustering
artificial single-element (Mahalanobis distance)
clusters or distort
centroids
Sample Size Too small a sample gives Minimum n = 100; ideally n > 200 for
unstable clusters hierarchical; n can be large for k-means
No Unique Solution Different methods/seeds Run multiple methods and compare; choose
may yield different cluster theoretically meaningful solution
structures
No Statistical Test for No definitive statistical test Use multiple criteria (elbow, silhouette, theory)
Number of Clusters for 'correct' k convergently
UNIT 15: INTERDEPENDENCE & INDEPENDENCE
TECHNIQUES

15. Dependence vs. Interdependence Techniques


The Core Distinction
DEPENDENCE TECHNIQUES: At least one variable is designated as the DEPENDENT
(criterion/outcome) variable, and others are INDEPENDENT (predictor) variables. The goal is to
predict/explain the DV from the IVs.
INTERDEPENDENCE TECHNIQUES: NO variable is designated as dependent or independent.
ALL variables are analyzed simultaneously to understand the underlying structure or groupings.
There is no Y to predict.

Complete Classification of Multivariate Techniques


Category Technique DV Type # DVs # IVs Primary
Purpose
DEPENDENCE Multiple Metric 1 Multiple metric Predict
Regression (continuous) continuous DV
from multiple
IVs
DEPENDENCE Logistic Non-metric 1 Metric + Non- Predict
Regression (binary/nominal) metric probability of
categorical
outcome
DEPENDENCE Discriminant Non-metric 1 Multiple metric Classify into
Analysis (categorical) (groups) groups;
identify
discriminating
variables
DEPENDENCE MANOVA Multiple metric Multiple 1+ non-metric Compare
groups on
multiple
continuous
DVs
simultaneously
DEPENDENCE Conjoint Metric/Non- 1 Non-metric Assess
Analysis metric (attributes) relative
importance of
product
attributes on
preference
DEPENDENCE Canonical Metric Multiple Multiple metric Relate two
Correlation sets of metric
Category Technique DV Type # DVs # IVs Primary
Purpose
variables to
each other
DEPENDENCE Structural Metric/latent Multiple Multiple Test complex
Equation causal models
Modeling (SEM) with latent
constructs;
path analysis
+ factor
analysis
DEPENDENCE ANOVA / Metric 1 1+ non-metric (+ Compare
ANCOVA covariate) group means;
control for
covariates
INTERDEPENDENCE Factor Analysis — — All metric Identify latent
structure;
reduce
variables to
factors
INTERDEPENDENCE Cluster Analysis — — All types Group objects
into
homogeneous
clusters
(market
segmentation)
INTERDEPENDENCE Multidimensional — — Similarity/distance Map objects in
Scaling (MDS) perceptual
space (brand
positioning
maps)
INTERDEPENDENCE Correspondence — — Categorical Graphically
Analysis represent
associations
between
categorical
variables

15.1 Dependence Techniques — Detailed Overview


1. Multiple Regression
Feature Details
DV One continuous (metric) variable
IVs Multiple metric variables (can include dummy-coded categorical)
Purpose Predict and explain variation in DV; identify significant predictors and their
relative importance
Feature Details
Output R², F-statistic, β coefficients, t-statistics, p-values
Example Predicting monthly sales (DV) from ad spend, price, distribution coverage, and
brand awareness (IVs)
Key Statistic Standardized β (beta) shows relative importance; R² shows overall model fit

2. Conjoint Analysis
Definition
Conjoint Analysis determines the relative IMPORTANCE of different product attributes in
influencing consumer preferences. Respondents evaluate product profiles (combinations of
attribute levels), and the analysis estimates 'part-worths' for each attribute level.
Marketing Application: Identifying the optimal product configuration that maximizes consumer
utility.

• Example: Laptop conjoint analysis with attributes: Brand (Apple/Dell/HP), Price (Rs
60k/80k/100k), RAM (8GB/16GB/32GB), Battery (8hr/12hr)
• Respondents rank/rate 16 laptop profiles → conjoint estimates that Brand accounts for 35% of
purchase decision, Price 30%, RAM 25%, Battery 10%
• Optimal laptop: Apple + Rs 80k + 16GB + 12hr battery → highest predicted utility

3. Structural Equation Modeling (SEM)


Definition
SEM is an advanced multivariate technique that simultaneously estimates MULTIPLE
interrelated relationships among variables (both observed and latent). It combines factor
analysis (measurement model) and path analysis (structural model).
It can test complex causal chains: e.g., Service Quality → Customer Satisfaction → Customer
Loyalty → Word-of-Mouth

Feature Detail
Measurement Model CFA component: Tests how well indicators measure latent constructs (validity
and reliability)
Structural Model Path analysis component: Tests hypothesized causal relationships between
constructs
Software AMOS, LISREL, SmartPLS (PLS-SEM), Mplus, R (lavaan)
Fit Indices CFI > 0.95, TLI > 0.95, RMSEA < 0.06, SRMR < 0.08 = good fit
Advantage over Handles measurement error explicitly; tests full theoretical models
Regression simultaneously
Example Testing if Brand Trust and Perceived Quality mediate the effect of Marketing
Mix on Brand Loyalty
4. Canonical Correlation Analysis
Definition
Canonical Correlation relates TWO SETS of metric variables to each other. It identifies the linear
combinations of each set (canonical variates) that maximize the correlation between the two
sets.

• Example: Relate a set of MARKETING MIX variables (ad spend, price, promotions, distribution) to
a set of MARKET PERFORMANCE variables (sales, market share, brand awareness, customer
satisfaction)
• Identifies which marketing mix dimensions most strongly drive which performance dimensions
• Key statistic: Canonical Correlation (Rc) — squared value gives shared variance between the two
canonical variates

5. MANOVA (Multivariate Analysis of Variance)


Definition
MANOVA extends ANOVA to test differences among groups on MULTIPLE dependent variables
SIMULTANEOUSLY. It controls the experimentwise Type I error rate better than running
separate ANOVAs.

Feature Detail
DV Multiple continuous variables analyzed simultaneously
IV One or more categorical group variables
Key Test Statistics Pillai's Trace, Wilks' Lambda, Hotelling's Trace, Roy's Largest Root
Follow-up If MANOVA is significant, run univariate ANOVAs + discriminant analysis to
identify which DVs drive differences
Example Does training program (3 types) affect employee performance, job satisfaction,
AND customer satisfaction scores simultaneously?
Advantage Controls Type I error; accounts for correlations among DVs

15.2 Interdependence Techniques — Detailed Overview


1. Multidimensional Scaling (MDS)
Definition
MDS is a technique that creates a PERCEPTUAL MAP — a visual representation of objects
(brands, products, companies) in a low-dimensional space (usually 2D) such that their distances
in the map reflect their perceived (dis)similarity.
Input: Similarity or dissimilarity ratings between pairs of objects. Output: Spatial map with
objects positioned relative to each other.
• Example: 10 soft drink brands rated on pairwise similarity → 2D perceptual map shows Coca-
Cola and Pepsi cluster closely; Red Bull and Monster cluster separately; positioning spaces
reveal opportunities for new brands
• Dimensions must be interpreted from context — axes may represent 'energy vs. refreshment' or
'premium vs. mass-market'
• Key statistic: STRESS value (measures how well the map represents actual (dis)similarities);
Stress < 0.10 = good fit; < 0.05 = excellent

2. Correspondence Analysis
Definition
Correspondence Analysis creates a joint perceptual map of ROWS and COLUMNS from a
cross-tabulation table of categorical data. It shows graphically which categories of one variable
are associated with which categories of another.

• Example: Mapping which consumer demographic profiles are associated with which brand
choices in a cross-tabulation
• Advantage: Works with categorical data; no metric assumptions required
• Output: 2D or 3D perceptual map where proximity = association between row and column
categories

Summary: Choosing the Right Multivariate Technique


Key Question Answer Recommended Technique
Is there a DV to predict? Yes — DV is continuous Multiple Regression
Is there a DV to predict? Yes — DV is binary (0/1) Logistic Regression
Is there a DV to predict? Yes — DV is categorical (3+ Discriminant Analysis (metric IVs) or
groups) Multinomial Logit (mixed IVs)
Is there a DV to predict? Yes — Multiple DVs MANOVA or SEM
Is there a DV to predict? Yes — Complex causal Structural Equation Modeling (SEM)
model with latent constructs
Is there a DV to predict? Yes — Relate two sets of Canonical Correlation
variables
Is there a DV to predict? No — Identify structure Factor Analysis
among variables
Is there a DV to predict? No — Group Cluster Analysis
objects/respondents into
clusters
Is there a DV to predict? No — Map perceptual space Multidimensional Scaling (MDS)
(brand positioning)
Is there a DV to predict? No — Show categorical Correspondence Analysis
associations visually
MASTER REVISION SUMMARY

Unit Topic Key Definition Key Test/Output


1 Theory & Theory=systematic framework; Construct=latent H0 vs H1; Directional vs
Constructs concept; Proposition=relational statement; Non-directional
Hypothesis=testable prediction
2 Research 6-step systematic process from problem definition Problem → Plan →
Process to decision-making Collect → Analyze →
Present → Decide
3 Research Blueprint for data collection: Exploratory (vague Focus groups / Surveys /
Design problems), Descriptive (describe market), Causal A/B tests
(cause-effect)
4 Measurement & Measurement=assigning numbers; Likert (interval scale); 5-
Scaling Scaling=continuum of responses; 4 scale types: point or 7-point
Nominal, Ordinal, Interval, Ratio
5 Sampling Selecting a subset from population; Probability SRS, Stratified, Cluster,
(random) vs Non-probability (judgment) Systematic vs
Convenience, Quota,
Snowball
6 Data Collection Gathering primary/secondary data via surveys, Primary=firsthand;
interviews, observation, experiments Secondary=existing
7 Data Edit→Code→Enter→Clean→Transform→Tabulate; GIGO prevention; mean
Preparation handles missing data and outliers substitution; z-score
outlier detection
8 Validity & Validity=accuracy (measuring right thing); KMO for FA; Cronbach's
Reliability Reliability=consistency (same result each time) alpha ≥0.70; Wilks' λ for
DA
9 Hypothesis Statistical test of H0 using p-value vs α; decision: p≤α → Reject H0; t-test,
Testing reject or fail to reject H0 chi-square, ANOVA,
correlation
10 ANOVA Compare 3+ group means; F=MSB/MSW; p<0.05 Post-hoc: Tukey's HSD;
→ at least one group differs Assumptions: normality,
homogeneity
11 Correlation & Correlation: strength/direction (r); Regression: R²=variance explained;
Regression prediction equation Y=a+bX+e β=predictor importance;
VIF<10
12 Discriminant & DA: classify into groups using metric predictors; Wilks' λ; Hit ratio; Odds
Logit Logit: predict binary outcome using any predictors Ratio; Hosmer-Lemeshow
test
13 Factor Analysis Reduce many variables to few latent factors; EFA KMO≥0.60;
vs CFA; Extraction → Rotation → Interpretation Eigenvalue≥1; Varimax
rotation; Communality;
Factor loadings≥0.40
14 Cluster Analysis Group objects into homogeneous clusters; Dendrogram; Elbow
Hierarchical (Ward's) vs Non-hierarchical (K- method; Silhouette;
means) Centroid profiles
Unit Topic Key Definition Key Test/Output
15 Dependence vs Dependence: has DV to predict; Interdependence: Regression/DA/Logit/SEM
Interdependence no DV — find structure/groupings vs Factor/Cluster/MDS

You might also like