0% found this document useful (0 votes)
4 views37 pages

Correlation and ANOVA Research Methods

The document covers various statistical concepts including correlation coefficients, ANOVA, and regression analysis. It provides definitions, methods for calculation, and examples of application in research. Additionally, it includes self-assessment questions and further readings for deeper understanding.

Uploaded by

corojo5125
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views37 pages

Correlation and ANOVA Research Methods

The document covers various statistical concepts including correlation coefficients, ANOVA, and regression analysis. It provides definitions, methods for calculation, and examples of application in research. Additionally, it includes self-assessment questions and further readings for deeper understanding.

Uploaded by

corojo5125
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Notes

Research Methodology

2. The coefficient of correlation

A. is the square of the coefficient of determination


B. is the square root of the coefficient of determination
C. is the same as r-square
D. can never be negative

3. When the values of two variables move in the opposite directions, correlation is said to be
............................
A. Linear
B. Non-linear
C. Positive
D. Negative

4. When the values of two variables move in the opposite directions, correlation is said to be
............................
A. Linear
B. Non-linear
C. Positive
D. Negative

5. Rank correlation coefficient was discovered by....................................

A. Fisher
B. Spearman
C. Karl Pearson

D. Bowley

6. Spearman’s Rank Correlation Coefficient is usually denoted by....................

A. K
B. A
C. S
D. R

7. Study of correlation among three or more variables simultaneously is called.............

A. Partial correlation
B. Multiple correlation
C. Nonsense correlation
D. Simple correlation

8. Which of these distributions is used for a testing hypothesis?


A. Normal Distribution
B. Chi-Squared Distribution
C. Gamma Distribution

Lovely Professional University 167


Notes

Unit 10: Test of Association

D. Poisson Distribution

9. The Variance of Chi Squared distribution is given as k.

A. True
B. False

10. On account of simple calculation involved, χ2 test is very frequently used by the statistician.

A. True
B. False

11-Zero correlation coefficient between two variables could mean


A. The variables are non-linearly related to each other
B. There is a cause and effect relationship between variables
C. That there is error of measurement in variables
D. None of the above is true

12-If all the scatter of points on two variables lie on a negatively stopped straight line, the
correlation coefficient between the variables would be
A. +1
B. -1
C. Zero
D. None of the above

13-A positive and a negative relationship may have the same strength.

A. True
B. False

14-A non-directional hypothesis predicts a negative correlation between two variables.


A. True
B. False

15-Spearman’s rank order correlation is a parametric statistic.

A. True
B. False

Answers for Self Assessment


1. C 2. B 3. D 4. D 5. C

6. D 7. B 8. B 9. B 10. A

11. A 12. B 13. A 14. B 15. B

Review Questions
1. Show that the coefficient of correlation, r, is independent of change of origin and scale.
[Link] that the coefficient of correlation lies between – 1 and + 1.
3. What is Spearman’s rank correlation? What are the advantages of the coefficient of rank
correlation over Karl Pearson’s coefficient of correlation?

168 Lovely Professional University


Notes

Research Methodology

[Link] can you conclude on the basis of the fact that the correlation between body weight and
annual income were high and positive?

5. From the data given below,

find out Karl Pearson's coefficient of correlation.

1. Suppose we have ranks of 8 students of [Link]. in Statistics and Mathematics. On the basis of
rank we would like to know that to what extent the knowledge of the student in Statistics
and Mathematics is related.
Rank in 52 60 58 39 41 53 47 34
Statistics
40 46 43 54 49 55 48 57
Rank in
Mathematics

2. Enumerate the steps in chi-square calculation.


3. In a study equal number of boys and girls were asked to express their preference for
lecture method and discussion method. The data are given below:

Further Readings
Abrams, M.A., Social Surveys and Social Action, London: Heinemann, 1951.
Arthur, Maurice, Philosophy of Scientific Investigation, Baltimore: John Hopkins
University Press, 1943.
RS. Bhardwaj, Business Statistics, Excel Books, New Delhi, 2008.
S.N. Murthy and U. Bhojanna, Business Research Methods, Excel Books, 2007

Web Links
[Link]
formula/
[Link]
[Link]
[Link]
[Link]

Lovely Professional University 169


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques


Dr. Atif Ghayas, Lovely Professional University

Unit11:Analysis of Variance (ANOVA) and Prediction Techniques


CONTENTS
Objectives
Introduction
11.1 Analysis of variance (ANOVA)
11.2 Reliability and Validity
11.3 Regression Analysis
Summary
Keywords
Self Assessment
Answers for Self Assessment
Review Questions
Further Readings

Objectives
After studying this unit, you will be able to:
 Explain the Concept of Analysis of Variance (ANOVA)
 Discuss reliability and validity

 Define the bivariate regression

 Carry outmultiple regression analysis

Introduction
ANOVA stands for "analysis of variance," and it's a statistical technique for testing a hypothesis
and determining how various groups react to one another by connecting independent and
dependent variables. ANOVA is a statistical test that compares the means of two groups to see
if there is a difference between [Link] is an advanced technique for the experimental treatment of
testing differences among all of the means.
The ANOVA technique allows us to do this simultaneous test and is thus regarded as a valuable
analytical tool in the hands of a researcher. Using this method, one can estimate if the samples
weretaken from populations with the same mean.
Regression analysis is a proven method for determining which variables have an impact on a
certain subject. Regression analysis allows you to confidently establish which elements are most
important, which factors may be ignored, and how these factors interact. Data is at the heart of
regression analysis. It aids businesses in comprehending the data they have and using it –
specifically, the correlations between data points – to make better decisions, ranging from sales
forecasting to inventory levels and supply and demand analysis. Regression analysis is frequently
referred to as one of the most important business analysis approaches.

11.1 Analysis of variance(ANOVA)


ANOVA is a statistical technique. It is used to test the equality of three or more sample means.
Based on the means, inference is drawn whether samples belongs to same population or not.

170 Lovely Professional University


Notes

Research Methodology

Notes: Conditions for using ANOVA

1. Data should be quantitative in nature.

2. Data normally distributed.

3. Samples drawn from a population follow random variation.

ANOVA can be discussed in two parts:


1. One-way classification
2. Two and three-way classification.

1. One-way ANOVA
Following are the steps followed in ANOVA:
1. Calculate the variance between samples.
2. Calculate the variance within samples.
3. Calculate F ratio using the formula. F = Variance between the samples/Variance within the
sample
4. Compare the value of F obtained above in (3) with the critical value of F such as 5% level of
significance for the applicable degree of freedom.
5. The difference in sample means is not significant when the calculated value of F is less than the
table value of F, and the null hypothesis is accepted. When the estimated value of F is greater than
the critical value of F, on the other hand, the difference in sample means is regarded significant, and
the null hypothesis is rejected.

Example: ANOVA is useful.


1. To compare the mileage achieved by different brands of automotive fuel.
2. Compare the first year earnings of graduates of half a dozen top business schools.

Application in Market Research Consider the following pricing experiment. For a new toffee box
introduced by Nutrine Company, three prices are explored. The price of three different types of
toffee boxes is 39, 44, and 49 dollars. The goal is to figure out how price levels affect sales. These
toffee boxes will be shown in five supermarkets. The sales are as follows:

What the manufacturer wants to know is: (1) whether the difference among the means is
significant? If the difference is not significant, then the sale must be due to chance. (2) Do the means
differ? (3) Can we conclude that the three samples are drawn from the same population or not?

Example: In a company there are four shop floors. Productivity rate for three methods of
incentives and gain sharing in each shop floor is presented in the following table. Analyze
whether various methods of incentives and gain sharing differ significantly at 5% and 1% F-
limits.

Lovely Professional University 171


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

Solution:
Step 1: Calculate mean of each of the three samples (i.e., x1, x2 and x3, i.e. different methods
of incentive gain sharing).
5+6+2+7
‾ = =5
4
4+3+2+3
‾ = =3
4
4+3+2+3
‾ = =3
4

‾ ‾ ‾
Step 2: Calculate mean of sample means i.e., ‾‾=

where, K denotes Number of samples= = 4(approximated)


Step 3: Calculate sum of squares (s.s.) for variance between and within the samples.

ss between = n (x − x) + n (x − x) + n (x − x)
ss within= Σ( − ‾ ) + Σ( − ‾ ) + Σ( −‾ )
The sum of squares (ss) for variance between samples is calculated by subtracting the
sample mean deviations from the mean of sample means () and computing the squares of
such deviations, which are then multiplied by the number of items or categories in the
samples to get their total. The sum of squares (ss) for variance within samples is calculated
by subtracting all sample item values from their respective sample averages, squaring the
deviations, and then adding them together. For our illustration then
ss between = 4(5 − 4) + 4(4 − 4) + 4(3 − 4)
= 4+0+4 = 8
{(5 − 5) + (6 − 5) + (2 − 5) + (7 − 5) } {(4 − 4) + (4 − 4) + (2 − 4) + (6 − 4) }
ss within = +
Σ(x − x ) Σ(x − x )
{(4 − 3) + (3 − 3) + (2 − 3) + (3 − 3) }
+
Σ(x − x )
= (0 + 1 + 9 + 4) + (0 + 0 + 4 + 4) + (1 + 0 + 1 + 0)
= 14 + 8 + 2
= 24

Step 4: ss of total variance which is equal to total of s.s. between and ss within and is
denoted by formula as follows:

Σ − ‾̅

Where
= 1.23
= 1.23

for our example, total ss will thus be:


[{(5 − 4) + (6 − 4) + (2 − 4) + (7 − 4) } + {(4 − 4) + (4 − 4) + (2 − 4) + (6 − 4) }
+ {(4 − 4) + (3 − 4) + (2 − 4) + (3 − 4) }]
= {(1 + 4 + 4 + 9) + (0 + 0 + 4 + 4) + (0 + 1 + 4 + 1)}
= 08 + 8 + 6 = 32

We will, however, get the same value if we simply total respective values of ss between and
ss within.
172 For our example, ss between
Lovely is 8Professional
and ss withinUniversity
is 24, thus ss of total variance is 32
(8+24). Step 5: Ascertain degrees of freedom and mean square (MS) between and within the
samples. Degrees of freedom (df) for between samples and within samples are computed
differently as follows. For between samples, df is (k-1), where k' represents number of
Notes

Research Methodology

2. Two-way ANOVA
The approach for calculating variance is identical to that used for one-way classification. The
following is an example of ANOVA two-way classification: Assume a company has four different
types of machines: A, B, C, and D. It has placed four of its employees on each machine for a given
amount of time, such as one week. The average production of each worker on each type of machine
was calculated at the end of one week. These data are given below:
Average Production by the MachineType

The firm is interested in knowing:


1. Whether the mean productivity of workers is significantly different.
2. Whether there is a significant difference in the mean productivity of different types of machines.

Example: Company ‘X’ wants its employees to undergo three different types of
training programme with a view to obtain improved productivity from them. After
the completion of the training programme, 16 new employees are assigned at
random to three training methods and the production performance were recorded.
The training managers’ problem is to find out if there are any differences in the
effectiveness of the training methods? The data recorded is as under
Daily Output of New Employees

Following steps are followed.


Following steps are followed.

1 Calculate Sample mean i.e. ‾

2 Calculate General mean i.e. ‾

3 Calculate variance between columns using the formula ‾ =


∑ ( ‾)
where =( + + − 3)

4 Calculate sample variance. It is calculated using formula:


∑( ‾)
Sample variance = where n is No. of observation under each
method.

5 Calculate variance within columns using the formula ‾ =

between column variance


6 Calculate F using the ratio F =
within column variance

7 Calculate the number of degree of freedom in the numerator F ratio using


equation, d. f = (No. of samples -1).

8 Calculate the number of degree of freedom in the denominator of F ratio


using the equation d.f = ( − )

9 Refer to table 8 find value.

10 Draw conclusions.
Solution:

Lovely Professional University 173


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

1 Sample mean is calculated as follows:


85 105 114
‾ = = 17, ‾ = = 21, ‾ = = 19
5 5 6

2 Grand mean
15 + 18 + 19 + 22 + 11 + 22 + 27 + 18 + 21 + 17 + 24 + 19 + 16 + 22 + 15 + 18
=
16
304
= = 19
16

3 Calculate variance between columns:

∑ ( − ) 40
‾ = = = 20
−1 3−1
4. Calculation sample variance:

∑( ‾) ∑( ‾) ∑( ‾)
Sample variance = = , = , =

70 62 60
= = 17.5, = = 15.5, = = 12
4 4 5

5. Within column variance ‾ = ∑

5−1 5−1 6−1


= × 17.5 + × 15.5 + × 12
16 − 3 16 − 3 16 − 3
4 4 5
= × 17.5 + × 15.5 + × 12
13 13 13

Within column variance = = 14.76


174 Lovely Professional University
Between column variance
6. F = = = 1.354
Within column variance .

7. d.f. of Numerator= (3 − 1) = 2.
Notes

Research Methodology

11.2 Reliability and Validity


There are two criteria to decide whether the scale selected is good or not. They are:
1. Reliability
2. Validity

Reliability Analysis
The degree to which the measurement method is error-free is referred to as reliability. Accuracy
and consistency are two aspects of reliability. If the scale produces the same findings when
repeated measurements are taken under the same conditions, it is said to be reliable.

Example: Attitude towards a product or brand preference.

Reliability can be ensured by using the same scale on the same set of respondents, using the same
method. However, in actual practice, this becomes difficult as:
1. Extent to which a scale produces consistent results
2. Test-retest Reliability: Respondents are administered scales at 2 different times under nearly
equivalent conditions
3. Alternative-form Reliability: 2 equivalent forms of a scale are constructed, then tested with the
same respondents at 2 different times
4. Internal Consistency Reliability:
(a) The consistency with which each item represents the construct of interest
(b) Used to assess the reliability of a summated scale
(c) Split-half Reliability
5. Items constituting the scale divided into 2 halves, and resulting half scores are correlated:
Coefficient alpha (most common test of reliability)
6. Average of all possible split-half coefficients resulting from different splitting of the scale items.

Validity Analysis
The paradigm of validity focused in the question "Are we measuring, what we think, we are
measuring?" Success of the scale lies in measuring "What is intended to be measured?" Of the two
attributes of scaling, validity is the most important.
There are several methods to check the validity of the scale used for measurement:

1. Construct Validity:A sales manager feels that there is a direct link between job satisfaction and
the degree to which a person is an extrovert, as well as the sales force's performance. As a
result, those who have high job satisfaction and outgoing personalities should perform well. If
they don't, the measure's construct validity is called into question.
2. Content Validity:The problem should be clearly defined by the researcher. Determine the
object to be measured. Create a scale that is appropriate for this purpose. Regardless of these
factors, the scale may be criticised for its lack of content validity. Face validity is another term
for content validity. The advent of new packaged foods is one example. When a new packaged
food is introduced, it represents a significant change in flavour. Hundreds of thousands of
people may be urged to try the new packaged meals. People may report that they liked the
new flavour overwhelmingly. Even with such a positive response, the product may
nevertheless fail when it is launched on a commercial basis. So, what's the issue? Perhaps a
vital question was overlooked.

Lovely Professional University 175


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

3. Predictive Validity: This pertains to "How best a researcher can guess the future performance
from the knowledge of attitude score"?
4. Criterion Validity:
(a) Examines whether measurement scale performs as expected in relation to other variables
selected as meaningful criteria, i.e., predicted and actual behavior should be similar.
(b) Addresses the question of what construct or characteristic the scale is actually measuring
5. Convergent Validity: Extent to which scale correlates positively with other measures of the
same construct.
6. Discriminant Validity: Extent to which a measure does not correlate with other constructs
from which it is supposed to differ.
7. Nomological Validity: Extent to which scale correlates in theoretically predicted ways with
measures of different but related constructs.

11.3 Regression Analysis


Regression is often put into two- bivariate and multiple regression analysis

Bivariate Regression
Bivariate Regression, often known as simple regression analysis, is a technique for determining the
strength of a relationship between two variables. The two variables are commonly referred to as X
and Y, with one acting as an independent (or explanatory) variable and the other as a dependent
variable (or outcome variable).
Bivariate Regression Analysis employs a linear regression line (since the relationship between the
variables is considered to be linear) to help measure how the two variables change together in order
to establish the relation.
For a bivariate data (Xi, Yi), i = 1, 2, ...... n, we can have either X or Y as independent variable. If X is
independent variable then we can estimate the average values of Y for a given value of X. The
relation used for such estimation is called regression of Y on X. If on the other hand Y is used for
estimating the average values of X, the relation will be called regression of X on Y. For a bivariate
data, there will always be two lines of regression. It will be shown later that these two lines are
different, i.e., one cannot be derived from the other by mere transfer of terms, because the
derivation of each line is dependent on a different set of assumptions.
The general form of the line of regression of Y on X is YCi = a + bXi, where YCi denotes the average
or predicted or calculated value of Y for a given value of X = Xi. This line has two constants, a and
b. The constant a is defined as the average value of Y when X = 0. Geometrically, it is the intercept
of the line on Y-axis. Further, the constant b, gives the average rate of change of Y per unit change
in X, is known as the regression coefficient. The above line is known if the values of a and b are
known. These values are estimated from the observed data (Xi, Yi), i = 1, 2, ...... n.

Notes: It is important to distinguish between YCi and Yi. Whereas Yi is the


observed value, YCi is a value calculated from the regression equation.

176 Lovely Professional University


Notes

Research Methodology

Using the regression YCi = a + bXi, we can obtain YC1 , YC2 , ...... YCn corresponding to the X
values X1 , X2 , ...... Xn respectively. The difference between the observed and calculated value for a
particular value of X say Xi is called error in estimation of the ith observation on the assumption of
a particular line of regression. There will be similar type of errors for all the n observations. We
denote by e i = Yi – YCi (i = 1, 2,.....n), the error in estimation of the ith observation. As is obvious
from Figure 9.4, ei will be positive if the observed point lies above the line and will be negative if
the observed point lies below the line. Therefore, in order to obtain a Figure of total error, ei¢s are
squared and added. Let S denote the sum of squares of these errors,
i.e.,S= ∑ =∑ ( − )

The regression line can, alternatively, be written as a deviation of Yi from YCi i.e. Yi – YCi = e i or
Yi = YCi + e i or Yi = a + bXi + e i. The component a + bXi is known as the deterministic component
and ei is random component. The value of S will be different for different lines of regression. A
different line of regression means a different pair of constants a and b. Thus, S is a function of a and
b. We want to find such values of a and b so that S is minimum. This method of finding the values
of a and b is known as the Method of Least Squares. Rewrite the above equation as S = S(Yi – a –
bXi)2 (YCi = a + bXi).
The necessary conditions for minima of S are

Equations (1) and (2) are a system of two simultaneous equations in two unknowns a and b, which
can be solved for the values of these unknowns. These equations are also known as normal
equations for the estimation of a and b. Substituting these values of a and b in the regression
equation YCi = a + bXi, we get the estimated line of regression of Y on X

Lovely Professional University 177


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

Expressions for the Estimation of a and b. Dividing both sides of the equation (1) by n, we have

178 Lovely Professional University


Notes

Research Methodology

Lovely Professional University 179


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

Dividing both sides of equation (13) by n, we have ‾= + ‾

180 Lovely Professional University


Notes

Research Methodology

Lovely Professional University 181


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

Remarks: It should be noted here that the two lines of regression are different because these have
been obtained in entirely two different ways. In case of regression of Y on X, it is assumed that the
values of X are given and the values of Y are estimated by minimisingS(Yi – YCi) 2 while in case of
regression of X on Y, the values of Y are assumed to be given and the values of X are estimated by
minimising S(Xi – XCi) 2 . Since these two lines have been estimated on the basis of different
assumptions, they are not reversible, i.e., it is not possible to obtain one line from the other by mere
transfer of terms. There is, however, one situation when these two lines will coincide. From the
study of correlation we may recall that when r = ± 1, there is perfect correlation between the
variables and all the points lie on a straight line. Therefore, both the lines of regression coincide and
hence they are also reversible in this case. By substituting r = ± 1 in equation (12) or (24) it can be
shown that the lines of regression in both the cases become.
−‾ − ‾

Further when = 0, equation (12) becomes ⊙ = ‾ and equation (24) becomes = ‾ . These are the
equations of lines parallel to -axis and -axis respectively. These lines also intersect at the point
( ‾ , ‾ ) and are mutually perpendicular at this point, as shown in Figure.

Multiple Regression Analysis

Multiple Regression Analysis is an extension of two variable regression analysis. In this analysis,
two of more independent variables are used to estimate the values of a dependent variable, instead
of one independent variable.
The objective of multiple regression analysis are:

1. To derive an equation which provides estimates of the dependent variable from values of
the two or more variables independent variables.
2. To obtain the measure of the error involved in using the regression equation as a basis of
estimation.
3. To obtain a measure of the proportion of variance in the dependent variable accounted for
or explained by the independent variables.
Multiple regression equation explains the average relationship between the given variables and the
relationship is used to estimate the dependent variable. Regression equation refers the equation for
estimating a dependent variable.

Example: Estimating dependent variable from the independent variables , … … … … ..


it is known as regression equation of on 2 … … … …Regression equation, when three variables
are involved, is given below:

= + + .

182 Lovely Professional University


Notes

Research Methodology

Where , =estimated value of the dependent variable =independent variables. =


(Constant) the intercept made by the regression plan it gives the value of the dependent variable,
when all the independent variables assume a value equal to zero.
and = Partial regression coefficients or net regression coefficients. . = measures the
amount by which a unit change in is expected to aflect
when is held constant.
Deviations Taken From Actual Means

. = . +
=( − ‾ )
=( − ‾ )
=( − ‾ )

and can be obtained by solving the following equations.

Σ = +
Σ = Σ + Σ
.
. =
.
− −
( − ‾ )= ( − ‾ )+ ( −‾ )
1− 1−

Regression equation of and and is:

− −
( − )= ( − ‾ )+ ( −‾ )
1− 1−

Summary
 ANOVA is a technique of statistics and it is applied to test the equality of three or more
sample means.
 It is an advanced technique for the experimental treatment of testing differences among all
of the means.
 The ANOVA allows to do this simultaneous test and is thus considered as a valuable
analytical tool in the hands of a researcher.
 Regression analysis allows you to confidently establish which elements are most
important, which factors may be ignored.
 Regression is a term used for predicting the value of one variable from the other.
 Least square method is used to fit the line.

Keywords
ANOVA: It is a statistical technique used to test the equality of three or more sample means.
Bivariate Regression: a technique for determining the strength of a relationship between two
variables.

Lovely Professional University 183


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

Regression Equation: If the coefficient of correlation calculated for bivariate data (Xi, Yi), i = 1, 2, n,
is reasonably high and a cause and effect type of relation is also believed to be existing between
them, the next logical step is to obtain a functional relation between these variables. This functional
relation is known as regression equation in statistics.

Reliability Analysis:the extent to which the measurement process is free from errors.

Internal Consistency in Reliability: The consistency with which each item represents the construct
of interest.

Validity Analysis: means "Are we measuring, what we think, we are measuring?"

Self Assessment
1-Analysis of variance is a statistical method of comparing the ________ of several populations.
A. standard deviations
B. variances
C. means
D. proportions

2- The one-way ANOVA is used to test statistical hypotheses concerning:


A. Variances
B. Group Means
C. Standard Deviations
D. None of these

3- ANOVAs cannot be used when testing data collected in educational research as it cannot be
applied to social science.

A. True
B. False

4- What type of data are best analysed in ANOVA?

A. Correlational
B. Random
C. Experimental
D. Simple

5-What is the definition of 'mean square'?

A. A sum of squares divided by its degrees of freedom


B. The square root of the mean
C. The square of the mean
D. A table of means with four cells

6-In regression, the equation that describes how the response variable (y) is related to the
explanatory variable (x) is:

A. the correlation model


B. the regression model
C. used to compute the correlation coefficient
D. None of these alternatives is correct.

184 Lovely Professional University


Notes

Research Methodology

7- In regression analysis, the variable that is being predicted is the

A. response, or dependent, variable


B. independent variable
C. intervening variable
D. is usually x

8- Regression equation is also named as


A. predication equation
B. estimating equation
C. line of average relationship
D. all the above

9- Regression coefficient is independent of


A. origin
B. scale
C. both origin and scale
D. neither origin nor scale.

10- The regression analysis measures ________________ between X and Y.


A. Dependence
B. Independence
C. Both a & b
D. None

11- The ________ sum of squares measures the variability of the sample treatment means
around the overall mean.
A. treatment
B. error
C. interaction
D. total

12- Which of the following is an assumption of one-way ANOVA comparing samples from
three or more experimental treatments?
A. All the response variables within the k populations follow a normal distributions.
B. The samples associated with each population are randomly selected and are independent
from all other samples.
C. The response variables within each of the k populations have equal variances.
D. All of the above.

13- As variability due to chance decreases, the value of F will


A. increase
B. stay the same
C. decrease
D. can’t tell from the given information

14-What do ANOVA calculate?


A. F ratios
B. Z Scores
C. T Scores
D. None

Lovely Professional University 185


Notes

Unit 11: Analysis of Variance (ANOVA) and Prediction Techniques

15- Which of the following assumptions must be met to use an ANOVA?


A. There is homogeneity of variance
B. The dependent variable must be interval or ratio
C. The data must be normally distributed
D. All of these

Answers for Self Assessment


1. C 2. B 3. B 4. C 5. A

6. C 7. A 8. D 9. A 10. A

11. A 12. D 13. A 14. A 15. D

Review Questions
1. What do you think as the reason behind the two lines of regression being different?
2. From the data given below:-

and find out the following:


i) The two regression equations.
ii)The most likely value of X when Y = 41.
iii)The most likely value of Y when X = - 45.

3-Obtain the equations of the two lines of regression for the data given below:

4-In the estimation of regression equation of two variables X and Y the following results were
obtained. X = 90, Y = 70, n = 10, Ȉx 2 =6360; Ȉy 2 = 2860, Ȉxy = 3900 Obtain the two regression
equations.
5- A test was given to five students taken at random from the fifth class of three schools of a town.
The individual scores are

Carry out the analysis of variance


6- Three varieties of coal were analysed by four chemists and the ash-content in the varieties was
found to be as under.

Carry out the analysis of variance.

186 Lovely Professional University


Notes

Research Methodology

7-What is analysis of variance?

8. Distinguish between t-test for difference between means and ANOVA.


9. What is multiple regression? How does it differ from bivariate regression?

FurtherReadings
Abrams, M.A, Social Surveys and Social Action, London: Heinemann, 1951.
Arthur, Maurice, Philosophy of Scientific Investigation, Baltimore: John Hopkins
University Press, 1943.
R.S. Bhardwaj, Business Statistics, Excel Books, New Delhi, 2008.
S.N. Murthy and U. Bhojanna, Business Research Methods, Excel Books, 2007.
A Parasuraman, Dhruv Grewal, Marketing Research, Biztantra
Paneerselvam, R, Research Methods, PHI.

Web Links
[Link]
[Link]
[Link]
[Link]
[Link]

Lovely Professional University 187


Notes

Dr. Atif Ghayas, Lovely Professional University Unit 12: Multivariate Analysis

Unit12: Multivariate Analysis


CONTENTS
Objectives
Introduction
12.1 Multivariate Analysis
12.2 Classification
12.3 Factor Analysis
12.4 Cluster Analysis
12.5 Discriminant Analysis
12.6 Multidimensional Scaling (MDS)
12.7 Conjoint Analysis
Summary
Keywords
Self Assessment
Review Questions
Answers for Self Assessment
Further Readings

Objectives
After studying this unit, you will be able to:
 Explain the concept of multivariate analysis

 Classify the multivariate analysis

 Define the Discriminant Analysis and Conjoint Analysis

 Discuss the Factor Analysis and Cluster Analysis

 State the Multidimensional Scaling (MDS)

Introduction
As the name indicates, multivariate analysis comprises a set of techniques dedicated to the analysis
of data sets with more than one variable. Several of these techniques were developed recently in
part because they require the computational capabilities of modern computers. Multivariate
analysis (MVA) is based on the statistical principle of multivariate statistics, which involves
observation and analysis of more than one statistical variable at a time. In design and analysis, the
technique is used to perform trade studies across multiple dimensions while taking into account
the effects of all variables on the responses of interest. Sometimes, the marketers will come across
situations, which are complex involving two or more variables. Hence, bivariate analysis deals with
this type of situation. Chi-Square is an example of bivariate analysis.

12.1 Multivariate Analysis


In multivariate analysis, the number of variables to be tackled are many.

188 Lovely Professional University


Notes

Research Methodology

Example: The demand for television sets may depend not only on price, but also on the
income of households, advertising expenditure incurred by TV manufacturer and other
similar factors. To solve this type of problem, multivariate analysis is required.

12.2 Classification
Multiple-variate analysis: This can be classified under the following heads:

A. Factor Analysis
B. Cluster Analysis
C. Discriminant Analysis
D. Multidimensional Scaling
E. Conjoint Analysis

12.3 Factor Analysis


The main purpose of Factor Analysis is to group large set of variable factors into fewer factors.
Each factor will account for one or more component. Each factor a combination of many variables.
There are two most commonly employed factor analysis procedures or methods. They are:
1. Principle component analysis
2. Common factor analysis.

When the objective is to summarise information from a large set of variables into fewer factors,
principle component factor analysis is used. On the other hand, if the researcher wants to analyse
the components of the main factor, common factor analysis is used.

Example: Common factor – Inconvenience inside a car. The components may be:
1. Leg room

2. Seat arrangement

3. Entering the rare seat

4. Inadequate dickey space

5. Door locking mechanism.

Principle Component Factor Analysis


Purposes: Customer feedback about a two-wheeler manufactured by a company.
Method: The MR Manager prepares a questionnaire to study the customer feedback. The researcher
has identified six variables or factors for this purpose. They are as follows:
1. Fuel efficiency (A)

2. Durability (Life) (B)

3. Comfort (C)

4. Spare parts availability (D)

5. Breakdown frequency (E)

6. Price (F)

The questionnaire may be administered to 5,000 respondents. The opinion of the customer is
gathered. Let us allot points 1 to 10 for the variables factors A to F. 1 is the lowest and 10 is the
highest. Let us assume that application of factor analysis has led to grouping the variables as
follows:
A, B, D, E into factor-1

Lovely Professional University 189


Notes

Unit 12: Multivariate Analysis

F into Factor -2

C into Factor - 3
Factor - 1 can be termed as Technical factor;

Factor - 2 can be termed as Price factor;

Factor - 3 can be termed as Personal factor.

For future analysis, while conducting a study to obtain customers’ opinion, three factors mentioned
above would be sufficient. One basic purpose of using factor analysis is to reduce the number of
independent variables in the study. By having too many independent variables, the M.R study will
suffer from following disadvantages:
1. Time for data collection is very high due to several independent variables.
2. Expenditure increases due to the time factor.
3. Computation time is more, resulting in delay.

4. There may be redundant independent variables.

Did you Know?


What is correspondence analysis?

Correspondence analysis is a descriptive/exploratory technique designed to analyze simple


two-way and multi-way tables containing some measure of correspondence between the rows
and columns.

The results provide information which is similar in nature to those produced by Factor Analysis
techniques, and they allow one to explore the structure of categorical variables included in the
table. The most common kind of table of this type is the two-way frequency cross-tabulation table.

In a typical correspondence analysis, a cross-tabulation table of frequencies is first standardized, so


that the relative frequencies across all cells sum to 1.0. One way to state the goal of a typical
analysis is to represent the entries in the table of relative frequencies in terms of the distances
between individual rows and/or columns in a low-dimensional space.

Example: Following are the data on the drinking habits of different employees in an
organization:

One may think of the 4 column values in each row of the table as coordinates in a 4-
dimensionalspace, and one could compute the (Euclidean) distances between the 5
row points in the 4dimensional space. The distances between the points in the 4-
dimensional space summarize all information about the similarities between the
rows in the table above. Now suppose one could find a lower-dimensional space, in
which to position the row points in a manner that retains all, or almost all, of the
information about the differences between the rows. You could then present all
information about the similarities between the rows (types of employees in this
case) in a simple 1, 2, or 3-dimensional graph. While this may not appear to be
particularly useful for small tables like the one shown above, one can easily imagine

190 Lovely Professional University


Notes

Research Methodology

how the presentation and interpretation of very large tables (e.g., differential
preference for 10 consumer items among 100 groups of respondents in a consumer
survey) could greatly benefit from the simplification that can be achieved via
correspondence analysis (e.g., represent the 10 consumer items in a two-
dimensional space).

Rotation in Factor Analysis


Rotation is the step-in factor analysis that permits you to identify meaningful factor names or
descriptions like these.

Linear Functions of Predictors


To identify with rotation, first consider a problem that doesn’t involve factor analysis. Suppose you
want to predict the grades of college students (all in the same college) in many dissimilar courses,
from their scores on general “verbal” and “math” skill tests. To build up predictive formulas, you
have a body of past data consisting of the grades of numerous hundred previous students in these
courses, plus the scores of those students on the math and verbal tests. To predict grades for
present and future students, you might use these data from past students to fit a series of two-
variable multiple regressions, each regression forecasting grade in one course from scores on the
two skill tests.
At present suppose a co-worker suggests summing each student’s verbal and math scores to obtain
a composite “academic skill” score I’ll call AS and taking the difference among each student’s
verbal and math scores to obtain a second variable I’ll call VMD (verbal-math difference). The co-
worker advises running the same set of regressions to predict grades in individual courses, except
using AS and VMD as predictors in each regression, instead of the original verbal and math scores.
In this instance, you would get exactly the same predictions of course grades from these two
families of regressions: one predicting grades in individual courses from verbal and math scores,
the other predicting the identical grades from AS and VMD scores. In fact, you would get the same
predictions if you formed composites of 3 math + 5 verbal and 5 verbal + 3 math and ran a series of
two-variable multiple regressions forecasting grades from these two composites. These examples
are all linear functions of the original verbal and math scores.
The vital point is that if you have m predictor variables, and you replace the m original predictors
by m linear functions of those predictors, you usually neither gain nor lose any information— you
could if you wish use the scores on the linear functions to rebuild the scores on the original
variables. But multiple regression uses whatever information you have in the optimum way (as
measured by the sum of squared errors in the current sample) to forecast a new variable (e.g.
grades in a particular course). Since the linear functions contain the same information as the
original variables, you get the similar predictions as before.
Specified that there are lots of ways to get exactly the same predictions, is there any advantage to
using one set of linear functions rather than another? Yes there is; one set might be simpler than
another. One particular pair of linear functions may enable many of the course grades to be
forecasted from just one variable (that is, one linear function) rather than from two. If we regard
regressions with less predictor variables as simpler, then we can ask this question: Out of all the
possible pairs of predictor variables that would give the same predictions, which is simplest to use,
in the logic of minimizing the number of predictor variables needed in the typical regression? The
pair of predictor variables maximizing some measure of minimalism could be said to have simple
structure. In this example involving grades, you might be able to predict grades in some courses
correctly from just a verbal test score and predict grades in other courses accurately from just a
math score. If so, then you would have achieved a “simpler structure” in your predictions than if
you had used both tests for each and every predictions.

Simple Structure in Factor Analysis


The points of the preceding section are relevant when the predictor variables are factors. Think of
the m factors F as a set of independent or predictor variables, and imagine of the p observed
variables X as a set of dependent or criterion variables. Think a set of p multiple regressions, each
predicting one of the variables from all m factors. The standardized coefficients in this set of
regressions structure a p x m matrix called the factor loading matrix. If we replaced the original
factors by a set of linear functions of those factors, we would get just the same predictions as before,
but the factor loading matrix would be different. So we can ask which, of the many possible sets of
linear functions we might use, produces the simplest factor loading matrix. Specially we will define

Lovely Professional University 191


Notes

Unit 12: Multivariate Analysis

simplicity as the number of zeros or near-zero entries in the factor loading matrix—the more zeros,
the simpler the structure. Rotation does not alter matrix C or U at all, but does transform the factor
loading matrix.
In the intense case of simple structure, each X-variable will have merely one large entry, so that all
the others can be ignored. But that would be a simpler structure than you would usually expect to
achieve; after all, in the real world each variable isn’t in general affected by only one other variable.
You then name the factors subjectively, based on an examination of their loadings.
In common factor analysis the procedure of rotation is in fact somewhat more abstract that I have
implied here, since you don’t actually know the individual scores of cases on factors. However, the
statistics for a multiple regression that is mainly relevant here—the multiple correlation and the
standardized regression slopes—can all be calculated just from the correlations of the variables and
factors involved. So we can base the calculations for rotation to simple structure on just those
correlations, devoid of using any individual scores.

A rotation which necessitates the factors to remain uncorrelated is an orthogonal rotation, while
others are oblique rotations. Oblique rotations regularly achieve greater simple structure, though at
the cost that you have to also consider the matrix of factor intercorrelations when interpreting
results. Manuals are usually clear which is which, but if there is ever any ambiguity, a simple rule is
that if there is any capability to print out a matrix of factor correlations, then the rotation is oblique,
as no such capacity is needed for orthogonal rotations.

12.4 Cluster Analysis


Cluster Analysis is used:
1. To classify persons or objects into small number of clusters or group.
2. To identify specific customer segment for the company’s brand.

Cluster Analysis is a technique used for classifying objects into groups. This can be used to sort
data (a number of people, companies, cities, brands or any other objects) into homogeneous groups
based on their characteristics.
The result of Cluster Analysis is a grouping of the data into groups called clusters. The researcher
can analyse the clusters for their characteristics and give the cluster, names based on these.
Where can Cluster Analysis be applied?
The marketing application of cluster analysis is in customer segmentation and estimation of
segment sizes. Industries, where this technique is useful include automobiles, retail stores,
insurance, B-to-B, durables and packaged goods. Some of the well-known frameworks in consumer
behaviour (like VALS) are based on value cluster analysis.
Cluster Analysis is applicable when:
1. An FMCG company wants to map the profile of its target audience in terms of life-style,
attitude and perceptions.
2. A consumer durable company wants to know the features and services a consumer takes into
account, when purchasing through catalogues.
3. A housing finance corporation wants to identify and cluster the basic characteristics, lifestyles
and mindset of persons who would be availing housing loans. Clustering can be done based
on parameters such as interest rates, documentation, processing fee, number of installments
etc.

Process
There are two ways in which Cluster Analysis can be carried out:
1. First, objects/respondents are segmented into a pre-decided number of clusters. In this case,
a method called non-hierarchical method can be used, which partitions data into the specified
number of clusters

2. The second method is called the hierarchical method.

192 Lovely Professional University


Notes

Research Methodology

The above two are basic approaches used in cluster analysis. This can be used to segment customer
groups for a brand or product category, or to segment retail stores into similar groups based on
selected variables.

Interpretation of Results
Ideally, the variables should be measured on an interval or ratio scale. This is because the clustering
techniques use the distance measure to find the closest objects to group into a cluster. An example
of its use can be clustering of towns similar to each other which will help decide where to locate
new retail stores.
If clusters of customers are found based on their attitudes towards new products and interest in
different kinds of activities, an estimate of the segment size for each segment of the population can
be obtained, by looking at the number of objects in each cluster.
Marketing strategies for each segment are fine-tuned based on the segment characteristics. For
instance, a segment of customers, like sports car, get a special promotional offer during specific
period.

Example: In cluster analysis, the following five steps to be used:


1. Selection of the sample to be clustered (buyers, products, employees).

2. Definition on which the measurement to be made (E.g.: product attributes, buyer


characteristics, employees’ qualification).

3. Computing the similarities among the entities.

4. Arrange the cluster in a hierarchy.

5. Cluster comparison and validation.

Did you know?

Names can also be given to clusters to describe each one. For example, there can be a
cluster called “neo-rich”. Segments are prioritized based on their estimated size.

Cluster Analysis on Three Dimensions


The example below shows Cluster Analysis based on three dimensions age, income and family size.
Cluster Analysis is used to segment the car-buying population in a Metro. For example “A” might
represent potential buyers of low end cars. Example: Maruti 800 (for common man). These are
people who are graduating from the two-wheeler market segment. Cluster “B” may represent mid-
population segment buying Zen, Santro, and Alto etc. Cluster “C” represents car buyers, who
belong to upper strata of society. Buyers of Lancer, Honda city etc. Cluster “D” represents the
super-rich cluster, i.e., Buyers of Benz, BMW, etc.

Lovely Professional University 193


Notes

Unit 12: Multivariate Analysis

Example: Suppose there are five attributes, 1 to 5, on which we are judging two
objects A and B. The existence of an attribute may be indicated by 1 and its absence
by 0. In this way, two objects are viewed as similar if they share common attributes.

One measure of simple matching S is given by:

+
=
+ + +

Where
a = No. of attributes possessed by brands A and B
b = No. of attributes possessed by brand A but not by brand B
c = No. of attributes possessed by brand B but not by brand A
d = No. of attributes not possessed by both brands.

Substituting, we get = = = 0.43


A and B’s association is to be the extent of 43%. It is now clear that object A possess attributes 1, 4,
and 7 while object B possess the attributes 3, 4 and 5. A glance at the above table will indicate that
objects A and B are similar in respect of 2 (0 & 0), 6 (0 & 0) and 4 (1 & 1). In respect of other
attributes, there is no similarity between A and B. Now we can arrive at a simple matching measure
by (a) counting up the total number of matches - either 0, 0 or 1, (b) dividing this number by the
total number of attributes.
Symbolically SAB = M/N
SAB = Similarity between A and B
M = Number of attributes held in common (0 or 1)
N = Total number of attributes
SAB = 3/7 = 0.43
i.e., A & B are similar to the extent of 43%.

SPSS Command for Cluster Analysis


Stage 1
Enter the input data along with variable and value labels in an SPSS file.
1. Click on STATISTICS at the SPSS menu bar.

2. Click on CLASSIFY followed by HIERARCHICAL CLUSTER.

3. Dialogue box will appear select all the variables which are required to be used in cluster
analysis. This can be done by clicking on the right arrow to transfer them from the variable
list on the left.

4. Click on METHOD. The dialogue box will open. Choose "Between Groups Linkage" as the
CLUSTER METHOD.

5. Click CONTINUE to return to main dialogue box.

6. Click STATISTICS on the main dialogue box. Choose "Agglomeration schedule" so that it will
appear in the final output click CONTINUE.

7. Choose DENDROGRAM then on the box called ICICLE, Choose "All Clusters" and "Vertical".

8. Click OK on the main dialogue box to get the output of the hierarchical cluster analysis.

194 Lovely Professional University


Notes

Research Methodology

Stage 2
This stage is used to know how many clusters are required. This stage is called K- MEANS
CLUSTERING.
1. Click CLASSIFY, followed by K- FANS CLUSTER desired.

2. Fill in the desired number of clusters that has been identified from stage 1.

3. Click OPTIONS on the main dialogue box. Select "Initial Cluster Centers". Then click
CONTINUE to return to the main dialogue box.

4. Click OK on the main dialogue box to get the output which has final clusters.

12.5 Discriminant Analysis


In this analysis, two or more groups are compared. In the final analysis, we need to find out
whether the groups differ one from another.

Example: Where discriminant analysis is used


1. Those who buy our brand and those who buy competitors’ brand.

2. Good salesman, poor salesman, medium salesman

3. Those who go to Food World to buy and those who buy in a Kirana shop.

4. Heavy user, medium user and light user of the product.

Suppose there is a comparison between the groups mentioned as above along with demographic
and socio-economic factors, then discriminant analysis can be used. One way of doing this is to
proceed and calculate the income, age, educational level, so that the profile of each group could be
determined. Comparing the two groups based on one variable alone would be informative but it
would not indicate the relative importance of each variable in distinguishing the groups. This is
because several variables within the group will have some correlation which means that one
variable is not independent of the other.
If we are interested in segmenting the market using income and education, we would be interested
in the total effect of two variables in combinations, and not their effects separately. Further, we
would be interested in determining which of the variables are more important or had a greater
impact. To summarize, we can say, that Discriminant Analysis can be used when we want to
consider the variables simultaneously to take into account their interrelationship.
Like regression, the value of dependent variable is calculated by using the data of independent
variable.

Z = b1x1 + b2x2 + b3x3+..............


Z = Discriminant score

b1 = Discriminant weight for variable

x = Independent variable

As can be seen in the above, each independent variable is multiplied by its corresponding
weightage.
This results in a single composite discriminant score for each individual. By taking the average of
discriminant score of the individuals within a certain group, we create a group mean. This is
known as centroid. If the analysis involves two groups, there are two centroids. This is very similar
to multiple regression, except that different types of variables are involved.

Application
A company manufacturing FMCG products introduces a sales contest among its marketing
executives to find out “How many distributors can be roped in to handle the company’s product”.
Assume that this contest runs for three months. Each marketing executive is given target regarding
number of new distributors and sales they can generate during the period. This target is fixed and

Lovely Professional University 195


Notes

Unit 12: Multivariate Analysis

based on the past sales achieved by them about which, the data is available in the company. It is
also announced that marketing executives who add 15 or more distributors will be given a Maruti
Omni-van as prize. Those who generate between 5 and 10 distributors will be given a two-wheeler
as the prize. Those who generate less than 5 distributors will get nothing. Now assume that 5
marketing executives won a Maruti van and 4 won a two-wheeler.
The company now wants to find out, “Which activities of the marketing executive made the
difference in terms of winning a prize and not winning the prize”. One can proceed in a number of
ways. The company could compare those who won the Maruti van against the others.
Alternatively, the company might compare those who won, one of the two prizes against those
who won nothing. It might compare each group against each of the other two.
Discriminant analysis will highlight the difference in activities performed by each group members
to get the prize. The activity might include:
1. More number of calls made to the distributors.

2. More personal visits to the distributors with advance appointments.

3. Use of better convincing skills.

Discriminant analysis answers the following questions:


1. What variable discriminates various groups as above; the number of groups could be
two or more? Dealing with more than two groups is called Multiple Discriminant
Analysis (M.D.A.).

2. Can discriminating variables be chosen to forecast the group to which the


brand/person/ place belong to?

3. Is it possible to estimate the size of different groups?

SPSS Commands for Discriminate Analysis


Input data has to be typed in an SPSS file.
1. Click on STATISTICS at the SPSS menu bar.

2. Click on CLASSIFY followed by DISCRIMINANT.

3. Dialogue box will appear. Select the GROUPING VARIABLE. This can be done by clicking on
the right arrow to transfer them from the variable list on the left to the grouping variable box
on the right.

4. Define the range of values by clicking on DEFINE RANGE. Enter Minimum and Maximum
value then click CONTINUE.

5. Select all the independent variable for discriminant analysis from the variable list by clicking
on the arrow that transfers them to box on the right.

6. Click on STATISTICS on the lower part of main dialogue box. This will open up a smaller
dialogue box.

7. Click on CLASSIFY on the lower part of the main dialogue box select SUMMARY TABLE
under the heading DISPLAY in a small dialogue box that appears.

8. Click OK to get the discriminant analysis output.

12.6 Multidimensional Scaling (MDS)


In addition to fulfilling the goals of detecting underlying structure and data reduction that is shares
with other methods, multidimensional scaling (MDS) provides the researcher with a spatial
representation of data that can facilitate interpretation and reveal relationships. Therefore, we can
define MDS as “a set of multivariate statistical methods for estimating the parameters in and
assessing the fit of various spatial distance models for proximity data.”
The spatial display of data provided by MDS is why it is also sometimes referred to as perceptual
mapping. MDS has much more flexibility about the types of data that can be used to generate the
solution. Almost any measures of similarity and dissimilarity can be used, depending on what your
statistical computer software will accept.

196 Lovely Professional University


Notes

Research Methodology

Types of MDS
In general, there are two types of MDS:
1. Metric

2. Non-metric

Metric MDS makes the assumption that the input data is either ratio or interval data, while the non-
metric model requires simply that the data be in the form of ranks. Therefore, the nonmetric model
has fewer restrictions than the metric model, but also less rigor. One technique to use if you are
unsure whether your data is ordinal or can be considered interval is to try both metric and non-
metric models. If the results are very close, the metric model may be used.
An advantage of the non-metric models is that they permit the researcher to categorize and
examine preference data, such as the kind obtained in marketing studies or other areas where
comparisons are useful.
Another technique, correspondence analysis, can work with categorical data, i.e., data at the
nominal level of measurement, however that technique will not be described here.

Notes:
Similarities and Differences between Factor Analysis and MDS
We have already seen that MDS can accept more different measures of similarity and
dissimilarity than factor analysis techniques can. In addition, there are some
differences in terminology. These differences reflect the origin of MDS in the field of
psychology. The measure corresponding to factors are called alternatively
dimensions or stimulus coordinates.
The output of MDS looks very similar to that of factor analysis and the determination
of the optimal number of dimensions is handled in much the same way.

Steps in using MDS

There are four basic steps in MDS:


1. Data collection and formation of the similarity/dissimilarity matrix

2. Extraction of stimulus coordinates

3. Decision about the number of stimulus coordinates that represent the data

4. Rotation and interpretation

Example: Let us say that you have a matrix of distances between a number of
major cities, such as you might find on the back of a road map. These distances can
be used as the input data to derive an MDS solution. When the results are mapped
in two dimensions, the solution will reproduce a conventional map, except that the
MDS plot might need to be rotated so that the north-south and east-west
dimensions conform to expectations. However, the once the rotation is completed,
the configuration of the cities will be spatially correct.

12.7 Conjoint Analysis


Conjoint analysis is concerned with the measurement of the joint effect of two or more attributes
that are important from the customers’ point of view. In a situation where the company would like
to know the most desirable attributes or their combination for a new product or service, the use of
conjoint analysis is most appropriate.

Example: An airline would like to know, which is the most desirable combination
of attributes to a frequent traveler: (a) Punctuality (b) Air fare (c) Quality of food
served on the flight and (d) Hospitality and empathy shown.

Lovely Professional University 197


Notes

Unit 12: Multivariate Analysis

Conjoint Analysis is a multivariate technique that captures the exact levels of utility that an
individual customer places on various attributes of the product offering. Conjoint Analysis enables
a direct comparison,

Example: A comparison between the utility of a price level of 400 versus 500, a
delivery period of 1 week versus 2 weeks, or an after-sales response of 24 hours versus 48
hours.

Once we know the utility levels for each attribute (and at individual levels as well), we can combine
these to find the best combination of attributes that gives the customer the highest utility, the
second best combination that gives the second highest utility, and so on. This information is then
used to design a product or service offering.

Application
Conjoint Analysis is extremely versatile, and the range of applications includes virtually in any
industry. New product or service design, including the concepts in the pre-prototyping stage can
specifically benefit from the conjoint applications.
Some examples of other areas where this technique can be used are:
1. Designing an automobile loan or insurance plan in the insurance industry,
2. Designing a complex machine for business customers.

Process
Design attributes for a product are first identified. For a shirt manufacturer, these could be design
such as designer shirts vs plain shirts, this price of 400 versus 800. The outlets can have
exclusive distribution or mass distribution. All possible combinations of these attribute levels are
then listed out. Each design combination will be ranked by customers and used as input data for
Conjoint Analysis. Then the utility of the products relative to price can be measured.
The output is a part-worth or utility for each level of each attribute. For example, the design may
get a utility level of 5 and plain, 7.5. Similarly, the exclusive distribution may have a part utility of
2, and mass distribution, 5.8. We then put together the part utilities and come up with a total utility
for any product combination we want to offer and compare that with the maximum utility
combination for this customer segment.
This process clarifies to the marketer about the product or service regarding the attributes that they
should focus on in the design.
If a retail store finds that the height of a shelf is an important attribute for selling at a particular
level, a well-designed shelf may result from this knowledge. Similarly, a designer of clocks will
benefit from knowing the utility attached by customers to the dial size, background colours, and
price range of the clocks.

Approach
From a discussion with the client, identify the design attributes to be studied and the levels at
which they can be offered. Then build a list of product concepts on offer. These product concepts
are then ranked by customers. Once this data is available, use Conjoint Analysis to derive the part
utilities of each attribute level. This is then used to predict the best product design for the given
customer segment. Use the SPSS Conjoint procedure to analyse the data.
There are three steps in conjoint analysis:
1. Identification of relevant products or service attributes.
2. Collection of data.
3. Estimation of worth for the attribute chosen.

For attributes selection, the market researcher can conduct interview with the customers directly.

198 Lovely Professional University


Notes

Research Methodology

Example: Example of conjoint analysis for a Laptop:


For a laptop, consider 3 attributes:
1. Weight (3 Kg or 5 Kg)
2. Battery life (2 hours or 4 hours)
3. Brand name (Lenovo or Dell)

SPSS Command for Conjoint Analysis


SPSS commands for conjoint Analysis. A data file is to be created containing all possible attribute
combination.
1. Ask each of the respondent to rank all the combination of attributes contained in the file. This
is nomenclated at DATA FILE 1. All the rankings should be entered in another file called
DATA FILE 2.

2. Now 2 files namely DATA FILE 1 and DATA FILE 2 are created.

3. A third file called SYNTAX file is to be opened. By using the FILE, OPEN command followed
by syntax.

4. Type the following - conjoint plan = DATA FILE 1 SAV/DATA' DATA FILE 2 SAV/

SCORES=SCORE 1 to Score number of ranking/FACTOR VARI (DISCRETE)/PLOT ALL


(Here 25 is the possible combination of attributes). Score is the term used for rankings. The no
of scores will be equal to number of rankings. We should use the word RANK in the syntax
instead of scores if Rankings are contained in the data file.
5. Click RUN from the menu of the syntax file that was created click all in the menu which
appears on the screen. If the syntax is correct, the output for conjoint will appear.
Task: Rank order the following combination of these characteristics:
1= Most preferred, 2 = Least preferred

One combination 3 kg, 4 hours, Dell clearly dominates and 5 kg, 2 hours, Lenovo
is leastpreferred.

Let us now take the average rank for 3 kg option = 4 + 3 + 2 + 1/4 = 2.5
For 5 kg option average rank is 5 + 8 + 7 + 6/4 = 6.5
For 4 hour option 5 + 3 + 7 + 1/4 = 4
For 2 hour option 4 + 8 + 2 + 6/4 = 5
For Dell 5 + 6 + 1 + 2/4 = 3.5
For Lenovo 5.5
Looking at the difference in average ranks, the most important characteristic to
this respondent is weight = 4, followed by brand name = 2 and battery life = 1.

Lovely Professional University 199


Notes

Unit 12: Multivariate Analysis

Summary
 Multivariate analysis is used if there are more than 2 variables.

 Some of the multi variate analysis are discriminant analysis, Factor analysis, Cluster analysis,
conjoint analysis, and multi-dimensional scaling.

 In discriminant analysis, it is verified whether the 2 groups differ from one another.

 Factor analysis is used to reduce large no of various factors into fewer variables cluster
analysis is used to segmenting the market or to identify the target group.

 Regression is a term used for predicting the value of one variable from the other.

 Least square method is used to fit the line.

 MDS as a set of multivariate statistical methods for estimating the parameters in and
assessing the fit of various spatial distance models for proximity data.
 The output of MDS looks very similar to that of factor analysis and the determination of the
optimal number of dimensions is handled in much the same way.

Keywords
Cluster Analysis: Cluster Analysis is a technique used for classifying objects into groups.

Conjoint Analysis: Conjoint analysis is concerned with the measurement of the joint effect of two
or more attributes that are important from the customers’ point of view.

Discriminant Analysis: In this analysis, two or more groups are compared. In the final analysis, we
need to find out whether the groups differ one from another.
Factor Analysis: Factor Analysis is the analysis whose main purpose is to group large set of
variable factors into fewer factors.

Multivariate Analysis: In multi variate analysis, the number of variables to be tackled are many.

Self Assessment
1-In discriminant analysis the averages for the independent variables for a group define the

A. centroid
B. median
C. mode
D. central tendency

2____ is a method for deriving the utility values that consumers attach to varying levels of a
product's attributes

A. Regression
B. Conjoint analysis
C. Correlation
D. T test

3-The conjoint analysis procedure is based on trade-offs respondents make when evaluating
alternatives.

A. True
B. False

4-What is the idea behind conjoint analysis?

200 Lovely Professional University


Notes

Research Methodology

A. Understanding Consumer Preferences


B. Manufacturing process determination
C. Building Social Media
D. Ignoring Price Increase

5-What is the first step in setting up a conjoint analysis?

A. Using data to improve your products


B. Choosing features, functions or attributes
C. Asking consumers to choose their top features
D. Collecting responses from consumers

6-In Conjoint Analysis, responses are collected from………

A. Researchers
B. Industries
C. Marketers
D. Consumers

7-Conjoint Analysis is a ……………… technique which is used to determine customers’


preferences:

A. Descriptive
B. Predictive
C. Inferential
D. None of These

8-Factor analysis requires that variables:

A. Are measured at nominal level


B. Are abstract concepts
C. Are not related to each other
D. Are related to each other

9-Factor analysis is a(n) _____ in that the entire set of interdependent relationships is examined.

A. KMO measure of sampling adequacy


B. orthogonal procedure
C. interdependence technique
D. varimax procedure

10-_____ are simple correlations between the variables and the factors.

A. Factor scores
B. Factor loadings
C. Correlation loadings
D. Both a and b are correct

11-A factor can be considered to be an underlying latent variable:

Lovely Professional University 201


Notes

Unit 12: Multivariate Analysis

A. on which people differ


B. that is explained by unknown variables
C. that cannot be defined
D. that is influenced by observed variables

12-The decision about how many factors to retain is based on:

A. personal choice
B. Kaiser’s rule
C. Scree test
D. Both Kaiser’s rule and Scree test

13-The goal of clustering is to-

A. Divide the data points into groups


B. Classify the data point into different classes
C. Predict the output values of input data points
D. All of the above

14-Which of the following is a bad characteristic of a dataset for clustering analysis-

A. Data points with outliers


B. Data points with different densities
C. Data points with non-convex shapes
D. All of the above

15-On which data type, we cannot perform cluster analysis

A. Time series data


B. Text data
C. Multimedia data
D. None

Review Questions
1. Which technique would you use to measure the joint effect of various attributes while
designing an automobile loan and why?
2. Do you think that the conjoint analysis will be useful in any manner for an airline? If yes how,
if no, give an example where you think the technique is of immense help.
3. In your opinion, what are the main advantages of cluster analysis?
4. Which analysis would you use in a situation when the objective is to summarize information
from a large set of variables into fewer factors? What will be the steps you would follow?
5. Which analysis would answer if it is possible to estimate the size of different groups?
6. Which analysis would you use to compare a good, bad and a mediocre doctor and why?
7. Analyse the weakness of principle component factor analysis.
8. Which multivariate analysis would you apply to identify specific customer segment for a
company’s brand and why?
9. Critically evaluate multidimensional scaling.

202 Lovely Professional University


Notes

Research Methodology

10. In your opinion what will be the disadvantages of having too many independent variables in
an MR study?
[Link] have been rated on their suitability for an advanced training course in computer
programming on the basis of six ratings given by their manager (rated 1=low to 20=high):

(a) Intellect

(b) Interest in doing the course

(c) Experience of computer programming

(d) Likelihood of them staying with the company

(e) Commitment to the company

(f) Loyalty to their team and two other ratings:

(g) Number of GCSEs

(h) Score on a computer programming aptitude test

The training department believe that these are really measuring only three things; intellect,
computer programming experience and loyalty, and want you to carry out a factor analysis
to explore that hypothesis. Describe the decisions you would have to make in carrying out a
factor analysis and what the results would be likely to tell you.
12. Six observations on two variables are available, as shown in the following table:

Obs. X1 X2

a 3 2

b 4 1

c 2 5

d 5 2

e 1 6

f 4 2
(a) Plot the observations in a scatter diagram. How many groups would you say there are,
and what are their members?

(b) Apply the nearest neighbor method and the squared Euclidean distance as a measure
of dissimilarity. Use a dendrogram to arrive at the number of groups and their
membership.

13. Six observations on two variables are available, as shown in the following table:

Obs. X1 X2

a -1 -2

b 0 0

c 2 2

d -2 -2

e 1 -1

f 1 2
(a) Plot the observations in a scatter diagram. How many groups would you say there are,
and what are their members?

(b) Apply the nearest neighbor method and the Euclidean distance as a measure of
dissimilarity.

Lovely Professional University 203

Common questions

Powered by AI

Cluster analysis aids in market segmentation by grouping consumers into clusters based on shared characteristics, thus identifying consumer segments with similar behavior or preferences. Using methods such as hierarchical or K-means clustering allows businesses to analyze and target these segments with tailored marketing strategies. This approach uncovers patterns in consumer behavior, such as purchasing habits, that might not be apparent otherwise, making it invaluable for strategic marketing decisions .

Discriminant analysis is preferred when distinguishing between two or more groups based on multiple correlated independent variables, as it considers the total effect and interrelationships of these variables simultaneously. This method is particularly useful in scenarios where the goal is to determine which factors contribute most significantly to group differentiation, such as segmenting markets based on income and education, or comparing customer preferences across different socio-economic groups .

Convergent validity ensures that a new measurement scale correlates positively with other established measures of the same construct, thereby confirming that the scale accurately assesses the intended concept. Discriminant validity, on the other hand, checks that the scale does not strongly correlate with measures of different, unrelated constructs, thus validating the scale's specificity. Together, these types of validity help to confirm that a new measurement tool is both comprehensive and specific, accurately reflecting the unique aspects of the construct it intends to measure .

The regression coefficient in a bivariate regression model indicates the average rate of change of the dependent variable with respect to changes in the independent variable. Accurately determining this coefficient helps in understanding the strength and direction of the relationship, allowing for precise predictions and strategic decision-making. For businesses, this insight can guide resource allocation, pricing strategies, and policy changes by portraying how future outputs may respond to changes in key input variables .

Bivariate regression can be used in two ways: predicting average values of a dependent variable (Y) given an independent variable (X) or vice versa. The regression of Y on X provides estimates of Y for different values of X using the equation YCi = a + bXi, where a is the intercept and b is the slope indicating Y's change rate per unit change in X. Conversely, when predicting X given Y, the roles are reversed, highlighting the different assumptions and derivations involved. These two regression lines will differ as they are based on different sets of assumptions related to the variables .

Criterion validity is crucial as it assesses whether a measurement scale predicts or reflects the expected behavior or trait, ensuring the scale's relevance and accuracy. It is evaluated by comparing the scale's outcomes with external criteria, such as predicting future performance based on current scores or evaluating actual behavior against predicted outcomes, thus confirming the measurement's practical applicability .

Challenges in implementing discriminant analysis for market segmentation include multicollinearity between predictor variables, non-normality of data, and unequal covariance among groups. These issues can distort results and affect the analysis’s reliability. To address them, one might use variable selection methods to manage multicollinearity, apply transformations or robust statistical techniques to handle non-normality, and ensure that sample sizes are large enough to accurately estimate covariances, thus enhancing the overall validity of the analysis .

Conjoint analysis aids in optimizing new product development by evaluating the joint effect of multiple product attributes as perceived by consumers. By determining the utility values each consumer places on various attributes, businesses can identify the most desirable product features and their ideal combinations. This analysis allows companies to design products or services that align closely with customer preferences, thereby maximizing perceived value and potential market success .

Ensuring the face validity of a new product involves clearly defining the problem and determining what specifically is being measured. For instance, when introducing a new packaged food with a significant flavor change, it is crucial to create a scale that assesses customer reactions accurately. A common pitfall is relying solely on initial consumer enjoyment without considering other critical factors such as long-term satisfaction or health implications, which might lead to commercial failure despite positive initial feedback .

Multidimensional scaling (MDS) plays a crucial role in visualizing complex data by transforming high-dimensional data into a lower-dimensional space for easier interpretation. It uses similarity or dissimilarity data to derive spatial representations that preserve the distances or differences between data points, such as cities or customer preferences. Practically, MDS can be applied to various fields, including marketing for visualizing brand positions or geography for mapping spatial data, making complex relationships visually comprehensible .

You might also like