0% found this document useful (0 votes)
3 views50 pages

UnitV

The document outlines a statistical analysis involving discriminant analysis, principal component analysis, and factor analysis on various datasets. It provides detailed procedures for calculating discriminant functions, extracting principal components, and evaluating sales staff performance through factor analysis. Results include linear discriminant functions, classification tables, and principal components explaining variance in the data.

Uploaded by

naga1962malli
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views50 pages

UnitV

The document outlines a statistical analysis involving discriminant analysis, principal component analysis, and factor analysis on various datasets. It provides detailed procedures for calculating discriminant functions, extracting principal components, and evaluating sales staff performance through factor analysis. Results include linear discriminant functions, classification tables, and principal components explaining variance in the data.

Uploaded by

naga1962malli
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DISCRIMINANT ANALYSIS

Problem:
Three psychological tests were given to 10 men and 9 women. The values are
X1 -Pictorial inconsistencies, X2 – Tool recognition and X3 – Vocabulary.
Men Women
X1 X2 X3 X1 X2 X3
15 24 14 13 12 21
17 32 26 14 14 26
15 29 23 12 21 21
13 10 16 12 10 16
20 26 28 11 16 16
15 26 21 12 14 18
15 26 22 10 18 24
13 26 22 10 13 23
14 30 17 12 19 23
17 30 27

i. Calculate discriminant function coefficient and write down the linear discriminant
function.
ii. Find the classification table using LDF.

Aim:
To find the Linear Discriminant Function and to find the classification table using it.

Procedure:
➢ From the menu, choose, Analyze → Classify → Discriminant…
➢ Select the grouping variable and move it to ‘Grouping Variable’ and also define
the range of the group by clicking ‘Define Range’ option and then click
‘Continue’.
➢ Select the all observation variables and move them to ‘Independents’ list.
➢ Click ‘Statistics’ and choose; Mean, Box’s M, Fisher’s, Within-Group covariance,
etc. like whatever options are necessary and click ‘Continue’.
➢ Click ‘Classify’ and choose the options like All groups equal, Within groups,
Casewise result, Summary table, Combined groups, etc. and click ‘Continue’.
➢ If needed, click ‘Save’ and choose ‘Predicted group membership’ and click
’Continue’.
➢ Now click ‘Ok’.

Step 1
Step 2

Step 3
Step 4

Step 5
Output:

Group Statistics
Grouping Valid N (listwise)
Mean Std. Deviation Unweighted Weighted
1 X1 15.40 2.119 10 10.000
X2 25.90 6.118 10 10.000
X3 21.60 4.742 10 10.000
2 X1 11.78 1.302 9 9.000
X2 15.22 3.563 9 9.000
X3 20.89 3.551 9 9.000
Total X1 13.68 2.540 19 19.000
X2 20.84 7.373 19 19.000
X3 21.26 4.121 19 19.000

Pooled Within-Groups Matricesa


X1 X2 X3
Covariance X1 3.174 2.403 4.199
X2 2.403 25.792 9.990
X3 4.199 9.990 17.841
a. The covariance matrix has 17 degrees of freedom.

Classification Function Coefficients


Grouping
1 2
X1 4.704 3.135
X2 .671 .167
X3 -.272 .340
(Constant) -42.669 -23.973
Fisher's linear discriminant functions

Classification Resultsa
Grouping Predicted Group Membership
1 2 Total
Original Count 1 9 1 10
dimension2

2 0 9 9
% 1 90.0 10.0 100.0
dimension2

2 .0 100.0 100.0
a. 94.7% of original grouped cases correctly classified.
Casewise Statistics

Case Number
Actual Group Predicted Group
Original 1 1 1
2 1 1
3 1 1
4 1 2**
5 1 1
6 1 1
7 1 1
8 1 1
9 1 1

dimension1
10 1 1
11 2 2
12 2 2
13 2 2
14 2 2
15 2 2
16 2 2
17 2 2
18 2 2
19 2 2
**. Misclassified case

Result:
✓ The Linear Discriminant Function is given by
1.569X1 + 0.504X2 – 0.612X3 – 18.696
✓ The Estimated Classification table is as follows:
Predicted group
Actual group No. of Samples
Men Women
Men 10 9 1

Women 9 0 9
PRINCIPAL COMPONENT ANALYSIS

Problem:
Consider the census dract data listed in the following table.

X1: Total X2: Total X3: Health Service


Dract Population Employment Employment
(in ‘000) (in ‘000) (in ’00)
1 5.935 2.265 2.27
2 1.523 0.597 0.75
3 2.599 1.237 1.11
4 4.009 1.649 0.81
5 4.687 2.312 2.50
6 8.044 3.641 4.51
7 2.766 1.244 1.03
8 6.538 2.618 2.39
9 6.451 3.147 5.52
10 3.314 1.606 2.18
11 3.777 2.119 2.83
12 1.530 0.798 0.84
13 2.768 1.336 1.75
14 6.585 2.763 1.91

Extract the Principal Component Analysis.

Aim:
To extract the Principal Component Analysis for the given data.

Procedure:
➢ From the menu, choose, Analyze → Dimension Reduction → Factor…
➢ Select the observations and move them to ‘Variables’.
➢ Click ‘Descriptives’, and choose the options, ‘Initial Solutions’ and ‘Coefficients’ and
whatever necessary and click ‘Continue’.
➢ Click ‘Extraction’, and choose the options like Correlation matrix, Unrotated factor
solution, etc.(whatever necessary) and change the eigen value as 0.001 and click
‘Continue’.
➢ If needed click ‘Scores’, and select the option ‘Display factor score coefficient
matrix’ and click ‘Continue’.
➢ Now click ‘Ok’.

Step 1
Step 2

Step 3
Output:

Correlation Matrix
X1 X2 X3
Correlation X1 1.000 .971 .740
X2 .971 1.000 .848
X3 .740 .848 1.000

Total Variance Explained


Component Initial Eigenvalues Extraction Sums of Squared Loadings
Total % of Variance Cumulative % Total % of Variance Cumulative %
1 2.709 90.310 90.310 2.709 90.310 90.310

dimension0
2 .278 9.280 99.590 .278 9.280 99.590
3 .012 .410 100.000 .012 .410 100.000
Extraction Method: Principal Component Analysis.

Component Matrixa
Component
1 2 3
X1 .954 -.292 .066
X2 .991 -.108 -.086
X3 .904 .426 .024
Extraction Method: Principal Component
Analysis.
a. 3 components extracted.

Result:
✓ The Principal Components are
Y1 = 0.954X1 + 0.991X2 + 0.904X3
Y2 = -0.292X1 - 0.108X2 + 0.426X3
Y3 = 0.066X1 - 0.086X2 + 0.024X3
✓ The percentage of variation by first component is 90.31%,
The percentage of variation by second component is 9.28% and
The percentage of variation by third component is 0.41%.
FACTOR ANALYSIS

A firm is attempting to evaluate the quality of its sales staff and is trying to find an
examination or series of tests that may reveal the potential for good performance in sales.
The firm has selected a random sample of 50 sales people and has evaluated each on three
measures of performance: growth of sales (X1), profitability of sales (X2) and new account
sales (X3). These measures have been converted to a scale, on which 100 indicates
“average” performance. Each of 50 individuals took each of 4 tests, which purported to
measure creativity (X4), mechanical reasoning (X5), abstract reasoning (X6) and
mathematical ability (X7) respectively. The data is given in the following table.
(a) Assume an orthogonal factor model for the standardized variables Zi = (Xi – μi) /
i = 1, 2, ... , 7. Obtain the principal component solution and the maximum likelihood
solution for m = 3 common factors.
(b) Obtain the rotated loadings for m = 3. Compare the two sets of rotated loadings.
Interpret the m = 3 factor solutions.
Index of: Score on:
X1 X2 X3 X4 X5 X6 X7
93.0 96.0 97.8 09 12 09 20
88.8 91.8 96.8 07 10 10 15
95.0 100.3 99.0 08 12 09 26
101.3 103.8 106.8 13 14 12 29
102.0 107.8 103.0 10 15 12 32
95.8 97.5 99.3 10 14 11 21
95.5 99.5 99.0 09 12 09 25
110.8 122.0 115.3 18 20 15 51
102.8 108.3 103.8 10 17 13 31
106.8 120.5 102.0 14 18 11 39
103.3 109.8 104.0 12 17 12 32
99.5 111.8 100.3 10 18 08 31
103.5 112.5 107.0 16 17 11 34
99.5 105.5 102.3 08 10 11 34
100.0 107.0 102.8 13 10 08 34
81.5 93.5 95.0 07 09 05 16
101.3 105.3 102.8 11 12 11 32
103.3 110.8 103.5 11 14 11 35
95.3 104.3 103.0 05 14 13 30
99.5 105.3 106.3 17 17 11 27
88.5 95.3 95.8 10 12 07 15
99.3 115.0 104.3 05 11 11 42
87.5 92.5 95.8 09 09 07 16
105.3 114.0 105.3 12 15 12 37
107.0 121.0 109.0 16 19 12 39
93.3 102.0 97.8 10 15 07 23
106.8 118.0 107.3 14 16 12 39
106.8 120.0 104.8 10 16 11 49
92.3 90.8 99.8 08 10 13 17
106.3 121.0 104.5 09 17 11 44
106.0 119.5 110.5 18 15 10 43
88.3 92.8 96.8 13 11 08 10
96.0 103.3 100.5 07 15 11 27
94.3 94.5 99.0 10 12 11 19
106.5 121.5 110.5 18 17 10 42
106.5 115.5 107.0 08 13 14 47
92.0 99.5 103.5 18 16 08 18
102.0 99.8 103.3 13 12 14 28
108.3 122.3 108.5 15 19 12 41
106.8 119.0 106.8 14 20 12 37
102.5 109.3 103.8 09 17 13 32
92.5 102.5 99.3 13 15 06 23
102.8 113.8 106.8 17 20 10 32
83.3 87.3 96.3 01 05 09 15
94.8 101.8 99.8 07 16 11 24
103.5 112.0 110.8 18 13 12 37
89.5 96.0 97.3 07 15 11 14
84.3 89.8 94.3 08 08 08 09
104.3 109.5 106.5 14 12 12 36
106.0 118.5 105.0 12 16 11 39
➢ From the menu, choose, Analyze → Dimension Reduction → Factor…
➢ Select the observations and move them to ‘Variables’.
➢ Click ‘Descriptives’, and choose the options, ‘Initial Solutions’ and ‘Coefficients’ and
whatever necessary and click ‘Continue’.
➢ Click ‘Extraction’, and choose the options like Correlation matrix, Unrotated factor
solution, etc.(whatever necessary) and change the eigen value as 0.001 and click
‘Continue’.
➢ If needed click ‘Scores’, and select the option ‘Display factor score coefficient
matrix’ and click ‘Continue’.
For rotated factor:
Click ‘Extraction’, and choose the options like Correlation matrix, VARIMAX factor
solution, etc.(whatever necessary)
Analyze → Dimension Reduction → Factor
Correlation Matrix

Y1 Y2 Y3 Y4 Y5 Y6 Y7

Correlation Y1 1.000 .926 .884 .572 .708 .674 .927

Y2 .926 1.000 .843 .542 .746 .465 .944

Y3 .884 .843 1.000 .700 .637 .641 .853

Y4 .572 .542 .700 1.000 .591 .147 .413

Y5 .708 .746 .637 .591 1.000 .386 .575

Y6 .674 .465 .641 .147 .386 1.000 .566

Y7 .927 .944 .853 .413 .575 .566 1.000

KMO and Bartlett's Test

Kaiser-Meyer-Olkin Measure of Sampling Adequacy. .616


Bartlett's Test of Sphericity Approx. Chi-Square 499.661

df 21

Sig. .000

Communalities

Initial Extraction

Y1 1.000 1.000
Y2 1.000 1.000
Y3 1.000 1.000
Y4 1.000 1.000
Y5 1.000 1.000
Y6 1.000 1.000
Y7 1.000 1.000

Extraction Method: Principal


Component Analysis.
Total Variance Explained

Initial Eigenvalues Extraction Sums of Squared Loadings

Component Total % of Variance Cumulative % Total % of Variance Cumulative %

1 5.035 71.923 71.923 5.035 71.923 71.923


2 .934 13.336 85.259 .934 13.336 85.259
3 .498 7.113 92.372 .498 7.113 92.372
4 .421 6.018 98.390 .421 6.018 98.390
5 .081 1.158 99.547 .081 1.158 99.547
6 .020 .291 99.838 .020 .291 99.838
7 .011 .162 100.000 .011 .162 100.000
Communalities

Initial Extraction

Y1 1.000 1.000
Y2 1.000 1.000
Y3 1.000 1.000
Y4 1.000 1.000
Y5 1.000 1.000
Y6 1.000 1.000
Y7 1.000 1.000

Extraction Method: Principal Component Analysis.

Component Matrixa

Component

1 2 3 4 5 6 7

Y1 .973 -.108 -.053 -.028 .180 -.048 -.056


Y2 .943 .028 -.312 .007 .000 .112 -.011
Y3 .945 .009 .144 -.211 -.200 -.022 -.043
Y4 .660 .646 .319 -.196 .074 .016 .032
Y5 .783 .285 .004 .549 -.050 -.028 .008
Y6 .649 -.621 .426 .100 .025 .034 .024
Y7 .914 -.194 -.306 -.160 -.014 -.053 .068
Extraction Method: Principal Component Analysis.

VARIMAX Rotation:

Component Matrixa

Component
1 2 3

Y1 .973 -.108 -.053


Y2 .943 .028 -.312
Y3 .945 .009 .144
Y4 .660 .646 .319
Y5 .783 .285 .004
Y6 .649 -.621 .426
Y7 .914 -.194 -.306

Extraction Method: Principal Component


Analysis.
a. 3 components extracted.
Communalities

Extraction

Y1 .961
Y2 .987
Y3 .913
Y4 .955
Y5 .695
Y6 .988
Y7 .967

Extraction Method:
Principal Component
Analysis.
Rotated Component Matrixa

Component

1 2 3

Y1 .779 .387 .452


Y2 .908 .356 .189
Y3 .616 .548 .484
Y4 .213 .952 .047
Y5 .552 .607 .146
Y6 .286 .060 .950
Y7 .909 .181 .328

Extraction Method: Principal Component


Analysis. Rotation Method: Varimax with
Kaiser Normalization.
a. Rotation converged in 4 iterations.

Maximum Likelihood method


Communalitiesa

Extraction

Y1 .961
Y2 .965
Y3 .912
Y4 .999
Y5 .552
Y6 .999
Y7 .963
Extraction Method: Maximum Likelihood.

Factor Matrixa
Factor
1 2 3
.842 -.075 .497
.690 .061 .696
.898 .050 .322
.753 .657 -.026
.656 .160 .311
.760 -.648 -.029
.673 -.115 .705
Extraction Method: Maximum Likelihood.
a. 3 factors extracted. 8 iterations required.

Total Variance Explained

Extraction Sums of Squared Loadings Rotation Sums of Squared Loadings

Factor Total % of Variance Cumulative % Total % of Variance Cumulative %

1 4.019 57.413 57.413 3.179 45.417 45.417


2 .903 12.898 70.312 1.719 24.563 69.980
3 1.430 20.430 90.742 1.453 20.762 90.742

Extraction Method: Maximum Likelihood.

Rotated Factor Matrixa


Factor
1 2 3
Y1 .794 .374 .437
Y2 .912 .316 .184
Y3 .652 .544 .437
Y4 .255 .966 .019
Y5 .541 .464 .208
Y6 .300 .054 .952
Y7 .918 .179 .296
Extraction Method: Maximum Likelihood.
Rotation Method: Varimax with Kaiser
Normalization.
a. Rotation converged in 4 iterations.
MULTIPLE COMPARISON TESTS

Problem:
Six different design for a digital computer circuit are being studied to compare the
amount of noise present. The following data have been obtained.

Circuit Design Noise Observed


1 19 20 19 30 8
2 80 61 73 56 80
3 47 26 25 35 50
4 95 46 83 78 97
5 110 100 95 112 108
6 25 30 24 26 20

i) Is the amount of noise present the same for all six designs? Use α = 0.05
ii) Also apply multiple comparison tests LSD, SNK, DMR and Tukey test to analyze
the above data.

Aim:
To analyze the given completely randomized design and to interpret the results.

Hypothesis:
H0: The amount of noise present is same in all six circuit designs.

Procedure:
➢ From the menu choose, Analyze → Compare Means → One-Way ANOVA…
➢ From the dialogue box appears, select the observation variable and move it to
‘Dependent List’ and select the grouping variable and move it to ‘Factor’ and
click ‘Post Hoc…’
➢ From the dialogue box appears now, select the options LSD, S-N-K, Tukey and
Duncan and then click ‘Continue’.
➢ Click ‘Ok’.
Step1

Step2
Output:

ANOVA
noise observed

Sum of Squares df Mean Square F Sig.


Between Groups 29275.067 5 5855.013 43.792 .000
Within Groups 3208.800 24 133.700

Total 32483.867 29

Multiple Comparisons
Dependent Variable:noise observed
(I) circuit (J) circuit Mean 95% Confidence Interval
design design Difference (I- Std. Lower Upper
J) Error Sig. Bound Bound
Tukey 1 2 -50.800* 7.313 .000 -73.41 -28.19
HSD 3 -17.400 7.313 .203 -40.01 5.21

dimension3
4 -60.600* 7.313 .000 -83.21 -37.99
5 -85.800* 7.313 .000 -108.41 -63.19
6 -5.800 7.313 .966 -28.41 16.81
2 1 50.800* 7.313 .000 28.19 73.41
3 33.400* 7.313 .002 10.79 56.01

dimension3
4 -9.800 7.313 .760 -32.41 12.81
5 -35.000* 7.313 .001 -57.61 -12.39
6 45.000* 7.313 .000 22.39 67.61
3 1 17.400 7.313 .203 -5.21 40.01

dimension2
2 -33.400* 7.313 .002 -56.01 -10.79

dimension3
4 -43.200* 7.313 .000 -65.81 -20.59
5 -68.400* 7.313 .000 -91.01 -45.79
6 11.600 7.313 .615 -11.01 34.21
4 1 60.600* 7.313 .000 37.99 83.21
2 9.800 7.313 .760 -12.81 32.41

dimension3
3 43.200* 7.313 .000 20.59 65.81
5 -25.200* 7.313 .023 -47.81 -2.59
6 54.800* 7.313 .000 32.19 77.41
5 1 85.800* 7.313 .000 63.19 108.41

dimension3
2 35.000* 7.313 .001 12.39 57.61
3 68.400* 7.313 .000 45.79 91.01
4 25.200* 7.313 .023 2.59 47.81
6 80.000* 7.313 .000 57.39 102.61
6 1 5.800 7.313 .966 -16.81 28.41
2 -45.000* 7.313 .000 -67.61 -22.39

dimension3
3 -11.600 7.313 .615 -34.21 11.01
4 -54.800* 7.313 .000 -77.41 -32.19
5 -80.000* 7.313 .000 -102.61 -57.39
LSD 1 2 -50.800* 7.313 .000 -65.89 -35.71
3 -17.400* 7.313 .026 -32.49 -2.31

dimension3
4 -60.600* 7.313 .000 -75.69 -45.51
5 -85.800* 7.313 .000 -100.89 -70.71
6 -5.800 7.313 .435 -20.89 9.29
2 1 50.800* 7.313 .000 35.71 65.89
3 33.400* 7.313 .000 18.31 48.49

dimension3
4 -9.800 7.313 .193 -24.89 5.29
5 -35.000* 7.313 .000 -50.09 -19.91
6 45.000* 7.313 .000 29.91 60.09
3 1 17.400* 7.313 .026 2.31 32.49
2 -33.400* 7.313 .000 -48.49 -18.31

dimension3
4 -43.200* 7.313 .000 -58.29 -28.11
5 -68.400* 7.313 .000 -83.49 -53.31
6 11.600 7.313 .126 -3.49 26.69
dimension2

4 1 60.600* 7.313 .000 45.51 75.69


2 9.800 7.313 .193 -5.29 24.89

dimension3
3 43.200* 7.313 .000 28.11 58.29
5 -25.200* 7.313 .002 -40.29 -10.11
6 54.800* 7.313 .000 39.71 69.89
5 1 85.800* 7.313 .000 70.71 100.89
2 35.000* 7.313 .000 19.91 50.09

dimension3
3 68.400* 7.313 .000 53.31 83.49
4 25.200* 7.313 .002 10.11 40.29
6 80.000* 7.313 .000 64.91 95.09
6 1 5.800 7.313 .435 -9.29 20.89
2 -45.000* 7.313 .000 -60.09 -29.91

dimension3
3 -11.600 7.313 .126 -26.69 3.49
4 -54.800* 7.313 .000 -69.89 -39.71
5 -80.000* 7.313 .000 -95.09 -64.91
*. The mean difference is significant at the 0.05 level.
noise observed
Subset for alpha = 0.05
circuit design
N 1 2 3 4

1 5 19.20

6 5 25.00

3 5 36.60

Student-Newman-Keulsa dimension1
2 5 70.00

4 5 79.80

5 5 105.00

Sig. .064 .193 1.000

1 5 19.20

6 5 25.00

3 5 36.60

Tukey HSDa dimension1 2 5 70.00

4 5 79.80

5 5 105.00

Sig. .203 .760 1.000

1 5 19.20

6 5 25.00 25.00

3 5 36.60

Duncana dimension1
2 5 70.00

4 5 79.80

5 5 105.00

Sig. .435 .126 .193 1.000

Means for groups in homogeneous subsets are displayed.


a. Uses Harmonic Mean Sample Size = 5.000.

Result:
✓ F-calculated value = 43.792 and p-value = 0.000
✓ Since p-value is less than 0.05, we reject the hypothesis.
✓ Hence, the amount of noise present is not same in all six circuit designs.
✓ Therefore, we performed Multiple Comparison Tests to test the hypothesis, H0: i =
j to identify the significant pairs under LSD, Tukey, S-N-K and Duncan and shown
below.
✓ LSD: The pairs of circuit designs (1,2), (1,3), (1,4), (1,5), (2,3), (2,5), (2,6), (3,4),
(3,5), (4,5), (4,6) and (5,6) are significant.
✓ DMR: The significant pairs of circuit designs are (1,5), (5,6), (3,5), (2,5), (4,5),
(1,4), (4,6), (3,4), (1,2), (2,6) and (1,3).
✓ Tukey: The pairs of circuit designs (1,2), (1,4), (1,5), (2,3), (2,5), (2,6), (3,4), (3,5),
(4,5), (4,6) and (5,6) are significant.
✓ SNK: The significant pairs of circuit designs are (2,3), (4,5), (2,6), (3,4), (2,5), (1,2),
(4,6), (3,5), (1,4), (5,6) and (1,5).
24 FACTORIAL EXPERIMENT

Problem:
Perform 24 factorial design for the following design.

Replications
Factors
1 2 3 4
1 32 43 27 19
m 47 41 48 45
n 26 36 24 18
mn 61 76 56 64
p 29 39 27 28
mp 51 34 40 48
np 36 31 32 30
mnp 76 65 70 63
k 35 42 56 35
mk 63 41 60 53
nk 80 68 75 67
mnk 100 68 87 66
pk 40 44 53 36
mpk 64 39 75 72
npk 105 99 74 73
mnpk 90 82 89 101

Aim:
To perform the analysis for the given 24 factorial experimental design.

Hypothesis:
H0: The main factor effects are homogeneous.
The first order interaction effects are homogeneous.
The second order interaction effects are homogeneous.
The third order interaction effects are homogeneous.
The replication effects are homogeneous.
Procedure:
➢ From the menu choose, Analyze → General Linear Model → Univariate...
➢ In the dialogue box appears, select the observation variable and move it to ‘Dependent
Variable’ and select replication and other factors and move them to ‘Fixed Factor(s)’.
➢ Click the button ‘Model’ and select the option ‘Custom’ from the dialogue box
appears.
➢ Move the factors available in ‘Factors & Covariates’ to ‘Model’ box as main effects,
interaction effects according to the design and then click ‘Continue’.
➢ Click ‘Ok’.

Step 1
Step 2

Output:

Between-Subjects Factors
N
Replication 1 16
2 16
3 16
4 16
m 0 32

1 32
n 0 32
1 32
p 0 32
1 32
k 0 32
1 32
Tests of Between-Subjects Effects
Dependent Variable:Observation
Source Type III Sum of
Squares df Mean Square F Sig.
Replication 493.312 3 164.437 1.816 .158
m 5184.000 1 5184.000 57.258 .000**
n 7267.562 1 7267.562 80.271 .000**
p 484.000 1 484.000 5.346 .025**
k 9264.062 1 9264.062 102.323 .000**
m*n 169.000 1 169.000 1.867 .179
m*p 1.562 1 1.562 .017 .896
m*k 900.000 1 900.000 9.941 .003**
n*p 196.000 1 196.000 2.165 .148
n*k 1914.062 1 1914.062 21.141 .000**
p*k 169.000 1 169.000 1.867 .179
m*n*p 33.062 1 33.062 .365 .549
m*n*k 1156.000 1 1156.000 12.768 .001**
n*p*k 4.000 1 4.000 .044 .834
m*p*k 10.562 1 10.562 .117 .734
m*n*p*k 39.063 1 39.063 .431 .515
Error 4074.188 45 90.538
Total 31359.438 63
a. R Squared = .870 (Adjusted R Squared = .818)
b. ** = Significant

Result:
✓ F – calculated values and p-values of all main effects, interaction effects and
replications are shown in the above table.
✓ Since p-values of M, N, P, K, MK, NK and MNK are less than 0.05, we reject the
corresponding hypothesizes.
✓ Hence, all main effects and interaction effects MK, NK, MNK are not
homogeneous.
✓ Since p-values of MN, MP, NP, PK, MNP, MKP, NKP and MNPK are greater
than 0.05, there is no evidence to reject the corresponding hypothesizes.
✓ Hence, the interaction effects MN, MP, NP, PK, MNP, MKP, NKP and MNPK
are homogeneous.
32 FACTORIAL EXPERIMENT

Problem:
The effective life of a tool installed in a GPP machine is thought to be effected by the
cutting speed and the tool angle. Three speeds and three angles are selected and a factorial
experiment with two replicates is performed. The resulting data are given below. Analyze
the given data.

Cutting Speed (in / min)


Tool Angle
(Degrees)
125 150 175

8 7 12
15
9 10 13

10 11 14
20
12 13 16

9 15 10
25
10 16 9

Aim:
To analyze and interpret the results for the given 32 factorial experiment.

Hypothesis:
H0: Tool life is equal at different level of tool angle.
Tool life is equal at different level of cutting speed.
There is no influence of tool angle on cutting speed.
There is no significant difference between the replication effects.
Procedure:
➢ From the menu choose, Analyze → General Linear Model → Univariate…

➢ Select observation’s variable and move it to ‘Dependent Variable’ and select the
replication and all factors and move it to ‘Fixed Factor(s)’.

➢ Click ‘Model’ and choose the option ‘Custom’.

➢ Select the variables available in ‘Factors & Covariates’ list and move them to
‘Model’ list as main effects and interaction of the available factors and click
‘Continue’.

➢ Click ‘Ok’.

Step 1
Step 2

Output:

Between-Subjects Factors

N
Replication 1 9

2 9
A 0 6

1 6
2 6
B 0 6

1 6
2 6
Tests of Between-Subjects Effects
Dependent Variable:Observation
Source Type III Sum of
Squares df Mean Square F Sig.
Replication 8.000 1 8.000 12.800 .007**
A 24.333 2 12.167 19.467 .001**
B 25.333 2 12.667 20.267 .001**
A*B 61.333 4 15.333 24.533 .000**
Error 5.000 8 .625
Total 124.000 17
a. R Squared = .960 (Adjusted R Squared = .914)
b. ** = Significant Value

Result:

✓ The F-values and p-values of Replication, main effects A, B and interaction effect
AB are shown in the above table.
✓ Since the p-values of replication, main effects A, B and interaction effect AB are
less than 0.05, we reject the null hypothesizes.
✓ Hence we conclude that,
1. There is significant difference between the replication effects.
2. Tool life is not equal at different levels of tool angle.
3. Tool life is not equal at different levels of cutting speed.
4. There is influence of tool angle on cutting speed.
33 FACTORIAL EXPERIMENT

Problem:
In this problem the measured variable was yield and the factors which might affect this
response were days, operators and concentration of solvent. Three days, three operators and
three concentrations were chosen. Each treatment combination was replicated three times
and the resulting data are given below. Analyze the data.

Days
1 2 3
Concentrations Operators
A B C A B C A B C
1.0 0.2 0.2 1.0 1.0 1.2 1.7 0.2 0.5
0.5 1.2 0.5 0.0 0.0 0.0 0.0 1.2 0.7 1.0
1.7 0.7 0.3 0.5 0.0 0.5 1.2 1.0 1.7
5.0 3.2 3.5 4.0 3.2 3.7 4.5 3.7 3.7
1.0 4.7 3.7 3.5 3.5 3.0 4.0 5.0 4.0 4.5
4.2 3.5 3.5 3.5 4.0 4.2 4.7 4.2 3.7
7.5 6.0 7.2 6.5 5.2 7.2 6.7 7.5 6.2
2.0 6.5 6.2 6.5 6.0 5.7 6.7 7.5 6.0 6.5
7.7 6.2 6.7 6.2 6.5 6.8 7.0 6.0 7.0

Aim:
To analyze and interpret the results for the given 33 factorial experiment.

Hypothesis:
H0: Main effects are homogeneous.
First order interaction effects are homogeneous.
Second order interaction effects are homogeneous.
There is no significant difference between the replication effects.
Procedure:
➢ From the menu choose, Analyze → General Linear Model → Univariate…

➢ Select observation’s variable and move it to ‘Dependent Variable’ and select the
replication and all factors and move it to ‘Fixed Factor(s)’.

➢ Click ‘Model’ and choose the option ‘Custom’.

➢ Select the variables available in ‘Factors & Covariates’ list and move them to
‘Model’ list as main effects and interaction of the available factors and click
‘Continue’.

➢ Click ‘Ok’.

Step 1
Step 2

Output:

Between-Subjects Factors
N
Replication 1 27
2 27
3 27
A 0 27
1 27
2 27
B 0 27
1 27
2 27
C 0 27
1 27
2 27
Tests of Between-Subjects Effects
Dependent Variable:Observation
Source Type III Sum of
Squares df Mean Square F Sig.
Replication .500 2 .250 1.378 .261
A 466.597 2 233.299 1286.870 .000**
B 3.377 2 1.688 9.312 .000**
C 6.077 2 3.039 16.761 .000**
A*B .435 4 .109 .600 .664
A*C .792 4 .198 1.093 .370
B*C 3.744 4 .936 5.163 .001**
A*B*C .922 8 .115 .636 .744
Error 9.427 52 .181
Total 491.871 80
a. R Squared = .981 (Adjusted R Squared = .971)
b. ** = Significant Values

Result:
✓ The F-values and p-values of Replication, main effects A, B and C and
interaction effects AB, AC, BC and ABC are shown in the above table.
✓ Since the p-values of main effects A, B, C and interaction effect BC are less than
0.05, we reject the null hypothesizes.
✓ Hence we conclude that,
1. The effect due to different concentrations is not homogeneous.
2. The effect due to different days is not homogeneous.
3. The effect due to different operators is not homogeneous.
4. The interaction effect due to days and operators is not homogeneous.
✓ Since the p-values of interaction effects AB, AC, ABC and replication effects
are greater than 0.05, there is no evidence to reject the null hypothesizes.
✓ Hence we conclude that,
1. The interaction effect due to concentrations and days is homogeneous.
2. The interaction effect due to concentrations and operators is
homogeneous.
3. The interaction effect due to concentrations, days and operators is
homogeneous.
4. There is no significant difference between the replication effects.
SPLIT PLOT ANALYSIS

Problem:
The data given below were the combined effect of oven temperature (T) and baking
time (B) and life of an electrical component. The data collected with their different levels of
factors are given below. Analyze the data using the split plot design.

Replications

I II III
Temperature
Baking Time

5 10 15 5 10 15 5 10 15

580 517 223 175 188 201 195 162 170 213

600 158 138 152 126 130 147 122 185 180

620 229 186 155 160 170 161 167 181 182

640 223 222 156 201 18 174 182 201 199

Aim:
To analyze and to interpret the given split plot design.

Hypothesis:
H0: Main plot effects are homogeneous.
Subplot effects are homogeneous.
Interaction effects are homogeneous.
Replication effects are homogeneous.
Procedure:
➢ From the menu choose, Analyze → General Linear Model → Univariate…

➢ Select the observation’s variable and move it to ‘Dependent Variable’ and select
other variables (replicate, main plot and sub plot) and move it to ‘Fixed Factor(s)’.

➢ Click ‘Model’ and choose the option ‘Custom’.

➢ Move the factors available in ‘Factors & Covariates’ to ‘Model’ box such that the
moving effects should be as main plots, sub plots, replicates, replicates vs main plot
and main plot vs sub plot and then click ‘Continue’.
➢ Click ‘Ok’.
➢ Now again do the first step and then click ‘Paste’ option.
➢ Type the following syntax in the ‘Syntax’ window which appears and click ‘Run’
button.
➢ /TEST replicate VS replicate*main_plot
/TEST main_plot VS replicate*main_plot.

Step 1
Step 2

Step 3
Output:

Between-Subjects Factors
Value Label N
Replicate 0 I 12
1 II 12
2 III 12
Temperature 0 580 9
1 600 9
2 620 9
3 640 9
Baking Time 0 5 12
1 10 12
2 15 12

Tests of Between-Subjects Effects


Dependent Variable:Observation
Source Type III Sum of
Squares df Mean Square F Sig.
Replicate 11573.556 2 5786.778 1.555 .241
Temperature 28956.111 3 9652.037 2.594 .089
Baking_Time 5342.056 2 2671.028 .718 .503
Temperature * Baking_Time 14762.389 6 2460.398 .661 .682
Replicate * Temperature 18570.889 6 3095.148 .832 .563
Error 59535.556 16 3720.972
Total 138740.556 35
a. R Squared = .571 (Adjusted R Squared = .061)

Custom Hypothesis Tests Index

1 Hypothesis Term Replicate

Error Term Replicate *


Temperature
2 Hypothesis Term Temperature

Error Term Replicate *


Temperature
Test Results
Dependent Variable:Observation
Source Sum of Squares df Mean Square F Sig.
Contrast 11573.556 2 5786.778 1.870 .234
Errora 18570.889 6 3095.148
a. Replicate * Temperature

Test Results
Dependent Variable:Observation
Source Sum of Squares df Mean Square F Sig.
Contrast 28956.111 3 9652.037 3.118 .110
Errora 18570.889 6 3095.148
a. Replicate * Temperature

Result:
✓ The F-values of replicate(1.87), temperature(3.118), baking time(0.718),
interaction effect of temperature vs baking time(0.661) and their
corresponding p-values are shown in their above respective tables.
✓ Since the p-values of replicate, main plot (temperature), sub plot (baking time)
and the interaction effect of main plot and sub plot are greater than 0.05, there
is no evidence to reject the null hypothesizes.
✓ Hence we conclude that,
1. Replication effects are homogeneous.
2. Effects due to different temperatures are homogeneous.
3. Effects due to different baking times are homogeneous.
4. Interaction effect due to temperatures and baking times are homogeneous.
LOGISTIC REGRESSION
PROBLEM
In this example, we are attempting to model the likelihood of early termination
from counseling in a sample of n=45 clients at a community mental health center.
The dependent variable in the model is ‘terminate’ (coded 1=terminated early,
0=did not terminate early), where the “did not terminate” group is the reference
(baseline) category and the “terminated early” group is the target category. Two
predictors in the model are categorical: gender identification (‘genderid’, coded
0=identified as male, 1=identified as female) and ‘income’ (ordinal variable, coded
1=low, 2=medium, 3=high). The reference category for ‘genderid’ is male
identification, whereas the reference category for ‘income’ is the low income
group. Finally two predictors are assumed continuous in the model: avoidance of
disclosure (‘avdisc’) and symptom severity (‘sympsev’).
AIM
To perform logistic regression for the given data
PROCEDURE
 From the menu, choose Analyze →Regression →Binary Logistic…
 Transfer the dependent variable, in dependent box and independent
variables into covariates
 Click the categorical button. You will be presented with the Logistic
Regression: Define Categorical Variables dialog box. Transfer the
independent, categorical variable, gender from covariates to categorical
covariates
 In the change contrast change the reference category to first (Whether you
choose last or first will depend on how you set up your data. In this case,
females are to be compared to males, with males acting as the reference
category (who were coded “0”). Therefore, first is chosen)
 Click the continue button. You will be returned to the Logistic Regression
dialogue box. Click the save button and choose Probabilities from Predicted
Values and click “Continue”.
 Click the options button. You will be presented with the Logistic
Regression: Options dialogue box. Select Classification plots, Hosmer-
Lemeshow goodness-of-fit and CI for exp(B). Then click “Continue”.
.
OUTPUT
INTERPRETATION
The Omnibus Tests for Model Coefficients contains results from the
likelihood ratio chi-square tests. These test whether a model including the
full set of predictors is a significant improvement in fit over the null
(intercept-only) model. In effect, it can be considered an omnibus test of the
null hypothesis that the regression slopes for all predictors in the model are
equal to zero. The results shown here indicate that the model fits the data
significantly better than a null model, χ²(4)=29.28, p<0.001.
The model summary table contains the -2Log likelihood and two “pseudo-R-
square” measures. Here the pseudo R² values are 0.478 and 0.640, which
says that this model is a good fit.
The Hosmer & Lemeshow test is another test that can be used to evaluate
global fit. Here p-value is 0.869, indicates non-significance. So it is a good
model fit. So, we can fit a logistic regression for it.
The classification table provides the frequencies and percentages reflecting
the degree to which the model correctly and incorrectly predicts category
membership on the dependent variable. We can see that 82.2% is correctly
classified.
From the Variables in the Equation table,
 Avoidance of disclosure is a positive and significant (B=0.364, S.E.
=0.125, p=0.004) predictor of the probability of early termination,
with the OR (Odds Ratio) indicating that for every one unit increase
on this predictor the odds of early termination change by a factor of
1.439 (meaning the odds are increasing). i.e., Odds ratio=1.439,
which indicates the males are likely to terminate their treatment early.
 Symptom severity is a negative and significant (B=-0.339,
S.E.=0.133, p=0.011) predictor of the probability of early
termination. The OR indicates that for every one unit increment on
the predictor, the odd of terminating increase by a factor of 0.713
(meaning that the odds are decreasing).
 Genderid is a non-significant predictor of early termination (B=-
1.479, S.E.=0.949,p=0.119).[Had the predictor been significant, then
the negative coefficient would be taken as an indicator that females
(coded 1) are less likely to terminate early than males.]

You might also like