0% found this document useful (0 votes)
7 views53 pages

Regression Analysis Model Building Guide

Uploaded by

kanizf933
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views53 pages

Regression Analysis Model Building Guide

Uploaded by

kanizf933
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

lOMoARcPSD|54931287

SM SBE13E Chapter 16

Methods Of Economic Statistics (香港中文大學)

Scan to open on Studocu

Studocu is not sponsored or endorsed by any college or university


Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])
lOMoARcPSD|54931287

Chapter 16
Regression Analysis: Model Building

Learning Objectives

1. Learn how the general linear model can be used to model problems involving curvilinear
relationships.

2. Understand the concept of interaction and how it can be accounted for in the general linear model.

3. Understand how an F test can be used to determine when to add or delete one or more variables.

4. Develop an appreciation for the complexities involved in solving larger regression analysis problems.

5. Understand how variable selection procedures can be used to choose a set of independent variables
for an estimated regression equation.

6. Learn how analysis of variance and experimental design problems can be analyzed using a regression
model.

7. Know how the Durbin-Watson test can be used to test for autocorrelation.

16 - 1
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Solutions:

1. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 362.1 362.13 6.85 0.059
x 1 362.1 362.13 6.85 0.059
Error 4 211.4 52.84
Total 5 573.5

Model Summary

S R-sq R-sq(adj) R-sq(pred)


7.26928 63.14% 53.93% 0.00%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -6.8 14.2 -0.48 0.658
x 1.230 0.470 2.62 0.059 1.00

Regression Equation

y = -6.8 + 1.230 x

b. Since the p-value corresponding to F = 6.85 is .059 >  the relationship is not significant.

c.

45

40

35

30
y

25

20

15

10
20 25 30 35 40 45

The scatter diagram suggests that a curvilinear relationship may be appropriate.

16 - 2
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

d. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 541.85 270.92 25.68 0.013
x 1 220.94 220.94 20.94 0.020
xsq 1 179.72 179.72 17.03 0.026
Error 3 31.65 10.55
Total 5 573.50

Model Summary

S R-sq R-sq(adj) R-sq(pred)


3.24822 94.48% 90.80% 76.85%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -168.9 39.8 -4.24 0.024
x 12.19 2.66 4.58 0.020 161.00
xsq -0.1770 0.0429 -4.13 0.026 161.00

Regression Equation

y = -168.9 + 12.19 x - 0.1770 xsq

e. Since the p-value corresponding to F = 25.68 is .013 <  the relationship is significant.

f. ŷ = -168.9 + 12.19(25) - 0.1770(25)2 = 25.225

2. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 59.39 59.39 4.76 0.117
x 1 59.39 59.39 4.76 0.117
Error 3 37.41 12.47
Total 4 96.80

Model Summary

S R-sq R-sq(adj) R-sq(pred)


3.53110 61.36% 48.48% 0.00%

16 - 3
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 9.32 4.20 2.22 0.113
x 0.424 0.194 2.18 0.117 1.00

Regression Equation

y = 9.32 + 0.424 x

The high p-value (.117) indicates a weak relationship; note that 61.4% of the variability in y has
been explained by x.

b. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 93.529 46.765 28.60 0.034
x 1 48.974 48.974 29.95 0.032
xsq 1 34.135 34.135 20.87 0.045
Error 2 3.271 1.635
Total 4 96.800

Model Summary

S R-sq R-sq(adj) R-sq(pred)


1.27880 96.62% 93.24% 81.98%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -8.10 4.10 -1.97 0.187
x 2.413 0.441 5.47 0.032 39.22
xsq -0.0480 0.0105 -4.57 0.045 39.22

Regression Equation

y = -8.10 + 2.413 x - 0.0480 xsq

At the .05 level of significance, the relationship is significant; the fit is excellent.

c. ŷ = -8.10 + 2.413(20) - 0.048(20)2 = 20.96

3. a. The scatter diagram shows some evidence of a possible linear relationship.

16 - 4
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

b. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 18.46 18.461 4.37 0.075
x 1 18.46 18.461 4.37 0.075
Error 7 29.54 4.220
Lack-of-Fit 5 16.87 3.374 0.53 0.753
Pure Error 2 12.67 6.333
Total 8 48.00

Model Summary

S R-sq R-sq(adj) R-sq(pred)


2.05423 38.46% 29.67% 0.00%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 2.32 1.89 1.23 0.258
x 0.637 0.304 2.09 0.075 1.00

Regression Equation

y = 2.32 + 0.637 x

c. The following standardized residual plot indicates that the constant variance assumption is
not satisfied.

2.0

1.5
St andardized Residual

1.0

0.5

0.0

-0.5

-1.0

-1.5
3 4 5 6 7 8
Fit t ed Value

d. The logarithmic transformation does not appear to eliminate the wedged-shaped pattern in the above
residual plot. The reciprocal transformation does, however, remove the wedge-shaped pattern.

16 - 5
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Neither transformation provides a good fit. The Minitab output for the reciprocal transformation and
the corresponding standardized residual plot are shown below.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 0.010501 0.010501 4.19 0.080
x 1 0.010501 0.010501 4.19 0.080
Error 7 0.017563 0.002509
Lack-of-Fit 5 0.007790 0.001558 0.32 0.869
Pure Error 2 0.009774 0.004887
Total 8 0.028064

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.0500902 37.42% 28.48% 2.80%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 0.2750 0.0460 5.98 0.001
x -0.01518 0.00742 -2.05 0.080 1.00

Regression Equation

1/y = 0.2750 - 0.01518 x

2.0

1.5
St andardized Residual

1.0

0.5

0.0

-0.5

-1.0

0.150 0.175 0.200 0.225 0.250


Fit t ed Value

16 - 6
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

4. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 33223 33223 31.86 0.005
Vehicle Speed 1 33223 33223 31.86 0.005
Error 4 4172 1043
Total 5 37395

Model Summary

S R-sq R-sq(adj) R-sq(pred)


32.2940 88.84% 86.06% 65.18%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 943.0 59.4 15.88 0.000
Vehicle Speed 8.71 1.54 5.64 0.005 1.00

Regression Equation

Traffic Flow = 943.0 + 8.71 Vehicle Speed

b. p-value = .005 <  = .01; reject H0

5. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 36643.4 18321.7 73.15 0.003
Vehicle Speed 1 5756.6 5756.6 22.98 0.017
Vehicle Speed SQ 1 3420.2 3420.2 13.65 0.034
Error 3 751.4 250.5
Total 5 37394.8

Model Summary

S R-sq R-sq(adj) R-sq(pred)


15.8264 97.99% 96.65% 92.96%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 433 141 3.06 0.055
Vehicle Speed 37.43 7.81 4.79 0.017 106.47
Vehicle Speed SQ -0.383 0.104 -3.70 0.034 106.47

16 - 7
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Regression Equation

Traffic Flow = 433 + 37.43 Vehicle Speed - 0.383 Vehicle Speed SQ

b. p-value = .003 <  = .05; reject H0.

c. The Minitab output follows:

Regression Equation

Traffic Flow = 433 + 37.43 Vehicle Speed - 0.383 Vehicle Speed SQ

Variable Setting
Vehicle Speed 38
Vehicle Speed SQ 1444

Fit SE Fit 95% CI 95% PI


1302.01 9.92840 (1270.41, 1333.61) (1242.55, 1361.47)

Analysis of Variance

The fitted value is 1302.01, with a standard deviation of 9.9284. The 95% confidence interval is
1270.41 to 1333.61; the 95% prediction interval is 1242.55 to 1361.47.

6. a. The scatter diagram follows:

1.8

1.6

1.4

1.2
Distance

0.8

0.6

0.4

0.2

0
5 10 15 20 25 30 35

Number

b. No; the relationship appears to be curvilinear.

c. Several possible models can be fitted to these data, as shown below:

ŷ = 2.90 - 0.185x + .00351x2 Ra2  .91

16 - 8
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

1
yˆ  0.0468  14.4   Ra2  .91
 x
7. a.

1200

1000

800
Mortgage ($)

600

400

200

0
700 750 800 850 900 950 1000 1050 1100
Ren ($)

The scatter diagram suggests that there is a curvilinear relationship, so a simple linear regression
model may not be appropriate.

b. The Minitab output follows.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 153962 153962 24.81 0.001
Rent ($) 1 153962 153962 24.81 0.001
Error 8 49653 6207
Total 9 203614

Model Summary

S R-sq R-sq(adj) R-sq(pred)


78.7819 75.61% 72.57% 63.68%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -198 188 -1.05 0.322
Rent ($) 1.070 0.215 4.98 0.001 1.00

Regression Equation

Mortgage ($) = -198 + 1.070 Rent ($)

16 - 9
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Fits and Diagnostics for Unusual Observations

Mortgage
Obs ($) Fit Resid Std Resid
1 539.0 700.8 -161.8 -2.17 R

R Large residual

Versus Fits
(response is Mortgage ($))

1.0

0.5
St andardized Residual

0.0

-0.5

-1.0

-1.5

-2.0

-2.5
600 700 800 900 1000
Fit t ed Value

The standardized residual plot suggests that a curvilinear regression model will be more appropriate
than a simple linear regression model.

c. The Minitab output follows.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 182947 91474 30.98 0.000
Rent ($) 1 22663 22663 7.68 0.028
RentSq 1 28986 28986 9.82 0.017
Error 7 20667 2952
Total 9 203614

Model Summary

S R-sq R-sq(adj) R-sq(pred)


54.3363 89.85% 86.95% 81.76%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 3966 1335 2.97 0.021
Rent ($) -8.26 2.98 -2.77 0.028 404.96
RentSq 0.00513 0.00164 3.13 0.017 404.96

Regression Equation

Mortgage ($) = 3966 - 8.26 Rent ($) + 0.00513 RentSq

16 - 10
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Fits and Diagnostics for Unusual Observations

Mortgage
Obs ($) Fit Resid Std Resid
1 539.0 646.8 -107.8 -2.23 R

R Large residual

d. The estimated regression equation in part (d) provides a better fit.

8. a. The scatter diagram follows:

4500
4000
3500
3000
2500
Price

2000
1500
1000
500
0
12 14 16 18 20
Rating

A simple linear regression model does not appear to be appropriate. There appears to be a curvilinear
relationship between the two variables.

b. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 12643604 6321802 14.15 0.001
Rating 1 3274143 3274143 7.33 0.019
RatingSq 1 3937445 3937445 8.82 0.012
Error 12 5359686 446641
Lack-of-Fit 4 883453 220863 0.39 0.807
Pure Error 8 4476233 559529
Total 14 18003290

Model Summary

S R-sq R-sq(adj) R-sq(pred)


668.312 70.23% 65.27% 44.69%

16 - 11
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 33829 13657 2.48 0.029
Rating -4571 1688 -2.71 0.019 296.05
RatingSq 153.6 51.7 2.97 0.012 296.05

Regression Equation

Price ($1000) = 33829 - 4571 Rating + 153.6 RatingSq

Fits and Diagnostics for Unusual Observations

Price
Obs ($1000) Fit Resid Std Resid
2 4000 2420 1580 2.76 R
15 70 361 -291 -0.82 X

R Large residual
X Unusual X

c. The Minitab output for the model using the base 10 log of Price as the dependent variable follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 3.7885 3.78851 53.81 0.000
Rating 1 3.7885 3.78851 53.81 0.000
Error 13 0.9153 0.07041
Lack-of-Fit 5 0.3268 0.06536 0.89 0.531
Pure Error 8 0.5885 0.07357
Total 14 4.7038

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.265348 80.54% 79.04% 72.28%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -2.208 0.658 -3.36 0.005
Rating 0.2857 0.0390 7.34 0.000 1.00

Regression Equation

logPrice = -2.208 + 0.2857 Rating

d. The model in part (c) is preferred because it provides a better fit.

9. a.

16 - 12
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

A simple linear regression model appears to be appropriate.

b.

Note the line drawn through the data. This line indicates a possible curvilinar relationship between
these two variables.

c. In the Minitab output that follows IndexSq denotes the square of the Cost-of-Living Index.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 609.96 203.319 27.70 0.000
Cost-of-Living Index 1 39.81 39.814 5.42 0.024
IndexSq 1 39.07 39.067 5.32 0.026
Income 1 261.45 261.446 35.62 0.000
Error 46 337.67 7.341
Total 49 947.62
Model Summary

16 - 13
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

S R-sq R-sq(adj) R-sq(pred)


2.70934 64.37% 62.04% 57.71%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 49.2 17.3 2.85 0.006
Cost-of-Living Index -0.673 0.289 -2.33 0.024 159.51
IndexSq 0.00282 0.00122 2.31 0.026 163.08
Income 0.4042 0.0677 5.97 0.000 2.05

Regression Equation

Creative Class (%) = 49.2 - 0.673 Cost-of-Living Index + 0.00282 IndexSq +


0.4042 Income

Fits and Diagnostics for Unusual Observations

Creative
Obs Class (%) Fit Resid Std Resid
20 20.60 31.08 -10.48 -3.94 R
21 44.40 45.16 -0.76 -0.35 X
37 24.20 30.68 -6.48 -2.48 R
39 38.80 41.66 -2.86 -1.38 X
40 41.60 36.19 5.41 2.09 R
43 43.70 41.61 2.09 0.90 X

R Large residual
X Unusual X

At the .05 level of significance there is overall significance. And, each of the three independent
variables (Cost-of-Living Index, IndexSq, and Income) is significant.

d. Cost-of-Living Index = 99, IndexSq = 9801, and Income = 42.984

Estimate = 49.2 - .673(99) +.00282(9801) + .4042(42.984) = 27.58595%

The primary concern of using this estimate is that the estimated regression equation was developed
for metropolitan areas with a population of 1,000,000 or more. But, the population for Tucson is so
close to 1,000,000 that the estimated regression equation should still provide a good estimate. Note
to Instructor: the actual value of the percentage of the workforce in creative fields for Tucson
reported by Kiplinger was 31.1%.

10. a. SSR = SST - SSE = 1030

MSR = 1030 MSE = 520/25 = 20.8 F = 1030/20.8 = 49.52

Using Excel or Minitab, the p-value corresponding to F = 49.52 is .000.

Because p-value ≤ α, x1 is significant.

(520  100) / 2
b. F  48.3
100 / 23

Using Excel or Minitab, the p-value corresponding to F = 48.3 is .000.

16 - 14
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

  Because p-value ≤ α, the addition of variables x2 and x3 is significant.

11. a. SSE = SST - SSR = 1805 - 1760 = 45

MSR = 1760/4 = 440 MSE =45/25 = 1.8

F = 440/1.8 = 244.44

Using Excel or Minitab, the p-value corresponding to F = 244.44 is .000.

  Because p-value ≤ α, the overall relationship is significant.

b. SSE(x1, x2, x3, x4) = 45

c. SSE(x2, x3) = 1805 - 1705 = 100

(100  45) / 2
d. F  15.28
1.8
Using Excel or Minitab, the p-value corresponding to F = 15.28 is .000.

  Because p-value ≤ α, x1 and x4 contribute significantly to the model.


12. a. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 78.25 78.2452 61.86 0.000
Putting Ave. 1 78.25 78.2452 61.86 0.000
Error 132 166.95 1.2648
Lack-of-Fit 98 137.21 1.4001 1.60 0.060
Pure Error 34 29.74 0.8747
Total 133 245.20

Model Summary

S R-sq R-sq(adj) R-sq(pred)


1.12463 31.91% 31.40% 29.66%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 47.95 3.12 15.36 0.000
Putting Ave. 0.806 0.102 7.87 0.000 1.00

Regression Equation

Scoring Ave. = 47.95 + 0.806 Putting Ave.

b. A portion of the Minitab output follows:

16 - 15
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 229.677 76.559 641.25 0.000
Putting Ave. 1 122.801 122.801 1028.58 0.000
Greens in Reg. 1 149.336 149.336 1250.83 0.000
Drive Accuracy 1 0.000 0.000 0.00 0.964
Error 130 15.521 0.119
Total 133 245.197

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.345528 93.67% 93.52% 93.08%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 57.16 1.06 53.95 0.000
Putting Ave. 1.0319 0.0322 32.07 0.000 1.04
Greens in Reg. -23.102 0.653 -35.37 0.000 1.05
Drive Accuracy -0.022 0.481 -0.05 0.964 1.02

Regression Equation

Scoring Ave. = 57.16 + 1.0319 Putting Ave. - 23.102 Greens in Reg. - 0.022
Drive Accuracy

c. SSE(reduced) =166.95

SSE(full) = 15.521

MSE(full) = .119

SSE(reduced) - SSE(full) 166.95 - 15.521


F  number of extra terms  2  634.19
MSE(full) .119

The p-value associated with F = 634.19 (2 degrees of freedom numerator and 130 denominator) is
0.00. With a p-value < α =.05, the addition of the two independent variables is statistically
significant.

13. a. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 2.62806E+11 2.62806E+11 20.78 0.000
Putting Ave. 1 2.62806E+11 2.62806E+11 20.78 0.000
Error 132 1.66934E+12 12646499715
Lack-of-Fit 98 1.27492E+12 13009364286 1.12 0.361
Pure Error 34 3.94420E+11 11600595953
Total 133 1.93214E+12

Model Summary

S R-sq R-sq(adj) R-sq(pred)


112457 13.60% 12.95% 10.52%
Coefficients

16 - 16
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Term Coef SE Coef T-Value P-Value VIF


Constant 1508010 312129 4.83 0.000
Putting Ave. -46709 10246 -4.56 0.000 1.00

Regression Equation

Earnings ($) = 1508010 - 46709 Putting Ave.

b. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 1.11859E+12 3.72863E+11 59.58 0.000
Putting Ave. 1 4.76239E+11 4.76239E+11 76.10 0.000
Greens in Reg. 1 8.55323E+11 8.55323E+11 136.67 0.000
Drive Accuracy 1 16625430493 16625430493 2.66 0.106
Error 130 8.13554E+11 6258106800
Total 133 1.93214E+12

Model Summary

S R-sq R-sq(adj) R-sq(pred)


79108.2 57.89% 56.92% 54.09%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 961320 242576 3.96 0.000
Putting Ave. -64259 7366 -8.72 0.000 1.04
Greens in Reg. 1748381 149552 11.69 0.000 1.05
Drive Accuracy -179313 110014 -1.63 0.106 1.02

Regression Equation

Earnings ($) = 961320 - 64259 Putting Ave. + 1748381 Greens in Reg. -


179313 Drive Accuracy

c. SSE(reduced) =1.66934E+12

SSE(full) = 8.13554E+11

MSE(full) = 6258106800

SSE(reduced) - SSE(full) 1.67E+12 - 8.14E+11


F  number of extra terms  2  68.37
MSE(full) 6258106800

The p-value associated with F = 68.37 (2 degrees of freedom numerator and 130 denominator) is
less than .001. With a p-value < α =.05, the addition of the two independent variables is statistically
significant.

d. A portion of the Minitab output follows:

16 - 17
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 1.12596E+12 1.12596E+12 184.36 0.000
Scoring Ave. 1 1.12596E+12 1.12596E+12 184.36 0.000
Error 132 8.06183E+11 6107449382
Lack-of-Fit 107 7.66712E+11 7165531932 4.54 0.000
Pure Error 25 39471401762 1578856070
Total 133 1.93214E+12

Model Summary

S R-sq R-sq(adj) R-sq(pred)


78150.2 58.28% 57.96% 55.74%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 4998308 361864 13.81 0.000
Scoring Ave. -67765 4991 -13.58 0.000 1.00

Regression Equation

Earnings ($) = 4998308 - 67765 Scoring Ave.

Because the equation developed in part (d) provides a better fit and is a simpler model, it is preferred
over the equation developed in part (b).

14. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 3379.6 1689.82 35.41 0.000
Age 1 2748.2 2748.19 57.58 0.000
Blood_Pressure 1 1607.7 1607.66 33.69 0.000
Error 17 811.3 47.72
Total 19 4190.9

Model Summary

S R-sq R-sq(adj) R-sq(pred)


6.90826 80.64% 78.36% 74.26%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -110.9 16.5 -6.74 0.000
Age 1.315 0.173 7.59 0.000 1.11
Blood_Pressure 0.2964 0.0511 5.80 0.000 1.11

Regression Equation

Risk = -110.9 + 1.315 Age + 0.2964 Blood_Pressure

Fits and Diagnostics for Unusual Observations

16 - 18
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Obs Risk Fit Resid Std Resid


17 8.00 25.05 -17.05 -2.54 R

R Large residual

b. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 4 3672.11 918.03 26.54 0.000
Age 1 130.29 130.29 3.77 0.071
Blood_Pressure 1 58.14 58.14 1.68 0.214
Smoker 1 287.77 287.77 8.32 0.011
AgePress 1 11.37 11.37 0.33 0.575
Error 15 518.84 34.59
Total 19 4190.95

Model Summary

S R-sq R-sq(adj) R-sq(pred)


5.88127 87.62% 84.32% 79.34%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -123.2 56.9 -2.16 0.047
Age 1.513 0.780 1.94 0.071 30.87
Blood_Pressure 0.448 0.346 1.30 0.214 69.91
Smoker 8.87 3.07 2.88 0.011 1.37
AgePress -0.00276 0.00481 -0.57 0.575 71.91

Regression Equation

Risk = -123.2 + 1.513 Age + 0.448 Blood_Pressure + 8.87 Smoker - 0.00276


AgePress

Fits and Diagnostics for Unusual Observations

Obs Risk Fit Resid Std Resid


17 8.00 20.91 -12.91 -2.34 R

R Large residual

SSE(reduced) - SSE(full) 811.3  518.84


c. F  number of extra terms  2  4.23
MSE(full) 34.59

The p-value associated with F = 4.23 (2 numerator and 15 denominator DF) is .000

Because p-value ≤ α = .05, the addition of the two terms is significant.

16 - 19
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

15. a. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 6.4044 6.40437 29.41 0.000
H/9 1 6.4044 6.40437 29.41 0.000
Error 48 10.4512 0.21773
Lack-of-Fit 47 10.4480 0.22230 69.47 0.095
Pure Error 1 0.0032 0.00320
Total 49 16.8556

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.466619 38.00% 36.70% 33.20%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -0.253 0.735 -0.34 0.732
H/9 0.4527 0.0835 5.42 0.000 1.00

Regression Equation

ERA = -0.253 + 0.4527 H/9

Fits and Diagnostics for Unusual Observations

Obs ERA Fit Resid Std Resid


6 2.540 3.655 -1.115 -2.41 R
27 4.840 3.699 1.141 2.47 R
38 4.900 4.668 0.232 0.54 X
43 4.220 3.254 0.966 2.13 R

R Large residual
X Unusual X

b. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 13.114 4.37123 53.74 0.000
H/9 1 7.037 7.03715 86.51 0.000
HR/9 1 2.845 2.84529 34.98 0.000
BB/9 1 3.662 3.66208 45.02 0.000
Error 46 3.742 0.08134
Total 49 16.856

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.285210 77.80% 76.35% 74.41%

16 - 20
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -2.564 0.538 -4.76 0.000
H/9 0.5121 0.0551 9.30 0.000 1.16
HR/9 0.980 0.166 5.91 0.000 1.04
BB/9 0.3400 0.0507 6.71 0.000 1.12

Regression Equation

ERA = -2.564 + 0.5121 H/9 + 0.980 HR/9 + 0.3400 BB/9

Fits and Diagnostics for Unusual Observations

Obs ERA Fit Resid Std Resid


2 2.5300 3.1820 -0.6520 -2.35 R
27 4.8400 4.0219 0.8181 2.97 R

R Large residual

c. SSE(reduced) = 10.4512

SSE(full) = 3.7419

MSE(full) = .0813

SSE(reduced) - SSE(full) 10.4512 - 3.7419


F  number of extra terms  2  41.26
MSE(full) .0813

The p-value associated with F = 41.26 (2 degrees of freedom numerator and 46 denominator) is
.000. With a p-value < α =.05, the addition of the two independent variables is statistically
significant.

16 - 21
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

16. a. The sample correlation coefficients are as follows:

Weeks Age Educ Married Head Tenure Manager


Age 0.577
0.000

Educ 0.007 0.100


0.962 0.490

Married -0.130 -0.209 -0.151


0.370 0.145 0.296

Head -0.205 0.027 -0.156 -0.449


0.153 0.854 0.280 0.001

Tenure 0.398 0.459 0.174 -0.057 -0.046


0.004 0.001 0.228 0.692 0.750

Manager -0.198 0.097 0.160 0.073 -0.200 -0.113


0.167 0.504 0.266 0.616 0.164 0.435

Sales -0.134 0.137 0.124 -0.148 -0.013 0.097 -0.156


0.354 0.343 0.393 0.306 0.926 0.504 0.279

Cell Contents: Pearson correlation


P-Value

The independent variable most highly correlated with Weeks is Age. The Minitab output
corresponding to using Age as the independent variable follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 9161 9161.4 24.01 0.000
Age 1 9161 9161.4 24.01 0.000
Error 48 18316 381.6
Lack-of-Fit 23 8961 389.6 1.04 0.459
Pure Error 25 9355 374.2
Total 49 27478

Model Summary

S R-sq R-sq(adj) R-sq(pred)


19.5342 33.34% 31.95% 28.70%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -8.9 11.0 -0.80 0.425
Age 1.509 0.308 4.90 0.000 1.00

Regression Equation

Weeks = -8.9 + 1.509 Age

16 - 22
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Fits and Diagnostics for Unusual Observations

Obs Weeks Fit Resid Std Resid


23 80.00 37.93 42.07 2.18 R
26 94.00 84.71 9.29 0.53 X
36 4.00 45.47 -41.47 -2.15 R
39 80.00 84.71 -4.71 -0.27 X
41 98.00 59.06 38.94 2.04 R

R Large residual
X Unusual X

b. The Minitab Stepwise Regression output follows.

Stepwise Selection of Terms

Candidate terms: Age, Educ, Married, Head, Tenure, Manager, Sales

----Step 1---- -----Step 2---- -----Step 3---- -----Step 4----


Coef P Coef P Coef P Coef P
Constant -8.9 -9.1 -0.1 -0.07
Age 1.509 0.000 1.574 0.000 1.609 0.000 1.725 0.000
Manager -20.06 0.029 -24.61 0.006 -28.67 0.001
Head -14.31 0.012 -15.09 0.005
Sales -17.42 0.008

S 19.5342 18.7494 17.6857 16.5069


R-sq 33.34% 39.87% 47.64% 55.38%
R-sq(adj) 31.95% 37.31% 44.22% 51.41%
R-sq(pred) 28.70% 33.31% 37.85% 44.89%
Mallows’ Cp 22.52 17.81 11.83 5.87

α to enter = 0.05, α to remove = 0.05

The results suggest a model using four independent variables: Age, Manager, Head, and Sales. The
corresponding Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 4 15216 3804.0 13.96 0.000
Age 1 11539 11539.4 42.35 0.000
Head 1 2365 2364.7 8.68 0.005
Manager 1 3400 3400.2 12.48 0.001
Sales 1 2127 2126.7 7.80 0.008
Error 45 12261 272.5
Total 49 27478

Model Summary

S R-sq R-sq(adj) R-sq(pred)


16.5069 55.38% 51.41% 44.89%

16 - 23
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -0.07 9.84 -0.01 0.994
Age 1.725 0.265 6.51 0.000 1.04
Head -15.09 5.12 -2.95 0.005 1.05
Manager -28.67 8.12 -3.53 0.001 1.09
Sales -17.42 6.24 -2.79 0.008 1.05

Regression Equation

Weeks = -0.07 + 1.725 Age - 15.09 Head - 28.67 Manager - 17.42 Sales

Fits and Diagnostics for Unusual Observations

Obs Weeks Fit Resid Std Resid


24 7.00 39.61 -32.61 -2.09 R
39 80.00 89.47 -9.47 -0.70 X

R Large residual
X Unusual X

c. The results using Minitab’s Forward Selection procedure are the same as the results using Minitab’s
Stepwise procedure in part (b).

d. The results using Minitab’s Backward Elimination procedure are shown below:

Backward Elimination of Terms

Candidate terms: Age, Educ, Married, Head, Tenure, Manager, Sales

-----Step 1---- -----Step 2---- -----Step 3---- -----Step 4----


Coef P Coef P Coef P Coef P
Constant 22.9 13.6 13.1 -0.07
Age 1.509 0.000 1.521 0.000 1.637 0.000 1.725 0.000
Educ -0.613 0.516
Married -10.74 0.081 -9.85 0.098 -9.76 0.099
Head -19.78 0.002 -19.01 0.002 -19.41 0.001 -15.09 0.005
Tenure 0.426 0.366 0.373 0.418
Manager -26.74 0.003 -27.71 0.001 -28.99 0.001 -28.67 0.001
Sales -18.56 0.005 -18.99 0.004 -18.97 0.004 -17.42 0.008

S 16.3497 16.2408 16.1794 16.5069


R-sq 59.14% 58.72% 58.08% 55.38%
R-sq(adj) 52.33% 52.96% 53.32% 51.41%
R-sq(pred) 43.80% 46.00% 46.60% 44.89%
Mallows’ Cp 8.00 6.43 5.09 5.87

α to remove = 0.05

These results also suggest using the model with four independent variables: Age, Head, Manager, and
Sales.

16 - 24
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

e. The results using Mintab’s Best-Subset procedure are shown below:

Response is Weeks

M M
a T a
r e n S
E r H n a a
A d i e u g l
R-Sq R-Sq Mallows g u e a r e e
Vars R-Sq (adj) (pred) Cp S e c d d e r s
1 33.3 32.0 28.7 22.5 19.534 X
1 15.8 14.0 10.0 40.6 21.954 X
2 39.9 37.3 33.3 17.8 18.749 X X
2 38.2 35.6 30.2 19.5 19.005 X X
3 47.6 44.2 37.8 11.8 17.686 X X X
3 46.8 43.3 38.1 12.7 17.831 X X X
4 55.4 51.4 44.9 5.9 16.507 X X X X
4 49.1 44.6 37.4 12.3 17.628 X X X X
5 58.1 53.3 46.6 5.1 16.179 X X X X X
5 56.0 51.0 43.9 7.3 16.582 X X X X X
6 58.7 53.0 46.0 6.4 16.241 X X X X X X
6 58.3 52.5 44.6 6.8 16.318 X X X X X X
7 59.1 52.3 43.8 8.0 16.350 X X X X X X X

The results suggest a model using five independent variables: Age, Married, Head, Manager, and Sales.

The corresponding Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 5 15959.5 3191.9 12.19 0.000
Age 1 9982.6 9982.6 38.13 0.000
Married 1 743.5 743.5 2.84 0.099
Head 1 3103.5 3103.5 11.86 0.001
Manager 1 3473.0 3473.0 13.27 0.001
Sales 1 2465.3 2465.3 9.42 0.004
Error 44 11518.0 261.8
Lack-of-Fit 37 8489.5 229.4 0.53 0.900
Pure Error 7 3028.5 432.6
Total 49 27477.5

Model Summary

S R-sq R-sq(adj) R-sq(pred)


16.1794 58.08% 53.32% 46.60%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 13.1 12.4 1.05 0.298
Age 1.637 0.265 6.18 0.000 1.08
Married -9.76 5.79 -1.69 0.099 1.35
Head -19.41 5.64 -3.44 0.001 1.32
Manager -28.99 7.96 -3.64 0.001 1.09
Sales -18.97 6.18 -3.07 0.004 1.08

16 - 25
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Regression Equation

Weeks = 13.1 + 1.637 Age - 9.76 Married - 19.41 Head - 28.99 Manager -
18.97 Sales

Fits and Diagnostics for Unusual Observations

Obs Weeks Fit Resid Std Resid


10 13.00 47.68 -34.68 -2.24 R
24 7.00 40.95 -33.95 -2.22 R

R Large residual

17. The output obtained using Minitab’s Best Subset Regression follows:

Response is Scoring Ave.

G D
r r
e P i
e u v
n t e D
s t r
i A i
i n c v
n g c e
u G
R A r r
e v a e
R-Sq R-Sq g e c e
Vars R-Sq (adj) (pred) Mallows Cp S . . y n
1 43.4 43.0 41.7 1034.5 1.0254 X
1 31.9 31.4 29.7 1270.7 1.1246 X
2 93.7 93.6 93.3 2.2 0.34421 X X
2 60.7 60.1 58.8 681.1 0.85801 X X
3 93.7 93.5 93.1 4.2 0.34553 X X X
3 93.7 93.5 93.1 4.2 0.34553 X X X
4 93.7 93.5 92.7 5.0 0.34524 X X X X

The Best Subset Regression output indicates that several models produce R-Sq values over .90, but a
model using two independent variables, Greens in Reg. and Putting Ave., may be a good choice. The
Minitab output for this model follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 229.676 114.838 969.26 0.000
Greens in Reg. 1 151.431 151.431 1278.11 0.000
Putting Ave. 1 123.274 123.274 1040.46 0.000
Error 131 15.521 0.118
Lack-of-Fit 130 15.483 0.119 3.10 0.429
Pure Error 1 0.038 0.038
Total 133 245.197

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.344209 93.67% 93.57% 93.26%

16 - 26
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 57.148 0.989 57.76 0.000
Greens in Reg. -23.106 0.646 -35.75 0.000 1.04
Putting Ave. 1.0320 0.0320 32.26 0.000 1.04

Regression Equation

Scoring Ave. = 57.148 - 23.106 Greens in Reg. + 1.0320 Putting Ave.

18. a. Because the independent variable most highly correlated with RPG is OBP, it will
provide the best one-variable estimated regression equation. The Minitab output
using OBP to predict RPG follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 72.11 72.1076 78.85 0.000
OBP 1 72.11 72.1076 78.85 0.000
Error 18 16.46 0.9145
Total 19 88.57

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.956308 81.41% 80.38% 75.42%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -4.05 1.01 -4.02 0.001
OBP 27.55 3.10 8.88 0.000 1.00

Regression Equation

RPG = -4.05 + 27.55 OBP

Fits and Diagnostics for Unusual Observations

Std
Obs RPG Fit Resid Resid
15 5.150 3.198 1.952 2.13 R

R Large residual

16 - 27
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

b. The output using Minitab’s Stepwise Regression procedure using Alpha-to-Enter = 0.05 and Alpha-
to-Remove = 0.05 follows:

Candidate terms: H, 2B, 3B, HR, RBI, BB, SO, SB, CS, OBP, SLG, AVG

-----Step 1---- -----Step 2----- -----Step 3-----


Coef P Coef P Coef P
Constant -4.05 -1.595 -0.981
OBP 27.55 0.000 17.16 0.000 25.14 0.000
HR 0.0711 0.001 0.0695 0.000
AVG -12.62 0.005

S 0.956308 0.692987 0.555522


R-sq 81.41% 90.78% 94.43%
R-sq(adj) 80.38% 89.70% 93.38%
R-sq(pred) 75.42% 87.39% 89.20%
Mallows’ Cp 66.65 26.99 12.79

α to enter = 0.05, α to remove = 0.05

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 83.631 27.8771 90.33 0.000
HR 1 7.912 7.9120 25.64 0.000
OBP 1 14.610 14.6099 47.34 0.000
AVG 1 3.226 3.2262 10.45 0.005
Error 16 4.938 0.3086
Total 19 88.569

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.555522 94.43% 93.38% 89.20%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -0.981 0.776 -1.26 0.224
HR 0.0695 0.0137 5.06 0.000 2.24
OBP 25.14 3.65 6.88 0.000 4.11
AVG -12.62 3.90 -3.23 0.005 2.76

Regression Equation

RPG = -0.981 + 0.0695 HR + 25.14 OBP - 12.62 AVG

Fits and Diagnostics for Unusual Observations

Std
Obs RPG Fit Resid Resid
15 5.150 4.192 0.958 2.39 R

R Large residual

16 - 28
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Using less sensitive values for Alpha-to-Enter and Alpha-to-Remove will provide a model with
additional independent variables. For example, the output using Minitab’s Stepwise Regression
procedure using Alpha-to-Enter = 0.10 and Alpha-to-Remove = 0.10 follows:

Stepwise Selection of Terms

Candidate terms: H, 2B, 3B, HR, RBI, BB, SO, SB, CS, OBP, SLG, AVG

-----Step 1---- -----Step 2----- -----Step 3----- -----Step 4-----


Coef P Coef P Coef P Coef P
Constant -4.05 -1.595 -0.981 -0.616
OBP 27.55 0.000 17.16 0.000 25.14 0.000 26.61 0.000
HR 0.0711 0.001 0.0695 0.000 0.0681 0.000
AVG -12.62 0.005 -16.50 0.001
3B 0.1823 0.079
BB

S 0.956308 0.692987 0.555522 0.515891


R-sq 81.41% 90.78% 94.43% 95.49%
R-sq(adj) 80.38% 89.70% 93.38% 94.29%
R-sq(pred) 75.42% 87.39% 89.20% 90.03%
Mallows’ Cp 66.65 26.99 12.79 10.04

------Step 5------
Coef P
Constant -0.909
OBP 32.18 0.000
HR 0.1088 0.000
AVG -21.51 0.000
3B 0.2439 0.010
BB -0.02231 0.011

S 0.420960
R-sq 97.20%
R-sq(adj) 96.20%
R-sq(pred) 93.86%
Mallows’ Cp 4.46

α to enter = 0.1, α to remove = 0.1

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 5 86.088 17.2176 97.16 0.000
3B 1 1.580 1.5798 8.92 0.010
HR 1 6.936 6.9361 39.14 0.000
BB 1 1.511 1.5112 8.53 0.011
OBP 1 15.667 15.6669 88.41 0.000
AVG 1 5.650 5.6496 31.88 0.000
Error 14 2.481 0.1772
Total 19 88.569

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.420960 97.20% 96.20% 93.86%

16 - 29
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -0.909 0.617 -1.47 0.163
3B 0.2439 0.0817 2.99 0.010 1.55
HR 0.1088 0.0174 6.26 0.000 6.26
BB -0.02231 0.00764 -2.92 0.011 8.02
OBP 32.18 3.42 9.40 0.000 6.28
AVG -21.51 3.81 -5.65 0.000 4.58

Regression Equation

RPG = -0.909 + 0.2439 3B + 0.1088 HR - 0.02231 BB + 32.18 OBP - 21.51 AVG

The following output using Minitab’s Best Subset procedure also confirms that a variety of models
will provide a good fit.

Response is RPG

R O S A
R-Sq R-Sq Mallows 2 3 H B B S S C B L V
Vars R-Sq (adj) (pred) Cp S H B B R I B O B S P G G
1 81.4 80.4 75.4 66.6 0.95631 X
1 78.9 77.7 73.1 77.9 1.0192 X
2 90.8 89.7 87.4 27.0 0.69299 X X
2 88.4 87.0 83.5 37.8 0.77872 X X
3 94.4 93.4 89.2 12.8 0.55552 X X X
3 94.4 93.3 87.9 13.0 0.55820 X X X
4 95.8 94.6 89.9 8.8 0.50014 X X X X
4 95.5 94.3 90.0 10.0 0.51589 X X X X
5 97.2 96.2 93.9 4.5 0.42096 X X X X X
5 97.2 96.2 93.8 4.6 0.42336 X X X X X
6 97.6 96.6 92.6 4.5 0.40042 X X X X X X
6 97.5 96.4 94.1 5.1 0.41198 X X X X X X
7 98.2 97.2 94.4 3.9 0.36245 X X X X X X X
7 98.2 97.1 94.4 4.1 0.36664 X X X X X X X
8 98.3 97.1 93.3 5.3 0.36471 X X X X X X X X
8 98.3 97.0 93.3 5.7 0.37269 X X X X X X X X
9 98.4 97.0 92.4 7.1 0.37506 X X X X X X X X X
9 98.4 96.9 91.3 7.3 0.38077 X X X X X X X X X
10 98.4 96.7 91.8 9.0 0.39477 X X X X X X X X X X
10 98.4 96.7 88.1 9.0 0.39496 X X X X X X X X X X
11 98.4 96.3 86.0 11.0 0.41758 X X X X X X X X X X X
11 98.4 96.2 88.6 11.0 0.41848 X X X X X X X X X X X
12 98.4 95.7 83.8 13.0 0.44629 X X X X X X X X X X X X

It would be difficult to make an argument that there is one best model given these results. The five
variable model identified using Minitab’s Stepwise Regression procedure with Alpha-to-Enter =
0.10 and Alpha-to-Remove = 0.10 seems like a reasonable choice. The Minitab regression output
corresponding to this model follows:

16 - 30
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 5 86.088 17.2176 97.16 0.000
3B 1 1.580 1.5798 8.92 0.010
HR 1 6.936 6.9361 39.14 0.000
BB 1 1.511 1.5112 8.53 0.011
OBP 1 15.667 15.6669 88.41 0.000
AVG 1 5.650 5.6496 31.88 0.000
Error 14 2.481 0.1772
Total 19 88.569

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.420960 97.20% 96.20% 93.86%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -0.909 0.617 -1.47 0.163
3B 0.2439 0.0817 2.99 0.010 1.55
HR 0.1088 0.0174 6.26 0.000 6.26
BB -0.02231 0.00764 -2.92 0.011 8.02
OBP 32.18 3.42 9.40 0.000 6.28
AVG -21.51 3.81 -5.65 0.000 4.58

Regression Equation

RPG = -0.909 + 0.2439 3B + 0.1088 HR - 0.02231 BB + 32.18 OBP - 21.51 AVG

The corresponding standardized residual plot follows:

1
St andardized Residual

-1

-2
1 2 3 4 5 6 7 8 9 10
Fit t ed Value

The standardized residual plot does not indicate any reason to question the usual assumptions
regarding the error term. Thus, the estimated regression equation using OBP, HR, AVG, 3B, and BB
appears to be a good choice.

16 - 31
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

19. See the solution to Exercise 14 in this chapter. The Minitab output using the best
subsets regression procedure follows:

Response is Risk

B
l
o
o
d P
_ r
P A A e
r g g s
e S e e s
s m P S S
s o r m m
A u k e o o
R-Sq R-Sq Mallows g r e s k k
Vars R-Sq (adj) (pred) Cp S e e r s e e
1 63.3 61.3 56.6 27.6 9.2430 X
1 54.8 52.3 44.5 37.7 10.262 X
2 80.6 78.4 74.3 9.0 6.9083 X X
2 79.8 77.4 72.4 10.0 7.0538 X X
3 87.9 85.6 81.4 2.4 5.6313 X X X
3 87.3 85.0 80.9 3.0 5.7566 X X X
4 88.4 85.3 80.5 3.8 5.6908 X X X X
4 87.9 84.7 78.6 4.4 5.8111 X X X X
5 88.5 84.4 76.4 5.6 5.8571 X X X X X
5 88.4 84.3 77.1 5.8 5.8876 X X X X X
6 89.1 84.0 73.6 7.0 5.9391 X X X X X X

This output suggests that the model involving Age, Blood Pressure, and Smoker is the preferred
model; the Minitab output for this model follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 3660.7 1220.25 36.82 0.000
Age 1 1394.8 1394.84 42.09 0.000
Blood_Pressure 1 1027.4 1027.35 31.00 0.000
Smoker 1 281.1 281.10 8.48 0.010
Error 16 530.2 33.14
Total 19 4190.9

Model Summary

S R-sq R-sq(adj) R-sq(pred)


5.75657 87.35% 84.98% 80.92%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant -91.8 15.2 -6.03 0.000
Age 1.077 0.166 6.49 0.000 1.46
Blood_Pressure 0.2518 0.0452 5.57 0.000 1.25
Smoker 8.74 3.00 2.91 0.010 1.36

16 - 32
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Regression Equation

Risk = -91.8 + 1.077 Age + 0.2518 Blood_Pressure + 8.74 Smoker

Fits and Diagnostics for Unusual Observations

Obs Risk Fit Resid Std Resid


17 8.00 21.11 -13.11 -2.42 R

R Large residual

20. The dummy variables are defined as follows:

x1 x2 x3 Treatment
0 0 0 A
1 0 0 B
0 1 0 C
0 0 1 D

E(y) = 0 + 1 x1 + 2 x2 + 3 x3

21. The dummy variables are defined as follows:

x1 x2 Treatment
0 0 1
1 0 2
0 1 3

x3 = 0 if block 1 and 1 if block 2

E(y) = 0 + 1x1 + 2x2 + 3x3

22. Factor A

x1 = 0 if level 1 and 1 if level 2

Factor B
x2 x3 Level
0 0 1
1 0 2
0 1 3

E(y) = 0 + 1 x1 + 2 x2 + 3 x1x2 + 4 x1x3

23. a. The dummy variables are defined as follows:

D1 D2 Mfg.
0 0 1
1 0 2
0 1 3

E(y) = 0 + 1 D1 + 2 D2

b. The Minitab output follows:

16 - 33
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 104.000 52.000 10.64 0.004
D1 1 50.000 50.000 10.23 0.011
D2 1 8.000 8.000 1.64 0.233
Error 9 44.000 4.889
Total 11 148.000

Model Summary

S R-sq R-sq(adj) R-sq(pred)


2.21108 70.27% 63.66% 47.15%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 23.00 1.11 20.80 0.000
D1 5.00 1.56 3.20 0.011 1.33
D2 -2.00 1.56 -1.28 0.233 1.33

Regression Equation

Time = 23.00 + 5.00 D1 - 2.00 D2

c. H 0 : 1 = 2 = 0

d. The p-value of .004 is less than  = .05; therefore, we can reject H0 and conclude that the mean time
to mix a batch of material is the same for each manufacturer.

24. a. The dummy variables are defined as follows:

D1 D2 D3 Paint
0 0 0 1
1 0 0 2
0 1 0 3
0 0 1 4

The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 330.00 110.00 2.54 0.093
D1 1 90.00 90.00 2.08 0.168
D2 1 22.50 22.50 0.52 0.481
D3 1 302.50 302.50 6.99 0.018
Error 16 692.00 43.25
Total 19 1022.00

16 - 34
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Model Summary

S R-sq R-sq(adj) R-sq(pred)


6.57647 32.29% 19.59% 0.00%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 133.00 2.94 45.22 0.000
D1 6.00 4.16 1.44 0.168 1.50
D2 3.00 4.16 0.72 0.481 1.50
D3 11.00 4.16 2.64 0.018 1.50

Regression Equation

Time = 133.00 + 6.00 D1 + 3.00 D2 + 11.00 D3

The appropriate hypothesis test is:

H 0 : 1 = 2 = 3 = 0

The p-value of .093 is greater than  = .05; therefore, at the 5% level of significance we
cannot reject H0.

b. Note: Estimating the mean drying for paint 2 using the estimated regression equations developed in
part (a) may not be the best approach because at the 5% level of significance, we cannot reject H0.
But, if we want to use the output, we would proceed as follows.

Time = 133 + 6(1) + 3(0) +11(0) = 139

25. X1 = 0 if computerized analyzer, 1 if electronic analyzer

X2 and X3 are defined as follows:

X2 X3 Car
0 0 1
1 0 2
0 1 3

The complete data set and the Minitab output are shown below:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 3 289.00 96.33 9.17 0.100
X1 1 216.00 216.00 20.57 0.045
X2 1 12.25 12.25 1.17 0.393
X3 1 72.25 72.25 6.88 0.120
Error 2 21.00 10.50
Total 5 310.00

16 - 35
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Model Summary

S R-sq R-sq(adj) R-sq(pred)


3.24037 93.23% 83.06% 39.03%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 52.00 2.65 19.65 0.003
X1 -12.00 2.65 -4.54 0.045 1.00
X2 3.50 3.24 1.08 0.393 1.33
X3 8.50 3.24 2.62 0.120 1.33

Regression Equation

Time = 52.00 - 12.00 X1 + 3.50 X2 + 8.50 X3

To test for any significant difference between the two analyzers we must test H0: 1 Since the p-
value corresponding to t = -4.54 is .045 <  = .05, we reject H0: 0 the time to do a tune-up is
not the same for the two analyzers.

26. Size = 0 if a small advertisement and 1 if a large advertisement.

DesignB and DesignC are defined as follows:

Advertisement
DesignB DesignC Design
0 0 A
1 0 B
0 1 C

LargeDesignB denotes the interaction between Large and DesignB

LargeDesignC denotes the interaction between Large and DesignC

The complete data set and the Minitab output are shown below:

Number Size DesignB DesignC LargeDesignB LargeDesignC


8 0 0 0 0 0
12 0 0 0 0 0
22 0 1 0 0 0
14 0 1 0 0 0
10 0 0 1 0 0
18 0 0 1 0 0
12 1 0 0 0 0
8 1 0 0 0 0
26 1 1 0 1 0
30 1 1 0 1 0
18 1 0 1 0 1
14 1 0 1 0 1

16 - 36
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 5 448.000 89.6000 5.60 0.029
Size 1 0.000 0.0000 0.00 1.000
DesignB 1 64.000 64.0000 4.00 0.092
DesignC 1 16.000 16.0000 1.00 0.356
LargeDesignB 1 50.000 50.0000 3.12 0.128
LargeDesignC 1 2.000 2.0000 0.12 0.736
Error 6 96.000 16.0000
Total 11 544.000

Model Summary

S R-sq R-sq(adj) R-sq(pred)


4 82.35% 67.65% 29.41%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 10.00 2.83 3.54 0.012
Size 0.00 4.00 0.00 1.000 3.00
DesignB 8.00 4.00 2.00 0.092 2.67
DesignC 4.00 4.00 1.00 0.356 2.67
LargeDesignB 10.00 5.66 1.77 0.128 3.33
LargeDesignC 2.00 5.66 0.35 0.736 3.33

Regression Equation

Number = 10.00 + 0.00 Size + 8.00 DesignB + 4.00 DesignC + 10.00


LargeDesignB + 2.00 LargeDesignC

Overall the model is significant because the p-value corresponding to F = 5.60 < α = .05.

Individually, none of the variables are significant using α = .05

The Minitab output using only Design B follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 294.0 294.00 11.76 0.006
DesignB 1 294.0 294.00 11.76 0.006
Error 10 250.0 25.00
Total 11 544.0

Model Summary

S R-sq R-sq(adj) R-sq(pred)


5 54.04% 49.45% 27.84%

16 - 37
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 12.50 1.77 7.07 0.000
DesignB 10.50 3.06 3.43 0.006 1.00

Regression Equation

Number = 12.50 + 10.50 DesignB

Fits and Diagnostics for Unusual Observations

Obs Number Fit Resid Std Resid


4 14.00 23.00 -9.00 -2.08 R

R Large residual

Thus, DesignB is significant using α = .05. However, the model involving just the interaction
between Large and DesignB also provides some interesting results:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 345.6 345.60 17.42 0.002
LargeDesignB 1 345.6 345.60 17.42 0.002
Error 10 198.4 19.84
Total 11 544.0

Model Summary

S R-sq R-sq(adj) R-sq(pred)


4.45421 63.53% 59.88% 50.91%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 13.60 1.41 9.66 0.000
LargeDesignB 14.40 3.45 4.17 0.002 1.00

Regression Equation

Number = 13.60 + 14.40 LargeDesignB

Here we see that the interaction term is significant. Thus, one might consider that the differences are
due to both design and a large advertisement. But, it is difficult to reach any definite conclusions
given the size of the data set. A larger sample size is needed to make any stronger conclusions about
the relationships among the variables.

16 - 38
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

27. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 107.700 107.700 295.39 0.000
Period 1 107.700 107.700 295.39 0.000
Error 18 6.563 0.365
Total 19 114.263

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.603819 94.26% 93.94% 92.85%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 81.989 0.280 292.30 0.000
Period 0.4024 0.0234 17.19 0.000 1.00

Regression Equation

Price ($) = 81.989 + 0.4024 Period

Fits and Diagnostics for Unusual Observations

Obs Price ($) Fit Resid Std Resid


5 82.840 84.002 -1.162 -2.02 R

R Large residual

Durbin-Watson Statistic

Durbin-Watson Statistic = 0.798118

b. The Durbin-Watson statistic is .798118. At the .05 level of significance, dL = 1.20 and dU =1.41.
Because d < dL, there is significant positive autocorrelation.

28. From Minitab, d = 1.59617. At the .05 level of significance, dL = 1.04 and dU = 1.77. Since dL  d
dU, the test is inconclusive.

16 - 39
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

29. a. A scatter diagram follows.

9
8
7
6
5
Yield

4
3
2
1
0
0 5 10 15 20 25 30 35
Years to Maturity

A simple linear regression model does not appear to be appropriate; there appears to be a curvilinear
relationship.

b. A portion of the Minitab output follows.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 68.30 34.1506 37.19 0.000
Years to Maturity 1 29.43 29.4337 32.05 0.000
YearsSq 1 14.43 14.4271 15.71 0.000
Error 37 33.97 0.9182
Lack-of-Fit 19 19.42 1.0220 1.26 0.312
Pure Error 18 14.56 0.8088
Total 39 102.28

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.958250 66.78% 64.99% 61.09%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 1.017 0.435 2.34 0.025
Years to Maturity 0.4606 0.0814 5.66 0.000 17.91
YearsSq -0.01025 0.00259 -3.96 0.000 17.91

Regression Equation

Yield = 1.017 + 0.4606 Years to Maturity - 0.01025 YearsSq

16 - 40
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Fits and Diagnostics for Unusual Observations

Std
Obs Yield Fit Resid Resid
30 8.204 6.080 2.124 2.38 R
38 5.903 5.646 0.257 0.32 X

R Large residual
X Unusual X

c. A portion of the Minitab output follows.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 68.48 68.4760 76.98 0.000
lnYears 1 68.48 68.4760 76.98 0.000
Error 38 33.80 0.8895
Lack-of-Fit 20 19.24 0.9621 1.19 0.358
Pure Error 18 14.56 0.8088
Total 39 102.28

Model Summary

S R-sq R-sq(adj) R-sq(pred)


0.943120 66.95% 66.08% 63.26%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 0.828 0.379 2.18 0.035
lnYears 1.563 0.178 8.77 0.000 1.00

Regression Equation

Yield = 0.828 + 1.563 lnYears

Fits and Diagnostics for Unusual Observations

Obs Yield Fit Resid Std Resid


29 0.767 0.828 -0.061 -0.07 X
30 8.204 5.904 2.300 2.55 R
39 1.816 0.828 0.988 1.14 X

R Large residual
X Unusual X

With R-Sq = 66.95%, the fit provided by the logarithmic model is slightly better than the fit
provided by the second-order model with R-Sq(adj) = 64.99%. But, looking at the estimated
regression lines corresponding to both models shown on the following scatter diagrams, the
logarithmic model appears to provide a better description of the relationship between years to
maturity and yield at higher values of years to maturity. For instance, the second-order model
implies that yield is starting to move downward whereas the logarithmic model shows yields
increasing but at a much slower rate.

16 - 41
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

9
8
7
6
5
Yield
4
3
y = ‐0.0103x2 + 0.4606x + 1.017
2 R² = 0.6678
1
0
0 5 10 15 20 25 30 35
Years to Maturity

9
8
7
6
5
Yield

4
3
y = 1.5626ln(x) + 0.8279
2 R² = 0.6695
1
0
0 5 10 15 20 25 30 35
Years to Maturity

16 - 42
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

30. a.

There appears to be a curvilinear relationship between weight and price.

b. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 3161747 1580874 26.82 0.000
Weight 1 833683 833683 14.14 0.002
WeightSq 1 674815 674815 11.45 0.004
Error 16 943263 58954
Lack-of-Fit 8 569380 71172 1.52 0.283
Pure Error 8 373883 46735
Total 18 4105011

Model Summary

S R-sq R-sq(adj) R-sq(pred)


242.804 77.02% 74.15% 70.43%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 11376 2565 4.43 0.000
Weight -728 194 -3.76 0.002 287.41
WeightSq 11.97 3.54 3.38 0.004 287.41

Regression Equation

Price = 11376 - 728 Weight + 11.97 WeightSq

16 - 43
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Fits and Diagnostics for Unusual Observations

Std
Obs Price Fit Resid Resid
2 1800 1148 652 2.81 R

R Large residual

The results obtained support the conclusion that there is a curvilinear relationship between weight
and price.

c. A portion of the Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 2944410 1472205 20.30 0.000
Type_Fitness 1 1005840 1005840 13.87 0.002
Type_Comfort 1 2821029 2821029 38.89 0.000
Error 16 1160601 72538
Total 18 4105011

Model Summary

S R-sq R-sq(adj) R-sq(pred)


269.328 71.73% 68.19% 62.73%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 1283.8 95.2 13.48 0.000
Type_Fitness -572 154 -3.72 0.002 1.20
Type_Comfort -907 145 -6.24 0.000 1.20

Regression Equation

Price = 1283.8 - 572 Type_Fitness - 907 Type_Comfort

Fits and Diagnostics for Unusual Observations

Obs Price Fit Resid Std Resid


1 1800 1284 516 2.05 R
2 1800 1284 516 2.05 R
8 650 1284 -634 -2.52 R

R Large residual

Type of bike appears to be a significant factor in predicting price. But, the estimated regression
equation developed in part (b) appears to provide a slightly better fit.

d. A portion of the Minitab output follows. In this output WxF denotes the interaction between the
weight of the bike and the dummy variable Type_Fitness and WxC denotes the interaction between
the weight of the bike and the dummy variable Type_Comfort.

16 - 44
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Regression Analysis: Price versus Weight, Type_Fitness, Type_Comfort, WxF,


WxC

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 5 3450170 690034 13.70 0.000
Weight 1 454593 454593 9.02 0.010
Type_Fitness 1 300713 300713 5.97 0.030
Type_Comfort 1 415469 415469 8.25 0.013
WxF 1 274999 274999 5.46 0.036
WxC 1 404778 404778 8.04 0.014
Error 13 654841 50372
Lack-of-Fit 6 282207 47035 0.88 0.552
Pure Error 7 372633 53233
Total 18 4105011

Model Summary

S R-sq R-sq(adj) R-sq(pred)


224.438 84.05% 77.91% 71.70%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 5924 1547 3.83 0.002
Weight -214.6 71.4 -3.00 0.010 45.74
Type_Fitness -6343 2596 -2.44 0.030 492.96
Type_Comfort -7232 2518 -2.87 0.013 516.81
WxF 261 112 2.34 0.036 537.48
WxC 266.4 94.0 2.83 0.014 762.67

Regression Equation

Price = 5924 - 214.6 Weight - 6343 Type_Fitness - 7232 Type_Comfort + 261


WxF + 266.4 WxC

Fits and Diagnostics for Unusual Observations

Std
Obs Price Fit Resid Resid
2 1800 1203 597 2.87 R

R Large residual

By taking into account the type of bike, the weight, and the interaction between these two factors,
this estimated regression equation provides an excellent fit.

16 - 45
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

31. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 4 2587.7 646.9 5.42 0.002
Industry 1 1180.3 1180.3 9.89 0.003
Public 1 154.8 154.8 1.30 0.263
Quality 1 586.3 586.3 4.91 0.033
Finished 1 577.3 577.3 4.84 0.035
Error 35 4176.3 119.3
Lack-of-Fit 23 2945.6 128.1 1.25 0.353
Pure Error 12 1230.8 102.6
Total 39 6764.0

Model Summary

S R-sq R-sq(adj) R-sq(pred)


10.9235 38.26% 31.20% 18.21%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 80.43 5.92 13.60 0.000
Industry 11.94 3.80 3.15 0.003 1.06
Public -4.82 4.23 -1.14 0.263 1.05
Quality -2.62 1.18 -2.22 0.033 1.06
Finished -4.07 1.85 -2.20 0.035 1.00

Regression Equation

Delay = 80.43 + 11.94 Industry - 4.82 Public - 2.62 Quality - 4.07


Finished

b. The low value of the adjusted coefficient of determination (31.2%) does not indicate a good fit.

16 - 46
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

c. The scatter diagram follows:

100

90

80

70
Delay

60

50

40

30
0 1 2 3 4 5

Finished

The scatter diagram suggests a curvilinear relationship between these two variables.

d. The output from Minitab’s best subsets procedure follows, where FinishedSq is the square of
Finished.

Response is Delay

F
i
I F n
n Q i i
d P u n s
u u a i h
s b l s e
t l i h d
R-Sq R-Sq Mallows r i t e S
Vars R-Sq (adj) (pred) Cp S y c y d q
1 15.9 13.7 6.5 34.0 12.234 X
1 9.9 7.6 0.0 39.0 12.662 X
2 37.8 34.4 27.0 17.8 10.667 X X
2 26.9 22.9 14.9 26.9 11.561 X X
3 51.1 47.1 39.7 8.7 9.5813 X X X
3 42.2 37.3 30.4 16.1 10.425 X X X
4 59.1 54.4 48.3 4.1 8.8960 X X X X
4 51.6 46.1 35.3 10.3 9.6712 X X X X
5 59.1 53.1 43.0 6.0 9.0149 X X X X X

The estimated regression equation using Industry, Quality, Finished, and FinishedSq has an adjusted
coefficient of determination of 54.4%.

16 - 47
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

32. The computer output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 1 1076 1076.1 7.19 0.011
Industry 1 1076 1076.1 7.19 0.011
Error 38 5688 149.7
Total 39 6764

Model Summary

S R-sq R-sq(adj) R-sq(pred)


12.2344 15.91% 13.70% 6.54%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 63.00 3.39 18.57 0.000
Industry 11.07 4.13 2.68 0.011 1.00

Regression Equation

Delay = 63.00 + 11.07 Industry

Fits and Diagnostics for Unusual Observations

Obs Delay Fit Resid Std Resid


5 91.00 63.00 28.00 2.38 R
38 46.00 74.07 -28.07 -2.34 R

R Large residual

Durbin-Watson Statistic

Durbin-Watson Statistic = 1.54936

At the .05 level of significance, dL = 1.44 and dU = 1.54. Since d = 1.54936 > dU, there is no
significant positive autocorrelation.

16 - 48
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

33. a. The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 1818.6 909.3 6.80 0.003
Industry 1 1378.6 1378.6 10.31 0.003
Quality 1 742.4 742.4 5.55 0.024
Error 37 4945.4 133.7
Lack-of-Fit 7 1057.8 151.1 1.17 0.351
Pure Error 30 3887.6 129.6
Total 39 6764.0

Model Summary

S R-sq R-sq(adj) R-sq(pred)


11.5611 26.89% 22.93% 14.90%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 70.63 4.56 15.50 0.000
Industry 12.74 3.97 3.21 0.003 1.03
Quality -2.92 1.24 -2.36 0.024 1.03

Regression Equation

Delay = 70.63 + 12.74 Industry - 2.92 Quality

Fits and Diagnostics for Unusual Observations

Obs Delay Fit Resid Std Resid


5 91.00 67.71 23.29 2.13 R
38 46.00 71.70 -25.70 -2.27 R

R Large residual

Durbin-Watson Statistic

Durbin-Watson Statistic = 1.42611

16 - 49
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

b. The residual plot as a function of the order in which the data are presented follows:

30

20

Residuals 10

-10

-20

-30
4 8 12 16 20 24 28 32 36 40
Order I n Which Dat a Are Present ed

There is no obvious pattern in the data indicative of positive autocorrelation.

c. At the .05 level of significance, dL = 1.39 and dU = 1.60. Since dL ≤ d ≤ dU, the test is inconclusive.

34. The dummy variables are defined as follows:

D1 D2 Type
0 0 Non
1 0 Light
0 1 Heavy

The Minitab output follows:

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 9.333 4.667 4.00 0.034
D1 1 4.000 4.000 3.43 0.078
D2 1 9.000 9.000 7.71 0.011
Error 21 24.500 1.167
Total 23 33.833

Model Summary

S R-sq R-sq(adj) R-sq(pred)


1.08012 27.59% 20.69% 5.42%

16 - 50
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Regression Analysis: Model Building

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 4.250 0.382 11.13 0.000
D1 1.000 0.540 1.85 0.078 1.33
D2 1.500 0.540 2.78 0.011 1.33

Regression Equation

Score = 4.250 + 1.000 D1 + 1.500 D2

Since the p-value = .034 is less than  = .05, there are significant differences between comfort levels
for the three types of browsers.

35. Let SizeMidsize = 1 if a midsize car, 0 otherwise; and SizeLarge = 1 if a large car, 0 otherwise.

A portion of the Minitab output follows.

Analysis of Variance

Source DF Adj SS Adj MS F-Value P-Value


Regression 2 1486.9 743.43 26.76 0.000
SizeMidsize 1 287.1 287.14 10.34 0.001
SizeLarge 1 1482.4 1482.43 53.36 0.000
Error 306 8501.0 27.78
Total 308 9987.9

Model Summary

S R-sq R-sq(adj) R-sq(pred)


5.27077 14.89% 14.33% 13.33%

Coefficients

Term Coef SE Coef T-Value P-Value VIF


Constant 31.814 0.464 68.55 0.000
SizeMidsize -2.145 0.667 -3.21 0.001 1.18
SizeLarge -6.051 0.828 -7.30 0.000 1.18

Regression Equation

Hwy MPG = 31.814 - 2.145 SizeMidsize - 6.051 SizeLarge

16 - 51
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])


lOMoARcPSD|54931287

Chapter 16

Fits and Diagnostics for Unusual Observations

Obs Hwy MPG Fit Resid Std Resid


34 44.000 31.814 12.186 2.32 R
35 44.000 31.814 12.186 2.32 R
36 44.000 31.814 12.186 2.32 R
92 46.000 31.814 14.186 2.70 R
121 20.000 31.814 -11.814 -2.25 R
122 18.000 31.814 -13.814 -2.63 R
125 19.000 31.814 -12.814 -2.44 R
126 19.000 31.814 -12.814 -2.44 R
156 42.000 29.669 12.331 2.35 R
188 48.000 29.669 18.331 3.49 R
239 18.000 29.669 -11.669 -2.22 R
240 19.000 29.669 -10.669 -2.03 R

R Large residual

Significant overall difference; p-value corresponding to F = 26.76 is .000 <  = .05. And, the p-
values for both of the independent variables are significant as well because both p-values values are
less than  = .05.

16 - 52
© 2017 Cengage Learning. All Rights Reserved.
May not be scanned, copied or duplicated, or posted to a publicly accessible website, in whole or in part.

Downloaded by Tanzina Tamanna (tanzinatamanna28@[Link])

You might also like