Multiple Choice Questions:
1. Is there evidence that the proportion of male faculty members who feel the university is
supportive of female and minority faculty is larger than the corresponding proportion of female
faculty members? To determine this, test the hypotheses H0: p1 = p2 vs Ha: p1> p2. What can we
say about the P-value of the hypothesis test?
a. The P-value is smaller than 0.001.
b. The P-value is between 0.001 and 0.01.
c. The P-value is between 0.01 and 0.05.
d. The P-value is larger than 0.05.
2. In a test of statistical hypotheses, what does the P-value tell us?
a. If the null hypothesis is true.
b. If the alternative hypothesis is true.
c. The largest level of significance at which the null hypothesis can be rejected.
d. The smallest level of significance at which the null hypothesis can be rejected.
3. In a study, fast-food menu items were analyzed for their fat content (measured in grams) and
calorie content. The goal is to predict the number of calories in a menu item from knowing its
fat content. The least-squares regression line was computed, and added to a scatter plot of
these data:
The equation of the least-squares regression line is:
Calories = 204 + 11.4 x (Fat)
The correlation between Calories and Fat is r = .979. Hence, r2 = .958.
Finally, the average number of calories in menu items is 660, and the average fat content in
menu items is 40 grams.
Which of the following is true?
a. The least-squares regression line passes through the point (40,660).
b. 95.8% of the data points fall on the least-squares regression line.
c. The data point denoted by *is highly influential.
d. All of the above.
1/6
Q.1: The following table shows the median number of hours of leisure time that Americans had each
week in various years.
Year Median Number of Leisure Hours Per Week
1973, 0 26.2
1980, 7 19.2
1987, 14 16.6
1993, 20 18.8
1997, 24 19.5
Source: Louis Harris and Associates
a) Make a scatter plot of the data, letting x represent the number of years since 1973, and
determine which model best fits the data.
b) Use a calculator to fit the linear and quadratic regression equation
c) Calculate the value of R 2 for both linear and quadratic and tell which model is appropriate and
why?
d) Use the function found in part (b) to estimate the number of leisure hours per week in 1978,
1990, and 2005.
Q.2 : Four training programs are being considered. The time required for each program is: Program A, 6
hours; Program B, 7 hours; Program C, 8 hours; and Program D, 9 hours. Twelve workers were randomly
assigned to the four programs and a week later their production output was recorded. The Minitab
output is given below.
Production 83 89 86 162 168 159 194 175 64 247 195 99
Training Hours 7 6 7 8 8 9 9 9 6 9 8 6
Regression Analysis: Production versus Training Hours
The regression equation is:
Production = - 167 + 40.5 Training Hours
Predictor Coef SE Coef T P
Constant -167.39 56.80 -2.95 0.015
Training Hours 40.540 7.322 5.54 0.000
S = 29.8921 R-Sq = 75.4% R-Sq (adj) = 72.9%
Analysis of Variance
Source DF SS MS F P
Regression 1 27392 27392 30.66 0.000
Residual Error 10 8935 894
Total 11 36327
a. What is the estimated production amount attributable to a single additional hour?
b. Find the expected production and residual values for the first trainee; i.e., 7 hours trained and
83 units produced. Interpret both of these values.
c. What is the intercept? What is its meaning?
2/6
d. Predict the production rate for 6.5 and 12 hours of training.
e. What percent of the variation in production is due to the number of training hours?
f. What is the standard deviation around the regression line?
Q.3: These data are crime-related and demographic statistics for 47 US states in 1960. The data were
collected from the FBI's Uniform Crime Report and other government agencies to determine how the
variable crime rate depends on the other variables measured in the study. There have been hypotheses
that the amount spent on the police force and unemployment rates could be indicators of crime. Scatter
plots of the two potential independent variables appear below.
Variables are:
R: Crime rate: # of offenses reported to police per million populations
Ex0: 1960 per capita expenditure on police by state and local government
U2: Unemployment rate of urban males per 1000 of age 35-39
Regression Analysis: R versus Ex0
The regression equation is R = 14.45 + 0.8948 Ex0
Predictor Coef SE Coef T P
Constant 14.45 12.67 1.14 0.260
Ex0 0.8948 0.1409 6.35 0.000
S = 28.39 R-Sq = 47.3% R-Sq(adj) = 46.1%
Analysis of Variance
Source DF SS MS F P
Regression 1 32533 32533 40.36 0.000
Residual Error 45 36276 806
Total 46 68809
Regression Analysis: R versus U2
3/6
The regression equation is R = 62.92 + 0.812 U2
Predictor Coef SE Coef T P
Constant 62.92 23.51 2.68 0.010
U2 0.8120 0.6719 1.21 0.233
S = 38.48 R-Sq = 3.1% R-Sq(adj) = 1.0%
Analysis of Variance
Source DF SS MS F P
Regression 1 2164 2164 1.46 0.233
Residual Error 45 66646 1481
Total 46 68809
a. Does there seem to be a strong linear relationship between either (Crime (R) and amount spent on
police (Ex0)) or (Crime (R) and Unemployment Rate (U2))? Explain the relationship (if any) that you
see between R and each of the independent variables. If you had to choose one of the X variables to
predict R, which one would you choose?
b. Using whichever variable you chose above, write out the regression equation (outputs can be found
above). Interpret both coefficients.
c. Is your chosen variable statistically significant or useful? Prove your answer.
d. Using only your chosen variable, utilize the regression equation to predict the Crime Rate for an area
that has an Unemployment Rate of 40 and spends $100 per capita on the police force.
e. What is the value for the standard error of estimate?
Q.4: A large mail-order house believes that there is a linear relationship between the weight of the mail
it receives and the number of orders to be filled, i.e. the weight of the mail that it receives influences the
number of orders to be filled. It would like to investigate the relationship in order to predict the number
of orders, based on the weight of the mail. From an operational perspective, knowledge of the number
of orders will help in the planning of the order-fulfillment process. A sample of 25 mail shipments is
selected that range from 200 to 700 pounds. The results of the linear regression equation are as follows:
Y = 0.1912 + 0.0297 X. R2 = 0.9731, tcal =28.84
a. Find the value of slope and y-intercept?
b. Predict the number of orders to be filled, if the weight of mail they received is 450 pounds.
Round to a whole number.
c. Predict the number of orders to be filled, if the weight of mail they received is 750 pounds.
Round to a whole number.
d. Determine the coefficient of correlation, r?
e. What is the interpretation of the coefficient of determination?
f. Set up the null and alternative hypotheses to test if there is a linear relationship between the
weight of mail in pounds and the numbers of orders to be filled. use 5% level of significance?
----------END----------
4/6