1.
Correlation and regression are concerned with
A. the relationship between two categorical variables.
B. the relationship between two quantitative variables.
C. the relationship between a quantitative explanatory variable and a categorical response variable.
D. the relationship between a categorical explanatory variable and a quantitative response variable.
2. A scatterplot is
A. one-dimensional graph of randomly scattered data.
B. two-dimensional graph of a straight line.
C. two-dimensional graph of a curved line.
D. two-dimensional graph of data values.
3. Which choice is not an appropriate description of Y ˆ in a regression equation?
A. Estimated response
B. Predicted response
C. Estimated average response
D. Observed response’
4. Which of the following can NOT be answered from a regression equation?
A. Predict the value of y at a particular value of x.
B. Estimate the slope between y and x.
C. Estimate whether the linear association is positive or negative.
D. Estimate whether the association is linear or non-linear
Explanation you need to look at the data itself to do so before the regression
5. In the simple linear regression equation, the symbol y ˆ represents the
A. average or predicted response
B. estimated intercept
C. estimated slope
D. explanatory variable
6. In the simple linear regression equation, the term b0 represents the
A. estimated or predicted response
B. estimated intercept
C. estimated slope
D. explanatory variable
7. In the simple linear regression equation, the term b1 represents the
A. estimated or predicted response
B. estimated intercept
C. estimated slope
D. explanatory variable response
8. One interpretation of the slope is
A. a student who scored 0 on the midterm would be predicted to score 50 on the final exam.
B. a student who scored 0 on the final exam would be predicted to score 50 on the midterm exam.
C. a student who scored 2 points higher than another student on the midterm would be predicted to
score 1 point higher than the other student on the final exam.
D. none of the above are an interpretation of the slope.
9. For which one of these relationships could we use a regression analysis? Only one choice is
correct.
A. Relationship between weight and height.
B. Relationship between political party membership and opinion about abortion.
C. Relationship between gender and whether person has a tattoo.
D. Relationship between eye color (blue, brown, etc.) and hair color (blond, etc.).
10. Two variables have a positive association when
A. the values of one variable tend to increase as the values of the other variable increase.
B. the values of one variable tend to decrease as the values of the other variable increase.
C. the values of one variable tend to increase regardless of how the values of the other variable change.
D. the values of both variables are always positive.
11. A scatter plot and regression line can be used for all of the following EXCEPT
A. to determine if any (x,y) pairs are outliers.
B. to predict y at a specific value of x.
C. to estimate the average y at a specific value of x.
D. to determine if a change in x causes a change in y.
12. Which of the following is a deterministic relationship?
A. The relationship between hair color and eye color.
B. The relationship between father's height and son's height.
C. The relationship between height in inches and height in centimeters.
D. The relationship between height as determined with a ruler and height as determined by a tape
measure.
13. Which one of the following is NOT appropriate for studying the relationship between two
quantitative variables?
A. Scatterplot
B. Bar chart
C. Correlation
D. Regression
14. Which graph shows a pattern that would be appropriately described by the equation y = b0+b1x
?
a. b. c. d.
15. Describe the type of association shown in the following scatterplot:
A. Positive linear association
B. Negative linear association
C. Positive curvilinear association
D. Negative curvilinear association
Mathematical and short Essay section
1. You were given the following regression computer output modeling the relationship
between sales (dependent variable ) and advertising (independent variable)
a) Interpret the intercept and slope coefficients, do you think they make sense.
b) Predict the value of Y when X=7 and interpret the results
5.653 million
2. A regression between foot length (response variable in cm) and height (explanatory
variable in inches) for 33 students resulted in the following regression equation:
ˆ y = 10.9 + 0.23 x
a) One student in the sample was 73 inches tall with a foot length of 29 cm. What is the
predicted foot length for this student?
27.69 cm
b) One student in the sample was 73 inches tall with a foot length of 29 cm. What is the
error/ residual for this student?
1.31 cm its is actual minus estimated
c) What is the estimated average foot length for students who are 70 inches tall?
27 cm
3. The scatterplot below shows student heights (y axis) versus father’s heights (x axis) for a
sample of 173 college students. The symbol “+” represents a male student and the
symbol “o” is represents a female student. Based on the scatterplot, what is the
problem with using a regression equation for all 173 students?
There are really two subgroups , that requires having separate regressions to have better
estimates
4. Below is a table of 11 student’s scores out of 100 on their Maths and
English tests. Plot a scatter graph from this data and explain it clealry
Scatter graphs are a tool that we use to display data with two variables. For
example, you might collect data how each of the people in your class performed in
their Maths and English tests. To plot a scatter graph from this data, you would firstly
draw a pair of axes with Maths grades on the xx-axis and English grades on the yy-
axis. Then, each student’s pair of grades forms a pair of coordinates that we can plot
on our Maths/English graph. The result of plotting all the students’ grades is a scatter
graph. Let’s see an example.
As stated, we put the Maths mark on the xx-axis and the English mark on the yy-
axis. It doesn’t matter which way round these things go, as long as you draw the
graph correctly.
Then, we plot each individual student’s Maths mark against their English mark in
the way that we normally plot coordinates. The resulting graph is on the right.
The aim of drawing a scatter graph is to determine if there is a link or relationship
between the two variables that have been plotted. If yes, then we say there
is correlation.
There are two types of correlation:
– Positive correlation – as one variable increases, the other one also increases.
– Negative correlation – as one variable increases, the other decreases.
Looking at this graph, we can see that, in general, as people’s grade in Maths
increase, their grades in English tend to decrease. So, there is negative
correlation between Maths and English grades.
We can also comment on how strong the correlation is. If all the points are very
closely aligned (either in negative or positive correlation), then we say that there
is strong correlation. If there is correlation but the points are quite spread out and
not clearly in line, we say that there is weak correlation. If the reality is somewhere
in between, then there is moderate correlation.
In the example above, we might say there is strong negative correlation. On the
contrary, the picture to the left is an example of weak positive correlation.
It helps a lot to have a clear ruler when doing this – it makes it a lot easier to make
sure your line is right where you want it. Then, once you’ve drawn the line, check
how many points fall on either side of the line. If the number is roughly the same for
both sides, then that’s a positive sign.
The line of best fit for the scatter graph above has been drawn in green onto the
picture on the left. Counting the points, we can see that there are 6 points underneath
the line and 5 above it, so that’s all good.
Now, what is the point of this? Well, a line of best fit is supposed to represent the
correlation of the data. In other words, the line of best fit gives us a clear outline
of the relationship between the two variables, and it gives us a tool to make
predictions about future data points. Note: a positive correlation will give a line of
best fit with positive gradient. The same goes for the negative case.
On this note, the question has asked us to predict the English mark of someone who
managed a mark of 60 in Maths. To do this, we draw a straight, vertical line (orange)
up from 60 on the Maths axis until we hit the line of best fit. Then, we draw a
horizontal line (also orange) across from that point to the English axis. It touches
that axis at 50, so 50 is the predicted English grade.