0% found this document useful (0 votes)
2 views12 pages

Endterm 2021

The document outlines the instructions and structure for a Statistics II endterm exam conducted by Prof. Filipa Reis on January 12, 2022, including rules for exam conduct, allowed materials, and a dataset analysis related to individuals' fitness levels in South Korea. It provides variable descriptions, summary statistics, multiple-choice questions, and open-ended questions focused on regression models and hypothesis testing. The exam is designed to assess students' understanding of statistical concepts and their application to real-world data.

Uploaded by

Mariana Diniz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views12 pages

Endterm 2021

The document outlines the instructions and structure for a Statistics II endterm exam conducted by Prof. Filipa Reis on January 12, 2022, including rules for exam conduct, allowed materials, and a dataset analysis related to individuals' fitness levels in South Korea. It provides variable descriptions, summary statistics, multiple-choice questions, and open-ended questions focused on regression models and hypothesis testing. The exam is designed to assess students' understanding of statistical concepts and their application to real-world data.

Uploaded by

Mariana Diniz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistics II - Endterm

Prof. Filipa Reis

January 12, 2022

Instructions:

• DO NOT OPEN THE EXAM UNTIL WE ANNOUNCE THAT IT IS TIME TO OPEN IT.
• Duration: 2 hours
• This is a closed book exam. Students are allowed TWO A4 two-sided cheat-sheets.
• Mobile phones, computers, tablets, and smartwatches are not allowed.
• Graphical calculators are allowed. The exam supervisor may inspect the calculators.
• Write your name and student ID on every ODD page.
• Write the answer to each question on the designated space.
• In open answer questions, state any necessary assumptions and show all the work/calculations.
This allows us to apply partial credits where appropriate.
• You can detach this page from the rest of the exam.
• Budgeting time is a helpful practice. Consider taking a few minutes to skim through the entire exam.
• Good Luck!

1
Name: Number:

Investigating the Determinants of Individual’s Fitness Level

Consider the data set collected in South Korea by the National Sports Promotion Agency. The data was
collected in order to investigate the factors associated to individuals’ fitness level. Data collection took place
at various fitness centers distributed across the country. Below you can find the variable description and
some summary statistics.

Table 1: Variable Description

vars desc
age Age in years
gender Female (F) or male (M)
height_cm Height in cm
weight_kg Weight in kg
body_fat Percentage of body fat
grip_force Grip force in kg
sit_up_count Number of sit ups
female =1 if person is female, 0 if male
body_fat_healthy =1 if the % of body fat is within the healthy range, 0 otherwise

Table 2: Summary Statistics - FULL SAMPLE

Statistic N Mean St. Dev. Min Pctl(25) Pctl(75) Max


age 13,393 36.775 13.626 21 25 48 64
height_cm 13,393 168.560 8.427 125.000 162.400 174.800 193.800
weight_kg 13,393 67.447 11.950 26.300 58.200 75.300 138.100
body_fat 13,393 23.240 7.257 3.000 18.000 28.000 78.400
grip_force 13,393 36.964 10.625 0.000 27.500 45.200 70.500
sit_up_count 13,393 39.771 14.277 0 30 50 80
female 13,393 0.368 0.482 0 0 1 1
body_fat_healthy 13,393 0.412 0.492 0 0 1 1

Table 3: Summary Statistics - FEMALES Table 4: Summary Statistics - MALES


Statistic N Mean St. Dev. Statistic N Mean St. Dev.
age 4,926 37.851 14.418 age 8,467 36.149 13.103
height_cm 4,926 160.485 5.649 height_cm 8,467 173.257 5.810
weight_kg 4,926 56.906 7.640 weight_kg 8,467 73.580 9.469
body_fat 4,926 28.486 6.225 body_fat 8,467 20.188 5.953
grip_force 4,926 25.818 4.695 grip_force 8,467 43.448 7.170
sit_up_counts 4,926 30.888 13.889 sit_up_counts 8,467 44.939 11.729
body_fat_healthy 4,926 0.338 0.473 body_fat_healthy 8,467 0.455 0.498

2
Name: Number:

Part I - Multiple Choice [10 points; 0.625 points per question]

For the following set of questions, select the most appropriate option. There are no penalties for wrong
answers.

1. What is the unit of analysis of this dataset?


a) A fitness center.
b) A person’s fitness level.
c) A person.
d) A determinant of an individual’s fitness level.
e) A — coefficient.

2. What is the population of interest?


a) South Korean fitness centers.
b) All South Korean adults.
c) All Korean people.
d) The fitness level of South Korean people.
e) The mean, µ, fitness level of South Korean people.

3. Is the selected sample representative of the target population?


a) Likely yes, because the sample is very large.
b) Likely not, because the data collection was made at fitness centers.
c) Likely yes, because it was collected by an official agency.
d) Likely not, because South Korea has over 50 million people thus the sample is not large enough.
e) It is impossible to tell from the available information.

For questions 4 to 7, consider the following statement regarding the variable sit_up_count: “The true mean
is greater than 25 sit ups”.

4. Which of the options below best matches the statement?


a) H0 : X̄ Ø 25 v.s. Ha : X̄ < 25
b) H0 : µ Ø 25 v.s. Ha : µ < 25
c) H0 : µ Æ 25 v.s. Ha : µ > 25
d) H0 : X̄ Æ 25 v.s. Ha : X̄ > 25
e) H0 : µ > 25 v.s. Ha : µ Æ 25

5. Setting – = 1%, the rejection region for this test is:


a) ] ≠ Œ, ≠2.58] U [2.58, +Œ[
b) ] ≠ Œ, ≠2.33]
c) [2.33, +Œ[
d) ] ≠ Œ, ≠2.33] U [2.33, +Œ[
e) [2.58, +Œ[

6. The test statistic for this test is:


a) 452.4
b) 72.6
c) 95.2
d) 2.58
e) 119.74

3
Name: Number:

7. The p-value for this test is:


a) approx. 1
b) approx. 0
c) approx. 0.05
d) approx. 0.01
e) None of the above.

For question 8 to 11, consider the following statement regarding the variable body_fat_healthy: “Most
(i.e. the majority of) South Koreans have a healthy body fat percentage.”

8. Which of the options below best matches the statement?


a) H0 : µ Ø 0.5 v.s. Ha : µ < 0.5
b) H0 : p̂ Ø 0.5 v.s. Ha : p̂ < 0.5
c) H0 : p Æ 0.5 v.s. Ha : p > 0.5
d) H0 : µ Æ 0.5 v.s. Ha : µ > 0.5
e) H0 : X̄ = 0.5 v.s. Ha : X̄ ”= 0.5

9. In order to statistically test this statement, which of the following is most appropriate?
a) Given that the sample size is large, the Central Limit Theorem applies. Thus, we should assume
the population distribution of the random variable body_fat_healthy to be normal.
b) Given that the random variable body_fat_healthy is a binary variable, its population distribution
is a Bernoulli distribution. Thus, knowing the sample size is irrelevant.
c) Given that the sample size is large, the Central Limit Theorem applies. Thus, the relevant sampling
distribution will be approximately normally distributed.
d) Given that the sample size is large, the Central Limit Theorem applies. Thus, the relevant sampling
distribution will be a Bernoulli distribution.
e) None of the above.

10. A statistical analysis of the available sample. . .


a) Gives strong support to the statement (p-value < 0.01).
b) Gives weak support to the statement (p-value < 0.1).
c) Does not support the statement (p-value < 0.05).
d) Does not support the statement (p-value > 0.1).
e) The available information is not sufficient to statistically test this statement.

11. When statistically testing this statement, a type II error would occur if we. . .
a) set – to 10% instead of setting – to 5%.
b) find that the — of the test is larger than –.
c) claim that most South Koreans have a healthy body fat percentage when in fact they do not.
d) claim that most South Koreans do not have a healthy body fat percentage when in fact they do.
e) None of the above.

4
Name: Number:

For questions 12 to 14 consider an hypothesis test comparing the number of sit ups that can be performed
by men v.s. women.

12. The most appropriate statistical test is:


a) An independent sample test based on the t distribution with 13393 degrees of freedom
b) An independent sample test based on the standard normal distribution.
c) A dependent sample test based on the t distribution with 13392 degrees of freedom
d) A dependent sample test based on the standard normal distribution
e) None of the above.

13. The statistical power of this test is the likelihood that a difference between the two groups will be
detected when. . .
a) it has already been found by previous research.
b) it exists in reality.
c) it has been hypothesized to exist.
d) it exists in different samples from the same population.
e) None of the above.

14. The estimated 95% confidence interval associated to this test is (≠14.51, ≠13.59). Thus:
a) At the 5% significance level, we reject the null hypothesis of equality of means.
b) At the 5% significance level, we fail to reject the null hypothesis that the mean of women is at
least as great as the mean of men.
c) At the 5% significance level, we fail to reject the null hypothesis of equality of means.
d) At the 5% significance level, we reject the alternative hypothesis that the mean of women differs
from the mean of men.
e) None of the above.

For the following set of questions, select the most appropriate option.

15. The correlation coefficient between age and sit_up_count is r = ≠0.55 with an associated test statistic
of t = ≠75.138. We can thus conclude that:
a) We fail to reject H0 : fl = 0
b) We fail to reject H0 : r = 0
c) We reject H0 : fl = 0
d) We reject H0 : r = 0
e) None of the above.

16. An hypothesis test comparing the variance of grip force of men and the variance of grip force of women,
results in a test statistic of 2.33. This value should be compared against the critical value from which
of the following distributions?
a) F8466,4925
b) F8467,4926
c) F4925,8466
d) F4926,8467
e) ‰213392

5
Name: Number:

Part II - Open Questions [10 points]

Consider the following three regression models exploring the factors related to the number of sit ups
(y) an individual is able to perform. Models are labeled (1), (2), and (3). For each model, the table
presents the estimated coefficients of each variable, their respective standard errors (in parentheses), and
stars indicating the p-value of the respective individual significance test.

Table 5:
Dependent variable:
sit_up_count
(1) (2) (3)
age ≠0.571úúú ≠0.543úúú ≠0.459úúú
(0.008) (0.006) (0.006)
female ≠13.127úúú ≠4.001úúú
(0.183) (0.308)
height_cm ≠0.230úúú
(0.019)
weight_kg 0.263úúú
(0.013)
body_fat ≠0.944úúú
(0.018)
Constant 60.755úúú 64.554úúú 101.208úúú
(0.298) (0.258) (3.023)
Observations 13,393 13,393 13,393
R2 0.297 0.492 0.592
Adjusted R2 0.297 0.492 0.592
Residual Std. Error 11.974 10.172 9.124
F Statistic 5,645.681úúú 6,496.168úúú 3,880.068úúú
Note: ú p<0.1; úú p<0.05; úúú p<0.01

1. [1.5 points] Consider model (1). What is the confidence interval for the mean number of
sit ups performed by people who are 25 years old?

6
Name: Number:

2. [5.5 points] Consider model (2).

a) [0.75 points] Write down the population model.

b) [0.75 points] Write down the estimated model.

c) [3 points] Interpret all the estimated coefficients and respective significance levels.

7
Name: Number:

d) [1 point] Interpret the coefficient of determination.

3. [3 points] Consider now model (3).

a) [1.5 points] Find the 95% confidence interval for the coefficient of weight_kg.

b) [1.5 point] What is the predicted number of sit ups for a 28 year old woman with 165 cm, 48 kg, and
20% of body fat?

8
Name: Number:

Normal Distribution (Standardized)

9
Name: Number:

t-distribution

10
Chi-square distribution (F (x)) Name:

df 0.005 0.01 0.025 0.05 0.1 0.9 0.95 0.975 0.99 0.995
1 0.000 0.000 0.001 0.004 0.016 2.706 3.841 5.024 6.635 7.879
2 0.010 0.020 0.051 0.103 0.211 4.605 5.991 7.378 9.210 10.597
3 0.072 0.115 0.216 0.352 0.584 6.251 7.815 9.348 11.345 12.838
4 0.207 0.297 0.484 0.711 1.064 7.779 9.488 11.143 13.277 14.860
5 0.412 0.554 0.831 1.145 1.610 9.236 11.070 12.833 15.086 16.750
6 0.676 0.872 1.237 1.635 2.204 10.645 12.592 14.449 16.812 18.548
7 0.989 1.239 1.690 2.167 2.833 12.017 14.067 16.013 18.475 20.278
8 1.344 1.646 2.180 2.733 3.490 13.362 15.507 17.535 20.090 21.955
9 1.735 2.088 2.700 3.325 4.168 14.684 16.919 19.023 21.666 23.589
10 2.156 2.558 3.247 3.940 4.865 15.987 18.307 20.483 23.209 25.188
11 2.603 3.053 3.816 4.575 5.578 17.275 19.675 21.920 24.725 26.757
12 3.074 3.571 4.404 5.226 6.304 18.549 21.026 23.337 26.217 28.300
13 3.565 4.107 5.009 5.892 7.042 19.812 22.362 24.736 27.688 29.819
14 4.075 4.660 5.629 6.571 7.790 21.064 23.685 26.119 29.141 31.319
15 4.601 5.229 6.262 7.261 8.547 22.307 24.996 27.488 30.578 32.801
16 5.142 5.812 6.908 7.962 9.312 23.542 26.296 28.845 32.000 34.267
17 5.697 6.408 7.564 8.672 10.085 24.769 27.587 30.191 33.409 35.718
18 6.265 7.015 8.231 9.390 10.865 25.989 28.869 31.526 34.805 37.156
19 6.844 7.633 8.907 10.117 11.651 27.204 30.144 32.852 36.191 38.582
20 7.434 8.260 9.591 10.851 12.443 28.412 31.410 34.170 37.566 39.997

11
25 10.520 11.524 13.120 14.611 16.473 34.382 37.652 40.646 44.314 46.928
30 13.787 14.953 16.791 18.493 20.599 40.256 43.773 46.979 50.892 53.672
Number:

50 27.991 29.707 32.357 34.764 37.689 63.167 67.505 71.420 76.154 79.490
100 67.328 70.065 74.222 77.929 82.358 118.498 124.342 129.561 135.807 140.169
1000 888.564 898.912 914.257 927.594 943.133 1057.724 1074.679 1089.531 1106.969 1118.948
1225 1101.261 1112.801 1129.895 1144.737 1162.010 1288.846 1307.537 1323.893 1343.081 1356.251
1250 1124.967 1136.632 1153.910 1168.910 1186.366 1314.491 1333.364 1349.879 1369.250 1382.545
1500 1362.674 1375.529 1394.555 1411.059 1430.249 1570.608 1591.215 1609.233 1630.353 1644.838
5000 4746.175 4770.310 4805.905 4836.659 4872.281 5128.576 5165.615 5197.884 5235.572 5261.338
10000 9639.480 9673.949 9724.718 9768.525 9819.195 10181.662 10233.749 10279.070 10331.934 10368.033
Name: Number:

F-distribution

12

You might also like