0% found this document useful (0 votes)
12 views14 pages

LUMS CS334 Data Science Exam 2b

The document outlines the instructions and structure for Exam-2b of the CS334 course at Lahore University of Management Sciences, focusing on data science principles. It includes sections for multiple choice questions, true/false questions, and long questions, with specific instructions on exam conduct and materials allowed. The exam consists of 60 total marks distributed across different question types, emphasizing the importance of clarity in responses and adherence to academic integrity.

Uploaded by

ayat.jpgs
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views14 pages

LUMS CS334 Data Science Exam 2b

The document outlines the instructions and structure for Exam-2b of the CS334 course at Lahore University of Management Sciences, focusing on data science principles. It includes sections for multiple choice questions, true/false questions, and long questions, with specific instructions on exam conduct and materials allowed. The exam consists of 60 total marks distributed across different question types, emphasizing the importance of clarity in responses and adherence to academic integrity.

Uploaded by

ayat.jpgs
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lahore University of Management Sciences

CS334: Principles & Techniques of Data Science (Fall 2024)


Instructor: Dr. Ihsan Ayyub Qazi

Exam-2b

Name: _____________________________________________________________________________

Roll Number: _______________________________________________________________________

Instructions:
●​ Read all the instructions. Do not open this exam booklet until directed to do so.
●​ This exam is close-book and closed-notes.
●​ If you find a question ambiguous, please write down any assumptions you make.
●​ You can use a calculator but you cannot borrow one during the exam.
●​ No phone or computer is permitted during the exam.
●​ Please write with a pen (and not a pencil).
●​ Write your solutions in the space provided. Anything written outside the provided
space/grid for answers will NOT be graded.
●​ You have been provided with a reference sheet at the end of the booklet containing
information regarding pandas functions, regular expressions, etc
●​ You cannot use your own reference/cheat sheet in the exam
●​ For rough work, you may use the extra pages provided at the end of this booklet
●​ You have 75 minutes to complete this exam.
●​ Please try to manage your time judiciously. Good Luck!

“I certify that I have neither received nor given unpermitted aid on this examination and that I have
reported all such incidents observed by me in which unpermitted aid is given.”

_________________________
Signature
2

Exam Sections and Distribution of Marks

Section Topic Total Marks


Marks Obtained

A Multiple Choice & Short Questions 20

B True/False Questions 10

C Long Questions 30

Total 60
Marks
3

A. Multiple Choice Questions [20 pts]

Please note that you are required to select only one correct option each of the following
MCQs. No marks will be awarded in case of cutting or overwriting.

Add your answers to the following table:

1 2 3 4 5 6 7 8 9 10

Q1) Which of the following scenarios illustrates Simpson's Paradox?

A.​ A correlation between exercise frequency and health outcomes disappears after
adjusting for age groups.
B.​ A high correlation between income and job satisfaction when causation is actually
absent.
C.​ An association between two variables that gets stronger as the sample size increases.
D.​ The inability to detect a trend due to a high level of random error.
E.​ None of the above.

Q2) In a pre-post study design, what is used as the comparison group?

A.​ A separate group of individuals who did not experience the intervention.
B.​ Participants enrolled in a similar but different intervention.
C.​ The same group of participants measured before the intervention was applied.
D.​ Randomly selected individuals who were not part of the study.
E.​ None of the above.

Q3) In a regression discontinuity design (RDD), which of the following statements is true?

A.​ The sample is randomized around the cutoff value.


B.​ Program participation is voluntary and not based on a specific rule.
C.​ Eligibility for the program is determined strictly by a cutoff score.
D.​ All participants are treated the same regardless of the cutoff.
E.​ None of the above.
4

Q4) A study assesses the impact of a new smoking ban on respiratory health by comparing two
cities over time. Which assumption is essential for the difference-in-differences (DiD) estimator
to yield valid results?

A.​ Respiratory health levels were identical in both cities before the policy.
B.​ Both cities must have the same health trends before the smoking ban.
C.​ The smoking ban had the same effect on every age group.
D.​ Changes in respiratory health levels would have been identical in both cities if the
smoking ban had not been applied.
E.​ None of the above.

Q5) Which of the following statements is true about a p-value of 0.10 for a given test?

A.​ There is about a 10% chance that our test will accept the null hypotheses
B.​ The probability under the null hypothesis that the test statistic is equal to the observed
test statistic is about 0.10
C.​ The probability of the null hypothesis being true is less than or equal to 0.10
D.​ The probability that the alternative hypothesis is true is about 0.10
E.​ None of the above

Q6) In a causal diagram with variables X, Y, and Z, suppose X is a common effect of Y and Z. To
determine the causal effect of Y on Z, which action is appropriate?

A.​ Condition on X to block non-causal associations.


B.​ Condition on Y to make X and Z independent.
C.​ Do not condition on X since it is a collider.
D.​ Condition on Z to block the association between X and Y.
E.​ Condition on all nodes.

Q7) Which statement is correct for experiments requiring cluster sampling?

A.​ Increasing between cluster variance would lead to a higher power of our test.
B.​ Cluster sampling is often used in randomized experiments to encourage spillover effects
C.​ If ICC is high, it is generally better to increase the size of clusters than the number of
clusters to increase the probability of distinguishing an estimated effect from zero (if the
true effect is of a given size).
D.​ If individuals within a cluster are independent, then there is effectively no difference in
power whether we sampled clusters or individuals.
E.​ All of the above are correct.
5

Q8) In a study evaluating the effects of a new sleep medication on sleep quality, Abdullah
hypothesizes that the average sleep quality scores for individuals taking the medication
(alternative hypothesis) differ from those of a control group receiving a placebo (null
hypothesis). The null hypothesis states that the average sleep quality scores for both groups are
the same, while the alternative hypothesis suggests that the average sleep quality in the
treatment group could either be better or worse than that of the control group. The two vertical
lines indicate the critical values of Abdullah’s experiment.

Using the following table along with the graph, estimate the power of his study:

A.​ 0.917
B.​ 0.109
C.​ 0.068
D.​ 0.083
E.​ None of the above
6

Q9) A researcher examines the effect of an honors scholarship on academic performance using
an RDD based on students' test scores. Which assumption is critical for the validity of the RDD?

A.​ Students with higher scores are more likely to benefit from the scholarship.
B.​ The scholarship impact on grades is identical for all students.
C.​ Students cannot influence their test scores to qualify for the scholarship.
D.​ Test scores follow a bell curve distribution.
E.​ None of the above.

Q10) Haider is conducting an experiment to study the impact of caffeine on the cognitive
performance of university students. He recruits students from LUMS and randomly assigns
them to either a control group or a treatment group. Which of the following steps could Haider
practically take to increase the power of his experiment?

1.​ Increase the effect size to ensure there is less overlap between the null and alternative
hypotheses.
2.​ Increase the sample size by sampling students from other universities as well to make
the distribution curves narrower.
3.​ Put a greater proportion of students in the treatment group than control from the sample
to ensure a greater effective sample size.
A.​ 1 and 2
B.​ 2 only
C.​ 3 only
D.​ 1, 2, and 3
E.​ None of the above

B. True / False Questions [10 pts]

Add your answers to the following table:

1 2 3 4 5 6 7 8 9 10

Q1) Randomized Control Trials (RCTs) are widely used for finding causal effects because they
can control for unobserved confounding.

Q2) When finding the cause effect of a variable A on B, we can always find individual-level
causal effects using randomized control trials.
7

Q3) For the given causal DAG, to observe the true causal effect of X on Y, we need to condition
on Z and M both.

Q4) A violation of the Stable Unit Treatment Value Assumption (SUTVA) means that the
treatment effect on one individual depends on whether others received the treatment,
complicating causal inference.

Q5) If the value of X is dependent on the value of Y, then conditioning on Y = 1 when calculating
the probability of X being equal to 1 must have the same effect as performing the intervention
do(Y = 1).

Q6) As the sample size increases, the standard error approaches zero.

Q7) In an experiment testing the effect of a new drug, failing to reject the null hypothesis (i.e.,
concluding that the drug has no effect) when the drug actually does have a significant effect is a
Type II error.

Q8) By 95% confidence level, we mean there is a 95% probability that the true population
parameter lies with the confidence interval.

Q9) If the sample size is large enough, the sample mean of a population with a finite variance
and a finite mean will approach a normal distribution.

Q10) Roshnik is comparing the average heights of students from two universities, SSE and
SDSB. After conducting a hypothesis test, she finds that the test statistic falls in the 83rd
percentile of the null distribution, and concludes there is no significant difference between the
groups at a 20% significance level. Roshnik’s conclusion may not be valid if she mistakenly used
a two-tailed test.
8

C. Long Questions [30 pts]

Q1) [15 pts]

Part 1 A nutritionist is analyzing two types of dietary plans for weight loss in patients with
different levels of obesity. She provides Plan X, a balanced diet with controlled portions, and
Plan Y, a high-protein, low-carb diet. Additionally, the patients differ in their Exercise Level,
which also impacts weight loss outcomes. The level of obesity and exercise determine the
weight loss plan adopted. Here are the success rates of each plan after a year:

Exercise Obesity Plan X Plan Y

Mild Obesity (84%) 168 / 210 (85%) 85 / 100


High Exercise
Severe Obesity (77.1%) 108 / 140 (77.5%) 93 / 120

Mild Obesity (66.7%) 40 / 60 (70%) 140 / 200


Low Exercise
Severe Obesity (50%) 30 / 60 (60%) 60 / 100

Combined Results (73.6%) 346 / 470 (72.7%) 378 / 520

a.​ If a new patient, who exercises regularly, approaches the nutritionist without which
class their level of obesity falls under, should she recommend Plan X or Plan Y? Provide
an explanation for your answer. [2 pts]
9

b.​ Construct a causal directed acyclic graph to represent the given scenario. [2 pts]

c.​ Which variable(s) should be conditioned on to remove potential confounding, in order


to find the direct impact of the dietary plan on the weight loss outcome? [1.5 pt]

d.​ Calculate the conditional average treatment effect of each plan by conditioning on the
variable(s) specified in (c). [3 pts]
10

Part 2 Refer to the following causal DAG for this part of the question:

a.​ State all causal paths and non-causal paths between ‘U’ and ‘Y’. [1.5 pts]

b.​ List all the d-separated nodes in the graph, without any conditioning. [3 pts]

c.​ List all the d-separated nodes in the graph, when conditioned on X2. [2 pts]
11

Q2) Note: This problem has been adapted to fit on paper, meaning that sample sizes and
numbers are smaller than they should be in practice. [15 pts]

You are investigating whether background music negatively affects exam performance. To
assess this, you conducted an observational study by analyzing the exam scores of two groups
of 5 students with similar characteristics: one group studied with background music (treatment
group), and the other studied without music (control group). Each exam score is out of 100, and
the scores are as follows:

●​ Treatment Group: 12, 35, 34, 51, 64


●​ Control Group: 43, 65, 74, 55, 80

To validate your hypothesis, you conducted a bootstrap analysis. The bootstrap samples and
summary statistics are provided in the tables below:
12

a.​ State the null and alternative hypotheses for this experiment. [2 pts]

b.​ What are the sample means of the original control and treatment groups? [2 pt]

c.​ Which of the two test statistics, ‘t-statistic’, and ‘z-statistic’ is more appropriate for
testing your null hypothesis? Justify your answer. [2]
13

d.​ Using the differences between the row means of the bootstrap replicates, calculated as
‘Treatment Row Mean - Control Row Mean’, simulate the null hypothesis and list all values
that form the null distribution. [2 pts]

e.​ Using the bootstrap estimates, construct a 90% confidence interval for the test statistic. [2
pts]

f.​ At a 10% significance level, would you accept or reject the null hypothesis based on these
results? [2 pts]
14

g.​ A colleague suggests that the variance of the sample means for both groups appears too
high. She proposes increasing the size of each bootstrap sample to address this issue. Do
you agree with her approach? Provide proper reasoning for your answer. [3 pts]

You might also like