0% found this document useful (0 votes)
4 views2 pages

Categorical Data Analysis Assignment

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views2 pages

Categorical Data Analysis Assignment

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Addis Ababa University

College of Natural and Computational Sciences


Department of Statistics
Categorical Data Analysis (Stat 723)– Assignment
Please Show ALL the necessary steps and interpretations!

Due Date: June 17, 2023

Name:

Student ID Number:

This assignment contains 2 pages (including this cover page) and 5 questions. Total of points is
100.
Good luck and Happy reading work!

Distribution of Marks
Question Points Score
1 20
2 24
3 16
4 35
5 5
Total: 100

1
Categorical Data Analysis (Stat 723) Assignment

1. To collect data in an introductory statistics course, I gave the students a questionnaire. One
question asked whether the student was a vegetarian. Of 25 students, 0 answered yes. They
were not a random sample, but use these data to illustrate inference for a proportion. Let π
denote the population proportion who would say yes. Consider H0 : π = 0.50 and H1 : π ̸= 0.50.
(a) (5 points) What happens when you conduct the Wald test, which uses the estimated stan-
dard error in the z test statistic?
(b) (5 points) Find the 95% Wald confidence interval for π. Is it believable?
(c) (5 points) Conduct the score test, which uses the null standard error in the z test statistic.
Report and interpret the p-value.
(d) (5 points) Verify that the 95% score confidence interval equals (0.0, 0.133).

2. For diagnostic testing, let X = true status (1 = disease, 2 = no disease) and Y = diagnosis (1
= positive, 2 = negative). Let π1 = P (Y = 1|X = 1) and π2 = P (Y = 1|X = 2). Let γ denote
the probability that a subject has the disease.
(a) (8 points) Given that the diagnosis is positive, use Bayes’ Theorem to show that the
probability a subject truly has the disease is P (X = 1|Y = 1) = π1 γ/[π1 γ + π2 (1 − γ)].
(b) (8 points) For mammograms for detecting breast cancer, suppose γ = 0.01, sensitivity =
π1 = 0.86, and specificity 1 − π2 = 0.88. Find the positive predictive value.
(c) (8 points) To better understand the answer in (b), find the joint probabilities for the 2 × 2
cross-classification of X and Y. Discuss their relative sizes in the two cells that refer to a
positive test result.

3. An article summarized results from the Nurses’ Health Study and the Health Professionals
Follow-Up Study. The article reported (with RR=relative risk) that ”Compared with nonregular
use, regular aspirin use was associated with lower risk of overall cancer (RR 0.97; 95% CI
0.94, 0.99), which was primarily due to a lower incidence of gastrointestinal cancers, especially
colorectal cancers (RR 0.81; 95% CI 0.75, 0.88).”
(a) (8 points) Identify the response variables and the explanatory variable for these two re-
sults. Explain how to interpret the confidence interval about colorectal cancers.
(b) (8 points) Would the association with overall cancer be considered
(i) significant or nonsignificant?
(ii) strong or weak? Explain.

4. Using the data in R (birthwt) fit an appropriate model to assess factors associated with low
birth weight status. (Use all relevant variables in the data)
(a) (5 points) Write down the model
(b) (5 points) Estimated the proposed model
(c) (5 points) Is the overall model significant?(Justify your answer)
(d) (10 points) Interpret the odds ratios for ALL variables.
(e) (5 points) Are ALL variables in the model significant?(Justify your answer)
(f) (5 points) Assess goodness of fit of then estimated model.

5. (a) (5 points) Discuss how the Fisher scoring algorithm works

Page 2 of 2

You might also like