Stats 506 Problem Set 2 Instructions
Stats 506 Problem Set 2 Instructions
Standard errors in the analysis of residential energy consumption data should be computed using replicate weights to account for the survey's complex design. Replicate weights help in estimating more accurate standard errors by allowing for the variation due to the survey design to be appropriately considered, thus providing a more robust estimation of variability in the point estimates .
Using survey weights in the final model of part e in Question 2 is crucial to obtaining unbiased estimates, as the NHANES data is derived from a complex survey design. To properly account for this, the svy command in Stata is used, specifically with "svyset sdmvpsu [pweight=wtmec2yr], strata(sdmvstra) vce(linearized)" to set up the survey weights so that analyses appropriately reflect the survey design .
Converting wide data to long format is important in analyses where operations like time-series analyses, repeated measures or longitudinal data analysis benefit from a structure where each observation is a single row. This conversion facilitates certain statistical techniques which require data to be in long format. Students can find resources and tutorials on how to perform this conversion on platforms like UCLA's Statistical Consulting Group and from Richard Williams’s presentation, both of which provide guidance on reshaping data using Stata .
The strategy to format and present logistic regression results for Question 2 involves creating a nicely formatted regression table that visually distinguishes the coefficients, standard errors, and significance levels. The formatting should ensure clarity, using bolding or colored text for significant variables, explanatory text describing model fitting process steps, and leveraging tools like Rmarkdown to incorporate tables directly into the PDF. Additionally, explanations of the inclusion of variables based on BIC improvement should accompany the table to contextualize the model's construction .
The key difference in executing the analysis for Questions 2 and 3 lies in the programming languages and tools used; Question 2 requires the use of Stata, while Question 3 requires the same steps to be implemented in R. This means the data reading, manipulation, and regression analyses are done using Stata commands for Question 2 and equivalent R functions for Question 3. For part d in Question 3, the 'margins' package or custom code in R must be used to compute marginal effects and adjusted predictions instead of the 'margins' command in Stata used in Question 2 .
The assignment mandates that students use commands or techniques not covered in class or course notes, encouraging them to utilize web resources to identify appropriate approaches. This requirement aims to develop students' ability to learn independently and be resourceful in problem-solving. Additionally, students are expected to attempt locating relevant information before seeking help, reinforcing self-directed learning as part of the exercise .
In integrating gender and race/ethnicity into the logistic regression model for Question 2, gender should be included if it improves the Bayesian Information Criterion (BIC). Indicators for each race/ethnicity category should be created, using the largest as the reference and combining 'Other Hispanic' and 'Other.' Each category is sequentially added to the model, retaining only those that result in improved BIC. Additionally, the poverty income ratio should also be added under the same criteria .
To properly submit the assignment for Problem Set 2 in Stats 506, Fall 2018, students must upload their assignment via Canvas by the due date, and if utilizing late days, they must indicate this with a comment on Canvas specifying the number of late days planned to use. A maximum of two late days is allowed. The submission should include a single PDF created from Rmarkdown with answers, tables, and graphs. The assignment should also be submitted as a .zip file containing the files: ps2.pdf or ps2.html, ps2.Rmd, ps2_q1.do, recs2015_usage.csv, ps2_q2.do, ps2_q2.log, ps2_q3.R, all of which should be executable without errors and located in the same working directory .
The expected output from executing the file ps2_q1.do is a comma-delimited file named recs2015_usage.csv, containing the computed national totals for residential energy consumption of various energy sources, along with their standard errors. These outputs should be read into the Rmarkdown document and used to produce a nicely formatted table with estimates and 95% confidence intervals included in the PDF report .