0% found this document useful (0 votes)
3 views5 pages

Semester Project

This document outlines a group project focused on applied linear regression and statistical analysis using a specific dataset. It details three main tasks: problem identification, scenario and KPI identification, and statistical examination, along with guidelines for a written report and presentation. The project emphasizes original analysis and interpretation, with strict submission instructions and grading criteria.

Uploaded by

d4f4pfgwwp
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views5 pages

Semester Project

This document outlines a group project focused on applied linear regression and statistical analysis using a specific dataset. It details three main tasks: problem identification, scenario and KPI identification, and statistical examination, along with guidelines for a written report and presentation. The project emphasizes original analysis and interpretation, with strict submission instructions and grading criteria.

Uploaded by

d4f4pfgwwp
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

GROUP PROJECT

Linear Regression & Statistical Analysis


Applied Statistics Project — Written Report & Presentation

1. Overview
This project introduces you to applied linear regression and statistical analysis using a real, messy
dataset. Each group has been assigned one dataset (see Section 4). Every group answers exactly the
same three tasks below — only the content of your dataset differs. This means your group's mark
depends entirely on the quality of your thinking and analysis, not on which dataset you were given.
The project has two deliverables:
• A written report addressing Task 1, Task 2 and Task 3 in full (100 marks).
• A 10-12 slide presentation summarising your report (graded separately — see Section 5).

2. Task 1 — Problem Identification


Total: 20 marks
1a. Study your assigned dataset (variable names, units, number of records). List all variables and
classify each as numerical (continuous/discrete) or categorical. [5 marks]
1b. Identify a real-world problem that a stakeholder in this field (e.g. a school, hospital, farm
manager, bank, city planner - depending on your dataset) would want solved using this data. State
the problem in 1-2 sentences. [5 marks]
1c. Formulate your problem as a specific, answerable research question, in the form: “To what
extent does [predictor(s)] influence/predict [outcome variable]?” [5 marks]
1d. State one hypothesis (H₁) and its null hypothesis (H₀) linked to your research question. [5
marks]

3. Task 2 — Scenario & KPI Identification


Total: 25 marks
2a. Write a short scenario (100-150 words) describing who the stakeholder is, what decision they
need to make, and why it matters to them. [8 marks]
2b. From your scenario, identify the Key Performance Indicator (KPI) - this is your dependent
variable (Y), the outcome to be predicted. State why this variable was chosen as the KPI. [7 marks]
2c. Identify at least three independent variables (predictors) from your dataset that could plausibly
explain changes in the KPI. For each one, justify why you expect it to be related, and predict
whether the relationship should be positive or negative. [10 marks]

4. Task 3 — Statistical Examination Task


Total: 55 marks
3a. Data cleaning: Check your dataset for missing values (NaN). State how many missing values
exist per variable, and explain and apply an appropriate method to handle them (e.g. mean/median
imputation, mode for categorical, or row deletion - justify your choice). [8 marks]

Page 1 of 5
3b. Descriptive statistics: Compute the mean, median, standard deviation, minimum and maximum
for your KPI and each chosen predictor. Present this as a table. [8 marks]
3c. Correlation analysis: Compute the correlation coefficient (r) between your KPI and each
predictor. Present a correlation table and comment on the strength/direction of each relationship. [8
marks]
3d. Simple linear regression: Choose the single predictor with the strongest correlation to your KPI.
Fit a simple linear regression model (Y = a + bX). State the regression equation, the R² value, and
interpret the slope (b) in context. [10 marks]
3e. Multiple linear regression: Using all your chosen predictors from Task 2c, fit a multiple linear
regression model. Report the regression equation, R², adjusted R², and the coefficient/p-value for
each predictor. State which predictors are statistically significant (p < 0.05). [12 marks]
3f. Model evaluation & interpretation: Compare the simple vs multiple regression models (which
explains more variance? which would you recommend to the stakeholder?). Using your final model,
write a 3-4 sentence recommendation back to the stakeholder from your Task 2a scenario. [9
marks]

Written Report Mark Distribution

Task Description Marks


Task 1 Problem Identification 20
Task 2 Scenario & KPI Identification 25
Task 3 Statistical Examination Task 55
TOTAL 100

Page 2 of 5
5. Presentation Guidelines
Each group presents their project as a 10-12 slide PowerPoint presentation, in 8-10 minutes,
followed by 2-3 minutes of questions. Every group follows the same slide structure below - only the
dataset content changes.

# Slide Content Maps to


1 Title Slide Group name/members, dataset title, project title —
What the dataset contains; number of records;
2 Dataset Overview Task 1a
variables listed with type (numerical/categorical)
Problem The real-world problem, stated in stakeholder
3 Task 1b
Statement terms
Research
4 Question & The formal research question, plus H0 and H1 Task 1c-1d
Hypothesis
Who the stakeholder is, what decision they face,
5 Scenario Task 2a
why it matters
The dependent variable and 3+ chosen predictors,
6 KPI & Predictors Task 2b-2c
with justification for each
Missing values found; method used to handle
7 Data Cleaning Task 3a
them
Descriptive
Summary table + correlation table; comment on
8 Statistics & Task 3b-3c
strongest relationship
Correlation
Simple Linear Equation, R2, one plot (scatter + line of best fit),
9 Task 3d
Regression plain-language interpretation
Multiple Linear Equation, R2/adjusted R2, significant predictors,
10 Task 3e-3f
Regression comparison with simple model
The 3-4 sentence recommendation back to the
11 Recommendation Task 3f
stakeholder
Limitations & 2-3 limitations of the data or model; what you
12 —
Next Steps would do with more time/data

Design Rules for the Slide Deck


• One idea per slide - no slide should carry more than approximately 40 words of running text;
move detail into the spoken explanation.
• Every regression/statistics slide needs a visual - a scatter plot with regression line, a correlation
heatmap/table, or a bar chart of coefficients. No slide of statistics should be text-only.
• Report numbers to 2 decimal places, and always attach units (e.g. “R² = 0.62”, not just “0.62”).
• Consistent labelling - your KPI must keep the same name/label across every slide it appears on.
• All group members must present at least one slide - agree this division in advance.

Page 3 of 5
Presentation Grading (separate from the written report)

Criterion Weight
Clarity of problem/scenario framing 15%
Correct and clearly explained statistics/regression 40%
Quality and accuracy of visuals 20%
Delivery, time management, ability to answer questions 15%
Adherence to the 12-slide structure 10%

Page 4 of 5
6. Submission Instructions
• Submit the written report (Tasks 1-3) as a single PDF or Word document.
• Submit the presentation as a PPT or PDF export of your slides.
• File naming: Group[Number]_[DatasetName]_Report / Group[Number]_[DatasetName]_Slides.
• All the documents should be zip into one folder and name it as GroupNumber_Session_Level
• Deadline: 3rd August 2026- 11:59pm.
• Late submissions will not be tolerated.

Academic Integrity
All analysis (descriptive statistics, correlation, and regression) must be run on your own group's
dataset. You may discuss statistical concepts with other groups, but your problem, scenario, KPI
selection, and written interpretation must be your own group's original work.
AI detection will be performed on all documents. Don’t use AI for any of the writeups, slides
and codes. You should own all the documents.
All the best

Page 5 of 5

You might also like