0% found this document useful (0 votes)
11 views4 pages

Titanic Survival Prediction Analysis

Uploaded by

johnthairu079
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views4 pages

Titanic Survival Prediction Analysis

Uploaded by

johnthairu079
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

P100/1625G/21

NG’ANG’A JOHN THAIRU


KARATINA UNIVERSITY
Predictive analytics in Business Intelligence
Technical Assignment II
Predictive Analytics Using Machine Learning

1. Dataset Description

1.1 Overview

The Titanic dataset contains information about passengers aboard the RMS Titanic, which sank
in 1912. The goal is to predict survival (Survived: 0 = No, 1 = Yes) based on passenger
attributes.

1.2 Features

Feature Description
Pclass Ticket class (1 = 1st, 2 = 2nd, 3 = 3rd)
Sex Gender (Male/Female)
Age Age in years
SibSp Number of siblings/spouses aboard
Parch Number of parents/children aboard
Fare Passenger fare
Embarke
Port of embarkation (C = Cherbourg, Q = Queenstown, S = Southampton)
d
Cabin Cabin number (many missing)

1.3 Data Preprocessing

 Handled missing values in Age, Embarked, and Fare.


 Engineered new features:
o FamilySize = SibSp + Parch + 1
o IsAlone (1 if no family aboard)
o Title (extracted from names, e.g., Mr, Mrs, Master)
o Has_Cabin (1 if cabin data available)

2. Exploratory Data Analysis (EDA) – Key Findings


2.1 Survival Rate by Features

Feature Survival Rate (%)


Sex
Female 74.2%
Male 18.9%
Pclass
1st Class 63.0%
2nd Class 47.3%
3rd Class 24.2%
Age
Children (<12) 59.0%
Adults (20-40) 38.1%

2.2 Visual Insights

 Women and children had significantly higher survival rates.


 1st-class passengers were more likely to survive.
 Passengers with cabins (likely higher class) had a 66.7% survival rate vs. 30% without.

3. Model Selection & Evaluation

3.1 Models Compared

Accurac
Model Precision Recall F1-Score ROC-AUC
y
Logistic Regression 79.3% 75.0% 70.6% 72.7% 84.3%
Decision Tree 81.0% 76.9% 73.2% 75.0% 80.3%
Random Forest 83.2% 80.0% 76.5% 78.2% 88.1%

3.2 Confusion Matrix (Random Forest)

Predicted Died Predicted Survived


Actual Died 92 14
Actual
19 54
Survived
 Precision (80%): When the model predicts survival, it’s correct 80% of the time.
 Recall (76.5%): The model captures 76.5% of actual survivors.

3.3 Cross-Validation Score

 Mean Accuracy (5-fold CV): 81.5% (±2.1%)


4. Feature Importance Analysis

Top 5 Features (Random Forest)

Importanc
Feature Interpretation
e
Sex_Code 52.0% Women (1) survived more than men (0)
Fare 18.0% Higher fare = higher survival
Age 8.4% Children had better survival
Pclass 4.6% 1st class > 2nd > 3rd
Title_Code 5.7% "Mrs" and "Miss" survived more

5. Final Insights & Recommendations

5.1 Key Takeaways

✔ Women, children, and 1st-class passengers had the highest survival rates.
✔ Random Forest (83.2% accuracy) performed best, followed by Decision Tree (81.0%).
✔ Sex and Fare were the strongest predictors of survival.

5.2 Recommendations for Improvement

Try XGBoost or Neural Networks for potential accuracy gains.


Engineer more features:

 Extract deck levels from Cabin (e.g., A, B, C).


 Create interaction terms (e.g., Sex × Pclass).
🔹 Address class imbalance (if any) using SMOTE.

5.3 Business Impact

 Safety protocols: Prioritize women, children, and high-paying passengers in emergencies.


 Historical analysis: Understand socio-economic biases in survival

Appendix

Code & Data Availability

 Kaggle Dataset:

Common questions

Powered by AI

'Sex' and 'Fare' were found to be the most important features in predicting survival. 'Sex' was the strongest predictor with a 52.0% feature importance, highlighting that women had a significantly higher likelihood of survival compared to men. 'Fare', with an 18.0% importance, suggested that passengers who paid higher fares were more likely to survive, indicating a correlation between wealth and survival probability .

The Random Forest model outperformed both Logistic Regression and Decision Tree models. It achieved an accuracy of 83.2%, which was higher than Logistic Regression at 79.3% and Decision Tree at 81.0%. In terms of precision, recall, F1-Score, and ROC-AUC, Random Forest showed superior general performance, with a precision of 80% and a recall of 76.5%. The model was more robust in capturing actual survivors and balancing sensitivity with specificity .

The survival analysis findings underscore significant socio-economic biases, reflecting historical inequalities. Women, children, and wealthier passengers (1st-class) enjoyed higher survival rates, indicating prioritization based on gender and class. Such biases highlight entrenched social hierarchies and economic disparities of the time, demonstrating how socio-economic status influenced crisis outcomes .

Engineering additional features such as interaction terms (e.g., Sex × Pclass) could uncover relationships between passenger characteristics that influence survival, potentially improving predictive accuracy. Extracting deck levels from cabin data could introduce a new variable correlated with passenger class and survival likelihood, offering better insights into location-based survival advantages or disadvantages, thereby enhancing model performance .

The cross-validation score indicated that the Random Forest model had a mean accuracy of 81.5% with a standard deviation of ±2.1%. This reflects the model's consistency in performance across different data subsets, suggesting it is reliably predicting survival outcomes and generalizing well to unseen data .

Key business recommendations include prioritizing women, children, and high-paying passengers during emergencies, reflecting effective crisis management strategies observed from survival likelihood patterns. Additionally, organizations can use historical data insights to understand socio-economic biases and improve equity in safety protocols, ensuring fairer treatment across different passenger classes in future scenarios .

The EDA revealed that survival rates were significantly higher for women and children. Specifically, women had a 74.2% survival rate compared to 18.9% for men. Ticket class also played a key role, with 1st-class passengers having a 63.0% survival rate, significantly higher than 47.3% for 2nd-class and 24.2% for 3rd-class passengers. These insights suggest a socio-economic bias in survival outcomes .

Feature engineering significantly improved predictive accuracy by creating new relevant attributes from existing data, such as FamilySize, IsAlone, Title, and Has_Cabin. These new features allowed the models to better capture underlying patterns related to survival. For example, FamilySize helped understand the impact of having companions on board, while Title provided social status insights. Additionally, Has_Cabin indicated access to more expensive accommodations linked to higher survival rates .

Handling missing values in Age, Embarked, and Fare was crucial to improving model reliability by reducing data inaccuracies and biases that could impact model predictions. Imputation of missing age values ensured robust age-related inference, while filling Embarked and Fare gaps helped in capturing embarkation trends and fare-based survival likelihoods. Additionally, preprocessing steps such as creating FamilySize and Has_Cabin helped in accommodating missing cabin data effectively .

Using XGBoost or Neural Networks might enhance model accuracy by leveraging their capability to handle complex non-linear relationships and interactions more effectively than Random Forests. XGBoost can provide higher precision with its boosting algorithm, especially in handling class imbalance. Neural Networks, with their layered architecture, could better capture intricate patterns in the data, potentially improving prediction outcomes through deeper learning .

You might also like