0% found this document useful (0 votes)
18 views31 pages

Data Science Internship Report at CodSoft

Report

Uploaded by

NITIN RAJ JAANU
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views31 pages

Data Science Internship Report at CodSoft

Report

Uploaded by

NITIN RAJ JAANU
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A REPORT OF PRACTICAL TRAINING - I

at

CodSoft

SUBMITTED IN PARTIAL FULFILLMENT OF THE REQUIREMENT FOR THE


AWARD OF THE DEGREE OF

Bachelor of Technology

in Computer Science & Engineering – Artificial Intelligence

Department of Engineering and Technology


Gurugram University, Gurugram, Haryana

July – Dec , 2025

Submitted By

Rohan
(Registration No.: 231030050048)
i
DEPARTMENT OF COMPUTER SCIENCE & ENGINEERING
GURUGRAM UNIVERSITY GURUGRAM

CANDIDATE'S DECLARATION

I Rohan hereby declare that I have undertaken 4 weeks Software Training at “

CodSoft ” during a period from 20th June to 20th July in partial fulfillment of

requirements for the award of degree of [Link] Computer Science and Engineering

(Artificial Intelligence) at Department of Engineering and Technology, Gurugram

University Gurugram. The work which is being presented in the training report

submitted to Department of Engineering and Technology, Gurugram University

Gurugram is an authentic record of training work.

Signature of the Student

The Software training Viva–Voce Examination of has been held


on and accepted.

Chairperson Signature of External Examiner


Department of Engineering and Technology

ii
ACKNOWLEDGEMENTS

I would like to express my sincere gratitude to my project supervisor for his


valuable guidance, continuous encouragement, and insightful suggestions
throughout the completion of this project. His expertise and support played a
crucial role in shaping the quality and direction of this work.

I am deeply thankful to the Head of the Department, Department of Engineering


& Technology, Gurugram University, for providing the necessary academic
environment, resources, and motivation that helped me accomplish this project
successfully.

My heartfelt thanks also go to all faculty members of the department for their
constant support, constructive feedback, and academic mentorship during my
entire course of study. Their dedication to teaching and research has greatly
contributed to my learning experience.

I would also like to extend my gratitude to CodSoft for providing exposure to


aerospace tools, simulation techniques, and technical training that significantly
contributed to the successful execution of this project.

iii
Abstract

This report presents the outcomes of a four-week Data Science internship


completed at CodSoft, during which four machine learning projects were
developed to demonstrate proficiency in data preprocessing, exploratory analysis,
model building, and performance evaluation.

The first project, Iris Flower Classification, involved predicting the species of iris
flowers based on sepal and petal measurements. Using classical machine learning
algorithms, the model achieved high accuracy, confirming the separability of the
three species in the dataset.

The second project, Titanic Survival Prediction, focused on predicting passenger


survival using demographic, socio-economic, and travel-related features. Through
data cleaning, feature encoding, and model fitting, the classifier produced strong
accuracy, demonstrating the usefulness of logistic regression and tree-based
models in binary classification tasks.

The third project, Credit Card Fraud Detection, presented a real-world challenge
due to extreme class imbalance. The dataset contained less than 1% fraudulent
transactions, requiring careful handling of sampling strategies and performance
metrics beyond accuracy. A Random Forest model was trained and evaluated using
precision, recall, and accuracy, successfully identifying anomalous patterns
indicative of fraudulent activity.

The fourth project, Movie Rating Prediction, addressed regression by estimating


IMDb ratings based on features such as year, duration, genre, and number of votes.
After experimenting with linear and ensemble models, the Random Forest
Regressor achieved a satisfactory R² score, illustrating the relationship between
movie attributes and audience ratings.

Overall, the internship provided significant hands-on experience across the


complete data science workflow, from dataset preparation to visualization, model
design, validation, and interpretation of results.

iv
LIST OF ABBREVIATIONS

Abbreviation Full Form


ML — Machine Learning
DS — Data Science
RF — Random Forest
LR — Logistic Regression
R² — Coefficient of Determination
TP — True Positive
TN — True Negative
FP — False Positive
FN — False Negative
DF — Data Frame
CSV — Comma Separated Values

v
LIST OF FIGURES

Chapter 4 Page

Fig 4.1 - Variations of sepal length, sepal width, petal length, petal xiii
width according to the species of iris flower

Fig 4.2 - Amount per transactions by class xi

Fig 4.3 – Distribution of percentage of movies per genre xiv

Fig 4.4 - Age distribution of passengers xvii

vi
List of Tables

Chapter 4 Pages

Table 4.1 – Iris Dataset ix

Table 4.2 – Info about Credit Card Fraud Detection Dataset xii

Table 4.3 – Movie Rating Dataset xv

Table 4.4 – Titanic Survival Dataset xviii

vii
Contents

Content Details Page No.


Certificate i
Declaration ii
Acknowledgements iii
Abstract iv
List of Abbreviations v
List of Figures vi
List of Tables vii
Contents i
Chapter 1 Introduction ii
Chapter 2 Software Training Undertaken iii-iv
Chapter 3 INDUSTRIAL TRAINING UNDERTAKEN v-vi

Chapter 4 PROJECT WORK vii-xviii

Chapter 5 RESULTS & DISCUSSION xix


Chapter 6 CONCLUSION & FUTURE SCOPE xx

i
🧩 CHAPTER 1: INTRODUCTION TO THE ORGANIZATION
1.1 About CodSoft
CodSoft is a leading skill-development and training organization that specializes in
industry-oriented educational programs. With a focus on bridging the gap between
academic concepts and industry needs, CodSoft designs internship programs that
allow students to gain exposure to real-world projects in fields such as Data Science,
Machine Learning, Full Stack Development, and Artificial Intelligence.
1.2 Vision & Mission
• Vision: To create a technologically skilled workforce ready for modern
industries.
• Mission: To provide practical, hands-on training opportunities that allow
learners to apply theoretical knowledge to real datasets and real-world
problems.
1.3 Internship Structure
The Data Science internship spanned 4 weeks, divided as:
• Week 1: Basic training on Python, Pandas, NumPy, data preprocessing.
• Week 2: Classification project — Iris Dataset.
• Week 3: Fraud Detection and Titanic Survival Projects.
• Week 4: Regression project — Movie Rating Prediction.
1.4 Relevance of Internship
In the modern data-driven world, businesses rely heavily on data science techniques
for decision-making. The internship provided exposure to the entire data science
pipeline:
• data collection
• preprocessing
• cleaning and feature engineering
• applying machine learning models
• interpreting results

ii
🧩 CHAPTER 2: SOFTWARE TRAINING UNDERTAKEN
2.1 Python Programming Fundamentals
During training, I strengthened my understanding of Python fundamentals required
for data science:
• Variables, data types
• Loops and functions
• List/dictionary operations
• File handling
• Data manipulation
2.2 Libraries and Tools Learned
2.2.1 NumPy
• Handling multi-dimensional arrays
• Mathematical operations
• Vectorization
• Statistics functions
2.2.2 Pandas
• Loading CSV/XLSX files
• Handling missing values
• Data filtering
• GroupBy operations
• Merging, joining, concatenation
2.2.3 Matplotlib & Seaborn
• Line charts, bar charts, histograms
• Heatmaps for correlation
• Pairplots for feature visualization
iii
• Customizing colors, labels, and styles
2.2.4 Scikit-Learn
• Dataset splitting
• Standardization & normalization
• Encoding categorical variables
• Model training (Regression/Classification)
• Evaluation metrics (Accuracy, Precision, Recall, R²)
2.3 Jupyter Notebook
The entire internship was executed in Jupyter Notebook, offering:
• Code + output integration
• Step-by-step workflow
• Easy visualization
• Markdown documentation

iv
🧩 CHAPTER 3: INDUSTRIAL TRAINING UNDERTAKEN
3.1 Internship Goals
The main objectives were:
• Understanding real datasets
• Applying data science techniques
• Implementing ML models from scratch
• Documenting & evaluating results
3.2 Workflow Followed
Every project followed this standardized structure:
3.2.1 Data Collection & Importing
Using Pandas to load .csv datasets.
3.2.2 Data Cleaning
• Handling null values
• Converting string categories to numbers
• Checking missing patterns
• Removing duplicates
3.2.3 Exploratory Data Analysis (EDA)
Using Seaborn & Matplotlib to:
• Visualize distributions
• Identify correlations
• Detect outliers
• Understand feature importance
3.2.4 Feature Engineering
• Label Encoding
• Standard Scaling
v
• Creating derived features
• Selecting most relevant variables
3.2.5 Model Training & Validation
Using Scikit-Learn models:
• Logistic Regression
• Random Forest
• Decision Trees
• KNN
• Linear Regression
3.2.6 Model Evaluation
• Confusion matrix
• Classification report
• Accuracy score
• R² score (for regression)

vi
🧩 CHAPTER 4: PROJECT WORK
Below are expanded project explanations (medium-length detail).

4.1 Iris Flower Classification (Classification)


4.1.1 Objective
Predict the species of an Iris flower based on:
• Sepal Length
• Sepal Width
• Petal Length
• Petal Width
4.1.2 Dataset Details
• Total samples: 150
• Classes: 3 (Setosa, Versicolor, Virginica)
• Balanced dataset
4.1.3 EDA Highlights
• Petal length strongly separates Setosa from others
• Versicolor & Virginica have overlapping distributions
• Pairplots show clear linear separability
4.1.4 Models Used
• K-Nearest Neighbors
• Decision Tree
• Logistic Regression
4.1.5 Results
• Achieved ~96% accuracy
• Confusion matrix showed near-perfect separation
vii
4.1.6 Conclusion
The Iris dataset is highly suitable for ML beginners due to distinct clusters and
balanced data.

Fig 4.1 - Variations of sepal length, sepal width, petal length, petal width
according to the species of iris flower

viii
Table 4.1 – Iris Dataset

ix
4.2 Credit Card Fraud Detection (Classification, Imbalanced Dataset)
4.2.1 Objective
Detect fraudulent transactions among a highly imbalanced dataset.
4.2.2 Dataset Characteristics
• Total transactions: 284,807
• Fraud cases: Only 492 (0.17%)
• Features: PCA-transformed V1–V28 + Amount + Time
4.2.3 Challenges
• Extreme imbalance causes misleading accuracy
• Need for sampling strategies
4.2.4 Approaches Used
• Undersampling majority class
• Oversampling fraud cases (SMOTE)
• Random Forest classifier
4.2.5 Evaluation Metrics
• Accuracy: ~93.5%
• But more important:
o Precision
o Recall
o F1-score
4.2.6 Insights
• High recall is crucial to detect fraud
• Random Forest performed best

x
Fig 4.2 - Amount per transactions by class

xi
Table 4.2 – Info about Credit Card Fraud Detection Dataset

xii
4.3 Movie Rating Prediction (Regression)
4.3.1 Objective
Predict IMDb rating of movies using:
• Year
• Duration
• Genre
• Votes
• Directors/Actors
4.3.2 EDA Insights
• Movie rating correlates with votes
• Average rating stable across decades
• Duration shows weak correlation
4.3.3 Models Used
• Linear Regression
• Random Forest Regressor
4.3.4 Metrics
• R²: 0.65 – 0.72
• Random Forest outperformed Linear Regression
4.3.5 Conclusion
Movie features moderately influence rating; advanced models may improve results.

xiii
Fig 4.3 – Distribution of percentage of movies per genre

xiv
Table 4.3 – Movie Rating Dataset

xv
4.4 Titanic Survival Prediction (Classification)
4.4.1 Objective
Predict whether a passenger survived based on:
• Age
• Gender
• Ticket class
• Fare
• Embarked location
4.4.2 Data Cleaning
• Filled missing ages using mean
• Encoded categorical variables
• Handled empty cabin values
4.4.3 Key Observations
• Females had higher survival rate
• First-class passengers survived more
• Younger passengers had better chances
4.4.4 Models Used
• Logistic Regression
• Random Forest
4.4.5 Performance
• Accuracy: ~79%
• Random Forest performed best

xvi
Fig 4.4 - Age distribution of passengers

xvii
Table 4.4 – Titanic Survival Dataset

xviii
🧩 CHAPTER 5: RESULTS & DISCUSSION
5.1 Summary of All Models

Project Model Metric Score

Iris Classification KNN Accuracy 96%

Titanic Survival Random Forest Accuracy 89%

Credit Card Fraud Detection Random Forest Precision/Recall High

Movie Rating Prediction RF Regressor R² ~0.70

5.2 Key Learnings


• Balanced datasets → high accuracy easily achievable
• Imbalanced datasets require advanced techniques
• Regression requires careful feature selection
• Visualization helps uncover data patterns
5.3 Discussion
• Classification problems performed strongly
• Fraud detection requires more recall-focused evaluation
• Regression performance indicates possible improvements with ensemble
methods

19
🧩 CHAPTER 6: CONCLUSION & FUTURE SCOPE
6.1 Conclusion
The internship provided hands-on exposure to the full data science lifecycle. Each
project strengthened my:
• Python programming
• Data preprocessing skills
• Understanding of ML algorithms
• Evaluation & interpretation abilities
6.2 Future Scope
Potential areas for enhancement:
• Deploying models using Flask/Streamlit
• Using deep learning models
• Hyperparameter optimization (GridSearchCV)
• Applying feature selection techniques
• Automating EDA and preprocessing with pipelines

20
REFERENCES

1. Jake VanderPlas, Python Data Science Handbook, O’Reilly Media,


2016.
2. UCI Machine Learning Repository — Iris Dataset:
[Link]
3. Kaggle Dataset — Credit Card Fraud Detection:
[Link]
4. Kaggle — Titanic: Machine Learning from Disaster:
[Link]
5. IMDb Official Dataset Documentation:
[Link]
6. Pedregosa et al., “Scikit-learn: Machine Learning in Python”, Journal
of Machine Learning Research, 2011.

21

You might also like