Course Title: Introduction to Data Science Semester VII
Course Code CIE 40
Marks
Teaching Hours/Week(L:T:P) 3:0:0 SEE 60
Marks
Total Hours of Pedagogy 45 hours Total 100
Marks
Credits 03 Exam 2.5
Hours
Examination nature(SEE) Theory
Course Learning Objectives : This Course () will enable students to,
To introduce fundamental concepts of data science
To understand data collection, preprocessing, and visualization
To learn basic statistical and analytical techniques
To provide hands-on exposure to simple tools like Python
MODULES Hours RBT Levels
MODULE: 1
Introduction to Data Science: Definition and importance of
data science, Data types and sources , Data science lifecycle ,Roles
of a data scientist
15 L1, L2,
L3, L4
MODULE: 2
Data Collection & Preprocessing
15 L2, L3, L4,
L5, L6
Data collection methods
Data cleaning (missing values, noise removal)
Data transformation and integration
Introduction to data formats (CSV, JSON)
Exploratory Data Analysis (EDA)
Descriptive statistics (mean, median, mode, variance)
Data visualization techniques
Charts: bar, pie, histogram, scatter plot
Introduction to correlation
MODULE: 3
Introduction to Machine Learning 15 L2, L3, L4,
L5, L6
Types: supervised and unsupervised learning
Basic algorithms:
o Linear regression
o Classification
o Clustering
Model evaluation basic
Tools & Applications
Introduction to Python for data science
Libraries: Pandas, NumPy, Matplotlib
Real-world applications of data science
Ethics and data privacy
Course Outcomes: At the end of the course, the student will be able to,
CO1: Understand the data in different forms
CO2: Apply different techniques to Explore Data Analysis and the Data Science Process
CO3: Analyze features election algorithms & design a recommender system.
CO4: Evaluated at a visualization tools and libraries and plot graphs.
CO5: Develop different charts and include mathematical expressions.
Suggested Learning Resources:
Textbooks:
1. Doing Data Science, CathyO’Neil and Rachel Schutt, O'Reilly Media, Inc O'Reilly
Media,Inc,2013
2. Data Visualization workshop, Tim Grobmann and Mario Dobler, Packt Publishing,
ISBN 9781800568112
Reference Book:
1. Mining of Massive Datasets, Anand Rajaraman and Jeffrey D. Ullman, Cambridge
University Press, 2010
2. DataSciencefromScratch,JoelGrus,ShroffPublisher/O’ReillyPublisherMedia A hand
book for data driven design by Andy krik
Web links and Video Lectures (e-Resources):
1. [Link]
2. [Link]