VASIREDDY VENKATADRI INSTITUTE OF TECHNOLOGY
NAMBUR-522508 , GUNTUR ,ANDHRA PRADESH, INDIA
YEAR : IV [Link] SEMESTER:I
COURSE NAME: DATA SCIENCE
COURSE CODE: OE4101
BRANCH: CSE
PREREQUISITE: STATISTICS, MATHEMATICS, DATA MINING
COURSE OBJECTIVES:
1. To gain knowledge in the basic concepts of Data Analysis
2. To acquire skills in data preparatory and preprocessing steps
3. To learn the tools and packages in Python for data science
4. To gain understanding in classification and Regression Model
5. To acquire knowledge in data interpretation and visualization techniques
COURSE OUTCOMES: Students will be able to:
Cognitive
Levels as Weightage
SN0 OUTCOME
per Bloom’s (%)
Taxonomy
Gain knowledge in the basic concepts of Data
CO1 L1, L2, L3 20
Analysis
Acquire skills in data preparatory and pre-
CO2 L1,L2,L3, L4 20
processing steps
Learn the tools and packages in Python for
CO3 L1,L2,L3, L4 20
data science
Gain understanding in classification and
CO4 L2, L3, L4 20
Regression Model
Acquire knowledge in data interpretation and
CO5 L1, L2, L4 20
visualization techniques
WEIGHTAGE OFBLOOM’S LEGENDS & PERCENTAGEOF QUESTIONS
IN EXAMINATIONS:
L1 (Remembering) = 30- 40%, L2 (Understanding) = 30 - 40%,
L3 (Applying) = 10-20 %, L4 (Analysing) = 10 - 20%,
Easy (%) = 15%-20%, Average (%)= 60% - 70%, Difficult (%)= 15% - 20%
TOTAL = L1 + L2 + L3 + L4 = 100%(on an average about 2minutes per mark)
Note: This specification weightage in above shall be treated as a general guideline
for students, teachers and paper setters. The actual distribution of marks in the
question paper may vary slightly.
DETAILED SYLLABUS:
UNIT I
Introduction: Need for data science – benefits and uses – facets of data – data
science process – setting their search goal – retrieving data – cleansing,
integrating, and transforming data – exploratory data analysis – build the models –
presenting and building applications.
UNIT II
Describing Data: Frequency distributions – Outliers – relative frequency
distributions – cumulative frequency distributions – frequency distributions for
nominal data – interpreting distributions – graphs –averages – mode – median –
mean – averages for qualitative and ranked data – describing variability – range –
variance – standard deviation – degrees of freedom – inter quartile range –
variability for qualitative and ranked data.
UNIT III
Python for Data Handling: Basics of NumPy arrays – aggregations – computations
on arrays – comparisons, masks, Boolean logic – fancy indexing – structured arrays –
Data manipulation with Pandas – data indexing and selection
– operating on data – missing data – hierarchical indexing – combining datasets –
aggregation and grouping – pivot tables.
UNIT IV
Describing Data II: Normal distributions – z scores – normal curve problems–
finding proportions – finding scores –more about z scores – correlation – scatter
plots – correlation coefficient for quantitative data –computational formula for
correlation coefficient – regression – regression line – least squares regression line
– standard error of estimate – interpretation of r2– multiple regression equations –
regression toward the mean.
UNIT V
Python for Data Visualization: Visualization with matplotlib – bor plot, line plots
– scatter plots – visualizing errors – density and contour plots – histograms,
binnings, and density –three-dimensional plotting – geographic data – data
analysis using StatsModels and seaborn – graph plotting using Plotly – interactive
data visualization using Bokeh.
Text Books:
1. David Cielen, Arno D. B. Meysman, and Mohamed Ali, “Introducing Data Sci-
ence”, Manning Publications, 2016. (First two chapters for Unit I)
2. Robert S. Witte and John S. Witte, “Statistics”, Eleventh Edition, Wiley Publica-
tions, 2017. (Chapters 1–7 for Units II and III)
3. Jake VanderPlas, “Python Data Science Handbook”, O’Reilly, 2016. (Chapters 2–
4 for Units IV and V)
Reference Books:
1. Allen B. Downey, “Think Stats: Exploratory Data Analysis in Python”, Green Tea
Press, 2014.
MICRO-SYLLABUS:
Unit Module Micro content
Need for data science Defining data science ,need of data science
Uses of data science in Commercial
Benefits and uses companies, Governmental organizations,
Nongovernmental organizations, Universities
Structured data, Unstructured data, Natural
language, Machine-generated data ,Graph-
Facets of data
based or network data, Audio, image, and
video, Streaming data
Setting the research goal, Retrieving data ,
Data preparation, Data exploration, Data
Data science process
modeling or model building, Presentation and
Automation
1 understanding the goals and context search
Setting their search goal Create a project charter
data stored within the company, data quality
Retrieving data checks ,prevent problems
Cleansing data, Correct errors as early as
Cleansing, integrating,
possible ,Combining data from different data
and transforming data sources, Transforming data
Simple graphs, combined graphs, link and
Exploratory data analysis brush, non graphical techniques
Model and variable selection ,Model execution
Build the models Model diagnostics and model comparison
Presenting and building Presenting findings and building applications
applications. on top of them
Unit Module Micro content
Frequency distributions for group data and
Frequency distributions ungrouped data, constructing Frequency
Distributions
Definition, handling outliers: Check for
Outliers Accuracy, Exclude from Summaries, Enhance
Understanding
Relative frequency Definition, Constructing Relative Frequency
distributions Distributions,
Definition, Constructing Cumulative,
Cumulative frequency
Frequency Distributions, Cumulative,
distributions Percentages, Percentile Ranks
2 Frequency distributions Ordered Qualitative Data, Relative and
for nominal data Cumulative Distributions for Qualitative Data
Interpreting Distributions
Interpreting distributions
Constructed by Others
Graphs For Quantitative Data, Histograms,
Graphs Frequency Polygon, Stem And Leaf Displays,
Typical Shapes, A Graph For Qualitative Data
mode , more Than One Mode, median, Finding
Averages median ,mean, formula for sample men,
formula for population mean
Averages for qualitative Averages for qualitative data ,Averages for
and ranked data Ranked Data
range , Shortcomings of Range , variance,
Weakness of Variance ,standard deviation,
Sum of Squares (SS) , Sum of Squares
Describing variability Formulas for Population , Sum of Squares
Formulas for Sample , Standard Deviation for
Population, Standard Deviation for Sample
,degrees of freedom, inter quartile range
Variability for qualitative Variability for Qualitative Data, Variability for
and ranked data. Ordered Qualitative and Ranked Data
Unit Module Micro content
NumPy Array Attributes ,Array Indexing, Array
Basics of NumPy arrays Slicing, Reshaping of Arrays, Array
Concatenation and Splitting
Min, Max, and Everything in Between,
Aggregations Summing the Values in an Array ,Minimum
and Maximum
Universal Functions ,Introducing Ufuncs ,
Exploring NumPy’s UFuncs ,Advanced Ufunc
Computations on arrays
Features, Introducing Broadcasting, Rules of
Broadcasting
Comparisons, masks, Comparison Operators as ufuncs ,Working
Boolean logic with Boolean Arrays ,Boolean Arrays as Masks
Fancy Indexing ,Exploring Fancy Indexing
Fancy indexing Combined Indexing ,Modifying Values with
Fancy Indexing
Structured Data: NumPy’s Structured Arrays
Structured arrays Creating Structured Arrays and Record Arrays
Introducing Pandas Objects ,The Pandas
Data manipulation with
3 Series Object ,The Pandas DataFrame Object,
Pandas
The Pandas Index Object
Data indexing and Data Indexing and Selection, Data Selection in
selection Series ,Data Selection in DataFrame
Operating on Data in Pandas, Index
Operating on data Preservation ,Index Alignment, Operations
between DataFrame and Series
Handling Missing Data ,Missing Data in
Missing data Pandas ,Operating on Null Values
Hierarchical Indexing, A Multiply Indexed
Series ,Methods of MultiIndex Creation,
Hierarchical indexing Indexing and Slicing a MultiIndex,
Rearranging Multi-Indices, Data Aggregations
on Multi-Indices
Concat and Append, Merge and Join,
Relational Algebra ,Categories of Joins,
Combining datasets Specification of the Merge Key, Specifying Set
Arithmetic for Joins, Overlapping Column
Names
Simple Aggregation in Pandas, GroupBy: Split,
Aggregation and grouping
Apply, Combine
Pivot Tables, Motivating Pivot Tables, Pivot
Pivot tables. Tables by Hand, Pivot Table Syntax
Unit Module Microcontent
The Normal Curve, Properties Of The Normal
4 Normal distributions
Curve, Different Normal Curves,
Definition, Converting To Z Scores, Standard
z scores Normal Curve, Standard Normal Table
Two Main Types Of Normal Curve Problems,
Normal curve problems Solve Problems Logically
Finding Proportions for One Score, Finding
Finding proportions Proportions between two Scores, Finding
Proportions beyond Two Scores
Finding scores Finding One Score, Finding Two Scores
z Scores for Non-normal Distributions,
More about z scores Standard Score, Transformed Standard
Scores
Three types of relationships: Positive
Correlation Relationship, Negative relationship, little or
No relationship
Definition , construction , categorizing
Positive, Negative, or Little or No Relationship
Scatter plots using scatter plot, Strong and Weak
Relationship, Perfect Relationship, Curvilinear
Relationship,
Correlation coefficient for Key Properties of r, Sign of r, Numerical Value
quantitative data of r , Interpretation of r
Computational formula Finding correlation coefficient using
for correlation coefficient computational formula
A Regression Line, Placement of Line,
Regression
Predictive Errors
Least squares regression Least Squares Regression Equation, Finding
line Values of b and a, Key Property ,Solving for Y′
Standard error of Finding the Standard Error of Estimate, Key
estimate Property
Interpretation of r 2 Finding r2, Small Values of r 2
Multiple regression Definition, finding multiple regression
equations equation, Common Features
Regression toward the
Definition, The Regression Fallacy
mean
Unit Module Microcontent
Visualization with Introduction to matplotlib, Importing
Matplotlib matplotlib ,Two Interfaces of matplotlib
Simple Line Plots, Adjusting the Plot: Line
Bar plot, line plots
Colors and Styles ,Adjusting the Plot: Axes
Limits, Labeling Plots,
Simple Scatter Plots, Scatter Plots with
Scatter plots [Link], Scatter Plots with [Link], plot
Versus scatter:
5 Visualizing Errors, Basic Errorbars,
Visualizing errors
Continuous Errors
Density and Contour Plots, Visualizing a
Density and contour plots Three-Dimensional Function
Histograms, binnings, and Histograms, Binnings, and Density ,Two-
density Dimensional Histograms and Binnings
Three-Dimensional Plotting in Matplotlib,
Three-dimensional Three-Dimensional Points and Lines ,Three-
plotting Dimensional Contour Plots ,Wireframes and
Surface Plots ,Surface Triangulations
Geographic Data with Basemap, Map
Geographic data Projections, Drawing a Map Background
,Plotting Data on Maps
Visualization with StatsModels, Visualization
Data analysis using
with Seaborn, Seaborn Versus Matplotlib,
StatsModels and seaborn
Exploring Seaborn Plots
Graph plotting using
Introduction to Plotly, visualization with Plotly
Plotly
Interactive data Introduction to Bokesh, visualization with
visualization using Bokeh. Bokeh
Code No: R20
IV B. TECH I SEMESTER REGULAR EXAMINATION MODEL PAPER
DATA SCIENCE
(CSE BRACNCH)
Time: 3 Hours Max. Marks: 7 0
Note : Answer ONE question from each unit (5 × 14 = 70 Marks)
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
UNIT-I CO BL
1. a) What is data science? Explain briefly the need for data [7M] CO1 L2
science with real time applications.
b) Explore the various steps in the data science process [7M] CO1 L2
and explain any three steps of it with suitable diagrams
and examples.
(OR)
2. a) Examine the different facets of data with the challenges [7M] CO1 L2
in their processing.
b) What is data cleaning? Outline common errors detected [7M] CO1 L1
in data cleaning process with possible solutions.
UNIT-II
3. a) What is frequency distribution? What are uses of [7M] CO2 L3
frequency distributions? Outline the different types of
frequency distributions with examples.
b) What are the outliers in the data? How to deal with [7M] CO2 L1
outliers.
(OR)
4. a) Illustrate the importance of variability. Describe [7M] CO2 L3
different measures of variability, with their uses in
describing data.
b) Write short note on the following: [7M] CO2 L1
i) Graphs for Qualitative data
ii) Averages for Qualitative data
iii) Variability for Qualitative data.
UNIT-III
5. a) What are universal functions in NumPy array? Explain [7M] CO3 L4
the different advanced features of universal functions.
b) How to handle missing data in pandas [7M] CO3 L3
(OR)
6. a) Demonstrate the use of structured arrays and record [7M] CO3 L4
arrays in NumPy
b) What is pivot table in pandas? Explain it clearly with [7M] CO3 L2
python code
UNIT-IV
7. a) What is normal curve? List out the properties of normal [7M] CO4 L1
Curve
b) Explain linear regression towards mean with example [7M] CO4 L3
(OR)
8. a) What is correlation coefficient? Outline the procedure for [7M] CO4 L1
finding correlation coefficient using computational
formula with example and corresponding python
program
b) Explain in detail about z scores for non-normal [7M] CO4 L3
Distribution
UNIT-V
9. a) Briefly explain about geographic data with basemap [7M] CO5 L1
with example programs.
b) Outline graph plotting using Plotly [7M] CO5 L2
(OR)
10. a) Showcase three- d i m e n s i o n a l drawing in matplotlib [7M] CO5 L2
with corresponding python code
b) Describe interactive data visualization using Bokeh [7M] CO5 L1
*****
THE ABOVE MODEL PAPER ATTAINMENTS OF BLOOM’S TEXONOMY AS
FOLLOWS
L1: 7*7 = 49= 35%
L2: 6*7 = 42 = 30%
L3: 5*7 = 35 = 25%
L4: 2*7 = 14 = 10%
SIGNATURES OF
COURSE COORDINATER MODULE COORDINATER HOD