MSc Data Science Curriculum Overview
MSc Data Science Curriculum Overview
Data Science will revolutionize every industry in the near future. India has the opportunity
to be the talent provider to the world for data science. Spurring data science-based innovation and
establishing data science-ready infrastructure will be critical for preparing India's jobs and skills
markets for a data science-based future. Keeping in mind the extraordinary importance of data
science, Loyola College has decided to start the Department of Data Science and offer [Link] (Data
Science) Programme from June 2019. At present, the department has four experienced staff
members, all with Ph.D qualification in their own field of specialization.
Be it financial services, healthcare, education or even security and governance, data science can be
utilised for the benefit of citizens and the country. The global economic impact associated with the
use, development, and adoption of data science from 2022 through 2027 is expected to be
whopping $1.84 trillion to $2.95 trillion!
The main objective of this first-of-its-kind [Link]. (Data Science) Programme is to enable
the students to get a very good exposure to the promising field of data science. The PG Programme
will lay a strong theoretical foundation that will enable the students to develop their own
customized data science algorithms needed for deriving insights from very large data sets, which
are now continuously generated, thanks to IoT, Social Media and Digitisation. Apart from the
regular class room interactions, the PG Programme will involve a lot of guest lectures by industry
experts, intensive lab work and discussion of several business case studies. The students will
undergo an internship program at the end of second semester and carry out a major project in the
fourth semester.
CONTENTS
VISION
MISSION
• To kindle in young minds, the spirit of social and environmental justice with a blend of
academic excellence and empathy.
CORE VALUES
• Cura Personalis
• Pursuit of Excellence
• Moral Rectitude
• Social Equity
• Fostering solidarity
• Global Vision
• Spiritual Quotient
2
VISION AND MISSION OF THE DEPARTMENT
VISION
MISSION
To provide a learning ambience and curiosity to explore new avenues with social responsibilities.
PEOs STATEMENTS
PEO1 LEARNING ENVIRONMENT AND LIFE LONG LEARNING
TEMPERAMENT
PEO2
To think innovatively, analyse scientifically and make decisions appropriately, for
handling contemporary global concerns through the knowledge earned in the
computational sciences curriculum.
ACADEMIC EXCELLENCE AND CORE COMPETENCY
PEO3 To excel in modern computational techniques and compete in higher studies/career, for
addressing contemporary challenging problems with ease.
SKILL DEVELOPMENT AND ENTREPRENEURSHIP
PEO4 To develop analytical, logical and critical problem-solving skills for executing
professional
Work and become experts/entrepreneurs in the field of computational sciences.
ENVIRONMENT AND SUSTAINABILITY
To identify real world problems concerning environment and other issues; and apply the
PEO5
expertise in the computational sciences, to face the challenges and provide sustainable
Solutions.
3
PROFESSIONALISM AND ETHICS WITH SOCIAL RESPONSIBILITY
POs STATEMENTS
DISCIPLINARY KNOWLEDGE, INFORMATION/DIGITAL LITERACY &LIFE-
LONG LEARNING:
PO1
To acquire scholarly knowledge for life-long learning of the respective discipline of
computational sciences and demonstrate digital literacy.
CRITICAL, ANALYTICAL & SCIENTIFIC THINKING IN PROBLEM-SOLVING
STATEMENTS
PSOs
PSO1 Ability to identify analyze and design solutions for data analytics problems using
fundamental
principles of mathematics, Statistics, computing sciences, and relevant domain disciplines
PSO2 Acquire the skills in handling data analytics programming tools towards problem solving
and
Solution analysis for domain specific problems.
Understand and commit to professional ethics and cyber regulations, responsibilities, and
norms of
PSO3
professional computing practices
Understand the role of statistical approaches and apply the same to solve the real-life
problems in
PSO4
the fields of data analytics.
Ability to apply the advanced concepts of Big Data that pave the way to create a platform
to gain
PSO5
analytical skills which impact business decisions and strategies
PSO6 Apply the research-based knowledge to analyse and solve advanced problems in data
analytics.
To become a skilled Data Scientist in industry, academia, or government and software
PSO7 tools for
data storage, analysis and visualization
5
[Link]. Restructured CBCS curriculum with effective from June-2022
PJ Project Architecture
Planning (7C)
Project Data
Engineering(6C)
Project Coding,
Testing &
Implementation(7C)
5
Correlation Rubrics
High Moderate Low No Correlation
3 2 1 0
8
LOYOLA COLLEGE (AUTONOMOUS), CHENNAI
DEPARTMENT OF DATA SCIENCE
(2022 - Restructured Curriculum)
9
III PDS3601 Elective 2A: Natural Language T ME 2 4
PDS3602 Processing
Elective 2B: Reinforcement Learning
III PDS3301 MEAN Stack L ID 3 6
III PSS3401 Soft Skill - SHE Dept T SK 1 2
III PSL3402 Service Learning Dept Course T SL 1 2
III PVA3403 Value Added Course T VA 1 2
10
COURSE DESCRIPTOR
Semester I
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE LEVEL
I Introduction to Data Science – 8 CO1 K1
Evolution of Data Science – CO2 K2
Data Science Roles – Stages in CO3 K3
a Data Science Project – CO4 K4
Applications of Data Science in CO5 K5
various fields – Data Security K6
Issues.
II Data Collection Strategies – 14 CO1 K1
Data Pre-Processing Overview CO2 K2
– Data Cleaning – Data CO3 K3
Integration and Transformation CO4 K4
– Data Reduction – Data CO5 K5
Discretization. K6
11
III Descriptive Statistics – Mean, 14 CO1 K1
Standard Deviation, Skewness CO2 K2
and Kurtosis – Box Plots – CO3 K3
Pivot Table – Heat Map – CO4 K4
Correlation Statistics – CO5 K5
ANOVA. K6
IV Simple and Multiple 12 CO1 K1
Regression – Model Evaluation CO2 K2
using Visualization – Residual CO3 K3
Plot – Distribution Plot – CO4 K4
Polynomial Regression and CO5 K5
Pipelines – Measures for In- K6
sample Evaluation – Prediction
and Decision Making.
V Model Evaluation 12 CO1 K1
Generalization Error – Out-of- CO2 K2
Sample Evaluation Metrics – CO3 K3
Cross CO4 K4
Validation – Overfitting – CO5 K5
Under Fitting and Model K6
Selection – Prediction by using
Ridge Regression –Testing
Multiple Parameters by using
Grid Search.
TEXT BOOKS:
1. Jojo Moolayil, “Smarter Decisions : The Intersection of IoT and Data Science”, PACKT, 2016.
2. Cathy O’Neil and Rachel Schutt , “Doing Data Science”, O'Reilly, 2015.
3. David Dietrich, Barry Heller, Beibei Yang, “Data Science and Big data Analytics”, EMC 2013
4. Raj, Pethuru, “Handbook of Research on Cloud Infrastructures for Big Data Analytics”, IGI
Global.
SUGGESTED READINGS:
1. Jojo Moolayil, “Smarter Decisions: The Intersection of IoT and Data Science”, PACKT, 2016.
2. Joel Grus, “Data Science from Scratch” O’REILLY, 2018.
3. Rafael A. Irizarry, “Introduction to Data Science”, Chapman & Hall, 2022
4. Gupta. S.C. & Kapoor,V.K. , Fundamentals of Mathematical Statistics, Sultan Chand & Sons
Pvt. Ltd. New Delhi,2002.
12
Website:
1. [Link]
2. [Link] and Lasso
3. [Link]
4. [Link]
underfitting
13
SEMESTER I
COURSE DESCRIPTION
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE
LEVEL
I Set Theory - Number system, Sets and their CO1 K1
operations, Relations and functions - Relations 6 CO2 K2
and their types, Functions and their types. CO3 K3
Quadratic Functions – Quadratic equations- CO4 K4
Minima, maxima, vertex, and slope. Gradients- CO5 K5
Gradient descents-Learning rate-Loss function. K6
14
dependence; Linear independence – CO3 K3
rank/dimension for vector space using gaussian CO4 K4
elimination. CO5 K5
K6
III Rank and Nullity of a matrix- The null space of a 6 CO1 K1
matrix - finding nullity and a basis -System of CO2 K2
linear equations-eigen values and eigen vectors- CO3 K3
Linearmapping-Linear transformation, Kernel and CO4 K4
Images - Linear transformations, ordered bases CO5 K5
and matrices; Image and kernel of linear K6
transformations.
TEXT BOOKS:
15
SUGGESTED READINGS:
16
Course Code PDS1503
Course Title STATISTICS AND PROBABILITY
Credits 4
Hours/Week 4
Category MC
Semester I
Regulation 2022
Course Overview:
Able to analyse basic characteristics of the features.
Can perform univariate and Bivariate analysis.
Able to apply Probability concepts.
Can understand the concepts related to Distribution Functions.
Enable to identify and apply appropriate Probability Distributions.
Course Objective:
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE LEVEL
I Sampling Techniques – Data 14 CO1 K1
Classification – Tabulation – CO2 K2
Frequency and graphic CO3 K3
Representation – Measures of CO4 K4
Central Tendency – Measures of CO5 K5
Variation – Quartiles and K6
Percentiles– Moments -
Skewness and Kurtosis.
II Scatter Diagram – Karl Pearson’s 15 CO1 K1
Correlation Coefficient – Rank CO2 K2
Correlation –Correlation Coefficient CO3 K3
for Bivariate Frequency Distribution – CO4 K4
Regression Coefficients – Fitting of CO5 K5
Regression Lines. K6
17
III Random Experiment – 15 CO1 K1
Sample Space – Events – CO2 K2
Axiomatic Definition of CO3 K3
probability –Addition CO4 K4
Theorem– Multiplication CO5 K5
Theorem – Baye’s Theorem- K6
Applications.
REFERENCES:
1. Gupta,[Link],V.K.:“FundamentalsofMathematicalStatistics”,Sultan&Chand&S
ons, NewDelhi,11th Ed, 2002.
2. Hastie, Trevor, etal. “The elements of Statistical Learning”, Springer, 2009.
3. Ross, S.M., “Introduction to Probability and Statistics”, Academic Foundation, 2011.
4. Papoulis,[Link],S.U.,“Probability,RandomVariablesandStochasticProcesses”,TMH,
2010
18
Website:
[Link]
[Link]
sample-space-sample-points-and-events/
[Link]
techniques/concepts
sample-space-sample-points-and-events/
19
Course Code PDS 1504
Course Title PYTHON FOR DATA SCIENCE
Credits 04
Hours/Week 05
Category Major Core(MC)–Theory
Semester I
Regulation 2022
Course Overview
1. Understand data structures and OOP concepts in Python
2. Explore the functionalities and applications of Numpy & Pandas packages
3. Provide hands on training in Data Wrangling
4. Apply Data Aggregation and Grouping operations on real time data sets.
5. Exposure to Data Visualization techniques made available by Python
Course Objectives
1. To develop Python programming skill with data science perspective
2. To perform Data Wrangling operations of different types
[Link] effectively perform Data Aggregation and Grouping operations & Data Visualization
Prerequisites Basic programming knowledge.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE LEVEL
I Installing and using Jupyter Notebook – 12 CO1 K1,K2,K3
Creating and executing Python Programs CO2 K4,K5,K6
– Statements – Expressions – Variables CO3
– Operators – Data Types – Type CO4
Conversions – Control Flow Statements CO5
– Exception Handling
II Functions – Data Structures: Lists, 12 CO1 K1,K2,K3
Dictionaries, Tuples, Sets – File CO2 K4,K5,K6
handling – Regular Expressions – CO3
Object-Oriented Programming CO4
CO5
III Functional Programming: Lambda, 12 CO1 K1,K2,K3
Iterators, Generators, List CO2 K4,K5,K6
Comprehensions – NumPy Arrays – CO3
Pandas Series – Pandas Dataframes CO4
CO5
IV Data Wrangling with Pandas – Querying 12 CO1 K1,K2,K3
DataFrames – Merging DataFrames – CO2 K4,K5,K6
20
Applying Functions to DataFrames – CO3
Aggregations with Pandas and NumPy CO4
CO5
V Matplotlib package – [Link] 12 CO1 K1,K2,K3
package: Scatter matrices, Lag Plots, CO2 K4,K5,K6
Autocorrelation Plots, Bootstrap Plots CO3
CO4
Seaborn Package: Stripplot, Swarmplot, CO5
Heatmap, Pairplot,Regression Plot –
Formatting – Customizing
Visualizations
REFERENCES:
1. Gowrishanker and Veena, “Introduction to Python Programming”, CRC Press, 2019
2. Stefanie Molin, “Data Analysis with Pandas”, Packt, 2019
3. Joel Grus, “Data Science from scratch”, O'Reilly, 2015
5. Jake Vanderplas, “Python Data Science Handbook: Essential Tools for Working with Data
“, 2012
1. [Link]
2. [Link]
3. [Link]
21
Course Code PDS1505
Course Title PYTHON OR DATA SCIENCE LAB
Credits 04
Hours/Week 04
Category Major Core (MC) – Lab
Semester I
Regulation 2022
Course Overview
This Lab course aims to acquire skills in Python Programming concepts like data structures,
object oriented programming, data wrangling, data aggregation, grouping operations and data
visualization techniques.
Course Objectives
1. To apply OOP concepts in Python to solve a variety of problems
2. To develop solutions using the functions in Numpy and Pandas packages
3. To perform data wrangling, data aggregation and grouping operations
4. To effectively build data visualizations for different contexts
Prerequisites Basic programming knowledge.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I 12 CO1 K1,K2,K3
1. Editing and executing Programs CO2 K4,K5,K6
involving Flow Controls. CO3
2. Editing and executing Programs CO4
involving Functions. CO5
3. Program in String Manipulations
22
and Binary Files CO2 K4,K5,K6
11. Combining and Merging Data Sets CO3
CO4
CO5
V 12. Program involving Regular 12 CO1 K1,K2,K3
Expressions CO2 K4,K5,K6
13. Data Aggregation and GroupWise CO3
Operations CO4
CO5
REFERENCES:
1. Gowrishanker and Veena, “Introduction to Python Programming”, CRC Press, 2019
2. Stefanie Molin, “Data Analysis with Pandas”, Packt, 2019
3. Joel Grus, “Data Science from scratch”, O'Reilly, 2015
5. Jake Vanderplas, “ Python Data Science Handbook: Essential Tools for Working with Data”,
2012
Website:
1. [Link]
2. [Link]
3. [Link]
23
Course Code PDS1506
Course Title Machine Learning
Credits 04
Hours/Week 05
Category Major Core(MC)–Theory
Semester I
Regulation 2022
Course Overview
1. This course provides the various types of machine learning algorithms.
2. Machine Learning focuses on the development of predictive models that learn automatically
3. This course covers complex Machine Learning algorithms used for solving real world
problems.
4. It enables better decision making, predictive analysis, visualization and pattern discovery.
Course Objectives
1. To understand a range of Machine learning algorithms along with their merits and demerits.
2. To learn the methodology and apply the machine learning algorithms to real world problems.
3. To implement visualization of solutions for effective understanding and decision making.
4. To explore the concepts of market basket analysis and recommendation systems.
Prerequisites Basic knowledge in data science algorithms
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I Introduction: 15 CO1 K1,K2,K3
Machine Learning Foundations – CO2 K4,K5,K6
Overview – Design of a Learning System CO3
– Types of Machine Learning – CO4
Supervised Learning and Unsupervised CO5
Learning – Applications of Machine
Learning – Tools Overview for ML.
II Supervised Learning – I: 15 CO1 K1,K2,K3
Simple Linear Regression – Multiple CO2 K4,K5,K6
Linear Regression – Polynomial CO3
Regression – Ridge Regression – Lasso CO4
Regression – Evaluating Regression CO5
Models – Model Selection – Bagging –
Ensemble Methods.
III Supervised Learning – II: 15 CO1 K1,K2,K3
Classification – Logistic Regression – CO2 K4,K5,K6
24
Decision Tree Regression and CO3
Classification – Random Forest CO4
Regression and Classification – Support CO5
Vector Machine Regression and
Classification - Evaluating Classification
Models.
IV Unsupervised Learning: 15 CO1 K1,K2,K3
Clustering – K-Means Clustering – CO2 K4,K5,K6
Density-Based Clustering – CO3
Dimensionality Reduction – Collaborative CO4
Filtering. CO5
V Association Rule Learning and 15 CO1 K1,K2,K3
Reinforcement Learning: CO2 K4,K5,K6
Association Rule Learning – Apriori – CO3
Eclat – Reinforcement Learning – Upper CO4
Confidence Bound – Thompson CO5
Sampling – Q-Learning.
Text Books
1. Kevin P. Murphy, “Machine Learning: A Probabilistic Perspective”, MIT Press, 2012.
2. Ethem Alpaydin, “Introduction to Machine Learning”, MIT Press, Third Edition, 2014.
3. Tom Mitchell, "Machine Learning", McGraw-Hill, 1997.
4. Sebastian Raschka, Vahid Mirjilili,” Python Machine Learning and deep learning”, 2nd
edition, kindle book, 2018
5. Carol Quadros,” Machine Learning with python, scikit-learn and Tensorflow”, Packet
Publishing, 2018
6. Gavin Hackeling,” Machine Learning with scikit-learn”, Packet publishing, O’Reily,
2018
7. Stanford Lectures of Prof. Andrew Ng on Machine Learning
8. Christopher Bishop, “Pattern Recognition and Machine Learning” Springer, 2007.
Suggested Readings
1. Samir Madhavan, 2016. Mastering Python for Data Science, PACKT Publishing.
2. Ethem Alpaydin, 2009. Introduction to Machine Learning, The MIT Press.
3. Jake VanderPlas, 2016. Python Data Science Handbook, O’REILLY.
4. Stanford Lectures of Prof. Andrew Ng.
5. NPTEL Lectures of Prof. [Link]
Web Resources
1. [Link]
2. [Link]
python-video
3. [Link]
25
CourseOutcomes (COs)and Cognitive LevelMapping
26
Course Code PDS 1507
Course Title Machine Learning Lab
Credits 04
Hours/Week 04
Category Major Core(MC)–Lab
Semester I
Regulation 2022
Course Overview
1. This course helps to understand a wide variety of machine learning algorithms
2. It helps to understand how to evaluate models generated from data
3. Machine learning techniques enable us to automatically extract features from data so as to
solve predictive tasks, such as speech recognition, object recognition, machine translation
4 . This course helps to design and implement various machine learning algorithms in a range of
real-world application.
Course Objectives
1. To be able to formulate machine learning problems agreeing to different applications
2. To understand a variety of machine learning algorithms along with their strengths and
weaknesses
3. To be able to apply machine learning algorithms to solve problems of moderate complexity
4. To apply the algorithms to a real-world problem, enhance the models learned and report on
the expected accuracy that can be achieved by applying the model.
Prerequisites Basic Knowledge in Programming
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I 1. Simple and Multiple Linear 12 CO1 K1,K2,K3
Regression CO2 K4,K5,K6
CO3
CO4
2. Polynomial Regression CO5
27
CO4
7. Random Forest Classification CO5
28
CourseOutcomes (COs)and Cognitive LevelMapping
29
Semester II
COURSE DESCRIPTION
Course Code PDS2501
Course Title STATISTICAL INFERENCE
Credits 3
Hours/Week 4
Category MC
Semester II
Regulation 2022
Course Overview:
1. Able to understand and apply basic concepts of Statistical Inference.
2. Able to understand and apply important results such as NP Lemma and LR test.
3. Enable to easily derive conclusions from Large Samples.
4. Can understand the concepts related to Small Sample tests.
5. Enable to identify problematic situations and apply Non-parametric Tests.
Course Objectives:
1. To study basic concepts of Statistical Inference.
2. To apply important results such as NP Lemma and LR test.
3. To study Large Sample Tests and Small Sample tests.
4. To identify problematic situations and apply Non-parametric Tests.
Pre requisites: Basic understanding of Statistics
SYLLABUS
U CONTENT HRS COs COG
N NITI
I VE
T LEVE
L
I Testing of Hypothesis - Statistical Hypothesis - Simple and 14 CO1 K1
composite hypothesis, Null and Alternative hypothesis - two CO2 K2
kinds of errors, level of significance, size and power of a CO3 K3
test most powerful test, Neyman-Pearson lemma with proof. CO4 K4
CO5 K5
K6
30
II Simple examples using Neyman Pearson lemma. 15 CO1 K1
Uniformly most powerful tests andun biased tests based on CO2 K2
normal Likelihoodratiotest(withoutproof)and its CO3 K3
properties. Application of LR test for single mean. CO4 K4
CO5 K5
K6
III Testofsignificanceformean(s),variance(s),proportion 15 CO1 K1
(s),correlationcoefficient(s)basedon Normal CO2 K2
distribution. CO3 K3
CO4 K4
CO5 K5
K6
IV Test of significance for mean(s), variance(s), correlation 15 CO1 K1
coefficient(s), regression coefficient, based on t, Chi-square CO2 K2
and F-distributions. Applications of Chi-square in test of CO3 K3
significance(independence of attributes, goodness off it).- CO4 K4
ANOVA-One way Classification-Two way Classification- CO5 K5
CRD-RBD-LSD. K6
REFERENCES:
1. Gupta,[Link],V.K.:“FundamentalsofMathematicalStatistics”,Sultan&Chand&Sons,
NewDelhi,11th Ed, 2002.
2. Rohatgi,V.K.:“Statistical Inference”,JohnWileyand sons,1984.
3. Hogg,R.V,[Link]:“Introductiontomathematicalstatistics”,PrenticeHall,England
, 1995.
4. [Link].S.N.:“ModernMathematicalstatistics”,JohnWileyandsons,1988.
31
1. Website:
2. [Link]://[Link]/tutorial/python-for-statistical-analysis/hypotesis-introduction/
3. [Link]
4. LEIm3
5. [Link]
32
Course Code PDS 2502
Course Title BIG DATA ANALYTICS THROUGH SPARK
Credits 03
Hours/Week 04
Category Major Core(MC)–Theory
Semester II
Regulation 2022
Course Overview
1. Understand the Big Data Platform and its Use cases
2. Provide Concepts and Interfacing with HDFS and Map Reduce
3. Provide hands on Spark programming and Eco System
4. Apply spark analytics on Structured, Unstructured Data.
5. Exposure to Data Analytics with Machine Learning Algorithm using Spark
Course Objectives
1. To develop dynamic RDD spark programming using Different dataset.
2. To perform Big Data analytics using Spark.
[Link] effectively build Model using Machine Learning Algorithms to analysis in Big data
Prerequisites Basic programming knowledge.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I Unit I: Introduction to Big Data and 12 CO1 K1,K2,K3
Hadoop CO2 K4,K5,K6
Big Data and its importance – Sources of CO3
Big Data – Characteristics of Big Data – CO4
Big Data Analytics – Big Data CO5
Applications, Hadoop Distributed File
System – Map Reduce Paradigm- Hadoop
Ecosystem
II Unit II: Spark Programming with 12 CO1 K1,K2,K3
Python CO2 K4,K5,K6
Apache Spark Ecosystem - Resilient CO3
Distributed Datasets – Spark Architecture CO4
-Loading and Storing Data – CO5
Transformations – Actions – Key-Value
Resilient Distributed Datasets – Local
Variables – Broadcast Variables –
Accumulators – Partitioning – Persistence.
.
33
III Unit III: Spark SQL 12 CO1 K1,K2,K3
Overview of Spark SQL – Spark Session CO2 K4,K5,K6
– Data Frames – Schema of a Data Frame CO3
– Operations supported by Data Frames – CO4
Filter, Join, GroupBy, Agg operations – CO5
Nesting the Operations – Temporary
Tables – Viewing and Querying
Temporary Tables.
IV Unit IV: Spark Streaming 12 CO1 K1,K2,K3
Use Cases for Real time Analytics – CO2 K4,K5,K6
Transferring, Summarizing,Analysing CO3
Real time data – Data Sources supported CO4
by Spark Streaming – Flat files, TCP/IP – CO5
Flume – Kafka – Kinesis – Streaming
Context –D DStreams operations.
V Unit V: Machine Learning with Spark 12 CO1 K1,K2,K3
Linear Regression – Decision Tree CO2 K4,K5,K6
Classification – Principal Component CO3
Analysis – Random Forest Classification CO4
– Text Pre-processing with TF-IDF – CO5
Naïve Bayes Classification – K-Means
Clustering – Recommendation Engines.
Text Books
1. Michael Berthold, David J. Hand, “Intelligent Data Analysis”, Springer, 2007. 2. Tom
2. White “ Hadoop: The Definitive Guide” Third Edition, O‟reilly Media, 2011
3. Tomasz Drabos, “Learning PySpark”, PACKT, 2017.
Suggested Readings
1. Padma Priya Chitturi, “Apache Spark for Data Science”, PACKT, 2017.
2. Holden Karau, “ Learning Spark”. PACKT, 2016.
3. Sandy Riza, “Advanced Analytics with Spark”, O’ Reilly, 2016.
4. Romeo Kienzler, “Mastering Apache Spark”, PACKT, 2017.
Web Resources
1. [Link]
2. [Link]
34
CourseOutcomes (COs)and Cognitive LevelMapping
35
Course Code PDS2503
Course Title BIG DATA ANALYTICS THROUGH SPARK - LAB
Credits 03
Hours/Week 04
Category Major Core (MC) – Lab
Semester II
Regulation 2022
Course Overview
This Lab course aims to acquire skills in Big Data Analytics Through Spark concepts
like creating RDD, various RDD operations, Spark with SQL, Spark Streaming, and
spark with Machine Learning
Course Objectives
1. To apply RDD concepts to solve the real-world problems
2. To develop dynamic RDD spark programming using Different dataset.
3. To perform analysis in Big data using various Methods
4. To effectively build Model using Machine learning Algorithms to analysis in big data
Prerequisites Basic programming knowledge.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I 1. Program involving Resilient 12 CO1 K1,K2,K3
Distributed Datasets CO2 K4,K5,K6
2. Program involving Transformations CO3
and Actions CO4
3. Program involving Key-Value CO5
Resilient Distributed Datasets
II 4. Program involving Local Variables, 12 CO1 K1,K2,K3
Broadcast Variables and CO2 K4,K5,K6
Accumulators CO3
5. Program involving Filter, Join, CO4
GroupBy, Agg operations CO5
6. Viewing and Querying Temporary
Tables
III 7. Transferring, Summarizing and 12 CO1 K1,K2,K3
Analysing Twitter data CO2 K4,K5,K6
8. Program involving Flume, Kafka CO3
and Kinesis CO4
9. Program involving DStreams and CO5
Dstream RDDs
36
IV 10. Linear Regression 12 CO1 K1,K2,K3
11. Decision Tree Classification CO2 K4,K5,K6
12. Principal Component Analysis CO3
CO4
CO5
V 13. Random Forest Classification 12 CO1 K1,K2,K3
14. Text Pre-processing with TF-IDF CO2 K4,K5,K6
15. Naïve Bayes Classification CO3
16. K-Means Clustering CO4
CO5
Text Books
1. Michael Berthold, David J. Hand, “Intelligent Data Analysis”, Springer, 2007. 2. Tom
2. White “ Hadoop: The Definitive Guide” Third Edition, O‟reilly Media, 2011
3. Tomasz Drabos, “Learning PySpark”, PACKT, 2017.
Suggested Readings
1. Padma Priya Chitturi, “Apache Spark for Data Science”, PACKT, 2017.
2. Holden Karau, “ Learning Spark”. PACKT, 2016.
3. Sandy Riza, “Advanced Analytics with Spark”, O’ Reilly, 2016.
4. Romeo Kienzler, “Mastering Apache Spark”, PACKT, 2017.
Web Resources
1. [Link]
2. [Link]
PDS 2504 BIG DATA ANALYTICS THROUGH SPARK (MC) COGNITIVE LEVEL
37
Course Code PDS 2504
Course Title NoSQL DATABASES
Credits 03
Hours/Week 04
Category Major Core (MC) – Theory
Semester II
Regulation 2022
Course Overview
NoSQL database course introduction, overview NoSQL databases (non-relational
databases). The four types of NoSQL databases (e.g. Document-oriented, Key-Value Pair,
Column-oriented and Graph) will be explored in detail
Course Objectives
1. Knowledge on SQL query language.
2. Knowledge on MongoDB query language.
3. Ability to comprehend the principles of NoSQL.
4. Understand the difference of NoSQL key value and Document database
[Link] the Column database and data modeling technique
Prerequisites: Basic Big Data knowledge.
SYLLABUS
UNI CONTENT HOURS COs COGNITIVE
T LEVEL
I Introduction of Relational Data Base 12 CO1 K1,K2,K3
K4,K5,K6
Creating a table Inserting, deleting, alter, Updating
Select command , Where clause, Aggregate
functions, Numeric functions, Constraints, keys,
Group By, Having, Sub Queries, Alias, Joins,
Operators, String Functions, Normalization.
38
Cassandra, HBASE, Neo4j use and deployment,
Application, RDBMS approach, Challenges
NoSQL approach, Key-Value and Document Data
Models, Column-Family Stores, Aggregate-
Oriented Databases. sharding, MapReduce on
databases. Distribution Models, Single Server,
Sharding, Master-Slave Replication, Peer-to-Peer
replication, Combining Sharding and Replication .
III KEY VALUE DATA STORES 12 CO2 K1,K2,K3
CO3 K4,K5,K6
NoSQL Key/Value databases using MongoDB, CO4
Document Databases, Document oriented CO5
Database Features, Consistency, Transactions,
Availability, Query Features, Scaling, Suitable
Use Cases, Event Logging, Content Management
Systems, Blogging Platforms, Web Analytics or
Real-Time Analytics, E-Commerce Applications,
Complex Transactions Spanning Different
Operations, Queries against Varying Aggregate
Structure.
39
Text Books
1. Sadalage, P. & Fowler,NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot
Persistence, Wiley Publications,1st Edition,2022.
Suggested Readings
1. Christopher [Link], Prabhakar Raghavan, Hinrich Schutze, An introduction to
Information Retrieval, Cambridge University Press
2. Daniel Abadi, Peter Boncz and Stavros Harizopoulas, The Design and Implementation of
Modern Column-Oriented Database Systems, Now Publishers.
3. Guy Harrison, Next Generation Database: NoSQL and big data, Apress.
Web Resources
1. [Link]
2. [Link]
3. [Link]
4. [Link]
CO2 To apply objects, load data, query data and performance K1,K2,K3
tune Column-oriented NoSQL databases.
CO3 To illustrate NoSQL database operations. K3,k4
40
Course Code PDS2505
Course Title NoSQL DATABASES – LAB
Credits 03
Hours/Week 04
Category Major Core (MC) – Lab
Semester II
Regulation 2022
Course Overview
1. Knowledge on MongoDB query language.
2. Ability to comprehend the principles of NoSQL.
3. Understand the difference of NoSQL key value database and Document database
4. Know the concept of Column database
5. Understand the data modelling technique
Course Objectives
1. Demonstrate competency in designing NoSQL database management systems.
[Link] competency in describing how NoSQL databases differ from relational
databases from a theoretical perspective
3. Demonstrate competency in selecting a particular NoSQL database for specific use cases.
Prerequisites Basic SQL and Big Data knowledge.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I 1. Query to create and drop database. 12 CO1 K1,K2,K3
[Link] to create, display and drop CO2 K4,K5,K6
collection CO3
[Link] to insert, query, update and CO4
delete a document CO5
II [Link]-value databases 12 CO1 K1,K2,K3
[Link] with column-family stores CO2 K4,K5,K6
(cassandra) CO3
[Link] databases (neo4j)
CO4
CO5
III [Link] function 12 CO1 K1,K2,K3
8 .Push and addtoset expression. CO2 K4,K5,K6
9. First and last expression. CO3
CO4
CO5
IV [Link] of existing database 12 CO1 K1,K2,K3
41
[Link] of existing database CO2 K4,K5,K6
CO3
CO4
CO5
V 12. Restore database from the backup 12 CO1 K1,K2,K3
[Link] python with mongodb and CO2 K4,K5,K6
inserting, retrieving, updating and CO3
deleting. CO4
CO5
Text Books
1. Practical mongodb by “shakuntala gupta edward navin sabharwal publisherapress
2. Nosql distilled by pramod sadalge, martin fowler
3. Nosql for dummies by a willy Brand
Suggested Readings
4. Christopher [Link], Prabhakar Raghavan, Hinrich Schutze, An introduction to Information
Retrieval, Cambridge University Press
5. Daniel Abadi, Peter Boncz and Stavros Harizopoulas, The Design and Implementation of
Modern Column-Oriented Database Systems, Now Publishers.
6. Guy Harrison, Next Generation Database: NoSQL and big data, Apress.
Web Resources
1. [Link]
2. [Link]
3. [Link]
4. [Link]
CO2 To apply objects, load data, query data and performance tune K3
Column-oriented NoSQL databases in Data Set.
42
Course Code PDS2601
Course Title MARKETING ANALYTICS
Credits 2
Hours/Week 4
Category ME
Semester II
Regulation 2022
Course Overview:
1. Analyse the various types of marketing data
2. Assess the quality of marketing data and make appropriate interpretations of meaning
according to data sources and intended uses.
3. Compare and contrast common data models used in marketing data systems.
4. Able to identify social media platforms for forming Marketing strategies
5. Identify Web resources for forming effective Marketing strategies
COURSE OBJECTIVES:
1. Recognize challenges in dealing with data sets in marketing.
2. Identify and apply appropriate algorithms for analyzing the social media and web data
3. Make choices for a model for new machine learning tasks.
SYLLABUS
UNIT CONTENT HRS Cos COGNITIVE
LEVEL
I Marketing Analytics Basics 8 CO1 K1
Introduction, Data for Marketing Analytics, CO2 K2
Business Intelligence, Analytics, and Data CO3 K3
Science, Exploratory Data Analysis, CO4 K4
Descriptive Analysis, Predictive Analytics, CO5 K5
Prescriptive Analytics. Price Analytics – K6
Goals, Bunding, Skimming, Promotions,
Discounting
II Customer Analytics 14 CO1 K1
Segmentation- Introduction, Benefit of CO2 K2
Customer Analytics, Factors Essential, CO3 K3
Segmentation Analytics, Cluster Analysis. CO4 K4
Nurturing Customers - Metrics for Tracking CO5 K5
Customer Experience, Logistic Regression K6
Analysis, Use of Logistic Regression as a
43
Classification Technique. Customer
Analytics -Customer Lifetime Value, Churn
Analytics
III Digital Marketing Analytics 14 CO1 K1
Traditional Vs Digital Marketing, Strategies, CO2 K2
Advertising: Concept of Display Advertising, CO3 K3
Display Ads, Buying Models – Cost per CO4 K4
Click, Milli, Lead, Acquisition, Fixed Cost. CO5 K5
Social Media Marketing: How to build a K6
successful business strategy.
TEXT BOOKS
44
REFERENCE BOOKS
1. Marketing Analytics: A practical guide to real marketing science, Mike Grigsby, Kogen
Page, ISBN 9780749474171
2. Marketing Metrices 3e, Bendle, Farris, Pferfery, Reibstein,
3. Cutting Edge Marketing Analytics: Real World Cases and Data Sets for Hands on
Learning, Raj Kumar Venkatesan, Paul Farris, Ronald T. Wilcox.
45
Course Code PDS2602
Course Title HEALTH ANALYTICS
Credits 2
Hours/Week 4
Category ME
Semester II
Regulation 2022
Course Overview:
Analyse the various types and sources of healthcare data, including clinical, operational,
claims, and patient generated data.
Assess the quality of healthcare data and make appropriate interpretations of meaning
according to data sources and intended uses.
Compare and contrast common data models used in healthcare data systems.
Able to identify common measures used in healthcare data analysis for predictive models.
Identify approaches for precision medicine and treatments for personalised services.
Course Objective:
To understand the basic sources of healthcare data.
To perform image analysis and sensor data analysis.
To derive and evaluate data mining and analysis from social media.
To frame advanced data analytic models through visual analytics.
To identify fraud detection in healthcare from different sources of data.
Pre requisites : Basic knowledge in Statistics and Data analysis
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE
LEVEL
I Introduction to Healthcare Data Analytics- 8 CO1 K1
Electronic Health Records– Components of CO2 K2
EHR- Coding Systems- Benefits of EHR- CO3 K3
Barrier to Adopting HER Challenges- CO4 K4
Phenotyping Algorithms. CO5 K5
K6
II Biomedical Image Analysis- Mining of Sensor 14 CO1 K1
Data in Healthcare- Biomedical Signal Analysis- CO2 K2
Genomic Data Analysis for Personalized CO3 K3
Medicine. CO4 K4
CO5 K5
K6
46
III Natural Language Processing and Data Mining 14 CO1 K1
for Clinical Text- Mining the Biomedical Social CO2 K2
Media Analytics for Healthcare. CO3 K3
CO4 K4
CO5 K5
K6
TEXT BOOKS:
1. Chandan K. Reddy and Charu C Aggarwal, “Healthcare data analytics”, Taylor &
Francis, 2015.
2. Ross [Link] and Edward [Link], “Healthcare Analytics”, T&F/Routledge, 2020.
3. Chandan [Link], “Healthcare Data Analytics”, CRC Press, 2020.
4. Vikas Kumar, “Healthcare Analytics made simple”, Packt, 2020.
SUGGESTED READINGS:
1. Hui Yang and Eva K. Lee, “Healthcare Analytics: From Data to Knowledge to Healthcare
Improvement, Wiley, 2016.
2. Tim O’reilly , “How data science is transforming Healthcare”, O’reilly,2022.
3. Laura B. Madsen, “Data driven healthcare”, Wiley,2022.
4. Jason Burke, “Health Analytics”, Wiley, 2020.
47
Course Outcomes (COs) and Cognitive Level Mapping
48
Course Code PDS2506
Course Title RESEARCH METHODOLOGY`
Credits 03
Hours/Week 03
Category Major Core (MC) – Theory
Semester II
Regulation 2022
Course Overview
This methodology achieving competence and proficiency in the theory of and practice to
research. This fundamental objective can be realized through helping these students to develop
the subject of their research, encourage the formation of higher level of trained intellectual
ability, critical analysis, rigour, and independence of thought, foster individual judgment, and
skill in the application of research theory and methods, and develop skills required in writing
research proposals, reports, and dissertation
Course Objectives
These methodologies include, but are not limited to, experimental, survey and content analysis.
Class discussions and instructor lectures Examination and will be able to describe basic
approaches to qualitative research. and identify and critique articles based on different research
methods and to the know the data visualization
Prerequisites No prerequisites
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I Motivation and objectives – Research methods vs. 9 CO1 K1,K2,K3
Methodology. Types of research – Descriptive vs. CO2 K4,K5,K6
Analytical, Applied vs. Fundamental, Quantitative CO3
vs. Qualitative, Conceptual vs. Empirical, concept CO4
of applied and basic research process, criteria of
CO5
good research.
II Defining and formulating the research problem, 9 CO1 K1,K2,K3
selecting the problem, necessity of defining the CO2 K4,K5,K6
problem, importance of literature review in CO3
defining a problem, literature review-primary and CO4
secondary sources, reviews, monograph, patents,
CO5
research databases, web as a source, searching the
web, critical literature review, identifying gap
areas from literature and research database.
49
III Interpretation of Findings, Technique of 9 CO1 K1,K2,K3
Interpretation, Precaution in Interpretation, CO2 K4,K5,K6
Significance of Report Writing, Different Steps in CO3
Writing Report, Layout of the Research Report, CO4
Research Papers; Writing Research Papers, Thesis,
CO5
Reports and Project Proposals; Formatting,
Appendices, Citation Formats and Style; General
Conventions, Issues, Plagiarism and Copyright.
IV Development of working hypotheses, Types of 9 CO1 K1,K2,K3
Errors, Level of Significance, Critical Region , CO2 K4,K5,K6
Power of a Test, Tests of Significance for Large CO3
Samples, Tests of Significance Small Samples, CO4
Confidence Intervals
CO5
Text Books
1. Garg, B.L., Karadia, R., Agarwal, F. and Agarwal, U.K., 2002. An introduction to Research
Methodology, RBSA Publishers.
2. Wadehra, B.L. 2000. Law relating to patents, trademarks, copyright designs and geographical
Indications. Universal Law Publishing
3. Research Methodology: a step-by-step guide for beginners, Kumar, Pearson Education.
4 Practical Research Methods, Dawson, C., UBSPD Pvt. Ltd.
Suggested Readings
1. Anthony, M., Graziano, A.M. and Raulin, M.L., 2009. Research Methods: A Process of
Inquiry, Allyn and Bacon.
2. Carlos, C.M., 2000. Intellectual property rights, the WTO and developing countries: the
TRIPS agreement and policy options. Zed Books, New York
Web Resources
[Link]://[Link]/
50
CourseOutcomes (COs)and Cognitive LevelMapping
51
Course Code PDS3501
Course Title MULTIVARIATE TECHNIQUES FOR DATA
ANALYTICS
Credits 4
Hours/Week 4
Category MC
Semester III
Regulation 2022
Course Overview:
1. To understand the relationships between the variables.
2. Descriptive statistics helps to understand the characteristics of the features involved in the
data.
3. Course enables one to group features in a data set.
4. Provides knowledge to form clusters of the observation in big data.
5. Learn techniques for dimension reduction and feature selection.
Course Objective:
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE
LEVEL
I Measurement Scales( Metric and Non-metric 8 CO1 K1
Measurement Scales) – Classification of CO2 K2
Multivariate Techniques( Dependence and Inter- CO3 K3
dependence Techniques) – Applications of CO4 K4
Multivariate Techniques in different disciplines. CO5 K5
K6
II Introduction to Factor Analysis – Meaning, 14 CO1 K1
Objectives and Assumptions – Designing a Factor CO2 K2
Analysis Study – Deriving Factors – Assessing CO3 K3
52
Overall Factors – Validation of Factor Analysis. CO4 K4
CO5 K5
K6
III Introduction to Cluster Analysis – Objectives and 14 CO1 K1
Assumptions – Research Design in Cluster CO2 K2
Analysis – Hierarchical and Non-hierarchical CO3 K3
Methods – Interpretation of Clusters – Validation CO4 K4
of Profiling of Clusters. CO5 K5
K6
IV Introduction to Discriminant Analysis – Concepts, 12 CO1 K1
Objectives and Applications – Procedure for CO2 K2
conducting Discriminant Analysis – Stepwise CO3 K3
Discriminant Analysis – Mahalanobis Procedure – CO4 K4
Logit Model. CO5 K5
K6
V Dimensionality Reduction – Deriving Orthogonal 12 CO1 K1
Projections – Lower Dimensional Subspaces – CO2 K2
Characterization through Singular Value CO3 K3
Decomposition and Eigenvalue Analysis – CO4 K4
Rayleigh Quotient – Kernel PCA – Functional CO5 K5
PCA. K6
TEXT BOOKS:
1. Joseph F Hair, William C Black etal , “Multivariate Data Analysis” , Pearson Education,
7th edition, 2013.
2. T. W. Anderson , “An Introduction to Multivariate Statistical Analysis, 3rd Edition”,
Wiley, 2003.
3. William r Dillon, John Wiley & sons, “Multivariate Analysis methods and applications”,
Wiley, 1984.
4. Naresh K Malhotra, Satyabhusan Dash, “Marketing Research An Applied Orientation”,
Pearson, 2011.
SUGGESTED READINGS:
53
Course Outcomes (COs) and Cognitive Level Mapping
54
Course Code PDS 3502
Course Title Deep Learning
Credits 04
Hours/Week 04
Category Major Core(MC)–Theory
Semester III
Regulation 2022
Course Overview
1. This course represents the computational challenges of building stable representations for
high-dimensional data, such as images, text and data.
2. Deep Learning covers the concept of various neural networks such as CNN and RNN.
3 . It helps to understand the concept of Boltzmann machine and computer vision
4. This course covers the fundamentals of deep learning, and the main research activities in this
field.
Course Objectives
1. Understand the context of neural networks and deep learning
2. Understand the data needs of deep learning
3. Have a working knowledge of neural networks and deep learning
4. Explore the parameters for neural networks
Prerequisites Basic Knowledge in linear algebra, and probability theory.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I Artificial Neural Networks: 12 CO1 K1,K2,K3
The Neuron – Activation Function – CO2 K4,K5,K6
Gradient Descent – Stochastic Gradient CO3
Descent – Back Propagation – Business CO4
Problem. CO5
II Convolutional Neural Networks: 12 CO1 K1,K2,K3
Convolution Operation – ReLU layer – CO2 K4,K5,K6
Pooling – Flattening – Full Conversion CO3
Layer – Softmax and Cross-Entropy. CO4
CO5
III Recurrent Neural Networks: 12 CO1 K1,K2,K3
RNN intuition – Tackling Vanishing CO2 K4,K5,K6
Gradient Problem – Long Short-Term CO3
Memory – Building a RNN – CO4
Evaluating the RNN – Improving the CO5
RNN – Tuning the RNN.
55
IV Boltzmann Machines: 12 CO1 K1,K2,K3
Components of Boltzmann Machine – CO2 K4,K5,K6
Search and Learning Problem – CO3
Applications of Boltzmann Machine – CO4
Restricted Boltzmann Machine – Deep CO5
Belief Networks – Deep Boltzmann
Machine
V Computer Vision: 12 CO1 K1,K2,K3
Viola-Jones Algorithm – Haar-like CO2 K4,K5,K6
Features – Integral Image – Training CO3
Classifiers – Adaptive Boosting – CO4
Cascading – Face Detection with Open CO5
CV.
Text Books
1. Francois Challot, “ Deep learning with Python”, Manning, 2017.
2. Deep Learning Illustrated: A Visual, Interactive Guide to Artificial Intelligence,By Jon
Krohn, Grant Beyleveld and Aglaé Bassens, September 2022.
3. Ian Goodfellow, “Deep Learning”, MIT Press, 2017.
Suggested Readings
1. Josh Patterson, “Deep Learning: A Practitioner’s Approach”, PACKT, 2017.
2. Dipayan Dev, “ Deep Learning with Hadoop”, PACKT, 2017.
3. Hugo Larochelle’s Video Lectures on Deep Learning
Web Resources
1. [Link]
2. [Link]
3. [Link]
56
Course Code PDS 3503
Course Title Deep Learning - Lab
Credits 03
Hours/Week 04
Category Major Core(MC)–Lab
Semester III
Regulation 2022
Course Overview
1. This course helps to install deep learning libraries such as tensorflow, keras and pytorch.
2. It covers the details of neural networks, CNN, RNN and its applications.
3. This course covers the various operations on images such as Segmentation,
Transformations, etc…
4. It helps to develop deep learning application using neural networks.
Course Objectives
1. To develop the neural networks for handling Sequence and Image data
2. To learn hyper-parameter tuning in neural networks, which are crucial to make deep
learning systems.
3. To perform various operations on image such as gradients, contours, etc.
4. To understand the neural network-based models with a wide range of exciting applications
Prerequisites Basic Knowledge in Python Programming & Machine Learning Techniques
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I 1. Setting up the Spyder IDE 12 CO1 K1,K2,K3
Environment and Executing a Python CO2 K4,K5,K6
Program CO3
2. Installing Keras, Tensorflow and CO4
Pytorch libraries and making use of CO5
them
3. Artificial Neural Networks
II 4. Convolutional Neural Networks with 12 CO1 K1,K2,K3
Images and Text data CO2 K4,K5,K6
CO3
5. Image Transformations CO4
CO5
III 6. Image Gradients and Edge Detection 12 CO1 K1,K2,K3
57
CO2 K4,K5,K6
7. Image Contours CO3
CO4
CO5
IV 8. Image Segmentation 12 CO1 K1,K2,K3
CO2 K4,K5,K6
9. Harris Corner Detection CO3
CO4
CO5
V 10. Face Detection using Haar 12 CO1 K1,K2,K3
Cascades CO2 K4,K5,K6
CO3
11. Chatbot Creation CO4
CO5
Text Books
1. Ian Goodfellow, Yoshua Bengio, Aaron Courville, Deep Learning, MIT Press, 2016,
ISBN: 0387848576
2. Yegnanarayana, B., Artificial Neural Networks PHI Learning Pvt. Ltd, 2009. 3. Golub,
G.,H., and Van Loan,C.,F., Matrix Computations, JHU Press,2013.
3. Satish Kumar, Neural Networks: A Classroom Approach, Tata McGraw-Hill Education,
2004.
Suggested Readings
1. Trevor Hastie, Robert Tibshirani, Jerome Friedman, The Elements of Statistical Learning:
Data Mining, Inference, and Prediction, Second Edition, Springer, 2009, ISBN:
0387848576
2. Christopher M. Bishop, Pattern Recognition and Machine Learning, Springer, 2006,
ISBN: 0387310738
3. Sebastian Raschka, Python Machine Learning, Packt Publishing, 2015, ISBN:
1783555130
Web Resources
1. CS231n: Convolutional Neural Networks for Visual Recognition, Stanford
2. CS224d: Deep Learning for Natural Language Processing, Stanford
3. CS285: Deep Reinforcement Learning, Berkeley
4. MIT 6.S094: Deep Learning for Self-Driving Cars, MIT
58
CourseOutcomes (COs)and Cognitive LevelMapping
59
Course Code PDS 3504
Course Title Cloud Computing
Credits 04
Hours/Week 04
Category Major Core(MC)–Theory
Semester III
Regulation 2022
Course Overview
1. This course helps to understand the concepts and techniques in cloud computing.
2. It provides in-depth knowledge on cloud computing, types of cloud services and models.
3. This course imparts knowledge on the concepts of serverless architecture and DevOps.
4. It also explains the various cloud applications and data analytics as a service in cloud.
Course Objectives
1. To identify the basic elements of cloud architecture.
2. To familiarize the different services and models in cloud with examples.
3. To learn the concept of serverless architecture and DevOps.
4. To understand the cloud data centers and cloud security.
Prerequisites Basic knowledge in Computer Science and Internet.
SYLLABUS
UNIT CONTENT HOURS COs COGNITIVE
LEVEL
I Unit – I: Introduction 12 CO1 K1,K2,K3
Overview of Cloud Computing –Essential CO2 K4,K5,K6
Characteristics of cloud computing -Cloud computing CO3
architecture, Cloud Reference Model (NIST CO4
Architecture) – Operational models such as private,
dedicated, virtual private, community, hybrid and CO5
public cloud – Service models such as IaaS, PaaS and
SaaS – Example cloud vendors – Google cloud
platform, Amazon AWS, Microsoft Azure and Open
Stack.
II Unit – II: Platform Engineering 12 CO1 K1,K2,K3
Cloud Native Design and Microservices– CO2 K4,K5,K6
Containerized - Dynamically orchestrated design – CO3
Continuous delivery - Support for a variety of client CO4
devices – Monolithic vs Microservices Architecture -
Characteristics of microservice architecture – 12 factor CO5
application design - Service discovery – Service
Registry.
60
III Unit – III: Serverless Architecture and DevOps 12 CO1 K1,K2,K3
Function as a Service (FaaS) - Backend as a Service CO2 K4,K5,K6
(BaaS) - Advantages of serverless architectures – AWS CO3
Lamda – AWS Fargate; Introduction to DevOps - The CO4
DevOps toolchain – DevOps Practices -Continuous
Integration (CI), Continuous Delivery (CD), CO5
Continuous Deployment – Quality Attributes for
DevOps – DevOps cloud models.
IV Unit- IV Cloud Data Centers & Cloud Security 12 CO1 K1,K2,K3
Historical Perspective, Data center Components, CO2 K4,K5,K6
Design Considerations, Power Calculations, Evolution CO3
of Data Centers, Cloud data storage – CloudTM. CO4
Security Considerations – CIA Triad – STRIDE Threat
Model - Cloud specific Cryptographic Techniques – CO5
Security by Design
V Unit V Data Analytics as a Service & Cloud 12 CO1 K1,K2,K3
applications CO2 K4,K5,K6
Hadoop as a service, Map Reduce on Cloud, Chubby CO3
locking Service; Amazon Simple Notification Service CO4
(Amazon SNS), multi-player online game hosting on
cloud resources, Building content delivery networks CO5
using clouds.
Text Books
1. Architecting Cloud Computing Solutions by Scott Goessling, Kevin L. Jackson, Publisher:
Packt Publishing, Release Date: May 2018
2. Software Architect's Handbook, by Joseph Ingeno, Published by Packt Publishing, 2018
3. Kai Hwang, Geoffrey Fox, Jack J. Dongarra, Morgan Kaufmann, “Distributed and Cloud
Computing: From Parallel Processing to the Internet of Things,” 1st Edition, 2011.
4. Gautham Shroff, “Enterprise Cloud Computing: Technology, Architecture, Applications”,
Cambridge press, 2010.
5. Learning Path: AWS Certified Machine Learning-Specialty ML, By Noah Gift, April 2022
6. Microservices: Flexible Software Architecture, by Eberhard Wolff, Publisher: Addison-
Wesley Professional, Release Date: October 2016
Suggested Readings
1. KrisJamsa,2014. Cloud computing SaaS, PaaS, Virtualization, Business, Mobile security
and more, 1st Edition, Jones & Batrlett Students Education.
2. Rajkumar Buyya, Christian Vecchiola, [Link], 2013. Mastering cloud computing,
1st Edition, Tata McGrawHill.
3. Arshdeep Bahhga and Vijay Madisetti, 2017. Cloud Computing Hands on Approach, 1st
Edition, University Press.
61
Web Resources
1. [Link]
2. [Link]
3. [Link]
62
Course Code PDS 3505
Course Title Cloud Computing – Lab
Credits 03
Hours/Week 04
Category Major Core(MC)–Lab
Semester III
Regulation 2022
Course Overview
1. This course provides the way to create web applications in the cloud environment.
2. It helps to enable the file sharing and deploying the web applications in the cloud.
3. This course helps to create Docker Artifactory and execute the push/pull commands.
4. It also helps to create pipeline for Git and serverless applications
Course Objectives
1. To explore cloud computing driven commercial systems such as Microsoft Azure,
Amazon AWS, and other cloud applications
2. To provide a foundation of the Cloud Computing enabling them to start using and
adopting Cloud Computing services and tools in their real-life scenarios
3. Formulate DevOps based design and development of cloud applications
4. To impart knowledge in applications of cloud computing
63
III 5. Installing and Configuring Dockers 12 CO1 K1,K2,K3
in local host and running multiple CO2 K4,K5,K6
images on a Docker Platform CO3
6. Create a Docker Repo or CO4
Artifactory and execute Push/Pull CO5
commands for modified docker
base images
IV 7. DevOps deployment of library 12 CO1 K1,K2,K3
automation etc. on the cloud CO2 K4,K5,K6
platform with one complete upgrade CO3
of the application CO4
8. Create One simple Pipeline for Git, CO5
Jenkins, and Docker in local mode
V 9. Serverless on AWS- sample 12 CO1 K1,K2,K3
application AWS Lambda/ AWS CO2 K4,K5,K6
Fargate CO3
10. Amazon Simple Notification CO4
Service - AWS SNS CO5
Text Books
1. John Rhoton and Risto Haukiojal, “Cloud Computing Architectured : Solution
Design Handbook”, Recursive Press, 2013.
2. Dinkar Sitaram, Geetha Manjunathan, “Moving to the Cloud: Developing Apps
in the new world of Cloud Computing”, Syngress, 2012
Suggested Readings
1. Rajkumar buyya, Christian vecchiola, S Thamarai Selvi , “Mastering cloud
computing”, Tata McGraw Hill Education Private Limited, 2013
2. Anthony T .Velte, Toby J. Velte, Robert Elsenpeter, “Cloud Computing a
Practical Approach”, Tata McGraw-HILL, 2010 Edition.
3. Barrie sosinsky, “Cloud computing bible, Wiley publishing
Web Resources
1. [Link]
2. [Link]
3. [Link]
4. [Link]
5. [Link]
64
Course Outcomes (COs)and Cognitive Level Mapping
65
Course Code PDS3601
Course Title NATURAL LANGUAGE PROCESSING
Credits 2
Hours/Week 4
Category ME
Semester III
Regulation 2022
Course Overview:
1. Get an overview of traditional NLP concepts and methods
2. Preprocess the text and text classification.
3. To perform language modelling and sequence tagging.
4. Enable one to perform sequence to sequence task.
5. Semantic and pragmatic analysis on text coherence.
Course Objective:
1. To incorporate basic data pre-processing procedures on text.
2. To develop appropriate language modelling.
3. To apply statistical tools to develop a model for prediction using probabilistic approach.
4. To perform topic analysis using semantic analysis.
SYLLABUS
66
III Probabilistic Models of Pronunciation and 14 CO1 K1
Spelling – Weighted Automata – N- Grams CO2 K2
– Corpus Analysis – Smoothing – Entropy - CO3 K3
Parts-of-Speech – Taggers – Rule based – CO4 K4
Hidden Markov Models – Speech CO5 K5
Recognition. K6
IV Basic Concepts of Syntax – Parsing 12 CO1 K1
Techniques – General Grammar rules for CO2 K2
Indian Languages – Context Free Grammar CO3 K3
– Parsing with Context Free Grammars – CO4 K4
Top Down Parser – Earley Algorithm – CO5 K5
Features and Unification - Lexicalised and K6
Probabilistic Parsing.
V Computational Representation – Meaning 12 CO1 K1
Structure of Language – Semantic Analysis CO2 K2
– Lexical Semantics – WordNet – CO3 K3
Pragmatics – Discourse – Reference CO4 K4
Resolution – Text Coherence – Dialogue CO5 K5
Conversational Agents. K6
TEXT BOOKS:
1 Daniel Jurafskey and James H. Martin “Speech and Language Processing”, Prentice Hall, 2009.
2. Christopher [Link] and Hinrich Schutze, “Foundation of Statistical Natural Language
Processing”, MITPress, 1999.
3. Ronald Hausser, “Foundations of Computational Linguistics”, Springer-Verleg, 1999.
4. James Allen, “Natural Language Understanding”, Benjamin/Cummings Publishing Co. 1995.
SUGGESTED READINGS:
1. James Pustejovsky and Amber stubbs, “Natural language Annotation for machine learning”,
Shroff Publishers,2012.
2. Daniel M. Bikel, “Multilingual Natural language processing”, Pearson, 2012.
3. Emily M. Bender, “Linguistic Fundamentals for Natural language processing”, Margon &
Claypool Pub., 2013.
4. Hobson Lane, “Natural language processing in action”, Manning Pub.,2013.
67
Course Outcomes (COs) and Cognitive Level Mapping
68
Course Code PDS 3602
Course Title Reinforcement Learning
Credits 02
Hours/Week 04
Category Major Elective(ME)–Theory
Semester III
Regulation 2022
Course Overview
1. Reinforcement Learning focuses on general-purpose formalism for automated decision-
making and AI.
2. Reinforcement learning aims to model the trial-and-error learning process that is needed in
many problem situations where explicit instructive signals are not available.
3. This course introduces the statistical learning techniques where an agent explicitly takes
actions and interacts with the world.
4. It enables to understand the importance and challenges of learning agents that make intelligent
decision-making is of vital importance today.
5. This course enables the key concepts of Reinforcement Learning, underlying classic and
modern algorithms in RL.
Course Objectives
1. To formalize problems as Markov Decision Processes.
2. To understand the algorithmic concepts in Temporal Difference Learning.
3. To learn the RL tasks and the core principals behind the RL, including policies and eligibility
traces.
4. To understand and work with function approximate solutions.
5. To Learn the policy gradient methods from vanilla to more complex cases
69
II Unit II: Temporal Difference 12 CO1 K1,K2,K3
Learning CO2 K4,K5,K6
Temporal Difference learning: TD CO3
prediction, Optimality of TD(0), SARSA, CO4
Q-learning, Games and after states, CO5
Maximization Bias and Double Learning.
Text Books
1. R. S. Sutton and A. G. Barto. Reinforcement Learning - An Introduction. MIT Press.2nd
Edition. 2018.
Suggested Readings
1. Li, Yuxi. "Deep reinforcement learning." arXiv preprint arXiv:1810.06339 (2018).
2. Wiering, Marco, and Martijn Van Otterlo. "Reinforcement learning." Adaptation,
learning, and optimization 12 (2012): 3.
3. Russell, Stuart J., and Peter Norvig. "Artificial intelligence: a modern approach."Pearson
Education Limited, 2016.
70
Web Resources
1. [Link]
2. [Link]
3. Video Lectures by Prof. David Silver
4. Video Lectures by Prof. [Link]
71
Course Code PDS 3701
Course Title MEAN Stack
Credits 03
Hours/Week 06
Category Major Core(MC)–Theory
Semester III
Regulation 2022
Course Overview
The aim of a MEAN stack developer is to build complete web applications including frontend,
backend, and database management. It possesses knowledge of every part of the development and
work across a number of tools and frameworks.
Course Objectives
1. To implement Forms, inputs and Services using Angular JS
2. To develop a simple web application using Nodejs; Angular JS and Express
3. To implement data models using Mongo DB
Prerequisites Basic knowledge on front end application using HTML5, CSS3, JavaScript along
with jQuery frame work.
SYLLABUS
UNI CONTENT HOURS Cos COGNITIVE
T LEVEL
I Introduction to Web Technology and Angular JS 12 CO1, K1,
Introduction to Web Technology - Angular CO2, K2,K3,K4, K5
JSModel-View-Controller – Expression -Directives CO3,
and Controllers - Angular JS Modules – Arrays – CO4
Working with ng-model – Working with Forms – CO5
Form Validation – Error Handling with Forms –
Nested Forms with ng-form – Other Form Controls.
72
II DIRECTIVES& BUILDING DATABASES 12 CO1 K1, K2,
Filters – Using Filters in Controllers and Services – CO3 K4.K5,K6
Angular JS Services – Internal Angular JS Services CO4
– Custom Angular JS Services - Directives – CO5
Alternatives to Custom Directives – Understanding
the Basic options – Interacting with Server –HTTP
Services – Building Database, Front End and Back
End
73
Text Book
1. Getting MEAN with Mongo, Express, Angular, and NodeBy Simon Holmes, Clive Herber ·
2022 Manning Publications
2. AgusKurniawan–“AngularJS Programming by Example”, First Edition, PE Press, 2014.
3. David Hows, Peter Membrey, EelcoPlugge – “MongoDB Basics”, Apress, 2014.
4. Ethan Brown, “Web Development with Node and Express”, Oreilly Publishers, First
Edition, 2014
Suggested Readings
1. Full Stack JavaScript Development With MEAN MongoDB, Express, AngularJS, and [Link]
By Colin J Ihrig, Adam Bretz · 2015 SitePoint Pty, Limited
Web Resources
1. [Link]
2. [Link]
3. [Link]
74
Course Code PDS2901
Course Title DATA VISUALIZATION THROUGH R
Credits 01
Hours/Week 02
Category Major Core (MC) – Theory
Semester II
Regulation 2022
Course Overview
This course introduces the basics of R and the practical knowledge of data cleaning,
reorganization, modeling, statistics, and analysis for research and visualization, particularly in
geospatial fields. The goal of the course is to introduce students to the use of R programming
for univariate and multivariate analysis and visualization, mapping and spatial analysis.
Course Objectives
1. Use RStudio to perform basic data analysis functions including Input/Output, basic
Exploratory Data Analysis (EDA), and graphical output.
2. Use RStudio to develop, test, and execute R script.
3. Use advanced R programming to import, clean, transform, and summarize data.
4. Use ggplot2 to visualize data in points, lines, area charts and smoothed curves.
5. Import and map spatial data using R sf and ggplot2 package
Prerequisites No prerequisites
SYLLABUS
75
III Creating Data Frames – Matrix-like 9 CO1 K1,K2,K3
Operations on a Data Frame – Merging CO2 K4,K5,K6
Data Frames – Applying functions to CO3
Data Frames – Factors and Tables – CO4
Common Functions used with Factors –
CO5
Working with Tables
Text Books
1. g gplot2, Elegant Graphics for Data Analysis (2nd Edition), by Hadley Wickham, Springer,
(2016)
2. R for Data Science, Import, Tidy, Transform, Visualize and Model Data, (1st Edition) by
Hadely Wickham and Garrett Grolemund, O’Reilly (2016)
3. Geocomputation with R by Robin Lovelace, Jakub Nowosad, Jannes Muenchow (2019).
Available at [Link]
4. patial Data Science with R by Robert J.
5.
76
Course Outcomes (COs)and Cognitive Level Mapping
77
Course Code PDS3701
Course Title INTER DISCIPLINARY: STATISTICS FOR COMPUTER SCIENCE
Credits 3
Hours/Week 6
Category IDE
Semester III
Regulation 2022
Course Overview:
1. Able to analyse basic characteristics of the features.
2. Can perform univariate and Bivariate analysis.
3. Enable decision making using testing of hypothesis.
4. Based on the relation of the features can be able to form factors.
5. Enable to perform dimension reduction and feature selection.
Course Objective:
1.
To perform Explanatory data analysis.
2.
To study the relationship between the features and develop a model.
3.
To apply statistical techniques and derive factors.
4.
To perform dimension reduction and feature selection and fine tune the precision of
the model
Pre requisites: Basic understanding of Statistics
SYLLABUS
UNIT CONTENT HRS COs COGNITIVE LEVEL
I Sampling Techniques – Data 14 CO1 K1
Classification – Tabulation – Frequency CO2 K2
and graphic Representation – Measures CO3 K3
of Central Tendency – Measures of CO4 K4
Variation – Quartiles and Percentiles– CO5 K5
Moments -Skewness and Kurtosis. K6
II Scatter Diagram – Karl Pearson’s Correlation 15 CO1 K1
Coefficient – Rank Correlation –Correlation CO2 K2
Coefficient for Bivariate Frequency CO3 K3
Distribution – Regression Coefficients – CO4 K4
Fitting of Regression Lines. CO5 K5
K6
78
III Statistical Tests of Significance - Test of 15 CO1 K1
significance for mean(s), variance(s), CO2 K2
correlation coefficient(s), regression CO3 K3
coefficient, based on t, Chi-square and F- CO4 K4
distributions. Applications of Chi-square in CO5 K5
test of significance (independence of K6
attributes, goodness off it).
REFERENCES:
1. Gupta,[Link],V.K.:“FundamentalsofMathematicalStatistics”,Sultan&Chand&S
ons, NewDelhi,11th Ed, 2002.
2. Joseph F Hair, William C Black etal , “Multivariate Data Analysis” , Pearson Education,
7th edition, 2013.
3. Joseph F Hair, William C Black etal , “Multivariate Data Analysis” , Pearson Education,
7th edition, 2013.
4. T. W. Anderson , “An Introduction to Multivariate Statistical Analysis, 3rd Edition”,
Wiley, 2003.
SUGGESTED READINGS:
79
1. [Link]
2. [Link]
techniques/concepts-of-
3. sample-space-sample-points-and-events/
4. [Link]
techniques/concepts-of-
5. sample-space-sample-points-and-events/
80
LOCF BASED DIRECT ASSESSMENTS
COGNITIVE LEVEL (CL) AND COURSE OUTCOME (CO) BASED CIA QUESTION PAPER FORMAT (PG)
SECTION Q. NO COGNITIVE LEVEL (CL)
K1 K2 K3 K4 K5 K6
A (5 x 1 = 5) 1(a) +
Answer ALL (b) +
(c) +
(d) +
(e) +
(5 x 1 = 5) 2(a) +
Answer ALL (b) +
(c) +
(d) +
(e) +
B (1 x 8 = 8) 3 +
Answer 1 out of 2 4 +
C (1 x 8 = 8) 5 +
Answer 1 out of 2 6 +
D (1 x 12 = 12) 7 +
Answer 1 out of 2 8 +
E (1 x 12 = 12) 9 +
Answer 1 out of 2 10 +
No. of CL based Questions with Max. marks 5 (5) 5 (5) 1 (8) 1 (8) 1 (12) 1 (12)
No. of CO based Questions with Max. marks CO1 CO2 CO3 CO4 CO5
10 (10) 1 (8) 1 (8) 1 (12) 1 (12)
Forms of questions of Section A shall be MCQ, Fill in the blanks, True or False, Match the following, Definition, Missing letters. Questions of Sections B, C, D
and E could be Open Choice/ built in choice/with sub sections. Component III shall be exclusively for cognitive levels K5 and K5 with 20 marks each. CIA shall be
conducted for 50 marks with 90 min duration.
COGNITIVE LEVEL (CL) AND COURSE OUTCOME (CO) BASED END SEMESTER EXAMINATION QUESTION PAPER FORMAT (PG)
SECTION Q. NO COGNITIVE LEVEL (CL)
K1 K2 K3 K4 K5 K6
A (5 x 1 = 5) 1(a) +
Answer ALL (b) +
(c) +
(d) +
(e) +
(5 x 1 = 5) 2(a) +
Answer ALL (b) +
(c) +
(d) +
(e) +
B (3 x 10 = 30) 3 +
Answer 3 out of 5 4 +
5 +
6 +
7 +
C (2 x 12.5 = 25) 8 +
Answer 2 out of 4 9 +
10 +
11 +
D (1 x 15 = 15) 12 +
Answer 1 out of 2 13 +
E (1 x 20 = 20) 14 +
Answer 1 out of 2 15 +
No. of CL based Questions with Max. marks 5 (5) 5 (5) 3 (30) 2 (25) 1 (15) 1 (20)
No. of CO based Questions with Max. marks CO1 CO2 CO3 CO4 CO5
10 (10) 3 (30) 2 (25) 1 (15) 1 (20)
IMPORTANT
Forms of questions of Section A shall be MCQ, Fill in the blanks, True or False, Match the following, Definition, Missing letters.
Questions of Sections B, C, D and E could be Open Choice/ built in choice/questions with sub divisions.
Maximum sub divisions in questions of Sections B, C shall be 2 and 4 in Sections D, E).