BIG DATAANALYTICS Semester 7
Coursc Code BCS714D CIE Marks 50
Teaching Hours/Week (L.T:P.S) 3:0:0:0 SEE Marks 50
Total Hours of Pedagogy 40 Total Marks 100
Credits 03 Exam Hours 3
Examination nature (SEE) Theory
Course objectives:
1 To implement MapReduce programs for processing big data.
2
To realize storage and processing of big data using MongoDB,Pig, Hive and Spark.
3 Toanalyze big data using machine learning techniques.
Teaching-Learning Process (General Instructions)
These are sample Strategies; that teachers can use to accelerate the attainment of the various course outcomes.
1. Lecturer method (L) needs not to be only a traditional lecture method, but alternativec cffective teaching
methods could be adopted to attain the outcomes.
2. Use of Video/Animation to explain functioning of various concepts.
3. Encourage collaborative (Group Learning) Learning in the class.
4
Ask at least three HOT (Higher order Thinking) questions in the class, which promotes critical thinking.
5 Discuss how every concept can be applied to the real world -and when that's possible, it helps improve
the students' understanding.
6 Use any of these methods: Chalk and board, Active Learning, Case Studies.
MODULE-1
Classificationof data, Characteristics, Evolution and definition of Big data, What is Big data, Why Big data,
Traditional Business Intelligence Vs Big Data,Typical data warehouse and Hadoop environment.
Big Data Analytics: What is Big data Analytics, Classification of Analytics, Importance of Big Data
Analytics, Technologies used in Big data Environments, Few Top Analytical Tools ,NoSQL, Hadoop.
TB1: Ch 1: 1.1. Ch2: 2.1-[Link],2.9-2. I1, Ch3: 3.2,3.5,3.8,3.12, Ch4: 4.1,4.2
MODULE-2
Introduction to Hadoop: Introducing hadoop, Why hadoop, Why not RDBMS, RDBMS Vs Hadoop. History
of Hadoop, Hadoop overview, Use case of Hadoop, HDFS (Hadoop Distributed File System), Processing data
with Hadoop. Managing resources and applications with Hadoop YARN(Yet Another Resource Negotiator).
Introduction to Map Reduce Programming: Introduction, Mapper, Reducer, Combiner, Partitioner.
Searching, Sorting, Compression.
TBI: Ch 5:5.1-,5.8, 5.10-5.12, Ch 8: 8.1 -8.8
MODULE-3
Introduction toMongoDB: What is MongoDB, Why MongoDB,Terms used in RDBMS and MongoDB, Data
Types in MongoDB, MongoDB Query Language.
TB1: Ch 6: 6.1-6.5
MODULE-4
Introduction to Hive: What is Hive, Hive Architecture, Hive data types, Hive file formats, Hive Query
Language (HQL), RC File implementation, User Defined Function (UDF).
Introduction to Pig: What is Pig, Anatomy of Pig, Pig on Hadoop. Pig Philosophy, Use case for Pig. Pig Latin
Overview, Data types in Pig, Running Pig, Execution Modes of Pig, HDFS Commands, Relational Operators,
Eval Function, Complex Data Types,Piggy Bank,User Defined Function, Pig Vs Hive.
TBI: Ch9:9.1-9.6,9.8, Ch 10: 10.1 - 10.15, 10.22
MODULE-5
Spark and Big Data Analytics: Spark, Introduction to Data Analysis with Spark.
Text, Web Content and Link Analytics: Introduction, Text Mining, Web Mining. Web Content and Web
Usage Analytics, Page Rank, Structure of Wch and Analyzing a Wch Graph.
TB2: ChS: [Link], Ch9: 9.|-94
Course outcomes (Course SkillSet):
At the end of the course,the student will be able to:
Illustrate Big Data concepts, tools and applications.
Devclop programs using HADOOP framework.
Use Hadoop Cluster to deploy Map Reduce jobs, PIG,HIVE and Spark programs.
Analyze the given data set to identify deep insights.
Assessment Details (both CIE and SEE)
The weightage of Continuous Internal Evaluation (CIE) is 50% and for Semester End Exam (SEE)is 50%.
The minimum passing mark for theCIE is 40% of the maximum marks (20 marks out of 50) and for the
SEE minimum passing mark is 35% of the maximum marks (18 out of 50 marks). A student shall be
deemed to have satisfied the academic requirements and earned the credits allotted to each subject/
course if the student secures a minimum of 40% (40 marks out of 100) in the sum total of the CIE
(Continuous Internal Evaluation) and SEE (Semester End Examination) taken together.
Continuous Internal Evaluation:
For the Assignment component of the CIE, there are 25 marks and for the Internal Assessment Test
component, there are 25 marks.
The first test will be administered after 40-50% of the syllabus has been covered, and the second test will
be administered after 85-90% of the syllabus has been covered
Any two assignment methods mentioned in the 220B2.4, if an assignment is project-based then only one
assignment for the course shall be planned. The teacher should not conduct two assignments at the end
of the semester if two assignments are planned.
For the course, CIE marks will be based on a scaled-down sum of two tests and other methods of
assessment.
Internal Assessment Test question paper is designed to attain the different levels of Bloom's
taxonomy as per the outcome defined for thecourse.
Semester-End Examination:
Theory SEE will be conducted by University as per the scheduled timetable, with common question
papers for the course (duration 03 hours).
The question paper will have ten questions. Each question isset for 20 marks.
There willbe 2 questions from each module. Each of the two questions under a module (with a
maximum of 3sub-questions), should have a mix of topics under that module.
The students have to answer 5 full questions, selecting one fullquestion from each module.
Marks scored shallbe proportionally reduced to 50 marks.
Suggested Learning Resources:
Books:
1 Seema Acharya and Subhashini Cheliappan "Big data and Analytics", VWiley India Publishers, 2nd Edition,
2019.
2. Rajkamal and Preeti Saxena, "Big Data Analytics, Introduction to Hadoop, Spark and Machine Learning",
McGraw Hill Publication,2019.
Reference Books:
1. Adam Shook and Donald Mine, "MapRcduce Design Patterns: Building Efective Algorithms and Analytics for
Hadoop and Other Systems" -O'Reily 2012
2 Tom White, "Hadoop: The Definitive Guide" 4th Edition, O'reilly Media, 2015.
3. Thomas Erl, Wajid Khattak, and Paul Buhler, Big Data Fundamentals: Concepts,
Drivers & Techniques,
Pearson India Education Service Pvt. Ltd., 1 Edition, 2016
4. John D. Kelleher, Brian Mac Namee, Aoife D'Arcy -Fundamentals of Machine Learning for Predictive Data
Analytics: Algorithms, Worked Examples, MIT Press 2020, 2nd Edition
Web links and Video Lectures (e-Resources):
[Link] k-g5W1mo37urlQ0dCZ.
[Link]
dex=4
1kg5W l mo37urlQ0dCZ&in
[Link] [Link] 92A
Activity Based Learning (Suggested Activities in Class)/ Practical Based learning
1. Implement MongoDB based application to store big data for data processing and analyzing the results [15
marks]
2 Install Hadoop and Implement the following file management such as Adding files and directories, Retrieving
files, Deleting files and directories and execute Map- Reduce based programs.[10]