Big Data
Introduction to Big Data-Big Data - Beyond The Hype, Big Data Skills And Sources Of Big
Data, Big Data Adoption, Research And Changing Nature Of Data Repositories, Data
Sharing And Reuse Practices And Their Implications For Repository Data Curation
Not Found this is basic question
Introduction of Big Data Programming-Hadoop
[Link]
10
[Link]
10
The ecosystem and stack
[Link]
10
The Hadoop Distributed File System (HDFS)
[Link]
14
Components of Hadoop
[Link]
6000071e202e28ae1a1ba75d
1
Design of HDFS ( Specific not found )
[Link]
25( fudche pn mcq ahe 16 parynt hdfs ahe)
Java interfaces to HDFS,
[Link]
10
Architecture overview (not found Fact theory kar )
Development Environment, not found
Hadoop distribution and basic commands (not found Fact theory kar )
Eclipse development, Ha point yenar ch nahi 1000% tri pn kadhaloi tu mhanayla nako
kadhala nahis ( karu nakos ch)
[Link]
10
The HDFS command line and web interfaces
[Link]
called-used-to-interact-with-hdfs
1(Most imp question ahe)
Not found more
The HDFS Java API Not found more
Analyzing the Data with Hadoop
[Link]
10
[Link]
5
Scaling Out,
[Link]
10
[Link]
5
Hadoop event stream processing
[Link]
10
complex event processing jast gavayla nahi theory kar hyace
[Link]
computing/?page=7
MCQ no. 62
MapReduce Introduction,
[Link]
10
Developing a Map Reduce Application
[Link]
10
[Link]
10
[Link]
model/
10
[Link]
model/?page=2
10
[Link]
model/?page=3
10
How Map Reduce Works,Specific point not fount
The MapReduce Anatomy of a Map Reduce Job run, Failures, Job Scheduling, Shuffle and
Sort, Task execution,
[Link]
model/?page=4
10
[Link]
model/?page=5
10
[Link]
model/?page=7
10
[Link]
model/?page=8
10 Exam veda saglech karyla pahijet
Map Reduce Types and Formats, not fount
Map Reduce Features
[Link]
10
[Link]
10
Real-World MapReduce, not found
Hadoop Environment: Setting up a Hadoop Cluster, Cluster specification, Cluster Setup
and Installation, Hadoop Configuration, Security in Hadoop, Administering Hadoop, HDFS –
Monitoring & Maintenance, Hadoop benchmarks,
not found
Apache Airflow/ETL Informatica: Introduction to Data warehousing and Data lakes
[Link]
10 (hyache specific ase apache airflow che nahi milat ahet tula notes ch karave lagel)
[Link]
10
[Link]
10
Designing Data warehousing for an ETL Data Pipeline not found
[Link]
125
[Link]
114
Designing Data Lakes for ETL Data Pipeline, not found
ETL vs ELT not found
[Link]
%20Extract%20Transform,is%20not%20involved%20in%20ELT.
Only theory ahe
Introduction to HIVE ( ha subject ch nav introduction ahe tyamule lai deep karyache
nahi basic karayche )
Programming with Hive
[Link]
10
[Link]
60
Data warehouse system for Hadoop, not found
Optimizing with Combiners and Practitioners, not found
Bucketing
[Link]
10
more common algorithms: sorting, indexing and searching, Ha point mla ky kalena gg pn
evdhe kalale ki ha practical point ahe mhanun varche 60 pn kar
[Link]
10
[Link]
10
[Link]
10
[Link]
10
[Link]
10
[Link]
10
Relational manipulation: map-side and reduce-side joins not found
Evolution,purpose and use, Case Studies on Ingestion and warehousing not found
HBase: Overview,
[Link]
10
[Link]
10
[Link]
10
[Link]
10
comparison and architecture, java client API, CRUD operations and security not found
Apache Spark: Overview,
[Link]
10
[Link]
10
APIs for large-scale data processing, Linking with Spark, Initializing Spar, not found
Resilient Distributed Datasets (RDDs),
[Link]
q=9aZHDjblmRk=
10
[Link]
q=bDFvcmdllQA=
10
[Link]
q=6wHWymi73nk=
10
External Datasets, not found
RDD Operations
[Link]
55
Passing Functions to Spark (may be practical oriented ahe purn), Job optimization, Working
with Key-Value Pairs, Shuffle operations, not found
RDD Persistence,
[Link]
notes ahe bg he
Removing Data, not found
Shared Variables,
[Link]
[Link]
1
EDA using PySpark,
[Link]
analysis-(eda)/
10
[Link]
analysis-(eda)/?page=2
10
[Link]
analysis-(eda)/?page=3
10
[Link]
analysis-(eda)/?page=4
10
[Link]
analysis-(eda)/?page=5
10
Deploying to a Cluster Spark Streaming, not found