0% found this document useful (0 votes)
5 views7 pages

Big Data and Hadoop Overview

The document provides an overview of Big Data, focusing on various aspects such as programming with Hadoop, the Hadoop ecosystem, and components like HDFS and MapReduce. It includes links to resources for further study and practice questions related to Hadoop and data processing. Additionally, it touches on related technologies like Apache Spark and data warehousing concepts, highlighting areas where specific information is missing or needs further exploration.

Uploaded by

hackers.iknow
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views7 pages

Big Data and Hadoop Overview

The document provides an overview of Big Data, focusing on various aspects such as programming with Hadoop, the Hadoop ecosystem, and components like HDFS and MapReduce. It includes links to resources for further study and practice questions related to Hadoop and data processing. Additionally, it touches on related technologies like Apache Spark and data warehousing concepts, highlighting areas where specific information is missing or needs further exploration.

Uploaded by

hackers.iknow
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Big Data

Introduction to Big Data-Big Data - Beyond The Hype, Big Data Skills And Sources Of Big
Data, Big Data Adoption, Research And Changing Nature Of Data Repositories, Data
Sharing And Reuse Practices And Their Implications For Repository Data Curation
Not Found this is basic question

Introduction of Big Data Programming-Hadoop


[Link]
10
[Link]
10
The ecosystem and stack
[Link]
10

The Hadoop Distributed File System (HDFS)


[Link]
14

Components of Hadoop
[Link]
6000071e202e28ae1a1ba75d
1
Design of HDFS ( Specific not found )
[Link]
25( fudche pn mcq ahe 16 parynt hdfs ahe)

Java interfaces to HDFS,


[Link]
10
Architecture overview (not found Fact theory kar )
Development Environment, not found
Hadoop distribution and basic commands (not found Fact theory kar )
Eclipse development, Ha point yenar ch nahi 1000% tri pn kadhaloi tu mhanayla nako
kadhala nahis ( karu nakos ch)
[Link]
10

The HDFS command line and web interfaces


[Link]
called-used-to-interact-with-hdfs
1(Most imp question ahe)
Not found more
The HDFS Java API Not found more
Analyzing the Data with Hadoop
[Link]
10
[Link]
5

Scaling Out,
[Link]
10
[Link]
5

Hadoop event stream processing


[Link]
10

complex event processing jast gavayla nahi theory kar hyace


[Link]
computing/?page=7
MCQ no. 62
MapReduce Introduction,
[Link]
10
Developing a Map Reduce Application
[Link]
10
[Link]
10

[Link]
model/
10
[Link]
model/?page=2
10
[Link]
model/?page=3
10

How Map Reduce Works,Specific point not fount


The MapReduce Anatomy of a Map Reduce Job run, Failures, Job Scheduling, Shuffle and
Sort, Task execution,
[Link]
model/?page=4
10

[Link]
model/?page=5
10
[Link]
model/?page=7
10
[Link]
model/?page=8
10 Exam veda saglech karyla pahijet
Map Reduce Types and Formats, not fount
Map Reduce Features
[Link]
10
[Link]
10
Real-World MapReduce, not found

Hadoop Environment: Setting up a Hadoop Cluster, Cluster specification, Cluster Setup


and Installation, Hadoop Configuration, Security in Hadoop, Administering Hadoop, HDFS –
Monitoring & Maintenance, Hadoop benchmarks,
not found

Apache Airflow/ETL Informatica: Introduction to Data warehousing and Data lakes


[Link]
10 (hyache specific ase apache airflow che nahi milat ahet tula notes ch karave lagel)
[Link]
10
[Link]
10

Designing Data warehousing for an ETL Data Pipeline not found


[Link]
125

[Link]
114
Designing Data Lakes for ETL Data Pipeline, not found
ETL vs ELT not found
[Link]
%20Extract%20Transform,is%20not%20involved%20in%20ELT.
Only theory ahe

Introduction to HIVE ( ha subject ch nav introduction ahe tyamule lai deep karyache
nahi basic karayche )
Programming with Hive
[Link]
10

[Link]
60
Data warehouse system for Hadoop, not found
Optimizing with Combiners and Practitioners, not found
Bucketing
[Link]
10
more common algorithms: sorting, indexing and searching, Ha point mla ky kalena gg pn
evdhe kalale ki ha practical point ahe mhanun varche 60 pn kar
[Link]
10
[Link]
10
[Link]
10
[Link]
10
[Link]
10
[Link]
10
Relational manipulation: map-side and reduce-side joins not found
Evolution,purpose and use, Case Studies on Ingestion and warehousing not found

HBase: Overview,
[Link]
10
[Link]
10
[Link]
10
[Link]
10
comparison and architecture, java client API, CRUD operations and security not found

Apache Spark: Overview,


[Link]
10
[Link]
10
APIs for large-scale data processing, Linking with Spark, Initializing Spar, not found

Resilient Distributed Datasets (RDDs),


[Link]
q=9aZHDjblmRk=
10
[Link]
q=bDFvcmdllQA=
10
[Link]
q=6wHWymi73nk=
10

External Datasets, not found


RDD Operations
[Link]
55
Passing Functions to Spark (may be practical oriented ahe purn), Job optimization, Working
with Key-Value Pairs, Shuffle operations, not found
RDD Persistence,
[Link]
notes ahe bg he

Removing Data, not found


Shared Variables,
[Link]
[Link]
1

EDA using PySpark,


[Link]
analysis-(eda)/
10
[Link]
analysis-(eda)/?page=2
10
[Link]
analysis-(eda)/?page=3
10
[Link]
analysis-(eda)/?page=4
10
[Link]
analysis-(eda)/?page=5
10

Deploying to a Cluster Spark Streaming, not found

You might also like