0% found this document useful (0 votes)
7 views12 pages

4 Topics

The document provides an overview of big data, including its storage, processing, and analysis techniques. It explains the importance of scalable architecture for managing large data sets and outlines various storage technologies such as data lakes and warehouses. Additionally, it discusses the stages of big data processing and highlights key analysis techniques that leverage machine learning and statistical methods to derive insights from data.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views12 pages

4 Topics

The document provides an overview of big data, including its storage, processing, and analysis techniques. It explains the importance of scalable architecture for managing large data sets and outlines various storage technologies such as data lakes and warehouses. Additionally, it discusses the stages of big data processing and highlights key analysis techniques that leverage machine learning and statistical methods to derive insights from data.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Topics

*Big Data Storage Concepts ( DONE)


*Big Data Processing Concepts ( DONE)
*Big Data Storage Technology, and ( DONE)
*Big Data Analysis Techniques ( DONE)
PA- REVISE IF IBA YUNG WAY AND FORMMATING MO BUT IF WANT NA
IPASA NA NG GANYAN GOW MA’AM KASO BAKA MAY IDADAGDAG KA PA
OR GUSTONG BAGUHIN OR ALISIN . HAHAHAHAHAHA GOODLUCK 

What is big data?


Big data refers to extremely large data sets, either structured or not, that professionals
analyze to discover trends, patterns, or behaviors. It's unique in that it has what
professionals describe as the three Vs—volume, velocity, and variety—in such large
amounts that traditional data management systems struggle to store or analyze the data
successfully. Therefore, scalable architecture must be available to manage, store, and
analyze big data sets.

What is big data storage?


Big data storage is a scalable architecture that allows businesses to collect, manage, and
analyze immense sets of data in real time. The design of big data storage solutions is
specifically tailored to address the speed, volume, and complexity of the data sets.
Some examples of big data storage options are:
 Data lakes are centralized storage solutions that process and secure data in its native
format without size limitations. They can enable different forms of smart analytics, such
as machine learning and visualizations.
 Data warehouses aggregate data sets from different sources into a single storage unit
for robust analysis, supporting data mining, artificial intelligence (AI), and more.
Unlike a data lake, data warehouses have a three-tier structure for storing data.
 Data pipelines gather raw data and transport it into repositories, such as lakes or
warehouses.

Data lakes, warehouses, and pipelines exist within several different storage options,
including:
 Cloud-based storage system is where a business outsources the storage of its data to a
vendor that operates a cloud storage system.
 Colocation storage is the process of a business renting space to store its servers rather
than having it on-site.
 On-premise storage is where a business manages its network and servers on-site. This
can include hardware, such as servers, that houses the data at an organization’s
premises.

What is big data storage used for?


The primary purpose of big data storage is to successfully store immense amounts of
data for future analysis and use. Big data is crucial for businesses and organizations,
from healthcare research to retailers and security, to make more efficient, informed, and
effective decisions. Without big data storage, businesses wouldn’t have the time,
money, or technology to store and manage big data sets successfully.

Because big data is valuable for processing and understanding patterns and trends, it
needs correct storage. Big data storage makes applying big data to business decisions
possible.

How does big data storage work?


Big data storage employs a system of commodity servers and high-capacity disks
capable of analyzing the data sets. For example, in a cloud storage scenario, the big data
sets exist in a server hosted in an off-site location that can be accessed through the
internet. Virtual machines provide the space for the data to live safely, and it’s possible
to quickly create more virtual machines when the amount of data grows past the
servers’ current capacity.

Who uses big data storage?


Professionals across multiple industries use big data storage to store, manage, and
analyze their data. These industries include health care, finance, government, education,
and retail. These industries benefit from big data storage because it provides a unique
opportunity to analyze data at a large scale, offering insights that wouldn’t be possible
otherwise, such as predictions for the future and customer behavior analysis.

Pros and cons of using big data storage


The pros and cons of big data storage typically relate to the volume of data being
handled. Here are some advantages of using big data storage:
 Data-driven. The large-scale data analysis allows businesses to become data-driven
using concrete data to help make decisions and better inform strategic planning.
 Make safe and informed decisions. Big data storage keeps data safe and lets
professionals apply analytical tools to the data sets, resulting in more informed
decision-making, better customer service, more flexibility in strategic planning, and
increased efficiency in operations.
 Flexible. Cloud-based storage is flexible and allows businesses to scale their needed
servers up or down without up-front investment.

On the other hand, here are some factors to consider when using big data storage:
 Costly. It’s expensive for a business to purchase the necessary space to store big data
sets, and the cost will only increase as more data becomes available. For example, if a
business opts to manage its servers on-site, it may face the risk of needing to purchase
more systems and the staff to run them.
What is Big Data Processing?

Big Data Processing is the collection of methodologies or frameworks enabling access


to enormous amounts of information and extracting meaningful insights. Initially, Big
Data Processing involves data acquisition and data cleaning. Once you have gathered
the quality data, you can further use it for Statistical Analysis or building Machine
Learning models for predictions.

5 Stages of Big Data Processing


 Data Extraction
 Data Transformation
 Data Loading
 Data Visualization/BI Analytics
 Machine Learning Application
Stage 1: Data Extraction

This initial step of Big Data Processing consists of collecting information from
diverse resources like enterprise applications, web pages, sensors, marketing tools,
transactional records, etc. Data processing professionals extract information through
many Unstructured and Structured Data Streams. For instance, in building a Data
Warehouse, extracting entails merging information from multiple sources,
subsequently verifying the information by removing incorrect data. To make future
decisions based on the outcomes, the data collected during the data collection phase of
Big Data Processing must be labeled and accurate. This stage establishes a
quantitative standard as well as a goal for improvement.

Stage 2: Data Transformation

The data transformation phase of Big Data Processing defines changing or modifying
data into required formats which helps in building different insights and
visualizations. There are many transformation techniques like Aggregation,
Normalization, Feature Selection, Binning and Clustering, and concept hierarchy
generation. Using these techniques for Big Data Processing, developers transform
Unstructured Data into Structured Data and Structured Data into a user-
understandable format. Business and Analytical operations become more efficient as a
result of the transformation, and firms can make better data-driven choices.

Stage 3: Data Loading

The converted data is transported to the centralized database system in the load stage
of Big Data Processing. Before loading the data, index the database and remove the
constraints to make the process more efficient. Using Big Data ETL, the process of
loading became automated, well-defined, consistent, and Batch-driven or Real-time.

Stage 4: Data Visualization/BI Analytics

Data Analytics tools and methods for Big Data Processing enable firms to visualize
huge datasets and create dashboards for gaining an overview of the entire business
operations. Business Intelligence (BI) Analytics answer fundamental business growth
and strategy questions. BI tools make predictions and what-if analyses on the
transformed data that help stakeholders understand the depth patterns in data and the
correlations between the attributes.
Stage 5: Machine Learning Application

The Machine Learning phase of Big Data Processing is primarily concerned with the
creation of models that can learn to evolve in response to the new input. The learning
algorithms allow for more quickly analyzing large amounts of data.

 The first type of Machine Learning is Supervised Learning, which uses


labelled data for training the models and predicting the outcomes. Data patterns
are used in Supervised learning to identify new information output for the
labels. This method is often used in applications that utilize historical data to
predict future outcomes.
 Unsupervised Learning is the second type where the data is unlabeled and
trained by the algorithm. Unsupervised Machine Learning is utilized against
information that doesn’t have any historical labels.
 Reinforcement Learning is the final type in which there is no primary data
that can be inserted as input to models. The algorithms have to figure out the
decisions on their own based on observations or situations that happen
surrounding them. The decisions are manipulated with a reward function so that
the models try to make the correct decisions.

The Machine Learning phase of Big Data Processing enables automatic recognition
patterns and can perform feature extraction in complicated unstructured information
without any human interference, making this a significant resource for Big Data
research.
Data Storage Technologies

Apache Hadoop, Apache HBase, and Snowflake are three big data storage technologies often
used in the data lake analytics paradigm.

Hadoop

Hadoop has gained considerable attention as it is one of the most common frameworks to
support big data analytics. A distributed processing framework based on open-source software,
Hadoop enables large data sets to be processed across clusters of computers. Large data sets were
initially intended to be processed and stored across clusters of commodity hardware.

HBase

With HBase, you can use a NoSQL database or complement Hadoop with a column-oriented
store. This database is designed to efficiently manage large tables with billions of rows and
millions of columns. The performance can be tuned by adjusting memory usage, the number of
servers, block size, and other settings.
Snowflake

Snowflake for Data Lake Analytics is an enterprise-grade cloud platform for advanced analytics
applications built on top of Apache Hadoop. It offers real-time access to historical and streaming
data from any source and format at any scale without requiring changes to existing applications
or workflows. It also enables users to quickly scale up their processing power as needed without
having to worry about infrastructure management tasks such as provisioning
Big Data Analysis Techniques

The global big data market revenues for software and services are expected to increase from $42
billion to $103 billion by the year 2027.1 Every day, 2.5 quintillion bytes of data are created, and
it’s only in the last two years that 90% of the world’s data has been generated.2 If that’s any
indication, there’s likely much more to come.
The world is driven by data, and it’s being analyzed every second, whether it’s through your
phone’s Google Maps, your Netflix habits, or what you’ve reserved in your online shopping cart
– in many ways, data is unavoidable and it’s disrupting almost every known market.3 The
business world is looking to data for market insights and ultimately, to generate growth and
revenue. Although data is becoming a game changer within the business arena, it’s important to
note that data is also being utilized by small businesses, corporate, and creatives alike. A global
survey from McKinsey revealed that when organizations use data, it benefits the customer and
the business by generating new data-driven services, developing new business models and
strategies, and selling data-based products and utilities.4 The incentive for investing and
implementing data analysis tools and techniques is huge, and businesses will need to adapt,
innovate, and strategize for the evolving digital marketplace.
Every day, 2.5 quintillion bytes of data are created, and it’s only in the last two years that 90% of
the world’s data has been generated.

What is data analysis?


Data analysis, or analytics (DA) is the process of examining data sets (within the form of text,
audio, and video), and drawing conclusions about the information they contain, more commonly
through specific systems, software, and methods. Data analytics technologies are used on an
industrial scale, across commercial business industries, as they enable organizations to make
calculated, informed business decisions.5
Globally, enterprises are harnessing the power of various data analysis techniques and using
them to reshape their business models.6 As technology develops, new analysis software emerges,
and as the Internet of Things (IoT) grows, the amount of data increases. Big data has evolved as
a product of our increasing expansion and connection, and with it, new forms of extracting, or
rather “mining”, data.

Six Big Data Analysis Techniques


Big data is characterized by the three V’s: the major volume of data, the velocity at which it’s
processed, and the wide variety of data.7 It’s because of the second descriptor, velocity, that data
analytics has expanded into the technological fields of machine learning and artificial
intelligence.8 Alongside the evolving computer-based analysis techniques data harnesses, the
analysis also relies on traditional statistical methods.9 Ultimately, how data analysis techniques
function within an organization is twofold; big data analysis is processed through the streaming
of data as it emerges, and then batch analysis of data as it builds – to look for behavioral patterns
and trends.10 As the generation of data increases, so will the various techniques that manage it.
As data becomes more insightful in its speed, scale, and depth, the more it fuels innovation.

The world is driven by data, and it’s being analyzed every second, whether it’s through your
phone’s Google Maps, your Netflix habits, or what you’ve reserved in your online shopping cart.

McKinsey’s big data report identifies a range of big data techniques and technologies, that draw
from various fields such as statistics, computer science, applied mathematics, and
economics.11 As these methods rely on diverse disciplines, the analytics tools can be applied to
both big data and other smaller datasets:
1. A/B testing
This data analysis technique involves comparing a control group with a variety of test groups, in
order to discern what treatments or changes will improve a given objective variable. McKinsey
gives the example of analyzing what copy, text, images, or layout will improve conversion rates
on an e-commerce site.12 Big data once again fits into this model as it can test huge numbers,
however, it can only be achieved if the groups are of a big enough size to gain meaningful
differences.
2. Data fusion and data integration

By combining a set of techniques that analyze and integrate data from multiple sources and
solutions, the insights are more efficient and potentially more accurate than if developed through
a single source of data.

3. Data mining

A common tool used within big data analytics, data mining extracts patterns from large data sets
by combining methods from statistics and machine learning, within database management. An
example would be when customer data is mined to determine which segments are most likely to
react to an offer.

4. Machine learning
Well known within the field of artificial intelligence, machine learning is also used for data
analysis. Emerging from computer science, it works with computer algorithms to produce
assumptions based on data.14 It provides predictions that would be impossible for human
analysts.
5. Natural language processing (NLP).
Known as a subspecialty of computer science, artificial intelligence, and linguistics, this data
analysis tool uses algorithms to analyze human (natural) language.15
6. Statistics.

This technique works to collect, organize, and interpret data, within surveys and experiments.

Other data analysis techniques include spatial analysis, predictive modeling, association rule
learning, network analysis, and many, many more. The technologies that process, manage, and
analyze this data are of an entirely different and expansive field, that similarly evolves and
develops over time. Techniques and technologies aside, any form or size of data is valuable.
Managed accurately and effectively, it can reveal a host of business, product, and market
insights. What does the future of data analysis look like? It’s hard to say with the tremendous
pace of analytics and technology progress, but undoubtedly data innovation is changing the face
of business and society in its holistic entirety.

SOURCE
[Link]
[Link]
[Link]
[Link]

You might also like