0% found this document useful (0 votes)
8 views14 pages

Big Data Overview: Growth & Characteristics

Chapter 8 discusses Big Data, emphasizing its rapid growth, evolution from traditional databases, and its role as a competitive advantage. It outlines the definition, composition, growth drivers, and core characteristics of Big Data, highlighting the Four V's: Volume, Velocity, Variety, and Veracity. The chapter also provides updates on the current state of global data volume and mobile databases, reflecting significant advancements and the increasing importance of data management.

Uploaded by

iryanissecond
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views14 pages

Big Data Overview: Growth & Characteristics

Chapter 8 discusses Big Data, emphasizing its rapid growth, evolution from traditional databases, and its role as a competitive advantage. It outlines the definition, composition, growth drivers, and core characteristics of Big Data, highlighting the Four V's: Volume, Velocity, Variety, and Veracity. The chapter also provides updates on the current state of global data volume and mobile databases, reflecting significant advancements and the increasing importance of data management.

Uploaded by

iryanissecond
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Summary of Chapter 8: Big Data

It emphasizes the rapid growth of data and its implications for processing and
decision-making. Key sections include the definition, composition, growth drivers,
real-world examples (from 2014), a projection to 2020, and core characteristics.
The document highlights how Big Data has evolved from traditional databases
and positions it as a competitive advantage in the modern era.
Definition and Evolution
Big Data is described as "a collection of data sets so large and complex that it
becomes difficult to process using on-hand database management tools or
traditional data processing applications." The chapter traces its evolution:
 1980s (Databases): Relational databases, gigabytes in size, low latency.
 1990s (Data Warehousing): Terabytes in size, custom hardware.
 Today: Massive scales requiring new approaches, with data doubling
rapidly and access becoming a key competitive edge.
Composition of Big Data
Big Data consists of both structured and unstructured information:
 10% Structured: Data in databases, representing about 10% of the
overall story.
 90% Unstructured: "Human information" such as emails, videos, tweets,
Facebook posts, call-center conversations, closed-circuit TV footage,
mobile phone calls, and website clicks.
Reasons for Growth
The growth of Big Data is driven by three main factors:
1. Increase in storage capacities.
2. Increase in processing power.
3. Availability of data (various types).
Examples of Data Generation (2014 Data)
The chapter illustrates the "data deluge" with what happened online in 60
seconds in 2014:
 50 billion messages sent (WhatsApp).
 204 million emails sent.
 3.3 million Facebook posts.
 4 million Google searches.
 4 million photos uploaded (Instagram).
 342,000 tweets (Twitter).
 1.4 million voice calls (Skype).
 120 hours of video uploaded (YouTube).
Growth Projection (from 2014 Perspective)
Big Data was expected to reach 40,000 exabytes (or 40 zettabytes/40 trillion
gigabytes) by 2020, marking explosive growth.
Gartner Definition
“Big data are high volume, high-velocity and high-variety information assets that
require new forms of processing to enable enhanced decision making, insight
discovery and process optimization.”
Characteristics: The Four V's of Big Data
The chapter outlines the core traits of Big Data using the "Four V's" framework:

V Description Details

Volum
Scale of Data Terabytes to exabytes of existing data to process.
e

Velocit Analysis of
Streaming data, milliseconds to seconds to respond.
y Streaming Data

Different Forms of
Variety Structured, unstructured, text, multimedia.
Data

Veracit Managing the reliability and predictability of


Uncertainty of Data
y inherently imprecise data types.

Management Issues and Closing Quote


The document touches on management challenges in handling Big Data and
ends with a quote from Kevin Weil: “It’s no longer hard to find the answer to a
given question. The hard part is finding the right question and as questions
evolve, we gain better insight into our ecosystem and our business.”
Updates as of December 2025
The chapter's data is from 2014, with projections to 2020 that underestimated
the actual explosion driven by AI, IoT, and cloud computing. Here's an update
based on current estimates:
Current Global Data Volume
As of 2025, the global datasphere (total data created, captured, and replicated)
is estimated at around 180-200 zettabytes (ZB), far exceeding the 2014
projection of 40 ZB by 2020. Sources like IDC and Statista project continued
tripling by 2029, with daily creation at about 2.5 quintillion bytes. This growth is
fueled by AI-generated content, streaming, and edge devices.
[Link]
175 Zettabytes By 2025
Updated Internet Activity in 60 Seconds (2025 Estimates)
Internet usage has surged since 2014, with platforms like TikTok and AI tools
adding to the volume. Based on reports from Domo's "Data Never Sleeps" and
similar analyses, here's what happens online every minute in 2025:
 Emails sent: 231.4 million (up from 204 million in 2014).
 Google searches: 5.9 million (up from 4 million).
 Texts sent: 16 million.
 Instagram stories/Reels posted and viewed: 695,000.
 TikTok videos uploaded: 16,000.
 YouTube videos streamed: Millions (exact figures vary, but video hours
uploaded exceed 500+ per minute).
 Online spending: $43.6 million (e-commerce boom).
 Other: Over 5 million app downloads, billions of social media interactions,
and massive AI queries.
These stats reflect a more diverse ecosystem, including short-form video and e-
commerce, compared to 2014's focus on emails and posts.
[Link]
What Happens in 60 Seconds on the Internet? | PCMag
Overall, the chapter's core concepts remain relevant, but the scale has grown
immensely, underscoring the need for advanced analytics and ethical data
management in 2025.

The Four V's of Big Data


The Four V's—Volume, Velocity, Variety, and Veracity—represent the fundamental
characteristics that define Big Data, as popularized by IBM and Gartner. These
dimensions highlight why traditional data processing methods fall short and why
specialized tools and strategies are needed for handling massive datasets.
Originally, Gartner introduced three V's in the early 2000s (Volume, Velocity,
Variety), but Veracity was later added to address data quality issues.
[Link]
Big Data #1 - the 4V's of Big Data - Vamsi Talks Tech
While the PDF from your query (dated around 2014) provides a concise overview,
here's a more detailed explanation based on established definitions, including
real-world examples and implications as of 2025. These V's help organizations
extract value from data for better decision-making, insights, and process
optimization.
1. Volume: The Scale of Data
Volume refers to the sheer amount of data generated and stored, often
measured in terabytes (TB), petabytes (PB), exabytes (EB), or even zettabytes
(ZB). In the PDF, it's described as "terabytes to exabytes of existing data to
process." This dimension emphasizes that Big Data involves datasets too large
for conventional databases to handle efficiently.
 Detailed Explanation: The explosion in data volume is driven by sources
like social media, IoT devices, sensors, and digital transactions. For
instance, as of 2025, the global datasphere is estimated at over 180 ZB,
with projections to reach 300 ZB by 2030. Traditional systems struggle
because processing such volumes requires distributed storage solutions
like Hadoop or cloud-based systems (e.g., AWS S3 or Google Cloud
Storage). High volume demands scalable infrastructure to avoid
bottlenecks in storage and analysis.
 Examples:
o Social media platforms like X (formerly Twitter) generate billions of
posts daily.
o Healthcare systems store petabytes of patient records, imaging
data, and genomic sequences.
o E-commerce giants like Amazon process exabytes of transaction
and user behavior data.
 Challenges and Solutions: Managing volume involves data
compression, archiving, and using NoSQL databases. Without proper
handling, organizations face high costs and inefficiency.
2. Velocity: The Speed of Data
Velocity describes the rate at which data is generated, processed, and analyzed.
The PDF notes it as "analysis of streaming data, milliseconds to seconds to
respond," highlighting the need for real-time or near-real-time processing.
 Detailed Explanation: In today's fast-paced world, data streams in
continuously from sources like stock tickers, social feeds, and sensors.
Velocity isn't just about generation speed but also the timeliness of
insights—delays can render data obsolete. For example, high-velocity data
requires tools that can handle streaming analytics, such as Apache Kafka
or Spark Streaming, to process information in real-time.
 Examples:
o Financial trading systems analyze market data in microseconds to
execute trades.
o Ride-sharing apps like Uber process GPS data from millions of
devices in seconds for route optimization.
o Social media monitoring during events (e.g., elections or disasters)
tracks sentiment in real-time.
 Challenges and Solutions: High velocity can overwhelm systems,
leading to latency. Solutions include edge computing (processing data
closer to the source) and machine learning models for predictive analytics.
3. Variety: The Diversity of Data
Variety encompasses the different types and formats of data, from structured
(e.g., databases) to unstructured (e.g., text, images) and semi-structured (e.g.,
JSON). The PDF defines it as "different forms of data: structured, unstructured,
text, multimedia."
 Detailed Explanation: About 80-90% of Big Data is unstructured, making
integration complex. Variety requires flexible schemas to handle disparate
sources without losing context. Tools like data lakes (e.g., Azure Data Lake)
allow storage of raw data in native formats, while ETL (Extract, Transform,
Load) processes unify them for analysis.
 Examples:
o A retail company might combine structured sales data with
unstructured customer reviews, social media posts, and video
surveillance.
o Autonomous vehicles process varied inputs like LiDAR scans
(structured), camera footage (unstructured), and sensor logs (semi-
structured).
o Media companies analyze text articles, audio podcasts, and video
content together.
 Challenges and Solutions: Incompatibility between formats can lead to
silos. Multi-model databases (e.g., MongoDB) and AI-driven data
integration help harmonize variety.
4. Veracity: The Reliability of Data
Veracity addresses the uncertainty, accuracy, and trustworthiness of data. As per
the PDF, it's about "managing the reliability and predictability of inherently
imprecise data types."
 Detailed Explanation: Not all data is clean or reliable—issues like bias,
noise, duplicates, or incompleteness can skew results. Veracity ensures
data quality through validation, cleansing, and governance. It's often
considered the most critical V because poor veracity leads to flawed
insights, with estimates showing that bad data costs businesses trillions
annually. In 2025, with AI-generated content proliferating, veracity also
involves detecting deepfakes and ensuring ethical sourcing.
 Examples:
o In marketing, verifying social media data to avoid bots influencing
campaigns.
o Healthcare analytics filtering out erroneous sensor readings from
wearables.
o Climate modeling cross-validating satellite data with ground sensors
for accuracy.
 Challenges and Solutions: Uncertainty arises from human error or
malicious inputs. Data quality tools (e.g., Talend or Informatica) and
blockchain for provenance help maintain veracity.
These Four V's are interconnected; for instance, high variety can amplify veracity
issues, while velocity demands quick volume management. Organizations
leveraging them effectively gain competitive advantages, such as personalized
services or predictive maintenance.

Summary of Chapter 8: Mobile Database


This chapter is a slide deck presentation on mobile databases, authored by Nor
Intan Shafini binti Nasaruddin from the Faculty of Computer and Mathematical
Sciences. It introduces the concept, explains its necessity, architecture,
characteristics, applications, and limitations. The content appears dated around
2012-2014 (based on references like smartphone stats from 2012), focusing on
foundational aspects of mobile data management in wireless environments. The
presentation emphasizes how mobile databases enable data access and
processing on the go, separate from central servers, while handling challenges
like connectivity issues.
Definition and Introduction
A mobile database is defined as a database accessible via mobile computing
devices over wireless networks. Key features include:
 Physical separation from the central database server.
 Ability to communicate with central servers or other mobile clients
remotely.
 Residence on mobile devices (e.g., smartphones, laptops).
 Capability to handle local queries offline, without constant connectivity.
This setup allows for decentralized data management, making it ideal for mobile
users.
Need for Mobile Databases
The presentation highlights the growing demand driven by mobility and
technology advancements (circa 2012 data):
 Over 1 billion smartphones in use worldwide, increasing data generation
on the move.
 Need for real-time data collection as events occur.
 Support for "3A" business: Anytime, Anywhere, Any Device.
 Powerful, lightweight devices and low-cost connectivity pave the way for
data-driven mobile apps.
 Users require offline work capabilities due to unreliable connections,
enabling access to any data from anywhere.
A flowchart on page 7 illustrates the evolution: From mobile devices and apps to
mobile DBMS (Database Management Systems) that support disconnected
operations.
Architecture
Mobile databases typically follow two models:
1. Client-Server Model (Page 9):
o Central server and database connect to a central DBMS.

o Mobile DBMS on devices (e.g., laptops, phones) interacts with


mobile DBs.
o Data flows from central to mobile for synchronization.

2. Peer-to-Peer Model (Page 10):


o Devices communicate directly via wireless networks (e.g., through a
tower).
o No central server; mobile DBMS and DBs on each peer device.

o Enables direct data sharing between mobiles.

These architectures support hybrid operations: local processing offline and


syncing when connected.
Characteristics of Mobile Environments
Mobile databases operate in constrained settings, with key traits:
 Restricted bandwidth of wireless networks.
 Limited power supply (e.g., battery life).
 Mobility (constant movement affects connectivity).
 Disconnections (frequent interruptions).
 Limited resources (e.g., storage, processing power on devices).
These require optimized designs for efficiency and resilience.
Applications
Mobile databases are applied across sectors for real-time, on-the-go data
handling:
 Business: Access to customer, competitor, and market trend info;
salespersons update sales/customer data remotely.
 Public Sector: US Army uses it for inventory and readiness tracking to
reduce logistical costs.
 Medical: Doctors retrieve patient medical histories from anywhere.
[Link]
[Link]
 News: Reporters update databases in real-time.
A mind map on page 14 connects these applications, showing interconnected
use cases.
Management Issues and Limitations
Management focuses on overcoming environmental constraints. Limitations
include:
1. Limited wireless bandwidth.
2. Slow wireless communication speeds.
3. Limited energy sources (battery power).
4. Difficulty in making devices theft-proof.
5. Vulnerability to physical damage.
6. Lower security compared to stationary systems.
These issues demand strategies like data compression, power-efficient queries,
encryption, and robust syncing protocols.
The presentation ends with a "Thank you" slide, reinforcing the motivational
"Keep Calm and continue your database journey" theme.
Relation to Big Data (From Previous Chapter)
This chapter complements the Big Data discussion (also Chapter 8, likely from
the same course). Mobile databases contribute to Big Data's "Four V's":
 Volume and Variety: Mobile devices generate vast unstructured data
(e.g., location, sensor inputs), feeding into Big Data ecosystems.
 Velocity: Real-time mobile updates support streaming analysis.
 Veracity: Disconnections and mobility introduce data uncertainty,
requiring verification.
Together, they highlight how mobile tech amplifies data deluge, with mobile DBs
as edge tools for Big Data collection.
Updates as of December 16, 2025
The core concepts remain relevant, but mobile databases have evolved
significantly since the 2010s, driven by 5G/6G, edge computing, AI, and cloud
integration. Global smartphone users now exceed 7 billion, with daily data
generation in zettabytes from mobiles alone.
Modern Examples and Technologies
 Popular Mobile DBs: SQLite (embedded in Android/iOS apps), Realm
(now MongoDB Realm for real-time sync), Firebase Realtime Database
(cloud-backed for offline support), and Couchbase Mobile (for peer-to-peer
syncing).
 Advancements: 5G enables ultra-low latency (under 1ms), addressing
bandwidth/speed limits. Edge computing (e.g., via AWS Wavelength)
processes data on-device or nearby, reducing disconnections.
 Security and Limitations: Modern solutions use end-to-end encryption
(e.g., Signal's protocol) and biometric auth to counter theft/vulnerability.
Battery tech (e.g., solid-state batteries) extends life, but energy efficiency
remains key—AI optimizes queries to save power.
 Applications Today: In healthcare, apps like Epic MyChart use mobile
DBs for secure patient data access. Logistics (e.g., UPS) employs them for
real-time tracking. AI-driven apps (e.g., TikTok) handle massive user data
offline before syncing.
 Growth Stats: By 2025, mobile data traffic is ~200 exabytes/month
globally, per Ericsson Mobility Report. IoT integration (e.g., wearables)
adds complexity, with hybrid cloud-mobile setups dominant.
Overall, mobile databases are now integral to Big Data pipelines, with emphasis
on privacy (e.g., GDPR compliance) and AI for predictive syncing.
1. What is a mobile database? How does it differ from a
traditional/centralized database?
Students love this definitional question—it's foundational and often worth many
marks.
 Answer: A mobile database is a database that resides on mobile devices
(e.g., smartphones, tablets) and can connect to a central server over a
wireless network. It is physically separate from the central database
server, allows communication with the server or other mobiles remotely,
and supports local queries without connectivity (offline mode).
Key differences from a traditional centralized database:

Traditional/Centralized
Aspect Mobile Database
Database

On mobile device
Location Single central server
(distributed/edge)

Supports offline/disconnected
Connectivity Requires constant connection
operations

Limited (battery, bandwidth, Unlimited (high power, wired


Resources
storage) networks)

Synchronizati Needs periodic syncing with


Real-time, no syncing needed
on central DB

More vulnerable (theft, Easier to secure in one


Security
physical damage) location

[Link]

[Link]
In 2025, modern examples like Firebase or Realm handle automatic syncing
seamlessly.
2. Explain the architecture of mobile databases. Compare client-server
and peer-to-peer models.
This often includes a diagram request in exams.
 Answer: Mobile databases use two main architectures:
o Client-Server: Mobile devices (clients) sync with a central server.
Common for most apps (e.g., sales data syncing to HQ).
o Peer-to-Peer (P2P): Devices communicate directly without a
central server—useful for field teams sharing data in low-
connectivity areas.
[Link]

[Link]
Today, hybrid models dominate (e.g., Couchbase Mobile combines both with
edge-to-cloud sync).
3. What are the characteristics and limitations of mobile
environments/databases?
 Characteristics (from slides):
o Restricted wireless bandwidth

o Limited power supply (battery)

o Mobility (frequent location changes)

o Frequent disconnections

o Limited resources (CPU, storage)

 Limitations:
1. Limited bandwidth and slow speeds

2. Battery constraints

3. Theft/vulnerability to physical damage

4. Lower security (easier to compromise on-device)

In 2025, 5G/6G mitigates bandwidth issues, and encrypted embedded DBs (e.g.,
SQLite with SQLCipher) improve security.
4. Why do we need mobile databases? Give real-world applications.
 Reasons:
o Mobility and "3A" (Anytime, Anywhere, Any device)
o Real-time data collection

o Offline access in poor connectivity

o Growth in smartphones (slides mention 1 billion in 2012; now over 7


billion in 2025!)
 Applications (from slides + modern):
o Business: Sales reps updating CRM on-the-go

o Medical: Doctors accessing patient records remotely

o Public Sector: Military inventory tracking

o News: Real-time reporting

o Modern: Ride-sharing (Uber's offline maps), healthcare apps


(telemedicine records)

[Link]

[Link]
5. Give examples of mobile database systems.
Common short-answer question.
 Classic: Sybase, Oracle Lite (from older texts)
 Current (2025):
o SQLite (embedded, offline-first—used in Android/iOS)

o Realm/MongoDB Realm (real-time sync)

o Firebase Realtime Database/Firestore (cloud-synced, offline


support)
o Couchbase Mobile (P2P + sync)

These connect to your Big Data chapter: Mobile DBs feed unstructured data (e.g.,
location, sensors) into Big Data pipelines.
Other frequent ones: "How does synchronization work?" or "How do mobile DBs
handle disconnections?" (Answer: Local caching + conflict resolution on
reconnect).

You might also like