A
Micro-Project Report On
“Big Data”
Submitted By
Pathar Jinakal & 236200316158
Semester – 3rd Division – 3b
Department Of Information Technology Government
Polytechnic Rajkot
October - 2024
1
CRETIFICATE
This is to certify that this micro-project report
entitled
“BIG DATA” submitted by “Pathar Jinkal ” ,
“236200316158”, “ 3rd sem” has satisfactorily
completed her term work in the subject of SEMINAR.
Date : 14-9-2024 Guided By: Amit Bhalaodiya
Place : Rajkot
INDEX
2
Introduction About Topic 4
history of technology in Big Data And technology tools
5
Working In Big Data 6
Use Of Big Data
7
Advantages In Big Data
8
Limitation And Disadvantages
9
Future Expansion In Big Data
10
3
Introduction About Topic
Big data refers to extremely large datasets that are too co
mplex or vast for traditional data processing tools to handl
e.
This data can come from various sources, like social media
, sensors, transactions, and more.
The key characteristics of big data are often referred to as
the three V's: volume (the amount of data), velocity (the s
peed at which data is generated), and variety (the differen
t types of data).
The goal of big data analytics is to extract meaningful insi
ghts and information from these massive datasets.
It involves using advanced tools and techniques like machi
ne learning, data mining, and statistical analysis to proces
s and analyze the data.
Here’s a quick intro to big data
Volume: The sheer amount of data generated, measured i
n petabytes or exabytes.
Velocity: The speed at which new data is created and mo
ved.
Variety: The different forms of data, ranging from structur
ed to unstructured.
Value: The potential insights and benefits that can be extr
acted from the data.
Veracity: The accuracy and reliability of the data.
4
Here’s a brief history of technology i
n big data
1940s: The concept of "information explosion" emerged,
highlighting the rapid growth of data1.
1960s-1970s: Mainframe computers were introduced, sig
nificantly increasing data storage capacities2.
1980s: The development of relational databases allowed f
or more efficient data organization and retrieval.
1990s: The internet boom led to an exponential increase i
n data generation and the need for better data manageme
nt tools.
2000s: Hadoop, an open-source framework for distributed
storage and processing of large datasets, was developed2.
2010s: Big data analytics became mainstream, with adva
ncements in machine learning, AI, and cloud computing.
2020s: Real data processing and the Internet of Things (Io
T) have further expanded the scope and capabilities of big
data technologies.
Here are some key big data technology t
ools
Apache Hadoop: An open-source framework for distribut
ed storage and processing of large
datasets.
Apache Spark: An open-source analytics engine for large
-scale data processing.
Apache Flink: A stream processing framework for real-
time data processing.
Google Cloud Platform: Provides various big data servic
es, including storage, analytics, and machine learning.
MongoDB: A NoSQL database designed for handling large
volumes of unstructured data.
Sisense: A business intelligence tool that allows for data i
ntegration and visualization.
RapidMiner: A data science platform for data preparation
, machine learning, and predictive analytics.
5
Working In Big Data
Data Collection: Gathering large amounts of data from v
arious sources like social media, sensors, transactions, an
d more.
Data Storage: Using scalable storage solutions like Hado
op Distributed File System (HDFS) or cloud storage to stor
e the vast data sets.
Data Processing: Leveraging tools like Apache Spark, Ha
doop MapReduce, or real-time processing frameworks like
Apache Flink to process and manage the data.
Data Analysis: Applying analytical techniques and machi
ne learning algorithms to uncover patterns, correlations, a
nd insights. This can be done using tools like R, Python, or
RapidMiner.
Data Visualization: Presenting the data in an understan
dable format using visualization tools like Tableau, Power
BI, or Sisense. This helps in making data-driven decisions.
These steps form a cycle where the insights gained can influenc
e further data collection and analysis, continuously improving t
he decision-making process.
6
Use Of Big Data
1. Healthcare: Analyzing large datasets from electronic heal
th records to improve patient care, predict disease outbre
aks, and optimize treatment plans.
2. Finance: Detecting fraud, managing risk, and guiding inve
stment strategies by analyzing transaction data and mark
et trends.
3. Retail: Understanding customer behavior, personalizing m
arketing efforts, and optimizing supply chains by analyzing
sales data and customer feedback.
4. Telecommunications: Managing network performance, i
mproving customer service, and predicting maintenance n
eeds by analyzing usage patterns and service data.
5. Transportation: Optimizing routes, managing fleets, and
predicting maintenance issues by analyzing vehicle data a
nd traffic patterns.
6. Manufacturing: Enhancing production efficiency, improvi
ng product quality, and predicting equipment failures by a
nalyzing production data and machine logs.
7. Public Sector: Enhancing public services, improving urba
n planning, and managing emergencies by analyzing data
from various public sources.
7
Advantages In Big Data
1. Better Decision Making: Data-driven insights help organ
izations make more informed and accurate decisions.
2. Enhanced Customer Experiences: Personalizing custo
mer interactions and offerings based on behavior and pref
erences.
3. Operational Efficiency: Streamlining processes and redu
cing costs through data analysis and automation.
4. Innovation and Product Development: Identifying tren
ds and customer needs to create new products and servic
es.
5. Predictive Analytics: Anticipating future trends and beh
aviors to stay ahead of the competition.
6. Risk Management: Detecting and mitigating risks more
effectively by analyzing patterns and anomalies
7. Competitive Advantage: Leveraging data to gain insight
s that competitors may not have.
8
Limitation And Disadvantage
1. Data Quality: Ensuring the accuracy and reliability of ma
ssive datasets can be challenging.
2. Complexity: Handling and analyzing big data requires so
phisticated tools and expertise.
3. Privacy Concerns: Collecting and processing large amou
nts of data can raise significant privacy and ethical issues.
4. Cost: Infrastructure and tools for storing and processing bi
g data can be expensive.
5. Data Overload: The sheer volume of data can make it di
fficult to extract meaningful insights without proper techni
ques.
6. Security Risks: Large datasets can be vulnerable to secu
rity breaches and cyber-attacks.
7. Integration: Integrating big data with existing systems a
nd workflows can be complex and time-consuming.
9
Future Expansion in Big Data
1. Artificial Intelligence (AI) and Machine Learning (ML
): AI and ML will play an increasingly important role in big
data analysis, helping businesses quickly and accurately
make sense of vast amounts of data1.
2. Edge Computing: Processing data closer to the source, r
ather than sending it to a centralized location, will reduce l
atency and improve real-time data analysis1.
3. Internet of Things (IoT): The proliferation of IoT devices
will generate even more data, providing deeper insights a
nd enabling smarter decision-making2.
4. Cloud Computing: Widespread migration to the cloud wil
l offer scalable and flexible solutions for storing and proces
sing big data3.
5. Data Observability: Enhanced tools for monitoring and
managing data quality will ensure more reliable and action
able insights3.
6. Advanced Analytics: New techniques and tools will conti
nue to evolve, allowing for more sophisticated analysis an
d predictive modeling.
10
11