DATA SCIENCE
UNIT V
UNIT-V Case Study – Prediction of Disease - Setting research goals - Data retrieval –
preparation - exploration - Disease profiling - presentation and automation.
Data science has become popular in the last few years due to its successful application in
making business decisions. Data scientists have been using data science techniques to
solve challenging real-world issues in healthcare, agriculture, manufacturing, automotive,
and many more. For this purpose, a data enthusiast needs to stay updated with the latest
technological advancements in AI.
Examples of Data Science Case Studies
• Hospitality: Airbnb focuses on growth by analyzing customer voice using data
science. Qantas uses predictive analytics to mitigate losses
• Healthcare: Novo Nordisk is Driving innovation with NLP. AstraZeneca harnesses
data for innovation in medicine
• Covid 19: Johnson and Johnson uses data science to fight the Pandemic
• E-commerce: Amazon uses data science to personalize shopping experiences and
improve customer satisfaction
• Supply chain management: UPS optimizes supply chain with big data analytics
• Meteorology: IMD leveraged data science to achieve a record 1.2m evacuation
before cyclone ''Fani''
• Entertainment Industry: Netflix uses data science to personalize the content and
improve recommendations. Spotify uses big data to deliver a rich user experience for
online music streaming
• Banking and Finance: HDFC utilizes Big Data Analytics to increase income and
enhance the banking experience
Data Science Case Studies [For Various Industries]
1. Data Science in Hospitality Industry
In the hospitality sector, data analytics assists hotels in better pricing strategies, customer
analysis, brand marketing, tracking market trends, and many more.
Airbnb focuses on growth by analysing customer voice using data science. A famous
example in this sector is the unicorn ''Airbnb'', a startup that focussed on data science
early to grow and adapt to the market faster. This company witnessed a 43000 percent
hypergrowth in as little as five years using data science. They included data science
techniques to process the data, translate this data for better understanding the voice of the
customer, and use the insights for decision making. They also scaled the approach to cover
all aspects of the organization. Airbnb uses statistics to analyse and aggregate individual
DATA SCIENCE
UNIT V
experiences to establish trends throughout the community. These analyzed trends using
data science techniques impact their business choices while helping them grow further.
Travel industry and data science
Predictive analytics benefits many parameters in the travel industry. These companies can
use recommendation engines with data science to achieve higher personalization and
improved user interactions. They can study and cross-sell products by recommending
relevant products to drive sales and increase revenue. Data science is also employed
in analysing social media posts for sentiment analysis, bringing invaluable travel-related
insights. Whether these views are positive, negative, or neutral can help these agencies
understand the user demographics, the expected experiences by their target audiences, and
so on. These insights are essential for developing aggressive pricing strategies to draw
customers and provide better customization to customers in the travel packages and allied
services. Travel agencies like Expedia and [Link] use predictive analytics to create
personalized recommendations, product development, and effective marketing of their
products. Not just travel agencies but airlines also benefit from the same approach.
Airlines frequently face losses due to flight cancellations, disruptions, and delays. Data
science helps them identify patterns and predict possible bottlenecks, thereby effectively
mitigating the losses and improving the overall customer traveling experience.
How Qantas uses predictive analytics to mitigate losses
Qantas, one of Australia's largest airlines, leverages data science to reduce losses caused
due to flight delays, disruptions, and cancellations. They also use it to provide a better
traveling experience for their customers by reducing the number and length of delays
caused due to huge air traffic, weather conditions, or difficulties arising in operations.
Back in 2016, when heavy storms badly struck Australia's east coast, only 15 out of 436
Qantas flights were cancelled due to their predictive analytics-based system against their
competitor Virgin Australia, which witnessed 70 cancelled flights out of 320.
2. Data Science in Healthcare
The Healthcare sector is immensely benefiting from the advancements in AI. Data
science, especially in medical imaging, has been helping healthcare professionals come up
with better diagnoses and effective treatments for patients. Similarly, several advanced
healthcare analytics tools have been developed to generate clinical insights for improving
patient care. These tools also assist in defining personalized medications for patients
reducing operating costs for clinics and hospitals. Apart from medical imaging or
computer vision, Natural Language Processing (NLP) is frequently used in the
healthcare domain to study the published textual research data.
DATA SCIENCE
UNIT V
A. Pharmaceutical
Driving innovation with NLP: Novo Nordisk. Novo Nordisk uses the Linguamatics NLP
platform from internal and external data sources for text mining purposes that include
scientific abstracts, patents, grants, news, tech transfer offices from universities
worldwide, and more. These NLP queries run across sources for the key therapeutic areas
of interest to the Novo Nordisk R&D community. Several NLP algorithms have been
developed for the topics of safety, efficacy, randomized controlled trials, patient
populations, dosing, and devices. Novo Nordisk employs a data pipeline to capitalize the
tools' success on real-world data and uses interactive dashboards and cloud services to
visualize this standardized structured information from the queries for exploring
commercial effectiveness, market situations, potential, and gaps in the product
documentation. Through data science, they are able to automate the process of generating
insights, save time and provide better insights for evidence-based decision making.
B. BioTech
How AstraZeneca harnesses data for innovation in medicine. AstraZeneca is a globally
known biotech company that leverages data using AI technology to discover and deliver
newer effective medicines faster. Within their R&D teams, they are using AI to decode the
big data to understand better diseases like cancer, respiratory disease, and heart, kidney,
and metabolic diseases to be effectively treated. Using data science, they can identify new
targets for innovative medications. In 2021, they selected the first two AI-generated drug
targets collaborating with BenevolentAI in Chronic Kidney Disease and Idiopathic
Pulmonary Fibrosis.
Data science is also helping AstraZeneca redesign better clinical trials, achieve
personalized medication strategies, and innovate the process of developing new medicines.
Their Center for Genomics Research uses data science and AI to analyze around two
million genomes by 2026. Apart from this, they are training their AI systems to check
these images for disease and biomarkers for effective medicines for imaging purposes.
This approach helps them analyze samples accurately and more effortlessly. Moreover, it
can cut the analysis time by around 30%.
AstraZeneca also utilizes AI and machine learning to optimize the process at different
stages and minimize the overall time for the clinical trials by analyzing the clinical trial
data. Summing up, they use data science to design smarter clinical trials, develop
innovative medicines, improve drug development and patient care strategies, and many
more.
C. Wearable Technology
DATA SCIENCE
UNIT V
Wearable technology is a multi-billion-dollar industry. With an increasing awareness
about fitness and nutrition, more individuals now prefer using fitness wearables to track
their routines and lifestyle choices.
Fitness wearables are convenient to use, assist users in tracking their health, and encourage
them to lead a healthier lifestyle. The medical devices in this domain are beneficial since
they help monitor the patient's condition and communicate in an emergency situation. The
regularly used fitness trackers and smartwatches from renowned companies like Garmin,
Apple, FitBit, etc., continuously collect physiological data of the individuals wearing
them. These wearable providers offer user-friendly dashboards to their customers
for analyzing and tracking progress in their fitness journey.
3. Covid 19 and Data Science
In the past two years of the Pandemic, the power of data science has been more evident
than ever. Different pharmaceutical companies across the globe could synthesize Covid
19 vaccines by analyzing the data to understand the trends and patterns of the outbreak.
Data science made it possible to track the virus in real-time, predict patterns, devise
effective strategies to fight the Pandemic, and many more.
How Johnson and Johnson uses data science to fight the Pandemic
The data science team at Johnson and Johnson leverages real-time data to track the
spread of the virus. They built a global surveillance dashboard (granulated to county level)
that helps them track the Pandemic's progress, predict potential hotspots of the virus, and
narrow down the likely place where they should test its investigational COVID-19 vaccine
candidate. The team works with in-country experts to determine whether official numbers
are accurate and find the most valid information about case numbers, hospitalizations,
mortality and testing rates, social compliance, and local policies to populate this
dashboard. The team also studies the data to build models that help the company identify
groups of individuals at risk of getting affected by the virus and explore effective
treatments to improve patient outcomes.
4. Data Science in E-commerce
In the e-commerce sector, big data analytics can assist in customer analysis, reduce
operational costs, forecast trends for better sales, provide personalized shopping
experiences to customers, and many more.
Amazon uses data science to personalize shopping experiences and improve customer
satisfaction. Amazon is a globally leading eCommerce platform that offers a wide range of
online shopping services. Due to this, Amazon generates a massive amount of data that can
DATA SCIENCE
UNIT V
be leveraged to understand consumer behavior and generate insights on competitors'
strategies. Amazon uses its data to provide recommendations to its users on different
products and services. With this approach, Amazon is able to persuade its consumers into
buying and making additional sales. This approach works well for Amazon as it earns 35%
of the revenue yearly with this technique. Additionally, Amazon collects consumer data
for faster order tracking and better deliveries.
Similarly, Amazon's virtual assistant, Alexa, can converse in different languages; uses
speakers and a camera to interact with the users. Amazon utilizes the audio commands
from users to improve Alexa and deliver a better user experience.
5. Data Science in Supply Chain Management
Predictive analytics and big data are driving innovation in the Supply chain domain. They
offer greater visibility into the company operations, reduce costs and overheads,
forecasting demands, predictive maintenance, product pricing, minimize supply chain
interruptions, route optimization, fleet management, drive better performance, and
more.
Optimizing supply chain with big data analytics: UPS
UPS is a renowned package delivery and supply chain management company. With
thousands of packages being delivered every day, on average, a UPS driver makes about
100 deliveries each business day. On-time and safe package delivery are crucial to UPS's
success. Hence, UPS offers an optimized navigation tool ''ORION'' (On-Road Integrated
Optimization and Navigation), which uses highly advanced big data processing algorithms.
This tool for UPS drivers provides route optimization concerning fuel, distance, and time.
UPS utilizes supply chain data analysis in all aspects of its shipping process. Data about
packages and deliveries are captured through radars and sensors. The deliveries and routes
are optimized using big data systems. Overall, this approach has helped UPS save 1.6
million gallons of gasoline in transportation every year, significantly reducing delivery
costs.
6. Data Science in Meteorology
Weather prediction is an interesting application of data science. Businesses like aviation,
agriculture and farming, construction, consumer goods, sporting events, and many more
are dependent on climatic conditions. The success of these businesses is closely tied to the
weather, as decisions are made after considering the weather predictions from the
meteorological department.
DATA SCIENCE
UNIT V
Besides, weather forecasts are extremely helpful for individuals to manage their allergic
conditions. One crucial application of weather forecasting is natural disaster prediction and
risk management.
Weather forecasts begin with a large amount of data collection related to the current
environmental conditions (wind speed, temperature, humidity, clouds captured at a
specific location and time) using sensors on IoT (Internet of Things) devices and satellite
imagery. This gathered data is then analyzed using the understanding of atmospheric
processes, and machine learning models are built to make predictions on upcoming
weather conditions like rainfall or snow prediction. Although data science cannot help
avoid natural calamities like floods, hurricanes, or forest fires. Tracking these natural
phenomena well ahead of their arrival is beneficial. Such predictions allow governments
sufficient time to take necessary steps and measures to ensure the safety of the population.
IMD leveraged data science to achieve a record 1.2m evacuation before cyclone
''Fani''
Most data scientist’s responsibilities rely on satellite images to make short-term
forecasts, decide whether a forecast is correct, and validate models. Machine Learning is
also used for pattern matching in this case. It can forecast future weather conditions if it
recognizes a past pattern. When employing dependable equipment, sensor data is helpful
to produce local forecasts about actual weather models. IMD used satellite pictures to
study the low-pressure zones forming off the Odisha coast (India). In April 2019, thirteen
days before cyclone ''Fani'' reached the area, IMD (India Meteorological Department)
warned that a massive storm was underway, and the authorities began preparing for safety
measures.
It was one of the most powerful cyclones to strike India in the recent 20 years, and a
record 1.2 million people were evacuated in less than 48 hours, thanks to the power of data
science.
7. Data Science in the Entertainment Industry
Due to the Pandemic, demand for OTT (Over-the-top) media platforms has grown
significantly. People prefer watching movies and web series or listening to the music of
their choice at leisure in the convenience of their homes. This sudden growth in demand
has given rise to stiff competition. Every platform now uses data analytics in different
capacities to provide better-personalized recommendations to its subscribers and improve
user experience.
How Netflix uses data science to personalize the content and improve
recommendations
DATA SCIENCE
UNIT V
Netflix is an extremely popular internet television platform with streamable content
offered in several languages and caters to various audiences. In 2006, when Netflix entered
this media streaming market, they were interested in increasing the efficiency of their
existing ''Cinematch'' platform by 10% and hence, offered a prize of $1 million to the
winning team. This approach was successful as they found a solution developed by the
BellKor team at the end of the competition that increased prediction accuracy by 10.06%.
Over 200 work hours and an ensemble of 107 algorithms provided this result. These
winning algorithms are now a part of the Netflix recommendation system.
Netflix also employs Ranking Algorithms to generate personalized recommendations of
movies and TV Shows appealing to its users.
Spotify uses big data to deliver a rich user experience for online music streaming
Personalized online music streaming is another area where data science is being
used. Spotify is a well-known on-demand music service provider launched in 2008, which
effectively leveraged big data to create personalized experiences for each user. It is a huge
platform with more than 24 million subscribers and hosts a database of nearly 20million
songs; they use the big data to offer a rich experience to its users. Spotify uses this big data
and various algorithms to train machine learning models to provide personalized content.
Spotify offers a "Discover Weekly" feature that generates a personalized playlist of fresh
unheard songs matching the user's taste every week. Using the Spotify "Wrapped" feature,
users get an overview of their most favorite or frequently listened songs during the entire
year in December. Spotify also leverages the data to run targeted ads to grow its business.
Thus, Spotify utilizes the user data, which is big data and some external data, to deliver a
high-quality user experience.
8. Data Science in Banking and Finance
Data science is extremely valuable in the Banking and Finance industry. Several high
priority aspects of Banking and Finance like credit risk modeling (possibility of repayment
of a loan), fraud detection (detection of malicious or irregularities in transactional patterns
using machine learning), identifying customer lifetime value (prediction of bank
performance based on existing and potential customers), customer segmentation (customer
profiling based on behavior and characteristics for personalization of offers and services).
Finally, data science is also used in real-time predictive analytics (computational
techniques to predict future events).
How HDFC utilizes Big Data Analytics to increase revenues and enhance the banking
experience
DATA SCIENCE
UNIT V
One of the major private banks in India, HDFC Bank, was an early adopter of AI. It started
with Big Data analytics in 2004, intending to grow its revenue and understand its
customers and markets better than its competitors. Back then, they were trendsetters by
setting up an enterprise data warehouse in the bank to be able to track the differentiation to
be given to customers based on their relationship value with HDFC Bank. Data science
and analytics have been crucial in helping HDFC bank segregate its customers and offer
customized personal or commercial banking services. The analytics engine and SaaS use
have been assisting the HDFC bank in cross-selling relevant offers to its customers. Apart
from the regular fraud prevention, it assists in keeping track of customer credit histories
and has also been the reason for the speedy loan approvals offered by the bank.
9. Data Science in Urban Planning and Smart Cities
Data Science can help the dream of smart cities come true! Everything, from traffic flow
to energy usage, can get optimized using data science techniques. You can use the data
fetched from multiple sources to understand trends and plan urban living in a sorted
manner.
The significant data science case study is traffic management in Pune city. The city
controls and modifies its traffic signals dynamically, tracking the traffic flow. Real-time
data gets fetched from the signals through cameras or sensors installed. Based on this
information, they do the traffic management. With this proactive approach, the traffic and
congestion situation in the city gets managed, and the traffic flow becomes sorted. A
similar case study is from Bhubaneswar, where the municipality has platforms for the
people to give suggestions and actively participate in decision-making. The government
goes through all the inputs provided before making any decisions, making rules or
arranging things that their residents actually need.
10. Data Science in Agricultural Yield Prediction
Have you ever wondered how helpful it can be if you can predict your agricultural yield?
That is exactly what data science is helping farmers with. They can get information about
the number of crops they can produce in a given area based on different environmental
factors and soil types. Using this information, the farmers can make informed decisions
about their yield and benefit the buyers and themselves in multiple ways.
DATA SCIENCE
UNIT V
Farmers across the globe and overseas use various data science techniques to understand
multiple aspects of their farms and crops. A famous example of data science in the
agricultural industry is the work done by Farmers Edge. It is a company in Canada that
takes real-time images of farms across the globe and combines them with related data. The
farmers use this data to make decisions relevant to their yield and improve their produce.
Similarly, farmers in countries like Ireland use satellite-based information to ditch
traditional methods and multiply their yield strategically.
11. Data Science in the Transportation Industry
Transportation keeps the world moving around. People and goods commute from one
place to another for various purposes, and it is fair to say that the world will come to a
standstill without efficient transportation. That is why it is crucial to keep the
transportation industry in the most smoothly working pattern, and data science helps a lot
in this. In the realm of technological progress, various devices such as traffic sensors,
monitoring display systems, mobility management devices, and numerous others
have emerged.
Many cities have already adapted to the multi-modal transportation system. They use GPS
trackers, geo-locations and CCTV cameras to monitor and manage their transportation
system. Uber is the perfect case study to understand the use of data science in the
transportation industry. They optimize their ride-sharing feature and track the delivery
routes through data analysis. Their data science approach enabled them to serve more than
100 million users, making transportation easy and convenient. Moreover, they also use the
data they fetch from users daily to offer cost-effective and quickly available rides.
12. Data Science in the Environmental Industry
Increasing pollution, global warming, climate changes and other poor environmental
impacts have forced the world to pay attention to environmental industry. Multiple
initiatives are being taken across the globe to preserve the environment and make the
DATA SCIENCE
UNIT V
world a better place. Though the industry recognition and the efforts are in
the initial stages, the impact is significant, and the growth is fast.
The popular use of data science in the environmental industry is by NASA and other
research organizations worldwide. NASA gets data related to the current climate
conditions, and this data gets used to create remedial policies that can make a
difference. Another way in which data science is actually helping researchers is they can
predict natural disasters well before time and save or at least reduce the potential damage
considerably. A similar case study is with the World Wildlife Fund. They use data science
to track data related to deforestation and help reduce the illegal cutting of trees. Hence, it
helps preserve the environment.
The Power of Data in Healthcare
The healthcare industry generates vast amounts of data every day. Electronic health records
(EHRs), medical imaging, genetic information, and wearable devices contribute to the
enormous data repository. Leveraging this wealth of data with data science techniques opens
the door to better patient care.
1. Early Disease Detection:Data science plays a pivotal role in early disease
detection. Machine learning models can analyze patient data and identify subtle
patterns that human physicians might overlook. For instance, predictive analytics
can help identify potential diabetes patients by analyzing glucose levels and
lifestyle data, enabling early intervention and preventive measures.
2. Personalized Treatment Plans:One of the key benefits of data science in
healthcare is the ability to tailor treatment plans to individual patients. Machine
learning algorithms can consider a patient's medical history, genetics, lifestyle,
and more to design a treatment plan optimized for their specific needs. This not
only improves patient outcomes but also reduces adverse effects of treatments.
3. Drug Discovery and Development:Data science is accelerating the drug
discovery process. By analyzing genetic and molecular data, researchers can
identify potential drug candidates more efficiently. This approach reduces the
time and cost involved in bringing new drugs to market.
Predictive Diagnostics in Action
Predictive diagnostics, empowered by data science, are transforming healthcare in numerous
ways. Here are some real-world examples:
DATA SCIENCE
UNIT V
1. Cancer Prediction:Cancer is a leading cause of death worldwide. Data science
is aiding in the early detection of cancer by analyzing patient data, such as
genetic markers, lifestyle factors, and previous medical history. By identifying
individuals at high risk, healthcare providers can offer tailored screening and
prevention programs.
2. Cardiovascular Disease Risk Assessment:Heart disease remains a major global
health concern. Data science algorithms can predict an individual's risk of heart
disease based on various factors, including age, blood pressure, cholesterol
levels, and family history. This enables patients to take preventive measures and
make lifestyle changes.
3. Infectious Disease Outbreak Prediction:Data science is crucial in monitoring
and predicting infectious disease outbreaks. Machine learning models analyze
data from various sources, such as social media, weather patterns, and healthcare
records, to identify potential outbreaks. Early detection allows for swift
containment and preventive measures.
4. Genetic Medicine:Personalized medicine is a burgeoning field, and data science
is at its core. Genetic information can be used to determine an individual's
susceptibility to certain diseases and to tailor treatment options accordingly.
Challenges in Implementing Predictive Diagnostics
While the potential of predictive diagnostics in healthcare is immense, there are several
challenges in its implementation:
1. Data Privacy and Security:Handling sensitive patient data requires robust
security measures to protect patient privacy and comply with data protection
regulations like HIPAA in the United States.
2. Data Quality:The accuracy and completeness of the data are critical. Inaccurate
or incomplete data can lead to erroneous predictions and diagnoses.
3. Integration of Systems:Healthcare systems often use different technologies and
data formats. Integrating these systems to create a seamless data flow is a
significant challenge.
4. Ethical Concerns:The ethical use of predictive diagnostics is a pressing issue.
Decisions based on algorithms can raise ethical questions, and healthcare
professionals must carefully consider the implications.
DATA SCIENCE
UNIT V
Data Science Process
Data science process consists of six stages :
1. Discovery or Setting the research goal
2. Retrieving data
3. Data preparation
4. Data exploration
5. Data modeling
6. Presentation and automation
• Fig. 1.3.1 shows data science design process.
Step 1: Discovery or Defining research goal
This step involves acquiring data from all the identified internal and external sources, which
helps to answer the business question.
• Step 2: Retrieving data
DATA SCIENCE
UNIT V
It collection of data which required for project. This is the process of gaining a business
understanding of the data user have and deciphering what each piece of data means. This
could entail determining exactly what data is required and the best methods for obtaining it.
This also entails determining what each of the data points means in terms of the company. If
we have given a data set from a client, for example, we shall need to know what each column
and row represents.
• Step 3: Data preparation
Data can have many inconsistencies like missing values, blank columns, an incorrect data
format, which needs to be cleaned. We need to process, explore and condition data before
modeling. The cleandata, gives the better predictions.
• Step 4: Data exploration
Data exploration is related to deeper understanding of data. Try to understand how variables
interact with each other, the distribution of the data and whether there are outliers. To achieve
this use descriptive statistics, visual techniques and simple modeling. This steps is also called
as Exploratory Data Analysis.
• Step 5: Data modeling
In this step, the actual model building process starts. Here, Data scientist distributes datasets
for training and testing. Techniques like association, classification and clustering are applied
to the training data set. The model, once prepared, is tested against the "testing" dataset.
• Step 6: Presentation and automation
Deliver the final baselined model with reports, code and technical documents in this stage.
Model is deployed into a real-time production environment after thorough testing. In this
stage, the key findings are communicated to all stakeholders. This helps to decide if the
project results are a success or a failure based on the inputs from the model.
Elasticsearch
Elasticsearch is an Apache Lucene-based search server. It was developed by Shay
Banon and published in 2010. It is now maintained by Elasticsearch BV. Its latest version is
7.0.0.
Elasticsearch is a real-time distributed and open source full-text search and analytics engine.
It is accessible from RESTful web service interface and uses schema less JSON (JavaScript
Object Notation) documents to store data. It is built on Java programming language and
DATA SCIENCE
UNIT V
hence Elasticsearch can run on different platforms. It enables users to explore very large
amount of data at very high speed.
General Features
The general features of Elasticsearch are as follows −
• Elasticsearch is scalable up to petabytes of structured and unstructured data.
• Elasticsearch can be used as a replacement of document stores like MongoDB and
RavenDB.
• Elasticsearch uses denormalization to improve the search performance.
• Elasticsearch is one of the popular enterprise search engines, and is currently being
used by many big organizations like Wikipedia, The Guardian, StackOverflow,
GitHub etc.
• Elasticsearch is an open source and available under the Apache license version 2.0.
Key Concepts
The key concepts of Elasticsearch are as follows −
Node
It refers to a single running instance of Elasticsearch. Single physical and virtual server
accommodates multiple nodes depending upon the capabilities of their physical resources like
RAM, storage and processing power.
Cluster
It is a collection of one or more nodes. Cluster provides collective indexing and search
capabilities across all the nodes for entire data.
Index
It is a collection of different type of documents and their properties. Index also uses the
concept of shards to improve the performance. For example, a set of document contains data
of a social networking application.
Document
It is a collection of fields in a specific manner defined in JSON format. Every document
belongs to a type and resides inside an index. Every document is associated with a unique
identifier called the UID.
DATA SCIENCE
UNIT V
Shard
Indexes are horizontally subdivided into shards. This means each shard contains all the
properties of document but contains less number of JSON objects than index. The horizontal
separation makes shard an independent node, which can be store in any node. Primary shard
is the original horizontal part of an index and then these primary shards are replicated into
replica shards.
Replicas
Elasticsearch allows a user to create replicas of their indexes and shards. Replication not only
helps in increasing the availability of data in case of failure, but also improves the
performance of searching by carrying out a parallel search operation in these replicas.
Advantages
• Elasticsearch is developed on Java, which makes it compatible on almost every
platform.
• Elasticsearch is real time, in other words after one second the added document is
searchable in this engine
• Elasticsearch is distributed, which makes it easy to scale and integrate in any big
organization.
• Creating full backups are easy by using the concept of gateway, which is present in
Elasticsearch.
• Handling multi-tenancy is very easy in Elasticsearch when compared to Apache Solr.
• Elasticsearch uses JSON objects as responses, which makes it possible to invoke the
Elasticsearch server with a large number of different programming languages.
• Elasticsearch supports almost every document type except those that do not support
text rendering.
Disadvantages
• Elasticsearch does not have multi-language support in terms of handling request and
response data (only possible in JSON) unlike in Apache Solr, where it is possible in
CSV, XML and JSON formats.
• Occasionally, Elasticsearch has a problem of Split brain situations.
Comparison between Elasticsearch and RDBMS
In Elasticsearch, index is similar to tables in RDBMS (Relation Database Management
System). Every table is a collection of rows just as every index is a collection of documents
in Elasticsearch.
DATA SCIENCE
UNIT V
The following table gives a direct comparison between these terms−
Elasticsearch RDBMS
Cluster Database
Shard Shard
Index Table
Field Column
Document Row
Reference Link:
[Link]
[Link]
=case+study+in+data+sciences&ots=c5f4Bn0eHN&sig=Y36BVcc4Xwoz6F3PFsO_VKFf0N
I
[Link]
case+study+in+data+sciences&ots=xWraYm7Wto&sig=dPzFHVHIlbv2hX08Aa6ue6faTOI