CHAPTER 3
LITERATURE REVIEW
3.1 Introduction
We all know that health is very important key features nowadays to all of us. We know
that many countries like India, Bangladesh and Pakistan is really struggling with diabetes
patients. In America people also struggling with this stage. So many Researchers start
contributing their efforts in this field. In below section we studied number of research
papers and tried to build some summary for our research work.
According to the paper, Data mining is sub branch of computer science. It the way
through which we find some info from a given huge data. Here every day new technology
comes into existence like manufactured intelligences, DBMS, ML, DL. work of data
mining is found structural data that will provide some reasonable information from a give
huge data. Here Authors proposed that algorithm like Bayesian and KNN to apply of
patient data and try to find prediction of diabetics based upon given features. [10].
Finally, Authors conclude that Authors used a large dataset to ensure better
prediction result. Here Authors give some recommendation to the patient that how to
control diabetes in the case of young age patient. Authors build a system which will
anticipate diabetic patient. Here knowledge base assistance plays vital role in prediction
system. Authors taken a dataset which has 2000 in counts which will give nearness levels
of diabetic’s patient. Here prediction is taken place with the help of Naive bayes and k-
nearest Neighbors and also, they compare on the basis of some performance parameters.
This developed system may be very useful for HealthCare Industries for finding pre
diabetes patients [11].
Here Authors Explained that we have several Machine learning techniques which are
used for better prediction over a big data set. We all know that due to complexity
prediction in health sector is challenging job for all data scientist but it is very important
for HealthCare sector. This paper discussed about different six machine learning
1
algorithms are utilized for our prediction system. Performance and accuracy of applied on
a dataset. Here Authors applied different comparison parameters. Here Authors tried to
prove which one gives better result in terms of Accuracy. Aim of this research to help out
doctors and practioners for finding early prediction of diabetes with ML techniques.
According to a yearly issue brief released by American Health Insurance Program
(AHIP) [12], there are several factors that influence individual marketplace premiums.
These include individuals’ income, geography, age, and other factors. Individual
coverage preferences and benefits also influence premiums. In addition to these, other
market factors such as underlying health care cost growth, phase out of temporary
premium stabilization programs, increased utilization of services, market competition,
and risk pool effects in certain states and markets hike premiums. Government
involvement such as individual coverage requirement, premium subsidies, and increased
awareness of coverage options and enrolment initiatives reduces the premiums.
In this paper, Premiums in Ohio are very different from premiums in New Jersey.
Insurance companies with a population of customers of both healthy and unhealthy
individuals are likely to survive in insurance markets and can provide effective costs to
customers. Kaiser Family Foundation (KFF) [13] analyzed premium changes based on
insurance companies exit from the market place. KFF showed that overall health
insurance market place is less impacted by exiting of a health insurance company when
there are many players in the market place. However, the states with fewer insurance
companies that provide health plans to customers had a higher impact on the insurance
premiums because the remaining companies in the market place can hike the premiums.
Other papers by KFF [14] mentioned medical costs for people without insurance are more
compared to people with insurance. This is partially due to billing methods of hospitals
and negotiation between insurance companies and hospitals. Henry [15] et al. proposed a
model based on the data from HMO network using multi variate analysis to predict, the
people more likely to utilize health services. Health Care Payment Learning &
2
Action Network (HCP LAN) [16] suggested alternative payment plan with customer
centric model based on services utilized.
According to the Paper [17], the calculations of premiums involve following steps:
Experience Period Index Rate: To come up with the final premiums, insurance
companies begin with the Experience Period Index Rate, which is the average allowed
claims per member per month (PMPM) for Essential Health Benefits (EHBs) in the
previous year.
Projected Index Rate: Using the Experience Period Index Rate, the issuer adjusts for
health costs trends, demographics, and benefits to come up with a projected rate for the
upcoming plan year. The rate is basically the anticipated average allowed claims PMPM.
Market Adjusted Index Rate: The Projected Index Rate is adjusted for the federal
reinsurance program, risk, and marketplace user fees to arrive at the Market Adjusted
Index Rate.
Plan Adjusted Index Rate: From the issuer’s Market Adjusted Index Rate, the Plan
Adjusted Index Rate for each of the issuer’s plans is obtained by adjusting for plan
Actuarial Value and cost-sharing design, benefits, administrative costs, and catastrophic
plan eligibility variation.
Consumer Adjusted Premium Rates: Consumer Adjusted Premium Rates for each plan
are the Plan Adjusted Index Rate adjusted for age, tobacco usage, family size, and
geography. The Consumer Adjusted Premium Rate is the final premium charged to a
consumer.
Insurance industries are dealing with a lot of problems of interest to the operational
research community. An important interest of the insurance industry is studying and then
determining the insurance data. Hence, especially the exact
prediction of claim cost is considerably important for determining the scheme premium at
last by preventing the total loss of customers. Non–life insurance data is way different
3
arising out of the regression data because of its severe and frequency characteristics at a
lower range that is, the distribution of the claim cost has a large point mass at zero and it
is highly right-skewed. This paper will be focusing on model averaging or attaching the
methods all together to
improve the prediction accuracy. As stated by some of the reputed statistical models and
evidence which has proved that a combination of the model is a strong and useful way to
ameliorate the predictive performance as a whole.
Cardon et al., (2020) built a model for insurance options and obtained the suggestion of
insurance plan choice where people are loss averse as the cost of medical insurance
scheme shifting can maximize prosperity by decreasing adverse choice. An expensive
surety plan does not uplift expenditure in loss reduction actions as much as it should.
Kelly et al., (2017) concluded that with an increase of 10% in the final product, the health
sector efficiency collapses, having a productive consequence on prosperity [11]. Pelgrin
et al., (2016) discussed the lifetime consequences of exogenic change in medical
insurance on the dynamic optimal allocation status (health and wealth), welfare, (medical
investment, leisure, and consumption) and results from point to productive consequences
of policy on wealth, welfare and medication as well as mid-life replacement away from
healthy peace in favor of more medical charges and accelerating fitness issues caused by
peaking wages. Stavrunova et al., (2014) examined the effect of personal medical policy
mandate on call for private medical policy in Australia was examined and it was found
that the policy had no remarkable effect on universal request for private medical policy in
Australia. Mladenovic et al., (2020) used an ANN, namely an adaptive neuro-fuzzy
inference system (ANFIS) for simplifying the prediction method of the medical insurance
costs. They achieved an RMSE score of 7464.631. Loh et al., (2014) played an important
role in gaining acceptance and popularity by his intelligent usage of techniques which
includes regression trees, least squares and the proportional hazard models.
Alamelu et al., (2011) made a beneficial effort to study the financial success of the
insurance companies in India mainly in terms of asset standards and management systems
. Tkachenko et al., (2018) proposed a novel method for medical
4
insurance cost prediction i.e., piecewise-linear approach using the SGTM neural-like
structure. They also trained multi-layer perceptron and Common SGTM neural-like
structure. They collected observations about insurance costs in USA regions. They
observed that the proposed method performed very well with MAPE percentage of
30.60400373 and MAE of 3453.293634. Shinde et al., (2020) studied various regression
models and neural network models like Support Vector Machine Multiple Linear
Regression, Boost, Random Forest Regressor and Deep Neural Networks. They found
that Deep Neural Networks was the optimal method with RMSE value of 0.0695 and
accuracy of 87.95. Chowdhury et al., (2020) has developed an efficient machine learning
technique Neural Network (NN) with Internet of Things (IoT) for health insurance cost
prediction. The proposed method Fitness dependent Randomized Whale Optimization
Algorithm (FR-WOA) was the best. Izonin et al., (2019) proposed a novel technique of
constructing a committee based on SGTM neural architecture with RBF kernel. They also
used common SGTM neural architecture, multilayer perceptron, adaptive boosting, and
stochastic gradient descent regressor.
A claim severity can be defined as the amount of loss associated with an insurance claim.
The average severity is calculated by dividing the total amount of losses that an insurance
company experiences by the number of claims that were made against policies that it
underwrites. Loss is the amount paid or to be paid to the claimants under their insurance
policy contracts. Currently, the details of computing a forecast of the paid claim loss is
complicated [22].
Insurance companies rely on actuaries and the models that actuaries create to predict
future claims, as well as the losses that those claims may result in. The models are
dependent on a number of factors, including the type of risk being insured against, the
demographic and geographic information of the individual or business that bought a
policy, and the number of claims that are made. Actuaries look at past experience data to
determine if any patterns exist, and then compare this data to the industry at large. Claim
severity loss forecasting has played a major role in determining auto insurance rate and
premiums [23].
5
It is important to obtain accurate estimate of the losses that could arise from an insurance
contract. Also, in an event a car accident occurs, an insurance policy holder will prefer a
fast and quality service from the insurance company when it comes to processing claims
and it can also take a considerable amount of time to settle claims in some cases [24].
To provide quality claim service to millions of policy holders protected by insurance
companies, and also to create an accurate forecast that predicts rates and premiums, it is
necessary for insurance companies to have automated systems that can accurately predict
claim severity loss given a set of input. This research work aims to predict the severity
loss value of an insurance claim using continuous and categorical features from
previously processed insurance claims.
In the domain of loss prediction model, the work in [25] attempts to use convolution
approach to estimate loss severity distribution; using convolution of normal and
exponential distribution for modelling a loss distribution of property insurance claims.
The work in [26] demonstrates actuarial applications that uses hierarchical models to fit
micro-level insurance data consisting of claims policy and payment files to predict loss
type and accident frequency in automobile. Researchers have also tried to use insurance
claims data to build loss prediction models used for financial risk assessment in
construction projects [27].
The research work uses regression analysis to explore the relationships among
independent risk factors such as natural disasters, geographic information, and model
construction and the dependent variable (percentage of loss) to build a loss prediction
model based on the insurance payout records. A similar work [28 ] uses regression
analysis to build loss estimation models for insurance companies that shows the
correlation between post-earthquake damage of structural components to direct financial
loss of residential buildings based on post-earthquake damage evidence and obtained a R-
Square value of 0.41 using quadratic regression function.
As far as we are aware, we are the first to publish results from a regression model that
directly predicts severity loss value of an insurance claim. However, the work in [28]
6
uses regression models (K-nearest neighbor, support vector regression and feed forward
neural networks) to predict the overall cost of energy consumed during various categories
of entertainment events. The work in also used regression models to predict real estate
property prices. Similarly, the work in [29] uses regression models (boosting, linear
regression, support vector machine) to predict stock prices.
In addition, the work in uses regression models to predict the quality of signal transmitted
by optical fiber across communication channels. Other works attempted to illustrates how
to quantify the inherent uncertainty in fitting claims severity distributions and estimates
the cost of high layered factors that contributes to the severity loss of a claim [30].