0% found this document useful (0 votes)
5 views238 pages

Book Setting

Uploaded by

Ronit Paul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views238 pages

Book Setting

Uploaded by

Ronit Paul
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Contents

 Page No.
 An Overview of Accelerated Life Testing Models with Emphasis on 1-16
Design and Planning
Hemima Ahmed1, Sanjeeva Kumar Jha2, Bhanita Das3
 Wrapped three parameter Lindley Distribution for Modelling Circular 17-25
Data
Smriti Sikha Sarma1, Bishal Gurung2 and Bhanita Das3
 A Comparative Analysis of Different Statistical Learning Models for 26-36
Estimation of the Non-Destructive Shelf-Life Prediction of Okra
(lady‘s finger)
Joy Deb1, Dibyojyoti Bhattacharjee 2
 The Impact of Climate Change and other Factors on Maize 37-48
Production in Tanzania: Autoregressive Distributed Lag Approach
Bahati Ilembo1 and Johanes Tibanywana2
 Statistical Inference of Stress-Strength Reliability Model using 49-58
Lindley Distribution
Obed Benaiah Jyrwa1, Hemima Ahmed2 and Bhanita Das3
 A Review on Various Method of Compounding using Series and 59-74
Parallel Structure
Magrisha Namsaw1 and Bhanita Das2
 Educational Status of Rani Khamar Village, Palasbari, Assam: A 75-84
Socio-Economic Survey-Based Study
Ananya Guha¹, Menaka Sikdar2
 A Brief Overview on Construction of Discrete Analogues of 85-98
Continuous Probability Distribution
Diksha Das1 and Bhanita Das2
 An intensive assessment of Precipitation, Humidity and Temperature 99-113
patterns through the analytical framework of Extreme Value Theory
Dipanjali Ray1, Tanusree Deb Roy2, and Sebul Islam Laskar3
 International trade relation of Assam with neighboring country 114-120
Bhutan: Land Custom Station (LCS) based analysis
Rupjyoti Bordoloi
 On Some Developments in the Gamma-G Type 2 Family of Distributions 121-131
1 2 3
Hmingthansanga , Sanjeeva Kumar Jha , and Bhanita Das

1
 Market Forecasting Using Stochastic Process: A Study on Bharat 132-157
Heavy Electricals Limited
Ronit Paul 1, Dr. Tanusree Deb Roy 2
 Topp-Leone Unit Lindley Distribution: Derivation, Some 158-167
Fundamental Properties And Estimation
Sahana Bhattacharjee
 A Semi-Circular Exponential Distribution Induced by Inverse 168-178
Stereographic Projection: Properties and Application
Imliyangba1, Bhanita Das2 and Seema Chettri3
 Cardiovascular Failure Risk Prediction: A Biostatistical Perspective 179-197
1 2 3
Dr. Suman Jaiswal , Dr. Vaishali Saxena and Bhavya Jain
 Archimedean Copula Functions: A Powerful Tool for Dependency 198-210
Modeling
Ruhiteswar Choudhury1, and Tanusree Deb Roy1
 Technology Addiction among Students: An Analysis Based on 211-222
Students of Different Colleges under Dibrugarh University- A Case
Study
Kuldeep Goswami1 and Sricharan Shah2
 A New Form of Two Parameters Skew-Logistic Distribution and its 223-236
Real Life Application
Jondeep Das1, Partha Jyoti Hazarika2, Dimpal Pathak3

2
An Overview of Accelerated Life Testing Models with
Emphasis on Design and Planning

Hemima Ahmed1, Sanjeeva Kumar Jha2, Bhanita Das3


1
Research Scholar, North-Eastern Hill University,
Shillong, Email ID ahmedhemim@[Link]
2
Associate Professor, North-Eastern Hill University,
Shillong, Email ID [Link]@[Link]
3
Assistant Professor, North-Eastern Hill University, Shillong, Email ID
bhanitadas83@[Link]

Abstract:
Accelerated Life Testing (ALT) is a foundational methodology in reliability
engineering designed to expedite failure data collection by subjecting products to
intensified stress conditions, such as elevated temperature, voltage, pressure, load, or
vibration, to accelerate failures compared to typical operating environments. This
enables accurate estimation of product durability and performance under standard
usage while optimizing resource utilization through reduced time, labour, material,
and cost investments. Driven by rapid technological advancements, shorter product
development cycles, and growing reliability demands, ALT has evolved to integrate
diverse stress-loading approaches, including constant- stress, step-stress, and
progressive-stress tests, alongside statistical frameworks such as parametric,
nonparametric, and Bayesian models. Optimization criteria like V- optimality and D-
optimality are strategically employed to design test plans that balance resource
constraints with precise reliability assessments. This paper reviews critical ALT
models and planning methodologies, emphasizing their practical applications in
modern engineering contexts and providing guidance for practitioners to select
tailored strategies aligned with specific product requirements, testing objectives, and
operational constraints
Keywords: Accelerated life test; step-stress model; distribution; competing risk;
prediction.

1. INTRODUCTION
1.1. MOTIVATION
In the early stages of product development, reliability tends to be low, and the
product lifespan is limited. During this phase, reliability tests often replicate real--
world usage scenarios. However, as product dependability improves, conducting
1
such tests becomes increasingly difficult, as intentionally inducing failures under
normal conditions is not feasible. Additionally, these tests may not provide sufficient
failure data within a reasonable timeframe. To address these challenges, reliability
experts have developed the Accelerated Test (AT) model. In this approach, failures
are induced more quickly by subjecting the test units to higher stress levels than
those encountered under normal operating conditions. The data collected under these
high-stress conditions are then used to predict product life under normal stress levels
and improve product reliability.
AT methods are classified into two categories: Accelerated Life Test (ALT) or
Partially Accelerated Life Test (PALT), and Accelerated Deterioration Test (ADT).
In ALT, the test items are exposed to accelerated conditions, and both failure and
censoring times are recorded. This data is used to develop an ALT model that
extrapolates the product's reliability under normal conditions. On the other hand,
PALT involves testing units under both accelerated and normal usage conditions.
Some units in a PALT are operated under normal conditions throughout the test
duration. However, ALT is typically preferred over PALT when testing under normal
conditions would require excessive time. ALT data can be used to estimate
acceleration factors, particularly when the relationships between test conditions and
product lifetime (such as log-linear or log-quadratic models) are known. In some
cases, ALT may prove more effective than PALT. ADT focuses on measuring
product degradation by evaluating performance or product characteristics. A failure
is identified when the degradation crosses a predetermined threshold, which is
modelled based on the degradation path.
Currently, ALT is the most widely used technique for quickly assessing product
reliability in real-world scenarios. However, before conducting an ALT, it is crucial
to develop an effective plan supported by relevant statistical theories to ensure that
the evaluation of product reliability is accurate, efficient, and cost-effective. ALT is
extensively applied across industries such as manufacturing, engineering,
electronics, medical sciences, and biological studies. In an ALT setup, while the
experiment duration is reduced, it is not directly controlled. To manage the total
experimental duration, various censoring schemes are employed. Common
censoring schemes used in ALT include type-I censoring, type-II censoring, hybrid
censoring, progressive censoring, and random censoring. Of these, type-I and type-II
censoring are the most commonly used. Type-I censoring (time censoring) involves
terminating the test after a predetermined time, with the number of failures being a
random variable. In contrast, type-II censoring (failure censoring) terminates the test
after a specified number of failures, with the test duration being a random variable.

1.2. LITERATURE
Over the past several decades, the field of ALT has garnered significant attention
from researchers, driven by factors such as escalating reliability requirements, rapid
technological advancements, the high costs associated with product failures,
shortened product life cycles, adherence to stringent standards, competitive market
2
pressures, and growing sustainability concerns. The evolution of ALT methodologies
began with fundamental constant-stress loading tests and has progressively advanced
to encompass more sophisticated techniques, including variable stress loading and
degradation-based assessments.
The inception of ALT research can be traced back to the 1950s, with pioneering
contributions from Chernoff [1] and Bessler [2], who introduced foundational
concepts in accelerated testing. Their work laid the groundwork for subsequent
developments in the field. Nelson [3] provides a thorough and comprehensive
resource covering background information, practical methodologies, fundamental
theory, and illustrative examples related to accelerated testing (AT). Building upon
this foundation, Meeker and Hahn [4] provided practical guidelines for designing
effective ALT plans, emphasizing statistical rigor and applicability in industrial
settings. Bhattacharyya [5] contributed to the literature by presenting an overview of
ALT, highlighting recent advancements and their implications for reliability
engineering. Vassiliou and Mettas [6] further enriched the discourse by
distinguishing between quantitative and qualitative AT approaches, thereby refining
the conceptual framework of accelerated testing methodologies. Subsequent studies
by Meeker et al. [7] delved into the intricacies of statistical models, test planning,
and data analysis within the context of AT, offering nuanced insights into the design
and interpretation of accelerated tests. Comprehensive reviews by Farrar et al. [8],
Escobar et al. [9], Limon et al. [10], and Chen et al. [11] have synthesized existing
knowledge, providing critical evaluations of various AT models and their practical
applications across diverse engineering domains. Meeker et al. [12] have also
contributed to this area. Collectively, these scholarly contributions have significantly
advanced the field of ALT, furnishing practitioners and researchers with robust
frameworks and methodologies to assess and enhance product reliability under
accelerated conditions.

1.3. OVERVIEW
The rest of this paper is organized as follows: Section 2 covers the basic
concepts of ALT. Section 3 discusses the main models used in ALT, such as the
Cumulative Exposure Model (CEM), Tampered Random Variable Model (TRVM),
Tampered Failure Rate Model (TFRM), and Cumulative Risk Model (CRM).
Section 4 explores the challenges posed by competing risks in ALT. Section 5
focuses on the challenges of predicting product reliability and failure probabilities
under accelerated conditions. Finally, Section 6 wraps up the paper by summarizing
the key points and suggesting areas for future research in ALT.

2. BASICS OF ALTS
Accelerated Life Testing (ALT) is a pivotal methodology in reliability
engineering, enabling the prediction of product lifespan under normal usage by
subjecting it to elevated stress conditions in a controlled environment. This approach
facilitates the identification of potential failure modes within a shortened timeframe,
3
thereby enhancing product design and quality assurance. Accurate simulation of
field failure mechanisms in laboratory settings is essential for effective ALT.
Utilizing tools such as Failure Modes and Effects Analysis (FMEA), historical
failure data, and engineering expertise aids in identifying critical stress variables that
influence product reliability. A well-structured test plan should consider factors
including the product's life distribution, appropriate stress levels, types of ALTs, and
the relationships between stress and life expectancy.

2.1. TYPES OF ALTS


Stress can be applied to ALTs in various ways, including constant-stress, step-
stress, and progressive-stress. Among these, constant-stress is the most commonly
used method in ALT. Nelson [3] investigated the advantages and disadvantages of
each approach. In constant-stress accelerated life tests (CS-ALTs), each unit is
subjected to a constant high stress until failure or the test is terminated. Several
studies have focused on CS-ALTs. For example, Liu and Tang [13] proposed a
sequential CS-ALT method along with associated Bayesian inference. Fan and Hsu
[14] investigated the reliability of CS-ALTs in a series system with many
components, using independent Weibull distributions with log-linear scale
parameters at the stress variable level. Abd El-Raheem [15], [16] studied the optimal
sample size allocation for CS-ALTs with type-I and type-II censoring, assuming an
extension of the exponential distribution. Kumar et al. [17] explored eight
frequentist estimation approaches, in addition to the maximum likelihood method,
for estimating parameters of the generalized inverse Lindley distribution under CS-
ALT, with an exponential distribution and progressive type-II censoring. Finally,
Feng et al. [18] examined reliability inference for a dual CS-ALT model.
Step-stress accelerated life tests (SS-ALTs) are a special case of ALTs that allow
the experimenter to change the stress level in stages throughout the experiment.
There are two main types of SS-ALTs: simple and multiple. In a simple SS-ALT, the
stress level is changed once during the test, whereas in a multiple SS-ALT, the stress
level changes more than once. Millar et al. [19] studied optimal simple step-stress
plans for accelerated life testing. Tang et al. [20] investigated the optimal design for
SS-ALTs using a two-parameter Weibull distribution under type-I censoring.
Balakrishnan et al. [21] studied point and interval estimation for a simple step-stress
model with type-II censoring, deriving exact distributions for the maximum
likelihood estimates (MLEs) of the unknown parameters. Mohie El-Din et al. [22]
examined SS-ALTs for the power generalized Weibull distribution under progressive
type-II censoring. Alotaibi et al. [23] considered the step-stress model for alpha
power Weibull lifetimes under progressively type-II censored samples. Ling and Hu
[24] established the optimal design of a simple SS-ALT for one-shot devices under
the Weibull distribution by minimizing the asymptotic variance of the MLE of
reliability at normal operating conditions, focusing on sample allocation and
inspection durations. In progressive-stress ALTs (PS-ALTs), the applied stress on test

4
units is continuously increasing in time. If an ALT has a continuous linearly
increasing stress, this test is called a ramp-stress test. Several authors have dealt with
this type of ALT. Abdel-Hamid [25] examined the progressive-stress model for the
exponentiated exponential distribution under type-II progressive hybrid censoring.
Mohie El-Din et al. [26] studied the classical and Bayesian inference on PS-ALTs
for the extension of the exponential distribution under progressive type-II censoring
and Mohie El-Din et al. [27] have discussed PS-ALTs for power generalized Weibull
distribution under progressive type-II censoring. Kumar Mahto et al. [28] considered
PS-ALT under progressive type-II censoring is considered when the lifetime of test
units follows the logistic exponential distribution
In ALTs, the test items run only at accelerated conditions. Partially accelerated
life tests (PALTs) are a type of the ALT scheme, that allows part of the test units to
be run at a normal condition throughout the entire testing period. This means that the
PALTs combine both ordinary and accelerated life tests. The two different types of
PALTs are constant-stress partially accelerated life test (CS-PALT) and step-stress
partially accelerated life test (SS-PALT). In CS-PALTs, the total test units are first
divided into two groups, the items of one of the groups are allocated to life tests
under a normal condition, and the items of the other group are allocated to life tests
under a stressed condition, each of the units is run at a constant level of stress until
the unit fails. Zarrin et al. [29] studied the estimation in CS-PALTs for Rayleigh
distribution using type-I censored data respectively. Ismail [30] have estimated the
generalized exponential distribution parameters and the acceleration factor under
CS-PALTs with type-II censoring. Under time constraints, Srivastava et al. [31]
considered the optimum CS-PALTs for the truncated logistic distribution. Bayes
inference in CS-PALTs for the generalized exponential distribution with progressive
censoring was examined by Jaheen et al. [32]. Abushal et al. [33] considered the CS-
PALTs for Pareto distribution under Progressive censored data. In SS-PALT, the test
for units starts at normal use conditions for a specified time and the surviving units
are then run under accelerated conditions until they fail. Several writers have
explored PALTs using the step-stress approach. Ismail [34] studied maximum
likelihood estimates of the model parameters under SS-PALT are obtained assuming
the Weibull distribution with type-II censored data. Ismail [35] have discussed the
statistical inference for SS-PALTs with an adaptive type-I progressively hybrid
censored data from Weibull distribution. Soliman et al. [36] considered the SS-
PALTs model in the estimation of inverse Weibull parameters under progressive
type-II censoring. Abushal [37] considered the SS-PALTs for the compound Weibull-
gamma distribution under type-I censoring. Inferences for Weibull-exponential
distribution based on progressive type-II censoring under SS-PALTs were discussed
by M. El-Sagheer et al. [38]. Akgul et al. [39] studied the classical and Bayesian
inferences in SS-PALTs for inverse Weibull distribution under type-I censoring.
Pandey et al. [40] investigated the statistical analysis for generalized progressive
hybrid censored data from Lindley distribution under an SS-PALT model.

5
2.2. ACCELERATING STRESS VARIABLES
The selection of accelerating stress variables depends on the component type
and its typical operating conditions. Mechanical systems, such as springs, shafts, and
ball bearings, often undergo stresses like vibration, random shock, and humidity to
precipitate failure. In contrast, electronic components commonly rely on
temperature, voltage, current, humidity, and UV radiation as key accelerating
factors. While these variables are frequently applied individually, combinations may
be tailored to replicate actual operating environments and failure mechanisms more
accurately.
A critical consideration in ALT is determining the optimal number of stress
variables to apply. Single-stress testing remains the most prevalent approach due to
its simplicity, well-documented methodologies, and proven predictive models.
However, recent advancements have spurred interest in multi-stress testing to better
mimic real-world conditions. Despite its potential, multi-stress testing introduces
challenges such as stress interaction effects where combined stresses may amplify or
mask failure modes unpredictably and a scarcity of robust models correlating life
expectancy with multiple stressors. Consequently, practitioners must balance the
enhanced realism of multi-stress tests against the complexity and uncertainty they
introduce, often favouring single-stress methods for their reliability and ease of
interpretation in early-stage analyses.

2.3. LIFETIME DISTRIBUTIONS


Lifetime distributions play a pivotal role in designing and interpreting ALT
studies. Since failures occur unpredictably over time, the time-to-failure is modeled
as a random variable governed by a cumulative distribution function (CDF).
Commonly used accelerated failure time (AFT) models include the Weibull,
lognormal, exponential, normal, and extreme value distributions. Traditional
approaches, particularly with the Weibull distribution, assume the shape parameter
remains constant across stress levels, as it depends on inherent material properties,
while the scale parameter varies with applied stress. However, recent studies
challenge this paradigm, suggesting that stress-dependent shape parameters may
arise due to inconsistencies in manufacturing processes, supplier variations, or rapid
innovations in material engineering.
To enhance ALT accuracy, modern methodologies emphasize selecting flexible
lifetime distributions capable of modeling both monotonic (increasing or decreasing)
and non-monotonic (e.g., bathtub-shaped) hazard functions. Researchers have
developed advanced statistical frameworks by merging or modifying existing
distributions, often introducing additional parameters to improve adaptability. These
extended models, such as the exponentiated Weibull, generalized exponential, alpha
6
power Weibull, and Burr-type distributions, provide enhanced adaptability in
modeling extreme-value patterns such as heavy or light tails and improving
statistical alignment with intricate failure datasets. Such innovations address the
limitations of classical models, enabling more robust reliability predictions under
diverse stress conditions.

2.4. CURRENT STATE OF ALT


In practice, CSALT with single-stress loading and time-censored data remains
the most widely adopted method. However, Multiple Stress ALT (MSALT) and
PSALT with time-censoring are gaining prominence in engineering fields, driven by
demands for higher product reliability, operational complexity, and advancements in
testing instrumentation. Statistical methodologies for ALT are broadly classified into
four categories: descriptive, parametric, non-parametric, and Bayesian statistics.
Parametric approaches dominate current applications, with maximum likelihood
estimation (MLE) for location-scale distributions and linear stress-life relationships
representing the most refined and frequently utilized framework. Nevertheless, non-
parametric and Bayesian methods are increasingly favoured to address challenges
such as ambiguous product life distribution identification and insufficient failure
data, particularly as ALT expands to diverse engineering components.
Planning ALT experiments is an inverse problem to data analysis. Engineers and
statisticians prioritize designing optimal test plans, a central research focus in ALT
theory. Effective planning requires alignment between five elements: stress loading
strategy, statistical model, data analysis technique, testing constraints, and design
objectives. Consequently, optimizing ALT plans involves three interconnected
dimensions. The first centers on test configurations, such as selecting stress profiles
(e.g., CSALT, SSALT, PSALT) and censoring schemes (type-I, type-II) for single-,
dual-, or multi-stress scenarios. The second dimension involves statistical
frameworks, integrating descriptive, parametric, non-parametric, or Bayesian models
to align with failure mechanisms and data characteristics. The third-dimension
addresses optimization criteria, balancing objectives like V-optimality, which
minimizes the variance of quantile life estimates under normal stress, and D-
optimality, which maximizes the determinant of the Fisher information matrix, while
incorporating constraints such as cost, time, and resource limitations. Among these,
the V-optimal CSALT plan with type-I censoring, applied to location-scale
distributions and linear stress-life models, remains the most extensively studied and
implemented framework. This approach balances simplicity, computational
tractability, and alignment with classical reliability assumptions, making it a
cornerstone of industrial ALT practices.

7
3. MODELS USED TO ANALYSE THE ALTS
To analyse data using a step-stress pattern, we require a model that connects
CDF of the product's lifetime under various stress levels to the CDF of the product's
lifetime under normal operating conditions. Several models have been suggested in
the literature to explain this relationship.
3.1. CUMULATIVE EXPOSURE MODEL (CEM)
The CEM, first introduced by Sediakin [41] and later expanded by Nelson [42],
is a widely adopted framework for step-stress ALT analysis. The CEM operates
under two core principles: (1) the remaining lifetime of a test unit depends solely on
the total cumulative exposure it has endured, irrespective of the stress application
sequence, and (2) at each constant stress level, failures follow the CDF specific to
that level, while ensuring continuity in the overall lifetime distribution across stress
transitions.
For a k-step stress ALT with predefined stress change times 𝜏1 < 𝜏2 < ⋯ <
𝜏𝑘−1 , Nelson‘s [42] formulation of the CEM defines the cumulative distribution
function 𝐹(𝑡) as a piecewise structure:
𝐹1 (𝑡), 𝑖𝑓 0 ≤ 𝑡 < 𝜏1
𝐹2 (𝑡 − 𝜏1 + 𝑕1 ), 𝑖𝑓 𝜏𝑘−1 ≤ 𝑡 < 𝜏2
𝐹(𝑡) = { (1)

𝐹𝑘 (𝑡 − 𝜏𝑘−1 + 𝑕𝑘−1 ), 𝜏𝑘−1 ≤ 𝑡 < ∞
That is, with 𝜏0 = 0 and 𝜏𝑘 = ∞, we have the general form as
𝐹(𝑡) = 𝐹𝑗 (𝑡 − 𝜏𝑗−1 + 𝑕𝑗−1 ), 𝜏𝑗−1 ≤ 𝑡 < 𝜏𝑗 , 𝑗 = 1,2, … , 𝑘 (2)
While the shifting parameter 𝑕0 = 0 and 𝑕𝑗 for 𝑖 = 1, 2, . . . , 𝑘 − 1, is the
solution of the equation
𝐹𝑗 (𝑕𝑗−1 ) = 𝐹𝑗−1 (𝜏𝑗−1 − 𝜏𝑗−2 + 𝑕𝑗−2 ). (3)
because of the continuity of the function 𝐹(𝑡) at 𝜏𝑗 .

3.2. TAMPERED RANDOM VARIABLE MODEL (TRVM)


The TRVM, proposed by Goel [43], provides a framework for analysing step-
stress accelerated life testing (SS-ALT). This model postulates that increasing the
stress level from 𝑠1 to 𝑠2 at a predefined time 𝜏1 proportionally scales the remaining
lifetime of a test unit by an acceleration factor 𝛽 −1, where 𝛽 > 1 depends on the
stress levels 𝑠1 and 𝑠2.
Formally, let 𝑇 denote the lifetime of a unit under stress 𝑠1. Under the TRVM
assumption,
the lifetime Y under a simple SSLT is defined as:
𝑇, 𝑖𝑓 𝑇 ≤ 𝜏1
𝑌={ −1 (𝑇 (4)
𝜏1 + 𝛽 − 𝜏1 ), 𝑖𝑓 𝑇 > 𝜏1
Here, 𝑌 is termed a tampered random variable because elevating the stress level
at 𝜏1 effectively alters the natural failure process. The parameter 𝛽 represents the
8
acceleration (or tampering) factor, 𝛽 −1 is the tampering coefficient, and 𝜏1 is the
tampering point. If 𝑇 has a probability density function 𝑓𝑇 (. ), then the PDF of 𝑌
becomes
𝑓𝑌 (𝑦), 𝑖𝑓 𝑦 ≤ 𝜏1
𝑓𝑌 (𝑦) = { (5)
𝛽𝑓𝑌 (𝜏1 + 𝛽(𝑇 − 𝜏1 )), 𝑖𝑓 𝑇 > 𝜏1

3.3. TAMPERED FAILURE RATE MODEL (TFRM)


The TFRM, developed by Bhattacharyya et al. [44], addresses simple SS-ALT.
This model posits that switching the stress level from 𝑠1 to 𝑠2 at a predetermined
time 𝜏1 scales the failure rate function (FRF) at stress level 𝑠1 by a constant factor 𝛽.
Formally, the failure rate under step-stress conditions, denoted 𝜆̃(𝑡), is defined as:
𝜆(𝑡), 𝑖𝑓 𝑡 ≤ 𝜏1
𝜆̃(𝑡) = { (6)
𝛽𝜆(𝑡), 𝑖𝑓 𝑡 > 𝜏1
Where 𝜆(𝑡) and 𝜆̃(𝑡) represent the failure rates at stress level 𝑠1 and under the
step-stress regimen, respectively. The scaling factor 𝛽, typically greater than 1,
depends on both stress levels 𝑠1, 𝑠2 and potentially the tampering point 𝜏1 .
If the CDF at stress level 𝑠1 is 𝐹(𝑡), the CDF of the lifetime under step-stress
conditions, 𝐹̃ (𝑡), becomes:
𝐹(𝑡), 𝑖𝑓 𝑡 ≤ 𝜏1
𝐹̃ (𝑡)=8 1−𝛽 𝛽 (7)
1 − (1 − 𝐹(𝜏1 )) (1 − 𝐹(𝑡)) , 𝑖𝑓 𝑡 > 𝜏1

3.4. CUMULATIVE RISK MODEL (CRM)


Existing models such as the CEM, TRVM, and TFRM assume instantaneous
effects when stress levels change, a simplification often inconsistent with real-world
behaviour. In practice, altering stress levels introduces a gradual transition period
before the full impact is observed. To address this limitation, Drop et al. [45]
proposed a step-stress model incorporating latency periods, later formalized as the
Cumulative Risk Model (CRM) by Kannan et al. [46]. Consider a step-stress
experiment where the stress level shifts from 𝑠1 to 𝑠2 at time 𝜏1 . Let 𝑕1 (𝑡) and
𝑕2 (𝑡)denote the continuous hazard functions for stress levels 𝑠1 and 𝑠2,
respectively. The CRM introduces a latency period 𝛿 during which the hazard
function transitions linearly between stress regimes, resulting in a three-phase
piecewise structure:
𝑕1 (𝑡), 𝑖𝑓 0 < 𝑡 ≤ 𝜏1
𝑕(𝑡) = >𝑎 + 𝑏𝑡, 𝑖𝑓 𝜏1 ≤ 𝑡 ≤ 𝜏1 + 𝛿 (8)
𝑕2 (𝑡), 𝑖𝑓 𝑡 ≥ 𝜏1 + 𝛿
The parameters a and b enforce continuity at the transition boundaries.
Specifically, 𝑎 + 𝑏𝜏1 = 𝑕1 (𝜏1 ) ensures seamless alignment with the initial hazard
rate at 𝜏1 , while 𝑎 + 𝑏𝜏2 = 𝑕2 (𝜏1 + 𝛿) guarantees consistency with the new stress

9
level‘s hazard rate at 𝜏1 + 𝛿. These constraints uniquely determine a and b, enabling
a smooth escalation of risk that mirrors gradual material degradation or system
adaptation under elevated stress. By modeling delayed stress effects through linear
interpolation, the CRM better reflects physical failure mechanisms compared to
abrupt hazard jumps assumed in classical models. This approach enhances predictive
accuracy for applications where instantaneous stress responses are unrealistic, such
as thermal cycling in electronics or fatigue testing in mechanical systems.

4. HANDLING COMPETING RISKS WITH ACCELERATED LIFE TESTING


In reliability analysis, a unit may fail due to multiple distinct causes, such as
heart disease, cancer, or kidney failure in medical contexts, or electrical, mechanical,
or environmental factors in engineering systems. These scenarios are termed
competing risk problems, where different failure modes ―compete‖ to trigger the
event. Analysing such data requires evaluating each cause while accounting for the
presence of others. Each observation must include both the failure time and the
specific cause to enable meaningful analysis.
Two primary methodologies are employed: parametric approaches, which
assume specific lifetime distributions (e.g., exponential, Weibull) for each failure
mode, and non-parametric approaches, which avoid distributional assumptions,
offering flexibility at the cost of reduced precision. While independence between
failure causes is often assumed for simplicity, real- world dependencies may exist.
However, modeling dependent risks introduces challenges, particularly in
distinguishing between correlated causes and their individual effects, leading to
identifiability issues in the underlying model.
A significant practical challenge arises from the long lifetimes of many products,
which delay failure observation and escalate testing costs. Accelerated Life Testing
(ALT) mitigates this by subjecting units to elevated stress conditions (e.g., extreme
temperatures, voltage), accelerating failures without altering the hierarchy of failure
modes. This approach enables rapid data collection, reduces expenses, and supports
timely reliability assessments. By combining competing risk models with ALT,
analysts can efficiently isolate and quantify the impact of individual failure
mechanisms, even under constrained time and resources.
In reliability analysis, two principal frameworks address competing risk data:
the latent failure time model introduced by Cox [47] and the cause-specific hazard
function approach proposed by Prentice et al. [48]. These methodologies differ
conceptually, with Cox‘s model treating failure times as latent variables and
Prentice‘s focusing on cause-specific hazards. Kundu [49] demonstrated that for
exponential and Weibull distributions, both approaches yield identical likelihood
functions, producing the same parameter estimates despite differing interpretations.
Extensive research over five decades has advanced competing risk
methodologies. Crowder [[50], [51]] dealt with the different specific issues related to
various competing risk models. Klein and Basu [52] pioneered integrating

10
competing risk models with ALT. Pascual [53] optimized ALT designs for Weibull-
distributed failures. Balakrishnan and Han [54] derived exact inference methods for
exponential step-stress models, later extended by Han and Kundu [55] to generalized
exponential distributions. Srivastava et al. [56] adapted these principles to the
Khamis-Higgins model, while Lu and Shi [57] incorporated progressive censoring
schemes. Ganguly and Kundu [58] introduced random stress-change times,
broadening model applicability. Despite advancements, Cox‘s [47] latent failure time
model remains widely adopted for its interpretative flexibility and alignment with
traditional survival analysis, underscoring its utility in resolving complex competing
risks as methodologies evolve.

5. PREDICTION PROBLEM FOR ACCELERATED LIFE TESTING


A core challenge in statistical analysis involves predicting future censored
observations using available data, with applications spanning fields such as
engineering, medicine, and actuarial sciences. Prior studies have focused on
statistical inference for failure-time distributions in step-stress models under various
censoring schemes, yet prediction of actual failure times remains underexplored.
This gap persists despite the practical relevance of forecasting unobserved failures,
particularly in ALT frameworks.
Within ALT, Basak [59], Basak et al. [60], and Basak et al. [61] pioneered
methods to predict failure times for censored units under exponential distributions,
utilizing progressive Type-I, Type-II, and hybrid censoring schemes in step-stress
models. Amleh et al. [62] expanded this work to the Lomax distribution under Type-
II censoring in step-stress ALT (SS-ALT) settings. More recently, Amleh et al. [63]
addressed prediction challenges for censored Weibull lifetimes by integrating a
simple step-stress model with the Khamis-Higgins model. These contributions
collectively advance methodologies for reliability forecasting, addressing critical
gaps in ALT prediction research.

6. CONCLUSION
Accelerated life testing has emerged as an indispensable tool in reliability
engineering, enabling efficient assessment of product durability and performance
under normal operating conditions through controlled exposure to elevated stress
levels. This paper has systematically reviewed key ALT methodologies, including
constant-stress, step-stress, and progressive-stress loading schemes, alongside
critical statistical frameworks such as the CEM, TRVM, TFRM, and CRM. These
models provide robust mechanisms for extrapolating failure data from accelerated
conditions to real-world usage scenarios while addressing challenges such as stress
transition effects and gradual degradation processes. The discussion highlighted the
importance of integrating competing risk models with ALT to account for multiple
failure modes, emphasizing the need for both parametric and non-parametric
approaches to balance precision and flexibility. Furthermore, prediction challenges
11
in ALT, particularly forecasting censored failure times, were addressed through
advanced methodologies leveraging progressive censoring schemes and distribution-
specific models, underscoring their practical relevance in reliability forecasting.
Future research should prioritize multi-stress testing frameworks to better
simulate com plex operational environments, develop advanced models for
dependent competing risks, and integrate machine learning techniques for enhanced
predictive accuracy. Additionally, validation studies comparing model predictions
with field data remain critical to refining ALT methodologies. As industries continue
to demand higher reliability standards amid rapid technological advancements, the
evolution of ALT strategies will play a pivotal role in optimizing product design,
reducing costs, and ensuring safety across diverse engineering applications.
Practitioners are encouraged to adopt a systematic approach in selecting ALT
models, aligning stress protocols, statistical methods, and optimization criteria with
specific product requirements and testing objectives.

REFERENCES:
1. Chernoff, H. (1962). Optimal accelerated life designs for estimation.
Technometrics 4(3), 381–408.
2. Bessler, S., Chernoff, H., & Marshall, A.W. (1962). An optimal sequential
accelerated life test. Technometrics 4(3), 367–379.
3. Nelson, W.B. (1990). Accelerated Testing: Statistical Models, Test Plans, and
Data Analysis. John Wiley & Sons.
4. Meeker, W.Q., & Hahn, G.J. (1985). How to Plan an Accelerated Life Test:
Some Practical Guidelines. ASQC Statistics Division.
5. Bhattacharyya, G.K. (1986). Accelerated life test: An overview and some
recent advances. In: Proceedings of the Thirty-First Conference, pp. 253–273.
6. Vassiliou, P., & Mettas, A. (2001). Understanding accelerated life-testing
analysis. In: Annual Reliability and Maintainability Symposium, Tutorial
Notes, pp. 1–21.
7. Meeker, W. (1991). Accelerated testing: Statistical models, test plans, and data
analyses. Technometrics 33, 236–238.
8. Farrar, C.R., & Duffey, T.A. (1999). Cornwell, P.J., Bement, M.T., et al.: A
review of methods for developing accelerated testing criteria. In: Proc. of the
17th International Modal Analysis Conference, Kissimmee, FL, pp. 608–614.
9. Escobar, L.A., & Meeker, W.Q. (2006). A review of accelerated test models.
Statistical science, 552–577.
10. Limon, S., Yadav, O.P., & Liao, H. (2017). A literature review on planning and
analysis of accelerated testing for reliability assessment. Quality and
Reliability Engineering International 33(8), 2361–2383.
11. Chen, W. H., Gao, L., Pan, J., Qian, P., & He, Q. C. (2018). Design of
accelerated life test plans-overview and prospect. Chinese Journal of
Mechanical Engineering 31, 1–15.

12
12. Meeker, W.Q., Escobar, L.A., & Pascual, F.G. (2022). Statistical Methods for
Reliability Data. John Wiley & Sons.
13. Liu, X., & Tang, L.C. (2009). A sequential constant-stress accelerated life
testing scheme and its Bayesian inference. Quality and Reliability Engineering
International 25(1), 91–109.
14. Fan, T.H., & Hsu, T.M. (2014), Constant stress accelerated life test on a
multiple-component series system under Weibull lifetime distributions.
Communications in statistics-theory and methods 43(10-12), 2370–2383.
15. Abd El-Raheem, A.M. (2019). Optimal plans and estimation of constant-stress
accelerated life tests for the extension of the exponential distribution under
type-I censoring. Journal of Testing and Evaluation 47(5), 3781–3821.
16. Abd El-Raheem, A. (2021). Optimal design of multiple constant-stress
accelerated life testing for the extension of the exponential distribution under
type-II censoring. Journal of Computational and Applied Mathematics 382,
113094.
17. Kumar, D., Nassar, M., Dey, S., & Alam, F.M.A. (2022). On estimation
procedures of constant stress accelerated life test for generalized inverse
Lindley distribution. Quality and Reliability Engineering International 38(1),
211–228.
18. Feng, X., Tang, J., Balakrishnan, N., & Tan, Q. (2024). Reliability inference
for dual constant-stress accelerated life test with exponential distribution and
progressively type-II censoring. Journal of Statistical Computation and
Simulation, 1–28.
19. Miller, R., & Nelson, W. (1983). Optimum simple step-stress plans for
accelerated life testing. IEEE Transactions on Reliability 32(1), 59–65.
20. Y. Tang, Q. Guan, P. Xu., & H. Xu. (2012). Optimum design for type-I step-
stress accelerated life tests of two-parameter Weibull distribution.
Communications in Statistics-Theory and Methods 41(21), 3863–3877.
21. Balakrishnan, N., Kundu, D., Ng, K.T., & Kannan, N. (2007). Point and
interval estimation for a simple step-stress model with type-II censoring.
Journal of Quality Technology 39(1), 35–47.
22. Mohie El-Din, M.M., Abd El-Raheem, A.M., & Abd El-Azeem, S.O. (2021).
On step-stress accelerated life testing for power generalized Weibull
distribution under progressive type-II censoring. Annals of Data Science 8,
629–644.
23. Alotaibi, R., Almetwally, E.M., Kumar, D., & Rezk H. (2022). optimal test
plan of step-stress model of alpha power Weibull lifetimes under progressively
type-II censored samples. Symmetry 14(9), 1801.
24. Ling, M.H., & Hu, X. (2020). Optimal design of simple step-stress accelerated
life tests for one-shot devices under Weibull distributions. Reliability
Engineering & System Safety 193, 106630.
25. Abdel-Hamid, A.H., & Abushal, T.A. (2015). Inference on progressive-stress
model for the exponentiated exponential distribution under type-II progressive
13
hybrid censoring. Journal of Statistical Computation and Simulation 85(6),
1165–1186.
26. Mohie El-Din, M.M., Abu-Youssef, S.E., Ali, N.S.A., & Abd El-Raheem,
A.M. (2017). Classical and Bayesian inference on progressive-stress
accelerated life testing for the extension of the exponential distribution under
progressive type-II censoring. Quality and Reliability Engineering
International 33(8), 2483–2496.
27. Mohie El-Din, M.M., Abd El-Raheem, A.M., & Abd El-Azeem, S.O. (2018).
On progressive-stress accelerated life testing for power generalized Weibull
distribution under progressive type-II censoring. Journal of Statistics
Applications & Probability Letters 5(3), 131–143.
28. Kumar Mahto, A., Dey, S., & Mani Tripathi, Y. (2020). Statistical inference on
progressive-stress accelerated life testing for the logistic exponential
distribution under progressive type-II censoring. Quality and Reliability
Engineering International 36(1), 112–124.
29. Zarrin, S., Kamal, M., & Saxena, S. (2012). Estimation in constant stress
partially accelerated life tests for Rayleigh distribution using type-I censoring.
Reliability: Theory & Applications 7(4 (27)), 41–52.
30. A.A. Ismail (2013). Estimating the generalized exponential distribution
parameters and the acceleration factor under constant-stress partially
accelerated life testing with type-II censoring. Strength of Materials 45, 693–
702
31. Srivastava, P., & Mittal, N. (2013). Optimum constant-stress partially
accelerated life tests for the truncated logistic distribution under time
constraint. International Journal of Operational Research/Nepal 2, 33–47.
32. Jaheen, Z., Moustafa, H., & Abd El-Monem, G. (2014). Bayes inference in
constant partially accelerated life tests for the generalized exponential
distribution with progressive censoring. Communications in Statistics-Theory
and Methods 43(14), 2973–2988.
33. Abushal, T. A., & Soliman, A.A. (2015). Estimating the Pareto parameters
under progressive censoring data for constant-partially accelerated life tests.
Journal of Statistical Computation and Simulation 85(5), 917–934 (2015)
34. Ismail, A.A., & Al Tamimi, A. (2017). Optimum constant-stress partially
accelerated life test plans using type-I censored data from the inverse Weibull
distribution. Strength of Materials 49, 847–855.
35. Ismail, A.A. (2016). Statistical inference for a step-stress partially-accelerated
life test model with an adaptive type-I progressively hybrid censored data
from Weibull distribution. Statistical Papers 57(2), 271–301.
36. Soliman, A.A., Ahmed, E.A., Abou-Elheggag, N., & Ahmed, S.M. (2017).
Step-stress partially accelerated life tests model in estimation of inverse
Weibull parameters under progressive type-II censoring. Appl. Math 11(5),
1369–1381.

14
37. Abushal, T. (2017). Estimations in step-stress partially accelerated life tests for
the compound Weibull-gamma distribution under type-I censoring. Journal of
Statistics Applications & Probability. 6, 271–283.
38. EL-Sagheer, R.M., Mahmoud, M.A., & Nagaty, H. (2019). Inferences for
Weibull-exponential distribution based on progressive type-II censoring under
step-stress partially accelerated life test model. Journal of Statistical Theory
and Practice 13, 1–19.
39. Akgul, F., Yu, K., & Senoglu, B. (2020). Classical and Bayesian inferences in
step-stress partially accelerated life tests for inverse Weibull distribution under
type-I censoring. Strength of Materials 52, 480–496.
40. Pandey, A., Kaushik, A., Singh, S.K., & Singh, U. (2021). Statistical analysis
for generalized progressive hybrid censored data from Lindley distribution
under step-stress partially accelerated life test model. Austrian Journal of
Statistics 50(1), 105–120.
41. N.M. Sedyakin (1966). On One Physical Principle in Reliability Theory.
Techn. Cybernetics 3, 80–87.
42. W. Nelson (1980). Accelerated Life Testing Step-Stress Models and Data
Analyses. IEEE transactions on reliability 29(2), 103–108.
43. Goel, P. K. (1972). Some Estimation Problems in the Study of Tampered
Random Variables, 2436–2436.
44. Bhattacharyya, G. K., & Soejoeti, Z. (1989). A Tampered Failure Rate Model
for Step-Stress Accelerated Life Test. Communications in statistics-Theory
and methods 18(5), 1627–1643.
45. Van Dorp, J.R., Mazzuchi, T.A., Fornell, G.E., & Pollock, L.R. (1996). A
Bayes Approach to Step-Stress Accelerated Life Testing. IEEE Transactions
on Reliability 45(3), 491–498.
46. Kannan, N., Kundu, D. & Balakrishnan, N. (2010). Survival models for step-
stress experiments with lagged effects. Advances in Degradation Modeling:
Applications to Reliability, Survival Analysis, and Finance, 355–369.
47. Cox, D. R. (1959). The analysis of exponentially distributed life-times with
two types of failure. Journal of the Royal Statistical Society Series B:
Statistical Methodology 21(2), 411–421.
48. Prentice, R.L., Kalbfleisch, J.D., Peterson Jr, A.V., Flournoy, N., Farewell,
V.T., & Breslow, N.E. (1978). The analysis of failure times in the presence of
competing risks. Biometrics, 541–554.
49. Kundu, D. (2004). Parameter estimation for partially complete time and type
of failure data. Biometrical Journal: Journal of Mathematical Methods in
Biosciences 46(2), 165–179.
50. Crowder, M. (2001) Classical Competing Risks. Chapman and Hall/CRC.
51. Crowder, M.J. (2012). Multivariate Survival Analysis and Competing Risks.
CRC Press.
52. Klein, J.P., & Basu, A.P. (1981). Accelerated life testing under competing
exponential failure distributions.
15
53. Pascual, F. (2008). Accelerated life test planning with independent Weibull
competing risks. IEEE Transactions on Reliability 57(3), 435–444.
54. Balakrishnan, N., & Han, D. (2008). Exact inference for a simple step-stress
model with competing risks for failure from exponential distribution under
type-II censoring. Journal of Statistical Planning and Inference 138(12), 4172–
4186.
55. Han, D., & Kundu, D. (2014). Inference for a step-stress model with
competing risks for failure from the generalized exponential distribution under
type-I censoring. IEEE Transactions on reliability 64(1), 31–43.
56. Srivastava, P.W., Shukla, R., & Sen, K. (2014). Optimum simple step-stress
test with competing risks for failure using Khamis-Higgins model under type-
II censoring. International Journal of Operational Research/Nepal 3, 75–88.
57. Liu, F., Shi, Y. (2017). Inference for a simple step-stress model with
progressively censored competing risks data from Weibull distribution.
Communications in Statistics-Theory and Methods 46(14), 7238–7255.
58. Ganguly, A., & Kundu, D. (2016). Analysis of simple step-stress model in
presence of competing risks. Journal of Statistical Computation and
Simulation 86(10), 1989–2006.
59. Basak, I. (2014). Prediction of times to failure of censored items for a simple
step-stress model with regular and progressive type-I censoring from the
exponential distribution. Communications in Statistics-Theory and Methods
43(10-12), 2322–2341.
60. Basak, I., & Balakrishnan, N. (2017). Prediction of censored exponential
lifetimes in a simple step-stress model under progressive type-II censoring.
Computational Statistics 32, 1665–1687.
61. Basak, I., & Balakrishnan, N. (2018). A Note on the Prediction of Censored
Exponential Lifetimes in a Simple Step-Stress Model with Type-II Censoring.
Calcutta Statistical Association Bulletin 70(1), 57–73.
62. Amleh, M.A., & Raqab, M.Z. (2021). Inference in simple step-stress
accelerated life tests for type-I censoring Lomax data. Journal of Statistical
Theory and Applications 20(2), 364–379 (2021)
63. Amleh, M., & Raqab, M.Z. (2022). Prediction of censored Weibull lifetimes in
a simple step-stress plan with Khamis-Higgins model. Statistics, Optimization
& Information Computing 10(3), 658–677.

16
Wrapped three parameter Lindley Distribution for
Modelling Circular Data
Smriti Sikha Sarma1, Bishal Gurung2 and Bhanita Das3
*1
Department of Statistics, North-Eastern Hill University,
Shillong, India, smritisarma914@[Link]
2
Department of Statistics, North-Eastern Hill University,
Shillong, India, bishalgurung@[Link]
3
Department of Statistics, North-Eastern Hill University,
Shillong, India, bhanitadas83@[Link]

Abstract:
In this study, a new circular probability model with three parameters namely, the
wrapped three parameter Lindley distribution is proposed. Behaviour of the density
function with variation in the values of the parameters is examined and expressions
for its characteristic function, trigonometric moments and other related descriptive
measures are derived. The method of maximum likelihood is used to estimate the
unknown parameters of this distribution. Finally, the proposed model is fitted to two
real-life circular data sets and the goodness-of-fit of this distribution is assessed in
comparison to wrapped normal, wrapped stable and weighted von Mises
distributions to elucidate the modelling potential of the proposed distribution.
Finally, the proposed distribution has been fitted to a real dataset and its goodness-
of-fits are demonstrated.
Keywords: Three-parameter Lindley distribution, wrapped distribution,
trigonometric moments, circular data.

1. INTRODUCTION:
Circular data also known as two-dimensional directional data, corresponds to
those occurrences where the observations are recorded in terms of radians or
degrees. Circular data, being directions have no magnitude, and therefore, are
conveniently represented as points on a circle of unit radius, centered at the origin or
as a unit vector in the plane, connecting the origin to the corresponding point
(Jammalamadaka and SenGupta 2001). Common instances of record, are evident in
various natural and physical sciences such as geology (orientations of cross-beds in
rivers, measured in degrees), meteorology (wind direction), biology (vanishing
angles of birds soon after their release) (Schmidt-Koenig 1963), etc. Analyzing the


Corresponding author: smritisarma914@[Link]

17
behavior of these situations requires a special class of distributions, typically known
as circular probability distributions.
The wrapping of linear distributions around a unit circle, generates a rich and
useful class of circular models known as wrapped distributions. In this approach, a
linear random variable (r.v.) is transformed into a circular random variable (r.v.) by
reducing its modulo (Mardia and Jupp 2000). Lévy (1939), first introduced this idea
and obtained wrapped variables from the corresponding symmetric as well as non-
symmetric distributions on the real line. Many authors, since then, have carried out
comprehensive work on wrapped distributions.
A targeted area of interest for the researchers, however, remains modelling
occurrences, where the data is rightly-skewed. In this endeavor, Jammalamadaka and
Kozubowski (2004) developed the wrapped exponential distribution obtained by
wrapping the classical exponential distributions on the real line around the unit
circle. Later, Roy and Adnan (2012) incorporated the concept of weights, and
established a new class of circular distributions namely wrapped weighted
exponential distribution. Recently, Joshi and Jose (2018) projected the wrapped
Lindley density by wrapping a Lindley density and exhibited its superiority over the
wrapped exponential distribution while modelling skewed situations. Empirical
observations have justified the extensive use of Lindley density as a probability
model against various other life-time distributions. The two-parameter Lindley
density (Shanker et al. 2013), being a generalization of the Lindley and exponential
density also offers similar advantage over the other existing distributions. Similarly,
the three parameter Lindley was introduced by Alkarni, S. H. (2015). Therefore, in
order to investigate its usefulness as a circular model, the wrapped version of the
three-parameter Lindley distribution is introduced through the classical wrapping
approach and is named as wrapped three-parameter Lindley distribution.
In this study we represent a new circular distribution called the wrapped three
parameter Lindley distribution (WTPLD) using the wrapping method which was
explained in Jammalamadaka and SenGupta (2001). We have discussed various
distributional properties of the proposed circular model. Additionally, we provide
two real datasets for analysing the suggested distribution's performance against three
alternative models.
2. Wrapped Three Parameter Lindley Distribution (WTPLD):
In this section, a new circular probability model is proposed, which is obtained
by wrapping a Three-parameter Lindley distribution defined on,  around the
circumference of a circle with unit radius. Alternatively, the distribution is also
obtainable by a two-component mixing of the wrapped exponential distribution and
the wrapped gamma distribution. A general definition of this distribution is provided
which subsequently introduces its density function.
Definition 1: Let 𝑋 be a random variable on the real line with density function 𝑓(𝑋).
The wrapped circular variable 𝜑 corresponding to X can be obtained by defining
(Jammalamadaka and SenGupta 2001) 𝜑 = 𝑋,𝑚𝑜𝑑 2𝜋-
18
Proposition 1 Let 𝜑 be a circular random variable following the wrapped three-
parameter Lindley probability law with parameters  𝛿 𝑎𝑛𝑑 𝛽 denoted by 𝜑 ∼
𝑊𝑇𝑃𝐿𝐷(𝜃, 𝛿, 𝛽). The probability density function of 𝜑 is then given by
𝜃2 𝛿 + 𝛽𝜑 2𝜋𝛽𝑟
𝑓(𝜑) = 𝑒 −𝜃𝜑 [ + ] , 0 ≤ 𝜑 < 2𝜋
𝜃𝛿 + 𝛽 1−𝑟 (1 − 𝑟)2
Proof:
Let X follow a three-parameter Lindley distribution (TPLD) (Alkarni, S. H.
(2015)). Then, the probability density function (p.d.f.) of X is defined by,
𝜃𝑚 𝛽
𝑔(𝑥; 𝜃, 𝛿, 𝛽) = (𝛿 + 𝛽𝑥)𝑒 −𝜃𝑥 , where 𝑥 > 0, 𝜃 > 0, 𝛽 > 0, 𝛿 > −
𝜃𝛿+𝛽 𝜃

In the wrapping approach, the circular (wrapped) three-parameter Lindley


variate 𝜑 is generated using the relation in definition (1) and accordingly the pdf of
WTPLD (𝜃, 𝛿, 𝛽) is obtained by wrapping 𝑓(𝑦) around the circumference of a circle
of unit radius as follows (Jammalamadaka and SenGupta 2001):
𝑓(𝜑) = ∑∞
𝑘=0 𝑓(𝜑 + 2𝜋𝑘)
𝜃𝑚
= 𝜃𝛿+𝛽 ∑∞
𝑘=0(𝛿 + 𝛽(𝜑 + 2𝜋𝑘))𝑒
−𝜃(𝜑+2𝜋𝑘)

Factor out 𝑒 −𝜃𝜑 and set 𝑟 = 𝑒 −2𝜋𝜃 to evaluate the geometric sums

𝜃2
𝑓(𝜑) = ∑ 𝑒 −𝜃𝜑 (𝛿 + 𝛽(𝜑 + 2𝜋𝑘)) 𝑟 𝑘
𝜃𝛿 + 𝛽
𝑘=0
1 𝑟
Use the standard sums ∑∞ 𝑘 ∞ 𝑘
𝑘=0 𝑟 = 1−𝑟 and ∑𝑘=0 𝑘𝑟 = (1−𝑟)𝑚 (0 < 𝑟 <
1 𝑖, 𝑒 𝜃 > 0)
𝜃2 𝛿 + 𝛽𝜑 2𝜋𝛽𝑟
𝑓(𝜑) = 𝑒 −𝜃𝜑 [ + ] , 0 ≤ 𝜑 < 2𝜋
𝜃𝛿 + 𝛽 1−𝑟 (1 − 𝑟)2

Fig. 1: Pdf plot of WTPLD (𝜃, 𝛿, 𝛽) distribution with different parameter values

19
3. Properties of Wrapped Three-Parameter Lindley Distribution:
In this section, the expressions for the characteristic function, trigonometric
moments, coefficient of skewness and kurtosis and the median of the WTPLD
(𝜃, 𝛿, 𝛽) are derived.

3.1. Characteristic Function:


The characteristic function of the WTPLD (𝜃, 𝛿, 𝛽) distribution is given by,
2𝜋
Φ𝜑 (𝑝) = 𝐸[𝑒 𝑖𝑝𝜑 ] = ∫ 𝑒 𝑖𝑝𝜙 𝑓(𝜙)𝑑𝜙
0
Let 𝑎 = 𝑖𝑝 − 𝜃. Then,
𝜃2 1 2𝜋
2𝜋𝛽𝑟 2𝜋
Φ𝜑 (𝑝) = 6 ∫ 𝑒 𝑎𝜙 (𝛿 + 𝛽𝜙)𝑑𝜙 + ∫ 𝑒 𝑎𝜙 𝑑𝜙7
𝜃𝛿 + 𝛽 1 − 𝑟 0 (1 − 𝑟)2 0
Now to compute the elementary integrals (with 𝑎 = 𝑖𝑝 − 𝜃)
2𝜋
𝑒 𝑎2𝜋 − 1
∫ 𝑒 𝑎𝜙 𝑑𝜙 =
0 𝑎
2𝜋 𝑎2𝜋
𝑒 −1
∫ 𝑒 𝑎𝜙 𝛿 𝑑𝜙 = 𝛿
0 𝑎
2𝜋 𝑎2𝜋
2𝜋𝑒 𝑒 𝑎2𝜋 1
∫ 𝑒 𝑎𝜙 𝛽𝜙 𝑑𝜙 = 𝛽 4 − 2 + 25
0 𝑎 𝑎 𝑎
Because, p is an integer we have 𝑒 2𝜋𝑖𝑝 = 1. Hence, 𝑒 𝑎2𝜋 = 𝑒 2𝜋(𝑖𝑝−𝜃) = 𝑒 −2𝜋𝜃 = 𝑟
Plugging 𝑒 𝑎2𝜋 = 𝑟 into the integrals gives
2𝜋
𝑟−1
∫ 𝑒 𝑎𝜙 𝑑𝜙 =
0 𝑎
2𝜋
𝑟−1 2𝜋𝑟 𝑟 1
∫ 𝑒 𝑎𝜙 (𝛿 + 𝛽𝜙)𝑑𝜙 = 𝛿 +𝛽( − 2 + 2*
0 𝑎 𝑎 𝑎 𝑎
After substituting the values in the equation Φ𝜑 (𝑝) and simplifying the
obtained form is as follows,
𝜃 2 (−(𝛽 + 𝛿𝜃) + 𝑖𝛿𝑝)
Φ𝜑 (𝑝) = ,𝑝 ∈ ℤ
(𝛽 + 𝛿𝜃)(𝑝 + 𝑖𝜃)2
An equivalent form is
𝜃2 𝑖𝛿𝑝
Φ𝜑 (𝑝) = 2
(−1 + *
(𝑝 + 𝑖𝜃) 𝛽 + 𝛿𝜃
3.2. Pth non-central trigonometric moment:
Pth non-central trigonometric moment of circular variable is the characteristic
function evaluated at integer p
𝜇𝑝′ = 𝐸[𝑒 𝑖𝑝𝜑 ] = Φ𝜑 (𝑝), 𝑝∈𝑍

20
Using the compact form obtained above and putting
𝛿𝑝
𝐷 = 𝛽 + 𝛿𝜃, 𝐴 = −1, 𝐵= 𝐷
And 𝑋 = 𝑝2 − 𝜃 2 , 𝑌 = 2𝑝𝜃
𝜃𝑚 (−𝐷+𝑖𝛿𝑝) 𝜃𝑚 (𝐴+𝑖𝐵)
Then, 𝜇𝑝′ = Φ𝜑 (𝑝) = =
𝐷(𝑝+𝑖𝜃)𝑚 𝑋+𝑖𝑌
The Pth non-central trigonometric moment of 𝜑 can be written as
𝜇𝑝′ = 𝛼𝑝 + 𝑖𝛽𝑝
𝛿𝑝 𝛿𝑝
𝜃𝑚 (−𝑋+ 𝑌) 𝜃𝑚 (−(𝑝𝑚 −𝜃𝑚 )+ 2𝑝𝜃)
𝛽+𝛿𝜃
Where, 𝛼𝑝 = 𝑚 𝑚
𝐷
(𝑝 +𝜃 ) 𝑚 = 𝑚 𝑚
(𝑝 +𝜃 ) 𝑚

𝛿𝑝 2 𝛿𝑝 2 2
𝜃 2 ( 𝑋 + 𝑌) 𝜃 (𝛽 + 𝛿𝜃 (𝑝 − 𝜃 ) + 2𝑝𝜃*
𝛽𝑝 = 𝐷 =
(𝑝2 + 𝜃 2 )2 (𝑝2 + 𝜃 2 )2
3.3. Mean resultant length:
The non-central moment 𝜇1′ and 𝜇2′ are given by,
𝜃 2 (−(𝛽 + 𝛿𝜃) + 𝑖𝛿)
𝜇1′ =
(𝛽 + 𝛿𝜃)(1 + 𝑖𝜃)2
𝜃𝑚 (−(𝛽+𝛿𝜃)+𝑖2𝛿)
and 𝜇2′ = (𝛽+𝛿𝜃)(2+𝑖𝜃)𝑚

𝜌 = |𝜇1′ | = √(ℜ(𝜇1′ ))2 + (ℑ(𝜇1′ ))2


2𝛿𝜃
𝜃 2 −(1 − 𝜃 2 ) +
𝛽 + 𝛿𝜃
ℜ(𝜇1′ ) = . 2 2
(𝛽 + 𝛿𝜃) (1 + 𝜃 )
𝛿(1 − 𝜃 2 )
𝜃2 + 2𝜃
′) 𝛽 + 𝛿𝜃
ℑ(𝜇1 = .
(𝛽 + 𝛿𝜃) (1 + 𝜃 2 )2

3.4. Circular Variance:


The circular variance of 𝜑 ∼ 𝑊𝑇𝑃𝐿𝐷(𝜃, 𝛿, 𝛽) denoted by V is obtained by the
relation
𝑉 = 1−𝜌

3.5. Skewness and Kurtosis:


The measures of skewness and kurtosis, denoted by 𝜉10 and 𝜉20 respectively is
defined as
̅̅̅̅
𝛽 𝛼𝑚 𝑜
̅̅̅̅−𝜌
𝜉10 = 𝑉 𝑛/𝑚
𝑚
and 𝜉20 = 𝑉𝑚

The expressions for the above measures are large but however, can be obtained by
using the results derived in the above equations.

21
4. Maximum Likelihood Estimation:
This section discusses the estimation of the parameters of WTPLD(𝜃, 𝛿, 𝛽)
through the maximum likelihood method.
Let 𝜙1 , 𝜙2 , … 𝜙𝑛 𝜖,0, 2𝜋) be a random sample of size n from the proposed
WTPLD(𝜃, 𝛿, 𝛽).
𝑛 𝑛 𝑛 𝑛
𝜃2 𝛿 + 𝛽𝜑𝑖 2𝜋𝛽𝑟
𝐿(𝜃, 𝛿, 𝛽) = ∏ 𝑓(𝜑𝑖 ) = 4 5 exp (−𝜃 ∑ 𝜑𝑖 + ∏ [ + ]
𝜃𝛿 + 𝛽 1−𝑟 (1 − 𝑟)2
𝑖=1 𝑖=1 𝑖=1
Then the log likelihood function is given by,
𝑛 𝑛

ℓ(𝜃, 𝛿, 𝛽) = 𝑛(2 log 𝜃 − log(𝜃𝛿 + 𝛽)) − 𝜃 ∑ 𝜑𝑖 + ∑ log 𝐴𝑖 (𝜃, 𝛿, 𝛽)


𝑖=1 𝑖=1
𝛿+𝛽𝜑𝑖 2𝜋𝛽𝑟 −2𝜋𝜃
where 𝐴𝑖 (𝜃, 𝛿, 𝛽) = 1−𝑟
+ (1−𝑟)𝑚
,𝑟 =𝑒
𝜕𝑟
𝜕𝜃
= −2𝜋𝑟, 𝑢 =1−𝑟
𝜕𝐴𝑖 1 𝜕𝐴𝑖 𝜑𝑖 2𝜋𝑟
𝜕𝜃
= 𝑢, 𝜕𝛽
= 𝑢
+ 𝑢𝑚
𝜕𝐴𝑖 𝛿+𝛽𝜑 1 2𝑟 𝛿+𝛽𝜑 1+𝑟
Now computation of 𝜕𝑟
= 𝑢𝑚 𝑖
+ 2𝜋𝛽 .𝑢𝑚 + + 2𝜋𝛽 . 𝑢𝑛 / 𝑢𝑛
/ = 𝑢𝑚 𝑖
Then by chain rule,
𝜕𝐴𝑖 𝜕𝐴𝑖 𝜕𝑟 𝛿 + 𝛽𝜑𝑖 1+𝑟
= = −2𝜋𝑟 [ 2
+ 2𝜋𝛽 ( 3 *]
𝜕𝜃 𝜕𝑟 𝜕𝜃 𝑢 𝑢
𝜕 𝛿 𝜕 𝜃 𝜕 1
Also, 𝑙𝑛(𝜃𝛿 + 𝛽) = , 𝑙𝑛(𝜃𝛿 + 𝛽) = , 𝑙𝑛(𝜃𝛿 + 𝛽) =
𝜕𝜃 𝜃𝛿+𝛽 𝜕𝛿 𝜃𝛿+𝛽 𝜕𝛽 𝜃𝛿+𝛽
𝜕ℓ 𝑛𝜃 1
Now, 𝑆𝛿 = =− + ∑𝑛𝑖=1
𝜕𝛿 𝜃𝛿+𝛽 𝑢𝐴𝑖
𝑛
𝑛 1 𝜑𝑖 2𝜋𝑟
𝑆𝛽 = − +∑ ( + 2 *
𝜃𝛿 + 𝛽 𝐴𝑖 𝑢 𝑢
𝑖=1
𝑛
𝜕ℓ 2𝑛 𝑛𝛿 1 𝜕𝐴𝑖
𝑆𝜃 = = − − 𝑆𝜑 + ∑
𝜕𝜃 𝜃 𝜃𝛿 + 𝛽 𝐴𝑖 𝜕𝜃
𝑖=1
These are three nonlinear equations, we have to solve the equations numerically.

5. Applications:
This section comprises of applying the wrapped three-parameter Lindley
distribution to two real life directional data set having an extended right tail. The fit
of this distribution is compared to that of the wrapped normal distribution, wrapped
stable distribution and weighted von-Mises with the help of the statistics- log
likelihood, AIC (Akaike information criterion) and BIC (Bayesian information
criterion). A comparatively smaller value of the AIC and BIC test statistic would
construe of a better fit to the data set.
The first data set considered here, is the Sun compass orientations of 50 starhead
topminnows, measured under heavily overcast conditions, which is procured from
Goodyear (1970, Figure 1D) and published in Fisher (1993), Appendix B4.

22
Table I: Log-likelihood and AIC, BIC values for the dataset 1
Distributions Log-Likelihood AIC BIC
WTPLD -91.8939 189.788 195.524
Wrapped Normal -91.8938 189.878 197.611
Wrapped Stable -90.2231 195.446 196.782
Weighted vM -181.740 371.481 379.129

Table I summarizes the findings of the fitting of each of the four distributions to the
data set. Smallest values of the AIC and BIC statistic for WTPLD(𝜃, 𝛿, 𝛽) clearly
demonstrate that the proposed distribution provides the best fit to the data under
consideration.

Fig. 2: Empirical and fitted cumulative distribution function plot for the dataset 1.

The second data set considered here, is the Groove and tool marks, and flute marks,
measured in the Murruin Creek area, Southwest of Yerranderie, New South Wales,
which is procured from Fisher & Powell (1989) and published in Fisher (1993),
Appendix B15.
Table II : Log-likelihood and AIC, BIC values for the dataset 2
Distributions LogLik AIC BIC
WTPLD -54.7511 125.502 130.253
Wrapped Normal -63.75661 131.5132 134.6803
Wrapped Stable -60.25120 126.5024 134.2529
Weighted vM -120.12112 248.2422 254.5763
Table II summarizes the findings of the fitting of the four distributions to the second
data set. Smaller values of the AIC and BIC statistic for WTPLD(𝜃, 𝛿, 𝛽) clearly

23
demonstrate that the proposed distribution provides the better fit to the data under
consideration.

Fig. 3: Empirical and fitted cumulative distribution function plot for the dataset 2

6. Conclusion:
Through this manuscript, a new circular probability model viz. wrapped three-
parameter Lindley distribution (WTPLD) is proposed, which is obtained by
wrapping a three-parameter Lindley distribution on around a unit circle.
Alternatively, the proposed distribution is also shown to be obtainable through a
finite mixture of wrapped exponential and wrapped gamma distribution. A few
special cases of the WTPLD are showcased and the pictorial behavior of the density
function for varying values of the parameters is illustrated. Expressions for the
characteristic function, trigonometric moments and other related descriptive
measures of this distribution are derived. Parameter estimation is carried out by the
method of maximum likelihood. Finally, the proposed distribution is applied to two
real-life directional data sets and the goodness-of-fit of the distribution is assessed
and compared to that of the wrapped exponential and wrapped Lindley distribution
with the help of the log-likelihood, AIC and BIC measures. Results indicate that
WTPLD is a better fit and more flexible as a model as against the others, for
modelling the situations where the directions having lower magnitude have higher
likelihood of occurrence.

24
References:
Alkarni, S. H. (2015). Extended power Lindley distribution: A new statistical
model for non-monotone survival data. European Journal of Statistics and
Probability, 3(3), 19-34.
Fisher, N.I. (1993). Statistical analysis of circular data. Cambridge: Cambridge
University Press.
Goodyear, C. P. (1970) Terrestrial and aquatic orientation in the starhead
topminnow, Fundulusnotti. Science. 168 (3931):603-605.
Jammalamadaka, R.S. & Sengupta, A. (2001) Topics in Circular Statistics.
World Scientific, Singapore.
Jammalamadaka, S. R., & Kozubowski, T. J. (2001). A wrapped exponential
circular model. In Proc. of AP Academy of Sciences, 5(1), 43-56.
Jammalamadaka, S. R., & Kozubowski, T. J. (2004). New families of wrapped
distributions for modeling skew circular data. Communications in
Statistics-Theory and Methods, 33(9), 2059-2074.
Joshi, S., & Jose, K. K. (2018). Wrapped lindley distribution. Communications
in Statistics-Theory and Methods, 47(5), 1013-1021.
Lévy, P. (1939). L'addition des variables aléatoires définies sur une
circonférence. Bulletin de la Société mathématique de France, 67, 1-41.
Lund, U., & Agostinelli, C. (2018). CircStats: circular statistics, from ―topics in
circular statistics‖. Available in: [Link] R-project. org/package=
CircStats. Accessed on: July, 13, 2021.
Lund, U. (2010). Circular: circular statistics. R package version
0.4. [Link] R-project. org/package= circular.
Mardia, K. V., & Jupp, P. E. (2009). Directional statistics. John Wiley & Sons.
Roy, S., & Adnan, M. A. S. (2012). Wrapped weighted exponential
distributions. Statistics & probability letters, 82(1), 77-83.
Schmidt-Koenig, K. (1963). On the role of the loft, the distance and site of
release in pigeon homing (the" cross-loft experiment"). The Biological
Bulletin, 125(1), 154-164.
Shanker, R., Sharma, S., & Shanker, R. (2013). A two-parameter Lindley
distribution for modeling waiting and survival times data. Applied
mathematics, 4(2), 363-368.

25
A Comparative Analysis of Different Statistical Learning
Models for Estimation of the Non-Destructive Shelf-Life
Prediction of Okra
(lady‘s finger)

Joy Deb1, Dibyojyoti Bhattacharjee 2


*1
Department of Statistics, Assam University, Silchar, jdeb48389@[Link]
2
Department of Statistics, Assam University, Silchar, [Link]@[Link]

Abstract:
The shelf life of fresh perishable fruits and vegetables is a critical determinant of
post-harvest quality, marketability and consumer acceptance. In the case of okra,
rapid deterioration under ambient conditions creates a major challenge for farmers,
retailers and consumers. This study proposes a non-destructive framework for
predicting the shelf life of okra using non- destructive attributes including weight
loss, dimensional shrinkage, and colour transformation. A dataset was developed
through controlled storage experiments and different statistical learning models such
as Linear Regression, Ridge Regression, Lasso Regression, Decision Tree, k-Nearest
Neighbors (k-NN), Support Vector Regression (SVR), Random Forest and XG-
Boost were applied for predictive analysis. Model performance was evaluated using
Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and the coefficient
of determination (R²), under both test data and cross-validation approaches. The
results revealed that SVR achieved the highest predictive accuracy (R² = 0.65),
followed by Random Forest and k-NN, whereas traditional linear models yielded
moderate performance. These findings demonstrate that non-destructive features can
effectively estimate okra shelf life, thereby supporting improved supply chain
management, reducing food waste, and ensuring consumer satisfaction. The
proposed framework provides a practical foundation for the application of machine
learning in post-harvest technology and may be extended to other perishable
horticultural commodities as well.
Key Words: Machine Learning, Non-Destructive Methods, Okra, Shelf-Life
Prediction, Statistical Models

1. INTRODUCTION
Okra (Abelmoschus esculentus L. Moench), commonly known as lady‘s finger,
is an important vegetable crop cultivated widely in tropical and subtropical regions.


Corresponding author: jdeb48389@[Link]

Corresponding author: magrishanam123@[Link]
26
It is valued for its nutritional richness which is a good source of vitamins, minerals,
dietary fiber, and bioactive compounds (Elkhalifa et al., 2021). However, like most
perishable vegetables, okra exhibits a short post-harvest life with rapid physiological
changes such as weight loss, shrinkage, discoloration, and textural degradation that
reduce its market acceptability and consumer preference (Jain et al., 2014).
Predicting shelf-life is particularly important for consumers, who prioritize
longer shelf-life along with superior nutritional content and freshness (Weng et al.,
2020). In the case of okra, whose quality deteriorates rapidly under ambient storage,
accurate estimation of shelf-life is important for ensuring that it remains edible for
consumption. Reliable prediction of shelf-life enables farmers, retailers, and other
stakeholders to make timely decisions across the supply chain, which implies
maximising nutritional retention, minimising post-harvest losses, and increasing
economic returns (Mohammed et al., 2022). In particular, for retailers, shelf-life
prediction allows for optimal product display and pricing strategies, which not only
reduce waste but also enhance consumer satisfaction and loyalty (Iqbal et al., 2024).
Traditionally, shelf-life evaluation of vegetables has been conducted through
destructive methods such as chemical assessment, biochemical markers or invasive
sampling of tissues (Chhetri., 2024). While these methods provide precise insights,
they are labour-intensive, time-consuming and often unsuitable for real-time
monitoring of large quantities of produce. Furthermore, destructive approaches
reduce the sample utility, making them less practical for small-scale farmers and
fresh produce industries (Karunathilake et al. 2023).
Non-destructive approaches offer a more efficient solution based on
morphological and colorimetric attributes that can be recorded without damaging the
sample. With the advancement of computational techniques, which provide the
ability to map these measurable features to the storage duration and predict the
remaining shelf life. Such techniques are advantageous due to their cost-
effectiveness, scalability and ability to generalise predictions across variations in the
samples (Xiao et al. 2023).
The present study explores the application of statistical learning models for the
non-destructive prediction of okra shelf life. By leveraging morphological and
colour parameters as predictors, we aim to develop a regression-based framework
for estimating the edible stage of okra over multiple storage days. This approach not
only reduces reliance on destructive testing but also enhances the practical
applicability of shelf-life assessment in real-world scenarios.

2. LITERATURE SURVEY
Postharvest quality assessment and shelf-life prediction of horticultural products
have received increasing research attention in recent years. Various machine learning
(ML) and deep learning approaches have been applied across fruit and vegetable
commodities to model respiration rate, weight loss, sensory quality, and other
physiological attributes under diverse storage and supply chain conditions.

27
Ogbaji and Iorliam (2020) conducted a laboratory experiment on okra fruits and
reported that the shelf life of treated samples ranged from 1 to 15 days, compared
with 1 to 7 days in untreated control fruits. Their dataset forms a useful basis for
predictive modelling in subsequent research on okra shelf life.
Hu et al. (2024a) investigated avocado storage behaviour, applying LASSO
classification to integrate temperature, ethylene production rate, and respiration rate
for predicting fruit edibility, achieving accuracies between 72–91%. Another
dominant research stream has addressed weight loss and freshness under supply
chain and cold storage conditions. In blueberries, mechanical texture parameters and
fruit diameter were employed in partial least squares regression (PLSR) to estimate
weight loss, with strong model performance investigated by Oh et al. (2024).
Likewise, Ktenioudaki et al. (2021) investigated storage condition variables, such as
temperature and humidity, which were modelled using boosted regression trees to
predict weight loss in supply chains. In grapes, CatBoost was implemented to
predict comprehensive quality parameters, including mass loss, firmness, titratable
acidity, soluble solids content, vitamin C, stem strength, and sensory traits,
conducted by Chen et al. (2024).
Fu et al. (2024) used linear regression to estimate storage days of passion fruit
based on pericarp hardness, thickness, and fruit mass. Spectral analysis has emerged
as a major non-destructive approach for predicting quality traits during storage.
Several studies employed visible near-infrared and short-wave NIR spectroscopy
combined with ML models. For apples, Hu et al. (2024b) applied PLSR with k-
means clustering on optimal VIS/NIR wavelengths to predict storage time. Further
studies extended this to shelf-life prediction, Joshi et al. (2022), using NIR
spectroscopy for predicting storage days. In tomatoes, Sharma et al. (2024) used
short-wave NIR spectral features with ANN to classify lycopene content, achieving
95% accuracy.
Finally, multi-parameter freshness modelling has been explored Owoyemi et al.
(2022) used support vector regression (SVR) to predict orange fruit acceptance
scores from pre-harvest, storage time, and environmental conditions. Kasampalis et
al. (2021) applied multiple linear regression (MLR) and SVR for bell peppers,
achieving prediction accuracy for freshness classification.
While numerous studies have investigated the shelf life of fruits and vegetables,
research on okra shelf-life prediction using non-destructive approaches remains
scarce. Existing works often depend on biochemical assessment, spectral analysis
and other destructive techniques. Moreover, prior machine learning applications
have predominantly focused on classification tasks with limited exploration of
regression-based prediction frameworks that capture the continuous progression of
shelf-life. Importantly, very few studies have conducted a systematic comparison
across multiple statistical learning models. Such comparison is crucial because
model performance can vary depending on the nature of predictors, the degree of
nonlinearity in postharvest changes, and the size or variability of datasets. By
benchmarking different models ranging from linear regressors to tree-based and
28
kernel-based approaches, this study identifies not only the most accurate predictors
of okra shelf-life but also highlights the statistical opportunities and limitations of
each method.

3. METHODOLOGY
3.1 Sample Collection and Experimental Design
Fresh okra samples were procured from local markets by ensuring uniform
maturity stage and absence of visible defects. The samples were stored under
controlled ambient conditions and monitored daily for a predetermined shelf-life
duration. A completely randomised design was adopted, where each fruit sample was
treated as an independent experimental unit. For each sample, morphological
parameters were systematically recorded, including length (cm), width (cm), weight
(g), aspect ratio (length/width), RGB values, and shelf-life days, which represent the
number of days until the fruit loses edible quality. Data were collected on a day-wise
basis to capture the progression of quality deterioration throughout storage.

3.2 Mathematical Interpretation of Shelf-Life Prediction


The shelf-life of fresh vegetables is influenced by a combination of physical and
biochemical changes that occur during storage. These changes can be captured
through measurable parameters such as weight, size (length and width), and colour
attributes, which act as reliable non-destructive indicators of deterioration (Jain et
al., 2014). Mathematical formulations of these parameters are essential because they
provide quantitative evidence of the degradation process, allowing for accurate
modelling and prediction of shelf-life.
3.3.1 Weight Loss Dynamics
The weight is directly proportional to the shelf-life of the okra. The weight
decline was quantified as
𝑤𝑖𝑘 −𝑤𝑖𝑡
𝑊𝐿𝑖 (𝑡) = 𝑤𝑖𝑘
× 100 Eq.(1)

With 𝑤𝑖0 is the initial weight and 𝑤𝑖𝑡 is the weight at day? 𝑡. This follows the
exponential decay model
𝑤𝑖𝑡 = 𝑤𝑖0 . 𝑒 −𝜆𝑡 + 𝜖
Where 𝜆 is the respiration-transpiration constant.

3.3.2 Dimensional Shrinkage


Moisture loss of the vegetable may also affect the size. It is quantified as
𝑙𝑖𝑡 = 𝑙𝑖0 − 𝛿𝑙 𝑡, 𝑑𝑖𝑡 = 𝑑𝑖0 − 𝛿𝑑 𝑡 Eq. (2)
Where, 𝛿𝑙 and 𝛿𝑑 denotes daily shrinkage rates of length and width.

29
3.3.3 Colour Transition
RGB shifts reflect chlorophyll degradation
Δ 𝑅𝑖 (𝑡) = 𝑅𝑖𝑡 − 𝑅𝑖0 , Δ 𝐺𝑖 (𝑡) = 𝐺𝑖𝑡 − 𝐺𝑖0 , Δ 𝐵𝑖 (𝑡) = 𝐵𝑖𝑡 − 𝐵𝑖0
Therefore, the colour index was computed as
𝑅
𝐶𝐼𝑖 (𝑡) = 𝐺𝑖𝑡 Eq. (3)
𝑖𝑡
An increasing 𝐶𝐼 indicates chlorophyll loss and maturity decline.

3.4 Statistical Learning Models


In this study, the dependent variable was defined as the shelf-life of okra (in
days). The independent variables included the set of non-destructive features
extracted during storage. To establish predictive relationships between these features
and the shelf-life, different statistical models were employed such as Multiple Linear
Regression (MLR), Ridge Regression, Lasso Regression, Decision Tree Regressor
(DTR), Random Forest Regressor (RFR), Support Vector Regressor (SVR), k-NN
Regressor, XG Boost Regressor (Rashvand et al., 2025). Linear models were chosen
for their interpretability and ability to handle multicollinearity, with Lasso
additionally performing feature selection. Tree-based models were included to
capture complex non-linear relationships and improve prediction accuracy. SVR was
selected for its robustness in high-dimensional feature spaces, while k-NN provided
a simple similarity-based approach for regression. However, for model development,
the dataset was divided into training and testing sets into 80:20 to evaluate
generalization ability. Additionally, 5-fold cross-validation was applied on the
training data to ensure robustness of model performance and to minimize overfitting.

3.5 Performance Evaluation


The predictive accuracy of the regression models was assessed using multiple
performance indicators. These evaluation metrics provide complementary insights
into how well the predicted shelf-life values matched the observed values.
∑𝑚 ̂𝑖 )𝑚
𝑖=𝑙(𝑦𝑖 −𝑦
𝑅2 = 1 − ∑𝑚 (𝑦 𝑚 Eq. (4)
𝑖=𝑙 𝑖 ̅)
−𝑦
1 1/2
𝑅𝑀𝑆𝐸 = 2𝑚 ∑𝑚
𝑖=1(𝑦𝑖 − 𝑦
̂)2
𝑖 3 Eq. (5)
1 𝑚
𝑀𝐴𝐸 = ∑ |𝑦 − 𝑦̂| Eq. (6)
𝑚 𝑖=1 𝑖 𝑖

3. RESULTS AND DISCUSSION


The experimental results of weight degradation, colour change, and dimensional
shrinkage across storage days are presented in Fig. 1–3. A consistent decline in
weight was observed in Fig. 1, indicating continuous moisture loss over time.
Similarly, the length and width of the samples showed a decreasing trend with minor
fluctuations showed in Fig. 2, reflecting dimensional shrinkage during degradation.
Conversely, the colour index increased steadily in Fig. 3, which confirms

30
progressive pigment transformation. These patterns are in agreement with previous
work conducted by Jain et al., 2014 where storage duration significantly influenced
physical and colorimetric traits of okra.

Fig. 1. Degradation of weight over storage days.

Fig. 2. Changes in dimensional shrinkage (length and width) over storage days.

31
Fig. 3. Changes in color index (R/G) over storage days.

To model these degradation dynamics, different regression approaches were


employed and evaluated using test data and cross-validation performance metrics
presented in Table I. Among the linear models, Lasso Regression performed slightly
better than Multiple Linear and Ridge Regression, achieving the highest 𝑅 2of 0.442
in test data, though overall predictive accuracy remained moderate. Decision Tree
regression performed poorly on the test set 𝑅 2=0.084, highlighting its sensitivity to
overfitting and limited generalization.
Table I. Performance metrics of different regression models for predicting
degradation traits under test data and cross-validation
Model Test Data Cross Validation
MAE RMSE R² MAE RMSE R²
Multiple Linear Regression 0.912 1.078 0.411 0.902 1.240 0.189
Ridge Regression 0.907 1.063 0.427 0.901 1.237 0.193
Lasso Regression 0.902 1.048 0.442 0.903 1.235 0.196
Decision Tree 0.935 1.344 0.084 0.786 1.156 0.299
k-NN Regressor 0.771 0.956 0.536 0.747 0.963 0.508
SVR 0.660 0.827 0.653 0.645 0.821 0.648
Random Forest 0.705 0.867 0.619 0.658 0.837 0.633
XGBoost 0.756 0.973 0.519 0.696 0.914 0.566

The non-linear models outperformed linear approaches. Support Vector


Regression (SVR) achieved the best performance with 𝑅 2= 0.65 in test data and
0.648 in cross-validation, indicating strong robustness and stability. Similarly,

32
Random Forest and k-NN Regressor also showed good predictive power, with
Random Forest achieving 𝑅 2= 0.619 in test data and k-NN 𝑅 2= 0.536. XG
Boost yielded competitive results 𝑅 2=0.519, though slightly lower than Random
Forest and SVR.
However, the error metrics (MAE and RMSE) provide important insights into
the practical reliability of predictions for okra shelf-life. For instance, SVR not only
explained the highest variance but also maintained low error values (MAE = 0.660,
RMSE = 0.827), reflecting its ability to generate shelf-life estimates with minimal
deviation from observed values. In contrast, models with lower 𝑅 2 often had higher
RMSE, which implies greater uncertainty in predicting the actual degradation stage
of okra. Considering the sensitivity of fresh produce supply chains, such error
minimization is crucial, as even small deviations can lead to misjudgment of
marketable quality and economic loss.
The consistent superiority of non-linear methods demonstrates that degradation
patterns are inherently complex and better captured by algorithms capable of
modeling non-linear relationships. Specifically, SVR balanced both explanatory
power and prediction precision, making it the most suitable model for predictive
analysis of okra shelf-life under ambient storage. However, From the Fig. 4, it is
evident that the predictions generally follow the increasing trend of actual shelf-life,
though some scatter around the ideal line is observed, particularly at higher shelf-life
values (4–5 days). This suggests that the employed SVR model successfully
captured the overall deterioration pattern of okra but with moderate deviations in
precision. The clustering of points indicates that the model consistently predicts
within a close range of the actual shelf-life which supports its utility for practical
applications in non-destructive quality assessment.

Fig.4. Predicted vs. Actual Shelf-life of Okra

33
4. CONCLUSION
This study demonstrated the utility of statistical learning models for non-
destructive shelf-life prediction of okra, showing that simple measurable parameters
such as weight loss, dimensional shrinkage and colour transitions can be effectively
mapped to storage duration. From a statistical standpoint, the research contributes by
systematically comparing linear and non-linear regression frameworks for modelling
degradation dynamics in perishable produce. The results highlight that non-linear
approaches such as SVR, Random Forest, k-NN capture complex deterioration
patterns more accurately than linear methods, underscoring the importance of model
flexibility in handling biological variability. This provides a methodological
contribution to applied statistics by illustrating how model selection critically affects
predictive accuracy in agricultural contexts and by demonstrating the practical
application of regression-based learning in food systems where non-destructive,
real-time assessment is essential.
For practical relevance, the findings establish that non-destructive statistical
modelling offers a cost-effective decision-support tool for farmers, retailers, and
supply chain actors, enabling timely interventions, reducing post-harvest waste, and
improving consumer satisfaction. Beyond okra, this framework can be adapted to
other perishable commodities, thereby strengthening the bridge between statistical
learning theory and applied agricultural science.
Future research should aim at enlarging the dataset across cultivars,
environments and handling conditions to improve generalizability. Additionally,
advanced deep learning architectures can be explored for more robust predictions
particularly Recurrent Neural Networks (RNNs) and Long Short-Term Memory
(LSTM) models for time-series shelf-life monitoring, and Convolutional Neural
Networks (CNNs) for image-based feature extraction of external quality attributes
such as texture and colour gradients. Hybrid models that integrate spectral,
morphological, and environmental data could further enhance predictive power,
making shelf-life estimation more precise and scalable in real-world agricultural
supply chains.

REFERENCE
1. Chen, Q., Li, J., Feng, J., & Qian, J. (2024). Dynamic comprehensive quality
assessment of post-harvest grape in different transportation chains using
SAHP-Catboost machine learning. Food Quality and Safety, 8, 1–11.
[Link]
2. Chhetri, K. B. (2024). Applications of artificial intelligence and machine
learning in food quality control and safety assessment. Food Engineering
Reviews, 16(1), 1–21. [Link]
3. Elkhalifa, A. E. O., Alshammari, E., Adnan, M., Alcantara, J. C.,
Awadelkareem, A. M., Eltoum, N. E., Mehmood, K., Panda, B. P., & Ashraf,
S. A. (2021). Okra (Abelmoschus esculentus) as a potential dietary medicine

34
with nutraceutical importance for sustainable health applications. Molecules,
26(3), 696. [Link]
4. Fu, D., Wang, J., Chen, Y., Hu, Z., & Tang, W. (2024). Predicting storage time
of passion fruit using pericarp change index: A study on size, mass and texture
parameters. Journal of Food Measurement and Characterization, 18(5), 3893–
3905. [Link]
5. Hu, J., Liu, D., Zhu, Y., Chen, Z., Zhang, X., Han, X., & Zhou, P. (2024a).
Establishing a maturity prediction model for respiratory fruits via ethylene-
regulated physiology: A case investigation of avocado. Food Bioscience, 59,
104097. [Link]
6. Hu, Y., Qiao, Y., Hou, B., Qu, Z., Zhang, P., Han, R., & Guo, J. (2024b).
Building models to evaluate internal comprehensive quality of apples and
predict storage time. Infrared Physics & Technology, 136, 105043.
[Link]
7. Iqbal, M. W., Malik, A. I., Ramzan, M. B., Memon, M. S., Mari, S. I., &
Habib, M. S. (2024). Consumer response to adjustable price and shelf-life of
fresh food products under effective preservation policy. Computers &
Industrial Engineering, 188, 109897.
[Link]
8. Jain, A., Sharma, S. R., Mittal, T. C., & Gupta, S. K. (2014). Storage studies of
okra under ambient storage conditions. International Journal of Innovative
Research in Technology (IJIRT), 1(4), 1–8.
9. Joshi, P., Pahariya, P., Al-Ani, M. F., & Choudhary, R. (2022). Monitoring and
prediction of sensory shelf-life in strawberry with ultraviolet-visible-near-
infrared (UV-VIS-NIR) spectroscopy. Applied Food Research, 2(2), 100123.
[Link]
10. Karunathilake, E. M. B. M., Le, A. T., Heo, S., Chung, Y. S., & Mansoor, S.
(2023). The path to smart farming: Innovations and opportunities in precision
agriculture. Agriculture, 13(8), 1593.
[Link]
11. Kasampalis, D. S., Tsouvaltzis, P., Ntouros, K., Gertsis, A., Gitas, I., &
Siomos, A. S. (2021). The use of digital imaging, chlorophyll fluorescence and
Vis/NIR spectroscopy in assessing the ripening stage and freshness status of
bell pepper fruit. Computers and Electronics in Agriculture, 187, 106265.
[Link]
12. Ktenioudaki, A., Esquerre, C. A., Do Nascimento Nunes, C. M., & O‘Donnell,
C. P. (2022). A decision support tool for shelf-life determination of
strawberries using hyperspectral imaging technology. Biosystems
Engineering, 221, 105–117.
[Link]
13. Mohammed, M., Riad, K., & Alqahtani, N. (2022). Design of a smart IoT-
based control system for remotely managing cold storage facilities. Sensors,
22(13), 4680. [Link]
35
14. Ogbaji, M., & Iorliam, I. B. (2020). Effect of Drumstick tree (Moringa
oleifera) and Neem (Azadirachta indica) leaf powders on shelf life and
physiological quality of okra fruits during storage. International Journal for
Research in Biology & Pharmacy, 6, 1–12.
15. Oh, H., Pottorff, M., Giongo, L., Mainland, C. M., Iorizzo, M., & Perkins-
Veazie, P. (2024). Exploring shelf-life predictability of appearance traits and
fruit texture in blueberry. Postharvest Biology and Technology, 208, 112643.
[Link]
16. Owoyemi, A., Porat, R., Lichter, A., Doron-Faigenboim, A., Jovani, O.,
Koenigstein, N., & Salzer, Y. (2022). Evaluation of the storage performance of
‗Valencia‘ oranges and generation of shelf-life prediction models.
Horticulturae, 8(7), 570. [Link]
17. Rashvand, M., Ren, Y., Sun, D.-W., Senge, J., Krupitzer, C., Fadiji, T., Sanzo
Miró, M., Shenfield, A., Watson, N. J., & Zhang, H. (2025). Artificial
intelligence for prediction of shelf-life of various food products: Recent
advances and ongoing challenges. Trends in Food Science & Technology, 159,
104989. [Link]
18. Sharma, A., Kumar, R., Kumar, N., & Saxena, V. (2024). Machine learning
driven portable Vis-SWNIR spectrophotometer for non-destructive
classification of raw tomatoes based on lycopene content. Vibrational
Spectroscopy, 130, 103628. [Link]
19. Weng, S., Yu, S., Dong, R., Pan, F., & Liang, D. (2020). Nondestructive
detection of storage time of strawberries using visible/near-infrared
hyperspectral imaging. International Journal of Food Properties, 23(1), 269–
281. [Link]
20. Xiao, F., Wang, H., Xu, Y., & Zhang, R. (2023). Fruit detection and
recognition based on deep learning for automatic harvesting: An overview and
review. Agronomy, 13(6), 1625. [Link]

36
The Impact of Climate Change and other Factors on
Maize Production in Tanzania: Autoregressive Distributed
Lag Approach

Bahati Ilembo1 and Johanes Tibanywana2


1
Department of Mathematics and Statistics Studies, Mzumbe University, Tanzania
Email ID : bmilembo@[Link]
2
Department of Mathematics and Statistics Studies, Mzumbe University, Tanzania
Email ID : tibanywanajohanes@[Link]

Abstract
The study examined the influence of climatic and non-climatic parameters on
the maize yield. Time series data on maize production and other variables were used
as recorded from 1961 – 2022. The study used the Autoregressive Distributed Lag
model (ARDL) to analyze the nexus between climatic and non-climatic variables in
influencing maize yield. The results showed that fertilizer and land use (non-climatic
factors) have a stronger and more consistent impact on maize yield than climatic
variables (precipitation and temperature). The ARDL (1,0,0,0,0) model results
showed that the lagged value of production per hectare has a positive and significant
effect on maize production, suggesting persistence in agricultural output over time.
Fertilizer use has a negative and significant coefficient, indicating that, in the short
run, higher fertilizer use is associated with a slight reduction in production per
hectare, possibly reflecting diminishing returns or inefficiencies in its application.
Precipitation and temperature both have insignificant effects. Land use, however,
has a positive and significant effect. In the long run, fertilizer use has a small but
significant negative effect, which could indicate inefficiency. Precipitation and
temperature remain statistically insignificant, suggesting they do not strongly affect
maize production. The land use area has a positive and significant long-run effect.
The findings suggest that policy efforts in Tanzania should prioritize expanding
access to and efficient use of agricultural inputs, particularly fertilizer and arable
land, as these factors have a stronger and more consistent impact on maize yield than
climatic variables.
Keywords: Autoregressive Distributed Lag, Climate change, Maize, Precipitation

1. INTRODUCTION
Climate change poses a significant threat to agriculture (Ibrahim [Link], 2025;
Sharma [Link], 2023), and the agricultural sector is especially vulnerable because of

37
its strong dependence on the climatic variables, mainly precipitation and temperature
(Singh, Arora & Babu, 2024; Waris [Link], 2023; Tirfi & Oyekale, 2021). Climate
affects crop production decisions and outcomes in agriculture. From very short-term
decisions about which crops to grow, when to plant or harvest a field, to longer-term
decisions about farm investments, climate can positively or negatively affect
agricultural systems (Shoko, Belete & Chaminuka, 2019). Production of cereal
crops, especially maize, is very responsive to changes in rainfall and temperature, as
climatic parameters influencing productivity (Tirfi & Oyekale, 2021). However, the
examination of both climatic and non-climatic factors becomes very important.
Maize is an important food and income earner for rural and urban dwellers in
Tanzania (Moshi, Nestory & Mlilile, 2023; Lyimo, Mduruma & De Groote, 2014)
and across Sub-Saharan Africa (Adam [Link], 2020). Despite its significance and
efforts done by the government and private sector, the yield of maize has remained
significantly below the average of less than 2 metric tonnes per hectare (Moshi,
Nestory & Mlilile, 2023). This threatens food insecurity and poverty for rural
people. Both climatic and non-climatic factors affect the maize yield and
productivity. Studies have shown that both intra- and interseasonal changes in
temperature and precipitation influence cereal yields (Rowhani et. Al, 2011), with
average annual rainfall having a positive relationship with aggregate maize
production (Moshi, Nestory & Mlilile, 2023). Similar results were also found by
Magwaza & Lyaro (2024), and that temperature had a long-run effect, while
precipitation had a short-run effect on maize production. The non-climatic factors
such as land cultivated maize, GDP per capita, fertilizer, and agricultural technology
have also been used to examine the effect on maize production (Singh, Arora &
Babu, 2024; Noorunnahar, Mila & Haque, 2023; Okunola, 2023; Nasrullah et. al,
2021).
The analysis of factors influencing maize production is well documented in the
literature, and the interest has been on the effects of climate change on maize
production (Aryal, 2025; Magwaza & Lyaro, 2024; Noorunnahar, Mila & Haque,
2023; Pickson, 2022; Maïga et. al., 2021; Tirfi & Oyekale, 2021). However,
literature is scant in the Tanzania context, especially on the use of the Autoregressive
Distributed Lag model (ARDL), which is ideal for developing countries like
Tanzania, where long-term consistent data on maize production, rainfall, fertilizer
use, and land cultivation for maize may be lacking. Few studies in Tanzania (Kitole,
Lihawa, and Nsidagi, 2023; Adam et. al, 2020; Nyaligwa [Link], 2017; Rowhani,
2011) examined factors affecting maize production, but none employed the ARDL in
the analysis. This is the point of departure for our paper, aiming at observing short-
run dynamics (immediate effects of weather shocks) as well as long-run equilibrium
relationships. The rest of the paper is organized as follows: Section 2 is on
methodology, followed by Section 3 on the results and discussion of the results.
Section 4 concludes and provides the policy implications.

38
2. MATERIALS AND METHODS
2.1 DATA AND DATA SOURCE
This study used time series data on maize yield (hg/ha) from 1961 – 2022 from
the FAOSTAT database, calculated annually with other climatic (precipitation and
temperature) and non-climatic variables (land use and fertilizer). The variables are
quantity production (ha), fertilizer (tonnes), precipitation (mm), land use (ha), and
temperature(0C). The other data links are [Link]
org/country/tanzania/climate-data-historical and [Link]
which offered data for the study variables. Maize was selected due to its potential
staple food in Tanzania as well as the Sub-Saharan region, as also put by Moshi,
Nestory & Mlilile (2023), Lyimo, Mduruma & De Groote (2014), as well as Adam
[Link]. (2020). It is also revealed as the primary food crop for most households in
Tanzania (Laudien et al., 2020). The dataset consisted of 62 observations meeting
the requirements of the generalization ability of time series analysis. As suggested
by Hyndman and Athanasopoulos (2018), the only theoretical limit to perform a time
series analysis is that there should be more observations than the number of
parameters in the forecasting model, and thus qualifies the use of the ARDL model
in our analysis.

2.2 MODEL FORMULATION


The ARDL model can be used to define the complementary relationship between
the response variable and the predictors in both the short and long run, in addition to
defining the integral relationship between the dependent and independent variables
(Pesaran, Smith & Shin, 1996). In addition to determining the magnitude of the
impacts of all dependent variables on the explanatory variable, the ARDL models are
standard least squares regression models containing lags in both the dependent and
independent variables. Such models have been used in econometrics in selecting
long-term relationships and common integration between variables over the years.
They can be expressed as ARDL. (𝑝, 𝑞1 , 𝑞2 , … . , 𝑞𝑘 ), where 𝑝 is the number of lags
in the dependent variable, 𝑞1 is the number of lags in the first explanatory variable
and 𝑞𝑘 is the number of lags in the kth explanatory variable. Accordingly, the model
takes the following form;
𝑝 𝑘 𝑞𝑗
𝑦𝑡 =∝ + ∑ 𝛾𝑖 𝑦𝑖 + ∑ ∑ 𝑋𝑗,𝑡−1 𝛽𝑗,𝑖 + 𝜀𝑡
𝑖=1 𝑗=1 𝑖=0
Explanatory variables 𝑋𝑗 with no lags (𝑞𝑗 = 0) are labelled stationary
regressions, whereas those containing lags are called dynamic regressions. To build
an ARDL model, the number of lags in each variable should first be determined. In
other words, the 𝑝, 𝑞1 , 𝑞2 , … . , 𝑞𝑘 must first be determined in the light of the criteria
set by Akaike AIC, Schwarz SC, and Hannan-Quinn (H-Q) as an alternative for the
adjusted coefficient of determination (Adjusted 𝑅 2) to select an appropriate model.
Since ARDL estimates the dynamic relationship between the response and
explanatory variables (Ahmed & Ibrahim, 2016), the model can be transformed into
39
a long-run form to capture the long-run supply response of the dependent variable to
changes in the explanatory variable. The long-run coefficients can be estimated as;
𝑞𝑗
∑𝑖=1 𝛽̂
𝑗,𝑖
∅𝑗 = 𝑝
1 − ∑𝑖=1 𝛾𝑖
Standard error associated with long-run coefficients can be calculated from the
standard error of the original regression with the help of the delta method (Shoman
& Hasan, 2013)

3. RESULTS AND DISCUSSION


Table I: Summary statistics
Variable Obs Mean Std. Dev. Min Max
productionha 62 13,051.29 4,867.25 4,808 31,359
fertilizert 62 36,400.99 37,483.36 1,149 171,354.80
precipitationm 62 982.26 113.97 760.64 1,238.62
temperature 62 22.65 0.37 21.87 23.29
landuseha 62 32,718.68 4,082.64 26,000 39,521.20
Source: Created by author (s)

Based on Table I, the dataset comprises 62 observations capturing agricultural


production, input use, and climatic variables (temperature and precipitation). On
average, production per hectare is 13,051.29 units, ranging from 4,808 to 31,359,
indicating significant variation in productivity. Fertilizer use shows the widest
disparity, with a mean of 36,400.99 tones and values ranging from as low as 1,149
tones to as high as 171,354.8 tones, reflecting differences in farming intensity and
input access. Precipitation averages 982.26 mm annually, with moderate variation
between 760.64 mm and 1,238.62 mm, while temperatures are relatively stable at
around 22.65°C, ranging only between 21.87°C and 23.29°C. Land use for
cultivation averages 32,718.68 hectares, with moderate differences across
observations, from 26,000 ha to 39,521.2 ha. Overall, the results suggest that while
climatic conditions remain fairly constant, agricultural outputs vary considerably,
likely influenced by differences in fertilizer use and cultivated land size.
Table II: Correlation matrix
Variable productionha fertilizert precipitationm temperature landuseha
productionha 1 0.3037 -0.0622 0.5146 0.5626
fertilizert 0.3037 1 0.0916 0.5823 0.8354
precipitationm -0.0622 0.0916 1 -0.2309 -0.1297
temperature 0.5146 0.5823 -0.2309 1 0.7809
landuseha 0.5626 0.8354 -0.1297 0.7809 1
Source: Created by author (s)
40
The correlation results indicate that production per hectare is moderately and
positively related to temperature and land use, suggesting that warmer areas and
larger cultivated areas tend to achieve higher yields, while its relationship with
fertilizer use is positive but weaker. Fertilizer use shows a strong positive correlation
with land use, meaning that larger farms tend to apply more fertilizer, and it is also
moderately related to temperature, implying that warmer regions may support more
intensive input application. Rainfall exhibits weak negative correlations with
production, temperature, and land use, indicating that higher rainfall is not
associated with higher yields in this dataset and may even be linked to slightly lower
values for these variables. Temperature strongly correlates with land use, and both
together are key factors associated with higher fertilizer use and production. Overall,
the results suggest that land size, temperature, and input intensity are more
influential in determining production levels than rainfall in this sample.
Table III: Pperron test
Variables At level First Differencing
Test Statistics P-value Test Statistics P-value
productionha -3.664 0.0046 - -
fertilizert -0.765 0.8293 -24.697 0.0000
precipitation mm -7.659 0.0000 - -
temperature -3.064 0.0294 - -
landuseha -0.300 0.9254 -7.811 0.0000
Source: Created by author (s)
The unit root test results indicate that production per hectare, precipitation, and
temperature are stationary at their levels, as evidenced by their statistically
significant test statistics with p-values below 0.05. This means these variables do not
exhibit a unit root and their statistical properties remain stable over time without
differencing. In contrast, fertilizer use and land use are non-stationary at their levels,
with high p-values indicating the presence of a unit root. However, after first
differencing, both variables become highly stationary, as shown by their large
negative test statistics and p-values of 0.0000. Overall, the findings suggest that
while most variables in the dataset are inherently stationary, fertilizer use and land
use require first differencing to achieve stationarity, an important step before
applying time-series models to avoid spurious results.
Table IV: Lag selection table
Sample: 1965 through 2022 Number of obs = 58
Lag LL LR df p FPE AIC HQIC SBIC
0 -2122.2 4.90E+25 73.3512 73.4204 73.5288
1 -1959.2 326.03 25 0 4.30E+23 68.5921 69.0072 69.6578
2 -1925.3 67.656 25 0 3.20E+23 68.2877 69.0488 70.2416
3 -1901.4 47.832 25 0.004 3.50E+23 68.3251 69.4321 71.1671
4 -1886.2 30.547 25 0.204 5.40E+23 68.6605 70.1134 72.3906
Source: Created by author (s)

41
The lag selection results, based on multiple information criteria, suggest that the
optimal lag length for the VAR model lies between 1 and 2. Specifically, the Final
Prediction Error (FPE) and Akaike Information Criterion (AIC) indicate lag 2 as
optimal, while the Hannan-Quinn (HQIC) and Schwarz Bayesian Information
Criterion (SBIC) suggest lag 1. Since FPE and AIC often prioritize model fit and are
commonly used in VAR analysis, lag 2 is likely the most appropriate choice. This
means the model should include two lags of the endogenous variables—production
per hectare, fertilizer use, precipitation, temperature, and land use—to best capture
the dynamic relationships in the data without overfitting.
Table V: ARDL model
ARDL (1, 0, 0, 0, 0) regression output
Model statistics
 Number of observations = 60
 F(5, 54) = 10.56
 Prob > F = 0.0000
 R-squared = 0.4943
 Adjusted R-squared = 0.4475
 Root MSE = 3596.1760

Std. P- Lowest Highest


Variable Coefficient t-statistic
Error value C.I C.I
[Link] 0.3806 0.1259 3.02 0.004 0.1282 0.6331
fertilizert -0.0577 0.0255 -2.26 0.028 -0.1088 -0.0066
precipitation mm 4.8788 4.7028 1.04 0.304 -4.5498 14.3074
temperature 84.5203 2302.9 0.04 0.971 -4532.52 4701.56
landuseha 0.8588 0.3091 2.78 0.008 0.239 1.4786
_cons -24562 48951.5 -0.5 0.618 -122704 73580.3
Source: Created by author (s)

The ARDL (1,0,0,0,0) model results show that the lagged value of production
per hectare has a positive and statistically significant effect on current production
(coefficient = 0.3806, p = 0.004), suggesting persistence in agricultural output over
time. Fertilizer use has a negative and significant coefficient (-0.0577, p = 0.028),
indicating that, in the short run, higher fertilizer use is associated with a slight
reduction in production per hectare, possibly reflecting diminishing returns or
inefficiencies in application. Precipitation and temperature both have statistically
insignificant effects (p-values = 0.304 and 0.971, respectively), suggesting they do
not exert a short-run influence on production in this dataset. Land use, however, has
a positive and significant effect (0.8588, p = 0.008), meaning increases in cultivated
land area contribute to higher production. The model explains about 49% of the
variation in production, and its overall significance indicates that the included
variables jointly influence agricultural output.
42
Table VI: Test for correlation
By Breusch-Godfrey LM test:
Null Hypothesis (H0): No serial correlation up to lag p
Alternative Hypothesis (H1): Serial correlation up to lag p

Lags(p) chi2 df Prob > F


1 0.451 1 0.5017
2 0.484 2 0.7849
Source: Created by author (s)

The Breusch-Godfrey LM test results suggest that there is no evidence of serial


correlation in the residuals of the model. At all tested lags (1 through 4), the chi-
square statistics are low and the corresponding p-values are well above 0.05,
meaning we fail to reject the null hypothesis of no serial correlation. This indicates
that the residuals do not exhibit systematic correlation over time, and the model‘s
error terms are likely independent.
By Durbin-Watson (DW) statistic:
d-statistics (6, 60) = 2.087733
Similarly, the Durbin-Watson (DW) statistic is 2.0877, which is very close to
the ideal value of 2, further confirming that there is no first-order autocorrelation in
the residuals. Overall, both tests indicate that the ARDL model‘s assumptions
regarding error independence are satisfied, and the model is free from serial
correlation issues.
Table VII: Cointegration bounds test

Variable Coefficient t-statistic Significance

[Link]~a 0.381 3.02 Significant at 1%


fertilizert -0.0577 -2.26 Significant at 5%
precipitat~m 4.879 1.04 Not significant
temperature 84.52 0.04 Not significant
landuseha 0.859 2.78 Significant at 1%
_cons -24562 -0.5 Not significant
Source: Created by author(s)

The model indicates that past production ([Link]~a) positively and


significantly affects current production at the 1% level, showing persistence over
time. Fertilizer use (fertilizert) has a small negative impact, significant at 5%,
suggesting possible over-application or inefficiency. Precipitation and temperature
do not have significant effects on production in this dataset. Land use area

43
(landuseha) has a positive and highly significant effect, highlighting the importance
of land allocation. The constant term is negative but not statistically significant.
Table VIII: Long-run equilibrium
Variable Coefficient t-statistic Significance
[Link]~a 0.381 3.02 Significant at 1%
fertilizert -0.0577 -2.26 Significant at 5%
precipitat~m 4.879 1.04 Not significant
temperature 84.52 0.04 Not significant
landuseha 0.859 2.78 Significant at 1%
_cons -24562 -0.5 Not significant
Source: Created by author (s)

In the long run, past production ([Link]~a) positively and significantly


influences current production, implying that production trends persist over time.
Fertilizer use (fertilizert) has a small but statistically significant negative effect,
which could indicate inefficiency or overuse in the long run. Precipitation and
temperature remain statistically insignificant, suggesting they do not strongly affect
production within this model. Land use area (landuseha) has a positive and
significant long-run effect, emphasizing the importance of land allocation in
sustaining production. The constant term is negative and not significant, indicating
that the explanatory variables largely account for production levels.
Table IX: Error Correction Model
Variable Coefficient t-statistic Significance
[Link]~a 0.381 3.02 Significant at 1%
fertilizert -0.0577 -2.26 Significant at 5%
precipitat~m 4.879 1.04 Not significant
temperature 84.52 0.04 Not significant
landuseha 0.859 2.78 Significant at 1%
_cons -24562 -0.5 Not significant
Source: Created by author (s)

The ECM results indicate how short-run deviations adjust towards the long-run
equilibrium. Lagged production ([Link]~a) has a positive and highly significant
effect, showing that past production influences current production in the short run.
Fertilizer use (fertilizert) negatively affects production and is significant at 5%,
suggesting potential inefficiency or over-application in the short run. Precipitation
and temperature remain statistically insignificant, implying their short-run changes
do not strongly affect production. Land use area (landuseha) positively and
significantly impacts production, highlighting the importance of land allocation. The
44
constant term is negative and not significant, meaning the included variables largely
explain short-run production variations.
Table X: Table for lagged ECM residuals
Model statistics:
 Number of observations: 60
 F(6, 53) = 4.22, p = 0.0015 (model overall significant)
 R-squared = 0.3230
 Adjusted R-squared = 0.2464
 Root MSE = 3658.7

Std. t- P- 95% Confidence


Variable Coefficient
Error statistic value Interval
[Link] -0.0944 0.1417 -0.67 0.508 [-0.3786, 0.1899]
d_fertilizert -0.0379 0.0235 -1.61 0.113 [-0.0851, 0.0093]
d_precipitationmm 4.3757 3.4574 1.27 0.211 [-2.5589, 11.3104]
d_temperature -915.62 1850.77 -0.49 0.623 [-4627.797, 2796.556]
d_landuseha -0.5079 1.3771 -0.37 0.714 [-3.2700, 2.2542]
L_ecm_resid -0.5618 0.2037 -2.76 0.008 [-0.9705, -0.1532]
_cons 1569.602 1955.62 0.8 0.426 [-2352.867, 5492.071]
Source: Created by author (s)
The error correction term (L_ecm_resid) is negative and statistically significant
at 1%, indicating that deviations from long-run equilibrium are corrected at a speed
of approximately 56% per period. Short-run changes in production, fertilizer use,
precipitation, temperature, and land use are not statistically significant, suggesting
limited immediate impact on production. The model explains about 32% of the
variation in short-run changes (R² = 0.3230), and multicollinearity is not a concern
(mean VIF = 1.53).

4. Conclusion
The study examined the influence of climatic and non-climatic parameters on
the maize yield using the Autoregressive Distributed Lag model (ARDL).
Preliminary results showed that while climatic conditions remain fairly constant,
maize yield (production) varied considerably, likely influenced by differences in
fertilizer use and cultivated land size. Further showed that land size, temperature,
and fertilizer are more influential in determining production levels than
precipitation. We also revealed that while most variables used were inherently
stationary, fertilizer use and land use required first differencing to achieve
stationarity, an important step before applying time-series models to avoid spurious
results. During the model fitting, we included two lags of the endogenous variables,
production per hectare, fertilizer use, precipitation, temperature, and land use to best
capture the dynamic relationships in the data without overfitting. We realized that

45
the model explained about 49% of the variation in production, and its overall
significance indicates that the included variables jointly influence maize output.
The ARDL (1,0,0,0,0) model results showed that the lagged value of production
per hectare has a positive and statistically significant effect on current maize
production (coefficient = 0.3806, p = 0.004), suggesting persistence in agricultural
output over time. Fertilizer use has a negative and significant coefficient (-0.0577, p
= 0.028), indicating that, in the short run, higher fertilizer use is associated with a
slight reduction in production per hectare, possibly reflecting diminishing returns or
inefficiencies in application. Precipitation and temperature both have statistically
insignificant effects (p-values = 0.304 and 0.971, respectively), suggesting they do
not exert a short-run influence on maize production. Land use, however, has a
positive and significant effect (0.8588, p = 0.008), meaning increases in cultivated
land area contribute to higher maize yield.
In the long run, past maize production ([Link]~a) positively and
significantly influences current maize production, implying that production trends
persist over time. Fertilizer use (fertilizert) has a small but statistically significant
negative effect, which could indicate inefficiency or overuse in the long run.
Precipitation and temperature remain statistically insignificant, suggesting they do
not strongly affect maize production in the long run. Land use area (landuseha) has a
positive and significant long-run effect, emphasizing the importance of land
allocation in sustaining maize production. On the other hand, the error correction
model (ECM) results indicate how short-run deviations adjust towards the long-run
equilibrium. Lagged production ([Link]~a) has a positive and highly significant
effect, showing that past production influences current production in the short run.
Fertilizer use (fertilizert) negatively affects production and is significant at 5%,
suggesting potential inefficiency or over-application in the short run. Precipitation
and temperature remain statistically insignificant, implying their short-run changes
do not strongly affect production. Land use area (landuseha) positively and
significantly impacts production, highlighting the importance of land allocation.
5. Policy Implications
The findings suggest that policy efforts in Tanzania should prioritize expanding
access to and efficient use of agricultural inputs, particularly fertilizer and arable
land, as these factors have a stronger and more consistent impact on maize yield than
climatic variables such as precipitation. Given that temperature also emerged as a
significant factor, strategies to promote climate-resilient farming practices, such as
heat-tolerant maize varieties, could further stabilize production. The moderate
explanatory power of the model (49%) indicates that while these factors are
influential, additional structural or institutional variables such as access to credit,
extension services, or market infrastructure may also play a role and should be
considered in future policy design. Ensuring accurate data collection and addressing
non-stationarity through proper econometric techniques, as demonstrated, is also
essential for making reliable policy decisions based on time-series evidence.

46
References:
1. Adam, R. I., Mmbando, F., Lupindu, O., Ubwe, R. M., Osanya, J., & Muindi,
P. (2020). Beyond maize production: gender relations along the maize value
chain in Tanzania. Journal of Gender, Agriculture and Food Safety, 5(2), 27-
41.
2. Ahmed, M., & Ibrahim, G. M. (2016). Employing the ARDL approach to
estimate the demand for fish in Egypt. The Egyptian Journal of Agricultural
Economics, 26(1), 297–324.
3. Aryal, A., Yadav, A., Goyal, A., Mannepalli, B. K., Deep, P., Kamalvanshi, V.,
& Kushwaha, S. (2025). Climate change and its effects on maize yield in
Nepal: An empirical analysis using the ARDL model. Journal of
Agrometeorology, 27(3), 344-348.
4. Eliw, M., Mottawea, A., & El-Shafei, A. (2019). Estimating Supply Response
of Some Strategic Crops in Egypt Using the ARDL Model. South Asian
Journal of Social Studies and Economics, 1-22.
5. Hyndman, R. J., & Athanasopoulos, G. (2018). Forecasting Principles and
Practice. Melbourne: OTexts.
6. Ibrahim, I. M., Yahaya, U., Alhassan, M. T., & Adeyemi, O. O. (2025).
Understanding the Effect of Climate Change and Productivity Factors on
Maize Production in Nigeria: ARDL Approach. African Journal of
Agricultural Science and Food Research, 20(1), 104-117.
7. Kitole, F. A., Lihawa, R. M., & Nsindagi, T. E. (2023). Agriculture
productivity and farmers‘ health in Tanzania: analysis of the maize
subsector. Global Social Welfare, 10(3), 197-206.
8. Laudien, R., Schauberger, B., Makowski, D., & Gornott, C. (2020). Robustly
forecasting maize yields in Tanzania based on climatic predictors. Scientific
reports, 10(1), 19650.
9. Lyimo, S., Mduruma, Z., & De Groote, H. (2014). The use of improved maize
varieties in Tanzania. African Journal of Agricultural Research, 9(7), 643-657.
10. Maïga, A., Bathily, M., Bamba, A., Mouleye, I. S., & Nimaga, M. S. (2021).
Analysis of the effects of climate change on maize production in Mali. Asian
Res J Agric, 14(4), 42-52.
11. Magwaza, C., & Lyaro, S. (2024). Impact of Rainfall and Temperature as
Aspects of Climate Change on Maize Production in Zimbabwe. Journal of
Water Resources, Engineering, Management & Policy, 1(2).
12. Moshi, A., Nestory, M., & Mlilile, B. (2023). Economic analysis of factors
affecting maize production in Tanzania: Time series analysis. Rural Planning
Journal, 25(1), 33-45.
13. Nasrullah, M., Rizwanullah, M., Yu, X., Jo, H., Sohail, M. T., & Liang, L.
(2021). An autoregressive distributed lag (ARDL) approach to study the
impact of climate change and other factors on rice production in South
Korea. Journal of water and climate change, 12(6), 2256-2270.

47
14. Noorunnahar, M., Mila, F. A., & Haque, F. T. I. (2023). Does the supply
response of maize suffer from climate change in Bangladesh? Empirical
evidence using the ARDL approach. Journal of Agriculture and Food
Research, 14, 100667.
15. Nyaligwa, L., Hussein, S., Laing, M., Ghebrehiwot, H., & Amelework, B. A.
(2017). Key maize production constraints and farmers‘ preferred traits in the
mid-altitude maize agroecologies of northern Tanzania. South African Journal
of Plant and Soil, 34(1), 47-53.
16. Okunola, A. (2023). Effect of Land Allocation on Agricultural Productivity in
Nigeria Using the ARDL Approach to Cointegration. Journal of Land and
Rural Studies, 11(2), 115-130.
17. Pesaran, M. H., Shin, Y., & Smith, R. P. (1999). Pooled mean group estimation
of dynamic heterogeneous panels. Journal of the American Statistical
Association, 94(446), 621-634.
18. Pickson, R. B., Gui, P., Chen, A., & Boateng, E. (2022). Empirical analysis of
rice and maize production under climate change in China. Environmental
Science and Pollution Research, 29(46), 70242-70261.
19. Rowhani, P., Lobell, D. B., Linderman, M., & Ramankutty, N. (2011). Climate
variability and crop production in Tanzania. Agricultural and forest
meteorology, 151(4), 449-460.
20. Sharma, R. K., Dhillon, J., Kumar, P., Bheemanahalli, R., Li, X., Cox, M. S.,
& Reddy, K. N. (2023). Climate trends and maize production nexus in
Mississippi: empirical evidence from ARDL modelling. Scientific
Reports, 13(1), 16641.
21. Shoko, R. R., Belete, A., & Chaminuka, P. (2019). Maize yield sensitivity to
climate variability in South Africa: application of the ARDL-ECM
approach. Journal of Agribusiness and Rural Development, 54(4), 363-371.
22. Shoman, A., & El-Latif, H. A. (2013). Analysis of the long-run equilibrium
relationship using unit root tests by combining the autoregressive and
distributed lag models. Economic Science, College of Administration and
Economics, Baghdad University, 9(34), 174–210.
23. Singh, A., Arora, K., & Babu, S. C. (2024). Examining the impact of climate
change on cereal production in India: Empirical evidence from the ARDL
modelling approach. Heliyon, 10(18).
24. Tirfi, A. G., & Oyekale, A. S. (2021). Maize output supply response to
climatic and other input variables in Ethiopia. Asian Journal of Agriculture
and Rural Development, 11(4), 320-326.
25. Waris, U., Tariq, S., Mehmood, U., & ul-Haq, Z. (2023). Exploring potential
impacts of climatic variability on the production of maize in Pakistan using
the ARDL approach. Acta Geophysica, 71(5), 2545-2561.

48
Statistical Inference of Stress-Strength Reliability Model
using Lindley Distribution

Obed Benaiah Jyrwa1, Hemima Ahmed2 and Bhanita Das3


1
Department of Statistics, North-Eastern Hill University, Shillong,
Email ID obed2benaiah@[Link]
2
Research Scholar, Department of Statistics, North-Eastern Hill University,
Shillong, Email ID ahmedhemim@[Link]
3
Assistant Professor, Department of Statistics, North-Eastern Hill University,
Shillong, bhanitadas83@[Link]
Abstract:
This paper investigates the stress-strength reliability model, 𝑅 = 𝑃(𝑌 < 𝑋),
where both stress Y and strength X follow Lindley distributions with distinct shape
parameters. A closed-form expression for R is derived, and its parameters are
estimated using maximum likelihood estimation (MLE) alongside bootstrap
methods. The asymptotic distribution of the MLE of R is established, facilitating the
construction of confidence intervals. Extensive simulation studies are conducted to
evaluate the performance of the proposed estimators across varying sample sizes,
demonstrating consistency and efficiency. The practical utility of the model is
illustrated through a real-data application involving survival times of head and neck
cancer patients.
Keywords: Stress-strength reliability; Lindley distribution; Maximum likelihood
estimation; Bootstrap methods.

1. INTRODUCTION
In reliability engineering and life data analysis, the stress-strength reliability
(SSR) has attracted significant interest because it is practically useful in the analysis
of mechanical system performance, electronic devices, and many industrial
processes. The stress-strength reliability measure, which is typically expressed as
𝑅 = 𝑃(𝑌 < 𝑋), is the probability that a strength X of a system will be larger than
the imposed stress Y. A higher value of R indicates greater reliability, implying the
system is more likely to withstand the stress and perform its intended function
without failure.
The history of research on this parameter is long, starting with the pioneering
work of Birnbaum (1956) and being further developed by Church and Harris (1970),
who introduced the term ―stress-strength‖. A wide range of statistical procedures has
since been established to estimate R, including parametric as well as non-parametric
methods. Many researchers have given interest in this area including Kotz,
Balakrishnan & Johnson (2003), who provided a comprehensive survey of SSR
49
using distributions like Weibull, gamma, and log-normal. Their contribution lies in
establishing theoretical properties, formulas for R, and guidance on estimator
performance under different censoring conditions. Kundu and Raqab (2009) made
significant contributions in Bayesian estimation of SSR, particularly under censored
samples. Their work on MCMC and posterior intervals for R enabled the estimation
even when traditional methods were ineffective. For some of the recent work on the
stress-strength model can be obtained in Kohansal (2019), Alslman and Helu (2022),
Singh et al. (2024), Abu-Moussa et al. (2021), Saini (2024) and Mahto et al. (2020).
Though the conventional studies in this area have traditionally used traditional
lifetime distributions, however, there has been increasing interest in investigating
other distributions with more flexibility and a better fit for actual data.
The Lindley distribution, proposed by D. V. Lindley in 1958, in the context of
Bayesian statistics, has turned out to be a potential choice for lifetime data
modelling. It is a mixture distribution of exponential and gamma and is part of the
exponential family. It involves a number of advantages such as having a
straightforward mathematical expression, being unimodal, and having an increasing
hazard rate function, thus being well suited for failure times and waiting times
modelling. Ghitany et al. (2008) continued to investigate its properties and show its
usability in real-life applications. The use of Lindley distribution in stress-strength
reliability analysis has remained relatively restricted. Ghitany et al. (2008, 2011)
evaluated the suitability of the Lindley distribution in real-world reliability data.
They developed estimators, derived properties like hazard rates, and extended
Lindley to weighted Lindley and power Lindley models, which offer better
flexibility in stress-strength applications.
The use of Lindley distribution in stress-strength reliability analysis has
remained relatively restricted. In this work, we examine statistical inference for the
reliability parameter 𝑅 = 𝑃(𝑌 < 𝑋), where X and Y are independent Lindley-
distributed variables. The maximum likelihood approach is employed to estimate R,
and its asymptotic distribution is derived to construct confidence intervals for the
reliability parameter. The maximum likelihood method is used to derive an estimator
for R, and its asymptotic distribution is established to facilitate the construction of
confidence intervals. In addition to asymptotic methods, bootstrap techniques,
including percentile bootstrap (boot-p) and bootstrap-t (boot-t), are employed to
enhance the accuracy of interval estimation, particularly in small-sample scenarios.
The rest of the paper is organized as follows. Section 2 introduces the model
assumptions and provides a brief overview of the Lindley distribution. Section 3
presents the maximum likelihood estimator (MLE) for the stress–strength reliability
parameter RR, along with its asymptotic distribution and associated confidence
intervals. Bootstrap methods for interval estimation are also discussed in this
section. A comprehensive simulation study is conducted in Section 4 to evaluate the
performance of the proposed estimators under various sample sizes. Section 5
illustrates the practical application of the methodology using a real dataset. Finally,
the paper is concluded in Section 6 with a summary of findings and suggestions for
future research.

50
2. THE MODEL ASSUMPTIONS
2.1. LINDLEY DISTRIBUTION
The Lindley distribution (LD) is a continuous probability distribution introduced
by D.V. Lindley in 1958 within a Bayesian framework. Originally proposed as a
counterexample in fiducial inference, it has since gained recognition for its
flexibility and practical utility. Today, the Lindley distribution is extensively applied
in various fields, including reliability engineering, survival analysis, and lifetime
data modeling, due to its mathematical tractability and ability to model real-world
phenomena effectively.
Mathematically, a random variable 𝑋 is identified to have a Lindley distribution
with a single parameter 𝜃 (> 0), i.e., if 𝑋~𝐿𝑖𝑛𝑑𝑙𝑒𝑦(𝜃), then its probability density
function (pdf) is given by
𝜃𝑚
𝑓(𝑥; 𝜃) = 𝜃+1 (1 + 𝑥)𝑒 −𝜃𝑥 ; 𝑥 > 0, 𝜃 > 0 (1)
And the pdf of inverse Lindley is
𝜃
𝜃𝑚 1+𝑥
𝑓(𝑥; 𝜃) = . / 𝑒 −𝑥 ; 𝑥 > 0, 𝜃 > 0
𝜃+1 𝑥 𝑛
and the cumulative distribution function (cdf)
𝜃𝑥
𝐹(𝑥; 𝜃) = 1 − .1 + 𝜃+1/ 𝑒 −𝜃𝑥 ; 𝑥 > 0, 𝜃 > 0. (2)
Figure 1 displays the probability density function and cumulative distribution
function of the Lindley distribution for various values of the shape parameter θ .
The respective reliability/survival function (sf) and hazard rate function (hrf) are
𝜃𝑥
𝑆(𝑥; 𝜃) = .1 + 𝜃+1/ 𝑒 −𝜃𝑥 ; 𝑥 > 0, 𝜃 > 0, (3)
and
𝜃𝑚 (1 + 𝑥)
𝑕(𝑥; 𝜃) = 1+𝜃(1+𝑥) ; 𝑥 > 0, 𝜃 > 0. (4)
Figure 2 displays the survival function and hazard rate function of the Lindley
distribution for various values of the shape parameter θ .

Fig 1. pdf and cdf plot of Lindley distribution


51
Fig 2. sf and hrf plot of Lindley distribution

2.2 STRESS-STRENGTH RELIABILITY


Let 𝑌 and 𝑋 be independent stress and strength random variables that follow
Lindley distribution with parameters 𝜃₁ and 𝜃2 , respectively. Then, the stress–
strength reliability 𝑅 is defined as

𝑅 = 𝑃(𝑌 < 𝑋) = ∫0 𝑓(𝑥, 𝜃1 )𝐹(𝑥, 𝜃2 )𝑑𝑥, (5)
where,
𝜃12
𝑓(𝑥, 𝜃1 ) = (1 + 𝑥)𝑒 −𝜃𝑙 𝑥 , 𝑥 > 0, 𝜃1 > 0
(𝜃1 + 1)
𝜃 𝑥
𝐹(𝑥, 𝜃2 ) = 1 − .1 + 𝜃 𝑚+1/ 𝑒 −𝜃𝑚 𝑥 , 𝑥 > 0, 𝜃2 > 0
𝑚
∞ 𝜃𝑙𝑚 𝜃 𝑥
Therefore, 𝑅 = ∫0 (𝜃𝑙 +1)
(1 + 𝑥)𝑒 −𝜃𝑙 𝑥 . 1 − .1 + 𝜃 𝑚+1/ 𝑒 −𝜃𝑚 𝑥 / 𝑑𝑥.
𝑚
After evaluating the integrals, we get
𝜃𝑙𝑚 1 2𝜃𝑚 +1 2𝜃𝑚
𝑅 = 1 − (𝜃 0𝜃 +𝜃 + (𝜃 𝑚 + (𝜃 𝑛1 (6)
𝑙 +1) 𝑙 𝑚 𝑙 +1)(𝜃𝑙 +𝜃𝑚 ) 𝑙 +1)(𝜃𝑙 +𝜃𝑚 )

3. ESTIMATION OF STRESS AND STRENGTH RELIABILITY


3.1. MAXIMUM LIKELIHOOD ESTIMATION
Suppose 𝑋1 , 𝑋2 , . . ., 𝑋𝑛 is a strength random sample from 𝐿𝐷(𝜃1 ) and 𝑌1 , 𝑌2 ,
. . ., 𝑌𝑚 is a stress random sample from 𝐿𝐷(𝜃2 ) distributions. Thus, the likelihood
function based on the observed sample is given by
𝑛 𝑚
𝜃12 𝜃22
𝐿 = (𝜃1 , 𝜃2 |𝑥, 𝑦) = ∏ (1 + 𝑥𝑖 )𝑒 −𝜃𝑙𝑥𝑖 ∏ (1 + 𝑦𝑗 )𝑒 −𝜃𝑚𝑦𝑖
(𝜃1 + 1) (𝜃2 + 1)
𝑖=1 𝑗=1
2𝑛 2𝑚 𝑛 𝑚
𝜃1 𝜃2 𝑛 𝑚
= ∏(1 + 𝑥𝑖 )𝑒 −𝜃𝑙 ∑𝑖=𝑙 𝑥𝑖 ∏(1 + 𝑦𝑗 )𝑒 −𝜃𝑚 ∑𝑗=𝑙 𝑦𝑗
(𝜃1 + 1)𝑛 (𝜃2 + 1)𝑚
𝑖=1 𝑗=1

52
The log-likelihood function is given by
log 𝐿 = 2𝑛 𝑙𝑜𝑔𝜃1 + 2𝑚 𝑙𝑜𝑔𝜃2 − 𝑛 log(𝜃1 + 1) − 𝑚 log(𝜃2 + 1) − 𝜃1 𝑠1
−𝜃2 𝑠2 + ∑𝑛𝑖=1(1 + 𝑥𝑖 ) + ∑𝑚
𝑗=1(1 + 𝑦𝑗 ) (7)
Where, 𝑠1 = ∑𝑛𝑖=1 𝑥𝑖 and 𝑠2 = ∑𝑚 𝑗=1 𝑦𝑗
The MLE of 𝜃1 and 𝜃2 , say 𝜃1 and 𝜃̂2 , respectively, can be obtained as the
̂
solutions of the following equations
𝜕𝑙𝑜𝑔𝐿 2𝑛 𝑛
𝜕𝜃
= 𝜃 − (𝜃 +1) − 𝑠1 = 0 (8)
𝑙 𝑙 𝑙
𝜕𝑙𝑜𝑔𝐿 2𝑚 𝑚
𝜕𝜃𝑚
= 𝜃𝑚
− (𝜃 − 𝑠2 =0 (9)
𝑚 +1)
From (9) and (10), we obtain the MLEs 𝜃̂1 and 𝜃̂2 as follows
𝑛 − 𝑠1 + √(𝑠1 − 𝑛) + 8𝑛𝑠1
𝜃̂1 = ,
2𝑠1
and
𝑚 − 𝑠2 + √(𝑠2 − 𝑚) + 8𝑚𝑠2
𝜃̂2 = .
2𝑠2
Using the invariance property of the MLE, the MLE 𝑅̂𝑚𝑙𝑒 of 𝑅 can be obtained
by substituting 𝜃̂𝑘 in place of 𝜃𝑘 (k = 1, 2) in (21). The 𝑅̂𝑚𝑙𝑒 is thus given by
̂𝑚
𝜃 1 ̂𝑚 +1
2𝜃 2𝜃̂𝑚
𝑅̂𝑚𝑙𝑒 = 1 − 𝑙 [ +
̂𝑙 +1) 𝜃
̂𝑙 +𝜃
̂𝑚 𝑚 +
̂𝑙 +1)(𝜃
̂𝑙 +𝜃
̂𝑚 ) 𝑛] (10)
̂𝑙 +1)(𝜃
̂𝑙 +𝜃
̂𝑚 )
(𝜃 (𝜃 (𝜃

3.2. ASYMPTOTIC DISTRIBUTION AND CONFIDENCE INTERVALS


In this subsection, we derived the asymptotic distributions of 𝜃̂𝑘 (k=1,2) as well
as of 𝑅̂𝑚𝑙𝑒 . Based on the asymptotic distribution, we also obtained the asymptotic
confidence interval for the parameters. For a large sample, the asymptotic
distribution of 𝜃̂𝑘 is defined by
√𝑛(𝜃̂𝑘 − 𝜃𝑘 ) → 𝑁(0, 𝐼(𝜃𝑘 )−1 )
𝜕𝑚 𝑙𝑜𝑔 𝑓(𝑥) 𝜃𝑚 +4𝜃𝑘 +2
Where 𝐼(𝜃𝑘 ) = −𝐸 ( 𝜕𝜃𝑘𝑚
* = 𝜃𝑘𝑚 (𝜃 𝑚,𝑘 = 1,2
𝑘 𝑘 +1)
Using delta method, the Fisher information matrix of 𝛩 = (𝜃1 , 𝜃2 ) is defined as
𝜕2𝐿 𝜕2𝐿
𝜕𝜃12 𝜕𝜃1 𝜕𝜃2 𝐼 0
I(Θ) = −E = [ 11 ]
𝜕2𝐿 𝜕2𝐿 0 𝐼22
2
[𝜕𝜃2 𝜕𝜃1 𝜕𝜃2 ]
𝜕 𝐿 𝑚 2𝑛 𝑛 𝜕 𝐿 2𝑚 𝑚 𝑚 𝜕 𝐿 𝜕 𝐿 𝑚 𝑚
Where 𝐼11 = 𝜕𝜃 𝑚 = − 𝜃 𝑚 + (𝜃 +1)𝑚 , 𝐼22 = 𝜕𝜃 𝑚 = − 𝜃 𝑚 + (𝜃 +1)𝑚 and 𝜕𝜃 𝜕𝜃 = 𝜕𝜃 𝜕𝜃 = 0
𝑙 𝑙 𝑙 𝑚 𝑚 𝑚 𝑙 𝑚 𝑚 𝑙

The asymptotic 100 (1 − 𝛼) % confidence intervals for 𝜃𝑘 (k = 1, 2) are given by


2
𝜃̂𝑘2 (𝜃̂𝑘 + 1)
>𝜃̂𝑘 ± 𝑧𝛼 √ ?
2 𝜃̂𝑘2 + 4𝜃̂𝑘 + 2

53
𝛼 𝑡𝑕
where, 𝑧𝛼 is the upper . 2 / percentile of a standard normal random variable.
𝑚
Similarly, the asymptotic distribution of stress–strength reliability as 𝑛 → ∞ and
𝑚 → ∞ is given by
𝑅̂ − 𝑅
→ 𝑁(0,1)
2 2
𝑅 𝑅
√ 1 + 2
𝑛 𝐼11 𝑚 𝐼22
𝑑𝑅
Where 𝑅1 = 𝑑𝜃 =
𝑙
𝑑𝑅
𝑅2 =
=
𝑑𝜃2
Although 𝑅̂𝑚𝑙𝑒 can be obtained in explicit form, the exact distribution of 𝑅̂ is
difficult to obtain. Due to this reason, we constructed the asymptotic confidence
interval for R. The asymptotic 100 (1 − 𝛼) % confidence interval for 𝑅 can be
easily obtained as
𝑅̂12 𝑅̂22
>𝑅̂𝑚𝑙𝑒 ± 𝑧𝛼 √ + ?
2 𝑛𝐼(𝜃 ̂1 ) 𝑛𝐼(𝜃̂2 )
Where 𝐼(𝜃̂𝑘 ) and 𝑅̂𝑘 , 𝑘 = 1, 2 are the MLE‘s of 𝐼(𝜃𝑘 ) and 𝑅𝑘 , respectively.

3.3. BOOTSTRAP CONFIDENCE INTERVAL


In this section, two bootstrap methods are considered: the parametric bootstrap
(Boot-p) and the bootstrap-t method (Boot-t). The procedures for constructing
confidence intervals under these methods are given below
3.3.1. Bootstrap-p
Step 1. Generate random samples 𝑥1 , 𝑥2 , . . . , 𝑥𝑛 from 𝐹(𝑥, 𝜃1 ) and 𝑦1 , 𝑦2 , . . . , 𝑦𝑚
from 𝐹(𝑦, 𝜃2 ), respectively. Calculate the MLEs of 𝜃1 and 𝜃2 , say 𝜃̂1 and 𝜃̂2 .
Step 2. Use 𝜃1 and 𝜃2 , say 𝜃̂1 and 𝜃̂2 to generate independent bootstrap samples
𝑥1∗ , 𝑥2∗ , . . . , 𝑥𝑛∗
from 𝐹(𝑥, 𝜃1 ) and 𝑦1∗ , 𝑦2∗ , . . . , 𝑦𝑛∗ from 𝐹(𝑦, 𝜃2 ). Calculate the MLEs of
unknown parameters based on the bootstrap samples, denoted by 𝜃̂1∗ and 𝜃̂2∗
Step 3. Calculate the bootstrap estimate of R in (10), and denote by 𝑅̂ ∗
Step 4. Repeat Steps 2 and 3 N times; then we have 𝑅̂(1)

, 𝑅̂(2)

, … . , 𝑅̂(𝑁)

.
Step 5. Let 𝜑(𝑥) = 𝑃( 𝑅 ∗ ≤ 𝑥) be the cdf of 𝑅 ∗ . Define 𝑅̂𝑏𝑝 (𝑥) = 𝜑−1 (𝑥) for
given x. en, two-side 100(1 − 𝛾)%percentile confidence intervals of 𝑅 are given by
𝛾 𝛾
4𝑅̂𝑏𝑝 . / , 𝑅̂𝑏𝑝 .1 − /5
2 2

54
3.3.2. Bootstrap-t
Step1. Same as the bootstrap-p.
Step 2. Same as the bootstrap-p
Step 3. Same as the bootstrap-p
Step 4. Obtain the 𝑡𝑅∗ statistics 𝑡𝑅∗ ( 𝑅̂ ∗ − 𝑅)/𝜍𝑅∗ .
(1) (2) (𝑀)
Step 5. Repeat Steps 2, 3, and 4 M times; then we have .𝑡𝑅∗ , 𝑡𝑅∗ , … , 𝑡𝑅∗ /.
Step 6. Let 𝜓(𝑥) = 𝑃(𝑡𝑅 ≤ 𝑥) be the cdf of 𝑡𝑅 . Define 𝑅𝑏𝑡 (𝑥) = 𝑅̂ +
∗ ∗

𝜓 −1 (𝑥)𝜍𝑅 for given 𝑥. Then, two side 100(1 − 𝛾)% bootstrap-t confidence
intervals of 𝑅 are given by
𝛾 𝛾
4𝑅̂𝑏𝑡 . / , 𝑅̂𝑏𝑡 .1 − /5
2 2

4. SIMULATION STUDY:
In this section, a simulation study is carried out to assess the performance of
MLEs for the stress-strength reliability model. We consider four sample sizes such
as (𝑛, 𝑚) = (10,20), (20,30), (30,40) 𝑎𝑛𝑑 (40, 50) for two cases of the true values
of the parameters and corresponding actual values of R. The two cases are as follows
Case I: 𝜃1 = 1, 𝜃2 = 2, R=0.7655
Case II: 𝜃1 = 0.5, 𝜃2 = 1.5, R=1
For each sample size the evaluation of the estimates was performed based on the
mean squared errors (MSEs), which are calculated utilizing the R package. Average
lengths for interval estimates (asymptotic, bootstrap) are evaluated. The study is
performed for 1000 replicates. For each replication, 1000 bootstrap samples are
used.
Table 1. Mean estimates of R (first row) with their MSEs (second row) and ACIs,
asysmptotic, bootstrap for R for the stress parameter 𝜃1 = 1 and
strength parameter 𝜃1 = 2, and R=0.7655 with varying n and m.

(𝒏, 𝒎) ̂ 𝑴𝑳
𝑹 ACI Boot p Boot t
(10,20) 0.75024 0.67047 0.60923 0.70112
(0.00764) (0.9145) (0.9208) (0.9105)
(20, 30) 0.75084 0.54378 0.62780 0.6598
(0.00673) (0.9138) (0.9256) (0.9301)
(30, 40) 0.76024 0.71365 0.83570 0.89012
(0.00517) (0.9219) (0.9405) (0.9502
(40, 50) 0.76140 0.69549 0.84529 0.7094
(0.00438) (9.3126) (0.9565) (0.9613)

55
Table 2. Mean estimates of R (first row) with their MSEs (second row) and ACIs
and asymptotic bootstrap for R for the stress parameter 𝜃1 = 0.5 and strength
parameter 𝜃1 = 1.5, and R=1 with varying n and m.
(𝑛, 𝑚) 𝑅̂𝑀𝐿 ACI Boot p Boot t
(10,20) 1.50214 0.55490 0.6167 0.56838
(0.1453) (0.9012) (0.9163) (0.9063)
(20, 30) 1.4217 0.69541 0.64376 0.76775
(0.10027) (0.9286) (0.9346) (0.93457)
(30, 40) 1.30136 0.74531 0.87634 0.84355
(0.09242) (0.94327) (0.9527) (0.95775)
(40, 50) 1.12087 0.85402 0.89533 0.87663
(0.07823) (0.96423) (0.9755) (0.9502)

The mean estimates and MSEs, ACIs ans bootstrap-p and bootstrap-t for the
above two cases are presented in Table 1 and 2 respectively. From the table values it
is observed that as the sample sizes increases the biases decreases to zero and also
the MSEs diminishes to zero with increase in the sample sizes. This shows
consistency and unbiasedness of the MLEs.

5. APPLICATION TO THE REAL DATA SET


This section presents a real data analysis to demonstrate the proposed methods.
The datasets represent survival times of head and neck cancer patients: one group
treated with radiotherapy (RT) and the other with combined chemotherapy and
radiotherapy (CT + RT). These data were originally reported by Efron (1988), and
Makkar et al. (2014) demonstrated that a lifetime model with an upside-down
bathtub-shaped hazard rate is well-suited for this type of problem. The datasets are
provided as follows:
Survival times of patients treated using RT: Data X
6.53, 7, 10.42, 14.48, 16.1, 22.7, 34, 41.55, 42, 45.28, 49.4, 53.62, 63, 64, 83, 84, 91, 108,
112, 129, 133, 133, 139, 140, 140, 146, 149, 154, 157, 160, 160, 165, 146, 149, 154, 157,
160, 160, 165, 173, 176, 218, 225, 241, 248, 273, 277, 297, 405, 417, 420, 440, 523, 583,
594, 1101, 1146, 1417.

Survival times of patients treated using CT+RT: Data Y


12.2, 23.56, 23.74, 25.87, 31.98, 37, 41.35, 47.38, 55.46, 58.36, 63.47, 68.46, 78.26,
74.47, 81.43, 84, 92, 94, 110, 112, 119, 127, 130, 133, 140, 146, 155, 159, 173, 179,
194, 195, 209, 249, 281, 319, 339, 432, 469, 519, 633, 725, 817, 1776

First, we assessed the suitability of the Lindley distribution for the given
datasets using the Akaike Information Criterion (AIC) and Bayesian Information
Criterion (BIC), Consistent Akaike Information Criterion (CAIC) and Hannan–
Quinn Information Criterion (HQIC). We then compared its performance with the
56
exponential distribution (ED). The computed statistics for both models, based on the
real datasets, are presented in Table 3. Clearly, the LD fits well with the data in
comparison to ED.
The maximum likelihood estimates, standard error (SE) and the confidence
intervals (CIs) of 𝜃1 , 𝜃2 and 𝑅 are summarized in Table 4.
Table 3: The model fitting summary for both the data-sets
Distribution Loglikelihood AIC BIC CAIC HQIC
Data LD -381.8739 765.7478 767.8082 765.8192 766.5503
1 ED -385.7031 773.4062 775.4666 773.4776 774.2088
Data LD -279.5784 561.1568 562.9409 561.2520 561.8184
2 ED -289.5822 581.1643 582.9485 581.2596 581.8260

Table 4: MLE, SE, CI (lower, upper) of 𝜃1 , 𝜃2 and 𝑅 based on real data-sets.


Parameter MLE SE CI Lower CI upper
𝜃1 0.008804171 0.0008174623 0.007201945 0.0104064
𝜃2 0.008909947 0.0009498221 0.007048296 0.0107716
𝑅 0.5044268 0.05240342 0.4017161 0.6071375

6. CONCLUSION
In this article, we consider one parameter Lindley distribution. The flexibility of
the distribution has been demonstrated by studying the behaviours of the pdf and
hazard rate functions. It has been found that the Lindley distribution gives a good fit
to the data of survival of head and neck cancer patients. Further, we addressed the
problem of estimating the stress–strength reliability 𝑅 = 𝑃(𝑋 > 𝑌) where X and Y
follow the Lindley distributions. The estimations of the stress-strength reliability
model with the corresponding ACIs using the maximum likelihood, two parametric
bootstraps are obtained. A simulation study is computerized to inspect the estimation
method for different sample sizes (𝑛, 𝑚).

REFERENCES
1. Birnbaum, Z. W., (1956). On a use of the Mann-Whitney statistic. In:
Proceedings of Third Berkeley Symposium on Mathematical Statistics and
Probability, Vol. 1, University of California Press, Berkeley, CA, pp. 13–17.
2. Church, J. D., & Harris, B., (1970). The estimation of reliability from stress-
strength relationships. Technometrics, 12:49–54.
3. Kotz, S., Lumelskii, Y., & Pensky, M., (2003). The Stress-Strength Model and
its Generalizations: Theory and Applications. Singapore: World Scientific
Press.
4. Kundu, D., & Raqab, M. Z., (2009). Estimation of R = PY < X for three-
parameter Weibull distribution. Statistics & Probability Letters, 79:1839–
1846.
57
5. Lindley, D. V., (1958). Fudicial distributions and Bayes‘ theorem. Journal of
the Royal Statistical Society, Series B: Statistical Methodology, 20:102–107.
6. Ghitany, M. E., Atieh, B., & Nadarajah, S., (2008). Lindley distribution and its
application. Mathematics and Computers in Simulation, 78:493–506.
7. Kohansal, A., (2019). On estimation of reliability in a multicomponent stress-
strength model for a Kumaraswamy distribution based on progressively
censored sample. Statistical Papers, 60(6), 2185-2224.
8. Alslman, M., & Helu, A., (2022). Estimation of the stress-strength reliability
for the inverse Weibull distribution under adaptive type-II progressive hybrid
censoring. Public Library of Science ONE, 17(11), e0277514.
9. Singh, K., Mahto, A. K., Tripathi, Y., & Wang, L., (2024). Inference for
reliability in a multicomponent stress–strength model for a unit inverse
Weibull distribution under type-II censoring. Quality Technology &
Quantitative Management, 21(2), 147-176.
10. Abu-Moussa, M. H., Abd-Elfattah, A. M., & Hafez, E. H., (2021). Estimation
of stress-strength parameter for Rayleigh distribution based on progressive
type-II censoring. Information Sciences Letters, 10(1), 101-110.
11. Saini, S., (2024). Estimation of multi-stress strength reliability under
progressive first failure censoring using generalized inverted exponential
distribution. Journal of Statistical Computation and Simulation, 94(14), 3177-
3209.
12. Mahto, A. K., Tripathi, Y. M., & Kı zı laslan, F., (2020). Estimation of
reliability in a multicomponent stress–strength model for a general class of
inverted exponentiated distributions under progressive censoring. Journal of
Statistical Theory and Practice, 14(4), 58.
13. Efron, B., (1988). Logistic regression, survival analysis, and the Kaplan–
Meier curve. Journal of the American Statistical Association, 83, 414–425.
14. Makkar, P., Srivastava, P. K., Singh, R. S., & Upadhyay S. K., (2014).
Bayesian survival analysis of head and neck cancer data using lognormal
model. Communications in Statistics: Theory and Methods, 43, 392–407.

58
A Review on Various Method of Compounding using
Series and Parallel Structure

Magrisha Namsaw1 and Bhanita Das2


1* Department of Statistics, North Eastern Hill University,
Email ID: magrishanam123@[Link]
2 Department of Statistics, North Eastern Hill University,
Email ID: bhanitadas83@[Link]

Abstract:
Any given system can have a series structure, parallel structure, series
arrangement of parallel structure or parallel arrangement of series structure. In this
article a brief review of the recent developed distribution by the method of
compounding which is motivated by the failure times of a system with different
structural arrangement is made. A brief discussion about the characteristics and their
application is done. The aim of this cahpter is to summarize on the recent developed
distribution and commend on the future work that can be done and continued.
Key Words: Compounding, Series and Parallel structure, System, Lifetime
distribution, Failure times, cumulative distribution function (cdf) and
probability density function (pdf).

1. Introduction
For many decades, Researchers have been working on finding and developing
different ways to introduce parameters so as to expand the existing family of
[Link] objective have always been to make the well known and existing
distribution more flexible with wider characteristics so that it can be used for
modelling data in many disciplines as they are very limited in their characteristics
and are unable to show wide flexibility. Even Weibull Distribution one of the very
flexible distribution in analyzing lifetime data is not capable of modelling when it
comes to non-monotonic failure rate function(such as uni-modal and bathtub
shaped).
Marshall and Olkin (1997) introduce a method of adding a parameters to the
lifetime distribution through compounding based on the failure time of a series or
parallel system with unknown components and many extension of their work has
been done since [Link] to Ross(2010), any system can be represented both
as series arrangement of parallel structures or as a parallel arrangement of structures.


Corresponding author: magrishanam123@[Link]

59
Based on failures time of such system, Researchers have also proposed different new
family of distribution by following the procedure ideas of Marshall and Olkin
(1997). Adamidis and Loukas(1998) extended their work and introduced the two-
parameter exponential geometric distribution and similarly many more distribution
have been developed. This method of compounding have caught the attention of
many researchers as they provide flexibility and versatility to the new developed
distribution.
The objective of the chapter is to present the recent developed distributions by
the method of compounding based on the failure time of a system having different
structural [Link], we can still come across many different structural
arrangement of a system other than what have been mention [Link], this article will
also motivate the researchers to explore this method of compounding for a system
with different structural arrangement. The rest of the chapter is organized as follows.
In section 2, we have reviewed on the recent developed distribution by compounding
based on the failure time of a system with different structure. In section 3, we
present the [Link] we will make a conclusion remark in section 4 and
discuss the future work.
2. Compounding of distribution via system structure approaches
The compounding of distributions in reliability analysis is closely linked to the
structural arrangement of a system and the corresponding failure mechanisms. We
consider a system composed of a finite number of subsystems, where the system‘s
lifetime depends on the structural interdependence of these subsystems. For
compounding purposes, it is assumed that the N subsystems are independent and
identically distributed (i.i.d.) at any given time, with each subsystem following a
suitable probability distribution. Here, the number of subsystems N is treated as a
discrete random variable defined on the domain {1,2,3,… } and may follow any
appropriate discrete distribution.
The conditional cumulative distribution function (cdf) of the N-subsystem
system is derived using the framework of order statistics, depending on whether the
system failure corresponds to the minimum value, maximum value, or a combination
of both. The key objective is to connect this conditional cdf with the marginal
behavior of individual subsystems by employing the concept of compounding,
which is given by
𝐹(𝑥) = ∑ 𝑃(𝜃)𝐹(𝑥|𝜃) Eq. (1)
Although multiple structural configurations exist, the most commonly analyzed
arrangements are the series structure, the parallel structure, and mixed
configurations obtained by combining these two. Accordingly, the compounding of
distributions will be discussed for the following system structures:
(i) The series structure
(ii) The Parallel structure
(iii) The series-Parallel structure
(iv) The Parallel-series structure

60
2.1 Compound distribution for series structure
A system is said to possess a series structure when all of its components or
subsystems are connected in sequence such that the proper functioning of the system
requires every subsystem to operate successfully. In this configuration, the failure of
any single subsystem results in the failure of the entire system. Consequently, the
system‘s lifetime can be represented as the minimum of the failure times among all
subsystems. For given N, let F(x;𝜏) be the cdf of the i.i.d. N-subsystem denoted by
𝑋1 , 𝑋2 , . . . . . 𝑋𝑁 and P (N;𝜃) be the pdf of the random variable N. Now, let 𝑌 =
𝑚𝑖𝑛,𝑋𝑖 -𝑁 𝑖=1 , then the conditional cdf of Y N is given by
𝑁
𝐹(𝑥 𝑁; 𝜏) = 1 − (1 − 𝐹(𝑥)) Eq. (2)
Hence, the required marginal cdf of Y, is given by
𝐹(𝑌; 𝜏, 𝜃) = ∑∞ ( ; 𝜏)𝐹(𝑥 𝑁; 𝜏) Eq. (3)
A detail review on the recent developed distribution by this method is presented
in Table I:
Table I: Distributions using series structure approach
[Link] Year Distribution Author(s)
1 2007 Exponential Poisson Kus(2007)
2 2008 Exponential logarithmic Tahmasi and Rezaei(2008)
3 2011 Weibull Geometric Ortega et al. (2011)
4 2012 Weibull Poisson Lu and Shi (2012)
5 2012 Lindley Geometric Zakerzadeh and Mahmoudi (2012)
6 2012 Exponentiated Exponential binomial Bakouch et al. (2012)
7 2013 Pareto Poisson-Lindley Asgharzadeh et al. (2013)
8 2013 Exponentiated Lomax Poisson Wallace et al. (2013)
9 2013 Generalized gamma power series Silva et al. (2013)
10 2013 Binomial Lifetime Alkarni (2013)
11 2013 Exponential Poisson-Lindley Barreto Souza and Bakouch (2013)
12 2014 Exponential-Weibull Cordeiro (2014)
13 2014 Poisson Lindley-G Asgharzadeh et al. (2014)
14 2015 Burr XII power series Silva and Cordeiro (2015)
15 2015 Modified Weibull geometric Wang and Elbatal (2015)
16 2015 Power series beta Weibull Ortega et al. (2015)
17 2016 Burr XII geometric type I Lanjoni (2016)]
18 2016 Generalized Inverse Weibull Power series Hassan et al. (2016)
19 2016 Exponential modified discrete Lindley Yilmaz et al. (2016)
20 2016 Gompertz-Power series Jafari and Tahmasebi (2016)
21 2016 Generalized gamma power series Silva et al. (2016)
22 2016 Additive Weibull-Geometric Elbatal et al. (2016)
Exponential-Discrete generalized
23 2016 Nekoukhou and Bidran (2016)
Exponential
Reflected generalized Topp Leone Power
24 2017 Condino and Domna (2017)
series
Exponentiated Transmuted Weibull
25 2017 Fattah et al. (2017)
Geometric

61
26 2017 Exponentiated Lomax Geometric Hassan et al. (2017)
27 2017 Exponential Pareto Power series Elbatal et al. (2017)
28 2017 Exponentiated Inverse Weibull-Geometric Chung (2017)
29 2017 Topp Leone Geometric Okasha (2017)
30 2017 Linear rate Power series Mahmoudi and Jafari (2017)
31 2017 Half logistic Poisson Muhammad and Yahaya (2017)
32 2017 Weibull-Conway Maxwell-Poisson Gupta and Huang (2017)
33 2017 Compounded Geometric-I Chowdhury (2017)
34 2018 Power Lomax POisson Hassan and Nasir (2018)
35 2018 Ailamujia Power series Rashid (2018)
36 2018 Exponentiated Burr XII Power series Nasir et al. (2018)
37 2018 Exponentiated Power Lindley Power series Alizadeh et al. (2018)
38 2018 Weibull Lindley Asgharzadeh et al. (2018)
39 2018 Generalized Lindley Power series Rashid et al. (2018)
40 2018 Quasi Xgamma-Poison Sen et al. (2018)
41 2019 min Weibull Burr Nasir et al. (2019)
42 2019 Exponentiated generalized Power series Nasiry et al. (2019)
43 2019 Inverse Weibull-geometric Chakrabarty and Chowdhury (2019)
44 2019 Topp Leone exponential Poisson Abouelmagd (2019)
45 2020 Generalized hazard rate Power series Roozegar et al.(2020)
46 2021 Ishita Power series Hassan et al. (2021)
47 2021 Inverse Gamma Power series Rivera et al. (2021)
48 2021 Lindley-Burr XII Power series Makubate et al. (2021)
49 2021 Poisson Generalized exponential G Morshedy et al. (2021)
50 2022 Inverse Exponentiated Lomax Power series Hassan et al.(2022)
51 2022 Inverse Lindley Power series Shakhatreh et al. (2022)
52 2022 Power Series generalised Power Weibull Abonongo et al. (2022)
53 2023 Inverse Power muth Power series Barrios-Blanco et al. (2023)
54 2023 Pareto-Poisson Elshahhat et al. (2023)
55 2023 Gamma Zero Truncated Poisson Niyomdecha et al. (2023)
56 2024 Unit Gombertz Power series Yousef et al. (2024)
57 2025 Zhang-Power Series Memarbashi et al. (2025)
58 2025 Alpha Power Modified Weibull-Geometric Namsaw et al. (2025)

62
2.2 Compound distribution for parallel structure
The compounding of a distribution for a parallel system structure follows an
approach similar to that used for the series structure. The key difference lies in the
failure mechanism: for a parallel structure, the system fails only when all subsystems
fail. Hence, the lifetime of the system is determined by the maximum of the failure
times among all subsystems.
For given N, let F(x;𝜏) be the cdf of the i.i.d. N-subsystem denoted by
𝑋1 , 𝑋2 , . . . . . 𝑋𝑁 and P (N;𝜃) be the pdf of the random variable N. Now, let 𝑌 =
𝑚𝑎𝑥,𝑋𝑖 -𝑁 𝑖=1 , then the conditional cdf of Y N is given by
𝑁
𝐹(𝑥 𝑁; 𝜏) = (𝐹(𝑥)) Eq. (4)
Hence, the required marginal cdf of Y, is given by
𝐹(𝑌; 𝜏, 𝜃) = ∑∞ ( ; 𝜏)𝐹(𝑥 𝑁; 𝜏) Eq. (5)
A detail review on the recent developed distribution by this method is presented
in Table II:
Table II: Distributions using parallel structure approach
[Link] Year Distribution Author(s)
1 2010 Generalised exponential geometric Silva et al. (2010)
2 2012 Exponential Truncated Poisson Rezaei and Tahmasi (2012)
Mahmoudi and
3 2013 Exponentiated Weibull-Poisson
Sepahdar(2013)
4 2015 Exponentiated Extended Weibull Power series Tahmasi and Jafari (2015)
5 2015 Burr XII geometric type II Lanjoni (2015)
6 2015 Generalized Gompertz Power series Tahmasebi and Jaferi (2015)
7 2016 Inverse Weibull Power Series Shafiei (2016)
8 2016 Generalized extended Weibull Power series Alkarni al. (2016)
9 2016 Poisson half-logistic Abdel Hamid (2016)
10 2016 Dagum-Poisson Oluyede et al. (2016)
11 2017 Complementary Compound Lindley Power series Rashid et al. (2017)
12 2017 Exponentiated Power Lindley geometric Alizadeh et al. (2017)
13 2017 Exponentiated Power Lindley Poisson Pararai et al. (2017)
14 2017 Complementary Exponentiated Burr XII Poisson Muhammad (2017)
15 2017 Complementary Lindley-Geometric Gui et al. (2017)
16 2017 Compounded Geometric Chowdhury (2017)
17 2018 Complementary Generalized Lindley Power series Ahmad et al. (2018)
18 2019 Geometric-Zero Truncated Poisson Akdogan et al. (2019)
Complementary Generalized Power Weibull Power
19 2020 Nasiru and Abubakari (2020)
series
20 2020 Lindley Power series Si et al. (2020)
21 2021 Generalized Lindley-Weibull Makubate et al. (2021)
22 2021 Power series power function Hassan and Assar (2021)
23 2021 Complementary Poisson Generalized Half-logistic Muhammad and Liu (2021)

63
24 2022 Generalized Exponential Logarithmic Hakamipour et al. (2022)
25 2022 Complementary Bell Weibull Algarni (2022)
26 2023 Geometric-Mixture Exponential Karakaya et al.(2023)
27 2023 Power Inverted Topp Leone Power series El-Saud et al. (2023)
28 2023 Unit Burr XII Power series Zayed et al. (2023)
29 2023 Unit Exponentiated Half logistic Power series Alghamdi et al. (2023)
30 2025 Compound Truncated Poisson log-normal Meraou et al. (2025)

2.3 Compound distribution for series-parallel structure(combination of series and


parallel structure)
Many practical systems are composed of combinations of series and parallel
configurations. To evaluate the reliability of such systems, the overall structure can
first be decomposed into homogeneous subsystems. Each subsystem is then treated
as an independent unit, and its reliability is computed individually. These units are
subsequently recombined, either in series or in parallel, to obtain the reliability of
the entire system. Such mixed configurations are frequently encountered in real-
world applications. In the present context, the proposed family is motivated by a
system composed of series-arranged subsystems, where each subsystem itself
consists of components organized in a parallel structure.
The proposed family is motivated by a system composed of series subsystems,
where each subsystem consists of components arranged in a parallel structure.
Specifically, we consider a system formed by k series subsystems that operate
independently at any given time. Here, k is taken as a realization of the random
variable K, which follows a discrete distribution with probability mass function P(K
= k,𝜃) (for example, zero-truncated Poisson, geometric, or power series
distributions).. Assume that a system j, j = 1, 2... k, includes 𝑙𝑗 parallel components,
again functioning independently at a given time with pmf P (𝐿𝑗 = 𝑙𝑗 , 𝛽) with failure
times, say 𝑋1𝑗 , 𝑋2𝑗 , . . . , 𝑋𝑙𝑗𝑗 which are independent and identically distributed (iid)
RVs subjecting to a continuous lifetime distribution with pdf f(x; 𝜏 ) and cdf F(x; 𝜏
)(e.g exponential, Weibull, Lomax, Burr type XII, ...). The 𝑗𝑡𝑕 system succeeds if at
least one of the system components succeeds. In other words, the jth system fails if
all system components fail, that is 𝑍𝑗 = 𝑚𝑎𝑥 .𝑋1𝑗 , 𝑋2𝑗 , . . . , 𝑋𝑙𝑗𝑗 /, j = 1, 2... k. Since the k
systems are series systems, and then this new series system can succeed when all the
included systems succeed. In other words, the system fails when at least one of the
included systems fails, that is,
𝑙 𝑘
𝑗
𝑋 = 𝑚𝑖𝑛(𝑚𝑎𝑥(𝑋𝑖𝑗 )𝑖=1 )𝑘𝑗=1 = 𝑚𝑖𝑛(𝑍𝑗 )𝑗=1 Eq. (6)
where 𝑋𝑖𝑗 , i = 1, ... , 𝑙𝑗 , j = 1, ... , k are iid RVs and 𝑙𝑗 is a realization of RV 𝐿𝑗 .

64
𝑗 𝑙
Now, For all j, j = 1, ... , k, let 𝑍𝑗 = 𝑚𝑎𝑥(𝑋𝑖𝑗 )𝑖=1 . Therefore, the conditional density
function of 𝑍𝑗 , given 𝐿𝑗 = 𝑙𝑗 , is given by
𝑙
𝐹(𝑧𝑗 |𝑙𝑗 ; 𝜏) = [𝐹𝑋 (𝑧𝑗 ; 𝜏)] 𝑗 Eq. (7)
The marginal CDF of 𝑍𝑗 is given by
𝐹(𝑧𝑗 ; 𝛽, 𝜏) = ∑∞ 𝑁=1 𝑓 (𝑧𝑗 |𝑙𝑗 ; 𝜏)𝑃(𝐿𝑗 = 𝑙𝑗 ) Eq. (8)
𝑙𝑗 𝑘
Since 𝑋 = 𝑚𝑖𝑛(𝑚𝑎𝑥(𝑋𝑖𝑗 )𝑖=1 )𝑘𝑗=1 =𝑚𝑖𝑛(𝑍𝑗 )𝑗=1 , then the conditional density
function of X, given K=k, is given by
𝑘
𝐹(𝑥|𝑘; 𝛽, 𝜏) = 1 − ,1 − 𝐹𝑍 (𝑥; 𝛽, 𝜏)-𝑘 = 1 − [1 − 𝐹(𝑧𝑗 ; 𝛽, 𝜏)] Eq. (9)
The marginal CDF of X is given by
𝐹(𝑥; 𝛽, 𝜏, 𝜃) = ∑∞𝑘=1 𝐹 (𝑥|𝑘; 𝛽, 𝜏)𝑃(𝑘 = 𝑘) Eq. (10)
This is the required CDF of the distribution.
Some of the distribution that have been developed using this method are
discussed below in Table III:
Table III: Distributions using series-parallel structure approach
[Link] Year Distribution Author(s)

1 2016 Unified Lifetime Rodrigues et al. (2016)


2 2016 Geometric Power Lindley Poisson Mansour et al. (2016)
2 2017 Doubly Poisson-exponential Abdel-hamid and Hashem (2017)
3 2019 Exponential geoemtric power series Goldoust (2019)
4 2022 Poisson-geometric Lomax Abushal and Abdel-Hamid (2022)
5 2022 Power series Exponential Power series Roozegar et al. (2022)

2.4 Compound distribution for parallel-series structure (combination of parallel and


series structure)
In this framework, the proposed family is inspired by a system consisting of
parallel subsystems, each of which is composed of components connected in a series
structure. The system is assumed to comprise k parallel subsystems that function
independently at a given time, where k is a realization of the random variable K. The
variable K is governed by a discrete distribution with probability mass function P
P(K = k,𝜃) (such as the zero-truncated Poisson, geometric, or power series
distribution). Assume that a system j, j = 1, 2... k includes 𝑙𝑗 series components,
again functioning independently at a given time with pmf P(𝐿𝑗 = 𝑙𝑗 ) with failure
times, say 𝑋1𝑗 , 𝑋2𝑗 , . . . , 𝑋𝑙𝑗𝑗 which are independent and identically distributed (iid)
RVs subjecting to a continuous lifetime distribution with pdf f(x; 𝜃 ) and cdf F(x; 𝜃
)(e.g exponential, Weibull, Lomax, Burr type XII, ...). The 𝑗𝑡𝑕 system succeeds if all
of the subsystem components succeed. In other words, the jth system fails if at least
one of the system components fail, that is, 𝑍𝑗 = 𝑚𝑖𝑛 .𝑋1𝑗 , 𝑋2𝑗 , . . . , 𝑋𝑙𝑗𝑗 /, j = 1, 2, ... ,
k. Since the k subsystems are parallel systems, then this new parallel subsystem can
65
succeed when at least one of the included subsystems succeed. In other words, the
system fails when all of the included subsystems fails, that is,
𝑙 𝑘
𝑗
𝑋 = 𝑚𝑎𝑥(𝑚𝑖𝑛(𝑋𝑖𝑗 )𝑖=1 )𝑘𝑗=1 = 𝑚𝑎𝑥((𝑍𝑗 ) Eq. (11)
𝑗=1
where 𝑋𝑖𝑗 , i = 1, ... , 𝑙𝑗 , j = 1, ... , k are iid RVs and 𝑙𝑗 is a realization of RV 𝐿𝑗 .
𝑙𝑗
Now, for all j, j = 1... k, let 𝑍𝑗 = 𝑚𝑖𝑛(𝑋𝑖𝑗 )𝑖=1. Therefore, the conditional density
function of 𝑍𝑗 , given 𝐿𝑗 = 𝑙𝑗 , is given by
𝑙
𝐹(𝑧𝑗 |𝑙𝑗 ; 𝜏) = 1 − [1 − 𝐹𝑋 (𝑧𝑗 ; 𝜏)] 𝑗 Eq. (12)
The marginal CDF of 𝑍𝑗 is given by
𝐹(𝑧𝑗 ; 𝛽, 𝜏) = ∑∞
𝑁=1 𝑓 (𝑧𝑗 |𝑙𝑗 ; 𝜏)𝑃(𝐿𝑗 = 𝑙𝑗 ) Eq. (13)
𝑙𝑗 𝑘
Since 𝑋 = 𝑚𝑎𝑥(𝑚𝑖𝑛(𝑋𝑖𝑗 )𝑖=1 )𝑘𝑗=1 =𝑚𝑎𝑥(𝑍𝑗 ) , then the conditional density
𝑗=1
function of X, given K=k, is given by
𝑘
𝐹(𝑥|𝑘; 𝛽, 𝜏) = ,𝐹𝑍 (𝑥; 𝛽, 𝜏)-𝑘 = 1 − [1 − 𝐹(𝑧𝑗 ; 𝛽, 𝜏)] Eq. (14)
The marginal CDF of X is given by
𝐹(𝑥; 𝛽, 𝜏, 𝜃) = ∑∞
𝑘=1 𝐹 (𝑥|𝑘; 𝛽, 𝜏)𝑃(𝑘 = 𝑘) Eq. (15)
This is the required CDF of the distribution.
In Table IV, we have listed some of the distribution that have been developed
using this method.
Table IV: Distributions using parallel-series structure approach
[Link] Year Distribution Author(s)
1 2018 Exponential geometric Power series Goldoust et al. (2018)
2 2018 Geometric-Poisson-Rayleigh Nadarajah et al. (2018)
3 2019 Lifetime Power series2 Goldoust et al. (2019)
4 2021 Exponentail doubly Poisson Hashem and Alyami (2021)
5 2022 Poisson-logarithmic half logistic Hashem et al. (2022)

3. Characteristics and Areas of Applications


3.1 The Hazard rate function
This approach of compounding have shown to have various shapes and
behaviours. Some of the hazard rate shapes that we have seen so far from all the
recent developed distribution are;
(i) Monotonically increasing
(ii) Monotonically decreasing
(iii) Bathtub shaped
(iv) Upside down bathtub shaped

66
3.2 Estimation method
Some of the estimation methods that have been used for parameter estimation
are;
(i) Maximum Likelihood
(ii) Moments
(iii) Least Squares
(iv) Weighted Least Squares
(v) Bayes under LINEX loss function
(vi) Bayes under GE loss function
For some of the compound distribution the estimation methods have also been
studied under the censoring scheme (ex. Progressive type-II censoring)

3.3 Areas of application


With the flexibility in their characteristics, this method of compound distribution
have shown to be applicable in various field. Some of the field of application is
given below with one specific example
(i) Physical Laboratory - Strength of glass fibre
(ii) Meteorology- Air temperature
(iii) Hospital- Cancer data (survival times of patients)
(iv) Engineering- Failure of mechanical components
(v) Offices-Bank (waiting time customers)

4. Concluding Remarks
In this chapter, we have present the recent developed distribution by the method
of compounding which follows the ideas of the failure time of the possible available
structure of a system by following the basic principles (minimum and maximum)
used in series and parallel [Link] have seen approaches with a system having
series structure, parallel structure, series-parallel structure or parallel-series structure
and we have discussed it all here with the recent developed distributions. However,
we can still come across many different structural arrangement of a system so we
offer the researchers to explore and illustrate those different approaches of the
system structural arrangement in compounding. The possible future work one can
explore are
(i) To compare different approach of compounding and illustrate the
application and usefulness;
(ii) To propose new compounding approach with new and different structural
arrangement of a system;
(iii) To prepare a review on the bi-variate distribution by the method of
compounding.

67
References
1) Abdel-Hamid, A. H. (2016). Properties, estimations and predictions for a
Poisson-half-logistic distribution based on progressively type-II censored
samples. Applied Mathematical Modelling, 40(15–16), 7164–7181.
2) Abdel-Hamid, A. H., & Hashem, A. F. (2017). A new lifetime distribution
for a series-parallel system: Properties, applications and estimations under
progressive type-II censoring. Journal of Statistical Computation and
Simulation, 87(5), 993–1024.
3) Abonongo, A. I. L., Luguterah, A., & Nasiru, S(2022) . Power series
generalised power Weibull class of distributions. Asian Journal of
Probability & Statistics, 17(2), 28-51.
4) Abouelmagd, T. H. M., Hamed, M. S., Handique, L., Goual, H., Ali, M. M.,
Yousof, H. M., & Korkmaz, M. C. (2019). A new class of distributions
based on the zero truncated Poisson distribution with properties and
applications. Journal of Nonlinear Sciences and Applications, 12(3), 152–
164.
5) Abushal, T. A., & Abdel-Hamid, A. H. (2022). Inference on a new
distribution under progressive-stress accelerated life tests and progressive
type-II censoring based on a series-parallel system. AIMS Mathematics,
7(1), 425–454.
6) Adamidis, K., & Loukas, S. (1998). A lifetime distribution with decreasing
failure rate. Statistics & Probability Letters, 39(1), 35–42.
7) Ahmad, Z., Rashid, A., & Jan, T. R. (2018). Generalized version of
complementary Lindley power series distribution. Pakistan Journal of
Statistics and Operation Research, 14(1), 139–155.
8) Akdogan, Y., Kus, C., Bı dram, H., & Kı nacı , İ . (2019). Geometric-zero
truncated Poisson distribution: Properties and applications. Gazi University
Journal of Science, 32(4), 1339–1354.
9) Algarni, A. (2022). Group acceptance sampling plan based on new
compounded three-parameter Weibull model. Axioms, 11(9), 438.
10) Alghamdi, S. M., Shrahili, M., Hassan, A. S., Mohamed, R. E., Elbatal, I., &
Elgarhy, M. (2023). Analysis of milk production and failure data: Using unit
exponentiated half logistic power series class of distributions. Symmetry,
15(3), 714.
11) Alizadeh, M., Bagheri, S. F., Samani, E. B., Ghobadi, S., & Nadarajah, S.
(2018). Exponentiated power Lindley power series class of distributions:
Theory and applications. Communications in Statistics—Simulation and
Computation, 47(9), 2499–2531.
12) Alizadeh, M., Bagheri, S. F., Alizadeh, M., & Nadarajah, S. (2017). A new
four-parameter lifetime distribution. Journal of Applied Statistics, 44(5),
767–797.
13) Alkarni, S. H. (2016). Generalized extended Weibull power series family of
distributions. Journal of Data Science, 14(3), 415–439.
68
14) Alkarni, S. H. (2013). A class of truncated binomial lifetime distributions.
Open Journal of Statistics, 3(5), 305–311.
15) Asgharzadeh, A., Bakouch, H. S., Nadarajah, S., & Esmaeili, L. (2014). A
new family of compound lifetime distributions, Kybernetika, 50(1), 142-
169.
16) Asgharzadeh, A., Bakouch, H. S., & Esmaeili, L. (2013). Pareto Poisson
Lindley distribution with applications. Journal of Applied Statistics, 40(8),
1717–1734.
17) Asgharzadeh, A., Nadarajah, S., & Sharafi, F. (2018). Weibull Lindley
distribution. REVSTAT—Statistical Journal, 16(1), 87–113.
18) Bakouch, H. S., Ristić , M. M., Asgharzadeh, A., Esmaily, L., & Al-
Zahrani, B. M. (2012). An exponentiated exponential binomial distribution
with application. Statistics & Probability Letters, 82(6), 1067–1081.
19) Barreto-Souza, W., & Bakouch, H. S. (2013). A new lifetime model with
decreasing failure rate. Statistics, 47(2), 465–476.
20) Barreto-Souza, W., de Morais, A. L., & Cordeiro, G. M. (2011). The
Weibull-geometric distribution. Journal of Statistical Computation and
Simulation, 81(5), 645–657.
21) Barrios-Blanco, L., Gallardo, D. I., Gómez, H. J., & Bourguignon, M.
(2023). A compound class of inverse-power Muth and power series
distributions. Axioms, 12(4), 383.
22) Chakrabarty, J. B., & Chowdhury, S. (2019). Compounded inverse Weibull
distributions: Properties, inference and applications. Communications in
Statistics—Simulation and Computation, 48(7), 2012–2033.
23) Chowdhury, S., Mukherjee, A., & Nanda, A. K. (2017). On compounded
geometric distributions and their applications. Communications in
Statistics—Simulation and Computation, 46(3), 1715–1734.
24) Chung, Y., Dey, D. K., & Jung, M. (2017). The exponentiated inverse
Weibull geometric distribution. Pakistan Journal of Statistics, 33(3).
25) Condino, F., & Domma, F. (2017). A new distribution function with
bounded support: The reflected generalized Topp-Leone power series
distribution. Metron, 75(1), 51–68.
26) Cordeiro, G. M., Ortega, E. M. M., & Lemonte, A. J. (2014). The
exponential–Weibull lifetime distribution. Journal of Statistical
Computation and Simulation, 84(12), 2592–2606.
27) El-Morshedy, M., Alshammari, F. S., Hamed, Y. S., Eliwa, M. S., &
Yousof, H. M. (2021). A new family of continuous probability distributions.
Entropy, 23(2), 194.
28) El-Saeed, A. R., Hassan, A. S., Elharoun, N. M., Al Mutairi, A., Khashab,
R. H., & Nassr, S. G. (2023). A class of power inverted Topp-Leone
distribution: Properties, different estimation methods & applications. Journal
of Radiation Research and Applied Sciences, 16(4), 100643.

69
29) Elbatal, I., Mansour, M. M., & Ahsanullah, M. (2016). The additive
Weibull-geometric distribution: Theory and applications. Journal of
Statistical Theory and Applications, 15(2), 125–141.
30) Elbatal, I., Zayed, M., Rasekhi, M., & Butt, N. S. (2017). The exponential
Pareto power series distribution: Theory and applications. Pakistan Journal
of Statistics and Operation Research, 13(3), 603–615.
31) Elshahhat, A., El-Sherpieny, E.-S. A., & Hassan, A. S. (2023). The Pareto–
Poisson distribution: Characteristics, estimations and engineering
applications. Sankhya A, 85(1), 1058–1099.
32) Fattah, A. A., Nadarajah, S., & Ahmed, A.-H. N. (2017). The exponentiated
transmuted Weibull geometric distribution with application in survival
analysis. Communications in Statistics—Simulation and Computation,
46(6), 4244–4263.
33) Goldoust, M., Rezaei, S., Si, Y., & Nadarajah, S. (2018). A lifetime
distribution motivated by parallel and series structures. Communications in
Statistics—Theory and Methods, 47(13), 3052–3072.
34) Goldoust, M., Rezaei, S., Si, Y., & Nadarajah, S. (2019). Lifetime
distributions motivated by series and parallel structures. Communications in
Statistics—Simulation and Computation, 48(2), 556–579.
35) Goldoust, M., Mohammadpour, A., Alizadeh, M., & Hamedani, G. (2019).
A new generalized family of lifetime distributions motivated by parallel and
series structures. Statistics, Optimization & Information Computing, 7(4),
779–801.
36) Gui, W., Zhang, H., & Guo, L. (2017). The complementary Lindley-
geometric distribution and its application in lifetime analysis. Sankhya B,
79(2), 316–335.
37) Gupta, R. C., & Huang, J. (2017). The Weibull–Conway–Maxwell–Poisson
distribution to analyze survival data. Journal of Computational and Applied
Mathematics, 311, 171–182.
38) Hakamipour, N., Zhang, Y., & Nadarajah, S. (2022). A new family of
compound exponentiated logarithmic distributions with applications to
lifetime data. Mathematica Slovaca, 72(5), 1337–1354.
39) Hashem, A. F., Kuş , C., Pekgör, A., & Abdel-Hamid, A. H. (2022).
Poisson–logarithmic half-logistic distribution with inference under a
progressive-stress model based on adaptive type-II progressive hybrid
censoring. Journal of the Egyptian Mathematical Society, 30(1), 15.
40) Hassan, A. S., Almetwally, E. M., Gamoura, S. C., Metwally, A. S. M., et
al. (2022). Inverse exponentiated Lomax power series distribution: Model,
estimation, and application. Journal of Mathematics, 2022, 1–17.
41) Hassan, A. S., Assar, M. S., & Ali, K. A. (2016). The compound family of
generalized inverse Weibull power series distributions. British Journal of
Applied Sciences & Technology, 14(3), 1–18.

70
42) Hassan, A. S., & Assar, S. M. (2021). A new class of power function
distribution: Properties and applications. Annals of Data Science, 8(1), 205–
225.
43) Hassan, A. S., & Nassr, S. G. (2018). Power Lomax Poisson distribution:
Properties and estimation. Journal of Data Science, 18(1), 105–128.
44) Hassan, A. S., & Abdelghafar, M. A. (2017). Exponentiated Lomax
geometric distribution: Properties and applications. Pakistan Journal of
Statistics and Operation Research, 13(3), 545–566.
45) Hassan, A., Akhtar, N., Rashid, A., & Iqbal, A. (2021). A new compound
lifetime distribution: Ishita power series distribution with properties and
applications. Journal of Statistics Applications & Probability, 10(3), 883–
896.
46) Jafari, A. A., & Tahmasebi, S. (2016). Gompertz–power series distributions.
Communications in Statistics—Theory and Methods, 45(13), 3761–3781.
47) Karakaya, K., Kı nacı , İ ., Kuş , C., & Akdoğ an, Y. (2023). A new
distribution with four parameters: Properties and applications. Sigma
Journal of Engineering and Natural Sciences, 41(2), 276–287.
48) Kuş , C. (2007). A new lifetime distribution. Computational Statistics &
Data Analysis, 51(9), 4497–4509.
49) Lanjoni, B. R., Ortega, E. M. M., & Cordeiro, G. M. (2016). Extended Burr
XII regression models: Theory and applications. Journal of Agricultural,
Biological, and Environmental Statistics, 21(2), 203–224.
50) Lu, W., & Shi, D. (2012). A new compounding life distribution: The
Weibull–Poisson distribution. Journal of Applied Statistics, 39(1), 21–38.
51) Mahmoudi, E., & Jafari, A. A. (2017). The compound class of linear failure
rate–power series distributions: Model, properties, and applications.
Communications in Statistics—Simulation and Computation, 46(2), 1414–
1440.
52) Mahmoudi, E., & Sepahdar, A. (2013). Exponentiated Weibull–Poisson
distribution: Model, properties and applications. Mathematics and
Computers in Simulation, 92, 76–97.
53) Makubate, B., Gabanakgosi, M., Chipepa, F., & Oluyede, B. (2021). A new
Lindley–Burr XII power series distribution: Model, properties and
applications. Heliyon, 7(6), e07234.
54) Mansour, M. M., Ahsanullah, M., Nofal, Z. M., & Khalaf, O. H. (2016).
Geometric power Lindley Poisson distribution: Properties and applications.
Journal of Statistical Theory and Applications, 15(4), 313–325.
55) Marshall, A. W., & Olkin, I. (1997). A new method for adding a parameter
to a family of distributions with application to the exponential and Weibull
families. Biometrika, 84(3), 641–652.
56) Mebarbashi, H., Etminan, J., & Chahkandi, M (2025). On Zhang-poer
series distributions with application to lifetime modeling. Journal of Mahani
Mathematical Research, 14(1), 505-525
71
57) Meraou, M. A., Raqab, M. Z., Kundu, D., & Alqallaf, F. A. (2024).
Inference for compound truncated Poisson log-normal model with
application to maximum precipitation data. Communications in Statistics—
Simulation and Computation, 53(1), 1–22.
58) Muhammad, M. (2017). The complementary exponentiated Burr XII
Poisson distribution: Model, properties and application. Journal of Statistics
Applications & Probability, 6(1), 33–48.
59) Muhammad, M., & Liu, L. (2021). A new three-parameter lifetime model:
The complementary Poisson generalized half logistic distribution. IEEE
Access, 9, 60089–60107.
60) Muhammad, M., & Yahaya, M. A. (2017). The half logistic–Poisson
distribution. Asian Journal of Mathematics and Applications, 2017, 1–15.
61) Nadarajah, S., Abdel-Hamid, A. H., & Hashem, A. F. (2018). Inference for a
geometric–Poisson–Rayleigh distribution under progressive-stress
accelerated life tests based on type-I progressive hybrid censoring with
binomial removals. Quality and Reliability Engineering International, 34(4),
649–680.
62) Namsaw, M., Das, B., Hazarika, P.J., & Alizadeh M. (2025). The alpha
power modified Weibull-geometric distribution: A comprehensive
mathematical framework with simulation, goodness-of-fit analysis and
informed decision making using real life data. Statistics, optimization &
Information Computing, 14(2), 1060-1087
63) Nasir, A., Jamal, F., Chesneau, C., & Shah, A. A. (2019). A new
compounded four-parameter lifetime model: Properties, cure rate model and
[Link]-01902847v2.
64) Nasir, A., Yousof, H. M., Jamal, F., & Korkmaz, M. C. (2018). The
exponentiated Burr XII power series distribution: Properties and
applications. Stats, 2(1), 15–31.
65) Nasiru, S., & Abubakari, A. G. (2020). Complementary generalized power
Weibull power series family of distributions: Estimation and application.
Eurasian Bulletin of Mathematics, 3(1), 20–37.
66) Nasiru, S., Mwita, P. N., & Ngesa, O. (2019). Exponentiated generalized
power series family of distributions. Annals of Data Science, 6(3), 463–489.
67) Nekoukhou, V., & Bidram, H. (2016). Exponential-discrete generalized
exponential distribution: A new compound model. Journal of Statistical
Theory and Applications, 15(2), 169–180.
68) Niyomdecha, A., & Srisuradetchai, P. (2023). A new compounding life
distribution: Gamma zero-truncated Poisson distribution (Doctoral
dissertation, Thammasat University).
69) Okasha, H. M. (2017). A new family of Topp and Leone geometric
distribution with reliability applications. Journal of Failure Analysis and
Prevention, 17(2), 477–489.

72
70) Oluyede, B. O., Motsewabagale, G., Huang, S., Warahena-Liyanage, G., &
Pararai, M. (2016). Dagum–Poisson distribution: Model, properties and
application. Electronic Journal of Applied Statistical Sciences, 9(1), 169–
188.
71) Ortega, E. M. M., Cordeiro, G. M., Campelo, A. K., Kattan, M. W., &
Cancho, V. G. (2015). A power series beta Weibull regression model for
predicting breast carcinoma. Statistics in Medicine, 34(8), 1366–1388.
72) Pararai, M., Warahena-Liyanage, G., & Oluyede, B. O. (2017).
Exponentiated power Lindley–Poisson distribution: Properties and
applications. Communications in Statistics—Theory and Methods, 46(10),
4726–4755.
73) Ramos, M. W. A., Marinho, P. R. D., da Silva, R. V., & Cordeiro, G. M.
(2013). The exponentiated Lomax Poisson distribution with an application
to lifetime data. Advances and Applications in Statistics, 34(2), 107–123.
74) Rashid, A., Ahmad, Z., & Jan, T. R. (2018). A new lifetime distribution for
series system: Model, properties and application. Journal of Modern Applied
Statistical Methods, 17(1), 43–67.
75) Rashid, A., Ahmad, Z., & Jan, T. R. (2017). Complementary compound
Lindley power series distribution with application. Journal of Reliability and
Statistical Studies, 10(2), 143–158.
76) Rashid, A., Jan, T. R., Bhat, A. H., & Ahmad, Z. (2018). A new compound
lifetime distribution: Model, characterization, estimation and application.
Journal of Applied Mathematics, Statistics and Informatics, 14(2), 45–57.
77) Rezaei, S., & Tahmasbi, R. (2012). A new lifetime distribution with
increasing failure rate: Exponential truncated Poisson. Journal of Basic and
Applied Scientific Research, 2(2), 1749–1762.
78) Rivera, P. A., Calderín-Ojeda, E., Gallardo, D. I., & Gómez, H. W. (2021).
A compound class of the inverse gamma and power series distributions.
Symmetry, 13(8), 1328.
79) Rodrigues, J., Cordeiro, G. M., de Castro, M., & Nadarajah, S. (2016). A
unified class of compound lifetime distributions. Communications in
Statistics—Theory and Methods, 45(8), 2323–2331.
80) Roozegar, R., Hamedani, G. G., Amiri, L., & Esfandiyari, F. (2020). A new
family of lifetime distributions: Theory, application and characterizations.
Annals of Data Science, 7(1), 109–138.
81) Roozegar, R., Nadarajah, S., & Mahmoudi, E. (2022). The power series
exponential power series distributions with applications to failure data sets.
Sankhya B, 84(2), 1–35.
82) Ross, S. M. (2014). Introduction to probability models (11th ed.). Academic
Press.
83) Sen, S., Korkmaz, M. C., & Yousof, H. M. (2018). The quasi Xgamma–
Poisson distribution: Properties and application. İ statistik: Journal of the
Turkish Statistical Association, 11(3), 65–76.
73
84) Shafiei, S., Darijani, S., & Saboori, H. (2016). Inverse Weibull power series
distributions: Properties and applications. Journal of Statistical Computation
and Simulation, 86(6), 1069–1084.
85) Shakhatreh, M. K., Dey, S., & Kumar, D. (2022). Inverse Lindley power
series distributions: A new compounding family and regression model with
censored data. Journal of Applied Statistics, 49(13), 3451–3476.
86) Si, Y., & Nadarajah, S. (2020). Lindley power series distributions. Sankhya
A, 82(1), 242–256.
87) Silva, R. B., Barreto-Souza, W., & Cordeiro, G. M. (2010). A new
distribution with decreasing, increasing and upside-down bathtub failure
rate. Computational Statistics & Data Analysis, 54(4), 935–944.
88) Silva, R. B., Bourguignon, M., & Cordeiro, G. M. (2016). A new
compounding family of distributions: The generalized gamma power series
distributions. Journal of Computational and Applied Mathematics, 303,
119–139.
89) Silva, R. B., & Cordeiro, G. M. (2015). The Burr XII power series
distributions: A new compounding family. Brazilian Journal of Probability
& statistics, 29 (3), 565-589.
90) Tahmasbi, R., & Rezaei, S. (2008). A two-parameter lifetime distribution
with decreasing failure rate. Computational Statistics & Data Analysis,
52(8), 3889–3901.
91) Tahmasebi, S., & Jafari, A. A. (2015). Generalized Gompertz–power series
distributions. Hacettepe Journal of Mathematics and Statistics, 45(5), 1579–
1604.
92) Tahmasebi, S., & Jafari, A. A. (2015). Exponentiated extended Weibull–
power series class of distributions. arXiv preprint, arXiv:1503.08653.
93) Wang, M., & Elbatal, I. (2015). The modified Weibull geometric
distribution. Metron, 73(3), 303–315.
94) Yilmaz, M., Hameldarbandi, M., & Kemaloglu, S. A. (2016). Exponential-
modified discrete Lindley distribution. SpringerPlus, 5(1), 1–21.
95) Yousef, M., Hassan, A., & Almetwally, E. (2024). Statistical inference for
the unit Gompertz power series distribution using ranked set sampling with
applications. Assiut University Journal of Multidisciplinary Scientific
Research, 53(1), 154–189.
96) Zakerzadeh, H., & Mahmoudi, E. (2012). A new two-parameter lifetime
distribution: Model and properties. arXiv preprint, arXiv:1204.4248.
97) Zayed, M. A., Hassan, A. S., Almetwally, E. M., Aboalkhair, A. M., Al-
Nefaie, A. H., Almongy, H. M., et al. (2023). A compound class of unit Burr
XII model: Theory, estimation, fuzzy, and application. Scientific
Programming, 2023, 1–15.

74
Educational Status of Rani Khamar Village, Palasbari,
Assam: A Socio-Economic Survey-Based Study
Ananya Guha¹, Menaka Sikdar2
*¹Research Scholar, Department of Statistics, Gauhati University, Email ID:
annyguha25@[Link]
2
Assistant Professor, Department of Statistics, Arya Vidyapeeth College (A), Email
ID: sikdarmenaka@[Link]

Abstract
Education is universally acknowledged as the foundation of socio-economic
development, yet rural areas in India continue to experience disparities in access and
attainment. This chapter presents the findings of a socio-economic and educational
survey conducted in Rani Khamar village, Palasbari, Assam. The study, carried out
in 2021, examined the educational status of households and explored its associations
with socio-economic indicators such as gender, income, and occupation. Data were
collected using a structured household survey and analyzed through diagrammatic
tools, chi-square tests, and confidence intervals. The results revealed that nearly 15
percent of the population remained uneducated, while household income and
occupation showed notable correlations with education. Interestingly, gender was
not significantly associated with educational status, but around 8 percent of
households reported that children were not enrolled in school. These findings
highlight the complex interplay between economic conditions and educational
participation in rural Assam and underscore the need for targeted policy
interventions.
Keywords: Assam, education, income, occupation, rural development, socio-
economic survey

I. INTRODUCTION
Education is often described as the cornerstone of socio-economic development.
It not only enhances the quality of individual life but also contributes to the
collective progress of a community. In rural India, however, access to education is
shaped by complex interactions between poverty, occupation, infrastructure, and
cultural perceptions. These inequalities are particularly pronounced in Assam, a state
characterized by ethnic diversity, agricultural dependence, and limited rural
infrastructure.


Corresponding author: annyguha25@[Link]

75
Village-level surveys provide micro-level insights into these challenges. By focusing
on households, they highlight both structural and behavioral factors that influence
educational attainment. Against this backdrop, the present study examines the
educational status of Rani Khamar village, Palasbari, Assam, and seeks to uncover
the socio-economic correlates of education within this community.

II. OBJECTIVES OF THE STUDY


The present survey was conducted in Rani Khamar village with the following
specific objectives:
1. To assess the educational status of individuals and households in the
village.
2. To examine the relationship between socio-economic factors such as
gender, occupation, and household income with educational attainment.
3. To analyse school enrolment patterns, including the proportion of children
attending or not attending school.
4. To assess household aspirations regarding the future education of their
children.
5. To study intergenerational trends by comparing parents‘ educational status
with that of their wards.

III. LITERATURE REVIEW


The discrepancies found at the village level in the study are consistent with a
district-level analysis of Assam by (Bhagwati & Sarma, 2024), which revealed
regional differences in educational status. (Phukan & Gogoi ,2013) studied rural
Assam and found that economic and infrastructural constraints significantly affect
students‘ educational opportunities. Previous research has consistently demonstrated
that socio-economic factors such as income, occupation, and parental education
significantly shape children‘s access to schooling in rural areas (Tilak, 2002).
Additionally, a study by (Sarma & Das, 2021) analyzed educational inequalities in
rural Assam, emphasizing how limited access to quality education facilities
exacerbates socio-economic disparities. Another study by (Dutta et al.,2025)
developed the Assam Rural Livelihood and Farming Scale (ARLFS), highlighting
the socio-economic challenges faced by rural households, which indirectly affect
educational attainment.
At a methodological level, socio-economic surveys often employ chi-square
tests to identify associations between categorical factors such as gender, income
group, or occupation and educational attainment. These micro-level statistical
investigations provide evidence for designing localized interventions. The present
study situates itself within this tradition, contributing to the empirical understanding
of rural education in Assam.

76
IV. METHODOLOGY
The household survey was conducted on 15th February,2022 in Rani Khamar
village, a rural settlement in Palasbari, Assam. A structured questionnaire was
administered to village households, covering demographic characteristics, income
levels, occupational status, and educational attainment of family members. Specific
attention was given to school enrollment of children and parental satisfaction with
local educational facilities. Simple random sampling technique has been applied in
our survey.
Data were analyzed using descriptive and inferential methods. Pie charts and bar
diagrams were constructed to illustrate the distribution of education across different
demographic and socio-economic categories. Inferential analysis was conducted
using chi-square and Fisher‘s exact tests to assess the association of gender, income,
and occupation with educational outcomes. Confidence intervals were calculated for
key proportions such as the percentage of uneducated individuals and the proportion
of households not sending children to school.

IV.I Chi- Square test for Independence of Attributes


It is applied to test the association between two attributes. Here we consider our
null hypothesis that the attributes are independent. The two attributes say A and B
are divided into m and n classes respectively and their frequencies are represented in
contingency table that has m rows and n columns. Let (Ai), i=1,2,…,m be number of
persons possessing the attribute Ai,(Bj), j=1,2,…,n be number of persons possessing
both the attribute Bj and (AiBj) be the number of persons possessing the attribute Ai
and Bj.
∑𝑚 𝑛
𝑖=1(𝐴𝑖 ) = ∑𝑗=1(𝐵𝑗 ) = 𝑁 , N is the total frequency.

Let (AiBj)0 be the expected number of persons possessing both the attribute Ai
and Bj.
(𝐴𝑖 )(𝐵𝑗 )
(𝐴𝑖 𝐵𝑗 )0 = 𝑁
,i=1,2,…,m, j=1,2,…,n
𝟐
((𝑨𝒊 𝑩𝒋 )−(𝑨𝒊 𝑩𝒋 ) *
𝟎
The test statistic is given by ꭓ = ∑𝒊,𝒋
2
(𝑨𝒊 𝑩𝒋 )
which is Chi-square variate
𝟎
with (m-1)×(n-1) degrees of freedom.

[Link] Fisher‘s Exact Test


Fisher‘s Exact Test is used to test the association between two attributes. Here
we consider our null hypothesis that the attributes are independent. It is generally
applied when the sample size is small, but it is valid for all the sample sizes. Chi-
square test for independence of attributes is invalid if the cell frequencies are less
than 5. In this type of case, we may apply Fisher‘s Exact Test which was proposed

77
by Ronald Fisher . It is useful when the attributes are classified into two different
ways. It is applied for 2×2 contingency table.
A b a+b

C d c+d

a+c b+d n=a+b+c+d

(𝑎+𝑏)!(𝑐+𝑑)!(𝑎+𝑐)!(𝑏+𝑑)!
p-value of this test is given by 𝑃= 𝑎!𝑏!𝑐!𝑑!𝑛!

[Link] 95% CONFIDENCE INTERVAL


The 95% confidence interval of a population ‗P‘ is given by:
𝑝𝑞
𝑝 ± 𝑍𝛼 √
𝑛
where Zα is the tabulated value for a α % level of significance. Also, p is the
unbiased estimator of the population proportion P.
𝑇𝑕𝑒 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑕𝑜𝑢𝑠𝑒𝑕𝑜𝑙𝑑𝑠 𝑝𝑜𝑠𝑠𝑒𝑠𝑠𝑖𝑛𝑔 𝑡𝑕𝑒 𝑎𝑡𝑡𝑟𝑖𝑏𝑢𝑡𝑒
p=sample proportion = 𝑆𝑖𝑧𝑒 𝑜𝑓 𝑡𝑕𝑒 𝑠𝑎𝑚𝑝𝑙𝑒
and q=1-p

V. RESULTS AND DISCUSSION


The study found that 15.24 percent of the total population in Rani Khamar
village were uneducated. The rest of the population had attained varying levels of
schooling, ranging from primary to higher education.

Educational Level (Percentage wise)


Uneducated
4.88
8.54 [VALUE] Below class 10

14.63 Studying or studied


HSLC
Studied upto HS
56.71

Fig. 1: A pie-chart representation of education status of the


people in the village area

78
The relationship between gender and educational status was examined using the
Chi-square test of independence. The hypotheses for the test were formulated as
follows:
Null Hypothesis (H₀ ):
Gender and educational status are independent of each other.
Alternative Hypothesis (H₁ ):
Gender and educational status are associated with each other.
Table I presents the cross-tabulation of educational status by gender with row
percentages, while Table II reports the Chi-square test results. As shown in Table I,
the distribution of education across categories appears relatively balanced between
male and female respondents.
The Chi-square test (Table II) yielded a value of 3.8134 with 3 degrees of
freedom and a corresponding p-value of 0.2823. Since the p-value is greater than the
conventional 0.05 level of significance, the null hypothesis cannot be rejected. This
indicates that there is no statistically significant association between gender and
educational status in the surveyed population. In other words, male and female
respondents had comparable educational opportunities in Rani Khamar village,
suggesting that gender was not a determining factor in educational attainment.
Table I: Cross-Tabulation of Educational Status by Gender
Gender
Female Male Total
Education Level Uneducated 17 (68.0%) 8 (32.0%) 25 (100%)
Below class 10 55 (59.1%) 38 (40.9%) 93 (100%)
Studying/Studied 11 (45.8%) 13 (54.2%) 24 (100%)
HSLC
HS and Above 10 (45.5%) 12 (54.5%) 22 (100%)
Total 93 71 164

Table II: Chi-Square Test of Association between Gender and Educational Status
Value df p-value
Pearson Chi-Square 3.8134 3 0.2823
N of valid Cases 164

To examine whether household head‘s education is associated with the monthly


family income, a statistical test of independence was conducted. Since some cell
frequencies were small, Fisher‘s exact test was used instead of the conventional Chi-
square test. For this analysis, education was categorized into two groups: Primary
education, which includes those who are uneducated or studied up to Class 10, and
Secondary education, which includes those who have studied HSLC, Higher

79
Secondary, or Graduate and above. This grouping ensured sufficient cell frequencies
for valid statistical testing.
The hypotheses were set as follows:
Null Hypothesis (H₀ ): Household head‘s educational level and monthly
family income are independent of each other.
Alternative Hypothesis (H₁ ): Household head‘s educational level and
monthly family income are not independent
Table III: Result of Fisher‘s exact test for Testing Association between Household
Head‘s Education and Monthly Family Income
Monthly family income (in Rs.)
Less than 20,000 Total p-value of
20,000 and above Fisher‘s exact test
Household head Primary 21 (100.0%) 0 (0.0%) 21 (100%)
educational Secondary 5 (62.5%) 3 (37.5%) 8 (100%) 0.0153
level
Total 27 2 33

The results indicate a p-value of 0.0153, which is well below the 5%


significance level. Therefore, the null hypothesis is rejected, and it can be concluded
that there is a statistically significant association between household head‘s
education and monthly family income. Descriptively, Table III shows that
households with higher income levels are more likely to have heads with higher
education, while lower-income households are concentrated in the ―Primary
education‖ category. This finding strongly supports the view that education and
income reinforce each other, with higher education leading to better earning capacity
and, conversely, higher income enabling access to greater educational opportunities.
To examine whether occupational status is associated with educational attainment, a
Chi-square test was performed on the cross-tabulated data of education levels and
occupation categories. For this analysis, education was grouped into two levels:
Primary education, which includes those who are uneducated or studied up to Class
10, and Secondary education, which includes those who have studied HSLC, Higher
Secondary, or Graduate and above. Similarly, occupation was classified into two
broad categories: the Primary sector, comprising farmers, cultivators, and manual
laborers, and the Secondary sector, which includes service holders, businesspersons,
teachers, shopkeepers, and other non-agricultural occupations. This categorization
provided a clear distinction between traditional agricultural/manual occupations and
non-agricultural/service-oriented occupations, making the analysis statistically
robust and easier to interpret.
The hypotheses were set as follows:
Null hypothesis (H₀ ): Occupational status is independent of education level.
80
Alternative hypothesis(H₁ ): Occupational status depends on education level.

Table IV: Result of Fisher‘s exact test for Testing Association between Respondents
Education and Occupation
Occupation
Primary Secondary Total p-value of
Sector Sector chi-sq test
Education Primary 85(72.0%) 33 (27.9%) 118(100%)
Secondary 10(21.7%) 36 (78.2%) 46 (100%) 0.0131
Total 95 58 164

The results indicate a p-value of 0.0131, which is below the conventional 5%


significance level at 1 df. This means that the null hypothesis can be rejected at the
5% level; hence, there is a statistically significant association between education
level and occupational status at this level of significance.
Descriptively, the results show that individuals with higher education are more
likely to be employed in service or business occupations (Secondary sector), while
those with primary education predominantly work in the primary sector. The
observed pattern aligns with expectations in practice.
The survey revealed that 91.89 percent of households sent their children to
school, while 8.11 percent did not. The reasons cited for non-enrollment included
financial constraints and lack of parental awareness.
Another important aspect of the study was estimating the proportion of
households that send their children to school. Out of the 37 surveyed households, 34
reported that their children attended school. This yields a sample proportion of
34
𝑝̂ = = 0.928
37
with q= 1-p= 0.072
To construct a 95% confidence interval for the true population proportion P, the
following formula was applied:
𝑝̂ (1 − 𝑝̂ )
𝑝̂ ± 𝑍𝛼⁄2 √
𝑛
Substituting the values:
0.928(0.072)
0.928 ± 1.96√
37
Gives the confidence interval: (0.9143, 0.9417)
Thus, with 95% confidence, it can be concluded that between 91.43% and
94.17% of the children in Rani Khamar attend school. This reinforces the
observation that schooling is nearly universal in the village, though a small
proportion of children remain outside the system.
81
During the COVID-19 pandemic, one of the survey questions focused on
whether children in the households of Rani Khamar attended online classes. Out of
the 34 surveyed households, 17 reported that their children had participated in online
classes. To estimate the true population proportion of children attending online
classes, a 95% confidence interval for the unknown population proportion P was
constructed and we obtained the confidence interval as (0.471, 0.529).
Thus, we may conclude with 95% confidence that between 47.1% and 52.9% of
the children in Rani Khamar attended online classes during the pandemic. This
finding is particularly important in highlighting digital access inequalities in rural
areas, where approximately half the children could not participate in online learning.
Approximately 27 percent of respondents expressed dissatisfaction with the
schooling system, citing inadequate infrastructure and lack of qualified teaching
staff.
The study also examined the relationship between the educational attainment of
parents and the educational levels achieved by their wards. Fig. 2, illustrates the
comparison of parents‘ maximum educational qualification with that of their
children.

16
14 14 13
12
10 9
8
6 7
6 5
4 4
2
2 1 1
0
Uneducated Studied below Studied upto Studied upto Graduate and
class 10 class 10 class 12 above

Parent's max education Ward's max education

Fig. 2: Parent–Ward Educational Status in Rani Khamar Village

The results indicate that, even among households where parents were
uneducated, a significant proportion still encouraged their children to pursue
schooling. Only one parent in the surveyed households had attained higher
educational qualifications, yet aspirations for higher education among children were
visible across the community. For example, 14 wards of parents who studied only up
to below Class 10 continued their education further, while 13 wards of parents with
Class 10 education achieved higher schooling themselves.
This suggests that parents with relatively lower educational attainment
nonetheless recognize the importance of education and actively support their wards
82
in pursuing higher levels of learning. The finding underlines the aspirational value of
education in rural households, even in the absence of direct parental experience with
higher education.
The survey also captured the aspirations of parents regarding the level of
education they wished their children to pursue in the future. Table 7 summarizes
these responses, while Fig. 3 provides a visual representation.
Table V: Future Education Desired for Children in Rani Khamar Village
Education Level Frequency Percent
Up to Primary 1 2.7
Up to Secondary 2 5.4
Up to Higher Secondary 3 8.1
Up to College Education 19 51.4
As far as Economically Possible 12 32.4
Total 37 100

19
20
15
Frequency

12
10
5 2 3
1
0
Upto Primary Upto Secondary Upto Higher Upto College As far as
Secondary Education Economically
Possible
Desired Education Level

Fig. 3: Future Educational Aspirations of Parents in Rani Khamar Village

As shown in Table V, the majority of parents (51.4%) expressed the desire for
their children to attain at least a college-level education. A smaller proportion (8.1%)
wished their children to study up to higher secondary, while only 2.7% and 5.4%
preferred education up to primary and secondary levels, respectively. A significant
segment of households (32.4%) indicated that the educational future of their children
would depend on their economic capacity. These findings suggest a strong demand
for higher education in the village, though financial limitations remain a decisive
factor. This highlights the importance of scholarships, fee waivers, and other
financial aid schemes in ensuring that parental aspirations for higher education can
be fulfilled.
83
VI. CONCLUSION
The study of Rani Khamar village reveals the deeply interwoven relationship
between education and socio-economic conditions in rural Assam. While gender
does not emerge as a significant determinant of educational outcomes, income and
occupation clearly influence access and attainment. Households with higher income
and service-oriented occupations are more likely to invest in education, while those
constrained by poverty remain disproportionately uneducated. Nevertheless, even
among less educated parents, aspirations for their children‘s higher education were
evident, reflecting the growing recognition of education as a pathway to mobility
and empowerment. The findings also underscore systemic barriers: inadequate
school infrastructure, digital divides during the pandemic, and financial constraints
continue to hinder equitable access. Overall, the study highlights both progress and
persistent gaps. Achieving inclusive rural education will require sustained policy
interventions, improved infrastructure, financial support mechanisms, and
community-level awareness to ensure that aspirations translate into tangible
educational opportunities for all households in villages like Rani Khamar.

REFERENCES
1) Gupta, S.C. and Kapoor V.K. (2021). Fundamental of Applied Statistics. Sultan
Chand & Sons; 4th Edition; New Delhi
2) Gupta, S.C. and Kapoor V.K.(2021). Fundamental of Mathematical Statistics.
Sultan Chand & Sons; 12th Edition; New Delhi
3) Bhagwati, L., & Sarma, D. (2024). Relative Educational Status Of Assam: A
District Level Study. Educational Administration: Theory and Practice, 30(5),
11063-11072.
4) Phukan, R. S., & Gogoi, K. P. (2013). Education of Students in the Rural Areas
of Assam. Social Science Journal of Gargaon College, (49-59).
5) Tilak, J. B. G. (2002). Education and poverty. Journal of Human Development,
3(2), 191–207.
6) Sarma, M., & Das, K. (2021). A study on inequalities in education created due
to rural-urban divide and class position with special reference to
Assam. Turkish Journal of Computer and Mathematics Education, 12(10),
6317-6321.
7) Dutta, S., Borah, B., Bora, L., Deka, R. J., Deka, S. D., Phangchopi, D., ... &
Shyam, J. (2025). Development and Validation of the Assam Rural Livelihood
and Farming Scale (ARLFS) for the Socio-Economic Assessment of Khuti
System Dwellers, Assam, India. Journal of Scientific Research and
Reports, 31(8), 737-748.

84
A Brief Overview on Construction of Discrete Analogues
of Continuous Probability Distribution
Diksha Das1 and Bhanita Das2
*1
Research Scholar, Department of Statistics, North Eastern Hill University,
Shillong, Meghalaya. Email id: 0210dikshadas@[Link]
2
Assistant Professor, Department of Statistics, North Eastern Hill University,
Shillong, Meghalaya. Email id: bhanitadas83@[Link]

ABSTRACT:
In our daily lives, we often encounter situations where variables are
fundamentally continuous. However, it can be either impossible or impractical to
obtain samples from a continuous distribution, resulting in observations being
recorded in a discrete form. This can occur due to the inherent characteristics of the
variable or the limitations of our measuring tools. Additionally, in some cases, it
may be more appropriate to use a discrete model to analyze observations of a
continuous variable. This highlights the importance of developing new discrete
distributions that offer greater efficiency and flexibility. While many discrete
distributions currently exist, classical ones do not always effectively model various
count data sets. Consequently, there is an increasing need to create discrete analogs
of existing continuous distributions. This process of converting continuous
distributions into their discrete counterparts is known as discretization. The
fundamental principle behind developing new discretized distributions is the
preservation of one or more characteristic properties of the continuous distribution.
In this article, an attempt is made to illuminate this potent concept, with special
emphasis on the two most common methods of discretization: the infinite series
approach and the survival function approach.
KEYWORDS: Discretization, discrete analogue, infinite series approach, survival
function, Marshall Olkin scheme.

1. INTRODUCTION
In many real-life practices, we may come across certain situations where sample
observations come from a continuous distribution but are not measured on a
continuous scale. Instead, these measurements may be recorded to a finite number of
decimal places, which can occur either due to the inherent nature of the data or
limitations of the measuring tools. Collecting samples from a continuous distribution
can often be challenging or inconvenient. Typically, observed values are measured to
only a limited number of decimal places, resulting in them exhibiting discrete


Corresponding author: 0210dikshadas@[Link]

85
characteristics, which means they do not truly represent all points on a continuum. In
some cases, continuous variables are measured using frequencies of non-overlapping
class intervals. This situation may arise from the accuracy of the measuring device
or a need to conserve space. The class intervals are organized to ensure that their
union covers the entire range of random variables. Under these circumstances, it is
generally more appropriate to analyze the sample observations using a discrete
model rather than a continuous one.
In reliability analysis, particularly in survival analysis, situations may arise
where the reliability function (or survival function) is treated as a function of a
discrete random variable, despite being fundamentally a function of a continuous
time random variable. For example, the reliability of a switching device is evaluated
based on how many times the switch is operated. The reliability of an airplane tire is
assessed based on the number of landings it has experienced. The lifespan of an
electric circuit is determined by how many breakdowns occur within a month.
Additionally, the shelf life of a particular drug is measured by the number of days
until it is no longer effective.
In survival analysis, the length of a patient's stay in an observation ward is
typically measured in days. For instance, the survival time of a patient with a brain
hemorrhage is tracked based on the duration of their observation, while the survival
time for leukemia patients is often measured in weeks. In all of these cases, lifetimes
are not measured on a continuous scale; instead, they are counted, making them
discrete random variables. These examples illustrate that while lifetimes are often
considered continuous, they are frequently treated as discrete random variables due
to the method of counting used for measurement.
The primary reasons for discretizing continuous distributions are twofold: first,
the discrete version of a continuous distribution offers a probability mass function
that can compete with traditional discrete distributions commonly used in count data
analysis; second, a discretized distribution removes the need for a continuous
distribution when dealing with strictly discrete data. As noted by Lai (2013),
discretizing a continuous lifetime model provides an interesting and intuitively
appealing method for deriving a discrete counterpart to the corresponding
continuous lifetime model.
The fundamental principle in developing new discrete models is to preserve one
or more characteristic properties of the existing continuous lifetime models. This
means that the characteristics of the continuous distribution should remain intact in
the corresponding discretized distribution. These properties may include the
probability density function, moment-generating function, survival function,
moments, and hazard rate function, among others. The range of the discretized
model is determined based on either the full range or a subset of the continuous
range.
The concept of discretization has led many statisticians to develop new discrete
models, known as discretized models, which adapt classical continuous distributions
to better represent discrete failure times and reliability data. A systematic review of
86
discretization methods was conducted by Bracquemond and Gaudoin (2003),
highlighting various techniques for generating discrete versions of continuous
lifetime distributions. Lai (2013) derived several discrete lifetime distributions from
continuous ones and discussed the challenges involved in this derivation.
Additionally, Chakraborty (2015a) provided a comprehensive survey of different
methods of discretization, detailing the various approaches found in the literature for
constructing discrete versions of continuous models. His study also illustrates
various discretized models related to each method, along with their applications.
In recent decades, many researchers have proposed various discrete distributions
by utilizing different discretization methods reported in the statistical literature.
According to Chakraborty (2015a), nine distinct discretization techniques have been
identified. While several techniques exist for creating novel discretized distributions,
two have garnered the most interest from researchers: the infinite series approach
and the survival function approach for discretization.
The infinite series approach maintains the structure of the probability density
function of the original continuous distribution, with the support of the discretized
variable being determined from either the full range or a subset of the corresponding
continuous range. In contrast, the survival function approach preserves the form of
the survival function of the continuous distribution, with the support of the discrete
analogue determined from the entire range of the corresponding continuous
distribution.
Numerous discrete distributions have been developed in the literature using two
primary methods of discretization. In recent decades, both methods have attracted
significant attention from researchers. This interest has motivated us to highlight
these two effective discretization techniques and conduct a brief study on them. In
this article, we will focus on the two most commonly used methods for developing
new discrete distributions: the infinite series approach and the survival function
approach.
The remainder of the article is structured as follows: Section 2 explores the
infinite series method for discretization, including examples of discrete distributions
using this approach. Section 3 discusses the survival function approach to
discretization, highlighting the Marshall-Olkin scheme and recent advancements in
discretized distributions through this method. Finally, Section 4 presents the
conclusions of the study. In this paper, the continuous random variable to be
discretized is denoted by X, while its discrete analogue is denoted by Y.

2. METHODOLOGY I: INFINITE SERIES APPROACH OF DISCRETIZATION


In this method, the probability density function (pdf) of a continuous distribution
is used to construct a discretized distribution. The denominator incorporates a
summation of the pdf represented as an infinite series. Therefore, this method is
referred to as the infinite series approach to discretization.

87
Let 𝑓𝑋 (𝑥) be the probability density function of the continuous random variable
𝑋, with its variable support on −∞ < 𝑥 < ∞. Then the form of the probability mass
function (pmf) of the discretized random variable Y is given by
𝑓𝑋 (𝑘)
𝑃(𝑌 = 𝑘) = ∑∞ ; 𝑘 = 0, ±1, ±2, … (1)
𝑗=−∞ 𝑓𝑋 (𝑗)

However, if the continuous random variable 𝑋 is defined on 0 < 𝑥 < ∞ then the
pmf of the discrete random variable 𝑌 is defined as
𝑓 (𝑘)
𝑃(𝑌 = 𝑘) = ∑∞ 𝑋 ; 𝑘 = 0,1,2, … (2)
𝑗=𝑘 𝑓𝑋 (𝑗)

This methodology was first time been used for the derivation of Good
distribution as proposed by Good (1953) and was applied to model the population
frequencies of species and the estimation of population parameters. An extensive
study of this distribution was later conducted by Kulesekara and Tonkyn (1992) and
Doray and Luong (1997). The discrete form of general Dirichlet series was
presented by Siromoney (1964) and was used to model frequency distribution of the
length of wet spells during the period 1932-1962 in a place called Tambaram in
southern India.
Recently, Nekoukhou et al. (2012) proposed the well-known discrete generalized
exponential distribution, Lekshmi and Sebastian (2014) developed Generalized
Discrete Laplace distribution, that generalizes the discrete skew Laplace distribution
as proposed by Kozubowski and Inusah (2006). The discrete lognormal distribution
was studied by Pellegrin et al. (2015) as a model for pore size distributions of
hypochlorite aged PES/PVP ultrafiltration membranes. The discrete Beta-
Exponential distribution was proposed by Nekoukhou et al. (2015), the two-
parameter discrete Lindley has been developed by Hussain et al. (2016) and Josmar
et al. (2017) developed the discrete Shanker distribution and many more discretized
distributions are present in the literature that have been constructed using the infinite
series approach of discretization.

2.1. Illustrations on discretized distributions using Methodology I


Some of the discretized distributions those have been developed using the
infinite series method of discretization are discussed below:

2.1.1. Discrete normal distribution


For a continuous random variable 𝑋 following 𝑁(𝜇, 𝜍) with its pdf denoted by

1 (𝑥 − 𝜇)2
𝑓𝑋 (𝑥; 𝜇, 𝜍) = exp 6 7 ; −∞ < 𝑥 < ∞
𝜍√2𝜋 2𝜍 2

88
where the parameters −∞ < 𝜇 < ∞ , 0 < 𝜍 < ∞ .The pmf of the corresponding
𝑚 𝑚
discrete Normal distribution after re-parametrization 𝑒 (1−2𝜇)⁄2𝜍 = 𝜆 and 𝑒 −1⁄𝜍 =
𝑞 is given by
𝜆𝑘 𝑞𝑘(𝑘−𝑙)⁄𝑚
𝑃(𝑌 = 𝑘) = ∑∞ 𝑗 𝑗(𝑗−𝑙)⁄𝑚 ; 𝑘 = 0, ±1, ±2, … (3)
𝑗=−∞ 𝜆 𝑞

where the parameters 𝜆 > 0, 0 < 𝑞 < 1.


Many researchers, viz., Lisman and Zuylen (1972), Kemp (1997), Liang (1999)
and Szablowski (2001) have discussed this form of discrete Normal distribution in
application to diverse fields.
Dasgupta (1993) also proposed discrete Normal distribution by consider the re-
parametrization 𝜆 = 𝑞 1⁄2 𝑎𝑛𝑑 𝑞 = 𝑒 −2𝛽 in Eq. (3), such that the pmf is of the form
𝑞𝑘⁄𝑚 𝑒 −𝛽𝑘(𝑘−𝑙)
𝑃(𝑌 = 𝑘) = ∑∞ 𝑗⁄𝑚 𝑒 −𝛽𝑗(𝑗−𝑙) ; 𝑘 = 0, ±1, ±2, … (4)
𝑗=−∞ 𝑞

2.1.2. Discrete exponential distribution


The pdf of 𝑋 following the continuous exponential distribution is defined as
𝑓𝑋 (𝑥; 𝜆) = 𝜆𝑒 −𝜆𝑥 ; 𝑥 > 0, 𝜆 > 0
The discrete exponential distribution is proposed by Sato et al. (1999), using this
technique in application to model defect count distribution in semiconductor
deposition equipment. The pmf of their discrete exponential distribution was defined
as
𝑃(𝑌 = 𝑘) = (1 − 𝑒 −𝜆 )𝑒 −𝜆𝑘 ; 𝑘 = 0,1,2, … ; 𝜆 > 0 (5)
which takes the form of the pmf of geometric distribution with parameter
𝑝 = 𝑒 −𝜆 .
2.1.3. Discrete Laplace (double exponential) distribution
The pdf of 𝑋 following the classical Laplace distribution with scale parameter
𝜍 > 0, is given by
𝑓𝑋 (𝑥; 𝜍) = (2𝜍)−1 𝑒 −|𝑥|/𝜍 ; −∞ < 𝑥 < ∞
The discrete Laplace distribution is proposed by Inusah and Kozubowski (2006)
and applied this distribution in modelling different currency exchange rate data. The
pmf of the discrete Laplace distribution with the parameter 𝑝 = exp(− 1⁄𝜍) ; 0 <
𝑝 < 1 , is given by
1−𝑝
𝑃(𝑌 = 𝑘; 𝑝) = 1+𝑝 𝑝|𝑘| ; 𝑘 = 0, ±1, ±2, … (6)
where ⌊ . ⌋ is the greatest integer function.
This distribution inherits many characteristics of its continuous counterpart
namely unimodality, infinite divisibility, maximum entropy distribution for given
absolute moment. Later, Andersen et al. (2013) applied this distribution for
estimating Y-STR haplotype frequencies.
89
2.1.4. Discrete generalized exponential distribution
The pdf of 𝑋 following generalized exponential distribution, with two
parameters 𝛼 > 0 𝑎𝑛𝑑 𝜆 > 0 (as shape and scale parameters respectively), is of the
form
𝛼−1 −𝜆𝑥
𝑓(𝑥; 𝛼, 𝜆) = 𝛼𝜆(1 − 𝑒 −𝜆𝑥 ) 𝑒 ;𝑥 > 0
The discrete generalized exponential distribution is developed by Nekoukhou et
al. (2012), using this discretization method and its pmf is takes the forms as given
below
𝛼−1
𝑃(𝑌 = 𝑘; 𝛼, 𝑝) = 𝑐𝑝𝑘−1 (1 − 𝑝𝑘 ) ; 𝑘 = 1,2,3, … (7)
𝑗 𝑗
𝛼−1 (−1) 𝑝 (𝛼−1)(𝛼−2)…(𝛼−𝑗)
where 𝑐 −1
= ∑∞
𝑗=0 . 𝑗 / 1−𝑝𝛽+𝑗 𝑎𝑛𝑑 .𝛼−1
𝑗
/ =
𝑗!
and the parameters 𝛼 > 0, 0 < 𝑝 = exp(−𝜆) < 1 , ∀ 𝜆 > 0.
They applied this distribution to model rank frequencies of graphemes in a
Slavic language called ―Slovene‖.

2.1.5. Two parameter discrete Lindley distribution


To increase the flexibility of the one parameter discrete Lindley distribution,
Hussain et al. (2016) developed a two-parameter discrete Lindley distribution by
introducing a continuous two-parameter Lindley distribution, which is the mixture
distribution of 𝐺𝑎𝑚𝑚𝑎(1, 𝜃) and 𝐺𝑎𝑚𝑚𝑎(2, 𝜃) with mixing probabilities 𝑝1 =
𝜃⁄𝜃 + 𝛽 𝑎𝑛𝑑 𝑝2 = 𝛽⁄𝜃 + 𝛽 . And its pdf is given by
𝑓𝑋 (𝑥; 𝜃, 𝛽) = ,𝜃 2⁄(𝜃 + 𝛽)-(1 + 𝛽𝑥),𝑒𝑥𝑝(−𝜃𝑥)-; 𝑥 ≥ 0
where the parameters 𝛽 ≥ 0 𝑎𝑛𝑑 𝜃 > 0 .
The two-parameter discrete Lindley (TDL) distribution was developed using the
infinite series method of discretization and its pmf is given
(1−𝑝)𝑚 (1+𝛽𝑘)𝑝𝑘
𝑃(𝑌 = 𝑘; 𝑝, 𝛽) = ,1+𝑝(𝛽−1)-
; 𝑘 = 0,1,2, … (8)
where the parameters 0 < 𝑝 = exp(−𝜃) < 1 , ∀ 𝜃 > 0 and 𝛽 > 0.
They applied this distribution to two real life datasets related to the number of
European red mites on apple leaves and number of strikes in UK coal mining
industries in four successive week periods during 1948-1959.

2.1.6. Discrete Quasi-Xgamma distribution


The pdf of the random variable 𝑋 following two-parameter Quasi Xgamma
distribution is given by
𝜃 𝜃𝑚
𝑓𝑋 (𝑥; 𝛼, 𝜃) =
𝛼+1
0𝛼 + 2 𝑥 2 1 𝑒 −𝜃𝑥 ;𝑥 > 0
where the parameters (𝛼, 𝜃) > 0.
Two types of discrete Quasi Xgamma distribution were proposed by Josmar et
al. (2020), based on infinite series discretization method and survival function

90
approach, termed as Type 1 discrete Quasi Xgamma (DQX1) distribution and Type 2
discrete Quasi Xgamma (DQX2) distribution respectively. The pmf of the DQX1
distribution with parameters (𝛼, 𝜃) > 0, is given by
𝑕(−𝜃𝑘 ,0) 𝜃𝑚
𝑃(𝑌 = 𝑘; 𝛼, 𝜃) = 𝑞(𝛼,𝜃)
0𝛼 + 2
𝑘 2 1 ; 𝑘 = 0,1,2, … (9)

where the parameters 𝛼 > 0, 𝜃 > 0, 𝑕(𝑎, 𝑏) = 𝑒 𝑎 − 𝑏 and


1 𝑕(𝜃,0)[2𝛼.𝑕(𝜃,1)𝑚 + 𝜃𝑚 𝑕(𝜃,−1)]
𝑞(𝑎, 𝑏) = 2 3 ; (𝑎, 𝑏) > 0.
2 𝑕(𝜃,1)𝑛
The usefulness of this model has been exhibited by using two real datasets
related to the number of corn borers and number of outbreaks of strikes in the UK
coal mining industries in four successive week periods during 1948-1959.

3. METHODOLOGY II:
SURVIVAL FUNCTION APPROACH OF DISCRETIZATION
Survival function approach of discretization uses the survival function of the
continuous distribution and this technique preserves the form of the survival
function of the continuous distribution, while constructing the corresponding
discrete analogue. This method of discretization is the most used due to the ease in
computing the form of probability mass function and deriving various structural
properties.
Let us assume a continuous random variable 𝑋 with the survival function 𝑆𝑋 (𝑘),
then the form of the probability mass function for the discrete random variable 𝑌,
defined as 𝑌 = ⌊𝑋⌋ = largest integer less than or equal to 𝑋, is given by
𝑃,𝑌 = 𝑘- = 𝑃,𝑘 < 𝑋 < 𝑘 + 1-
= 𝑃,𝑋 > 𝑘- − 𝑃,𝑋 ≥ 𝑘 + 1-
= 𝑆𝑋 (𝑘) − 𝑆𝑋 (𝑘 + 1); 𝑘 = 0,1,2, ⋯ (10)
Under this method the form of the survival function (sf) is preserved on its
integer part, that is 𝑆𝑋 (𝑘) = 𝑆𝑌 (𝑘), where 𝑘 is an integer. So given any continuous
distribution it is possible to generate the discrete analogue of the corresponding
distribution using the form as in Eq. (10). Nakagawa and Osaki (1975) were the first
to use this method to develop the discrete version of the well-known continuous
Weibull distribution, called as discrete Weibull distribution.
3.1. Marshall-Olkin scheme followed by survival function approach of
discretization
Marshall and Olkin (1997) discussed a general method of generating a new
family of distributions, called as Marshall-Olkin family of distribution, by adding a
parameter in the existing distribution. Starting with the survival function 𝑆𝑋 (𝑥) of
the existing distribution, the survival function of the new Marshall-Olkin family is
given by
𝛼 𝑆 (𝑥)
𝕊(𝑥; 𝛼) = ,1−𝛼̅ 𝑋𝑆 (𝑥)- ; 𝛼>0 (11)
𝑋
where 𝛼̅ = 1 − 𝛼.
91
Then using the survival function approach of discretization, the probability mass
function of the discrete analogue of the new continuous Marshall-Olkin family of
distribution is constructed and is given by
𝑃(𝑌 = 𝑘; 𝛼) = 𝕊(𝑘; 𝛼) − 𝕊(𝑘 + 1 ; 𝛼)
𝛼*𝑆𝑋 (𝑘)−𝑆𝑋 (𝑘+1)+
= *1−𝛼̅ ̅ 𝑆𝑋 (𝑘+1)+
𝑆𝑋 (𝑘)+*1−𝛼
;𝑘 = 0 ,1, 2, … (12)

A new generalization of the geometric distribution by using this scheme of


discretization was discussed by Gómez (2010). Recently many researchers are
developing new discrete distributions using this scheme. The resultant new
distribution known as the Marshall–Olkin (MO) extended distribution is more
flexible than the original distribution and possesses the original distribution as a
unique feature.
Though several methods are available for discretization, even then the survival
function approach of discretization is the mostly used. The popularity of this method
is due to the compact form of the survival function and ease to conduct the
computations. Over the past few decades many researchers have constructed various
discretized distributions using this method of discretization. Recently, many well-
known continuous distributions have been studied using this approach of
discretization. Notable among them are discrete Burr and discrete Pareto
distributions were proposed by Krishna and Pundir (2009), discrete inverse Weibull
by Jazi et al. (2010), discrete Gamma by Chakraborty and Chakravarty (2012a),
discrete Logistic distribution by Chakraborty and Chakravarty (2012b), discrete
Generalized Gamma developed by Chakraborty (2015b), and discrete generalized
inverse Weibull distribution by Para and Jan (2019).
Most recently, Almetwally et al. (2020) developed the discrete Marshall-Olkin
Generalized Exponential distribution, Chakraborty et al. (2021) proposed the
discrete Gumbel distribution suitable for different skewed data sets, Ibrahim and
Almetwally (2021) developed discrete Marshall-Olkin Lomax distribution,
Almetwally et al. (2022) proposed discrete Marshall–Olkin inverted Topp–Leone
distribution in application to COVID-19 data sets in different countries. Most
recently, Das and Das (2023) introduced discrete version of Fréchet–Weibull
distribution in application to count data from several fields, suitable to datasets with
over-dispersion under-dispersion characteristics. Das et. al. (2023) developed
Discretized Generalized Gompertz distribution suitable for increasing, constant and
bathtub shaped hazard rate function. Here we shall discuss a few recent
developments of discretized distributions using the survival function approach of
discretization.

3.2. Illustrations on discretized distributions using Methodology II


Some of the discretized distributions those have been developed using the
survival function approach of discretization are discussed below:
92
3.2.1. Discretized Marshall-Olkin Generalized Exponential distribution
Almetwally et al. (2020) developed the discrete version of Marshall-Olkin
Generalized Exponential distribution with its survival function given by
𝛼
𝜆01 − (1 − 𝑒 −𝜃𝑥 ) 1
𝑆𝑋 (𝑥) = ; 𝑥>0
𝜆 + (1 − 𝜆)(1 − 𝑒 −𝜃𝑥 )𝛼
where the parameters (𝛼, 𝜆, 𝜃) > 0.
They constructed the probability mass function of discrete Marshall-Olkin
Generalized Exponential distribution using the survival function approach of
discretization, given by
𝛼 𝛼
𝜆01−(1−𝑒 −𝜃𝑘 ) 1 𝜆01−(1−𝑒 −𝜃(𝑘+𝑙) ) 1
𝑃,𝑌 = 𝑘- = ̅ (1−𝑒 −𝜃𝑘 )𝛼
− ̅ (1−𝑒 −𝜃(𝑘+𝑙) )𝛼
; 𝑘 = 0,1,2, … (13)
𝜆+𝜆 𝜆+𝜆

where the parameters (𝛼, 𝜆, 𝜃) > 0 and applied this to model the daily new cases
of COVID-19 in the case of Egypt.

3.2.2. Discretized Marshall-Olkin Weibull distribution


Opone et al. (2021) proposed the three-parameter discrete Marshall-Olkin
Weibull distribution in application to overdispersed and under dispersed datasets.
They first developed the continuous Marshall-Olkin Weibull distribution with is
survival function given by
𝛽
𝛼𝑒 −𝜃𝑥
𝑆𝑋 (𝑥) = 𝛽
; 𝑥 > 0, (𝛼, 𝜃, 𝛽) > 0
1 − 𝛼̅𝑒 −𝜃𝑥
Then the pmf of the discrete Marshall-Olkin Weibull distribution was
constructed using the survival function approach of discretization given by
𝛽 𝛽
𝛼2𝛾 𝑘 −𝛾(𝑘+𝑙) 3
𝑃,𝑌 = 𝑘- = 𝛽 𝛽 ; 𝑘 = 0,1,2, … (14)
̅ 𝛾(𝑘+𝑙) 3
̅ 𝛾𝑘 321−𝛼
21−𝛼

where the parameters 𝛼 > 0, 𝛽 > 0,0 < 𝛾 = 𝑒𝑥𝑝(−𝜃) < 1.


The pmf of this distribution can take various forms such as decreasing, left-
skewed unimodal, right-skewed unimodal and symmetric, which make the
distribution applicable to versatile situations.

3.2.3. Discretized Gumbel distribution


If 𝑋 has Gumbel (Type I) distribution which is one of the most referred extreme
value distributions, then its survival function for the parameters (𝜇, 𝜍) is given as
𝑆𝑋 (𝑥) = 1 − 𝑒𝑥𝑝[−𝑒 −(𝑥−𝜇)⁄𝜍 ] ; 0<𝑥<∞

93
Then the discrete analogue of this distribution was developed by Chakraborty et
al. (2021) using this method of discretization. Then the pmf of discrete Gumbel
distribution is given by
𝑘+𝑙 𝑘+𝑙
𝑃,𝑌 = 𝑘- = 𝑒 −𝛼𝑝 − 𝑒 −𝛼𝑝 ; 𝑘 = 0,1,2, … (15)
where the parameters 𝑝 = 𝑒 −1⁄𝜍 < 1 and 𝛼 = 𝑝−𝜇 > 0.
Depending upon the choice of the parameters this distribution can be positively
skewed or negatively skewed and possess long-tail.

3.2.4. Discretized Marshall-Olkin inverted Topp–Leone distribution


Almetwally et al. (2022) introduced the Marshall–Olkin inverted Topp–Leone
distribution first and then extended this new continuous distribution to propose its
discrete analogue, called as discrete Marshall–Olkin inverted Topp–Leone
(DMOITL) distribution. If 𝑋 has the MOITL distribution, then its survival function
is given by
𝛼(1 + 2𝑥)𝜗
𝑆𝑋 (𝑥) = ;0 < 𝑥 < ∞
(1 + 𝑥)2𝜗 − 𝛼̅(1 + 2𝑥)𝜗
where the parameters (𝛼, 𝜗) > 0. And the probability mass function of the
corresponding DMOITL distribution is given by
𝛼(1+2𝑘)𝜗 𝛼(3+2𝑘)𝜗
𝑃,𝑌 = 𝑘- = (1+𝑘)𝑚𝜗 ̅ (1+2𝑘)𝜗
− (2+𝑘)𝑚𝜗 ̅ (3+2𝑘)𝜗
; 𝑘 = 0,1,2, … (16)
−𝛼 −𝛼

They modeled this distribution to data sets comprising of newly reported


instances of COVID-19 in different countries, including Italy, Puerto Rico, and
Singapore.

3.2.5. Discretized Fréchet–Weibull distribution


If 𝑋 has Fréchet–Weibull distribution with parameter (𝛼, 𝛽, 𝑚, 𝛾) and its
survival function given by
𝑚 𝛼𝛾
𝑆𝑋 (𝑥) = 1 − 𝑒𝑥𝑝 {−𝛽 𝛼 . / } ; 0 < 𝑥 < ∞
𝑥
Using the survival method of discretization and substituting 0 < 𝑞 =
𝑒𝑥𝑝(−𝛽𝛼 ) < 1 and 𝜆 = 𝛼𝛾, the probability mass function of discrete version
Fréchet–Weibull distribution is given by
𝑚 𝜆 𝑚 𝜆
. / . /
𝑃,𝑌 = 𝑘- = 𝑞 𝑘+𝑙 − 𝑞 𝑘 ; 𝑘 = 0,1,2, … (17)
Das and Das (2023) introduced discretized Fréchet–Weibull distribution, which
depend upon various flexible properties such as it has increasing, decreasing and up-
side down bathtub shaped hazard rate, which are not much observed for count
distributions and it can exhibit both positively and negatively skewed forms of
probability mass function.

94
3.2.6. Discretized Generalized Gompertz distribution
The survival function of 𝑋 following the generalized Gompertz distribution with
the parameters (𝜆, 𝛼, 𝜃) is given as
𝜃
𝜆
𝑆𝑋 (𝑥) = 1 − [1 − 𝑒𝑥𝑝 {− (𝑒 𝛼𝑥 − 1)}] ; 𝑥 > 0
𝛼
The Discretized Generalized Gompertz distribution is proposed by Das et al.
(2023) which is a very versatile distribution, applicable to model with various
lifetime data having increasing, constant and bathtub shaped hazard rate functions.
Its probability mass function is given by
−𝑙 𝛼(𝑘+𝑙) 𝜃 −𝑙 𝛼𝑘 𝜃
𝑃,𝑌 = 𝑘- = 01 − 𝑞 𝛼 (𝑒 −1)
1 − 01 − 𝑞 𝛼 (𝑒 −1) 1 ; 𝑘 = 0,1,2, … (18)
where the parameters 0 < 𝑞 = 𝑒𝑥𝑝(−𝜃) < 1, 𝛼 ≥ 0 and 𝜃 > 0.

4. CONCLUSION
When a continuous lifetime random variable is measured or recorded at count
points and hence observed as discrete random variable, it is more relevant to use an
appropriate discrete distribution rather than corresponding it to a continuous model.
Holland (1975) asserts that, ―when only an approximating discrete random variable
is observable, estimation procedures employing the hypothetical continuous random
variables are sometime biased and hence a discrete distribution is more appropriate
for an observed data‖. Further, according to Roknabadi et al. (2009), there is a need
to focus on more realistic discrete life time distributions. ―Discretization of
continuous distribution may be looked upon as a filtering process which may help in
reducing of noise present in the data‖, added Chakraborty (2015a).
The process of creating new discrete random variables that are the discrete
equivalent of their corresponding continuous random variables is known as
discretization. The probability mass function of the new discrete variable can be
obtained by using the appropriate discretization technique. Under the process of
discretization, there will always be some loss of information and accuracy. However,
as a researcher, one must be careful to strike a balance between the need to discretize
and the resulting accuracy and information loss. Furthermore, the selection criteria
recommended by Bracquemond and Gaudoin (2003) for selecting the best
discretization technique must be adhered appropriately.
In recent decades, a number of researchers have been prompted to develop new
discrete distributions by the idea of discretizing a continuous distribution using
various techniques. Even if there are a lot of discretized distributions in the
literature, there is still scope to develop more that are relevant to other data sets.
Moreover, since all the methods of discretization did not receive the same attention
from the researchers, hence there are vast scope to explore those methods of
discretization also.
According to Chakraborty (2015a), future study in this area could focus on
preserving multiple characteristics of the continuous distribution while developing
the discretized version. In this paper it is aimed to put light on this emerging
95
dynamic research topic. Not every discretization technique has been covered here.
The aforementioned discretized distributions' numerous statistical characteristics and
application domains are likewise not thoroughly explained in this article. They can
be obtained, nevertheless, from the corresponding references.

REFERENCES
1. Almetwally, E. M., Hisham, M. A., and Hany, A. S. (2020). Managing risk of
spreading ―COVID-19‖ in Egypt: Modelling using a discrete Marshall-Olkin
generalized exponential distribution. International Journal of Probability and
Statistics, 9(2): 33-41.
2. Almetwally, E. M., Abdo, D. A., Hafez, E.H., Jawa, T. M., Ahmed, N. S., and
Almongy, H. M. (2022). The new discrete distribution with application to
COVID-19 Data. Results in Physics, 32(1): 1-12.
3. Andersen, M. M., Eriksen, P. S., and Morling, N. (2013). A gentle introduction
to the discrete Laplace method for estimating Y-STR haplotype frequencies.
arXiv:1304.2129v4.
4. Bracquemond, C., and Gaudoin, O. (2003). A survey on discrete life time
distributions. International Journal of Reliability, Quality and Safety
Engineering, 10(1): 69-98.
5. Chakraborty, S. (2015a). Generating Discrete Analogues of Continuous
Probability Distributions- A Survey of Methods and Constructions. Journal of
Statistical Distributions and Applications, 2(6): 1-30.
6. Chakraborty, S. (2015b). A new discrete distribution related to generalized
gamma distribution and its properties. Communications in Statistics - Theory
and Methods, 44(8): 1691–1705.
7. Chakraborty, S., and Chakravarty, D. (2012a). Discrete gamma distributions:
properties and parameter estimations. Communications in Statistics - Theory
and Methods, 41(18): 3301–3324
8. Chakraborty, S., and Chakravarty, D. (2012b). A new discrete probability
distribution with integer support on (− ∞, ∞). Communications in Statistics -
Theory and Methods, 45(2): 492–505
9. Chakraborty, S., Chakravarty, D., Josmar M., and Wesley B. (2021). A discrete
analog of Gumbel distribution: properties, parameter estimation and
application. Journal of Applied Statistics, 48(4): 712-737.
10. Das, D., and Das, B. (2023). Discretized Fréchet–Weibull Distribution:
Properties and Application. Journal of the Indian Society for Probability and
Statistics, 24(1): 1-40.
11. Das, D., Das, B., and Hazarika, P. (2023). Discretized version of Generalized
Gompertz Distribution with Application to Real life data. International Journal
of Agricultural and Statistical Sciences, 19(1): 441-456.
12. Dasgupta, R. (1993). Cauchy equation on discrete domain and some
characterizations. Theory of Probability and its Application, 38(2): 318-328.

96
13. Doray, L. G., and Luong, A. (1997). Efficient estimators for the Good family.
Communications in Statistics - Simulation and Computation, 26 (1): 1075-
1088.
14. Good, I. J. (1953). The population frequencies of species and the estimation of
population parameters. Biometrika, 40(3/4): 237–264.
15. Gómez, D. E. (2010). Another generalization of the geometric distribution.
TEST, 19(2): 399-415.
16. Holland, B. S. (1975). Some Results on the discretization of continuous
probability distributions. Technometrics, 17(3): 333-339.
17. Hussain, T., Aslam, M., and Ahmad, M. (2016). A Two Parameter Discrete
Lindley Distribution. Revista Colombiana de Estadistica, 39(1): 45-61.
18. Ibrahim, G. M., and Almetwally, E. M. (2021). Discrete Marshall-Olkin
Distribution Application of COVID-19. Journal of Scientific & Technical
Research, 32(5): 25381-25390.
19. Inusah, S., and Kozubowski, T. J. (2006). A discrete analogue of the Laplace
distribution. Journal of Statistical Planning and Inference, 136(3): 1090-1102.
20. Jazi, M. A., and Lai, C. D., Alamatsaz, M. H. (2010). A discrete inverse
Weibull distribution and estimation of its parameters. Statistical Methodology,
7(2): 121-132.
21. Josmar, M., Wesley, B. D. S., and Ricardo, P. O. (2017). On the discrete
Shanker distribution. Chilean Journal of Statistics, 136(3): 6-14.
22. Josmar, M., Wesley, B., Oliveira, R. P., and André, F. B. M. (2020). On the
Discrete Quasi Xgamma Distribution. Methodology and Computing in
Applied Probability, 22(2): 747-775.
23. Kemp, A. W. (1997). Characterization of a discrete normal distribution.
Journal of Statistical Planning and Inference, 63(1): 223-229.
24. Kozubowski, T. J., and Inusah, S. (2006). A skew Laplace distribution on
Integers. Annals of the Institute of Statistical Mathematics, 58(1): 555–571
25. Krishna, H., and Pundir, P. (2009). Discrete Burr and discrete Pareto
distributions. Statistical Methodology, 6(2): 177-188.
26. Kulasekara, K. B., and Tonkyn, D. W. (1992). A new discrete distribution with
application to survival, dispersal and dispersion. Communications in Statistics
- Simulation and Computation, 21 (2): 499-518.
27. Lai, C. D. (2013). Issues concerning constructions of discrete lifetime models.
Quality Technology & Quantitative Management, 10(2): 251-262.
28. Lekshmi, S., and Sebastian, S. (2014). A skewed generalized discrete Laplace
distribution. International Journal of Mathematics and Statistics Invention,
2(3): 95-102.
29. Liang, T. C. (1999). Monotone empirical Bayes tests for a discrete normal
distribution. Statistics & Probability Letters, 44(3): 241-249.
30. Lisman, J. H. C., and Zuylen, V. M. C. A. (1972). Note on the generation of
the most probable frequency distribution. Statistica Neerlandica, Netherlands
Society for Statistics and Operations Research, 26(1): 19-23.
97
31. Marshall, A. W., and Olkin, I. (1997). A new method for adding a parameter to
a family of distributions with applications to the exponential and Weibull
families. Biometrika, 84(3): 641-652.
32. Nakagawa, T., and Osaki, S. (1975). The discrete Weibull distribution. IEEE
Transactions on Reliability, 24(5): 300-301.
33. Nekoukhou, V., Alamatsaz, M. H., and Bidram, H. (2012). A discrete analog of
the generalized exponential distribution. Communications in Statistics-Theory
and Methods, 41(11): 2000-2013.
34. Nekoukhou, V., Alamatsaz, M. H., Bidram, H., and Aghajani, A. H. (2015).
Discrete Beta-Exponential Distribution. Communications in Statistics -
Theory and Methods, 44(10): 2079-2091.
35. Opone, F. C., Izekor, E. A., Akata, I. U., and Osagiede, F. E. U. (2020). A
Discrete Analogue of the Continuous Marshall-Olkin Weibull Distribution
with Application to Count Data. Earthline Journal of Mathematical Sciences,
5(2): 415-428.
36. Para, B. A., and Jan, T. R. (2019). On three parameters discrete generalized
inverse Weibull distribution: properties and applications. Annals of Data
Science, 6(3): 549-570.
37. Pellegrin, B., Mezzari, F., Hanafi, Y., Szymczyk, A., Remigy, J. C., and
Causserand, C. (2015). Filtration performance and pore size distribution of
hypochlorite aged PES/PVP ultrafiltration membranes. Journal of Membrane
Science, 474 (1): 175-186.
38. Roknabadi, R., Borzadaran, A. H. M., and Khorashadizadeh, G. R. M. (2009).
Some aspects of discrete telescopic hazard rate function in telescopic families.
Economic Quality Control, 24(1): 35-42.
39. Sato, H., Ikota M., Sugimoto, A., and Masuda, H. (1999). A new defect
distribution meteorology with a consistent discrete exponential formula and its
applications. IEEE Transactions on Semiconductor Manufacturing, 12(4):
409–418.
40. Siromoney, G. (1964). The general Dirichlet‘s series distribution. Journal of
Indian Statistical Association, 2(3): 69-74.
41. Szablowski, P. J. (2001). ‗Discrete normal distribution and its relationship
with Jacobi Theta functions‘, Statistics & Probability Letters, 52(3): 289-299.

98
An intensive assessment of Precipitation, Humidity and
Temperature patterns through the analytical framework of
Extreme Value Theory
Dipanjali Ray1, Tanusree Deb Roy2, and Sebul Islam Laskar3
*1
Research Scholar, Department of Statistics, Assam University, Silchar,
dipanjali0298@[Link]
2
Assistant Professor, Department of Statistics, Assam University, Silchar,
[Link]@[Link]
3
India Meteorological Department (IMD), New Delhi, India, drsebul@[Link]

Abstract:
The aim of this study is to rigorously assess the suitability of different
probability distributions by applying a range of goodness-of-fit tests, with the goal
of formulating a standardized selection framework. The analysis is conducted using
climatic parameters, including total monthly precipitation, monthly mean maximum
and minimum temperatures, and humidity levels at 12 pm and 3 pm, recorded at
Guwahati station, one of the largest observational sites in Assam, over the period
1985 to 2022. Three extreme value distributions are employed to analyse the
probabilistic patterns of temperature, humidity, and rainfall at the Guwahati station.
The parameters of these distributions are estimated using the Maximum Likelihood
Estimation (MLE) technique, providing an accurate and reliable framework for the
analysis. Using various criteria of goodness of fit test i.e. the Kolmogorov-Smirnov
(K-S) test, Bayesian information criteria (BIC), Akaike information criterion (AIC),
the most suitable probability distribution is determined. Based on the extreme event
analysis of the selected key climatic parameters for the Guwahati station, it was
observed that approximately 65% of these variables obey to the Weibull distribution,
while the remaining 35% align with the Gumbel distribution in case of monthly data.
In season wise data, 60% is followed by Weibull distribution and 30% is covered by
Gumbel distribution.
Key Words: Rainfall, Humidity at 12 pm and 3 pm, maximum temperature,
minimum temperature, Gumbel, Frechet and Weibull distribution.

1. Introduction:
Climate represents a complex dynamical system shaped by a multitude of
interacting forces. While it is influenced by vast external drivers such as solar


Corresponding author: Email: dipanjali0298@[Link]

99
radiation and terrestrial topography, it is ultimately governed by an intricate array of
factors. Only a limited subset of these variables can be directly quantified, including
air temperature, atmospheric pressure, humidity, precipitation, solar radiation, and
wind. The remaining influences, due to their complexity, are often treated as
background variability or stochastic noise within the system (Ray et al. 2025).
Among the various climatic parameters, Temperature emerges as the most pivotal
and influential factor (Yáñez-López et al., 2012), In recent decades, rising
temperatures across various regions have posed significant socio-economic
challenges, with extreme heat exhibiting region-specific sensitivity to shifting
thresholds (Ray et al., 2025). Alongside temperature, precipitation and humidity are
also vital climatic parameters, intricately linked through their mutual influence on
atmospheric processes. The sun‘s radiation (Temperature) initiates the global
hydrological cycle by heating the Earth‘s surface, causing water to evaporate from
oceans and land. This moisture (Humidity) is transported by atmospheric winds,
condenses into clouds, and precipitates as Rainfall (the largest form of precipitation)
or snow. The water then flows through rivers back to the oceans, completing the
continuous cycle (Trenberth, 2011).
In light of the critical role these climatic parameters play, the maximum &
minimum temperature, humidity at two different times i.e. at 12 pm and 3 pm and
rainfall will be explored through the frame of distributional approach for the largest
city of Assam and one of the largest metropolis cities in Northeastern India is
Guwahati station. Guwahati City, situated in the northeastern region of India, lies
approximately between latitudes 26°4′ 45″ N and 26°14′ N, and longitudes
91°33′ E to 91°52′ 6″ E. Straddling both banks of the Brahmaputra River, the city
encompasses an area of about 328 square kilometres. Characterized by a subtropical
climate, the region experiences hot, humid summers, intense monsoonal rainfall, and
mild winters (Sarma et al., 2020). Guwahati Station has long held strategic
importance as a vibrant canter of tourism, trade, commerce, administration, and
political activity. Parts of the region are recognized among the rainiest in the world.
However, this climatic intensity, coupled with the city's topography, contributes to
its pronounced vulnerability to natural hazards such as floods, landslides, and
riverbank erosion (Hemani & Das, 2016).
A wide range of studies have utilized probability distribution functions to
analyse extreme events related to rainfall, maximum and minimum temperatures,
and humidity across multiple dimensions and contexts. In 1998, Duan et al. have
discussed for historical daily precipitation data for the eastern United States, the
Weibull distribution offers a better fit (Duan et al., 1998). In 1981, Swift and
Schreuder have considered Log-normal, Gamma, Weibull, SB and beta distribution
on daily precipitation amount considering high precipitation zone and concluded that
SB distribution gives better fit on that particular zone (Swift & Schreuder, 1981). In
2013, Hasan and Kassim tried to fit Generalized extreme value distribution in
temperature data and studied the stochastic trend in Malaysia (Hasan & Kassim,
2013). In 2015, Shanshoury have discussed about various techniques in three types
100
of extreme value distribution to fit temperature data in Dabba Region (El-
shanshoury, 2015). In 2018, Vivekanandan have studied thoroughly one day series
maximum rainfall, maximum temperature and minimum temperature by adopting
two types of extreme value distribution (type 1 & type 2) considering various
criteria‘s of goodness of fit (Vivekanandan, 2018). In 2021, Norrulashikin et al. have
discussed about the fitting of probability density function namely, Beta, Burr,
Gamm, Log normal and Weibull in meteorological data such as temperature,
maximum wind speed and precipitation from 1985-2009 (Norrulashikin et al., 2021).
In 2021, Gurung et al. have proposed generalized form of Gumbel distribution i.e.
type 1 extreme value distribution to model annual rainfall in India (Gurung et al.,
2021). In 2018, Ju et al. have presented work on energy efficient hot air drying by
controlling relative humidity base on Weibull model (Ju et al., 2018). In 2022,
Karakya and Jain tried to model temperature and relative humidity using Weibull,
Normal, Logistic, Gamma and log normal distribution (Karakaya & Jain, 2022).
Numerous studies have investigated the application of probability distribution
fitting to meteorological parameters across diverse regions and contexts. Within this
framework, the use of extreme value distributions is particularly critical, given the
inherently volatile nature of weather variables, which often exhibit sudden spikes or
unusually low values. To effectively characterize these extremes, the present study
focuses on fitting three specific types of extreme value distributions to selected
meteorological parameters in Guwahati station, as these models are best suited to
capture the statistical behaviours of such deviations.

2. Methodology:
Source of Data:
The quantitative information monthly mean maximum temperature & minimum
temperature, total monthly rainfall and humidity at 12 pm and 3 pm has been
collected for Guwahati station of Assam from National Data Centre (NDC), India
Meteorological Department (IMD), Pune ([Link] for the period
of 1985-2022.
Extreme value distributions (EVDs) play a critical role in modelling and
analysing rare or extreme phenomena. The theoretical framework underpinning
these distributions was established by Leonard Tippett (1902–1985), whose
pioneering contributions form the basis of contemporary methods in extreme value
analysis. The theory identifies three distinct types of EVDs, each designed to
represent the statistical behaviour of extreme values, either maxima or minima under
different conditions (Ray et al., 2025).
Gumbel or Type 1 EVD
If X is governed by a Gumbel distribution, or Type I Extreme Value Distribution
(EVD), characterized by a location parameter μ and a scale parameter  , the
probability density function (PDF) for modelling minima or maxima can be defined
as follows.
101
x  
x  x  x
1 
f x  e
1 
 e 
e ; x  0 ,   0,   0 or f x  e  e 
e ; x  0 ,   0,   0 (1)
 

Frechet or Type 2 EVD


If X follows Frechet distribution with shape parameter c and scale parameter 
with a unit location parameter then the probability density function can be defined as
f x  c xc1 ex ,
c
x  0, c  0 ,   0 (2)
Weibull or Type 3 EVD
If X follows Weibull distribution with shape parameter α and scale parameter
β , then the probability density function can be defined as

 x
  
f  x;  ;     x 1e    ; x  0,   0,   0 (3)

Method of Estimation and Goodness of fit test
The estimation of parameter of these three extreme value distributions can be
done by the method of Maximum Likelihood Estimation (MLE).
Let X1, X 2 … X n be the sample of size n from a particular distribution with
density function f ( x,  ) then the likelihood function of sample values x1 , x2 , …
xn is given by
n
L  f ( x1,  ), f ( x2 ,  )... f ( xn ,  )   f ( xi ,  ) (4)
i 1
The principle of maximum likelihood consists in finding an estimator for
unknown parameter   ( 1 ,  2 ,..., k ) , which maximize the likelihood function L.
Now taking log L and differentiating with respect to unknown parameter and
equating to zero one gets value of  which maximizes L for variations in  , which
is called Maximum Likelihood Estimator if (Gupta & Kapoor, 1997)
d ( L) d 2 ( L) (5)
0 and 0
d d 2
To evaluate the suitability of the selected probability distributions for the
meteorological data, the Kolmogorov-Smirnov (K-S) test is initially employed. The
K-S test statistic quantifies the maximum discrepancy between the empirical sample
cumulative distribution function (CDF) and the theoretical CDF, and can be
expressed as follows:
𝑘 𝑘−1
𝐷𝑛 = max1≤𝑘≤𝑛 2|𝑛 − 𝐹̂ (𝑥(𝑘) )| , |𝐹̂ (𝑥(𝑘) ) − 𝑛 |3 (6)

102
Here 𝐹̂ (𝑥(𝑘) ) denotes is the estimated value of the cumulative distribution
function (CDF), and {𝑥(1) , … , 𝑥(𝑛) } represents the ordered observations in ascending
sequence. If the test statistics 𝐷𝑛 exceeds the 𝐷𝑛 (𝛼) critical value, the null
hypothesis will be rejected at 𝛼 significance level, where 𝐷𝑛 (𝛼) corresponds to the
critical threshold for the K-S statistic (Moccia et al. 2021)
In order to compare these distributions, there are some criterions like -2log(L),
AIC (Akaike Information Criterion), BIC (Baysian Information Criterion) to test the
goodness of fit. Lower the value of -2log(L), AIC, BIC corresponds to the best
distribution.
AIC (Akaike Information Criterion):
AIC   2 log L   2k

  (7)
Where L   L x1, x2 ,..., xn ;   is the maximum likelihood function for the
 

   

estimated model.  is MLE of 1, 2 ,...n  for which likelihood function is
maximum, k is number of parameters of a model, n is the size of data. AIC balances
the lack of fit and model complexity, so smaller the value of AIC indicates the better
model. Minimum value of BIC indicates the higher posterior probability (Akaike,
1969; Schwarz, 1978; Ozonur et al., 2021; Mohamad & Adam, 2022).
BIC (Baysian Information Criterion):
BIC  k log n   2 log L 

  (8)
3. Result and Discussion:
In this section, the best fit extreme value distribution is elaborately discussed for
total monthly mean rainfall (mm), monthly mean maximum temperature and
minimum temperature (0C), monthly mean relative humidity (%) at 12 pm and 3 pm
monthly wise and season wise for Guwahati station. For season wise extreme value
analysis, the whole data is divided into four seasons i.e. Winter (January, February),
Pre-Monsoon (March, April, May, June), Monsoon (July, August, September) and
Post-Monsoon (October, November, December).
Table 1. Estimates of parameters of extreme probability distribution fitting of
monthly mean maximum temperature in Guwahati station

Guwahati Station
Month Distribution estimates
location scale shape
January Weibull 24.5141 22.4232
Gumbel 23.3946 1.0965
Frechet 0.3149 0.0417
103
February Weibull 27.5961 18.5882
Gumbel 26.0472 1.4977
Frechet 0.3041 0.0373
March Weibul 31.1214 21.2640
Gumbel 29.5394 1.6171
Frechet 0.2931 0.0329
April Weibul 31.8032 18.1072
Gumbel 30.2787 1.5624
Frechet 0.2912 0.0322
May Weibull 32.2380 33.9910
Gumbel 31.1382 1.3214
Frechet 0.2893 0.0315
June Weibull 32.9870 32.6804
Gumbel 32.0562 0.7958
Frechet 0.2873 0.0308
July Weibull 33.2771 29.3995
Gumbel 32.1279 1.0749
Frechet 0.2868 0.0306
August Weibull 33.6156 30.5448
Gumbel 32.6132 0.9198
Frechet 0.2855 0.0302
September Weibull 32.7737 35.3299
Gumbel 31.8417 0.8511
Frechet 0.2878 0.0310
October Weibull 31.7433 25.5403
Gumbel 30.5146 1.1424
Frechet 0.2909 0.0321
November Weibull 28.8511 30.7414
Gumbel 27.9061 0.8695
Frechet 0.2990 0.0353
December Weibull 25.7308 28.7710
Gumbel 24.8034 0.9564
Frechet 0.3097 0.0396

104
Season wise
Winter Weibull 26.3537 13.1417
Gumbel 24.4411 1.7401
Frechet
Pre-Monsoon Weibull 31.7655 21.0083
Gumbel 30.2454 1.6593
Frechet 0.2912 0.0322
Monsoon Weibull 33.1850 29.8868
Gumbel 32.1372 0.9540
Frechet 0.2869 0.0306
Post-Mon Weibull 29.4635 11.8565
Gumbel 26.9677 2.3738
Frechet 0.2997 0.0354

Table 2. Goodness of fit test of probability distribution fitting of monthly mean


maximum temperature in Guwahati station

Guwahati Station
Goodness of fit
Month Distribution -2logL AIC BIC KS p-value
January Weibull 121.2942 125.2942 128.5694 0.0834 0.9541
Gumbel 122.7805 126.7805 130.0557 0.1112 0.7357
Frechet 722.5943 726.5943 729.8695 0.5996 2.72E-12
February Weibull 145.3266 149.3266 152.6018 0.1023 0.8218
Gumbel 147.1697 151.1697 154.4449 0.1221 0.6226
Frechet 742.4236 746.4236 779.0025 0.6019 2.20E-12
March Weibul 144.8786 148.8786 152.1538 0.0884 0.9276
Gumbel 151.4851 155.4851 158.7603 0.1611 0.2776
Frechet 764.0014 768.0014 771.2766 0.6044 1.75E-12
April Weibul 152.0297 156.0297 159.3048 0.1808 0.1669
Gumbel 146.9416 150.9416 154.2168 0.1887 0.1336
Frechet 767.8280 771.8280 775.1032 0.6044 1.76E-12
May Weibull 113.2518 117.2518 120.5270 0.1213 0.6305
Gumbel 131.2883 135.2883 138.5634 0.1772 0.184
Frechet 771.7273 775.7273 779.0025 0.6057 1.55E-12
June Weibull 113.1558 117.1558 120.4310 0.1383 0.4615

105
Gumbel 101.5223 105.5223 108.7975 0.0986 0.8541
Frechet 775.9367 779.9367 783.2119 0.6060 1.51E-12
July Weibull 124.0374 128.0374 131.3126 0.1053 0.793
Gumbel 122.5238 126.5238 129.7990 0.1047 0.7996
Frechet 776.9954 780.9954 784.2706 0.6061 1.50E-12
August Weibull 118.6948 122.6948 125.9700 0.1171 0.6751
Gumbel 110.2958 114.2958 117.5709 0.1225 0.6183
Frechet 779.1106 783.1106 786.3858 0.6063 1.48E-12
September Weibull 108.9775 112.9775 116.2527 0.1242 0.6008
Gumbel 105.1422 109.1422 112.4174 0.0845 0.9489
Frechet 774.8631 778.8631 782.1383 0.6061 1.51E-12
October Weibull -65.2469 134.4937 137.7689 0.1317 0.5248
Gumbel -63.4889 130.9779 134.2530 0.0767 0.9786
Frechet 772.3721 772.3721 775.6473 0.6052 1.64E-12
November Weibull -54.9861 113.9721 117.2473 0.1248 0.5951
Gumbel -53.2394 110.4788 113.7540 0.0771 0.9777
Frechet 752.2477 756.2477 759.5229 0.6037 1.88E-12
December Weibull 106.1505 110.1505 113.4257 0.1458 0.3947
Gumbel 110.6211 114.6211 117.8963 0.1681 0.2329
Frechet 732.0126 736.0126 739.2878 0.6011 2.36E-12
Season wise
Winter Weibull 330.9932 334.9932 339.6547 0.1585 0.0439
Gumbel 320.5230 324.5230 329.1844 0.0651 0.904
Frechet 1465.2850 1469.2850 1473.9460 0.6000 2.20E-16
Pre-Mon Weibull -216.2527 436.5055 441.9779 0.0856 0.3739
Gumbel -226.7900 457.5799 463.0523 0.1565 0.00751
Frechet 2303.5970 1469.2850 1474.7570 0.6044 2.20E-16
Monsoon Weibull 481.5998 485.5999 491.6476 0.0951 0.1281
Gumbel 454.9882 458.9882 465.0360 0.0696 0.4539
Frechet 3106.9190 3110.9190 3116.9670 0.6059 2.20E-16
Post-Mon Weibull 549.1832 553.1832 558.6556 0.0776 0.4986
Gumbel 549.4730 553.4729 558.9453 0.0893 0.3236
Frechet 2253.5290 2257.5290 2263.0010 0.6021 2.20E-16

106
In Table 1, all parameter estimates are represented for three distributions
monthly and season wise. In Table 2, Probability distribution fitting of monthly
mean maximum temperature in Guwahati station for monthly and season wise is
presented. It elaborates parameters for each distribution and comparison criteria‘s
i.e. -2logL (log likelihood), AIC (Akaike Information Criterion), BIC (Bayesian
Information Criterion), K-S distance value and p-value. Fitting of best probability
distribution on a data set may be characterized as finding the distribution with the
lowest values of -2logL, AIC, BIC, and K-S values.
In the month of January, the observation points in January month are following
both Weibull (0.9541) and Gumbel (0.7357) distribution among three extreme value
distribution. Since both are having greater than 0.05, p-value of K-S test i.e. it
satisfies the null hypothesis of K-S test and null hypothesis of the K-S test is that the
data points follow the required distribution. But it is observed that by taking into
consideration of other criteria‘s i.e. -2logL, AIC, BIC and K-S distance, the Weibull
distribution fits better since it has least -2logL (121.2942), AIC (125.2942), BIC
(128.5694) and K-S distance (0.834).
The monthly mean maximum temperature of Guwahati station follows 58.33%
of Weibull distribution and 41.66% of Gumbel distribution monthly wise and season
wise it follows 50% of Weibull distribution and 50% of Gumbel distribution.
Similarly, the fitting of extreme value distribution can be done for Monthly
mean minimum temperature, total mean monthly rainfall, monthly mean relative
humidity at 12pm and 3pm of Guwahati station and it is done for all parameters of
remain station also.
Table 3. Parameter estimates and Goodness of fit test of probability distribution of
monthly mean minimum Temperature in Guwahati station

Month Distribution estimates


location scale shape
Jan Gumbel 10.5823 0.7809
Feb Weibull 13.3868 12.2513
March Weibuil 16.9786 19.4137
April Gumbel 19.8753 0.8128
May Weibull 23.1774 37.6125
June Weibull 25.4098 64.3403
July Weibull 26.0891 62.2112
Aug Weibull 26.0460 56.8702
Sept Weibull 25.1454 50.3479
Oct Weibull 22.6211 25.4787
Nov Weibull 17.6304 18.9104

107
Dec Gumbel 12.0820 0.8968
Winter Gumbel 11.2885 1.1803
Pre-Mon Weibull 21.0771 8.8317
Mon Weibull 25.7448 43.4497
Post-Mon Gumbel 15.2784 3.5700
Goodness of fit
-2logL AIC BIC KS p-value
Jan Gumbel 96.0628 100.0628 103.3379 0.0923 0.9023
Feb Weibull 117.8065 121.8065 125.0817 0.0866 0.9381
March Weibuil 103.3005 107.3005 110.5757 0.1241 0.6017
April Gumbel 99.2190 103.2190 106.4942 0.1422 0.4259
May Weibull 77.6468 81.6468 84.9220 0.1095 0.7525
June Weibull 50.4050 54.4050 57.6802 0.0852 0.9457
July Weibull 52.2757 56.2757 59.5509 0.0802 0.9674
Aug Weibull 60.6819 64.6819 67.9571 0.0889 0.9246
Sept Weibull 63.7710 67.7711 71.0462 0.1844 0.1507
Oct Weibull 104.9140 108.9140 112.1892 0.0932 0.896
Nov Weibull 112.3872 116.3872 119.6623 0.1118 0.7295
Dec Gumbel 106.7157 110.7157 113.9909 0.1530 0.3361
Winter Gumbel 260.7062 264.7061 269.3676 0.0748 0.789
Pre-Mon Weibull 545.6904 549.6905 555.1629 0.1203 0.0737
Mon Weibull 310.6206 314.6206 320.6684 0.0640 0.5628
Post-Mon Gumbel 646.3690 650.3690 655.8414 0.1191 0.07878

Table 4. Parameter estimates and Goodness of fit test of probability distribution of


monthly total mean Rainfall in Guwahati station

Month Distribution estimates


location scale shape
Jan Gumbel 5.7015 8.4914
Feb Gumbel 11.7239 12.7867
March Weibull 56.5877 1.3049
April Gumbel 132.5933 83.0751

108
May Weibull 279.3060 2.1867
June Weibull 2.8549 339.3131
July Gumbel 248.8695 89.8328
Aug Weibull 270.3478 2.31954
Sept Gumbel 149.3745 73.3686
Oct Gumbel 78.6785 62.8008
Nov Weibull 0.3705 0.3705
Dec Weibull 0.6858 0.2413
Winter Gumbel 8.4552 10.9346
Pre-Mon Weibull 170.1745 1.2388
Mon Weibull 291.8392 2.3422
Post-Mon Weibull 13.3164 0.3108
Goodness of fit
-2logL AIC BIC KS p-value
Jan Gumbel 289.4652 293.4652 296.7404 0.1681 0.2331
Feb Gumbel 317.9188 321.9187 325.1939 0.1173 0.6723
March Weibull 372.7302 376.7303 380.0055 0.0723 0.9803
April Gumbel 454.5254 458.5254 461.8005 0.0932 0.8655
May Weibull 466.7906 470.7906 474.0658 0.1054 0.7533
June Weibull 467.1984 471.1984 474.4735 0.0865 0.9152
July Gumbel 462.7566 466.7565 470.0317 0.0612 0.9988
Aug Weibull 461.8864 465.8865 469.1616 0.0726 0.9881
Sept Gumbel 444.8068 448.8068 452.0820 0.0561 0.9998
Oct Gumbel 435.2548 439.2548 442.5299 0.1425 0.4230
Nov Weibull 188.2986 192.2986 195.5738 0.1949 0.1115
Dec Weibull 10.2842 14.2842 17.5593 0.2829 0.0046
Winter Gumbel 615.1228 619.1229 623.7843 0.1374 0.1135
Pre-Mon Weibull 1376.6242 1380.6240 1386.0970 0.0559 0.8690
Mon Weibull 1867.8854 1871.8850 1877.9330 0.0493 0.8545
Post-Mon Weibull 746.5738 750.5739 756.0463 0.1678 0.0033

109
Table 5. Parameter estimates and Goodness of fit test of probability distribution of
monthly mean relative Humidity in Guwahati station at 12pm

Months Distribution estimates


location scale shape
Jan Weibull 73.2640 16.7997
Feb Gumbel 55.8699 4.5534
March Weibull 53.9630 10.1944
April Weibull 65.6050 11.1196
May Weibull 73.1008 18.1295
June Weibull 23.7448 79.3822
July Gumbel 78.9132 2.2435
Aug Weibull 81.8708 29.6199
Sept Gumbel 80.5615 2.2325
Oct Weibull 80.5336 34.9541
Nov Weibull 78.3515 35.2327
Dec Weibull 77.9301 23.1495
Winter Weibull 68.2787 9.2670
Pre-Mon Weibull 65.8812 7.4318
Mon Weibull 81.4953 26.4965
Post-Mon Weibull 79.0724 27.7252
Goodness of fit
-2logL AIC BIC KS p-value
Jan Weibull 223.6380 227.6381 230.9132 0.1748 0.1958
Feb Gumbel 232.0512 236.0513 239.3265 0.1048 0.7982
March Weibull 241.9160 245.9160 249.1912 0.0789 0.9719
April Weibull 243.8600 247.8600 251.0818 0.1167 0.6951
May Weibull 220.3100 224.3100 227.5852 0.1077 0.7706
June Weibull 201.7064 205.7064 208.9816 0.1690 0.2280
July Gumbel 176.8920 180.8920 184.1672 0.1477 0.3784
Aug Weibull 192.9596 196.9596 200.2347 0.1405 0.4410
Sept Gumbel 177.0245 181.0245 184.2997 0.1765 0.1871
Oct Weibull 181.0169 185.0169 188.2921 0.1277 0.5653
Nov Weibull 168.3984 172.3984 175.6202 0.1429 0.4362
Dec Weibull 208.7238 212.7237 215.9989 0.1268 0.5742
Winter Weibull 531.2938 535.2938 539.9553 0.0850 0.6426
Pre-Mon Weibull 834.2194 838.2194 843.6741 0.0745 0.5574
Mon Weibull 790.9180 794.9180 800.9658 0.1244 0.0181
Post-Mon Weibull 582.4064 586.4064 591.8612 0.1151 0.1004

110
Table 6. Parameter estimates and Goodness of fit test of probability distribution of
monthly mean relative Humidity in Guwahati station at 3 pm
Months Distribution estimates
location scale shape
Jan Weibull 88.6565 47.8045
Feb Gumbel 76.3177 3.7943
March Gumbel 66.1856 5.7842
April Weibull 76.7247 18.4328
May Gumbel 77.5429 2.7802
June Weibull 84.5207 34.2342
July Weibull 85.5728 35.2285
Aug Weibull 84.4145 39.7415
Sept Weibull 84.1660 47.3393
Oct Gumbel 81.0562 2.3988
Nov Weibull 83.9704 48.1541
Dec Weibull 88.0812 40.3433
Winter Weibull 85.4720 18.1215
Pre-Mon Weibull 76.8994 14.5367
Mon Weibull 84.7001 36.1276
Post-Mon Gumbel 82.6307 2.7965
Goodness of fit
-2logL AIC BIC KS p-value
Jan Weibull 165.4548 169.4548 172.7300 0.1277 0.5652
Feb Gumbel 215.2858 219.2858 222.5610 0.1376 0.4680
March Gumbel 245.3206 249.3205 252.5957 0.1456 0.3966
April Weibull 218.6036 222.6035 225.8254 0.1004 0.8501
May Gumbel 195.7086 199.7086 202.9838 0.1123 0.7239
June Weibull 188.6489 192.6489 195.9240 0.1129 0.7178
July Weibull 181.5091 185.5091 188.7843 0.1910 0.1249
Aug Weibull 172.9855 176.9855 180.2606 0.1789 0.1755
Sept Weibull 163.3646 167.3646 170.6398 0.1241 0.6019
Oct Gumbel 180.7574 184.7574 188.0325 0.1494 0.3647
Nov Weibull 157.3556 161.3556 164.6308 0.1511 0.3507
Dec Weibull 175.1843 179.1843 182.4595 0.1184 0.6607
Winter Weibull 472.7550 476.7549 481.4164 0.1213 0.2135
Pre-Mon Weibull 722.9864 726.9864 732.4411 0.0761 0.5298
Mon Weibull 721.8254 725.8254 731.8732 0.1368 0.0068
Post-Mon Gumbel 581.9286 585.9287 591.4011 0.1208 0.0720

111
From table 3 to 5, the estimation of parameters and goodness of fit have been
done for the fitting of best extreme value distribution among three for the remaining
parameters. Similarly, same analysis is completed for other stations. From that,
Weibull distribution fits 65 % combining all parameters of weather in Guwahati
station among three extreme value distributions and Gumbel fits 35%.

4. Conclusion and Future work:


The pattern of extreme events of climate parameters are studied by the three
types of extreme value distribution namely extreme value type 1 or Gumbel
distribution, extreme value type 2 or Frechet distribution and extreme value type 3
or Weibull distribution which have been tried to fit in each parameter for each
station both monthly and season wise data in the third objective. The parameter is
estimated by the method of MLE. To compare the fitting of these distributions some
goodness of fit criteria are taken into consideration such as log Likelihood (-2logL),
Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC),
Kolmogorov Smirnov test (K-S test). Considering all the parameters i.e. monthly
mean maximum temperature and minimum temperature, total monthly mean rainfall,
humidity at 12 pm and 3 pm of Guwahati station, it has been observed that 65% of
these parameters are followed by Weibull distribution and 35% is followed by
Gumbel distribution in monthly data and in season wise Weibull distribution covers
60% and 30% is followed by Gumbel distribution.
In several meteorological study, the extreme value distribution is considered to
study the various parameter of weather when it is about the fitting of distribution
function. But, besides these extreme value distributions in numerous cases other
form of distribution or extended form of these classical extreme value distributions
can been attempted in meteorological study. In future by adopting various
distribution rather than classical extreme value distribution, a depth study can be
possible for weather parameters.
REFERENCES
1) Akaike, H. (1969), Fitting autoregressive models for prediction. Annals of the
Institute of Statistical Mathematics 21, 243-247
2) Duan, J., Selker J., & Grant, G. E. (1998). Evaluation of Probability Density
Functions in Precipitation Models for the Pacific Northwest. JAWRA Journal
of the American Water Resources Association, 34(3), 617–627
3) El-shanshoury, G. (2015). Assessing the Adequacy of Probability Distributions
for Estimating the Extreme Events of Air Temperature in Dabaa Region, Arab
Journal of Nuclear Science and Applications, 48(2), 104-120
4) Gupta, S. C. and Kapoor, V. K. (1997). Fundamentals of Mathematical
Statistics. Sultan Chand and Sons, New Delhi: 11.23-12.2
5) Gurung, B., Sarkar K. P., Singh K. N., & Lama A. (2021). Modelling annual
maximum temperature of India: A distributional approach. Theoretical and
Applied Climatology, 145, 979–988.
112
6) Hasan, H., & Kassim S. (2013). Modeling Annual Extreme Temperature
Using Generalized Extreme Value Distribution: A Case Study in Malaysia.
AIP Conference Proceedings, 1522, 1195–1203.
7) Hemani, S., & Das, A. K. (2016). City profile: Guwahati. Cities, 50, 137–157.
8) Ju HY, Zhao SH, Mujumdar AS, Fang XM, Gao ZJ, Zheng ZA, & Xiao HW
(2018). Energy efficient improvements in hot air drying by controlling relative
humidity based on Weibull and Bi-Di models. Food and Bioproducts
Processing, 111, 20–29.
9) Karakaya, K., & Jain, V. (2022). Role of Probability Models for Enhancing the
Information of Temperature and Relative Humidity in Hoshangbad. J. Agri.
Bio. Appl. Stats, 1(1), 51-61.
10) Moccia, B., Mineo, C., Ridolfi, E., Russo, F. & Napolitano, F. (2021).
Probability distributions of daily rainfall extremes in Lazio and Sicily, Italy
and design rainfall inference, Journal of Hydrology: Regional Studies,
33(2021), 100771
11) Mohamed, J., Adam, B. M. (2022). Modeling of magnitude and frequency of
extreme rainfall in Somalia, Modelling Earth Systems and Environment, 8,
4277-4294
12) Norrulashikin, S. M., Yusof, F., Nor, S. R. M., & Kamisan, N. A. B. (2021).
Best Fitted Distribution For Meteorological Data In Kuala Krai. Journal of
Statistical Modeling & Analytics (JOSMA), 3(1), Article 1.
13) Ozonur, D., Pobocikova, I., & Souza, A. (2021). Statistical analysis of
monthly rainfall in Central West Brazil using probability distributions.
Modeling Earth Systems and Environment, 7, 1979-1989.
14) Ray, D., Roy, T. D. & Laskar, S. I. (2025). A statistical perspective on Assam‘s
temperature pattern from 1985-2022‖, MAUSAM, 76(2):471–484
15) Sarma, C. P., Dey, A., & Krishna, A. M. (2020). Influence of digital elevation
models on the simulation of rainfall-induced landslides in the hillslopes of
Guwahati, India. Engineering Geology, 268.
16) Schwarz, G. (1978), Estimating the dimension of a model. Annals of
Statistics, 6, 461-464
17) Swift Jr, L. W. & Schreuder, H. T. (1981). Fitting daily precipitation amounts
using the SB distribution. Monthly Weather Review, 109(12), 2535-2540.
18) Trenberth, K. E. (2011). Changes in precipitation with climate
change. Climate research, 47(1-2), 123-138.
19) Vivekanandan, N. (2018). Comparison of probability distributions in extreme
value analysis of rainfall and temperature data. Environmental Earth Sciences,
77(5), 201.
20) Yáñez-López, R., Torres-Pacheco, I., Guevara-González, R. G., Hernández-
Zul, M. I., Quijano-Carranza, J. A., & Rico-García, E. (2012). The effect of
climate change on plant diseases. African Journal of Biotechnology, 11(10),
Article 10.

113
International trade relation of Assam with neighboring
country Bhutan: Land Custom Station (LCS) based
analysis
Rupjyoti Bordoloi
Department of Economics,
Rabindranath Thakur Vishwavidyalaya

Abstract:
The present study has focused on the trade volume and it growth pattern
between India and Bhutan through the Land Custom Stations located in Assam. The
present study is based on secondary sources. Descriptive statistical methods have
been incorporated to analyzed the data for receiving the target of the study. There
are three active land custom stations along the Assam-Bhutan Border, named
Hatisar in Chirang district, Darranga and Kamardwisa earlier Rangapani in Baksa
district. All the three land custom stations are located under Bodoland Territorial
Region (BTR) of Assam. Among the custom stations Hatisar station is more active
and well with the Chirang of Assam. The exported items are dominated by
industrial manufactured items and petroleum products like HSDO, SKO and Mild
Steel. The amount of total trade volume has increased over the years in both the
countries. Lack of trade related transparency and in-adequate infrastructure in
border districts of both the countries are becoming major constraints for seamless
movement of goods.
Key Words: Land Custom Station, export, import, trade,

1.1 Introduction:
Trade has always played a central role in the economic growth and development
of nations, fostering cooperation and mutual prosperity among countries (Ismail et
al. 2015). International trade has given an opportunity for generating foreign
exchange, employment opportunity, new skills and technology, market and various
other opportunities for improving both trade partners. The border trade serves not
only as a gateway but also fosters the development of the borderland economy.
Crucially, these border corridors stimulate the local economy by facilitating frequent
transactions of goods and services, which can boost economic growth in both India
and Bhutan. (Lama, 2023) The WTO (World Trade Organization) is vital for
managing cross-border trade between countries. Its goal is to ensure that goods and
services can move between nations smoothly and predictably by making trade
agreements. (Lama, 2023) India has a strong trade relation with its neighboring
country Bhutan from the time immemorial. Border trade has been performed
114
through Land Custom Stations (LCS) connected to the border areas. According to
Ministry of Finance, Govt. of India, Land Custom Station (LCS) is a place at the
land border notified under section 7 of the Customs Act 1962 for enabling trade with
a neighbouring country. There are a total of 122 notified Land Customs Stations
strategically located across India‘s borders for trade with seven countries viz.
Afghanistan, Pakistan, China, Myanmar, Bhutan, Bangladesh and Nepal. (Ministry
of Finance, GoI. 2024)
The border between India and Bhutan stretches for 699 kilometers, and it shares
267 km boundaries with Assam. (MHA, GoI.) The sharing with international
boundary has given an opportunity for opening trade with those neighboring
countries. In our present study we have considered one foreign country i.e. Bhutan
with Assam a border state of India. According to Indo-Bhutan Treaty of 1949, India
share special bilateral relations formulated a free trade regime between the two
countries for friendship and cooperation in trade and development. The trade issue
between India and Bhutan is guided by the Agreement on Trade, Commerce and
transit between the government of the Republic of India and The Royal government
of Bhutan. The trade between India and Bhutan primarily occurs through the border
corridors of West Bengal and Assam. Presently, more than 13 land corridors are
active through which the transaction of goods and services has been done (Ministry
of Finance, Royal Government of Bhutan) In Assam, out of 13 Land Custom
Stations (LCS), 3 are functional and other 10 LCSs non-functional with Bhutan.
Among the North Eastern states, Assam has the largest economy in terms of
state gross domestic product. Assam is situated in 24 N0 –28 N0 and 90 E0 and 96 E0
in India. It has a land of 78,438 square km with a population of 31.1 million
according to the census 2011. It has 2.4 percent of county‘s total landmass and gives
shelter nearly 3 percent of the country‘s total population. The state has sharing
border with two foreign countries Bangladesh and Bhutan. Exported items of the
state included rice, cotton, oilseeds, dried fish, timber, lac, munjeet, black pepper,
elephants, ivory, cotton textiles, eri and muga silk, brass and so on, imports included
salt, wool and blankets, beads, ponies and several other animal husbandry products.
(Mohan et al. 2017)

1.2 Objectives of the Study:


The present study has primarily focused on the nature of traded items of Assam
with the Bhutan. Further the study will examine the size and volume of traded
amount with Bhutan through Land Customs Station named Darranga, Hatisar and
Kamardwisa under the Customs Preventive Commissionerates Shillong, Meghalaya.
Besides the study has focused on the tariff rates imposed by India and Bhutan on the
import and exported items during several time period. Finally special emphasis has
been put into the problems as well as prospects of trade with Bhutan.

115
2. Review of Literature
Mohan et al. (2017) have aimed to understand the cross-border trading practices
in Chirang District, Assam (India) in with Bhutan. Their study was primary data
based specifically, the border village of Dadgiri (Assam, India) and Gelephu town of
Bhutan to understanding of cross-border market practices. They observe that the
operational governing dynamics of market provides an enriching perspective on
gauging the social, economic and political landscape of the Indo-Bhutan border area
beyond the Chirang district of Assam, India. Lama (2023) has conducted the study is
to trace sustainable economic development, diplomatic relationships, and foster
mutual prosperity. Beside this the study revolves around conducting a
comprehensive analysis of the diverse trade routes, infrastructure, and policies
governing the movement of goods and services between India and Bhutan. The
study has observed that the expansion of trade between India and Bhutan through
various border corridors holds immense potential for fostering economic growth in
both border regions. Taneza et al. (2019) has mentioned that India is Bhutan‘s largest
export market, the biggest source of its imports and one of the top foreign investors
in the country. Cooperation in hydropower projects is one of the most significant
examples of win-win cooperation between India and Bhutan. They have
recommended several policies for improving bilateral trade between countries like
implementation of automated customs systems, electronic exchange of data,
automated risk management; automated border procedures; electronic, electronic
single windows and other related digital customs and trade facilitation initiatives.
Das et al. (2022) have identifies 19 priority border points among the northeastern
states of Assam, Manipur, Meghalaya, Mizoram, Nagaland, Sikkim, and Tripura. It
discusses the challenges of trade facilitation components, including ―at-the-border‖
infrastructure, customs, nontariff measures, and transport facilitation. They
recommended the development of inland container depots, connecting them to select
border points in the northeast of India.
3. Methodology
3.1 Selection of the study area
The present study has focused on trade of Assam with one of its neighboring
country named Bhutan. In terms of economic health the state has earned highest
amount of domestic product in comparison to the other states of the region. It has
diverse types of agricultural and horticultural products, manufactured items, services
that have crossed the border of states and acclaimed international recognition in
global market. The uniqueness of agricultural product and global market of it has
attracted attention for doing a detailed analysis about the position of state in trade
business.
3.2 Data Source and Techniques
The official export-import data have been collected from Customs Offices in
India and Bhutan. Secondary information has been collected from different websites,
published materials, journals, newspapers, and other sources and presented in tabular
116
and graph format. Tabulation and bar diagrams are used to achieve the targeted
objectives of the study. Insightful argument has been used to analysis information
for reaching certain conclusion of the entire study.
4. Analysis & Findings
4.1 Trade with India an overview
Bhutan is a land locked country where most of its trade activities are carried
out though India with the rest of the world. After formation of multilateral trade
agreement i.e. South Asian Association of Regional Cooperation (SAARC) in
1985 the India‘s trade relation with Bhutan and other member countries become
very smooth and seamless. The significant change in trade practices among
countries has resulted after the formation of World Trade Organization in 1995 at
international level. The free trade regime has started between the countries through
Indo-Bhutan Agreement on Trade in 2016. India becomes the largest foreign Direct
Investor in Bhutan especially in hydro-power projects, banking and information
technology. Similarly the bilateral free trade agreement has increased the trade
volume between countries over the years. It has been reflected through the
following table and bar diagram given below.
Table no. 1: Trade direction of India with Bhutan
(Amount in Crore)
Year Export Import Total Trade Trade Balance (X-M)
2019-20 5235 2871 8106 2364
2020-21 5193 3214 8407 1979
2021-22 6605 4054 10659 2551
2022-23 8663 4310 12973 4353
2023-24 7980 2807 10787 5173
Source: Handbook of Statistics on Indian Economy, 2023-24
Fig: 1. India‘s Export & Import trade with Bhutan(Amount in Crore)
14000

12000
4310
10000
2807
4054
8000
Import
2871 3214
6000 Export

8663
4000 7980
6605
5235 5193
2000

0
2019-20 2020-21 2021-22 2022-23 2023-24

Source: Table no.1


117
The above table and diagram has showed that the volume of bilateral trade has
been increasing between these two nations over the period of time. In this trade
relation it has been cleared that India is gaining through exporting more and more
amounts of output over the years. Similarly Bhutan has improved the trade volume
with our country that has shown through increasing its average imported value.
Thus the trade relation between counties is strong and effective for both the
countries.
4.2 Trade through LCS located in Assam
In our study we have focused the trade volume that has taken place through the
Land Custom Stations located in Assam. There are three active land custom
stations along the Assam-Bhutan Border, named Hatisar in Chirang district,
Darranga and Kamardwisa earlier Rangapani in Baksa district. All the three land
custom stations are located under Bodoland Territorial Region (BTR) of Assam.
Among the custom stations Hatisar station is more active and has connected
through Motorable Road from Samthiabari to Gaylegphug (Bhutan) via Runikhata
and Deosiri in the District of Chirang of Assam. Through this route the highest
amount of commodities are traded between the countries. Another Bhutanese
boarder town Samdrup-Thankar (in Bhutan) has been connected to the Darranga
custom station located in the District of Baksa of BTR, Assam through Rangia -
Tamulpur Motorable road. Kamardwisa earlier Rangapani land custom station has
Nganglam of Bhutan is notified post of entry and exit point for trade with our
country under mentioned in the custom act 1962.
It is well known that Assam is famous for tea more than 50 percent of tea has
been exported to the world market from Assam. Apart from tea, the most traded
exported items are HSDO (High Speed Diesel Oil) Mild Steel (MS) like pipes,
plates, structural steel, TMT bars, bricks, empty bottles, motor spirit, oranges,
rectified spirit, rice and Superior Kerosene Oil (SKO) through the selected custom
stations. The imported items through these custom stations from Bhutan are ferro
alloys, Plasters, Nitrogen, Bhutanese liquor, cardamom, animal feed, block board,
plywood, boulder stone and cement.
Table no.2: Land Custom Station wise Export and Import amount between Assam
and Bhutan (Amount in crore)

Name of Land Custom Station (LCS)


Year Hatisar Darranga Kamardwisa
Export Import Export Import Export Import
2010-11 49.16 3.94 NA NA NA NA
2012-13 129.94 2.58 NA NA NA NA
2014-15 217.27 20.52 101.86 0.37 NA NA
2023-24 697.51 113.97 736.42 359.61 458.31 121.04
Source: Industries and Commerce, Govt. of Assam

118
The exported items are dominated by industrial manufactured items and
petroleum products. The amount of total trade volume has increased over the years
in both the countries. It has been observed that the trade related information from
all the land custom stations is not readily available over the year. That has further
established that the lack of adequate transparency about the trade activities in the
stations. Through the various earlier studies it has been clear that the infrastructure
facilities are not up to the mark for seamless trade practices in both the countries.
The boarder district of Assam are lagging behind due its geographical location and
distances from the main land of our county. The absence of all-weather road
connectivity, banking facilities, the absence of proper information and the habit of
record keeping, difficult time consuming border entry procedure of vehicles often
creates problems in free movement of goods between countries. Besides these the
practice of informal trade practices in land custom station has become challenge
for both the country.

Conclusion
Active participation in international trade is necessary for achieving rapid
economic development of a country. The trade relations between India and Bhutan
are healthy and greater chances of further improvement in coming days. Certain
bottlenecks like geographic terrain of Bhutan are itself a big challenge for
constructing modern infrastructure. Construction of modern infrastructure with
sufficient number of technical persons in the land custom station and more
digitalization procedure can further help to improve the trade volume between
countries. Besides trade, the exchange of culture and traditions, costumes and design
through government sponsoring agencies, certain trade fairs will certainly boost the
economy of both the countries.

Reference
1. Chakma, R., Yarso A. S. (2022). ―Indo-Bangladesh Border Trade with Special
Reference to Assam Border: Problems and Prospects‖, International Journal of
Mechanical Engineering, vol. 7 (5) Pp- 695-705
2. Mohan, D., Raghunath, T., Medipally, S (2017). ―Governing Dynamics of
Cross-Border Trade: A Case Study from the Indo-Bhutan Border Region.‖
South Asia Democratic Forum (SADF), No.4 Pp – 1-25
3. Uttam Lama (2023). ―India and Bhutan: Challenges and Opportunities in Cross
Border Trade‖ Working Paper no 567, Institute for Social and Economic
Change Dr. V K R V Rao Road, Nagarabhavi Post, Bangalore - 560072,
Karnataka, India
119
4. Ismail, Normaz Wana; Mahyideen, Jamilah Mohd (2015). ―The impact of
infrastructure on trade and economic growth in selected economies in Asia‖
ADBI Working Paper, No. 553, Asian Development Bank Institute (ADBI),
Tokyo.
5. Ministry of Finance (2024). ―Bridging Borders & Connecting Nations, India's
Land Custom Stations‖, Central Board of Indirect Taxes & Customs (CBIC),
Govt. of India.
6. Taneja. N., Bimal. S., Nadeem T. Roy R. (2019). ―India-Bhutan Economic
Relations‖ Indian Council for Research on International Economic Relations,
Working Paper 384
7. Das B. S., and Chattopadhyay S. (2022). ―Identifying Challenges and
Improving Trade Facilitation in the States of Northeast India‖ ADBI Working
Paper, No. 95, Asian Development Bank 6 ADB Avenue, Mandaluyong City,
1550 Metro Manila, Philippines.

120
On Some Developments in the Gamma-G
Type 2 Family of Distributions

Hmingthansanga1, Sanjeeva Kumar Jha2, and Bhanita Das3


1
North-Eastern Hill University, Shillong, Meghalaya, 793022, Email ID:
[Link]@[Link]
2
North-Eastern Hill University, Shillong, Meghalaya, 793022, Email ID:
skjha@[Link]
3*
North-Eastern Hill University, Shillong, Meghalaya, 793022, Email ID:
bhanitadas83@[Link]

Abstract:
The recent literature has seen an increase in the development of generalized
statistical models in order to improve model performance in applications in various
fields of study. The present book chapter reviews and presents some developments
in a family of generalized distributions known in the literature as the gamma-g type
2 family of distributions which is derived as a modified form of the T-X families. A
brief introduction to the development of the family followed by several models
within this family that have been proposed by different authors are presented, and
the relevance in real data applications in their respective studies is also explored.

1. Introduction
Analyzing data sets and understanding their structure and probabilistic nature is
an important practice for modern research across diverse areas of study, which
necessitates for the developments in improved probability models for various
applications. Consequently, generalized families of distributions have been of key
interest in the recent literature on statistical distributions. Based on the foundation of
earlier developments such as (Pearson 1894), (Burr 1942) and (Pareto 1896--1897),
to name a few, recent developments have shifted to the generalization of existing
distributions using mathematical methods that may involve combining distributions
or simply adding one or more parameters to existing distributions, as also stated in
(Lee, Famoye, and Alzaatreh 2013). Some examples of generalized families include
Marshall and Olkin family by (Marshall and Olkin 1997), exponentiated family
studied by (Mudholkar and Srivastava 1993) and (Gupta, Gupta, and Gupta 1998),
beta generated class by (Eugene, Lee, and Famoye 2002), transmuted family by


Corresponding author: Email ID: [Link]@[Link]

121
(Shaw and Buckley 2009) and cubic transmuted family by (Granzotto, Louzada, and
Balakrishnan 2017). These generalizations aim to increase precision of existing
models and to improve their flexibility and precision in real-world applications.

2. T-X Method of Generating New Families of Distributions


The Transformed-Transformer (T-X) method can be considered as a
generalization of the beta generated class by (Eugene, Lee, and Famoye 2002). The
beta generated class is characterized by the following cdf G(x),

(𝑥)
𝐺(𝑥) = ∫ 𝑟 (𝑡)𝑑𝑡, 𝑥∈𝐷 (2.1)
0

where 𝑟(. ) is the pdf of the two-parameter beta distribution and 𝐹(𝑥) is the cdf
of some baseline distribution. The beta Gumbel distribution by (Nadarajah and Kotz
2004) and the beta-Pareto distribution (Akinsete, Famoye, and Lee 2008) are two
examples from the beta generated class.
As presented in (Alzaatreh, Lee, and Famoye 2013), the the T-X method of
generating new families of distributions utilizes two random variables, the
transformed random variable T and the transformer random variable X to form the
T-X families. Consider the transformer random variable or baseline cdf 𝐹(𝑥) and
pdf 𝑓(𝑥), also let 𝑅(𝑡) and 𝑟(𝑡) be the cdf and pdf of the transformed random
variable T respectively. Then the T-X families of distributions has the following cdf,
𝑊, (𝑥)-
𝐺(𝑥) = ∫ 𝑟 (𝑡)𝑑𝑡, 𝑥∈𝐷 (2.2)
0
where 𝑊,𝐹(𝑥)- is a function of the baseline cdf 𝐹(𝑥), sometimes referred to as
the link function and it satisfies the following conditions:
• 𝑊,𝐹(𝑥)- ∈ ,𝑎, 𝑏- for 𝑇 ∈ ,𝑎, 𝑏-.
• 𝑊,𝐹(𝑥)- is differentiable and monotonically non-decreasing.
• 𝑊,𝐹(𝑥)- → 𝑎 as 𝑥 → −∞ and 𝑊,𝐹(𝑥)- → 𝑏 as 𝑥 → ∞.
By considering different T and X distributions with a suitable link function, new
families and sub-models can be derived from the T-X families. Implementing the
parameters from the transformed variable and the transformer variable, it can give
rise to flexible families of distributions. More details on the developments in the T-
X families can be found in works such as (Khadim et al. 2022). It is easy to see that
the T-X method generalizes the beta generated class given in Eq. (2.1), it also
generalizes some other families such as the Kumaraswamy generated class by
(Gauss M. Cordeiro and Castro 2011). This method can be used to derive various
distributions as a transformed random variable of the T distribution, resulting in a
model with additional parameters that may result in more flexibility in the functional
forms of the resulting distribution. Hence, the generalized family can be of great
utility in modeling complicated data structures with various shapes, skewness and

122
hazard rates, that we come across in real-world applications and studies. Some of the
members of the T-X families are given in Table 1.
Some T-X families of distributions.
Distribution of T 𝑊,𝐹(𝑥)- Members of T-X family
Weibull 𝐹(𝑥) Weibull-G ((Bourguignon, Silva, and
,1 − 𝐹(𝑥)- Cordeiro 2022))
logistic 𝑙𝑛*−𝑙𝑛,1 − 𝐹(𝑥)-+ logistic-X ((Tahir et al. 2016))
support ∈ (0, ∞) ,1 − 𝐹(𝑥)- Weighted T-X ((Ahmad et al. 2020))
−𝑙𝑛 8 9
𝑒 (𝑥)
1− (𝑥)
exponential(1) 𝑒 −1 new exponential-X ((Shah et al. 2022))
−𝑙𝑛 8 9
𝑒 − ,1 − 𝐹(𝑥)-
beta prime 𝐹(𝑥) odd beta prime-G ((Suleiman et al.
,1 − 𝐹(𝑥)- 2023))
Maxwell 𝐹𝛽 (𝑥) generalized odd Maxwell-G ((Ishaq et al.
,1 − 𝐹𝛽 (𝑥)- 2024))
exponentiated log 𝐹(𝑥) exponentiated log logistic ((Ihtisham et
logistic al. 2024))
exponential(1) 𝑒 − 𝐹(𝑥),𝑒 − 1 + 𝐹(𝑥)- new modified-X ((Alshawarbeh 2024))
−𝑙𝑛 8 9
𝑒

3. The Gamma-G Type 2 Family of Distributions


Ristić and Balakrishnan (2012) proposed a generator which is described using the
survival function 𝐺‾ = 1 − 𝐺(𝑥) where 𝐺(𝑥) is the cdf of this family. Taking the
gamma parameter to be 𝛿 > 0, for some baseline pdf 𝑓(𝑥) and cdf 𝐹(𝑥) this gamma
generated family has the survival function given by
−𝑙𝑛 (𝑥)
1
𝐺‾ = ∫ 𝑡 𝛿−1 𝑒 −𝑡 𝑑𝑡, 𝛿 > 0; 𝑥 ∈ (3.1)
(𝛿) 0
where (. ) is the complete gamma function.
So, the corresponding cdf is
1
𝐺(𝑥) = 1 − 𝛾(𝛿, 𝑙𝑛,*𝐹(𝑥)+−1 -), 𝛿 > 0; 𝑥 ∈ (3.2)
(𝛿)
where 𝛾(. , . ) is the lower incomplete gamma function given by
𝑥
𝛾(𝑠, 𝑥) = ∫ 𝑡 𝑠−1 𝑒 −𝑡 𝑑𝑡.
0
The corresponding pdf of this family is given by
1
𝑔(𝑥) = ,−𝑙𝑛𝐹(𝑥)-𝛿−1 𝑓(𝑥). 𝛿 > 0; 𝑥 > 0 (3.3)
(𝛿)
From the expression in Eq. (3.1), we can see that it is a modified form of the T-
X method. This family of distributions can give rise to several flexible generalized
models for some baseline distribution. Several new distributions have been proposed
in the literature using this method, which are presented in the following subsections.
Some general studies on the family can be found in works such as (Gauss M.
Cordeiro and Bourguignon 2016) and (Ghosh and Hamedani 2018).
123
Some authors may consider the two-parameter gamma distribution for this
family where the survival function of the family is given as
−𝑙𝑛 (𝑥)
1 𝛿−1 −𝜃
𝑡
𝐺‾ = ∫ 𝑡 𝑒 𝑑𝑡. 𝛿, 𝜃 > 0; 𝑥 ∈ (3.4)
(𝛿)𝜃 𝛿 0
The pdf corresponding to equation (3.4) can also be obtained.
This family is also called the Ristić –Balakrishnan family by some authors.

3.1 The Gamma-Exponentiated Exponential Distribution


The gamma-exponentiated exponential distribution (Ristić and Balakrishnan
2012) is a generalization of the baseline exponentiated exponential distribution as a
Ristić -Balakrishnan gamma generated family.
When the baseline distribution of the Ristić -Balakrishnan gamma generated
family with gamma parameter 𝛿 > 0 is taken to be the exponentiated exponential
distribution, we get the pdf of the gamma-exponentiated exponential distribution as,
𝜆𝛼 𝛿 −𝜆𝑥 𝛼−1 𝛿−1 (3.1.1)
𝑔(𝑥) = 𝑒 (1 − 𝑒 −𝜆𝑥 ) (−𝑙𝑛[1 − 𝑒 −𝜆𝑥 ]) , 𝛿, 𝜆, 𝛼 > 0; 𝑥 > 0
(𝛿)
The study demonstrates the flexibility of the new model by using two real-life
data sets of number of successive failures and survival times.

3.2 The Gamma-Exponentiated Weibull Distribution


The gamma-exponentiated Weibull distribution by (Gustavo, Pinho, and
Cordeiro 2012) is a generalization of the exponentiated Weibull distribution. Taking
𝛿 > 0 to be the gamma parameter, the pdf of the gamma-exponentiated Weibull
distribution is given by,
𝑘𝛼 𝛿 𝑥 𝑘−1 𝑥 𝑘 𝑥 𝑘 𝛼−1
𝑔(𝑥) = . / 𝑒𝑥𝑝 {− . / } [1 − 𝑒𝑥𝑝 {− . / }]
𝜆 (𝛿) 𝜆 𝜆 𝜆
(3.2.1)
𝑥 𝑘 𝛿−1
× {−𝑙𝑛 [1 − 𝑒𝑥𝑝 {− . / }]} , 𝛿, 𝑘, 𝛼, 𝜆 > 0; 𝑥 > 0
𝜆

The real data set application is demonstrated using data set of daily minimum
wind speed.
A same study is also done by (Castellares and Lemonte 2016) where the same
distribution is referred to as the gamma dual Weibull distribuiton and a convergent
expansion of the density is derived. Real data demonstration is performed using a
data set of remission times of bladder cancer patients.

3.3 The Gamma Log-Logistic Weibull Distribution


The gamma log-logistic Weibull distribution proposed and studied by (Foya et
al. 2017) is a generalization of the log-logistic Weibull distribution as a member of
The Ristić and Balakrishnan gamma generated family. Taking the gamma
parameter to be 𝛿 > 0 and > 0 , we have the pdf of the model as follows,
124
1 𝛽
𝑔(𝑥) = 𝛿
(1 + 𝑥 𝑐 )−1 𝑒 −𝛼𝑥 [(1 + 𝑥 𝑐 )−1 𝑐𝑥 𝑐−1 + 𝛼𝛽𝑥 𝛽−1 ]
(𝛿)𝜃
𝛽 𝛿−1
× .−𝑙𝑛 01 − (1 + 𝑥 𝑐 )−1 𝑒 −𝛼𝑥 1/ (3.3.1)
𝛽 1/𝜃−1
× 01 − (1 + 𝑥 𝑐 )−1 𝑒 −𝛼𝑥 1 . 𝑐, 𝛼, 𝛽, 𝛿, 𝜃 > 0; 𝑥 > 0
The proposed distribution has several existing distributions as sub-models. The
flexibility of the distribution is demonstrated by using strengths of glass fibers data
set.

3.4 The Ristić and Balakrishnan Lindley-Poisson Distribution


The Ristić and Balakrishnan Lindley-Poisson distribution proposed and studied
by (Fagbamigbe et al. 2018) is a generalization of the Lindley-Poisson distribution.
For the Lindley-Poisson baseline distribution, consider 𝛿 > 0 to be the gamma
parameter, we have the pdf of the Ristić and Balakrishnan Lindley-Poisson
distribution as follows,
𝛿−1
1 + 𝜃 + 𝜃𝑥
1 1 − 𝑒𝑥𝑝 2𝜆 .1 − 1 + 𝜃 / 𝑒 −𝜃𝑥 3
𝑔(𝑥) = :−𝑙𝑛 < =;
(𝛿) 1 − 𝑒𝜆
(3.4.1)
2 (1 −𝜃𝑥 1 + 𝜃 + 𝜃𝑥 −𝜃𝑥
𝜆𝜃 + 𝑥)𝑒 𝑒𝑥𝑝 2𝜆 .1 − 1 + 𝜃 / 𝑒 3
× . 𝛿, 𝜃, 𝜆 > 0; 𝑥 > 0
(1 + 𝜃)(𝑒 𝜆 ) − 1

The flexibility of the proposed model is demonstrated using some real-life


failure times data set.

3.5 The Ristić -Balakrishnan Extended Exponential Distribution


The Ristić -Balakrishnan extended exponential distribution by (Silva, Andrade,
and Bourguignon 2018) is a generalization of the extended exponential distribution.
With the gamma parameter 𝑎 > 0, it has the following pdf,

𝛼 2 (1 + 𝛽𝑥)𝑒 −𝛼𝑥 𝑎−1


𝑔(𝑥) = 𝜌 (𝑥), 𝑎, 𝛼, 𝛽 > 0; 𝑥 > 0 (3.5.1)
(𝛼 + 𝛽) (𝑎)
Where, 𝜌(𝑥) = 𝑙𝑛(𝛼 + 𝛽) − 𝑙𝑛,𝛼 + 𝛽 − (𝛼 + 𝛽 + 𝛼𝛽𝑥)𝑒 −𝛼𝑥 -
A real-life data application of the proposed model is demonstrated using a
failure times data set.

3.6 The Gamma Log-Logistic Erlang Truncated Exponential Distribution


The gamma log-logistic Erlang truncated exponential distribution by (Jimoh et
al. 2019) it is a generalized model of lifetime of components that follow the log-

125
logistic and Erlang-turncated exponential distributions. With the gamma parameter
𝛿 > 0, the model has the following pdf,
1 −𝜆 𝛿−1
𝑔(𝑥) = .−𝑙𝑛 01 − (1 + 𝑥 𝑐 )−1 𝑒 −𝛽(1−𝑒 )𝑥 1/
(𝛿)
−𝜆 (3.6.1)
𝑒 −𝛽(1−𝑒 )𝑥
× {𝑐𝑥 𝑐−1 + (1 + 𝑥 𝑐 )𝛽(1 − 𝑒 −𝜆 )}. 𝑐, 𝛽, 𝜆, 𝛿 > 0; 𝑥 > 0
(1 + 𝑥 𝑐 )2
The proposed distribution is expected to be very flexible and it has several
existing distributions as special cases. In the study, the flexibility of the model is
demonstrated using real-life data sets of failure times and waiting times.

3.7 The Gamma Odd Burr III-G Family


The gamma odd Burr III-G family proposed and studied by (Peter et al. 2021) is
an extension of the odd Burr III-G family as a Ristić -Balakrishnan gamma
generated family. For gamma parameter 𝛿 > 0, and some baseline pdf 𝑓(𝑥) and cdf
𝐹(𝑥), the family has the following pdf,

𝛼 𝛿−1
𝛼𝛽 𝛿 ,1 − 𝐹(𝑥)-𝛼−1 1 − 𝐹(𝑥)
𝑔(𝑥) = 8𝑙𝑛 61 + 4 5 79
(𝛿) ,𝐹(𝑥)- 𝛼+1 𝐹(𝑥)
𝛼 −𝛽−1
(3.7.1)
1 − 𝐹(𝑥)
× 61 + 4 5 7 𝑓(𝑥), 𝛿, 𝛼, 𝛽 > 0; 𝑥 > 0
𝐹(𝑥)
The study proposed and studied some flexible distributions from this family, viz.
the gamma odd Burr III-log logistic distribution, the gamma odd Burr III-log
Kumaraswamy distribution and the gamma odd Burr III-beta distribution along with
their real data set applications.

3.8 The Gamma Odd Burr X-G Family of Distributions


The gamma odd Burr X-G family of distributions by (Tlhaloganyang, Sengweni,
and Oluyede 2022) is an extension of the odd Burr X-G family as the Ristić -
Balakrishnan gamma generated family. Let 𝛿 > 0 be the gamma parameter, then for
some baseline pdf f(x) and cdf F(x), we have the pdf of the gamma odd Burr X-G
family as,
2
2𝜃𝑓(𝑥)𝐹(𝑥) 𝐹(𝑥)
𝑔(𝑥) = 𝑒𝑥𝑝 {− 4 5 }
(𝛿),1 − 𝐹(𝑥)-3 ,1 − 𝐹(𝑥)-
2 𝜃−1
𝐹(𝑥)
× {1 − 𝑒𝑥𝑝 [− 4 5 ]} (3.8.1)
,1 − 𝐹(𝑥)-
2 𝛿−1
𝐹(𝑥)
× [−𝜃𝑙𝑛 {1 − 𝑒𝑥𝑝 [− 4 5 ]}] . 𝛿, 𝜃 > 0; 𝑥 > 0
,1 − 𝐹(𝑥)-
This family can give rise to different flexible distributions for different choices
of the baseline distribution.

126
The study proposed the gamma odd Burr X-Weibull distribution, the gamma
odd Burr X-log-logistic distribution and the gamma odd Burr X-uniform distribution
as distributions form this family. A real data set application of the gamma odd Burr
X-Weibull distribution is also demonstrated using taxes revenue and active repair
times data sets in the referenced study.
3.9 The Ristić -Balakrishnan-Topp-Leone-Gompertz-G Family of Distributions
The Ristić -Balakrishnan-Topp-Leone-Gompertz-G family of distributions
proposed and studied by (Pu, Moakofi, and Oluyede 2023) is a generalization of the
Topp-Leone-Gompertz-G family. For gamma parameter 𝛿 > 0, any baseline pdf
𝑓(𝑥) and corresponding cdf 𝐹(𝑥), it has the following pdf,
𝑏 𝛿−1
2𝑏 2 −𝜃
𝑔(𝑥) = ,1
4−𝑙𝑛 [1 − 𝑒𝑥𝑝 { (1 − − 𝐹(𝑥)- )}] 5
(𝛿) 𝜃
𝑏−1
2
× [1 − 𝑒𝑥𝑝 { (1 − ,1 − 𝐹(𝑥)-−𝜃 )}] (3.9.1)
𝜃
2
× 𝑒𝑥𝑝 { (1 − ,1 − 𝐹(𝑥)-−𝜃 )}
𝜃
× ,1 − 𝐹(𝑥)-−𝜃−1 𝑓(𝑥). 𝛿, 𝑏, 𝜃 > 0; 𝑥 > 0
As special cases of this distribution, the Ristić -Balakrishnan-Topp-Leone-
Gompertz-Burr XII distribution, the Ristić -Balakrishnan-Topp-Leone-Gompertz-
Weibull distribution and the Ristić -Balakrishnan-Topp-Leone-Gompertz-Uniform
distribution are also studied. Ristić -Balakrishnan-Topp-Leone-Gompertz-log-
logistic distribution as a special case of the Ristić -Balakrishnan-Topp-Leone-
Gompertz-Burr XII distribution is used for modeling survival times of guinea pigs
injected with tubercle bacilli and active repair times for airborne communication
transceivers data.

3.10 The Gamma-Topp-Leone-Type II-Exponentiated Half Logistic-G Family of


Distributions
The gamma-Topp-Leone-type II-exponentiated half logistic-G family of
distributions by (Oluyede and Moakofi 2023) is an extension of the Topp-Leone-
type II-exponentiated half logistic-G family as a member of the Ristić and
Balakrishnan gamma generated family. With a baseline Topp-Leone-type II-
exponentiated half logistic-G family with parameters 𝑏 > 0 and 𝑎 > 0, and the
gamma parameter 𝛿 > 0, we have the pdf of the gamma-Topp-Leone-type II-
exponentiated half logistic-G family as follows,
4𝑎𝑏 𝛿−1
𝑔(𝑥) = [−𝑙𝑛(,1 − (𝑥)-𝑏 )]
(𝛿)
× ,1 − (𝑥)-𝑏−1 ,1 − 𝐹(𝑥)-2𝑎−1 (3.10.1)
𝑓(𝑥)
× , 𝑎, 𝑏, 𝛿 > 0; 𝑥 > 0
,1 + 𝐹(𝑥)-2(𝑎+1)−1
127
1− (𝑥) 2𝑎
where (𝑥) = . / and 𝐹(𝑥) and 𝑓(𝑥) are the baseline cdf and pdf
1+ (𝑥)
respectively.
Some special cases viz. the gamma-Topp-Leone-type II-exponentiated half
logistic-Weibull distribution, the gamma-Topp-Leone-type II-exponentiated half
logistic-log logistic distribution and the gamma-Topp-Leone-type II-exponentiated
half logistic-Kumaraswamy distribution are presented and the application of the
gamma-Topp-Leone-type II-exponentiated half logistic-Weibull distribution is
demonstrated using COVID-19, vehicle fatalities and remission times data sets.

4 Conclusion
As presented in the chapter, several distributions have been derived from the
gamma-g type 2 family by different authors, exploring their properties and
applications. We can see from the studies that the models have improved fit on
various real data sets like lifetime and waiting time data sets. The improvement in
the model flexibility comes with a price of complexity in the model functional forms
like the pdf and cdf, so, generalized distributions are often implemented using
numerical optimization methods using software packages in R, python, matlab, etc.
Different methods of numerical computation-based inference can be explored for
such models. Various mathematical methods can also be explored to further improve
the models and also to develop better inferential methods. The models presented can
be further studied from their respective references.

REFERENCES
Ahmad, Zubair, Eisa Mahmoudi, Sanku Dey, and Saima K Khosa. 2020.
―Modeling Vehicle Insurance Loss Data Using a New Member of t-x Family
of Distributions.‖ Journal of Statistical Theory and Applications 19 (2):
133–47. [Link]
Akinsete, Alfred, Felix Famoye, and Carl Lee. 2008. ―The Beta-Pareto
Distribution.‖ Statistics 42 (6): 547–63. [Link]
880801983876.
Alshawarbeh, Etaf. 2024. ―A New Modified-x Family of Distributions with
Applications in Modeling Biomedical Data.‖ Alexandria Engineering
Journal 93: 189–206. [Link]
pii/S1110016824002266.
Alzaatreh, Ayman, Carl Lee, and Felix Famoye. 2013. ―A New Method for
Generating Families of Continuous Distributions.‖ Metron 71 (1): 63–79.
[Link]
Bourguignon, Marcelo, Rodrigo B. Silva, and Gauss M. Cordeiro. 2022. ―The
Weibull-g Family of Probability Distributions.‖ Journal of Data Science 12
(1): 53–68. [Link] .
Burr, Irving W. 1942. ―Cumulative Frequency Functions.‖ The Annals of
Mathematical Statistics 13 (2): 215–32.
128
Castellares, Fredy, and Artur J Lemonte. 2016. ―On the Gamma Dual Weibull
Model.‖ American Journal of Mathematical and Management Sciences 35
(2): 124–32. [Link]
Cordeiro, Gauss M, and Marcelo Bourguignon. 2016. ―New Results on the
Ristić –Balakrishnan Family of Distributions.‖ Communications in
Statistics-Theory and Methods 45 (23): 6969–88.
[Link] 10926.2014.972573.
Cordeiro, Gauss M., and Mário de Castro. 2011. ―A New Family of Generalized
Distributions.‖ Journal of Statistical Computation and Simulation 81 (7):
883–98.
Eugene, Nicholas, Carl Lee, and Felix Famoye. 2002. ―Beta-Normal
Distribution and Its Applications.‖ Communications in Statistics-Theory and
Methods 31 (4): 497–512. [Link]
Fagbamigbe, Adeniyi Francis, Pinkie Melamu, Broderick Olusegun Oluyede,
and Boikanyo Makubate. 2018. ―The Ristić and Balakrishnan Lindley-
Poisson Distribution: Model, Theory and Application.‖ Afrika Statistika 13
(4): 1837–64. [Link]
Foya, Susan, Broderick O Oluyede, Adeniyi F Fagbamigbe, and Boikanyo
Makubate. 2017. ―The Gamma Log-Logistic Weibull Distribution: Model,
Properties and Application.‖ Electronic Journal of Applied Statistical
Analysis 10 (1): 206–41. [Link]
Ghosh, Indranil, and Gholamhossein Hamedani. 2018. ―On the Ristic—
Balakrishnan Distribution: Bivariate Extension and Characterizations.‖
Journal of Statistical Theory and Practice 12: 436–49. [Link]
10.48550/arXiv.1711.00158.
Granzotto, DCT, Francisco Louzada, and N Balakrishnan. 2017. ―Cubic Rank
Transmuted Distributions: Inferential Issues and Applications.‖ Journal of
Statistical Computation and Simulation 87 (14): 2760–78. [Link]
10.1080/00949655.2017.1344239.
Gupta, Ramesh C., Pushpa L Gupta, and Rameshwar D. Gupta. 1998.
―Modeling Failure Time Data by Lehman Alternatives.‖ Communications in
Statistics - Theory and Methods 27 (4): 887–904. [Link]
10.1080/03610929808832134.
Gustavo, Luis, B. Pinho, and Gauss Moutinho Cordeiro. 2012. ―The Gamma-
Exponentiated Weibull Distribution.‖ Journal of Statistical Theory and
Applications 11 (4): 379–95. [Link]
ID:14223003.
Ihtisham, Shumaila, Sadaf Manzoor, Alamgir, Osama Abdulaziz Alamri, and
Muhammad Nouman Qureshi. 2024. ―Pareto Exponentiated Log-Logistic
Distribution (PELL) with an Application to Covid-19 Data.‖ AIP Advances
14 (1): 015052. [Link]
Ishaq, Aliyu Ismail, Uthumporn Panitanarak, Alfred Adewole Abiodun, Ahmad
Abubakar Suleiman, and Hanita Daud. 2024. ―The Generalized Odd
129
Maxwell-Kumaraswamy Distribution: Its Properties and Applications.‖
Contemporary Mathematics 5 (1): 711–42. [Link]
[Link]/CM/article/view/2888.
Jimoh, Hameed, Broderick Olusegun OLUYEDE, Divine Wanduku, and
Boikanyo Makubate. 2019. ―The Gamma Log-Logistic Erlang Truncated
Exponential Distribution with Applications.‖ Afrika Statistika 14 (4): 2141–
64. [Link]
Khadim, Aneeqa, Aamir Saghir, Tassadaq Hussain, and Mohammad Shakil.
2022. ―Some New Developments and Review on TX Family of
Distributions.‖ Journal of Statistics Applications & Probability 11 (3): 739–
57. [Link]
Lee, Carl, Felix Famoye, and Ayman Y. Alzaatreh. 2013. ―Methods for
Generating Families of Univariate Continuous Distributions in the Recent
Decades.‖ WIREs Computational Statistics 5 (3): 219–38. [Link]
[Link]/doi/abs/10.1002/wics.1255.
Marshall, Albert W., and Ingram Olkin. 1997. ―A New Method for Adding a
Parameter to a Family of Distributions with Application to the Exponential
and Weibull Families.‖ Biometrika 84 (3): 641–52. [Link]
stable/2337585.
Mudholkar, Govind S, and Deo Kumar Srivastava. 1993. ―Exponentiated
Weibull Family for Analyzing Bathtub Failure-Rate Data.‖ IEEE
Transactions on Reliability 42 (2): 299–302. [Link]
229504.
Nadarajah, Saralees, and Samuel Kotz. 2004. ―The Beta Gumbel Distribution.‖
Mathematical Problems in Engineering 2004 (4): 323–410.
[Link]
Oluyede, Broderick, and Thatayaone Moakofi. 2023. ―The Gamma-Topp-
Leone-Type II-Exponentiated Half Logistic-g Family of Distributions with
Applications.‖ Stats 6 (2): 706–33. [Link]
Pareto, Vilfredo. 1896--1897. Cours d‘Économie Politique Professée à
l‘université de Lausanne. Vol. 1–2. Lausanne, Switzerland: F. Rouge.
[Link]
Pearson, Karl. 1894. ―Contributions to the Mathematical Theory of Evolution.‖
Philosophical Transactions of the Royal Society of London. A 185: 71–110.
Peter, Peter Ookame, Broderick Oluyede, Huybrechts Frazier Bindele,
Nkumbuludzi Ndwapi, and Onkabetse Mabikwa. 2021. ―The Gamma Odd
Burr III-g Family of Distributions: Model, Properties and Applications.‖
Revista Colombiana de Estad stica 44 (2): 331–68.
[Link] rce.v44n2.89320.
Pu, Shusen, Thatayaone Moakofi, and Broderick Oluyede. 2023. ―The Ristić –
Balakrishnan–Topp–Leone–Gompertz-g Family of Distributions with
Applications.‖ Journal of Statistical Theory and Applications 22 (1): 116–
50. [Link]
130
Ristić , Miroslav M, and Narayanaswamy Balakrishnan. 2012. ―The Gamma-
Exponentiated Exponential Distribution.‖ Journal of Statistical Computation
and Simulation 82 (8): 1191–1206. [Link]
2011.574633.
Shah, Zubir, Amjad Ali, Muhammad Hamraz, Dost Muhammad Khan, Zardad
Khan, M. EL-Morshedy, Afrah Al-Bossly, and Zahra Almaspoor. 2022. ―A
New Member of t-x Family with Applications in Different Sectors.‖ Journal
of Mathematics 2022 (1): 1453451. [Link]
abs/10.1155/2022/1453451.
Shaw, William T., and Ian R. C. Buckley. 2009. ―The Alchemy of Probability
Distributions: Beyond Gram-Charlier Expansions, and a Skew-Kurtotic-
Normal Distribution from a Rank Transmutation Map.‖ [Link]
abs/0901.0434.
Silva, Frank Gomes, Thiago Alexandro Nascimento de Andrade, and Marcelo
Bourguignon. 2018. ―Ristić -Balakrishnan Extended Exponential
Distribution.‖ Acta Scientiarum. Technology 40. [Link]
actascitechnol.v40i1.34963.
Suleiman, Ahmad Abubakar, Hanita Daud, Mahmod Othman, Aliyu Ismail
Ishaq, Rachmah Indawati, Mohd Lazim Abdullah, and Abdullah Husin.
2023. ―The Odd Beta Prime-g Family of Probability Distributions:
Properties and Applications to Engineering and Environmental Data.‖
Computer Sciences &Amp; Mathematics Forum 7 (1). [Link]
com/2813-0324/7/1/20.
Tahir, M. H., Gauss M. Cordeiro, Ayman Alzaatreh, M. Mansoor, and M.
Zubair and. 2016. ―The Logistic-x Family of Distributions and Its
Applications.‖ Communications in Statistics - Theory and Methods 45 (24):
7326–49. [Link]
Tlhaloganyang, Bakang Percy, Whatmore Sengweni, and Broderick Oluyede.
2022. ―The the Gamma Odd Burr XG Family of Distributions with
Applications.‖ Pakistan Journal of Statistics and Operation Research, 721–
46. [Link]

131
Market Forecasting Using Stochastic Process: A Study on
Bharat Heavy Electricals Limited
Ronit Paul 1, Dr. Tanusree Deb Roy 2
*1
Research Scholar, Assam University, Silchar, Department of Statistics, Assam,
India, paulronit111@[Link]
2
Assistant Professor, Assam University, Silchar, Department of Statistics, Assam,
India, [Link]@[Link]

Abstract:
Geometric Brownian Motion (GBM) and Hidden Markov Models (HMM) are
the two main predictive modelling approaches that are examined in this study in
relation to stock price forecasts. Based on Bharat Heavy Electricals Limited (BHEL)
historical stock price data from July 2023 to March 2024, the study evaluates these
models' effectiveness. Using independent and unpredictable stock price fluctuations
as its foundation, GBM is based on stochastic processes. HMM is superior at
managing the uncertainty associated with financial markets because it can identify
changes in the market regime and temporal dependencies. Root Mean Square Error
(RMSE), Average Absolute Error (AAE), and Absolute Percentage Error (APE) are
the three performance indicators that the study uses to assess both models. Results
demonstrate that HMM performs better than GBM, producing predictions that are
more dependable and accurate for all stock price categories. Predictive accuracy is
improved by the HMM model's flexibility in changing market situations, especially
in dynamic contexts. GBM is less good at managing closing price predictions, even
if it is still beneficial for modelling continuous price movements. To further the
increase forecast accuracy, future research may combine machine learning methods
with GBM and HMM. According to the study's findings, GBM is still useful in some
circumstances but HMM is the better model for stock price prediction because of its
higher accuracy in capturing market dynamics.
Key Words: Predictive Modelling, Geometric Brownian Motion, Hidden Markov
Model, Stock Price Forecasting, Financial Market Dynamics

1. INTRODUCTION
Stocks, or equities, represent ownership in a company, granting holders a
share of profits and assets. Stock trading involves buying and selling these shares,
with prices influenced by factors such as company performance, economic
conditions, and market sentiment. Predictive modeling, critical for maximizing


Corresponding author: Email ID: paulronit111@[Link]

132
returns and minimizing risks, helps investors analyze stock movements. Traditional
models like ARIMA, though widely used, have limitations in non-stationary data.
Geometric Brownian Motion (GBM) models stock prices as stochastic processes,
while Hidden Markov Models (HMMs) excel at capturing market regime shifts and
uncertainties. HMMs, combined with machine learning techniques like neural
networks and support vector machines, improve forecast accuracy by handling
complex, non-linear relationships in financial data. GBM remains a fundamental
tool, but HMMs are particularly useful for predicting market dynamics due to their
state-transition framework. Together, these models provide valuable insights into
stock price forecasting.

2. LITERATURE REVIEW
Predictive modeling in stock market forecasting has seen significant
advancements from the early days of neural network models by White (1988) to the
more sophisticated fusion models by Hassan et al. (2007). These models have
evolved to incorporate a variety of techniques including genetic algorithms, neural
networks, and hidden Markov models (HMMs) to enhance predictive accuracy.
Predictive modeling is foundational for understanding the evolution and
advancements in stock market forecasting methodologies (White, 1988; Hassan et
al., 2007).
The accuracy and computational complexity of the IBCO-BP model are higher
than those of the other back propagation models. According to Khashei and Bijari
(2010), a neural network's performance isn't good enough for some real-time data.
As a result, they proposed a brand-new hybrid kind of artificial neural network
employing ARIMA models. When compared to the neural artificial network model
alone, the suggested approach yielded superior predictions for three different real
datasets. The ARIMA model stock index forecasting and the back propagation neural
network model were compared by Yao et al. (1999).
Geometric Brownian Motion (GBM) has been a fundamental tool in financial
modeling, particularly for predicting stock prices. It is favoured for its ability to
model stock prices as a stochastic process, reflecting the random nature of financial
markets. Agustini et al. (2018) developed a prediction model using Brownian motion
for various stock indices within the Jakarta Corporate Index. Their model
demonstrated high accuracy, with a mean absolute percentage error (MAPE) of less
than 20%. Similarly, Rathnayaka et al. (2014) compared GBM with ARIMA models
using data from the Colombo Stock Exchange (CSE) in Sri Lanka, finding that
GBM provided more significant forecasts than traditional models. Hidden Markov
Models (HMMs) have gained popularity in financial forecasting due to their ability
to capture the dynamics of regime transitions in financial markets. HMMs can model
different market regimes, characterized by unique statistical properties such as
changes in volatility and correlation structures. Hamilton (1989) was instrumental in
introducing regime-switching models, laying the groundwork for integrating regime
shifts into risk prediction models.
133
Rabiner (1989) emphasized the use of the Viterbi Algorithm (Hidden Markov
Model) alongside GBM for stock prediction. The Viterbi Algorithm is effective in
determining the most probable sequence of hidden states that account for observed
events, providing valuable insights into underlying market conditions. This
combined approach captures the complex dynamics of the stock market more
effectively than using GBM alone. By modeling the observed stock prices using
GBM and inferring the hidden market states with the Viterbi Algorithm within the
HMM framework, the approach significantly enhances stock price predictions.
Among the popular methods for predicting market risk is the application of
hidden Markov models (HMMs). HMMs can capture the dynamics of regime
transitions in financial markets, making them increasingly popular. The groundwork
for integrating regime shifts into risk prediction models was laid by Hamilton
(1989), who proposed regime-switching models. Different market regimes may be
distinguished using these models based on unique statistical characteristics such as
changes in correlation structures and volatility clustering. These models allow for a
more nuanced understanding of market behavior, accommodating the complexity
and unpredictability of financial markets. Additionally, HMMs provide a framework
for probabilistically determining the likelihood of transitioning between different
market states, enhancing the accuracy and robustness of risk predictions.
Hence, predictive modeling in stock market forecasting has evolved
significantly, incorporating techniques like genetic algorithms, neural networks, and
hidden Markov models (HMMs) to improve accuracy. Despite advancements,
challenges remain in neural network performance for real-time data, leading to the
development of hybrid models such as ARIMA-ANN. Geometric Brownian Motion
(GBM) is a fundamental tool for modeling stock prices, offering high accuracy in
various studies. HMMs have gained popularity for capturing regime transitions in
financial markets, enhancing the predictive capabilities of models. The Viterbi
Algorithm within HMMs further refines stock price predictions by determining the
most probable sequence of hidden states, reflecting underlying market conditions.
These advancements allow for a more nuanced understanding of market behavior,
accommodating its complexity and unpredictability.

3. METHODOLOGY AND DATA SOURCE


The methodology section contains basic four subsections. The first subsection
describes the data that were used to build the models. Each subsection outlines the
general theories and procedures for constructing the models, followed by a detailed
description of how these models were specifically fitted to the given dataset. The
overall performance of each of the models was checked by the analysis of the
residuals and four different error measures, namely the absolute percentage error
(APE), the average absolute error (AAE) and the root-mean-square error (RMSE)
134
(Nguyet Nguyen and Wakefield 2018). The formula to calculate these errors are as
follows:
𝐴𝑃𝐸
𝑁
1 𝑟𝑖 − 𝑟̅𝑖
= ∑ (1)
𝑟̅ 𝑁
𝑖=1
This metric measures the average percentage error between the predicted and
actual values. It normalizes the error by the mean of the actual values, providing a
relative measure of the prediction accuracy (Makridakis et al., 1998).
𝐴𝐴𝐸
𝑁
𝑟𝑖 − 𝑟̅𝑖
=∑ (2)
𝑁
𝑖=1
This metric calculates the mean of the absolute differences between predicted
and actual values. It provides an average measure of the magnitude of the errors
without considering their direction (Hyndman & Koehler, 2006).
𝑅𝑀𝑆𝐸
𝑁
1 𝑟𝑖 − 𝑟̅𝑖
= √ ∑ (3)
𝑁 𝑁
𝑖=1

This metric measures the square root of the average squared differences between
predicted and actual values. RMSE gives a higher weight to larger errors, making it
sensitive to outliers, and is useful for understanding the model's accuracy in terms of
the magnitude of errors (Chai & Draxler, 2014).

Data Source
In the energy and infrastructure industries, Bharat Heavy Electricals Limited
(BHEL) is one of the biggest engineering and manufacturing companies in India.
Being owned by the Indian government, it is a public sector undertaking (PSU).
BHEL was founded in 1964 and is a key player in the economic growth of India,
particularly in the power sector where it produces equipment for power production
such as boilers, generators, and turbines (Bharat Heavy Electricals Limited, dated
06/10/2024.). BHEL is a major player in the Indian market, having made a
substantial contribution to the nation's infrastructure and power generation
capacities. It has played a significant role in advancing India's production of large
electrical equipment. The fact that the Indian government is a significant shareholder
highlights BHEL's strategic significance to the Indian economy (Ministry of Heavy
Industries & Public Enterprises, dated 06/10/2024.).
The period of data gathering spans from July 14, 2023, to March 31, 2024, the
launch day of Chandrayaan-3. This period is notable because of the developments
around Bharat Heavy Electricals Limited (BHEL), which collaborates closely with
the Indian Space Research Organization (ISRO). The data that has been collected
135
over the course of several months is crucial for evaluating how advances pertaining
to space impact BHEL's commercial operations and strategic objectives, therefore
illustrating the link between India's space achievements and its broader industrial
undertakings. During this time, BHEL and ISRO's history underwent a significant
turning point, and their contributions to the development of India's space technology
and engineering industries are highlighted.
Financial market websites like Yahoo Finance provide historical and current
financial data for BHEL, including stock performance, for financial analysis, market
research, and scholarly reasons. This website offers extensive financial data, which
is necessary for carrying out in-depth financial research and comprehending market
patterns pertaining to BHEL. The data includes daily closing share prices, market
capitalization, and financial reports (Yahoo Finance, dated 06/10/2024.).

Geometric Brownian Motion


Analyzing the literature, the Geometric Brownian Motion was frequently used
by the researchers to model financial market risks and predict better estimates in the
presence of market volatility.
A process that generates some outcomes which are time-dependent but cannot
be said ahead of time is known as a stochastic process. A stochastic process
*𝑊(𝑡): 0 ≤ 𝑡 ≤ 𝑇+ is a standard Brownian motion on ,0, 𝑇- if
i. 𝑊(0) = 0
ii. It has independent increments. That is, for any 𝑡1 , 𝑡2 , … , 𝑡𝑛 , 𝑊(𝑡2 ) −
𝑊(𝑡1 ), 𝑊(𝑡3 ) − 𝑊(𝑡2 ), … , 𝑊(𝑡𝑛 ) − 𝑊(𝑡𝑛−1 ) are independent random
variables.
iii. For every 0 ≤ 𝑡 ≤ 𝑇, 𝑊(𝑡) − 𝑊(𝑠)~𝑁(0, 𝑡 − 𝑠).
A stochastic process *𝑋(𝑡): 0 < 𝑡 < 𝑇+ is said to be a general Brownian motion
𝑋(𝑡)−𝜇𝑡
with a drift parameter μ and diffusion coefficient 𝜍 2 if is a standard
𝜍
2
Brownian motion, 𝑊(𝑡) which can be written as 𝑋(𝑡)~𝐵𝑀(𝜇, 𝜍 ).
The general Brownian motion still follows first two properties of the standard
Brownian motion. However, the third property is modified as
𝑋(𝑡) − 𝑋(𝑠)~𝑁(𝜇(𝑡 − 𝑠)), 𝜍 2 (𝑡 − 𝑠)) for any 0 ≤ 𝑠 < 𝑡 ≤ 𝑇.

Geometric Brownian Motion (GBM) Model


If 𝑋(𝑡)~𝐵𝑀(𝜇, 𝜍 2 ) then X(t) satisfies the stochastic differential equation (Yang
and Aldous 2015)
𝑑𝑋(𝑡)
= 𝜇𝑡
+ 𝜍𝑑𝑊(𝑡), (4)
where, 𝑊(𝑡) is the standard Brownian motion or Wiener process. If the
stochastic process is defined as 𝑋(𝑡) = log 𝑆(𝑡) then 𝑑𝑆(𝑡) = 𝜇𝑆(𝑡)𝑑𝑡 +
𝜍𝑆(𝑡)𝑑𝑊(𝑡) is the stochastic differential equation for the stock price random
process.
136
For a given time 𝑡 > 0, the standard model for stock price prediction can be
given from the stochastic differential equation by integration
𝑡 𝑡
𝑆(𝑡) = 𝑆(0) + 𝜇 ∫ 𝑆(𝑟)𝑑𝑟 + 𝜍 ∫ 𝑆(𝑟)𝑑𝑊(𝑟) (5)
0 0
A more explicit formula can be derived using Ito‘s formula (Ševcovic et al.
2011) to the function 𝐹(log 𝑆(𝑡), 𝑡)
𝑑𝐹 =
𝑚
  1 
0 𝑡 + 𝜇  𝑆(𝑡)
+ 2 𝜍 2  𝑚 𝑆(𝑡)1 𝑑𝑡 +

.𝜍 / 𝑑𝑊(𝑡), (𝑎)
 𝑆(𝑡)
which results,
1 1 −1
𝑑 log 𝑆(𝑡) = 𝑑𝑆(𝑡) + { 2 } *𝑑𝑆(𝑡)+2
𝑆(𝑡) 2 𝑠 (𝑡)
1 −1
= 𝜇𝑑𝑡 + 𝜍𝑑𝑊(𝑡) + 2 *𝜇𝑆(𝑡)𝑑𝑡 + 𝜍𝑆(𝑡)𝑑𝑊(𝑡)+2
2 𝑠 (𝑡)
1 2
= (𝜇 − 𝜍 * 𝑑𝑡 + 𝜍𝑑𝑊(𝑡)
2
For time 𝑡 > 0, the differential can be written as
1
log 𝑆(𝑡) = log 𝑆(0) + (𝜇 − 𝜍 2 * 𝑡 + 𝜍𝑊(𝑡)
2
1
.𝜇− 𝜍 𝑚 /𝑑𝑡+𝜍𝑊(𝑡)
𝑂𝑟, 𝑆(𝑡) = 𝑆(0)𝑒 2 (6)

2.2.2 Geometric Brownian Motion GBM(𝛍, 𝛔𝟐 ) Simulation:


For a given time set 𝑡0 = 0 < 𝑡1 < 𝑡2 < ⋯ < 𝑡𝑛 , the stock price 𝑆(𝑡) at time
𝑡0 , 𝑡1 , … , 𝑡𝑛 can be generated by
1 𝑚
.𝜇− 𝜍 /(𝑡𝑖+𝑙 −𝑡𝑖 )+𝜍
√(𝑡𝑖+𝑙 −𝑡𝑖 ). 𝑍𝑖+𝑙
𝑆(𝑡𝑖+1 ) = 𝑆(𝑡𝑖 )𝑒 2 , (7)
where, 𝑍1 , 𝑍2 , … , 𝑍𝑖 are independent and identically distributed standard normal
and 𝑖 = ̅̅̅̅̅̅̅̅̅̅̅̅
0, (𝑛 − 1). In this case, the time interval 𝑡𝑖+1 − 𝑡𝑖 = 1 for all 𝑖 =
̅̅̅̅̅̅̅̅̅̅̅̅
0, (𝑛 − 1), since the task is to predict the next-day price, the model becomes;
1
.𝜇− 𝜍 𝑚 /+𝜍𝑍
𝑆(𝑡𝑖+1 ) = 𝑆(𝑡𝑖 )𝑒 2 𝑖+𝑙
(8)
Using the model, a large number of prices are simulated, and the average of
these simulations is taken to predict the next-day price. A total of 10 predictions
have been made using this model. A fixed window of 174 past observed stock prices
have been used to predict each of the next-day prices. As a result, the window's end
price was updated with the real price and the training dataset was transferred. In a
subsequent part, the model's diagnostics and outcomes are covered.

Viterbi Algorithm
The Viterbi Algorithm has seen limited application in financial market risk
analysis. This study aims to identify a more effective stochastic model for assessing
137
financial market risks and to provide improved predictions in the presence of market
volatility. By exploring the capabilities of the Viterbi algorithm within this context,
the research seeks to enhance the accuracy of risk predictions and offer more reliable
estimates under volatile market conditions.
The Viterbi Algorithm was proposed in 1967 (Viterbi, 1967) as a method of
decoding convolutional codes. Since that time, it has been recognized as an
attractive solution to a variety of digital estimation problems, somewhat as the
Kalman filter has been adapted to a variety of analog estimation problems. Like the
Kalman filter, the Viterbi Algorithm tracks the state of a Stochastic Process with a
recursive method that is optimum in certain sense, and that lends itself readily to
implementation and analysis. However, the underlying process is assumed to be
finite-state Markov rather than Gaussian, which leads to marked differences in
structure.
The Viterbi algorithm is a dynamic programming algorithm used to find the
most likely sequence of hidden states that results in a sequence of observed events,
particularly in the context of Hidden Markov Models (HMMs) (Viterbi, 1967). This
algorithm is especially useful in stock prediction when the market is assumed to
exhibit distinct hidden states (e.g., bull and bear markets) that influence observed
stock prices (Viterbi, 1967).
The HMM consists of:
i. A set of states, each associated with a probability distribution.
ii. A sequence of observations.
iii. Transition probabilities between states.
iv. Emission probabilities, which define the likelihood of an observation
given a state.
Given the observation sequence 𝑂 = (𝑜1 , 𝑜2 , , … , 𝑜𝑇 ) and the HMM
parameters, the Viterbi algorithm computes the most probable sequence of states
𝑄 = (𝑞1 , 𝑞2 , … , 𝑞𝑇 , ).

Initialization:
For each state 𝑠𝑖 :
1 (𝑠𝑖 ) = 𝜋𝑖 . 𝑏𝑖 (𝑜1 ) (9)
where, 𝜋𝑖 is the initial probability of state 𝑠𝑖 and 𝑏𝑖 (𝑜1 ) is the emission
probability of the first observation 𝑜1 from state 𝑠𝑖 .
Recursion:
For each time step 𝑡 = 2,3, … , 𝑇 and each state 𝑠𝑗 :
 𝑡 (𝑠𝑗 ) = max[ 𝑡−1 (𝑠𝑖 ). 𝑎𝑖𝑗 ]𝑏𝑗 (𝑜𝑡 ) (10)
𝑠𝑖

where 𝑎𝑖𝑗 is the transition probability from state 𝑠𝑖 to state 𝑠𝑗 and 𝑏𝑗 (𝑜𝑡 ) is the
emission probability of observation 𝑜𝑡 from state 𝑠𝑗 .

138
Termination:
Identify the probability of the most likely sequence:
𝑃∗ = max, 𝑡 (𝑠𝑖 )- (11)
𝑠𝑖
and the corresponding final state:
𝑞𝑇∗ = 𝑎𝑟𝑔max, 𝑡 (𝑠𝑖 )- (12)
𝑠𝑖
Path Backtracking:
Backtrack through the states to find the most probable sequence:
𝑞𝑇∗ = 𝑡+1 (𝑞𝑡+1

) (13)
Where, 𝑡 (𝑠𝑗 ) = 𝑎𝑟𝑔 max𝑠𝑖 [ 𝑡−1 (𝑠𝑖 ). 𝑎𝑖𝑗 ]
Using the Viterbi algorithm, we can identify the most likely sequence of hidden
market states that explain the observed stock prices. This helps in understanding
market regimes and predicting future stock prices based on the identified states
(Yang & Aldous, 2015).
In the stock prediction model, it is assumed that stock prices follow distinct
market states with unique statistical properties. By applying the Viterbi algorithm,
these hidden states can be inferred from historical stock prices and used to predict
future prices. The model involves training on historical data to estimate the HMM
parameters, including transition probabilities, emission probabilities, and initial state
probabilities. The Viterbi algorithm is then used to decode the most likely sequence
of hidden states and predict future stock prices accordingly.
In this study, a fixed window of past observed stock prices is used to train the
HMM and apply the Viterbi algorithm to predict the next-day price. This process is
repeated for multiple predictions, updating the training dataset with actual prices as
new observations become available. The results and diagnostics of this model are
discussed in the next chapter.
When choosing between a Hidden Markov Model (HMM) and Geometric
Brownian Motion (GBM) for stock prediction, the decision depends on the
characteristics of the stock data and the research objectives. HMMs are
advantageous when stock prices exhibit distinct regimes or market conditions, such
as bull and bear markets, each with different statistical properties. HMMs can
effectively model these regime shifts and capture the temporal dependencies within
the stock price sequences (Hassan & Nath, 2005; Bulla & Bulla, 2006). In contrast,
GBM is a continuous-time stochastic process that assumes stock prices follow a
random walk with drift and volatility, making it suitable for modeling continuous
price evolution over time. GBM is widely used in financial modeling for its
simplicity and its basis in the efficient market hypothesis (Black & Scholes, 1973;
Merton, 1973). Therefore, for stock prediction involving regime changes and
discrete market states, HMM is more appropriate, while GBM is better suited for
modeling continuous stock price movements.

139
4. RESULT
A comparison of predicted and actual stock prices over a 10-day period, along
with graphical representations, is presented.
Detailed results for each model are provided, including statistical diagnostics
and performance metrics. The final section compares the overall performance of the
two models, highlighting the strengths and weaknesses of each.
The prediction error is quantified and assessed using the formula:
𝑎𝑐𝑡𝑢𝑎𝑙 − 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑒𝑑
𝐸𝑟𝑟𝑜𝑟 = 𝑎𝑐𝑡𝑢𝑎𝑙
x100 (12)

Analysis and Findings using Geometric Brownian Motion


The Geometric Brownian Motion model, as outlined in Equation 8 {𝑆(𝑡𝑖+1 ) =
𝑙
.𝜇− 𝜍 𝑚 /+𝜍𝑍
𝑆(𝑡𝑖 )𝑒 𝑚 𝑖+𝑙
+ is used to predict stock prices over 10 days. Table 1 displays
the predicted and actual values, along with individual differences. The
corresponding graphical representations are shown in Figure 1. The blue line
represents the actual stock price, while the orange line represents the predicted stock
price for BHEL.

Table 1 Prediction by Geometric Brownian motion


Open High Low Close

Date Actual Predicted Diff. Actual Predicted Diff. Actual Predicted Diff. Actual Predicted Diff.

1-Apr 249.00 251.47 -2.47 254.85 252.12 2.73 248.50 244.23 4.27 253.75 245.57 8.18

2-Apr 253.75 250.82 2.93 254.90 251.48 3.42 249.80 243.67 6.13 252.2 244.87 7.33

3-Apr 250.55 255.23 -4.68 254.65 255.88 -1.23 248.65 247.49 1.16 251.8 249.69 2.11

4-Apr 253.00 255.44 -2.44 256.90 256.08 0.28 247.45 247.67 -0.22 251.5 249.91 1.59

5-Apr 251.50 255.81 -4.31 255.90 256.45 -0.55 247.75 247.99 -0.24 254.95 250.31 4.64

8-Apr 255.75 260.76 -5.01 258.30 261.38 -3.08 254.15 252.28 1.87 256.45 255.73 0.72

9-Apr 257.20 262.11 -4.91 259.90 262.73 -2.83 253.35 253.44 -0.09 255.75 257.21 -1.46

10-Apr 256.80 258.43 -1.63 265.30 259.06 6.24 255.95 250.26 5.69 262.5 253.18 9.32

12-Apr 258.60 256.45 2.15 269.20 257.09 12.11 258.00 248.55 9.45 262.5 251.01 11.49

15-Apr 254.20 255.17 -0.97 261.95 255.82 6.13 252.55 247.44 5.11 256.5 249.62 6.88

Note: Difference is written as Diff.

140
VALUES VALUES

245
250
255
260
265
270
275
245
250
255
260
1-APR 265
1-Apr
2-APR
2-Apr
3-APR 3-Apr
4-APR 4-Apr
5-APR 5-Apr
6-APR 6-Apr
7-APR 7-Apr

8-APR 8-Apr

DATE

DATE
9-APR 9-Apr

10-APR 10-Apr

141
11-Apr

HIGH
11-APR
OPEN

12-Apr
12-APR
13-Apr
13-APR

Figure 1.2: High


14-Apr
Figure 1.1: Open

14-APR
15-Apr
15-APR
Figure 1: Geometric Brownian Motion prediction

Actual

Actual
Predicted

Predicted
Figure 1.3: Low

CLOSE
265

260

255
VALUES

250 Actual

245
Predicted

240

10-APR

11-APR

12-APR

13-APR

14-APR

15-APR
1-APR

2-APR

3-APR

4-APR

5-APR

6-APR

7-APR

8-APR

9-APR

DATE

Figure 1.4: Close

LOW
260

255
VALUES

250
Actual
245 Predicted

240

DATE

The data from April 1 to April 15 shows that the predicted values often differ
from the actual values, with discrepancies ranging from -5.01 to 2.93. The model
tends to overestimate the actual values more frequently. The differences are

142
inconsistent, indicating variability in the prediction model's accuracy. The smallest
difference of -0.97 on April 15 suggests that some predictions were relatively close,
but overall, the model needs adjustments to improve its reliability and accuracy.
From April 1 to April 15, the predicted high values generally differed from the
actual highs, with discrepancies ranging from -3.08 to 12.11. The model frequently
underestimated the actual highs, with notable underestimations on April 10 (6.24)
and April 12 (12.11). Overestimations also occurred, but to a lesser extent. These
inconsistencies highlight a significant variability in the prediction model's accuracy,
suggesting a need for further refinement to improve its reliability and predictive
accuracy.
From April 1 to April 15, the predicted low values show varying degrees of
accuracy compared to the actual lows, with differences ranging from -0.24 to 9.45.
The model often underestimated the actual lows, especially on April 2 (6.13), April
10 (5.69), and April 12 (9.45). There were instances of slight overestimations, such
as on April 4 (-0.22) and April 5 (-0.24). Overall, these discrepancies indicate
variability in the model's accuracy, emphasizing the need for further refinement to
enhance its predictive reliability.
From April 1 to April 15, the predicted closing values show varying degrees of
accuracy compared to the actual closes, with differences ranging from -1.46 to
11.49. The model generally underestimated the actual closing values, particularly on
April 1 (8.18), April 2 (7.33), April 10 (9.32), and April 12 (11.49). There were a
few instances where the model slightly overestimated, such as on April 9 (-1.46).
These inconsistencies highlight a significant variability in the prediction model's
accuracy, indicating a need for further refinement to enhance its reliability and
predictive accuracy.
Statistical Diagnostic of GBM
The performance of the Geometric Brownian Motion model was evaluated using
three error measures: Absolute Percentage Error (APE), Average Absolute Error
(AAE), and Root Mean Square Error (RMSE), as defined in Equations 1, 2, and 3.
The results are summarized in Table 1.1.
Table 1.1 Prediction error by Geometric Brownian Motion
OPEN HIGH LOW CLOSE
Mod AP AA RMS AP AA RMS AP AA RMS AP AA RMS
el E E E E E E E E E E E E
GB 1.2 3.5 3.44 1.4 3.5 5.12 1.3 3.5 4.57 2.0 3.2 6.47
M 4 5 9 8 5 7 9 7

143
The performance metrics for the Geometric Brownian Motion (GBM) model
indicate that it generally performs well in predicting Open, High, and Low values,
with relatively low Mean Absolute Percentage Errors (APE) and Root Mean Squared
Errors (RMSE). However, the model struggles more with predicting the Close
values, as evidenced by the higher APE (2.09) and RMSE (6.47). These
discrepancies suggest that while the model's absolute errors for Close values are
somewhat lower (AAE of 3.27), the percentage errors are significantly higher.
Overall, the GBM model needs further refinement to improve its accuracy,
especially for predicting the Close values.
Analysis and Findings using Hidden Markov Model Result
The Hidden Markov Model, as described in equation 13 {𝑞𝑇∗ = 𝑡+1 (𝑞𝑡+1
∗ )},
is
used to predict stock prices over the same 10-day period. Table 2 present the
predicted and actual values, along with individual differences. The graphical
representations are shown in Figure 2.

Table 2 Prediction by Hidden Markov Model


Open High Low Close

Date Actual Predicted Diff. Actual Predicted Diff. Actual Predicted Diff. Actual Predicted Diff.

1-Apr 249.00 253.05 -4.05 254.85 253.70 1.15 248.50 245.60 2.90 253.75 247.30 6.45

2-Apr 253.75 253.05 0.70 254.90 253.70 1.20 249.80 245.60 4.20 252.20 247.30 4.90

3-Apr 250.55 249.00 1.55 254.65 254.85 -0.20 248.65 248.50 0.15 251.80 253.75 -1.95

4-Apr 253.00 253.75 -0.75 256.90 254.90 2.00 247.45 249.80 -2.35 251.50 252.20 -0.70

5-Apr 251.5 250.55 0.95 255.90 254.65 1.25 247.75 248.65 -0.90 254.95 251.80 3.15

8-Apr 255.75 253.00 2.75 258.30 256.90 1.40 254.15 247.45 6.70 256.45 251.50 4.95

9-Apr 257.20 251.50 5.70 259.90 255.90 4.00 253.35 247.75 5.60 255.75 254.95 0.80

10-Apr 256.80 255.75 1.05 265.30 255.90 9.40 255.95 254.15 1.80 262.50 256.45 6.05

12-Apr 258.60 257.20 1.40 269.20 259.90 9.30 258.00 253.35 4.65 262.50 255.75 6.75

15-Apr 254.20 256.80 -2.60 261.95 265.30 -3.35 252.55 255.95 -3.40 256.50 262.50 -6.00

Note: Difference is written as Diff.

144
VALUES VALUES

250
255
260
265
270
275
244
248
250
252
254
256
258
260

246
1-APR 1-APR
2-APR 2-APR
3-APR 3-APR
4-APR 4-APR
5-APR 5-APR
6-APR 6-APR
7-APR 7-APR
8-APR 8-APR

DATE
DATE
9-APR 9-APR

145
10-APR 10-APR

HIGH
OPEN

11-APR 11-APR
12-APR 12-APR
13-APR 13-APR

Figure 2.2: High


Figure 2.1: Open
Figure 2: Hidden Markov Model prediction

14-APR 14-APR
15-APR 15-APR

Actual
Actual

Predicted
Predicted
Figure 2.3: Low

LOW
260

255
VALUES

250
Actual
245
Predicted
240

10-APR
11-APR
12-APR
13-APR
14-APR
15-APR
1-APR
2-APR
3-APR
4-APR
5-APR
6-APR
7-APR
8-APR
9-APR

DATE

Figure 2.4: Close

CLOSE
265

260
VALUES

255
Actual
250
Predicted
245
10-APR
11-APR
12-APR
13-APR
14-APR
15-APR
8-APR
1-APR
2-APR
3-APR
4-APR
5-APR
6-APR
7-APR

9-APR

DATE

The data from April 1 to April 15 shows that the prediction model for open
values performs quite well, frequently producing results that are very close to the
actual values. For instance, the predictions on April 2 (0.7), April 5 (0.95), and April

146
10 (1.05) are impressively accurate. Even on days with larger differences, such as
April 9 (5.7), the model still provides valuable insights that are beneficial for making
informed decisions. The overall performance reflects a robust and reliable model,
and with continued refinement, it has the potential to achieve even higher accuracy
and consistency in predicting open values.
The data from April 1 to April 15 demonstrates that the prediction model for
high values performs admirably, often producing results that are quite close to the
actual values. Notably, on several occasions, such as April 1 (1.15) and April 2 (1.2),
the model's predictions were remarkably accurate. Even on days with larger
discrepancies, the model provides valuable insights that are useful for making
informed decisions. The overall performance reflects a strong foundation, and with
further fine-tuning, the model has great potential to achieve even higher accuracy
and reliability in predicting high values.
The data from April 1 to April 15 indicates that the prediction model for low
values performs commendably, often producing results that closely match the actual
values. The model demonstrates impressive accuracy on several days, such as April
3 (0.15) and April 10 (1.8). Even when there are differences, like on April 1 (2.9)
and April 2 (4.2), the model still provides valuable insights. While some days show
larger discrepancies, such as April 8 (6.7) and April 9 (5.6), the overall performance
is robust. With continued refinement, the model shows great potential to achieve
even higher accuracy and reliability in predicting low values.
The data from April 1 to April 15 indicates that the prediction model for close
values performs well, with several instances of close matches between predicted and
actual values. Notably, on April 4 (-0.7) and April 9 (0.8), the model's predictions
were very accurate. Although there are larger discrepancies on some days, such as
April 1 (6.45) and April 12 (6.75), the model still provides valuable predictive
insights that can guide decision-making. The overall performance is strong, and with
further refinement, the model has the potential to achieve even greater accuracy and
reliability in predicting close values.

3.2.1 Statistical Diagnostic of HMM


Similar to the GBM, the Hidden Markov Model's performance was assessed
using APE, AAE, and RMSE. The results are provided in Table 2.1..
Table 2.1 Prediction error by Hidden Markov Model.
OPEN HIGH LOW CLOSE
Model APE AAE RMSE APE AAE RMSE APE AAE RMSE APE AAE RMSE
HMM 0.85 3.60 2.66 1.26 3.58 4.61 1.29 3.56 3.81 1.68 3.60 4.72

147
The performance metrics for the Hidden Markov Model (HMM) across the
Open, High, Low, and Close values are impressive, indicating its strong predictive
capabilities. The model achieves a low Mean Absolute Percentage Error (APE)
across all categories, with particularly notable performance in predicting Open
values (0.85). The Average Absolute Error (AAE) and Root Mean Squared Error
(RMSE) are consistently low, demonstrating the model's accuracy and reliability.
Specifically, the AAE for all categories hovers around 3.6, and the RMSE remains
below 5, showcasing the model's robustness. The HMM model proves to be a
valuable tool, providing accurate predictions that can significantly aid in decision-
making processes. With these solid performance metrics, the model is well-
positioned for further refinement and application in predictive analytics.

Comparison
In this section, a comparison is made between the results obtained from the two
models. Table 3 and Figure 3 illustrate the comparison between the actual stock
prices and the predictions made by both models.
Table 3 Results comparison from the models – Geometric Brownian Motion (GBM)
and Hidden Markov Model (HMM)
Open High Low Close

Date Actual GBM HMM Actual GBM HMM Actual GBM HMM Actual GBM HMM

1-Apr 249.00 251.47 253.05 254.85 252.12 253.70 248.50 244.23 245.60 253.75 245.57 247.30

2-Apr 253.75 250.82 253.05 254.90 251.48 253.70 249.80 243.67 245.60 252.20 244.87 247.30

3-Apr 250.55 255.23 249.00 254.65 255.88 254.85 248.65 247.49 248.50 251.80 249.69 253.75

4-Apr 253.00 255.44 253.75 256.90 256.08 254.90 247.45 247.67 249.80 251.50 249.91 252.20

5-Apr 251.50 255.81 250.55 255.90 256.45 254.65 247.75 247.99 248.65 254.95 250.31 251.80

8-Apr 255.75 260.76 253.00 258.30 261.38 256.90 254.15 252.28 247.45 256.45 255.73 251.50

9-Apr 257.20 262.11 251.50 259.90 262.73 255.90 253.35 253.44 247.75 255.75 257.21 254.95

10-Apr 256.80 258.43 255.75 265.30 259.06 255.90 255.95 250.26 254.15 262.50 253.18 256.45

12-Apr 258.60 256.45 257.20 269.20 257.09 259.90 258.00 248.55 253.35 262.50 251.01 255.75

15-Apr 254.20 255.17 256.80 261.95 255.82 265.30 252.55 247.44 255.95 256.50 249.62 262.50

Note: Difference is written as Diff.

148
Figure 3: Comparison of Geometric Brownian motion and Hidden
Markov Model

Figure 3.1: Open

CLOSE
265

260
VALUES

255

250 Actual

245 GBM

240 HMM
1-APR
2-APR
3-APR
4-APR
5-APR
6-APR
7-APR
8-APR
9-APR
10-APR
11-APR
12-APR
13-APR
14-APR
15-APR
DATE

Figure 3.2: High

LOW
275
270
265
VALUES

260
255 Actual
250
GBM
245
240 HMM
4-APR
1-APR
2-APR
3-APR

5-APR
6-APR
7-APR
8-APR
9-APR
10-APR
11-APR
12-APR
13-APR
14-APR
15-APR

DATE

149
Figure 3.3: Low

HIGH
270

265
VALUES

260
Actual
255 GBM

250 HMM

11-APR
10-APR

12-APR
13-APR
14-APR
15-APR
3-APR
1-APR
2-APR

4-APR
5-APR
6-APR
7-APR
8-APR
9-APR

DATE

Figure 3.4: Close

OPEN
265

260
VALUES

255
Actual
250 GBM
HMM
245
10-APR
11-APR
12-APR
13-APR
14-APR
15-APR
5-APR
1-APR
2-APR
3-APR
4-APR

6-APR
7-APR
8-APR
9-APR

DATE

The comparison of the Geometric Brownian Motion (GBM) and Hidden Markov
Model (HMM) for predicting Open values from April 1 to April 15 shows that the
HMM generally provides more accurate predictions. The HMM model had smaller
differences from the actual values on most days, except for April 1 and April 9.
150
Notably, on April 2, April 3, April 4, April 5, April 10, and April 12, the HMM
predictions were closer to the actual values compared to the GBM predictions. The
GBM model performed slightly better on April 15, with a smaller difference from
the actual value compared to the HMM. Overall, the HMM model demonstrated
better performance, indicating it may be more reliable for predicting Open values in
this context.
The comparison of the Geometric Brownian Motion (GBM) and Hidden Markov
Model (HMM) for predicting high values from April 1 to April 15 reveals that the
HMM generally provides more accurate predictions. The HMM model consistently
shows smaller differences from the actual values on most days, such as April 1, April
2, April 3, April 5, and April 8. The GBM model, while performing slightly better on
April 4 and April 15, has larger discrepancies on several occasions. Notably, both
models exhibit significant discrepancies on April 9, April 10, and April 12. Overall,
the HMM model demonstrates better performance and reliability in predicting high
values, indicating its potential for more accurate future predictions.
The comparison of the Geometric Brownian Motion (GBM) and Hidden Markov
Model (HMM) for predicting low values from April 1 to April 15 shows that both
models offer valuable predictions, with the HMM generally providing more accurate
results. The HMM model has smaller differences from the actual values on several
days, such as April 1, April 2, April 3, April 5, and April 10. For example, on April
3, the HMM prediction was very close to the actual value (248.5 vs. 248.65).
However, on certain days, such as April 4 and April 15, the GBM model performed
better. Despite some larger discrepancies, both models exhibit strengths, with the
HMM showing overall superior performance in predicting low values for most
dates, indicating its greater reliability and accuracy.
The comparison of the Geometric Brownian Motion (GBM) and Hidden Markov
Model (HMM) for predicting close values from April 1 to April 15 shows that both
models have their strengths, but the HMM generally provides more accurate
predictions. The HMM model consistently shows smaller differences from the actual
values on most days, such as April 1, April 2, April 4, April 5, and April 8. For
instance, on April 4, the HMM prediction (252.2) was very close to the actual value
(251.5). On certain days, such as April 9 and April 15, the GBM model performed
slightly better.
Overall, the HMM model demonstrates better performance and reliability in
predicting close values, indicating its potential for more accurate future predictions.

Table 3.1. Error measures comparison from the two models (GBM and HMM)
OPEN HIGH LOW CLOSE
Model APE AAE RMSE APE AAE RMSE APE AAE RMSE APE AAE RMSE
GBM 1.24 3.55 3.44 1.49 3.58 5.12 1.35 3.57 4.57 2.09 3.27 6.47
HMM 0.85 3.60 2.66 1.26 3.58 4.61 1.29 3.56 3.81 1.68 3.60 4.72

151
The comparison of the Geometric Brownian Motion (GBM) and Hidden Markov
Model (HMM) for predicting Open, High, Low, and Close values indicates that the
HMM generally outperforms GBM. The HMM model shows lower Mean Absolute
Percentage Error (APE) and Root Mean Squared Error (RMSE) across all
categories, indicating higher accuracy and reliability. Specifically, for Open values,
the HMM has a lower APE (0.85) and RMSE (2.66) compared to GBM (1.24 APE
and 3.44 RMSE). For High values, HMM again has a better APE (1.26) and RMSE
(4.61) than GBM (1.49 APE and 5.12 RMSE). The trend continues for Low values,
where HMM's APE (1.29) and RMSE (3.81) are superior to GBM's (1.35 APE and
4.57 RMSE). For Close values, HMM shows a significant improvement with an
APE of 1.68 and RMSE of 4.72, compared to GBM's APE of 2.09 and RMSE of
6.47. Overall, the HMM model provides more precise and consistent predictions,
making it a more reliable tool for predicting stock prices.

5. CONCLUSION
Predictive modeling is an essential mathematical approach used in business,
engineering, and finance, among other fields, to provide projections based on past
data. Predicting stock prices is one of the most important uses of predictive
modeling among its various applications. Due to their broad application and
theoretical stability, two well-known stochastic models—Geometric Brownian
Motion (GBM) and Hidden Markov Model (HMM)—are selected. GBM, frequently
used to predict stock prices, is based on the idea that stock returns are independent
and normally distributed over time. This model is a mainstay of financial modeling,
distinguished by its mathematical beauty and simplicity. On the other hand, HMMs
excel at capturing different market regimes and transitions, making them ideal for
modeling temporal dependencies and regime shifts in financial time series data.
Historical stock price data from Bharat Heavy Electricals Limited (BHEL), a
significant participant in India's infrastructure and power production industries, is
used in the study. The data covers the period from July 14, 2023, to March 31, 2024,
characterized by several noteworthy events, such as the launch of Chandrayaan-3.
This time frame is chosen based on significant events likely to affect BHEL's stock
performance. BHEL provides the necessary historical and current financial data,
including daily closing share prices, from financial market websites like Yahoo
Finance. The Hidden Markov Model (HMM) and Geometric Brownian Motion
(GBM) are the two prediction models that are assessed.
This study primarily focuses on Geometric Brownian Motion (GBM) and
Hidden Markov Models (HMM), exploring the core concepts of stochastic processes
and their application in predicting stock movements. The objective is to improve
financial analysis and decision-making by gaining a thorough understanding of these
processes and their implications in financial modeling, thereby advancing predictive
modeling tools.
Stock prices based on a continuous-time stochastic process are simulated using
the Geometric Brownian Motion model. Predictions are produced over a 10-day
152
period after parameters for the model are determined using historical data. The
model's efficacy is assessed by comparing expected and actual prices. Any
differences are then examined to determine the model's limitations. The most likely
sequence of hidden states is decoded, and future stock values are predicted by the
Hidden Markov Model, which captures the dynamics of market regimes and
transitions using the Viterbi algorithm. The advantage of this model is that it can
identify and adjust to various market situations, providing a more sophisticated
understanding of stock price fluctuations.
Three important error metrics are used to evaluate the performance of both
models: Root Mean Square Error (RMSE), Average Absolute Error (AAE), and
Absolute Percentage Error (APE). These measurements provide a thorough
assessment of the predicted accuracy of each model. The findings show that, in
comparison to the GBM model, the HMM model often provides more precise and
trustworthy predictions. The HMM model has an APE of 0.85, an AAE of 3.60, and
an RMSE of 2.66 for Open Values, whereas the GBM model has an APE of 1.24, an
AAE of 3.55, and an RMSE of 3.44. For High Values, HMM has an APE of 1.26,
AAE 3.58, and RMSE 4.61, while GBM has an APE of 1.49, RMSE 5.12, and AAE
3.58. HMM exhibits an APE of 1.29, AAE 3.56, and RMSE 3.81 for Low Values,
whereas GBM has an APE of 1.35, AAE 3.57, and RMSE 4.57. In terms of Close
Values, HMM has an APE of 1.68, AAE 3.60, and RMSE 4.72, whereas GBM has
an APE of 2.09, AAE 3.27, and RMSE 6.47.
The comparison demonstrates the advantage of the HMM model in terms of
dependability and accuracy of prediction. Its ability to accurately capture the
dynamic nature of financial markets is evident from the HMM model's consistent
performance above the GBM model on all measures. Predictions are made with
more accuracy and dependability due to the HMM's excellent capacity to identify
and simulate regime changes. However, even though it is useful, the GBM model
has more errors and inconsistencies, especially in closing value prediction.
This study highlights the significance of choosing suitable models for stock
price prediction based on the unique properties of financial data. Despite its ease of
use and attractiveness as a theory, the Geometric Brownian Motion model performs
less well in this situation, especially when it comes to closing prices. It is still a
useful tool for simulating ongoing price changes. However, the Hidden Markov
Model outperforms others in terms of accuracy and resilience, making it a more
useful tool for predicting stock values due to its ability to identify changes in the
market and regime.
Both models might benefit from more calibration and refinement using larger
datasets and additional variables for future studies and real-world applications.
Prediction accuracy may be improved by using GBM and HMM components or
combining them with other cutting-edge machine learning methods. Furthermore,
integrating continuous learning and real-time analytic frameworks may aid in
adjusting the models to shifting market conditions, enhancing their applicability in
the real world.
153
In conclusion, due to the increased accuracy and resilience in capturing market
dynamics, the Hidden Markov Model is advised for stock price prediction. However,
the Geometric Brownian Motion model remains useful, particularly in situations that
fit its basic assumptions. This work advances predictive modeling methods for
financial analysis and decision-making, improving our understanding of stochastic
processes in financial modeling. The results demonstrate how advanced models,
such as HMM, may enhance financial forecasts and aid in better investment
decision-making.

6. FUTURE SCOPE OF THE STUDY


Future research could focus on further calibrating GBM and HMM using larger
datasets and additional variables like macroeconomic indicators and sentiment
analysis. Exploring hybrid models that combine elements of GBM, Hidden Markov
Model and different methods of Hidden Markov Model, as well as developing
frameworks for real-time analytics and continuous learning, could enhance
prediction accuracy. Additionally, improving regime detection, applying these
models to different markets and financial instruments, and extending their use in risk
management can provide more comprehensive insights and robust financial
forecasting tools.

7. REFERENCES
Adebiyi, Ayodele Ariyo, Aderemi Oluyinka Adewumi, and Charles Korede Ayo.
(2014). Comparison of Arima and artificial neural networks models for
stock price prediction. Journal of Applied Mathematics 2014: 614342.
Agustini, W. Farida, Ika Restu Affianti, and Endah R. M. Putri (2018). Stock
price prediction using geometric brownian motion. Journal of Physics:
Conference Series 974: 012047.
Artzner, P., Delbaen, F., Eber, J. M., & Heath, D. (1999). "Coherent measures of
risk." Mathematical Finance, 9(3), 203-228.
Bachelier, L. (1900). Théorie de la spéculation. Annales Scientifiques de l'École
Normale Supérieure, 3(17), 21-86.
Bulla, J., & Bulla, I. (2006). Stylized facts of financial time series and hidden
semi-Markov models. Computational Statistics & Data Analysis, 51(4),
2192-2209. DOI: 10.1016/[Link].2006.09.033.
Chai, T., & Draxler, R. R. (2014). Root mean square error (RMSE) or mean
absolute error (MAE)?–Arguments against avoiding RMSE in the
literature. Geoscientific Model Development, 7(3), 1247-1250.
Chen, Y., Hao, Y., & Chen, J. (2013). The applications of soft computing in
financial engineering: A review. Applied Soft Computing, 13(3), 1497-
1506.
Dmouj, Abdelmoula (2006_. Stock Price Modelling: Theory and Practice.
Masters‘s thesis, Vrije Universiteit, Amsterdam, The Netherlands.
154
Fama, E. F. (1965). "The Behavior of Stock-Market Prices." The Journal of
Business, 38(1), 34-105.
Forney, G. D. (1973). The Viterbi algorithm. Proceedings of the IEEE, 61(3),
268-278.
Hamilton, J. D. (1989). "A new approach to the economic analysis of
nonstationary time series and the business cycle." Econometrica: Journal
of the Econometric Society, 357-384.
Hamilton, J. D. (1990). Analysis of time series subject to changes in regime.
Journal of Econometrics, 45(1-2), 39-70.
Hassan, M. R., & Nath, B. (2005). Stock market forecasting using hidden
Markov model: A new approach. In Proceedings of the 5th International
Conference on Intelligent Systems Design and Applications (ISDA'05)
(pp. 192-196). IEEE.
Hassan, M. R., Nath, B., & Kirley, M. (2007). A fusion model of HMM, ANN,
and GA for stock market forecasting. Expert Systems with Applications,
33(1), 171-180.
Hassan, M. R., Nath, B., & Kirley, M. (2007). A fusion model of HMM, ANN
and GA for stock market forecasting. Expert Systems with Applications,
33(1), 171-180.
Hassan, Md Rafiul, and Baikunth Nath (2005). Stock market forecasting using
hidden markov model: A new approach. Paper presented at the
International Conference on Intelligent Systems Design and Applications
(ISDA‘05), Pretoria, South Africa, December 3–5, Piscataway: IEEE, pp.
192–96.
Hassan, Md Rafiul, Baikunth Nath, and Michael Kirley (2007). A fusion model
of hmm, ann and ga for stock market forecasting. Expert Systems with
Applications 33: 71–80.
Huang, W., Nakamori, Y., & Wang, S. Y. (2005). Forecasting stock market
movement direction with support vector machine. Computers &
Operations Research, 32(10), 2513-2522.
Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast
accuracy. International Journal of Forecasting, 22(4), 679-688.
Introduction to Stochastic Process Theory and its Uses by Palle Thoft-
Christensen Ph. D. & Michael J. Baker [Link]. (Eng) Chapter 9, page no-
145 – 169.
Islam, M. R., & Nguyen, N. (2020). Comparison of Financial Models for Stock
Price Prediction. Journal of Risk and Financial Management, 2020, 13,
181. doi:10.3390/jrfm13080181
Jurafsky, D., & Martin, J. H. (2009). Chapter 3: Hidden Markov Models (pp. 33-
41). In Speech and language processing (2nd ed.). Prentice Hall.
Khashei, Mehdi, and Mehdi Bijari (2010). An artificial neural network (p, d, q)
model for timeseries forecasting. Expert Systems with Applications 37:
479–89.
155
Kim, K. J. (2003). Financial time series forecasting using support vector
machines. Neurocomputing, 55(1-2), 307-319.
Kim, S., Shephard, N., & Chib, S. (1998). "Stochastic volatility: Likelihood
inference and comparison with ARCH models." The Review of Economic
Studies, 65(3), 361-393.
Lee, Kyungjoo, Sehwan Yoo, and John Jongdae (2007). Neural network model
versus sarima model in forecasting korean stock price index (kospi).
Issues in Information System 8: 372–8.
Makridakis, S., Wheelwright, S. C., & Hyndman, R. J. (1998). Forecasting:
methods and applications. John Wiley & Sons.
Malkiel, B. G. (2003). The Efficient Market Hypothesis and Its Critics. Journal
of Economic Perspectives, 17(1), 59-82.
Merh, Nitin, Vinod P. Saxena, and Kamal Raj Pardasani (2010). A comparison
between hybrid approaches of ann and arima for indian stock trend
forecasting. Business Intelligence Journal 3: 23–43.
Meyler, A., Kenny, G., & Quinn, T. (1998). "Forecasting Irish Inflation using
ARIMA Models." Central Bank and Financial Services Authority of
Ireland Technical Paper Series, 3/RT/98.
Meyler, Aidan, Geoff Kenny, and Terry Quinn (1998). Forecasting Irish Inflation
Using Arima Models. Dublin: Central Bank of Ireland.
Padi, T. R., Dar, G. F., & Sarode, R. (2023). Markov Modelling of Indian Stock
Market Prices with Reference to State Bank of India. International Journal
of Statistics and Reliability Engineering Vol. 9(3), pp. 481-492, 2022.
Paul, R., & Deb Roy, T. (2025). Forecasting stock trends of Bharat Heavy
Electricals Limited: An Application of Geometric Brownian Motion.
International Journal of Statistics and Reliability Engineering, 12(1), 11–
17
Rabiner, L. R. (1989). A Tutorial on Hidden Markov Models and Selected
Applications in Speech Recognition. Proceedings of the IEEE, 77(2), 257-
286. DOI: 10.1109/5.18626.
Rathnayaka, R. M. Kapila Tharanga, Wei Jianguo, and DMK N. Seneviratna
(2014). Geometric brownian motion with ito‘s lemma approach to
evaluate market fluctuations: A case study on colombo stock exchange.
Paper presented at the 2014 International Conference on Behavioral,
Economic, and Socio-Cultural Computing (BESC2014), Shanghai, China,
October 30–November 2, Piscataway: IEEE, pp. 1–6.
Robert, C. Y., & West, K. D. (1997). "Gaussian priors, black-litterman, and
expected returns." The Journal of Finance, 52(1), 35-64.
Saidane, M. (2022). A New Viterbi-Based Decoding Strategy for Market Risk
Tracking: an Application to the Tunisian Foreign Debt Portfolio During
2010–2012. Statistika Statistics and Economy Journal, Volume(Issue),
454-470.

156
Sajja, N. (2012). Analyzing Inherent Ambiguities in the Context of Stochastic
Processes. International Journal of Computer Applications, 49(8), 38-42.
Ševč ovič , D., Stehlíková, B., & Mikula, K. (2011). Analytical and Numerical
Methods for Pricing Financial Derivatives. Nova Science Publishers.
Tambi, Mahesh Kumar (2005). Forecasting exchange rate: A univariate out of
sample approach. The IUP Journal of Bank Management 4: 60–74.
Tsai, C. F., & Hsiao, Y. C. (2010). Combining multiple feature selection methods
for stock prediction: Union, intersection, and multi-intersection
approaches. Decision Support Systems, 50(1), 258-269.
Viterbi, A. J. (1967). Error bounds for convolutional codes and an
asymptotically optimum decoding algorithm. IEEE Transactions on
Information Theory, 13(2), 260-269. doi:10.1109/TIT.1967.1054010.
Welch, L. R. (2003). Hidden Markov models and the Baum-Welch algorithm.
IEEE Information Theory Society Newsletter, 53(4), 10-13.
White, H. (1988). Economic prediction using neural networks: The case of IBM
daily stock returns. Proceedings of the IEEE International Conference on
Neural Networks, 451-458.
White, H. (1988). Economic prediction using neural networks: The case of IBM
daily stock returns. In IEEE International Conference on Neural Networks
(Vol. 2, pp. 451-458). IEEE.
Yang, H., & Aldous, D. J. (2015). Hidden Markov Models with Applications in
Finance. Handbook of Statistics, 35, 641-675. DOI:
10.1016/[Link].2015.08.002.
Yao, Jingtao, Chew Lim Tan, and Hean-Lee Poh (1999). Neural networks for
technical analysis: A study on klci. International Journal of Theoretical
and Applied Finance 2: 221–41.
Zhang, G. P., Patuwo, B. E., & Hu, M. Y. (1998). Forecasting with artificial
neural networks: The state of the art. International Journal of Forecasting,
14(1), 35-62.
Zhang, G., & Wu, M. (2009). Stock market prediction with back propagation
neural networks and improved bacterial chemotaxis optimization.
Proceedings of the International Conference on Computational
Intelligence and Software Engineering, 1-4.
Zhang, Yudong, and Lenan Wu. (2009). Stock market prediction of S & P 500
via combination of improved bco approach and bp neural network. Expert
Systems with Applications 36: 8849–54.

157
Topp-Leone Unit Lindley Distribution: Derivation, Some
Fundamental Properties And Estimation
Sahana Bhattacharjee
Department Of Statistics, Gauhati University, Assam, Email ID:
[Link]@[Link]

Abstract:
In this paper, a new distribution called the Topp-Leone Unit Lindley distribution
has been developed by considering the Unit Lindley distribution as the baseline
distribution in the Topp Leone family of distributions. This distribution is also
expressible as the weighted sum of the exponentiated-Unit Lindley random variables
and it assumes values in the range 0 to 1 and thus, can be considered as a viable
alternative to distributions arising in the unit interval. The behaviour of the density,
distribution function and hazard rate function curve are studied and some
fundamental properties of the proposed distribution are explored. It is seen that as
compared to the Unit Lindley distribution which has an increasing hazard rate
function, the Topp-Leone Unit Lindley distribution which has an extra parameter has
different shapes of the hazard function based on the values of parameters. This
shows the flexibility of the proposed distribution. Finally, the parameter estimation
is performed using the maximum likelihood method and a simulation study is
carried out to assess the consistency of the estimates.
Key Words: Hazard rate function, Maximum likelihood estimation, Simulation
study, Topp Leone family of distributions, Unit Lindley distribution

1. INTRODUCTION
Probability distributions are one of the many indispensable tools in statistics
which are useful in modelling uncertainty of real-world phenomena and drawing
meaningful conclusions by estimating the model parameters and testing appropriate
hypotheses related to the population (Handique, 2018). Lifetime distributions are
probabilistic models which are used to describe the probability of survival or failure
of an article at a given time. Such models are often used in various fields such as
reliability engineering, demography, bio-medical science etc. Many lifetime
distributions have been developed over the last few decades which have been
applied to real life data on survival or failure time. In the recent few years, the
development of more flexible distributions by adding one or more extra parameter to
the baseline distribution has drawn the attention of statisticians working in the field
of distribution theory. These are broadly referred to as the generalized class of

158
probability distributions. Tahir and Nadarajah (2015) summarised how additional
parameters can be inducted to the well-established continuous univariate generalized
(or -G) families. Some of the popular G-families which exist in the literature are
Marshall- Olkin extended (MOE) family (Marshall and Olkin, 1997), beta-G family
(Eugene et al., 2002), Transmuted family (Shaw and Buckley, 2007), Gamma—G
family (Zografos and Balakrishnan, 2009), Kumaraswamy-G family (Cordeiro and
de-Castro, 2011), Transformed-transformer T-X family (Alza-Atreh et al., 2013),
Exponentiated-G family (Cordeiro et a., 2013).
The Topp-Leone family of distributions was introduced by Al-Shomrani et al.
(2016). It is also expressible as the weighted sum of the exponentiated-G
distributions. Till date, many sub-models of the Topp-Leone family of distributions
have been developed by various authors, such as the Topp-Leone Exponential
Distribution (Al-Shomrani et al., 2016), Topp-Leone Extended Exponential
Distribution (Khaoula et al., 2022), Topp-Leone Weibull Distribution (Tuoyo, 2021)
to name a few. However, not many authors have taken up the work on applying the
Topp-Leone generalization to a finitely bounded distribution. This motivated me to
apply the Topp-Leone generalization to the one- parameter Unit Lindley distribution
proposed by Mazucheli et al., 2019, which has support in the range (0,1) and explore
some of its fundamental properties. In this paper, an attempt has been made to
develop the Topp-Leone Unit Lindley distribution having support in the unit interval
and explore some of its fundamental properties, along with estimation of parameters
of the distribution using the maximum likelihood method.
This paper is organized as follows: Section 1 introduces the topic and motivation
behind the study and reviews the existing literature. Section 2 displays the derivation
of the density and cumulative distribution of the proposed Topp-Leone Unit Lindley
distribution and their behavioural study via graphs. Some of its fundamental
distributional properties like moments, hazard rate function is explored in Section 3.
Section 4 contains the maximum likelihood estimation of the parameters of the
distribution and a simulation study to assess the large sample performance of the
estimates so obtained. Section 5 presents the concluding remarks and future scopes
of the current study.

2. Derivation of the Topp-Leone Unit Lindley distribution


The Topp-Leone family of distributions was proposed by Al-Shomrani et al.
(2016), by using the form of the distribution function of the Topp-Leone distribution
(Nadarajah and Kotz, 2003) given by
𝐹𝑇𝐿 (𝑥) = 𝑥 𝛼 (2 − 𝑥)𝛼 , 0 < 𝑥 < 1; 𝛼 > 0 (1)
The distribution function of the Topp-Leone family of distributions was obtained
by replacing x by 𝐹(𝑥) in the RHS of (1) as
𝐹𝑇𝐿−𝐺 (𝑥) = ,𝐹(𝑥)-𝛼 ,2 − 𝐹(𝑥)-𝛼 , 𝑥 ∈ ℜ; 𝛼 > 0 (2)
where 𝐹(𝑥) is the distribution function of the baseline distribution.
The density function of the Topp-Leone family of distributions is given by
𝑓𝑇𝐿−𝐺 (𝑥) = 2𝛼𝑓(𝑥),1 − 𝐹(𝑥)-,𝐹(𝑥)-𝛼−1 ,2 − 𝐹(𝑥)-𝛼−1 , 𝑥 ∈ ℜ; 𝛼 > 0 (3)
159
The series representation of the density function is
2𝑗+1
𝑓𝑇𝐿−𝐺 (𝑥) = ∑∞ 𝑗=0 ∑𝑚=0 𝑏(𝑗, 𝑚)𝑕𝑚+1 (𝑥) (4)
(−1)𝑗+𝑚 2𝛤(𝛼+1)
Where 𝑏(𝑗, 𝑚) = 𝑗!𝛤(𝛼−𝑗)(𝑚+1)
(2𝑗+1
𝑚
) and 𝑕𝑚+1 (𝑥) = (𝑚 + 1)𝑓(𝑥)*𝐹(𝑥)+𝑚
is the density of the exponentiated-G distribution with the power parameter m. Thus,
the Topp-leone family of distributions is expressible as the weighted sum of the
exponentiated-G distributions.
The Unit Lindley distribution was developed by Mazucheli et al. (2019) as an
alternative to the popular distributions arising in the unit interval such as the two-
parameter Beta distribution, Kumaraswamy distribution, Topp-Leone distribution
etc. The density function and the distribution function of the one-parameter Unit
Lindley distribution is
𝜃𝑥
𝜃𝑚 1
𝑓(𝑥) = 1+𝜃 (1−𝑥)𝑛 𝑒 −𝑙−𝑥 , 0 < 𝑥 < 1, 𝜃 > 0 (5)
and
𝜃𝑥
𝜃𝑥
𝐹(𝑥) = 1 − {.1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } (6)
Considering the one-parameter Unit Lindley as the baseline distribution and
substituting its p.d.f. from (5) in (3), the density function of the Topp-Leone Unit
Lindley distribution is obtained as
𝜃𝑥 2 𝑚𝜃𝑥 𝛼
2𝛼𝜃 𝑚 1 − 𝜃𝑥
𝑔𝑇𝐿𝑈𝐿 (𝑥) = 𝑛 𝑒 𝑙−𝑥 [1 − {1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] [{1 − .1 −
1+𝜃 (1−𝑥)
𝑚𝜃𝑥 𝛼
𝛼−1 𝛼 𝛼−1
𝜃𝑥 2 𝜃𝑥 2 𝑚𝜃𝑥

(1+𝜃)(𝑥−1)
/ 𝑒 𝑙−𝑥 } ] [2 − {1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] ,0 < 𝑥 < 1 (7)
The c.d.f. of the Topp-Leone Unit Lindley distribution is derived by substituting
the c.d.f of the Unit Lindley distribution from (6) in (2) and its final form is shown
below:
2 𝑚𝜃𝑥 𝛼
𝜃𝑥
𝐺𝑇𝐿𝑈𝐿 (𝑥) = {1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } (8)
The density function 𝑔𝑇𝐿𝑈𝐿 (𝑥) can also be expressed as the weighted sum of the
exponentiated-Unit Lindley densities as
2𝑗+1
𝑔𝑇𝐿𝑈𝐿 (𝑥) = ∑∞
𝑗=0 ∑𝑚=0 𝑏(𝑗, 𝑚)𝑕𝑚+1 (𝑥) (9)
(−1)𝑗+𝑚 2𝛤(𝛼+1) 2𝑗+1
Where 𝑏(𝑗, 𝑚) = ( 𝑚 ) is the weight, and𝑕𝑚+1 (𝑥) =
𝑗!𝛤(𝛼−𝑗)(𝑚+1)
𝜃𝑥 𝜃𝑥 𝑚
𝜃𝑚 1 − 𝜃𝑥
(𝑚 + 1)
1+𝜃 (1−𝑥) 𝑛 𝑒 𝑙−𝑥 [1 − {.1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 }] is the exponentiated-Unit
Lindley distribution with power parameter m.
Fig. 1 and Fig. 2 display the density function curve and the distribution function
curve of the Topp-Leone Unit Lindley distribution for different combinations of the
parameters.

160
Fig. 1 p.d.f. Curve of the Topp-Leone Unit Lindley distribution

Fig. 2 c.d.f. Curve of the Topp-Leone Unit Lindley distribution

It is seen that the peakedness of the p.d.f. curve decreases with a decrease in the
value of α . Also, the p.d.f. curve of the Topp-Leone Unit Lindley distribution is
positively skewed, in contrast to the Unit Lindley distribution which is negatively
skewed.
3. Moments and Hazard Rate Function
3.1 Moments

161
The rth order raw moment of the Topp-Leone Unit Lindley distribution is given
by
𝜃𝑥 2 𝑚𝜃𝑥 𝛼
2𝛼𝜃 𝑚 𝑥𝑟 − 𝜃𝑥
𝑛 𝑒 𝑙−𝑥 [1 − {1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ]
1 1+𝜃 (1−𝑥)
𝜇𝑟′ = 𝐸(𝑋 𝑟 ) = ∫0 𝛼 𝛼−1 𝛼 𝛼−1
𝜃𝑥 2 𝑚𝜃𝑥
𝜃𝑥 2 𝑚𝜃𝑥
[{1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] [2 − {1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] 𝑑𝑥
(10)

Putting 𝑘 = 1,2,3,4 in (10), we obtain the first four raw moments of the Topp-
Leone Unit Lindley distribution. Since the integral in (10) cannot be evaluated
analytically, their approximations are obtained using the RStudio software, version
4.2.2 via the integrate() function. The skewness (Sk) and kurtosis (Ks) of the
distribution are computed using the formulae shown below:
𝜇𝑛′ −3𝜇𝑚′ 𝜇𝑙′ +2𝜇𝑙′𝑛 𝜇𝑜′ −4𝜇𝑛′ 𝜇𝑙′ +6𝜇𝑚′ 𝜇𝑙′𝑚 −3𝜇𝑙′𝑜
𝑆𝑘 = 𝑛 , 𝑠 = 𝑛
𝜇𝑚𝑚 𝜇𝑚𝑚
Fig. 3 and Fig. 4 show the plots of mean and variance of the Topp-Leone Unit
Lindley distribution as functions of α and θ .

Fig. 3 (a) Plot of mean as a function of α for θ =0.5 (b) Plot of mean as a function of
θ for α =0.5 (c) Plot of variance as a function of α for θ =0.5 (d) Plot of variance as
a function of θ for α =0.5
It is seen that both the mean and variance decrease with an increase in θ ,
whereas the mean initially increases fast with an increase in α , after which it slowly
starts dropping with a further increase in α . As for the variance, it initially increases

162
with an increase in α , after which it drops and then increases again with further
increase in α . The behaviour is contrary for the variance of the Unit Lindley
distribution, which initially increases and then steeply decreases with θ (Mazucheli,
2019).

Fig. 4 (a) Plot of skewness as a function of α for θ =0.5 (b) Plot of skewness as a
function of θ for α =0.5 (c) Plot of kurtosis as a function of α for θ =0.5 (d) Plot of
kurtosis as a function of θ for α =0.5

It is evident from Fig. 4 that both the skewness and kurtosis decrease in an
exponential manner with an increase in α . On the other hand, the skewness and
kurtosis initially increase, followed by a brief period of decrease, after which it again
starts increasing with an increase in θ . On the contrary, the skewness of the Unit
Lindley distribution exponentially increases with an increase in θ and the kurtosis
initially decreases then increases with θ (Mazucheli, 2019).
3.2 Hazard Rate Function

163
The hazard rate function measures the instantaneous risk of an event occurring
at a particular time, given that it has survived up to that time. It is denoted by 𝑕(𝑡)
and is calculated as
𝑓(𝑡)
𝑕(𝑡) =
1 − 𝐹(𝑡)
The hazard rate function of the Topp-Leone Unit Lindley distribution is
𝜃𝑥 2 𝑚𝜃𝑥 𝛼
𝛼−1
2𝛼𝜃 𝑚 1 𝜃𝑥
𝑕 𝑇𝐿𝑈𝐿 (𝑥) = 𝑒 −𝑙−𝑥 [{1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] [2 −
1+𝜃 (1−𝑥)𝑛
𝑚𝜃𝑥 𝛼
𝛼−1
𝜃𝑥 2
{1 − .1 − (1+𝜃)(𝑥−1)/ 𝑒 −𝑙−𝑥 } ] (11)

Fig. 5 shows the hazard rate function plots of the Topp-Leone Unit Lindley
distribution for different values of α (for fixed θ ) and different values of θ (for
fixed α ).

Fig. 5 (a) Hazard rate function plot for different values of α for θ =0.5 (b) Hazard
rate function plot for different values of θ for α =0.5

It is clear from the above figure that the hazard rate function exhibits a
negatively skewed behaviour for both θ and α , whereas the hazard rate function of
the Unit Lindley distribution displays an exponentially increasing pattern with
variation in θ (Mazucheli, 2019).
4. Maximum Likelihood Estimation of parameters and simulation study

164
4.1 Maximum Likelihood Estimation of parameters
The maximum likelihood estimators of parameters of a distribution are obtained
by maximizing the log-likelihood function of the sample with respect to variations in
the parameters.
Let 𝑥1 , 𝑥2 , … , 𝑥𝑛 be a sample from the Topp-Leone Unit Lindley distribution.
Then the log-likelihood function of the sample is given by
log 𝐿 =
𝑚𝜃𝑥𝑖 𝛼
2𝛼𝜃 𝑚 𝑥𝑖 𝜃𝑥 2 −
𝑛 log . / − 3 ∑𝑛𝑖=1 log(1 − 𝑥𝑖 ) − 𝜃 ∑𝑛𝑖=1 + ∑𝑛𝑖=1 log 41 − 81 − .1 − (1+𝜃)(𝑥𝑖 / 𝑒 𝑙−𝑥𝑖
9 5+
1+𝜃 1−𝑥𝑖 𝑖 −1)
𝑚𝜃𝑥𝑖 𝛼
𝜃𝑥𝑖 2 −
(𝛼 − 1) ∑𝑛𝑖=1 log 81 − .1 − / 𝑒 𝑙−𝑥𝑖 9 + (𝛼 −
(1+𝜃)(𝑥𝑖 −1)
𝑚𝜃𝑥𝑖 𝛼
𝜃𝑥 2 −
1) ∑𝑛𝑖=1 log 42 − 81 − .1 − (1+𝜃)(𝑥𝑖 / 𝑒 𝑙−𝑥𝑖
9 5 (12)
𝑖 −1)

The maximum likelihood estimators are obtained by solving the following


equations:
𝜕 𝜕
𝜕𝜃
log 𝐿 = 0 and 𝜕𝛼 log 𝐿 = 0
Since the equations cannot be solved analytically owing to the non-closed form
of the expressions, the solutions to them are obtained using the RStudio software,
version 4.2.2 via the optim() function.

4.2 Simulation study


A Monte Carlo simulation is conducted in this section to first evaluate and then
assess the finite sample behaviour of the maximum likelihood estimates (MLE) of
the parameters; their average bias (AB) and Root Mean Square Error (RMSE).
Samples of sizes 𝑛 = 50, 100, 250, 500, 800 are generated from the Topp-Leone
Unit Lindley distribution and the simulation experiment was replicated 10000 times.
To generate random variables from the Topp-Leone Unit Lindley distribution, the
following steps have been undertaken:
Step 1: A random variable from Uniform (0,1) distribution is generated, say Ui.
Step 2: Ui is equated to the c.d.f. 𝐺𝑇𝐿𝑈𝐿 (𝑥𝑖 ) in (2.8) to obtain the equation
2 𝛼
𝜃𝑥𝑖 2𝜃𝑥𝑖

𝑈𝑖 = 81 − (1 − * 𝑒 1−𝑥𝑖9
(1 + 𝜃)(𝑥𝑖 − 1)
Step 3: The above equation is solved for 𝑥𝑖 , which is a random variable from the
Topp-Leone Unit Lindley distribution. Since the above equation is difficult to be
solved analytically, the solution is obtained using the RStudio software, version
4.2.2 via the inverse() function of the GoFKernel package.
Step 4: Steps 1-3 are repeated the desired number of times to get a sample of the
required size.
Table 1 demonstrates the MLE of the parameters of the Topp-Leone Unit
Lindley distribution along with their AB and RMSE:
Table 1 MLE of the parameters of the Topp-Leone Unit Lindley distribution and
their AB and RMSE
165
α θ
Parameters n MLE AB RMSE MLE AB RMSE
50 0.338 0.112 0.015 0.055 -0.002 0.0007
100 0.167 0.054 0.005 0.027 -0.001 0.0002
α = 0.8 250 0.066 0.021 0.0013 0.010 -0.0005 0.00003
θ = 0.2 500 0.033 0.010 0.0004 0.005 -0.0002 0.00001
800 0.020 0.006 0.0002 0.003 -0.0001 0.000006
50 0.816 0.090 0.013 0.771 -0.136 0.020
100 0.404 0.042 0.004 0.381 -0.072 0.007
α = 1.2 250 0.161 0.016 0.001 0.151 -0.030 0.002
θ = 1.5 500 0.079 0.007 0.0003 0.074 -0.015 0.0006
800 0.049 0.004 0.0001 0.046 -0.009 0.0003

It is clear from Table 1 that the absolute value of the average bias of both α and
θ decrease with an increase in n, although the bias of θ is negative for both
combination of the parameters. As expected, the RMSE of both α and θ decrease
with an increase in n. This verifies the consistency of the MLE of the parameters
based on the generated samples.

5. Concluding Remarks and Future Scope


The Topp-Leone Unit Lindley distribution is established in this paper, which is
obtained by considering the Unit Lindley distribution as the baseline distribution in
the Topp-Leone family of distributions. It is found to be expressible as the weighted
sum of the exponentiated-Unit Lindley distribution. Some basic statistical properties
of this distribution are explored. It is seen that the behaviour of the moments and
hazard rate function of the baseline distribution changes with respect to its parameter
θ when the new parameter α is introduced. Moreover, the simulation study
confirms that the maximum likelihood estimates of the parameters of the distribution
are consistent in nature, thereby indicating towards the accuracy of the estimation
procedure.
As future objectives, derivation of various other properties of the Topp-Leone
Unit Lindley distribution such as the distribution of the order statistics, stress-
strength reliability, estimation of parameters using other estimation methods and
application of the distribution to real-life data can be taken up.

REFERENCES:
1. Al-Shomrani, A., Arif, O., Shawky, A., Hanif, S., & Shahbaz, M.Q. (2016).
Topp–Leone family of distributions: some properties and application. Pakistan
Journal of Statistics and Operation Research, 12(3), 443–451.
2. Alza-Atreh, A., Lee, C., & Famoye, F. (2013). A new method for generating
families of continuous distributions. Metron, 71, 63-79.
3. Cordeiro, G.M., & De Castro, M. (2011). A new family of generalized
distributions. Journal of Statistical Computation and Simulation, 81, 883-898.

166
4. Eugene, N., Lee, C., & Famoye, F. (2002). Beta-normal distribution and its
applications. Commun. Stat. Theory Methods, 31, 497-512.
5. Handique, L. (2018). New families of lifetime distributions and their
properties (Doctoral thesis). Dibrugarh, Assam, India: Dibrugarh University,
Department of Statistics.
6. Khaoula, A., Seddik-Ameur, N., Ahmad, A.E.A., & Khaleel, M. A. (2022).
The Topp-Leone Extended Exponential Distribution: Estimation Methods and
Applications. Pakistan Journal of Statistics and Operation Research, 18(4),
817-836.
7. Marshall, A., & Olkin, I. (1997). A new method for adding a parameter to a
family of distributions with applications to the exponential and Weibull
families. Biometrika, 84, 641-652.
8. Mazucheli, J., Menezes, A.F.B., & Chakraborty, S. (2019). On the one
parameter unit-Lindley distribution and its associated regression model for
proportion data. Journal of Applied Statistics. 46(4), 700-714.
9. Nadarajah, S., & Kotz, S. (2003). Moments of some J-shaped distributions.
Journal of Applied Statistics, 30, 311-317.
10. Shaw, W., & Buckley, I. (2007). The alchemy of probability distributions:
Beyond Gram-Charlier expansions and a skewkurtotic- normal distribution
from a rank transmutation map. (Research report) London, UK: King's
College, Department of Mathematics.
11. Tahir, M.H., & Nadarajah, S. (2015). Parameter induction in continuous
univariate distributions: Well-established G families. Anais da Academia
Brasileira de Ciências (Annals of the Brazilian Academy of Sciences), 87,
539-568.
12. Tuoyo, D. O., Opone, F. C., & Ekhosuehi, N. (2021). The Topp-Leone Weibull
Distribution: Its Properties and Applications. Earthline Journal of
Mathematical Sciences, 7(2), 381-401.
13. Zografos, K., & Balakrishnan, N. (2009). On families of beta- and generalized
gamma-generated distributions and associated inference. Statistical
Methodology, 6, 344-362.

167
A Semi-Circular Exponential Distribution Induced by
Inverse Stereographic Projection:
Properties and Application
Imliyangba1*, Bhanita Das2 and Seema Chettri3
1
Department of Statistics, NEHU, Shillong, 793022, India, imlijamir93@[Link]
2
Department of Statistics, NEHU, Shillong, 793022, India,
bhanitadas83@[Link]
3
Department of Statistics, NEHU, Shillong, 793022, India,
seemachettri094@[Link]

Abstract:
The semi-circular linear exponential distribution is a continuous random
variable introduced in the context of circular statistics using the method of modified
inverse stereographic projection (ISP). In this paper, we derived the density function
and distribution function of the proposed distribution. The graphical representations
with different parameter values are discussed. The first two trigonometric moments
are derived and the new distribution is extended to construct stereographic l-axial
linear exponential distribution. Method of maximum likelihood estimation is used
for estimating the unknown parameters and a simulation study is carried out to check
the consistency of the estimator. A real-life semi-circular data set is used to
demonstrate the applicability of the proposed distribution.
Keywords: Inverse Stereographic Projection, Semi-Circular Distributions, Graphical
Representations, Trigonometric Moments, Maximum Likelihood
Estimation.

1. Introduction:
In real-life, huge amounts of data exist related to many medical, ecological, and
environmental variables which are often measured in terms of angles wherein their
range is defined in ,0, 𝜋). To deal with these kinds of data, researchers have done
numerous works on directional statistics and have introduced several semi-circular
distributions. For example, when an aircraft is lost, but its departure points and its
initial headings are known, a random variable with values on a semi-circle is
sufficient for such a problem. Similarly, when a sea turtle emerges from the sea in
search of a nesting site on a dry land, a semi-circular random variable is sufficient to
model such data. These data are called axial or semi-circular data.
Jones (1968) and Guardiola (2004) pointed out that the angular data does not
require full circular models for modelling, in some cases. Since then, several

*Corresponding Author: imlijamir93@[Link]

168
researchers (Kim, 2008; Yedlapalli et al., 2013; Pramesti 2015; Yedlapalli et al.,
2016; Subrahmanyam et al., 2017; Abuzaid, 2018; Rambli et al., 2019; Yedlapalli et
al., 2020; Iftikhar & Hanif, 2022; Yedlapalli et al., 2023), have contributed in the
field of circular statistics, particularly semi-circular distributions. Recently, Awad &
Mohammed (2024) introduced the stereographic semi-circular three parameter
Lindley distribution for modelling the semi-circular data. Still there exist a gap on
the study of semi-circular distributions which needs more attention in the field of
directional statistics. The aim of the present paper is to contribute towards filing up
this gap. In this article, the authors are motivated to propose a new semi-circular
distribution, denominated as stereographic semi-circular linear exponential (SSLE)
distribution by inducing the method of modified ISP. Explicit expressions for
trigonometric moments are derived and the proposed distribution is extended to
stereographic l-axial linear exponential distribution.
The rest of the paper is organized as follows: Section 2 describes the
methodology of the modified ISP. In section 3, the stereographic semi-circular linear
exponential distribution is introduced. The probability density function ( pdf) and the
cumulative distribution function (cdf) of the proposed distribution are also
presented. Section 4 provides the graphical (linear and circular) representations of
the SSLE distribution and in section 5, we derived the first two trigonometric
moments of the SSLE distribution and it is extended to stereographic l-axial linear
exponential distribution for modeling axial data in section 6. The estimation of
parameter using method of maximum likelihood is discussed in section 7. In section
8, a simulation study is carried out to show that the obtained maximum likelihood
estimator is consistent. An application to a real-life semi-circular data set is
presented in Section 9 and final conclusions and discussions are given in section 10.
2. METHODOLOGY OF MODIFIED INVERSE STEREOGRAPHIC PROJECTION:
Yedlapalli et al., (2013) stated that the probability distributions (both circular
and linear) can be generated by applying stereographic projection, which yields a
one-to-one correspondence between the points on the unit circle and those on the
real line. ISP is defined by a one-to-one mapping given by
𝜃
𝑇(𝜃) = 𝑥 = 𝑢 + 𝑣 tan ( *
2
where, 𝑥 𝜖 (−∞, ∞), 𝜃 𝜖 (−𝜋, 𝜋), 𝑢 𝜖 , 𝑣 > 0
If x is random variable defined on the interval (−∞, ∞) with pdf as 𝑓(𝑥) and
cdf as 𝐹(𝑥), then
𝑥−𝑢
⟹ 𝑇 −1 (𝑥) = 𝜃 = 2 tan−1 . /
𝑣
is a random point on a unit circle. Let 𝐺(𝜃) and 𝑔(𝜃) represent the cdf and pdf
of this random point 𝜃, respectively, then 𝐺(𝜃) and 𝑔(𝜃) can be derived in terms of
𝐹(𝑥) and 𝑓(𝑥) using the following transformation. For 𝑣 > 0,
𝜃
𝐺(𝜃) = 𝐹 (𝑢 + 𝑣 tan ( ** = 𝐹(𝑥(𝜃))
2

169
𝜃
1 + tan2 .2 / 𝜃
𝑔(𝜃) = 𝑣 : ; 𝑓 (𝑢 + 𝑣 tan ( **
2 2
When 𝑢 = 0 and 𝑣 = 1, the ISP is modified and the above pdf i.e.,
𝑔(𝜃) becomes
𝜃
1 + tan2 .2 / 𝜃 1 𝜃 𝜃
𝑔(𝜃) = : ; 𝑓 (tan ( ** = sec 2 ( * 𝑓 (tan ( ** (1)
2 2 2 2 2
This is the pdf of the stereographic semi-circular distribution induced by
modified ISP.

3. SSLE DISTRIBUTION USING MODIFIED ISP:


Sah (2021) introduced a linear exponential distribution which is based on the
product of linear and exponential functions. The pdf and cdf of the linear
exponential distribution are given below:
𝜆2
𝑓(𝑥) = 4 5 *𝜆 + 𝑥+𝑒 −𝜆𝑥 (2)
1 + 𝜆2
1 + 𝜆𝑥 + 𝜆2 −𝜆𝑥
𝐹(𝑥) = 1 − 8 9𝑒
(1 + 𝜆2 )
Where, 𝑥 > 0, 𝜆 > 0
Therefore, using equation (1) and equation (2), the pdf of the SSLE distribution
can be derived as
1 𝜃 𝜆2 𝜃 𝜃
−𝜆 tan( *
𝑔(𝜃) = sec 2 ( * 4 5 {𝜆 + tan ( *} 𝑒 2 (3)
2 2 1 + 𝜆2 2
And the corresponding cdf is defined as
𝜃
1 + 𝜆 tan . / + 𝜆2 −𝜆 tan(𝜃*
𝐺(𝜃) = 1 − > 2 ?𝑒 2 (4)
(1 + 𝜆2 )

where, 𝜃 𝜖 ,0, 𝜋), 𝜆 > 0.

4. GRAPHICAL REPRESENTATIONS OF SSLE DISTRIBUTION:


The following figures represents the pdf of the proposed distribution with
different values of the parameter. Fig. 1 displays the graph of linear representations
for SSLE distribution where the values of parameter in Fig. 1(b) are higher than Fig.
1(a).

170
(a) (b)
Fig. 1: Pdf of SSLE distribution for 𝜆

Whereas, the circular representations for SSLE distribution are given in Fig. 2
where the values of 𝜆 in Fig. 2(a) is lower than that of Fig. 2(b).

5. Trigonometric Moments:
In this section we derived the characteristic function and trigonometric moments
for SSLE distribution. The characteristic function of a semi-circular distribution is
defined as:
𝜋
𝑖𝑝𝜃
𝜙(𝜃) = 𝐸(𝑒 ) = ∫ 𝑒 𝑖𝑝𝜃 𝑔(𝜃) 𝑑𝜃
0

(a) (b)
Fig. 2: Circular representations of SSLE distribution

Therefore, the characteristic function of SSLE distribution is obtained as:


1 𝜆2 𝜋
2
𝜃 𝜃 𝜃
−𝜆 tan( * 𝑖𝑝𝜃
2 𝑒
𝜙(𝜃) = 4 5 ∫ sec ( * {𝜆 + tan ( *} 𝑒 𝑑𝜃
2 1 + 𝜆2 0 2 2
According to the definition of the trigonometric moments 𝜙𝑝 = 𝛼𝑝 + 𝑖𝛽𝑝 ; 𝑝 =
𝜋
±1, ±2, … with 𝛼𝑝 = 𝐸,𝑐𝑜𝑠(𝑝𝜃)- = ∫0 𝑐𝑜𝑠(𝑝𝜃)𝑔(𝜃) 𝑑𝜃 and
𝜋 𝑡𝑕
𝛽𝑝 = 𝐸,𝑠𝑖𝑛(𝑝𝜃)- = ∫0 𝑠𝑖𝑛(𝑝𝜃)𝑔(𝜃) 𝑑𝜃 , being the 𝑝 order cosine and sine moments of

171
the random angle 𝜃, respectively and are required to study the population
characteristics.
The trigonometric moments 𝛼𝑝 = ∫0𝜋 𝑐𝑜𝑠(𝑝𝜃)𝑔(𝜃) 𝑑𝜃 and 𝛽𝑝 = ∫0𝜋 𝑠𝑖𝑛(𝑝𝜃)𝑔(𝜃) 𝑑𝜃 , for
𝑝 = 1, 2 of SSLE distribution can be derived by applying Meijer's G-function
(Gradshteyn & Ryzhik, 2007) which are given as follows
1
𝜆3 1 31 𝜆2 −
2 ,− 𝜆2 1 31 𝜆2 −1
𝛼1 = 1 − 𝐺13 ( | 𝐺 ( | 1+ (5)
(1 + 𝜆2 ) √𝜋 4 1 1 (1 + 𝜆2 ) √𝜋 13 4 −1, 0,
− , 0, 2
2 2
1
𝜆3 1 31 𝜆2 0 𝜆2 1 31 𝜆2 −
2 ,
𝛽1 = 𝐺13 ( | 1+ + 𝐺13 ( | (6)
(1 + 𝜆 ) √𝜋
2 4 0, 0, (1 + 𝜆 ) √𝜋
2 4 1 1
2 − , 0,
2 2

𝛼2 =
3
4𝜆𝑛 1 𝜆𝑚 − 4𝜆𝑚 1 𝑚 −2
31 𝜆
1 + (1+𝜆𝑚 ) 𝐺 31 : | 1 2 1; + (1+𝜆𝑚 ) 𝐺13 4 |−1, 0, 15 −
√𝜋 13 4 √𝜋 4
− , 0, 2
2 2
1
4𝜆𝑛 1 𝜆𝑚
− 4𝜆𝑚 1 𝜆𝑚 −1
2
𝐺 31
(1+𝜆𝑚 ) √𝜋 13
: | 1
31
1; − (1+𝜆𝑚 ) √𝜋 𝐺13 4 |−1, 0, 15 (7)
4 4
− , 0, 2
2 2
1
2𝜆3 1 31 𝜆2 0 2𝜆2 1 31 𝜆2 −
2 ,
𝛽2 = 𝐺 ( | 1+ + 𝐺 ( |
(1 + 𝜆2 ) √𝜋 13 4 0, 0, (1 + 𝜆2 ) √𝜋 13 4 1 1
2 − , 0,
2 2
3
4𝜆3 1 31 𝜆2 −1 4𝜆2 1 31 𝜆2 −
2 ,
− 𝐺 ( | 1+ − 𝐺 ( | (8)
(1 + 𝜆2 ) √𝜋 13 4 0, 0, (1 + 𝜆2 ) √𝜋 13 4 1 1
2 − , 0,
2 2
where,

𝑢2𝑣+2𝑄−2 𝑢2 𝜇2 1−𝑣
∫ 𝑥 2𝑣−1 (𝑢2 + 𝑥 2 )𝑄−1 𝑒 −𝜇𝑥 𝑑𝑥 = 31
𝐺13 ( | 1+
0 2√𝜋Γ(1 − 𝑄) 4 1 − 𝑄 − 𝑣, 0,
2
𝜋 𝑚 𝑚 1−𝑣
31 𝑢 𝜇
for |arg 𝑢𝜋| < 2
, 𝑅𝑒𝜇 > 0 and 𝑅𝑒𝑣 > 0 and 𝐺13 4
4
|1 − 𝑄 − 𝑣, 0, 15 is called
2
the Meijer‘s G-function.
Similarly, we can obtain the higher order of the trigonometric moments for the
proposed distribution by increasing the value of 𝑝 in 𝛼𝑝 and 𝛽𝑝 . The mean direction
of SSLE distribution is defined as:
𝛽1
𝜇 = 𝑡𝑎𝑛−1 ( *
𝛼1
The mean resultant length 𝜌 is:
𝜌 = √𝛼12 + 𝛽12
he circular variance and circular standard deviation are as follows:
𝜈 = 1 − 𝜌 = 1 − √𝛼12 + 𝛽12 and 𝜍 = √𝑙𝑜𝑔(𝛼12 + 𝛽12 )

172
The circular skewness and circular kurtosis of SSLE distribution are obtained as:
𝛽2∗
𝛾1 = 3
(1 − 𝜌)2
𝛼2∗ − 𝜌4
𝛾2 =
(1 − 𝜌)2
where 𝛼1 and 𝛽1 are defined in equations 5 and 6 respectively. Mardia & Jupp
(2000) had derived the expressions for central trigonometric moments 𝛼𝑝∗ and 𝛽𝑝∗
from the values of the characteristic function 𝜙𝑝 at 𝑝 = 1 and 𝑝 = 2.
Table I. Population characteristics of SSLE distribution for various values of 𝜆.
𝜆 0.5 1.0 1.5 2.0 2.5 5.0
Mean direction 𝜇 0.0354 0.1075 0.2287 0.3568 0.3871 0.6810
Resultant
𝜌 0.9997 0.9961 0.9916 0.9845 0.9818 0.9474
length
𝛼1 0.9991 0.9903 0.9658 0.9225 0.9091 0.7360
𝛼2 0.9962 0.9617 0.8671 0.7082 0.6619 0.1303
Trigonometric
moments 𝛽1 0.0354 0.1068 0.2248 0.3439 0.3706 0.5965
𝛽2 0.0706 0.2097 0.4272 0.6167 0.6510 0.7938

𝛼1 0.9997 0.9961 0.9916 0.9845 0.9818 0.9474
Central ∗
𝛼2 0.9987 0.9843 0.9666 0.9390 0.9284 0.8036
trigonometric ∗
moments 𝛽1 0 0 0 0 0 0

𝛽2 -1.9789 -0.0004 0.0004 0.0026 0.0027 0.0369
Circular
𝑣 0.0007 0.0087 0.0187 0.0347 0.0408 0.1197
variance
Skewness 𝛾1 -0.0042 -0.0435 0.0295 0.1129 0.1987 0.4744
Kurtosis 𝛾2 -2.9110 -4.6062 -0.0002 -0.0004 -0.0006 -0.0022

It is clear from Table I that:


 The mean direction increased with the increase in the values of 𝜆 whereas,
the resultant length decreased with increasing values of λ .
 The circular variance also increased when the values of λ increased.

6. STEREOGRAPHIC-L-AXIAL LINEAR EXPONENTIAL DISTRIBUTION:


The SSLE distribution is extended to the l-axial distribution, which is applicable
2𝜋
to any arc of arbitrary length say 𝑙 for 𝑙 ∈ 𝑁. It is possible to extend the SSLE
distribution to construct the stereographic-l-axial linear exponential distribution. We
considered the density function of SSLE distribution to construct the stereographic-
2𝜃
l-axial linear exponential distribution using the transformation 𝜙 = 𝑙 , 𝑙 = 1,2, …,.The

173
probability density function of stereographic-l-axial linear exponential distribution is
defined as:
1 𝜙𝑙 𝜆2 𝜙𝑙 𝜙𝑙
−𝜆 tan( *
𝑔(𝜙) = sec 2 ( * 4 5 {𝜆 + tan ( *} 𝑒 4 (9)
2 4 1 + 𝜆2 4
2𝜋
0<𝜙< 𝑙
, 𝜆 > 0, 𝛼 > 0 and 𝑙 = 1,2, …,
Case 1: When 𝑙 = 1, in equation (9), we get the density function
1 𝜙 𝜆2 𝜙 𝜙
−𝜆 tan( *
𝑔(𝜙) = sec 2 ( * 4 5 {𝜆 + tan ( *} 𝑒 4
2 4 1 + 𝜆2 4
0 < 𝜙 < 2𝜋, 𝜆 > 0 and 𝛼 > 0.
Case 2: When 𝑙 = 2 in equation (9), we get:
1 𝜙 𝜆2 𝜙 𝜙
−𝜆 tan( *
𝑔(𝜙) = sec 2 ( * 4 5 {𝜆 + tan ( *} 𝑒 2
2 2 1 + 𝜆2 2
Which is the same as that of SSLE distribution.
Case 3: When 𝑙 = 4 in equation (9), we get:
1 𝜆2
𝑔(𝜙) = sec 2(𝜙) 4 5 *𝜆 + tan(𝜙)+𝑒 −𝜆 tan(𝜙)
2 1 + 𝜆2

7. Maximum Likelihood Estimation:


Among many techniques for estimating the parameters, a common method of
parameter estimation is the maximum likelihood estimation. The maximum
likelihood estimation determines many desirable properties which includes
consistency, invariance, asymptotic efficiency, etc.
Then, the likelihood function of SSLE distribution is given by
𝑛 𝑛
𝜆2 −𝜆 ∑𝑛
𝜃𝑖
𝑖=𝑙 tan( 2 *
1 𝜃𝑖
𝐿=4 5 𝑒 ∑ {( * (𝜆 + tan ( **}
(1 + 𝜆 )
2 1 + cos 𝜃𝑖 2
𝑖=1

Therefore, the log-likelihood function is


𝑛
𝜃𝑖
log 𝐿 = 2𝑛 log 𝜆 − 𝑛 log (1 + 𝜆2 ) − 𝜆 ∑ tan ( *
2
𝑖=1
𝑛
1 𝜃𝑖
+ ∑ {log ( * + log (𝜆 + tan ( **} (10)
1 + cos 𝜃𝑖 2
𝑖=1

Differentiating partially w.r.t. 𝜆 we get


𝑛 𝑛
𝜕 2𝑛 2𝑛𝜆 𝜃𝑖 1
log 𝐿 = − − ∑ tan ( * + ∑ > ? (11)
𝜕𝜆 𝜆 (1 +𝜆 2 ) 2 𝜃𝑖
𝑖=1 𝑖=1 𝜆 + tan . /
2

174
Equating the above equations i.e., equations (11) to zero, we get the estimates of
𝜆.
𝜕
log 𝐿 = 0
𝜕𝜆
Since the maximum likelihood equations cannot be solved analytically,
therefore, a numerical technique is to be employed to get a solution for 𝜆. We use
statistical packages in R to get the MLE of the unknown parameter which is
explained through simulation study.

8. SIMULATION STUDY:
A simulation study was performed to obtain the MLE of the unknown
parameters i.e., 𝜆. Further, a random sample of different sizes (25, 50, 100, 200,
400, 500, 600, 800 and 1000) is generated for different values of 𝜆, and replicated
the program 𝑁 = 1000 times to get the MLE 𝜆̂. For 𝜆 = 3.8
Table II. Estimated values of 𝜆, and average values of Bias and MSE for 𝜆̂
n 𝜆 Bias MSE
25 3.9283 0.005132 0.000658
50 3.8890 0.001780 0.000158
100 3.8716 0.000716 0.000051
200 3.8707 0.000354 0.000025
400 3.8683 0.000171 0.000012
500 3.8418 0.000084 0.000003
600 3.8233 0.000039 0.000001
800 3.8139 0.000017 0.000000
1000 3.8024 0.000002 0.000000

From Table II, it is observed that, as the sample size increases, the estimated
values of the parameters approach very close to the true values of parameters used in
the simulation. Moreover, we can also see that the Bias and MSE values for the
estimated parameters decrease and tend to zero as the sample size increases. This
shows the adequacy of the estimation technique.

9. APPLICATION TO SEMI-CIRCULAR DATASET:


The construction of a new distribution leads to the next logical step, i.e., to fit
the new distribution as a model to the live data. The reason behind this step is to find
out how best the new distribution will fit a given live data set and compare its
relative performance with respect to the other distributions in the class. Therefore, in
this section, we will consider a dataset which is obtained from the images of the
posterior segment of the eyes of 23 patients. The images were taken using the
anterior segment optical coherence tomography (AS-OCT) at a glaucoma clinic at

175
the University of Malaya Medical Centre, Malaysia, to study the applicability of the
SSLE distribution. The data set is available in Abuzaid (2018).
Fig. 3 shows the circular plot of the images of posterior segments of the eyes of
23 patients.
Table III. Summary of statistics for SSLE, SCQL, SCSD and SCED distributions for
eye data
Distributions Parameters Mle -L AIC CAIC BIC HQIC P-value
𝜆
SSLE
1.0189 13.344 32.687 33.951 36.094 33.544 0.1753
𝜆 0.7228
SCQL 𝜇 0.1041 28.666 63.332 64.596 66.739 64.189 0.0002
𝜍 0.7016
SCSD 𝛼 1.2178 21.414 44.829 45.019 45.964 45.114 0.0059
𝜆 0.8986
SCED 22.915 49.830 50.431 52.102 50.402 0.0016
𝛼 0.0006

Fig. 3: Circular plot of the posterior corneal curvature measurement

Fig. 4 shows the fitted densities of SSLE distribution along with all competing
models. Various statistics, such as, negative log-likelihood value (-L), AIC, CAIC,
BIC, HQIC and P-value were calculated and the results are shown in Table III. We
can also observe that the (-L), AIC, CAIC, BIC and HQIC have the lowest values for
the SSLE distribution, and p-value for SSLE distribution has the highest value, as
compared to that of stereographic semi-circular Quasi Lindley (SCQL),
stereographic semi-circular Shanker (SCSD) and semi-circular exponential (SCED)
distributions. Hence, we can say that the proposed distribution fits well for the
given dataset in comparison to other considered distributions.

176
Fig. 4: Fitted densities of SSLE distribution and other distributions on the posterior
corneal curvature measurement

10. Conclusion:
In this article, we have discussed the method of modified inverse stereographic
projection and introduced a new linear exponential distribution, named as
stereographic semi-circular linear exponential distribution. The density function,
distribution function, along with the linear and circular representations of SSLE
distribution are discussed graphically. The maximum likelihood estimation method
was used to obtain the estimated parameters. A simulation study was carried out,
showing the adequacy of the estimation technique. The flexibility of the SSLE
distribution was shown by applying it to a real-life dataset which proved that the
proposed distribution is suitable to fit the angular data.

11. references
1. Abuzaid, A.H. (2018). A half circular distribution for modeling the posterior
corneal curvature. Communications in Statistics-Theory and Methods,
47(13), 3118-3124.
2. Awad, N.T., & Mohammed, S.F. (2024). On stereographic semi-circular three
parameter lindley distribution. International Journal of Engineering and
Information Systems (IJEAIS), 8(1), 84-91.
3. Gradshteyn, I.S., & Ryzhik, I.M. (2007). Table of Integrals, Series and
Products. 7th edition, Academic Press.
4. Guardiola, J.H. (2004). The Semicircular Normal Distribution, Ph.D.
Dissertation, Baylor University, Institute of Statistics.
5. Iftikhar, A., Ali, A., & Hanif, M. (2022). Half circular modified Burr− III
distribution, application with different estimation methods. PLoS ONE,
17(5), e0261901.

177
6. Jones, T. A. (1968). Statistical analysis of orientation data. Journal of
Sedimentary Petrology, 38, 61-67.
7. Kim, H.M. (2008). New family of the t distributions for modeling
semicircular data. Communications of the Korean Statistical Society, 15(5),
667-674.
8. Mardia, K.V., & Jupp, P.E. (2000). Directional Statistics. Chichester: John
Wiley.
9. Pramesti, G. (2015). The stereographic semicircular chi square models. Far
East Journal of Theoretical Statistics, 51(3), 119-128.
10. Rambli, A., Mohamed, I., Shimizu, K., & Ramli, N.M. (2019). A half-circular
distribution on a circle. Sains Malaysiana, 48(4), 887–892.
11. Sah, B.K. (2021). Linear-exponential distribution. International Journal of
Statistics and Applied Mathematics, 6(4), 109-115.
12. Subrahmanyam, P.S., Rao, A.V.D., & Girija, S.V.S. (2017). On stereographic
semicircular new weibull pareto model. IJIRST–International Journal for
Innovative Research in Science & Technology, 3(11), 267–276.
13. Yedlapalli, P., Girija, S.V.S., & Rao, A.V.D. (2013). On construction of
stereographic semicircular models. Journal of Applied Probability and
Statistics, 8(1), 75-90.
14. Yedlapalli, P., Girija, S.V.S., & Rao, A.V.D. (2016). The semicircular
reflected gamma distribution. i-manager‘s Journal on Mathematics, 5l(1l),
39-46.
15. Yedlapalli, P., Girija, S.V.S., Rao, A.V.D., & Sastry, K.L.N. (2020). A new
family of semicircular and circular arc tan-exponential type distributions.
Thai Journal of Mathematics, 18(2), 775-781.
16. Yedlapalli, P., Kishore, G.N.V., Boulila, W., Koubaa, A., & Mlaiki, N. (2023).
Toward enhanced geological analysis: a novel approach based on transmuted
semicircular distribution. Symmetry, 15(11), 2030.

178
Cardiovascular Failure Risk Prediction: A Biostatistical
Perspective

Dr. Suman Jaiswal1, Dr. Vaishali Saxena2 and Bhavya Jain3


1
Assistant Professor, Department of Statistics
2
Biostatistician, CDSCO(HQ), New Delhi
3
Associate Consultant, EY GDS, Gurgaon

Abstract
Globally, cardiovascular failure continues to be a major cause of death,
highlighting the importance of early risk assessment and prevention. This study
predicts the risk of cardiovascular failure in at-risk people using sophisticated
biostatistical techniques. We created predictive models utilising logistic regression,
machine learning algorithms, and survival analysis using a large dataset of patient
health indicators, including blood pressure, cholesterol levels, and lifestyle factors.
The results show how statistical algorithms can reliably predict cardiovascular
failure and identify important risk variables. In order to improve patient outcomes,
our research offers clinicians insightful information that supports early diagnosis and
individualised treatment plans.
The main goal of this study is to forecast cardiovascular failure by analysing
cardiovascular and physical risk factors and creating regression and survival models.
Kaggle datasets were used for this study, and they were cleaned and quality-checked
before analysis. In addition to statistical tests (T-tests, F-tests, and Chi-square tests),
regression models, survival analysis, and machine learning algorithms, exploratory
data analysis (EDA) techniques such as bar graphs, scatterplots, heatmaps, and box
plots were employed. Regression and survival models were created to find the best
fit for predicting the key characteristics that contribute to cardiovascular failure. The
findings offer insightful information about the statistical correlations and hypothesis
testing of important cardiovascular risk factors.

Key Points
1. Heart Failure Risk: The chance of being diagnosed with heart failure is 1 in
5 by the age of 40, and the risk rises with age.
2. Global Impact: Approximately 5.8 million Americans and 23 million people
globally suffer from heart failure, making it a serious public health issue.


Corresponding author: Assistant Professor, Department of Statistics

179
3. Key Risk Factors: Heart failure is far more likely to occur in people with
ischaemic heart disease, hypertension, smoking, obesity, and diabetes, and
these conditions are also associated with worse outcomes.
4. Mortality Rates: Although heart failure treatment has improved, the
condition still has a high 5-year mortality rate, similar to many
malignancies.
5. Cardiovascular failure must be diagnosed quickly in order to save patients'
lives and stop further harm.

Introduction
Heart failure — sometimes known as congestive heart failure — occurs when
the heart muscle doesn't pump blood as well as it should. Affecting one or both sides
of the heart, this serious, incurable condition is most often caused by other heart
conditions and diseases, such as coronary heart disease (CAD), high blood pressure,
diabetes; or conditions like inflation of heart muscles etc. Coronary arteries are
responsible for the supply of blood to the heart itself. When this happens, blood
often backs up and fluid can build up in the lungs, causing shortness of breath.
Certain heart conditions, such as narrowed arteries in the heart (coronary artery
disease) or high blood pressure, gradually leave the heart too weak or stiff to fill and
pump blood properly. Tests that may be done to diagnose heart failure may include:
• Blood tests
• Chest X-ray
• Electrocardiogram (ECG)
• Echocardiogram
• CT scan of the heart
• Heart MRI scan
• Myocardial biopsy
Initially, traditional investigation techniques were used for the identification of
heart disease, however, they were found complex. Owing to the non-availability of
medical diagnosing tools and medical experts specifically in undeveloped countries,
diagnosis and cure of heart disease are very complex. However, the precise and
appropriate diagnosis of heart disease is very imperative to prevent the patient from
more damage.
According to an article by Nature Reviews Cardiology, ―HF represents a
considerable burden to the health-care system, responsible for costs of more than
$39 billion annually in the USA alone, and high rates of hospitalizations,
readmissions, and outpatient visits.‖ Proper treatment can improve the signs and
symptoms of heart failure and may help some people live longer. Lifestyle changes
— such as losing weight, exercising, reducing salt (sodium) in your diet and

180
managing stress — can improve your quality of life. However, heart failure can be
life threatening. People with heart failure may have severe symptoms, and some may
need a heart transplant or a ventricular assist device (VAD). One way to prevent
heart failure is to prevent and control conditions that can cause it, such as coronary
artery disease, high blood pressure, diabetes and obesity.
The prediction of heart disease risk factors using clinical and statistical methods
has been receiving much attention in the recent decade. But few published studies
employed NLP techniques to investigate this research issue on the basis of textual
medical records.
To overcome the issues in conventional invasive-based methods for the
identification of heart disease, researchers attempted to develop different non-
invasive smart healthcare systems based on predictive machine learning techniques
namely: Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Naïve Bayes
(NB), and Decision Tree (DT), etc. As a result, the death ratio of heart disease
patients has decreased.
This paper is an illustration, which details the efforts to the risk factor challenge
task responsible for Cardiovascular Failure. Different models were developed, which
integrated a variety of methodological approaches, such as dictionary-based
keyword spotting, rules and supervised learning, for the detection of a variety of
heart disease risk factors. The developed system has been trained and tested on the 2
heart disease datasets which are available online on the Kaggle. All the processing
and computations were performed using Google Collab, MS Excel, R. Python has
been used as a tool for implementing all the classifiers. The main packages and
libraries used include pandas, NumPy and matplotlib. The main contribution of our
proposed work is given as:
1. To analyse all the physical and cardiovascular factors responsible for
cardiovascular failure.
2. To find correlation between different factors and design desired models
for them.
3. To conduct hypothesis testing to analyse different health factors
variations.
4. To apply machine learning algorithms and survival models to get the
maximum accuracy. The rest of the paper is organized as follows: Section
2 details the methods that were employed to deal with the research
including the datasets used, processing and cleaning of data followed with
EDAs and Hypothesis Testing used. In Section 3 the results are discussed
along with the discussion on the research including Data Figures, Tables
and Figures, survival models, regression models etc. Section 4 reflects the
conclusions. References are given in Section 5. Section 6 contains
Recommendations.
181
2. Material and Methodology
The main purpose of the research is to predict the occurrence of heart disease for
early detection of the disease using the risk factors in a short time. Different data
mining techniques and machine learning algorithms, k Nearest Neighbor (KNN),
Decision tree, Artificial Neural Network (ANN), Random Forest to predict the heart
disease based on some health parameters.
Data visualizations, pre-processing of data and plot graphs were also performed.
Dataset
Dataset is a collection of various types of data stored in a digital format. Data is
the key component of any Learning project. The quality of data is as important as the
quantity even if you have implemented great algorithms for machine learning
models. According to The State of Data Science 2020 report, data preparation and
understanding is one of the most important and time-consuming tasks of the
Machine Learning project lifecycle. Survey shows that most Data Scientists and AI
developers spend nearly 70% of their time analysing datasets. The remaining time is
spent on other processes such as model selection, training, testing, and deployment.
Looking at the significance of the dataset, two datasets i.e., Physical factors
disease dataset S1 and cardiovascular factors heart disease dataset (S2) are used,
which are available online at the Kaggle repository for conducting their research
studies. The S1 consists of 299 instances, where each instance has distinct 13
attributes (one independent and the rest 12 dependents). The dataset is composed of
two classes, on the basis of severity of heart disease. The S2 is composed of 299
instances in which 96 instances belong to the severe ones while the rest of 203 slight
ones. The description of attributes of both the datasets is different, S1 contains the
data corresponding to the physical health factors and S2 contains the data with
respect to cardiovascular health factors.
Exploratory Data Analysis
EDA focuses more narrowly on checking assumptions required for model fitting
and hypothesis testing. It also checks while handling missing values and making
transformations of variables as needed. EDA builds a robust understanding of the
data, and issues associated with either the info or process. It‘s a scientific approach
to getting the story of the data.
Pre-Processing of Data
Data pre-processing is the process of transforming raw data into meaningful
patterns. It is very crucial for a good representation of data, makes visualization
easier and increases the accuracy and speed of the machine learning algorithms that
train on the data. There are four stages of data processing: cleaning, integration,
reduction, and transformation. Various pre-processing approaches such as missing

182
values removal, standard scalar, and Min–Max scalar are used on the dataset in order
to make it more effective for classification.

Data Analysis and Data Visualizations


Data Visualization is the graphical representation of information and data. By
using visual elements like charts, graphs, and maps, data visualization tools provide
an accessible way to see and understand trends, outliers, and patterns in data. For the
comparison purpose, we have simple and multivariate line, bar and pie charts, box
plots. along with correlation plots. There are different types of graphs which can be
made depending upon the requirement.
BAR PLOT: A bar plot is a plot that presents categorical data with rectangular bars
with lengths proportional to the values that they represent. A bar plot shows
comparisons among discrete categories.
PIE CHART: A pie chart is a circular graphic used in statistical analysis., it is
divided into slices to illustrate numerical proportions.
BOX PLOTS: A box plot is a simple way of representing statistical data on a plot,
drawn to represent the second and third quartiles, usually with a vertical line inside
to indicate the median value. The lower and upper quartiles are shown as horizontal
lines either side of the rectangle.
AREA CHART: An area chart or area graph displays graphically quantitative data.
The area between axis and line are commonly emphasized with colours, textures and
hatchings. Commonly one compares two or more quantities with an area chart.
LINE CHART: A line chart or line graph, also known as curve chart, is a type of
chart which displays information as a series of data points called 'markers' connected
by straight line segments.
Statistical Tools Used
• Statistical hypothesis test (One sample & two sample)
• Linear & Multiple Regression Modelling
• Survival Analysis Models.
• Normality Assumption: This assumption states that if we collect many
independent random samples from a population and calculate some value of interest
(like the sample mean) and then create a histogram to visualize the distribution of
sample means, we should observe a perfect bell curve.

3. Results and Discussions


The following section includes a detailed analysis of visualizations, hypothesis
testing, models and algorithms.

183
i. Comparative Study of Physical and Cardiovascular Health Factors:

Medium/ Physical Health Cardiovascular Health


Factor of
Study
Factors [Link]: the age of the individual under
1. Age: the age of the individual
considera [Link]: the gender of the individual 2. Sex: gender of the individual
tion [Link]: frequency of Chest Pain (type) 3. Anaemia: Decrease of red blood
[Link] : resting blood pressure (in cells or haemoglobin (Boolean)
mmHg) 4. Creatinine Phos: Level of the
[Link] : cholesterol in mg/dl fetched CPK enzyme in the blood (mcg/L)
via BMI sensor 5. Ejection fraction: Percentage of
[Link] : (fasting blood sugar > 120 blood leaving the heart at each
mg/dl) (1 = true; 0 = false contraction (percentage)
7. thalachh: maximum heart rate 6. Platelets: Platelets in the blood
achieved. (kilo platelets/mL)
8. exng : exercise induced angina 1. Serum_creatinine: Level of
(1 = yes; 0 = no) serum creatinine in the blood
(mg/dL)
2. Serum_sodium : Level of serum
sodium in the blood (mEq/L)

Correlati
on Heat
Map

Correlation Heat Map between


different physical health factors with Correlation between different
the output (severity and causality of cardiovascular health factors with
CVFs) the output (severity and causality of
CVFs)

184
Age
Factor

Age v/s Severity/Causality of Case Age v/s Severity/Causality of Case:


Percentage of Super Senior people Percentage of Super Senior people
suffered from cardiovascular failure suffered from cardiovascular failure
are: 83.33 are: 59.62
Percentage of Senior people Percentage of Senior people
suffered from cardiovascular failure suffered from cardiovascular failure
are: 47.24 are:26.59
Percentage of Adults people Percentage of Adults people
suffered from cardiovascular failure suffered from cardiovascular failure
are: 69.90 are: 25.68

Multiple
Health
Factors
Analysis

Comparative study between Variation of serum creatinine v/s


cholesterol, resting blood pressure ejection fraction
and maximum heart rate achieved.

185
ii. Visualizations & Insights

[Link]. Factor Type of Graphical Variation Inferences &


Factor Comments

1. Chest Pain Physical Trend of severity of


Health chest pain, starting
Factor from minimum
to maximum severity.
The severity level of
chest pain is a crucial
indicator of
cardiovascular disease.
Mild discomfort may be
due to less critical
Chest Pain Severity Variation
issues, while severe
pain may signify a
higher risk of a
heartattac k or other
serious conditions.
Prompt medical
attention is vital for
accurate diagnosis and
timely intervention to
mitigate cardiovascular
risks.
2 Resting Physical First Quartile =
Blood Health 120mm/Hg
Pressure Factor Second Quartile
=130mm/Hg
Third Quartile =
140mm/Hg
IQR=140 mm/Hg -
120mm/Hg =20 mm/Hg
LOWER
BOUNDARY = 90
Resting Blood pressure recorded in mm/Hg
mm/Hg UPPER
BOUNDARY = 170
mm/Hg
Indicating the median
value of BP as
130mm/Hg with
minimum and
maximum level of 90
and 170 mm/Hg.

186
3 Creatinine CV Variations in creatinine
Phospho- Health phosphokinase (CPK)
kinase Factor levels can indicate
muscle damage,
including the heart.
Elevated CPK may
suggest cardiovascular
issues, such as heart
attack. Monitoring CPK
levels helps assess
Variation of Creatinine cardiac health, aiding in
Phosphokinase early detection and
with population count management of
potential cardiovascular
problems. The
minimum value
reported is
29 and the maximum
value is 7861. The
normal range of CPK,
however is
10 to 120 micrograms
per litre.
4 Presence CV Anemia,
of Health characterized by low
Anaemia Factor red blood cell count,
can impact
cardiovascular health.
Insufficient oxygen
delivery due to
anaemia may strain
the heart, potentially
leading to cardiac
Division of data on basis of complications.
presence of Anaemia Monitoring and
addressing anemia is
crucial for
maintaining overall
cardiovascular well-
being and preventing
associated health
risks.

187
5 Cholestrol Physical Elevated cholesterol
Levels Health levels, increase the
Factor risk of CV diseases.
Total cholesterol level
is less than 200 milli
grams per decilitre
(mg/dL) are
considered desirable
Cholesterol Level Variation in for adults. A reading
population between 200 and 239
mg/dL is considered.
Border line high and
a reading of 240 mg/
dL and above is
considered high.
Healthy lifestyle and
medication promotes
better heart health.
6 Blood Physical High sugar level in
Sugar Health the blood, often
levels Factor associated with
excessive sugar
consumption, can
contribute to
cardiovascular health
a. Blood Sugar Analysis Pie issues.
Chart Elevated sugar levels
may lead to
inflammation, insulin
resistance and
increased risk of heart
disease.
Out of 165 patients of
heart failure, only 23
b. Distribution of heart failure were diabetic and rest
patients our nondiabetic
patients.
Maintaining a
balanced diet &
managing sugar
intake can reduce the
risk of complications.

188
7 Platelets CV The mean value of
level Health dataset for platelets is
Factor
263358.03 due to
presence of extreme
outliers with highly
extreme values of
platelets which is not a
good sign. Platelets
play a crucial role in
a. Division of Platelets Recorded cardiovascular health
by forming blood clots
to prevent excessive
bleeding. However,
abnormal clotting can
lead to arterial
blockages, causing
heart attacks and
strokes.

b. Variation of platelets in
recorded dataset

8 Serum CV EXANG
Sodium Health enhances blood flow.
Levels Factor The plot is also pointing
towards following
negatively skewed
normal distribution with
mean as 149.77 and
standard deviation of
22.90.
The minimum value of
maximum heart rate
achieved is 71 and
maximum is 202.
Serum Sodium Variation Chart

189
9 Exercised Physical EXANG
included Health enhances blood flow. The
Angina Factor plot is also pointing
towards following
negatively skewed normal
distribution with mean as
149.77 and standard
deviation of 22.9.
The minimum value of
maximum heart rate
achieved is 71 and
Trends of Exercised Induced Angina maximum is 202.

10 Serum Physical According to graph, serum


Creatinine Health creatinine level is
Factor following a fairly normal
distribution with mean
1.39 and
standard
deviation of
1.03 The minimum value
of serum
creatinine observed is 0.5
and maximum value
observed is 9.4.
Frequency plot for serum Monitoring
creatinine creatinine helps assess
kidney function, crucial
for maintaining
cardiovascular
health. Early detection
allows timely intervention
and prevention of
major heart problem
at earliest stage.

11 Maximum Physical It is crucial as it reflects


Heart Health upper limit of one‘s
Rate Factor heart‘s capacity during
exercise. Monitoring
helps reduce risk of
heart related issues.
The first quartile is
obtained at 135.
Variation of Maximum Heart Rate The second quartile is
obtained at 153. The
third quartile is
obtained at 166.
190
12 Ejection CV Ejection fraction is a
Fraction Health cardiac measure
Factor indicating % of blood
pumped out of heart
during contraction. A
normal ejection fraction
range is between 52 and
72 per cent for men and
between 54 and 74 per
cent for women. An
ejection fraction that is
higher or lower may be
Frequency plot for ejection fraction a sign of heart failure or
recorded an under lying heart
condition.

(iii) Hypothesis Testing

191
Since t-value lies in rejection region, also p- value is less than 0.05; hence we
have sufficient evidence to reject H0 at 5% level of significance with 540 as degree
of freedoms for two tailed test. The mean age group of individulas in both the cases
is different.
192
4. Variance ratio amnalysis for age group of physical and cardovascular factors.
Hypothesis:
1
=1
H0: ratio of variances of physical and cardiovascular factor for age is 1 i.e.
2
1
1
H1: ratio of variances of physical and cardiovascular factor for age is 1 i.e.
2
F=0.58399, num df=298, denom df=298, p value=4.003 e-06 alternative
hypothesis: true ratio of variances is not equal to 1; 95% CI: 0.4651526 0.7331762;
sample setimates: ratio of varainces =0.5839853 p-value is calculated is 4.8812; we
have sufficient evidence to reject H0. The ratio of varainces of physical and
cardiovascular factor for age is not 1 at 5% level of significance with 298 as df for
two tailed test.
iv. Regression Models
Under Physical Health
Model 1: model=lm(Output~resting blood pressure)

Removal of cholesterol from the model has increased the adjusted R square
value and now both the parameters are significant to the model and removal of either
will decrease the R- square value With an R square value of 0.1898, the linear model
193
of output v/s resting blood pressure illustrates the best model to go for under
physical health risk factors.
To analyse it further for a verified conclusion, goodness of fit test is the answer
to the question. The calculated p-value is 11.53 which is greater than the table value
of 0.05. Hence, the model (output~ resting blood pressure is a better fit than any
other model.
Model 2: y=-0.004225*x+1.108124

Adjusted R square came out to be maximum for this model and now the 3
parameters are significant to the model and removal of any will decrease the R-
square level.
Model 3: Most Fitted Multiple Regression Model Under Quantitative
Cardiovascular Factors Responsible For CVDs

194
V. Data Prediction Algorithms & Accuracy
Machine-Learning Algorithm Accuracy
Under Physical Accuracy Under Cardiovascular Health

The Surv() function takes two times and output as input and creates an object
which serves as the input of survfir() function. We pass-1 in survfit() function to
ensure that we are telling the function to fit the model on basis of the survival object
and have an interrupt.

195
Here, we are interested in ―time‖ and ―output‖ in the analysis. Time represents
the survival time of patients. Since patients survive, we will consider their status as
dead or non-dead (censored) and some have left the survey due to any reason.
The surv() function takes two times and output as input and creates an object
which serves as the input of survfir() function. We pass ** sex in survfit() function
to ensure that we are telling the function to fit the model on basis of categorical
function sex of the survival object and have an interrupt.

Conclusion:
Our study employed advanced statistical techniques to explore critical
associations and correlations, illuminating complex relationships between biological
factors and cardiovascular failure risk. By analyzing essential physiological
markers—such as blood pressure, cholesterol, creatinine phosphokinase, platelet
count, and blood glucose levels—we identified significant statistical correlations
that offer valuable insights into cardiovascular health risk factors.
Key findings reaffirm the predictive power of traditional risk factors, including
elevated blood pressure, cholesterol, and blood sugar levels, while also underscoring
their dynamic interactions in influencing cardiovascular health. Notably, the study
highlights lesserrecognized markers, such as creatinine phosphokinase and platelet
count, as potential biomarkers of cardiovascular risk, suggesting their importance in
risk assessment models and early detection strategies.
This research contributes to the interdisciplinary field of cardiovascular studies
by integrating statistical modeling and biological analysis, providing a holistic
perspective that deepens our understanding of the biological pathways influencing
cardiovascular health. Such an approach bridges gaps between statistics and biology,
offering a multi-dimensional framework that surpasses the insights achievable
through singular disciplinary focus.
These results underscore the importance of interdisciplinary research in
addressing complex health conditions and highlight the potential of personalized
healthcare strategies. Our findings may guide the development of targeted public
health interventions, ultimately helping to mitigate the growing burden of
cardiovascular disease and improve outcomes for at-risk populations at both
individual and community levels.

References
1. Akhtar, N., & Taqdees, S. (2021). Heart Disease Prediction. ResearchGate,
11.
2. Ali, L., Niamat, A., Khan, J. A., Golilarz, N. A., Xingzhong, X., Noor, A., &
Bukhari, S. A. C. (2019). An optimized stacked support vector machines
based expert system for the effective prediction of heart failure. IEEE
access, 7, 54007-54014.
3. Allen, L. A., Stevenson, L. W., Grady, K. L., Goldstein, N. E., Matlock, D.
D., Arnold, R. M., & Spertus, J. A. (2012). Decision making in advanced
196
heart failure: a scientific statement from the American Heart Association.
Circulation, 125(15), 1928-1952.
4. Bui, A. L., Horwich, T. B., & Fonarow, G. C. (2011). Epidemiology and risk
profile of heart failure. Nature Reviews Cardiology, 8(1), 30-41.
5. Kahramanli, H., & Allahverdi, N. (2008). Design of a hybrid system for the
diabetes and heart diseases. Expert systems with applications, 35(1-2), 82-
89.
6. Levy, D., Kenchaiah, S., Larson, M. G., Benjamin, E. J., Kupka, M. J., Ho,
K. K., & Vasan, R. S. (2002). Long-term trends in the incidence of and
survival with heart failure. New England Journal of Medicine, 347(18),
1397-1402.
7. Meshref, H. (2019). Cardiovascular disease diagnosis: A machine learning
interpretation approach. International Journal of Advanced Computer
Science and Applications, 10(12).
8. Muhammad, Y., Tahir, M., Hayat, M., & Chong, K. T. (2020). Early and
accurate detection and diagnosis of heart disease using intelligent
computational model. Scientific reports, 10(1), 19747.
9. Redfield, M. M. (2000). Epidemiology and pathophysiology of heart failure.
Current cardiology reports, 2(3), 179-180.
10. Yang, H., & Garibaldi, J. M. (2015). A hybrid model for automatic
identification of risk factors for heart disease. Journal of biomedical
informatics, 58, S171-S182.

197
Archimedean Copula Functions: A Powerful Tool for
Dependency Modeling
Ruhiteswar Choudhury1, and Tanusree Deb Roy1
1
Department of Statistics, Assam University, Silchar - 788011, India.

Abstract
This paper reviews Archimedean copulas: their definition, construction via
generator functions, widely used Archimedean families (Clayton, Frank, Gumbel,
Joe, Ali-Mikhail- Haq and related families), dependence/tail properties, inference
and goodness-of-fit (GOF) methods, extensions (nested / hierarchical Archimedean
copulas), and common applica- tions. We provide key formulas, tail-dependence
expressions (where available), practi- cal estimation procedures (Kendall‘s-tau
inversion, maximum likelihood, IFM, rank-based semiparametric methods), and
summarize the state of the literature with core citations and recommended further
reading.

1 Introduction
Archimedean copulas represent a fundamental class of copula functions widely
utilized in statistical modeling to capture dependence structures among multivariate
random variables. Copulas, introduced by Sklar (1959), provide a flexible
framework for modeling joint distributions by separating marginal distributions from
their dependence structure. Archimedean copulas, in particular, are distinguished by
their construction via a single generator function, which simplifies their
mathematical representation and facilitates their application in high-dimensional
settings Nelsen (2006). These copulas are especially val- ued for their ability to
model a wide range of dependence patterns, including positive and negative tail
dependencies, making them indispensable in fields such as finance, hydrology, and
risk management Genest and Favre (2007).
The defining feature of Archimedean copulas is their generator function, a
continuous, decreasing, and convex function that maps the unit interval to the non-
negative real num- bers. This generator enables the construction of copulas with
desirable properties, such as symmetry and associativity, which are particularly
useful in applications requiring tractable multivariate models McNeil and
Nesˇ lehova´ (2009). Prominent examples of Archimedean copulas include the
Clayton, Gumbel, and Frank copulas, each characterized by distinct dependence
behaviors. For instance, the Clayton copula is adept at capturing lower tail
dependence, crucial for modeling financial risks during market downturns Patton


Corresponding author

198
(2006), while the Gumbel copula excels in upper tail dependence, relevant for
extreme event anal- ysis in hydrology Salvadori et al. (2007).
The theoretical underpinnings of Archimedean copulas have been extensively
explored, with significant contributions to their characterization and estimation.
Genest and MacKay (1986) provided key insights into the properties of their
generators, while Joe (1997) ad- vanced methods for parameter estimation and
model selection. Recent developments have extended their applicability to vine
copula constructions, enabling more flexible modeling of complex dependencies
Bedford and Cooke (2002). Moreover, Archimedean copulas have been adapted to
handle time-varying dependencies, enhancing their utility in dynamic financial
models Patton (2006).
This review aims to synthesize the theoretical foundations, estimation
techniques, and practical applications of Archimedean copulas, highlighting their
versatility and limita- tions. By examining their role in diverse disciplines, we
underscore their significance in advancing multivariate statistical modeling.

2 Literature Review
Archimedean copulas, defined by their single generator function, have become
integral to multivariate dependence modeling since Sklar (1959) established the
theoretical founda- tion of copulas, enabling the separation of marginal distributions
from joint dependence structures. Early developments shaped specific Archimedean
families: Gumbel (1960) linked bivariate exponential distributions to extreme-value
theory, laying the groundwork for the Gumbel copula, which excels in modeling
upper-tail dependence, critical for ex- treme event analysis in hydrology Salvadori et
al. (2007). Clayton (1978) introduced a model for association in bivariate life tables,
now recognized as the Clayton copula, which captures lower-tail dependence and
found early applications in epidemiology and survival analysis. Frank (1979)
proposed the Frank copula, a symmetric family without strong tail dependence, ideal
for moderate, symmetric associations.
Constructive representations advanced the field significantly. Marshall and
Olkin (1988) developed mixture-based constructions using Laplace transforms,
facilitating sampling al- gorithms and generator characterizations for Archimedean
copulas. Oakes (1989) con- nected frailty models in survival analysis to these
copulas, particularly Clayton variants, through shared random effects, enhancing
their utility in reliability and life-data appli- cations Balakrishnan and Lai (2009).
The algebraic simplicity of Archimedean copulas- owing to their symmetric and
associative properties-has made them suitable for multivari- ate extensions McNeil
and Nešlehová (2009).
Inference techniques for Archimedean copulas saw major progress in the 1990s.
Genest and Rivest (1993) pioneered semiparametric and rank-based estimation
methods, which Genest et al. (1995) formalized into the widely used Genest–
Ghoudi–Rivest estimator for dependence parameters, offering a robust alternative to
maximum likelihood estimation. Joe (1997) provided a comprehensive treatment of
199
multivariate dependence, establishing a standard reference for theoretical and
applied researchers. Nelsen (2006), in its canonical textbook, detailed construction
methods, properties, and practical examples, making Archimedean copulas
accessible to practitioners.
Recent advancements have focused on high-dimensional applications and
diagnostic tools. McNeil and Nešlehová (2009) clarified validity conditions for
multivariate Archimedean copulas through d-monotonicity and ℓ 1-norm symmetric
distributions, while McNeil and Nešlehová (2012) addressed nested structures for
hierarchical dependence modeling. Hofert (2008, 2012) developed efficient
sampling and estimation algorithms, including simulated- likelihood and diagonal-
based estimators, suitable for high dimensions, with implementations in R packages
like copula and nacopula Hofert et al. (2012). Genest et al. (2009) and Kojadinovic
et al. (2011) introduced goodness-of-fit tests, leveraging empirical copula processes
and bootstrap methods for robust model validation.
Applications have expanded across domains. Patton (2012) highlighted dynamic
Archimedean copulas for financial time-series modeling, while Genest and Favre
(2007) and Aas et al. (2009) emphasized their role in hydrology and insurance,
comparing them to vine copulas for flexibility. Charpentier and Segers (2014) and
McNeil (2010) advanced tail-dependence estimation, crucial for finance and risk
management. Recent works (2015–2024) have fo- cused on scalable algorithms,
dynamic models, and domain-specific applications in relia- bility, environmental
extremes, and survival analysis Balakrishnan and Lai (2009); Segers et al. (2014),
reinforcing Archimedean copulas‘ versatility despite limitations in capturing
asymmetric dependencies compared to vine structures Aas et al. (2009).
This review underscores the evolution of Archimedean copulas from theoretical
con- structs to practical tools, driven by advances in estimation, computation, and
application, while noting their constraints in complex dependence modeling.

3 Tail Dependency
To understand the data more precisely to model the appropriate copula, the tail
dependence needs to be checked (i.e., whether the data is lower tail or upper tail
dependence or no tail dependence). The lower tail dependency indicates the chance
of joint failure of both the variables are simultaneous and viceversa.
The general expression for lower tail dependency is –
𝜆𝐿 = lim+ 𝑃 .𝑌 ≤ 𝐹𝑌−1 (𝑞)|𝑋 ≤ 𝐹𝑋−1 (𝑞)/
𝑞→0
𝑃 .𝑌 ≤ 𝐹𝑌−1 (𝑞), 𝑋 ≤ 𝐹𝑋−1 (𝑞)/
= lim+
𝑞→0 𝑃(𝑋 ≤ 𝐹𝑋−1 (𝑞))
Similarly, the expression for upper tail dependency is –
𝜆𝑈 = lim+ 𝑃 .𝑌 ≥ 𝐹𝑌−1 (𝑞)|𝑋 ≥ 𝐹𝑋−1 (𝑞)/
𝑞→0

200
𝑃 .𝑌 ≥ 𝐹𝑌−1 (𝑞), 𝑋 ≥ 𝐹𝑋−1 (𝑞)/
= lim+
𝑞→0 𝑃(𝑋 ≥ 𝐹𝑋−1 (𝑞))
Here, since 𝑋 is a random variable, ⇒ 𝐹𝑋 (𝑥) ∼ 𝑈(0,1), it can be defined as
$𝑈 = 𝐹𝑋 (𝑥) & 𝑉 = 𝐹𝑌 (𝑦) where, 𝑈, 𝑉 ∼ 𝑈,0,1-. Hence –
1
𝑃 .𝑋 > 𝐹𝑋−1 (𝑞)/ = ∫ 𝑓(𝑢)𝑑𝑢 = 1 − 𝑞
𝑞

4 Archimedean Copula Families


Archimedean copulas are characterized by their generator function ϕ : [0, 1] →
[0, ∞], enabling the construction of diverse dependence structures. Below, we detail
the Clayton, Gumbel, Frank, and Joe copulas, including their equations, tail
dependence, and usability, with extensions like BB1 and BB7 for added flexibility
Joe (1997).
 Clayton Copula: Clayton (1978) The two-dimensional Clayton Copula function
is –

 Copula Generator function:


1
𝐶(𝑢, 𝑣) = (𝑢−𝜃 + 𝑣 −𝜃 − 1)𝜃 ; 𝜃 ≥ 0
𝑙
 Tail Dependence: Lower tail dependence, 𝜆𝐿 = 2−𝜃 , no upper tail
dependence.
 Usability: Ideal for modeling lower-tail dependence, such as financial asset
re- turns during market downturns or survival times in epidemiology, where
joint extreme low values are of interest Patton (2006); Balakrishnan and Lai
(2009).

Figure 1: Scatter Plot of the Clayton copula at different dependence parameter


values.

201
 Gumbel Copula: The two-dimensional Gumbel Copula function Gumbel (1960) is:
 Gumbel copula:
𝑙
 𝐶(𝑢, 𝑣, 𝜃) = exp [−[(− ln(𝑢))−𝜃 + (− ln(𝑣))−𝜃 ]𝜃 ] , 𝜃 ≥ 1
𝑙
 Tail Dependence: Upper tail dependence, 𝜆𝑈 = 2 − 2𝜃 , no lower tail depen-
dence.
 Usability: Suited for applications involving joint extreme high values, such
as extreme rainfall or flood events in hydrology Salvadori et al. (2007).
 Frank Copula: The two-dimensional Frank Copula function is Frank (1979)
 Frank Copula function: for positive dependence, negative for negative).
 Tail Dependence: No tail dependence, symmetric dependence structure.
 Usability: Effective for modeling moderate, symmetric dependencies, such
as in insurance or reliability studies with balanced associations Genest and
Favre (2007).
 Joe Copula Joe (1997):
 Generator: 𝜙(𝑡) = − log(1 − (1 − 𝑡)𝜃 ), 𝜃 ≥ 1
𝑙
 Copula: 𝐶(𝑢1 , 𝑢2 , … , 𝑢𝑑 ) = 1 − (1 − exp(− ∑𝑑𝑖=1(− log(1 − (1 − 𝑢𝑖 )𝜃 )) ))𝜃
𝑙
 Tail Dependence: Upper tail dependence, 𝜆𝑈 = 2 − 2𝜃 , no lower tail
dependence.
 Usability: Similar to Gumbel, useful for upper-tail dependence in risk
management and extreme value analysis, with slightly different dependence
patterns Joe (1997).

Figure 2: Scatter Plot of the Gumbel copula at different dependence parameter


values.

202
Extended families like BB1 and BB7, introduced by Joe and Hu (2015),
combine features (e.g., both upper and lower tail dependence) for greater flexibility
in modeling asymmetric dependencies, particularly in finance Patton (2012). These
copulas are constructed via transformations of existing generators, requiring
multiple parameters for enhanced versatility Durante and Sempi (2015).

5 Choosing an Appropriate Copula


The selection of an appropriate copula is a pivotal step in multivariate statistical
modeling, as copula functions are intricately tied to the tail dependence structure of
the underlying data. The upper tail dependence coefficient (λ U ) and lower tail
dependence coefficient (λ L) serve as critical indicators of the propensity for extreme
events to co-occur across variables. These coefficients guide the selection of a
copula that accurately captures the dependence structure, ensuring robust modeling
outcomes in applications such as finance, hydrology, insurance, and reliability
analysis.
For data exhibiting moderate lower tail dependence, where λ L > 0, the Clayton
cop- ula is often an effective choice. This copula excels at modeling asymmetric
dependence concentrated in the lower tail, making it particularly suitable for
scenarios where extreme negative events—such as simultaneous market downturns
or low hydrological flows—are likely to cluster. Conversely, when moderate upper
tail dependence is present, i.e., λ U > 0, the Joe copula is well-suited. Its strength lies
in capturing dependence in the upper tail,

Figure 3: Scatter Plot of the Frank copula at different dependence parameter values.

203
Which is valuable for modeling phenomena where extreme positive events, such
as syn- chronized asset price surges or high flood levels, tend to occur together.
In cases where the data exhibit strong lower tail dependence but weak upper tail
depen- dence, the Gumbel copula may be a more appropriate choice. The Gumbel
copula is adept at modeling asymmetric dependence with a focus on the upper tail,
but its structure can also accommodate scenarios where lower tail dependence is
more pronounced, making it versatile for applications like extreme event analysis in
hydrology or finance. On the other hand, when neither upper nor lower tail
dependence is evident, i.e., λ U = 0 and λ L = 0, the Frank copula emerges as an
optimal choice. The Frank copula is particularly effective for modeling symmetric
dependence structures without pronounced tail behavior, making it suitable for
datasets where extreme co-movements are minimal, such as in certain insur- ance or
reliability contexts.
The process of selecting an appropriate copula extends beyond evaluating tail
dependence coefficients. Empirical tools, such as rank-transformed scatterplots, tail
dependence diagnostics, or goodness-of-fit tests (e.g., those available in R‘s copula
package), are essential for assessing the suitability of a chosen copula. Parameter
estimation methods, including maximum likelihood estimation or inference
functions for margins, play a crucial role in calibrating the copula to the data.
Moreover, the choice of copula should be informed by the specific application
context, as different domains may prioritize distinct aspects of dependence. For
instance, financial risk models may emphasize lower tail dependence, while
hydrological models may focus on upper tail behavior.

Figure 4: Scatter Plot of the Joe copula at different dependence parameter values.

204
6 Theoretical Foundations
Archimedean copulas are defined by a generator function ϕ : [0, 1] → [0, ∞],
which is continuous, strictly decreasing, convex, and satisfies ϕ (1) = 0. For a d-
dimensional Archimedean copula, the joint cumulative distribution function is given
by:
C(u1, . . . , ud) = ϕ − 1 (ϕ (u1) + · · · + ϕ (ud)) , (6)

where ϕ − 1 is the inverse of the generator, and the copula is valid if ϕ is d-


monotone, i.e., it has non-negative derivatives up to order d − 2 and (− 1)kϕ (k)(t) ≥
0 for k = 0, . . . , d − 2 McNeil and ešlehová (2009). The generator‘s properties
ensure symmetry and associativity, making Archimedean copulas computationally
tractable but restrictive for asymmetric dependencies Nelsen (2006). Marshall and
Olkin (1988) provided a mixture representation, where the copula arises from a
Laplace transform of a positive random variable, facilitating sampling via:

Ui = ϕ − 1(Vi/W ), i = 1, . . . , d, (7)

where Vi ∼ Exp(1) and W follows the distribution associated with ϕ ‘s Laplace


transform Hofert (2008). McNeil and Nesˇ lehova´ (2012) extended this to nested
Archimedean cop- ulas, allowing hierarchical dependence structures, though
identifiability constraints limit their flexibility.

7 Parameter Estimation Techniques


Parameter estimation for Archimedean copulas typically involves the
dependence parame- ter θ . Key methods include:
Maximum Likelihood Estimation (MLE)
The log-likelihood for a sample {(ui1, . . . , uid)}n is maximized:
𝑛

𝑙(𝜃) = ∑ log 𝑐(𝑢𝑖1 , 𝑢𝑖2 , … , 𝑢𝑖𝑑 ; 𝜃)


𝑖=1
where c is the copula density, derived as c(u1, . . . , ud) = ϕ − 1(d)(Σ ϕ (ui)) Qd
ϕ ′ (ui)
Joe (1997). MLE is computationally intensive in high dimensions but precise for
known marginals Hofert (2012).
Semiparametric Estimation
Genest et al. (1995) proposed a rank-based estimator using Kendall‘s tau, τ ,
related to the generator via:
1
𝜙(𝑡)
𝜏 = 1 + 4 ∫ ′ 𝑑𝑡
0 𝜙 (𝑡)
This Genest–Ghoudi–Rivest estimator is robust to marginal misspecification and
widely used for its simplicity Genest and Rivest (1993).

205
Simulated Likelihood and Diagonal-Based Estimators
For high dimensions, Hofert (2012) introduced simulated-likelihood methods
and diagonal- based estimators, leveraging the copula‘s diagonal section
𝛿(𝑢) = 𝐶(𝑢, … , 𝑢)
to estimate θ , reducing computational complexity.

Inference function for margins


The IFM method is based on two-stage algorithms for estimating the best fitted
parameter. At first, estimating the parameters of the distribution based on Marginal
distribution. Fur- ther, putting the estimated parameters of the marginal distributions
in the copula density, to find the dependence parameter of the model. For more
explanation refer to Joe and Xu(1996) about the IFM method.
These methods balance accuracy and computational feasibility, with
semiparametric approaches favored for robustness and MLE for precision when
marginals are known Gen- est and Favre (2007) and for many financial data IFM
method should be implemented.

8 Goodness-of-Fit Tests
Goodness-of-fit tests assess whether an Archimedean copula fits the data.
Genest et al. (2009) reviewed key methods, including:

Empirical Copula Process


1
Compares the empirical copula 𝐶𝑛 (𝑢1 , … , 𝑢𝑑 ) = ∑𝑛𝑖=1 ∏𝑑𝑗=1{𝑈𝑖𝑗 ≤ 𝑢𝑖𝑗 }to the
𝑛
fitted copula Cθ . The test statistic, e.g., Cramé r-von Mises, is:

𝑆𝑛 = ∫ 𝑛*𝐶𝑛 (𝑢) − 𝐶𝜃 (𝑢)+2 𝑑𝐶𝑛 (𝑢)


,0,1-𝑑

P-values are computed via bootstrap due to the statistic‘s non-standard


distribution Genest et al. (2009).
Kendall Transform Tests
Transform data using the Kendall function K(t) = P (C(U1, . . . , Ud) ≤ t), testing
deviations from the fitted 𝜈 Genest et al. (2009).
Fast Large-Sample Tests
Kojadinovic et al. (2011) developed computationally efficient tests for large
samples, using rank-based statistics and multiplier bootstrap methods for improved
accuracy.

206
Bootstrap refinements by Segers et al. (2014) enhance small-sample
performance, mak- ing these tests practical for high-dimensional data Genest and
Favre (2007).
9 Applications and Limitations
Archimedean copulas have become a cornerstone of multivariate statistical
modeling, with significant applications in fields such as finance, hydrology,
reliability engineering, and survival analysis. Their capacity to model complex
dependence structures, particularly in the tails of distributions, makes them
exceptionally suited for addressing intricate problems where traditional correlation-
based approaches are inadequate.
In finance, Archimedean copulas have proven highly effective for modeling
dependencies among asset returns, especially during periods of market turbulence.
For example, Patton (2012) utilized dynamic Clayton copulas to capture time-
varying dependencies in asset returns, adeptly modeling lower-tail risks during
financial crises. This capability enhances risk management and portfolio
optimization by providing a more nuanced understanding of systemic risks. In
hydrology, Archimedean copulas have been pivotal for analyzing extreme
environmental phenomena. Salvadori et al. (2007) applied Gumbel copulas to model
the joint distribution of extreme rainfall events, capitalizing on their ability to
capture upper-tail dependence to improve flood risk assessments and inform
infrastructure design. Similarly, in reliability engineering and survival analysis,
Clayton and Frank copulas have been widely employed to model dependent failure
times. Balakrishnan and Lai (2009) demonstrated their effectiveness in quantifying
joint failure risks in complex systems, offering valuable insights for engineering
reliability and actuarial applications.
The practical adoption of Archimedean copulas has been greatly supported by
advanced software tools. The R package copula, developed by Hofert et al. (2012),
pro- vides a robust framework for parameter estimation, simulation, and goodness-
of-fit testing, enabling seamless integration into research and applied workflows.
Complementary tools in Python, such as the copulae library, and MATLAB-based
packages further enhance their accessibility, facilitating their use across academic
and industrial settings. Nevertheless, Archimedean copulas face certain limitations.
Their dependence on a single generator function constrains their ability to model
high-dimensional datasets with diverse dependence structures. While nested or
hierarchical Archimedean copulas offer a partial solution, they introduce significant
computational complexity and require meticulous parameter calibration. Estimating
parameters in dynamic or non-stationary environments re- mains challenging,
particularly with sparse or noisy data. Additionally, the exchangeability assumption
inherent in many Archimedean copulas may not hold in cases of asymmetric
dependencies. Computationally, fitting and simulating high-dimensional copula
models can be resource-intensive, posing challenges for real-time applications.
In conclusion, Archimedean copulas provide a powerful and versatile framework
for modeling multivariate dependence, with established applications in finance,
207
hydrology, and reliability analysis, as evidenced by key studies Patton (2012);
Salvadori et al. (2007); Balakrishnan and Lai (2009); Hofert et al. (2012). However,
their limitations in scalability, flexibility, and computational efficiency highlight the
need for ongoing research into hybrid models, such as those integrating vine copulas
or machine learning techniques, to address the demands of increasingly complex
datasets.
Limitations include their symmetry and restricted tail dependence, which Aas et
al. (2009) noted are less flexible than vine copulas for asymmetric structures. High-
dimensional validity requires strict generator conditions, limiting applicability
McNeil and Nešlehová (2009). Recent advances (2015–2024) address these via
nested structures and dynamic models but cannot fully overcome the inherent
constraints Durante and Sempi (2015).

10 Conclusion
Archimedean copulas constitute a cornerstone of multivariate dependence
modeling, offer- ing profound insights across domains such as finance, hydrology,
insurance, and environ- mental science. Anchored by their generator function and its
inverse, these copulas excel at capturing intricate dependence structures, particularly
in tail regions, where conventional correlation measures prove inadequate.
Established methodologies, including maximum likelihood estimation, simulation
algorithms, and goodness-of-fit assessments, provide ro- bust tools for analyzing
complex, high-dimensional datasets where independence assump- tions are
untenable.
Challenges remain, notably in scaling to ultra-high-dimensional settings,
accommo- dating non-stationary processes, and integrating machine learning
techniques to optimize generator selection and dependence calibration. Future
advancements may lie in hybrid frameworks, merging Archimedean copulas with
vine structures or neural network-based methods, paving the way for enhanced real-
time risk assessment and predictive modeling in an increasingly data-driven
landscape.
Ultimately, Archimedean copulas harmonize theoretical rigor with practical
applicabil- ity, serving as a vital instrument for understanding interconnected
systems—from global financial markets to climatic and networked infrastructures.
This review synthesizes their current capabilities and advocates for continued
innovation, encouraging their application to foster more resilient, data-informed
decision-making in a deeply interconnected world.

References
Sklar, A. (1959). Fonctions de re´partition a` n dimensions et leurs marges.
Publications de l‘Institut Statistique de l‘Universite´ de Paris, 8, 229–231.
Gumbel, E. J. (1960). Bivariate exponential distributions. Journal of the American
Statis- tical Association, 55(292), 698–707.

208
Clayton, D. G. (1978). A model for association in bivariate life tables and its
application in epidemiology. Biometrika, 65(1), 141–151.
Frank, M. J. (1979). On the simultaneous associativity of F(x,y) and x+y-F(x,y).
Aequa- tiones Mathematicae, 19(1), 194–226.
Marshall, A. W., and Olkin, I. (1988). Families of multivariate distributions.
Journal of the American Statistical Association, 83(403), 834–[Link], D.
(1989). Bivariate survival models induced by frailties. Journal of the
American Statistical Association, 84(406), 487–493.
Genest, C., and Rivest, L.-P. (1993). Statistical inference procedures for bivariate
Archimedean copulas. Journal of the American Statistical Association,
88(423), 1034– 1043.
Genest, C., Ghoudi, K., and Rivest, L.-P. (1995). A semiparametric estimation
procedure of dependence parameters in multivariate families of
distributions. Biometrika, 82(3), 543–552.
Joe, H. (1997). Multivariate Models and Dependence Concepts. Chapman and
Hall/CRC. Nelsen, R. B. (2006). An Introduction to Copulas. Springer.
Genest, C., Ré millard, B., and Beaudoin, D. (2009). Goodness-of-fit tests for
copulas: A review and a power study. Insurance: Mathematics and
Economics, 44(2), 199–213.
McNeil, A. J., and Nešlehová , J. (2009). Multivariate Archimedean copulas, d-
monotone functions and ℓ 1-norm symmetric distributions. The Annals of
Statistics, 37(5B), 3059– 3097.
Hofert, M. (2008). Sampling Archimedean copulas. Computational Statistics &
Data Anal- ysis, 52(12), 5163–5174.
Hofert, M. (2012). Estimators for Archimedean copulas in high dimensions.
Technomet- rics, 54(3), 243–256.
Genest, C., and Favre, A.-C. (2007). Everything you always wanted to know about
copula modeling but were afraid to ask. Journal of Hydrologic Engineering,
12(4), 347–368.
McNeil, A. J. (2010). Estimation of tail dependence and related measures in
Archimedean copulas. ASTIN Bulletin, 40(1), 113–129.
Patton, A. J. (2012). Copula methods for forecasting multivariate time series. In
Handbook of Economic Forecasting, 2, 899–960.
Hofert, M., Kojadinovic, I., Maechler, M., and Yan, J. (2012). copula: Multivariate
depen- dence with copulas. R package version 0.999-7.
McNeil, A. J., and Nešlehová , J. (2012). Nested Archimedean copulas and related
con- structions. Journal of Statistical Planning and Inference, 142(7), 1838–
1852.
Aas, K., Czado, C., Frigessi, A., and Bakken, H. (2009). Pair-copula constructions
and comparisons with Archimedean families. Insurance: Mathematics and
Economics, 44(2), 242–259.
Balakrishnan, N., and Lai, C.-D. (2009). Continuous Bivariate Distributions.
Springer.
209
Kojadinovic, I., Yan, J., and Holmes, M. (2011). Fast large-sample goodness-of-fit
tests for copulas. Statistica Sinica, 21(2), 841–871.
Charpentier, A., and Segers, J. (2014). Tails of multivariate Archimedean copulas.
Journal of Multivariate Analysis, 123, 252–264.
Bedford, T., and Cooke, R. M. (2002). Vines: A new graphical model for
dependent ran- dom variables. The Annals of Statistics, 30(4), 1031–1068.
Genest, C., and MacKay, J. (1986). The joy of copulas: Bivariate distributions
with uni- form marginals. The American Statistician, 40(4), 280–283.
Patton, A. J. (2006). Modelling asymmetric exchange rate dependence.
International Eco- nomic Review, 47(2), 527–556.
Salvadori, G., De Michele, C., Kottegoda, N. T., and Rosso, R. (2007). Extremes
in Nature: An Approach Using Copulas. Springer.
Segers, J., Schmid, F., and Stǎricǎ, V. (2014). Bootstrapping and rank-based
inference for copula models. Journal of Statistical Computation and
Simulation, 84(9), 1997–2014.
Durante, F., and Sempi, C. (2015). Principles of Copula Theory. Chapman and
Hall/CRC.
Joe, H., and Hu, T. (2015). Extensions of Archimedean copulas and their
applications. In
Copulae in Mathematical and Quantitative Finance, 123–142. Springer.

210
Technology Addiction among Students: An Analysis
Based on Students of Different Colleges under Dibrugarh
University- A Case Study
Kuldeep Goswami1 and Sricharan Shah2
1,2
Department of Statistics, Dibrugarh University, Dibrugarh-786004, Assam, India

Abstract
This study aims to explore phone addiction and its impact on students from
various colleges of Dibrugarh University, Dibrugarh, Assam, India, using statistical
methods. The analysis is based on a dataset of 283 students from different districts
who participated in the youth festival. The findings of the study reveal that students
from certain regions or districts of Assam, India are more prone to technology
addiction, particularly to mobile phones, compared to others.
Keywords: Phone addiction, Association of attributes, Independence of attributes,
Chi-square test

1. Introduction
Phone addiction, often referred to as internet addiction, Internet Use Disorder
(IUD), or Internet Addiction Disorder (IAD), is a relatively new phenomenon. It is
typically characterized as a serious issue where individuals struggle to control their
use of various technologies, particularly the internet, smartphones, tablets, and social
media platforms such as Facebook, Twitter, and Instagram [5, 6].
Mobile phones are one of the greatest inventions of the 20th century, and it's
hard to imagine life without them. It‘s clear that mobile phones offer numerous
benefits, making communication easier than ever before. In addition to calling,
mobile phones provide various functions, such as listening to music, chatting, and
playing games. In today‘s high-tech world, mobile phones are equipped with all the
necessary features, allowing people to chat for hours whenever they have time, from
day to day. However, excessive mobile phone use can have a negative impact on our
health. Do we truly understand the dangers of the radiation emitted by cell phones?
These waves can be harmful to our physical health, especially to the brain and heart.
Recent studies show that prolonged mobile phone use can seriously damage the
brain. For instance, talking on the phone for long periods can lead to headaches,
which are caused by the waves emitted by the device. Moreover, prolonged exposure


Corresponding author: email: charan.shah90@[Link]

211
to these waves may result in poor memory. Additionally, keeping a cell phone close
to the body, particularly under a pillow while sleeping, can cause heart problems due
to the strength of the radiation. Mobile phones also allow us to enjoy music and
videos anytime, thanks to headphones. While this is convenient, prolonged use of
earphones can damage our hearing and potentially lead to deafness. Excessive
mobile phone use, particularly among teens and young adults, can also lead to
restlessness and difficulty falling asleep. In conclusion, overuse of mobile phones
can negatively impact our brain, disrupt sleep, and damage our hearing [5, 6].
The diagnosis of phone addiction may vary across different countries, but
surveys conducted in the US and Europe indicates that between 1.5% and 8.2% of
the population suffers from internet addiction. Phone addiction is increasingly
recognized as a significant health issue in many other countries as well, including
Australia, China, Japan, India, Italy, Korea, and Taiwan [5, 6].

2. Objective of the Present Study


After reviewing the literature [5], it was decided that a survey would be
conducted among college students to understand their phone usage and its impact. A
questionnaire was developed for this purpose. We planned to carry out the survey
among the college students of Dibrugarh University, which has affiliated colleges
across seven districts in upper Assam. These districts vary in nature, with some
being industrialized and others rural, leading to differences in the lifestyles and
behaviors of the people. Consequently, the student behaviors from different districts
also show variation.
Every year, Dibrugarh University organizes a youth festival across its affiliated
colleges. In 2025, the youth festival was held at Swahid Peoli Phukan College in
Namti, Sivasagar from January 21‒ 24, 2025. Over 600 students from various
colleges participated in the event. We collected data from 283 participating students
through the distribution of questionnaires. It is anticipated that the feedback gathered
will provide insights into students' attitudes toward information technology and help
draw conclusions regarding phone addiction. Although this is a case study, it will
provide a broader understanding of phone addiction among the youth of Assam.
For data collection, we employed Convenient Sampling, a method in which sample
selection is based on the discretion or judgement of the researcher. This is sometimes
referred to as purposive or judgement sampling. In this method, the researcher
selects a sample of typical units from the population that are considered
representative or average. While this sampling technique is useful for opinion
surveys, it is not ideal for generalizing results, as it may introduce bias and prejudice
due to the researcher‘s subjective judgement. However, if the researcher is
experienced and skilled, this method can still yield valuable insights. One limitation
of this approach is that it does not allow for the calculation of the precision of
estimates from the sample data [4].

212
3. Methodology
The total number of students who attended the Youth Festival, held at Swahid
Peoli Phukan College in Namti, Sivasagar, is considered the population, with a total
population size of 1,061. A sample was selected from this population, and both the
sample size and confidence level were determined using statistical software,
specifically Raosoft [2], as detailed in Appendix B.
For analysis, various statistical tools and techniques were used, including Yule‘s
Coefficient of Association of Attributes [3] and the χ ²-test for the Independence of
Attributes [3]. To represent categorical data, a two-dimensional pie (or circular)
diagram was used, where each "slice" represents the proportion of the total
phenomenon attributed to each class or group. In these diagrams, the area of each
slice is proportional to the square of its radius. When making comparisons, pie
charts should be interpreted based on percentages, not absolute values [1].

4. Data Analysis and Findings


4.1. Diagrammatic Representation of Analysis
The pie diagram and bar graph representations of the various questions in the
questionnaire (see Appendix A) are provided below along with brief explanations.
These visual representations help to illustrate the responses and trends observed
from the survey data. The pie diagrams offer a clear view of the proportional
distribution of answers for categorical questions, while the bar graphs allow for easy
comparison of responses across different categories. The accompanying explanations
in Appendix C provide further context and interpretation of these visual data
representations, helping to enhance the understanding of the survey results.

1. Are you related to any NGO? 2. Do you check your phone in the morning for
messages and answering?
0% 20%
Y 25%0%
80% Y
N 75%
N
Figure 1
Figure 1 shows that most of the students Figure 2
use phone for personal communication. Figure 2 shows that most of the students check phone
after waking up, so they are somehow addicted to smart
phone.
3. During eating do you check or use your 4. On the way to college and home do you use phone or
phone? check for phone alert?
0% 42% 32% 0%
Y Y
58% 68% N
N No Response
Figure 3 Figure 4
In Figure 3, we see that more than 50% From Figure 4, we see that 68% students check their

213
students are conscious regarding their phone on the way to college and home which means that
eating habits still we see that some students they are too much addicted to phone and do not want to
have a tendency to check their phone miss a single call or message. We can conclude that this
during eating which can have negative breaks the attention on the present thing and can also say
effects on nutrition. that many road accidents takes place due to
unconsciousness to the surroundings.
5. Do you have a smart phone? 6. Do you spend time on virtual affairs such as chat or
romance?
0%
Y 0% 35%
20% Y
70% N 65% N
No Response
No Response

Figure 5 Figure 6
Figure 5 shows that 70% students use Figure 6 shows another interesting view that is 35%
smart phone, well in this generation we can students spends time on virtual affairs such as chat or
say that the mobile phones are an important romance, which is a natural phenomena seen in human
accessories of human being and if it is a being. As there is a saying that ―too much of anything is
smart phone it works wonders, and these harmful‖, the same way too much of addiction regarding
phones has influenced the students too. A this will affect the relationship with the real life partners
smart phone has many features, one can and may create distances too. But most of them are not
perform any kind of work within no engaged in such kind of activities.
minutes, any information regarding the
north pole to the south pole, any message,
and any call etc. efficiently.
7. Do you gamble in internet? 8. Do you have information overloads?

0% Y 1% 26% Y
34% N
66% N
No Response 73%
No Response
Figure 7
Figure 7 shows that 66% i.e., most of the Figure 8
students do not gamble. They are sensitive Figure 8 shows, most of the students do not piled
to spending. information unnecessarily. They asses information
whenever they require.
9. Are you obsessed with social media? 10. Are you obsessed with entertainment?
1% 1%
Y 31% Y
51% 48% N
68% N
No Response

Figure 9
Figure 10
Figure 9 shows, Half the students are using
Figure 10 shows, Most of the students are using phone
social media. This indicates that they are for listening music or watching videos, cinema etc.
very much aware about the things that are
going on in the society. This shows the
consciousness of the students also.

214
11. Total number of messages in a day you receive and answer:

Receive Answer
33%
Receive 4% 6% 0-150 4% 2% 0-150
67%
Answer
150-
90% 300 94% 150-300
300 &
Above
300 & Above
Figure 11

Figure 12 Figure 13
From Figure 11‒ 13, about 90% students receive and answer up to 150 messages in a day which is not
abnormal in the present day.
12. Total number of phone calls in a day you receive and answer:
Figure 15
Receive Answer
Answer 14% 0-10
Receive
41% 20% 0-10
59% 32% 27% 59%
10-20
10-20

48% 20 &
20 & Above
Above
Figure 14
Figure 16
From Figure 14‒ 16, less than 50% students do not make excess calls that are they use judiciously.
13. Total time you spent on phone in a day: 14. Total time you spend on online playing in a day:

24% 6% 6% 4%
0-5
Less than 1 hour
46%
5-10
30% within1 and 2 hour
84%
10-15
More than 2 hour
15-20

Figure 17 Figure 18
Figure 17 shows time spent on phone Figure 18 shows most of the students do not play online
varies for students. 24% use it for less than games.
an hour, 30% use it up to 2 hours and a
significant portion 46% use it for more
than 2 hours.

215
15. Total number of online purchase orders in a month:

0-5
8% 3%
8% 5-10

10-15

81% 15-20

Figure 19
Figure 19 shows most of the students purchase things on online but the number is less that is they are
not obsessed to the amazing items available on e-market.
16. Total numbers of e-mails received and sent in a week:

Receive Answer
0-100
1% 0%
0-100
Receive 2%1%
100-
28% 99%
100-200 200
72% 97%
Answer 200 &
Above
200 &
Above

Figure 20 Figure 21 Figure 22


Figure 20‒ 22 shows almost all students receive and answer less than 100 e-mails per week.
1
17. Since how many years you are using 8. Average phone bill (in Rs.) in a month:
mobile phone?

0-5
2% Below 200
3% 2% 5-10

10-15
35% 27% 200 - 500
58% 52%
15-20 21%
No Response More than 500

Figure 23
From Figure 23, it is observed that half of
the students are using mobile phones after Figure 24
passing HSLC examination. Figure 24 shows half of the students spent less than
rupees 200 per month in phone. The spending on phone
is moderate.

216
Districts-wise distribution of male and female responses

250
Responses of Male and

200
54% 51% 42% 56% 37%
65% 85%
150 Female No
46% 49% 58% 44% 63%
Female

35% 15%
100 31% 32% Female Yes
36% 46% 53% 43% 45%
50 69% 68% Male No
64% 54% 47% 57% 55%
0 Male Yes
Jorhat Dibrugarh Tinsukia Lakhimpur Sivsagar Golaghat Dhemaji

Districts

Figure 25
Figure 25 shows that more than 50% male students are addicted to phone and 40% female students are
addicted to phone of different districts.
4.2. Yule‘s Coefficient of Association of Attributes
The coefficient of association of attributes is a statistical measure used to
determine the strength and direction of the association between two binary variables.
It quantifies the relationship between attributes by comparing the frequency of
occurrences in different categories. This method is particularly useful in analyzing
categorical data and helps in understanding how two variables are related, which can
be valuable in identifying patterns or trends in the data. In this study, Yule's
Coefficient is employed to assess the relationship between various attributes in the
dataset, providing deeper insights into how different factors are associated with
phone addiction among students.
Let A = positive response of male
B = positive response of female
H : there is no positive responses between male and female
0

H : there is a positive responses between male and female


1

Table 1: Calculation of coefficient of association of attributes


A
A α Total
B
B (AB) = 1299 (α B) = 1140 (B) = 2439
β ( Aβ ) = 1479 (α β ) = 1320 ( β ) = 2799
Total (A) = 2778 (α ) = 2460 N = 5238

Q
 AB αβ   Aβ (αB)  1299)(1320)  (1479)(1140  0.00842
 AB αβ   Aβ (αB) 1299)(1320)  (1479)(1140
Based on the above computation using Yule's Coefficient of Association of
Attributes, we obtained a value of Q  0.00842 , which is approximately 0. This
suggests that there is no significant association between the positive responses of
males and females. In other words, the positive responses of males from different
districts are independent of the positive responses of females from different districts,
and vice versa.
217
4.3. 𝝌𝟐 -test for Independence of Attributes
The χ ²-test for the independence of attributes is a statistical method used to
determine if there is a significant relationship between two categorical variables. The
test compares the observed frequencies of occurrences in each category with the
expected frequencies if the variables were independent. The χ ² statistic is calculated
by summing the squared differences between observed and expected frequencies,
divided by the expected frequencies. This test is widely used in research to analyze
the association between categorical data and helps to identify patterns or
dependencies in various fields, including social science, marketing, and medical
research. In this study, the χ ²-test is employed to assess the relationship between
different attributes and their impact on phone addiction among students.
Let us setup the null hypothesis
H 0 : no distinction is made in responses on the basis of districts under Dibrugarh
University and Sex
against the alternative hypothesis
H1 : there exists a distinction in responses on the basis of districts under
Dibrugarh University and Sex

Table 2: Calculation of Chi-square test of independence of attributes


(𝑶𝒊 − 𝑬𝒊 )𝟐
Districts Sex Responses 𝑶𝒊 𝑬𝒊 𝑶𝒊 − 𝑬𝒊 (𝑶𝒊 − 𝑬𝒊 )𝟐
𝑬𝒊
Yes 64 63.23 0.77 0.5929 0.009377
Male
No 36 44.46 -8.46 71.5716 1.609798
Jorhat
Yes 46 37.85 8.15 66.4225 1.754888
Female
No 54 54.46 -0.46 0.2116 0.003885
Yes 54 63.23 -9.23 85.1929 1.347349
Male
No 46 44.46 1.54 2.3716 0.053342
Dibrugarh
Yes 35 37.85 -2.85 8.1225 0.214597
Female
No 65 54.46 10.54 111.0916 2.039875
Yes 47 63.23 -16.23 263.4129 4.165948
Male
No 53 44.46 8.54 72.9316 1.640387
Tinsukia
Yes 15 37.85 -22.85 522.1225 13.79452
Female
No 85 54.46 30.54 932.6916 17.12618
Yes 57 63.23 -6.23 38.8129 0.613837
Male
No 43 44.46 -1.46 2.1316 0.047944
Lakhimpur
Yes 49 37.85 11.15 124.3225 3.28461
Female
No 51 54.46 -3.46 11.9716 0.219824
Yes 55 63.23 -8.23 67.7329 1.071215
Male
No 45 44.46 0.54 0.2916 0.006559
Sivsagar
Yes 58 54.46 20.15 406.0225 10.72715
Female
No 42 54.46 -12.46 155.2516 2.850746

218
Yes 79 63.23 15.77 248.6929 3.933147
Male
No 31 44.46 -13.46 181.1716 4.074935
Golaghat
Yes 44 37.85 6.15 37.8225 0.999273
Female
No 56 54.46 1.54 2.3716 0.043548
Yes 68 63.23 4.77 22.7529 0.359843
Male
No 32 44.46 -12.46 155.2516 3.491939
Dhemaji
Yes 63 37.85 25.15 632.5225 16.71129
Female
No 37 37.85 -17.46 304.8516 5.597716
𝝌𝟐 = 97.79372

The tabulated 𝜒 2 at (𝑛 − 1) i.e., (28 − 1) = 27 degrees of freedom is 40.113.


Using the χ ²-test for the independence of attributes, we obtained a calculated χ ²
value of 97.79372, which is greater than the tabulated value of 40.113 at 27 degrees
of freedom (28-1) and a 5% level of significance. This difference is statistically
significant, leading to the rejection of the null hypothesis. Therefore, we can
conclude that there is a significant difference in responses based on both districts
under Dibrugarh University and gender. This suggests that mobile phone usage
varies between districts, with some regions having higher mobile usage than others.

5. Conclusion
From our empirical investigation, the following results were observed using
statistical techniques:
 The graphical representation clearly indicates that most students are
addicted to phones, primarily for personal communication.
 Using the coefficient of association of attributes, we obtained a Q value of
0.00842 (approximately 0), which suggests that the association between the
positive responses of males and females is independent. In other words, the
positive responses of males from different districts are not dependent on the
positive responses of females from different districts, and vice versa.
Similarly, the negative responses of males and females are also independent
across different districts.
 Applying the χ ²-test for the independence of attributes, we calculated a χ ²
value of 97.79372, which is greater than the tabulated value of 40.113 at 27
degrees of freedom (28-1) and a 5% level of significance. This result is
statistically significant, leading to the rejection of the null hypothesis. Thus,
we can conclude that there is a significant difference in responses based on
both districts under Dibrugarh University and gender. In other words,
mobile usage patterns differ between districts, with mobile users in one
district behaving differently from those in another district. Similarly, male
and female mobile users in each district also show different patterns,
indicating that some regions or districts have higher mobile usage than
others.

219
Technology significantly benefits humanity, but its overuse or misuse can lead to
negative consequences. From the above discussion, it is evident that there is a
moderate tendency of phone addiction among college students. The analysis also
shows that 50% of the students surveyed are heavy users of technology. While
technology is essential in modern life, it should not overpower us. To maintain a
balanced life, a judicious use of phones is recommended. Overall, it can be
concluded that a significant number of students across different regions or districts
are addicted to mobile phones.

References
[1] Rajkhow, Deba Pallav: Essential Statistics for Commerce and Applied
Subjects: A Textbook of Business Statistics for [Link], 3rd Semester,
Dibrugarh University (2013).
[2] Raosoft Software for determination of sample size.
[3] Gupta, S. C. & Kapoor, V. K.: Fundamentals of Mathematical Statistics,
Sultan Chand & Sons (2008).
[4] Gupta, S. C. & Kapoor , V. K.: Fundamentals of Applied Statistics by, Sultan
Chand and Sons (1997).
[5] Science Reporter, page numbers 12-16, August, 2014.
[6] Goswami, Priya Dev & Ahmed, Nazimuddin: Phone Addiction: Its Effect
And Impact On College Students- A Case Study, International Journal of
Management and Social Science Research Review, Vol-1, Issue – 29, Nov -
2016 Page 61, IJMSRR, E- ISSN - 2349-6746, ISSN -2349-6738(2016).

220
Apendix: A
Questionnaire
Survey on Phone Addiction among College Students

Serial No:
Date:…………………
(Tick the boxes)

(A) Name: Mr/Ms


…………………………..........................................................................
(B) Name of the college:
………………………...................................................................
(C) Class: ..................... (D) Age: ................…. (E) Male Female

1. Are you related to any NGO? Yes No


2. Do you check your phone in the morning for messages Yes No
and answering?
3. During eating do you check or use your phone? Yes No
4. On the way to college and home do you use phone or Yes No
check for phone alert?
5. Do you have a smart phone? Yes No
6. Do you spend time on virtual affairs such as chat or Yes No
romance?
7. Do you gamble in internet? Yes No
8. Do you have information overloads? Yes No
9. Are you obsessed with social media? Yes No
10. Are you obsessed with entertainment? Yes No
11. Total number of messages in a day you receive and answer
12. Total number of phone calls in a day you receive and answer
13. Total time you spent on phone in a day:
Less than 1 hour Within 1 - 2 hours More than 2 hours
14. Total time you spend on online playing in a day:
15. Total number of online purchase orders in a month:
16. Total numbers of e-mails received and sent in a week.
17. Since how many years you are using mobile phone?
18. Average phone bill (in Rs.) in a month:
Below 200 200 to 500 More than 500

Signature of interviewer

221
Appendix: B

Sample size calculator

What margin of error can you accept? 5% The margin of error is the amount of error that
5% is a common choice you can tolerate. If 90% of respondents
answer yes, while 10% answer no, you may be
able to tolerate a larger amount of error than if
the respondents are split 50-50 or 45-55.
Lower margin of error requires a larger
sample size.
What confidence level do you need? 95% The confidence level is the amount of
Typical choices are 90%, 95%, or 99% uncertainty you can tolerate. Suppose that you
have 20 yes-no questions in your survey. With
a confidence level of 95%, you would expect
that for one of the questions (1 in 20), the
percentage of people who answer yes would
be more than the margin of error away from
the true answer. The true answer is the
percentage you would get if you exhaustively
interviewed everyone. Higher confidence level
requires a larger sample size.
What is the population size? 1061 How many people are there to choose your
If you don't know, use 20000 random sample from? The sample size doesn't
change much for populations larger than
20,000.
What is the response distribution? 50% For each question, what do you expect the
Leave this as 50% results will be? If the sample is skewed highly
one way or the other, the population probably
is, too. If you don't know, use 50%, which
gives the largest sample size.
Your recommended sample size is 283 This is the minimum recommended size of
your survey. If you create a sample of this
many people and get responses from
everyone, you're more likely to get a correct
answer than you would from a large sample
where only a small percentage of the sample
responds to your survey.

222
A New Form of Two Parameters Skew-Logistic
Distribution and its Real Life Application
Jondeep Das1, Partha Jyoti Hazarika 2*, Dimpal Pathak3
1
Department of Statistics, Bhattadev University, Bajali, Assam, India.
2*
Department of Statistics, Dibrugarh University, Dibrugarh, Assam, India.
3
Department of Statistics, Bahona College, Jorhat, Assam, India.

Abstract
Many real-world datasets exhibit skewness that symmetric models like the
normal or logistic distributions cannot capture. This paper proposes a new Two-
Parameter Skew Logistic distribution that extends the skew logistic model by
introducing two shape parameters, offering greater flexibility in modeling
asymmetry. Key properties such as the probability density function, cumulative
distribution function, and moments are derived, and numerical results for mean,
variance, skewness, and kurtosis are presented. Parameters are estimated using the
maximum likelihood method, and the model‘s performance is illustrated using a real
dataset on lake acidity. Comparative analysis through AIC and likelihood ratio tests
shows that the TPSL distribution provides a better fit than existing logistic-based
models, confirming its usefulness for modeling asymmetric data.
Keywords : Logistic distribution, Skew logistic, Maximum Likelihood Estimation,
LR test.

1. Introduction:
The normal distribution, also known as the Gaussian distribution, is a
continuous probability distribution that is widely used in statistics and mathematics.
It is a continuous probability distribution that is symmetric about its mean, with the
probability density function taking the form of a bell-shaped curve. But in many
real-world applications, data can be skewed to the left or to the right, and a normal
distribution assumption may not be appropriate. To deal with that kind of situations,
Azzalini (1985) introduced the skew normal (SN) distribution that extends the
normal distribution to allow for skewness in the data. The probability density
function (pdf) of the SN distribution is given by
 SN x;    2 xx, z  R, (1)
where,   R,  . and . represents the pdf and cumulative distribution
function (cdf) of the standard normal distribution, respectively.
Following Azzalini‘s (1985) model, a number of research works has so far been
carried out to present different skew normal distributions derived from the
underlying symmetric one to model the asymmetric behavior of empirical data

223
which are suitable under different conditions. For example Henze (1986), Azzalini
and Dalla-Valle (1996), Branco and Dey (2001), Arnold and Beaver (2002) and
others studied skew normal distribution of Azzalini (1985) from different
perspectives. Besides, the small-sample behavior of different estimators of the
logistic regression model parameters with skew-normally distributed explanatory
variables was studied by Matin and Bagui (2008).
The skew normal distribution has an interesting characteristic in that it acts like
a half normal distribution as the skewing parameter approaches infinity. To deal with
such type of situations, Arellano-Valle et al. (2004) introduced Skew Generalized
Normal (SGN) distribution which is considered as generalization of Skew Normal
distribution of Azzalini (1985). The pdf of the SGN distribution is given by
 1 x 
f ( x; 1 , 2 )  2  ( x)    ; x  R, (2)
 1   x2 
 2 
where, 1  R , 2  0 ,  (.) and  (.) are the pdf and cdf of standard normal
distribution respectively. Later on, Choudhury & Matin (2011) introduced another
more flexible distribution namely extended skew generalized normal (ESGN)
distribution which generalizes both skew normal of Azzalini (1985) as well as skew
generalized normal of Arellano-Valle et al. (2004).
Borrowing the idea of existing form of skew normal distribution of Azzalini
(1985), many researchers have developed skewed distributions with some other
symmetric distributions, viz., Logistic, Laplace, Uniform, Cauchy etc. Skew
Logistic (SLG) distribution was introduced by Wahed & Ali (2001) under the
Azzalini‘s (1985) mechanism of skew normal distribution considering the pdf and
cdf of standard logistic distribution instead of standard normal distribution. The pdf
of SLG distribution is given by
2 ex
f ( x;  )  ; x R ,   R , (3)
1  e  1  e 
x 2  x

where,  is the skewness parameter. Nadarajah (2009) extended the skew


logistic distribution by introducing a scale parameter and studied its distributional
properties. Further, Gupta & Kundu (2010) studied the generalization of the skew
logistic distribution of Nadarajah (2009). Deriving the skewed version of the
generalized logistic distribution of Rathie & Swamee (2006), Rathie & Coutinho
(2011) discussed its suitability in real life using data modeling. Under the
Balakrishnan (2002) mechanism, Asgharzadeh et al. (2016) introduced an extension
of the skew logistic distribution.
In this article a family of skew logistic distribution has been introduced under
the mechanism of skew generalized normal distribution of Arellano-Valle et al.
(2004) introducing one additional parameter which generalizes the skew logistic
distribution and logistic distribution. The newly introduced skew distribution is more
flexible than that of skew logistic distribution.

224
The rest of the article is organized as follows. In section 2, the new skew
logistic distribution has been introduced along with the visualization of density and
its properties. Section 3 presents some important distributional properties of the
studied distribution. Mathematical optimization of the new distribution using the
maximum likelihood procedure has been included in section 4. The numerical
example of some real-life applications of the new distribution is provided in section
5. Finally, section 6 provides hypothesis testing of the distribution while conclusions
appear in section 7, followed by the references.

2. A Two-Parameter Skew Logistic Distribution


This section introduces a new form of two parameter skew logistic distribution
and investigates its basic properties.
Definition: A random variable X follows Two Parameter Skew Logistic distribution,
denoted by TPSL( 1 , 2 ) , if it has the density function
2 ex (4)
f ( x; 1 , 2 )  x
; x R , 1  R ,  2  0
 1

1  e  x 2 
 1  e
1  x 2 2 

 
1
Using Tylor Series Expansion for (1  z) one can write single series representation
of the trimodal skew logistic density in (4) as
𝜆𝑙 𝑗
−𝑥61+ 7
−1
2 ∑∞
𝑗=0 . 𝑗 / 𝑒
√1+𝑥 𝑚 𝜆𝑚
, 𝑥 ≥0
(1 + 𝑒 −𝑥 )2
𝑓(𝑥) = (5)
𝜆𝑙 𝜆𝑚 𝑗
−𝑥61− − 7
−1
2 ∑∞
𝑗=0 . 𝑗 / 𝑒
√1+𝑥 𝑚 𝜆𝑚 √1+𝑥 𝑚 𝜆𝑚
, 𝑥<0
{ (1 + 𝑒 −𝑥 )2
By expanding the terms of equation (5), the double series representation can be
written as
𝜆𝑙 𝑗
−𝑥<1+𝑘+ =
−1 −2 √𝑙+𝑥𝑚 𝜆𝑚
2 ∑∞
𝑗=0 . 𝑗 / ( 𝑘 )𝑒 , 𝑥 ≥0
𝑓(𝑥) = (6)
𝜆𝑙 𝜆𝑚 𝑗
−𝑥<1+𝑘+ + =
∞ −1 −2 √𝑙+𝑥𝑚 𝜆𝑚 √𝑙+𝑥𝑚 𝜆𝑚
{ 2 ∑𝑗=0 . 𝑗 / ( 𝑘 )𝑒 , 𝑥<0

2.1. Special case of TPSL( 1 , 2 ) distribution:


 If 1  0 then X ~ Logistic (0,1).
 If 2  0 then X ~ SLG (1 ).
 If 1  0, 2  0 then X ~ Logistic (0,1).
 If 1  0, 2  1 then X ~ Logistic (0,1).
 If X ~ TPSL(1 , 2 ) then  X ~ TPSL(1 ,2 )
225
2.2 Plots of density function
The density functions of TPSL( 1 , 2 ) for different choice of 1 and 2 is
plotted in Figure 1. Figure 1 illustrates the behavior of the probability density
function (PDF) of the Two-Parameter Skew Logistic (TPSL) distribution for
different combinations of the shape parameters 1 and 2 . The plots demonstrate
that the parameters jointly control the shape, skewness, and tail thickness of the
distribution. For small or equal values of parameters, the density remains nearly
symmetric and resembles the standard logistic form. As the values of parameters
diverge, the distribution becomes asymmetric, shifting the peak either to the left or
right, depending on the sign and magnitude of λ 1. Positive values of λ 1 generate
right-skewed densities with longer right tails, while negative values yield left-
skewed shapes with heavier left tails. Furthermore, λ 2 influences the steepness and
flatness of the curve, controlling the concentration of probability around the mode.
Overall, the graphical representations confirm that the proposed TPSL distribution is
highly flexible, capable of representing a wide variety of symmetric and asymmetric
patterns, making it suitable for modeling diverse types of skewed data encountered
in real-world applications.
fX x; 1 2 f X x; 1 2
0.35
0.4

0.30

0.3 0.25

0.20
0.2
0.15

0.10
0.1
0.05

x x
5 5 5 5

1 2, 2 0.6 1 2, 2 5.62 1 4.4, 2 1.6 1 4.4, 2 5.62

f X x; 1 2
f X x; 1 2
0.35
0.35
0.30
0.30
0.25
0.25
0.20
0.20
0.15
0.15

0.10
0.10

0.05 0.05

x x
5 5 5 5

1 2, 2 0 1 0, 2 0 1 2.9, 2 3.6 1 2, 2 5.62

Figure 1: The density of TPSL with different values of λ1 and λ2

226
3. Distributional Properties of TPSL( 1 , 2 )
In this section some of the distributional properties related to the two parameter
skew logistic studied distribution have been discussed.

3.1 Cumulative distribution function of TPSL( 1 , 2 ) :


Theorem 1: The cdf of TPSL( 1 , 2 ) distribution can be expressed as

2 ex
x
F ( x; 1 , 2 )    x 1
dx (7 )
 
1 e  2 

x 1  x 2
2

 1  e 
 

This immediately leads to the following properties:


 F ( x 1  0, 2 ) obtains density function of standard logistic distribution.
 F ( x 1 , 2 )  1  F ( x  1 , 2 ).

3.2 Moments of TPSL(1 , 2 ) :


Theorem 2: The rth order moment of TPSL( 1 , 2 ) distribution is given by


2 ex
E X r   X
r
 x 1
dx (8)
 
1 e  2 

x 1  x 22

 1  e 
 
And hence

2 ex
EX   x  x 1
dx  1 , (9)
 
1 e  2 

x 1  x 2
2

 1  e 
 

2 ex
E X 2   x
2
 x 1
dx   2 , (10)
 
1 e  2 

x 1  x 2 2
 1  e 
 
But the equation (8), (9) and (10) contain very complicated mathematical form.
 
Hence numerical values for E  X  and E X are calculated for different values of
2

1 and 2 and hence variance of TPSL( 1 , 2 ) distribution are calculated using the

227
 
formula Var ( X )  E X 2  E ( X ) in Table 1. Similarly value of the coefficient
2

of skewness and kurtosis for the new distribution are reported in Table 2.

Table 1: Mean and Variance of TPSL( 1 , 2 ) for different values of 1 and 2


𝜆1 𝜆2 Mean Variance
-1.50 2.00 -0.6109 2.9166
-2.75 1.52 -1.0293 2.2305
-3.50 0.50 -1.2841 1.6411
0.50 0.50 0.3654 3.1563
1.00 1.00 0.5478 2.9898
1.50 2.50 0.5624 2.9736

Table 2: Skewness and Kurtosis of TPSL( 1 , 2 ) for different values of 1 and 2 .

𝜆1 𝜆2 Skewness Kurtosis
-1.50 2.00 -0.596 3.715
-2.75 1.52 -0.841 3.982
-3.50 0.50 -1.126 4.214
0.50 0.50 0.274 3.376
1.00 1.00 0.558 3.683
1.50 2.50 0.861 4.027

4. Parameter Estimation of TPSL( 1 , 2 )


4.1 Location and Scale Extension :
In this section the location and scale extension of TPSL distribution has been
considered. Considering the transformation Y     X , a location scale
generalized two parameter skew logistic distribution has been obtained which pdf is
given as
 x  
 
2e   
f ( x; 1 , 2 ,  ,  )  ; x R , 1  R ,   R, 2  0 ,   0.
  x  
  1 
    
2  
 
2
 x    x  
1    (11)
1  e      1  e    2


   
   
 
 

228
4.2 Maximum Likelihood Estimation:
Suppose, z1, z2, z3 ,..., zn are independently and identically distributed random
variable drawn from the flexible alpha skew normal distribution, then log-likelihood
function for    ,  , 1 , 2  is obtained as
  x  
  1 
    
  
 
2
 x    x  
 1    2
1 n n n
 
l ( )  n log 2   zi     n log   2 log 1  e       log 1  e   
 (12)
 i 1 i 1
  i 1
   
 
 

Differentiate equation (12) with respect to parameters    ,  , 1 , 2  , the


likelihood equation becomes
 l ( ) n n
exp  ( zi   )   n
1
2
   
A(  ,  , 1 , 2 ) 2 zi    2  zi 2 12  A(  ,  , 1 , 2 )1
  2  ,
  i 1  1  exp  ( zi   )   1  exp  zi   1  1  zi     2 
2
i 1
 
 l ( ) n n
 zi    2 n  exp  ( z   )   z    
        i i

  i 1   2   2 i  1  exp ( zi   )  
1 

n
1
 1
 3

  3 A(  ,  , 1 , 2 )   3 zi  3 zi   zi 12
2 2 3
 
  ,
1  exp    z     1   zi    2 2    12 B(  ,  , 1 , 2 )  zi    
i 1

  i 1
    

 l ( ) n
B(  ,  , 1 , 2 )  zi   
 ,
 1 i 1 1  exp 
   zi   1  1   zi     2  
2

  

 l ( )  1
  
 
n
1
   2  A(  ,  , 1 , 2 ) zi  3zi   3zi    1 ,
3 2 2 3

 2   2  
i 1 1  exp   z     1   z        
 2 
  i 1 i

where,
exp   zi   1  1  zi     2  exp   zi   1  1  zi     2 
2 2

A(  ,  , 1 , 2 )    & B(  ,  , 1 , 2 )   

 1  zi    2 2 
3
2  1  zi     2
2

Some numerical optimization routines are required to obtain the simultaneous
solution of the desired estimate of the parameters.

4. Simulation Study
A simulation study was implemented to evaluate the behaviour of the estimated
parameters of the TPSL( ,  , 1 , 2 ) distribution using the Metropolis-Hastings
(MH) algorithm. Results are based on the 1000 samples with three different samples
sized n  100,300 and 500. The GenSA function in R software maximised the
likelihood function for each generated sample. The calculated statistics are presented

229
in terms of the biases and mean square errors of the estimates. The formula for the
same are given as

      
2

   
Bias ˆ  E ˆ   and MSE ˆ  V ˆ  Bias ˆ
The study observed that the parameter recovery was satisfactory for the selected
sample sizes. The simulation study also noted that the value of bias and MSE of the
estimated parameters was decreased with an increasing number of sample size,
indicating the excellent behaviour of the MLEs of the model, see Table 3-4. Tables
3-4 reports the results.

Table 3: Results of Simulation


𝜇 = 0, 𝛽 = 1
𝜇 𝛽 𝜆1 𝜆2
𝜆2 𝜆1 n Bias MSE Bias MSE Bias MSE Bias MSE
100 0.4639 0.3296 -0.3783 0.1724 0.2504 0.0912 0.2404 0.1035
300 0.4003 0.2977 -0.3310 0.1197 0.2439 0.0862 0.1709 0.0962
-1 500 0.3628 0.2428 -0.3177 0.1175 0.2411 0.0745 0.1414 0.0893
100 0.3078 0.2318 -0.4319 0.1148 0.6532 0.0876 0.3926 0.1111
300 0.2987 0.1165 0.3518 0.1079 -0.4987 0.0533 0.3819 0.0438
0.5 1 500 0.2934 0.0654 0.2265 0.0136 0.3659 0.0510 0.2146 0.0344
100 -0.5340 0.2234 -0.3813 0.1928 -0.4653 0.2343 0.1201 0.1592
300 -0.5195 0.2123 -0.3692 0.1578 -0.4005 0.2208 0.1142 0.1433
-1 500 0.0437 0.2002 -0.3371 0.1276 0.3992 0.2116 0.1056 0.1371
100 -0.5125 0.2220 -0.4107 0.2694 -0.4894 0.2426 0.3515 0.2192
300 0.4599 0.2176 0.3819 0.2388 0.4831 0.2044 0.2690 0.1821
1.5 1 500 0.4639 0.3296 -0.3783 0.1724 0.2504 0.0912 0.2404 0.1035
100 -0.4330 0.2182 -0.0577 0.0176 -0.3122 0.1668 -0.3707 0.2148
300 -0.4109 0.2030 -0.0281 0.0065 -0.3016 0.1241 -0.3495 0.2009
-1 500 -0.4029 0.2006 -0.0083 0.0026 -0.3010 0.111 -0.2905 0.1654
100 0.3022 0.0429 -0.1288 0.0345 -0.5640 0.0432 -0.3452 0.1097
300 0.2112 0.0411 -0.1139 0.0311 -0.3343 0.0239 -0.2378 0.0976
2 1 500 0.2011 0.0235 0.0987 0.0265 -0.2898 0.0129 0.2199 0.0154

230
Table 4: Results of Simulation
𝜇 = 1, 𝛽 = 2
𝜇 𝛽 𝜆1 𝜆2
𝜆2 𝜆1 n Bias MSE Bias MSE Bias MSE Bias MSE
100 0.3219 0.2291 -0.4199 0.3199 0.4991 0.3298 0.3100 0.4184
300 0.2182 0.2166 0.3999 0.2901 0.4712 0.3100 -0.2990 0.4000
-1 500 -0.2011 0.1981 0.3811 0.2811 -0.3614 0.2766 0.2900 0.3911
100 0.1845 0.3618 -0.2814 0.4185 0.2615 -0.8621 -0.3318 0.0921
300 -0.1719 0.3172 0.2617 0.4210 0.2510 0.7514 0.2341 0.0715
0.5 1 500 0.1710 0.2001 -0.2221 0.3871 -0.2511 0.6199 0.2301 0.0702
100 0.2889 0.3354 -0.6110 0.4519 0.7310 0.3317 -0.4188 0.2001
300 0.2801 0.3311 0.5996 0.4102 0.6932 0.3180 -0.3777 0.0988
-1 500 0.2675 0.3112 0.4980 0.3711 0.6151 0.2900 -0.3510 0.0888
100 -0.1991 0.0800 -0.8311 0.6107 -0.4229 0.3889 -0.5192 0.2739
300 -0.1821 0.0810 0.7301 0.5868 -0.3981 0.3312 0.4913 0.2677
1.5 1 500 0.1809 0.0721 0.7009 0.4722 0.3706 0.2221 0.4660 0.2449
100 0.1504 0.0329 -0.4301 0.2198 0.2005 0.0870 0.5599 0.3490
300 -0.0341 0.0301 0.4160 0.1998 -0.0669 0.0611 -0.6957 0.2198
-1 500 0.0288 0.0241 -0.3300 0.1997 0.666 0.0399 0.4141 0.1765
100 0.3771 0.2741 -0.2890 0.3441 0.1651 0.2441 0.1442 0.0317
300 -0.3199 0.2699 -0.2714 0.3201 0.1715 0.2211 0.1312 0.0300
2 1 500 0.3001 0.2500 0.2699 0.3187 0.1700 0.1851 0.0954 0.0217

5. Real Life Application


In this section, the applicability of TPSL distribution with one real life data set
has been illustrated. The logistic distribution LG ( ,  ) , skew logistic distribution
SLG( ,  ,  ) , alpha skew logistic distribution ASLG(,  , ) has been fitted as
well as compared these models with our newly introduced model i.e., TPSL(,  , 1 , 2 ) .
The values of the fitted models have been carried out by maximum likelihood
methods using GenSA package in R software. For comparing the newly introduced
model with the existing skew models, Akaike Information Criterion (AIC) which is
the analytical measure has been considered.
A data set of acidity index measured in a sample of 155 lakes in the
Northeastern United States has been considered which was previously analyzed as a
mixture of Gaussian distributions on the log scale by Crawford (1994). Different non
parametric plots for the said dataset are presented in Figure 2. The graphical
summaries presented in Figure 2 provide an overview of the distributional

231
characteristics of the dataset. The box plot indicates a moderate spread of values
with no extreme outliers, suggesting a relatively consistent data pattern. The
histogram shows a slightly bimodal distribution, implying the presence of two
dominant value ranges within the sample. The violin plot further confirms this
bimodal tendency, displaying two distinct peaks around the lower and higher ends of
the distribution. The scatter plot of scores suggests that the data follow an
approximately increasing trend, consistent with a non-random structure. The strip
plot demonstrates a fairly uniform scattering of individual observations without
notable clustering. Finally, the kernel density plot provides a smooth visualization of
the underlying probability density, again reflecting bimodality with peaks near 3.5
and 6.5. Overall, the data exhibit moderate variability and potential sub-group
structure, indicating that further modeling may need to account for mixed or
multimodal behavior. Table 3 obtained the maximum likelihood estimate of the fitted
models with the value of log-likelihood and AIC.

Figure 2: Non parametric plots for the acidity index dataset.

232
Table 5: Results of MLE‘s, log-likelihood and AIC for fitted models

Distribution     1 2 log L AIC


LG ( ,  ) 5.023 0.631 --- --- --- --- -232.796 469.592
SLG (,  ,  ) 3.845 0.941 7.681 --- --- --- -210.306 426.612
ASLG(,  ,  ) 4.537 0.579 --- 0.272 --- --- -226.179 458.358
TPSL( ,  , 1 , 2 ) 3.857 0.935 --- --- 14.251 7.434 -206.855 419.710

From the Table 5 it has been clear that the values of AIC are less than for
TPSL(  ,  , 1 , 2 ) distribution in comparison to the other regular distributions.
Besides, the profile likelihood plots provided in Figure 3 indicate that all parameters
are identifiable, with varying degrees of estimation precision. The observed
asymmetries, particularly for the shape parameters λ ₁ and λ ₂ , emphasize the
necessity of profile likelihood methods for reliable statistical inference in this four-
parameter distribution framework. So, it may be concluded that the
TPSL(  ,  , 1 , 2 ) distribution provides best fit to the data set in terms of AIC.

Figure 3: Log-likelihood profiles for acidity index dataset.

233
6. Hypothesis Testing
It has been observed that, LG ( ,  ) , SLG ( ,  ,  ) , ASLG(,  ,  ) and
TPSL(  ,  , 1 , 2 ) are nested models therefore to discriminate between them
likelihood ratio (LR) test has been used. LR test has been carried out to test the
following hypothesis:
1. H 0 : 1  0, 2  0, i.e., the sample is drawn from LG ( ,  ) ; against the
alternative H1 : 1  0, 2  0, i.e., the sample is drawn from
TPSL(  ,  , 1 , 2 ) .
2. H 0 : 1  0, 2  0, i.e., the sample is drawn from SLG ( ,  ,  ) ; against
the alternative H1 : 1  0, 2  0, i.e., the sample is drawn from
TPSL(  ,  ,  ,  ) .
1 2

3. H 0 : 1  0, 2  0, i.e., the sample is drawn from ASLG(,  ,  ) ;


against the alternative H1 : 1  0, 2  0, i.e., the sample is drawn from
TPSL(  ,  , 1 , 2 ) .

The results of the LR test statistic for different hypothesis have been obtained in
table 4.
Table 6: The values of LR test statistic for different hypothesis
LR test Degrees of Critical
Hypothesis statistic Freedom values at 5%
values
H 0 : 1  0, 2  0 vs H 0 : 1  0, 2  0 51.882 2 5.991
H 0 : 1  0, 2  0 vs H 0 : 1  0, 2  0 6.902 1 3.841
H 0 : 1  0, 2  0 vs H 0 : 1  0, 2  0 38.648 1 3.841

From the Table 6, it has been observed that, the value of LR test statistics is
greater than that of tabulated critical value for the entire three hypotheses at 5%
level of significance. Therefore it may be concluded that the sampled data comes
from the Two Parameter Skew Logistic ( TPSL(  ,  , 1 , 2 ) ) distribution.

7. Conclusion
In this article a new family of skew logistic distribution namely Two Parameter
Skew Logistic (TPSL) distribution has been introduced along with its important
distributional properties. This new distribution is nothing but an extension of skew
logistic distribution with an additional parameter showing more flexibility than that
of some existing families of skew logistic distribution. The problems in estimating
the TPSL distribution have been studied by using maximum likelihood methods. To
show the usefulness as well as flexibility of the TPSL distribution in comparison to
some other models, one real life data set has been considered where it has been

234
observed that the TPSL distribution better fits the data than the other distributions
considered in terms of the value of AIC. Therefore, TPSL distribution is the most
appropriate distribution for the specific data under consideration. Besides, LR test
has been carried out to discriminate between some nested models and it has been
found that the sampled data under consideration comes from the TPSL distribution.

References
1. Arellano-Valle, R. B., Gómez, H. W., & Quintana, F. A. (2004). A new class
of skew-normal distributions. Communications in statistics-Theory and
Methods, 33(7), 1465-1480.
2. Arnold, B. C., Beaver, R. J., Azzalini, A., Balakrishnan, N., Bhaumik, A.,
Dey, D. K., ... & Beaver, R. J. (2002). Skewed multivariate models related to
hidden truncation and/or selective reporting. Test, 11, 7-54.
3. Asgharzadeh, A., Esmaeili, L., & Nadarajah, S. (2016). Balakrishnan skew
logistic distribution. Communications in Statistics-Theory and
Methods, 45(2), 444-464.
4. Azzalini, A. (1985). A class of distributions which includes the normal
ones. Scandinavian journal of statistics, 171-178.
5. Azzalini, A., & Valle, A. D. (1996). The multivariate skew-normal
distribution. Biometrika, 83(4), 715-726.
6. Balakrishnan, N. (2002). Skewed multivariate models related to hidden
truncation and/or selective reporting-Discussion. Test, 11(1), 37-39.
7. Branco, M. D., & Dey, D. K. (2001). A general class of multivariate skew-
elliptical distributions. Journal of Multivariate Analysis, 79(1), 99-113.
8. Choudhury, K., & Matin, M. A. (2011). Extended skew generalized normal
distribution. Metron, 69, 265-278.
9. Crawford, S. L. (1994). An application of the Laplace method to finite
mixture distributions. Journal of the American Statistical
Association, 89(425), 259-267.
10. Gupta, R. D., & Kundu, D. (2010). Generalized logistic
distributions. Journal of Applied Statistical Science, 18(1), 51.
11. Henze, N. (1986). A probabilistic representation of the 'skew-normal'
distribution. Scandinavian journal of statistics, 271-275.
12. Matin, M. A. and Bagui, S. C. (2008) Small-sample performance of
estimators and tests in logistic regression with skew-normally distributed
explanatory variables, Journal of Statistical Theory and Applications, 7 (2),
151–167.

235
13. Nadarajah, S. (2009). The skew logistic distribution. AStA Advances in
Statistical Analysis, 93, 187-203.
14. Rathie, P. N., & Coutinho, M. (2011). A new skew generalized logistic
distribution and approximations to skew normal distribution. Aligarh J
Statist, 31, 1-12.
15. Rathie, P. N., & Swamee, P. K. (2006). On a new invertible generalized
logistic distribution approximation to normal distribution (No. 07).
Technical Research Report.
16. Wahed, A., & Ali, M. M. (2001). The skew-logistic distribution. J. Statist.
Res, 35(2), 71-80.

236

You might also like