MRA Project L-1
MRA Project L-1
BY
PG/18/0340
INTRODUCTION
Poisson regression is a widely used method for analysing count data. It is typically applied to
explore the relationship between a dependent count variable and one or more independent
variables. Count variables are non-negative integers and often represent phenomena such as
the number of product defects, software bugs, road accidents, monthly equipment failures,
Poisson regression model are generally estimated through the Maximum Likelihood
In numerous research disciplines, modelling count data has gained considerable importance.
While the Poisson regression model (PRM) is frequently used due to its simplicity,
alternative models are also employed, especially when the data exhibit over-dispersion or
under-dispersion. These include the negative binomial (NB) model, the bell model, and the
Conway–Maxwell–Poisson model. A major limitation of the PRM is its assumption that the
mean and variance of the response variable are equal, a condition not often met in practice.
When count data display greater variability than expected under the Poisson distribution, the
negative binomial model becomes a preferable alternative due to its flexibility (Abonazel et
al., 2023).
A defining feature of sample survey theory is its reliance on auxiliary information to enhance
the precision of estimators. Auxiliary variables are instrumental in several key survey design
estimate the population mean or total. When there is a strong positive correlation between the
1
auxiliary and study variables, ratio and regression estimators prove to be effective tools. In
However, the presence of extreme values or outliers can reduce the effectiveness of these
Poisson regression is a fundamental statistical method used to model count data, which
frequently arises in disciplines such as healthcare, economics, and the social sciences. This
modelling approach is based on the assumption that the response variable follows a Poisson
distribution, characterized by equal mean and variance. However, real-world count data often
deviate from this assumption, displaying either overdispersion—where the variance exceeds
the mean—or under dispersion—where the variance is lower than the mean.
Accurately estimating the mean of the response variable is a central objective in Poisson
regression analysis. Traditionally, parameter estimation is carried out using the maximum
likelihood estimator (MLE), which has been widely adopted due to its theoretical properties.
Nevertheless, the MLE can be sensitive to certain data issues, such as the presence of extreme
values, over-dispersed data, or small sample sizes, which can compromise its reliability.
estimates. Among these are robust regression methods and modified versions of MLE, which
are specifically designed to mitigate the influence of outliers and accommodate data
dispersion more effectively. These innovative approaches seek to enhance the overall
When dealing with count data and large means, linear regression models become less
practical despite the fact that the Poisson distribution approximates normality. This is
2
primarily because linear regression can yield negative predictions, which are not meaningful
for count data. Moreover, hypothesis testing in linear models assumes constant variance of
the dependent variable—an assumption that does not hold for count data. Consequently,
Poisson regression is the most appropriate and frequently utilized method in the applied
Estimating the mean in Poisson regression models is a fundamental task in statistical analysis,
especially in domains such as healthcare, economics, and the social sciences. However,
standard mean estimators derived from Poisson regression often face significant challenges,
data are over-dispersed, contain outliers, or are based on small samples. These weaknesses
can compromise the reliability and accuracy of estimates, thereby affecting policy
In most survey-based research, estimating the population mean or total is a primary objective.
the precision of mean estimators (e.g., Upadhyaya et al., 1985; Upadhyaya & Singh, 1999;
Shahzad, 2016; Shahzad et al., 2019, 2021; Rather et al., 2022; Bhushan et al., 2023;
Koçyigit & Rather, 2023; Bhushan & Kumar, 2024). When there is a strong correlation
between the auxiliary variable and the study variable, ratio-type and regression-based
estimators offer effective solutions. However, traditional methods tend to lose efficiency
To counter the influence of such anomalies, Kadilar et al. (2007) introduced a robust
regression technique using the Huber-M approach. Oral and Kadilar (2011a, 2011b) further
extended this by enhancing mean estimation using the basic simple random sampling (SRS)
3
method and refining the maximum likelihood approach. Abid et al. (2016) contributed by
applying non-traditional location parameters, while Koç et al. (2022) developed regression-
Modelling count data using linear regression becomes problematic when the average count is
high, or the Poisson distribution approximates a normal one. Linear models, which link
unsuitable for non-negative count data. Cameron and Trivedi (1998) conducted a
comprehensive study on count data regression, reaffirming that regression models remain the
Despite these various advancements, current Poisson regression-based estimators still face
irregularities. This study aims to address these issues by introducing a novel estimator
designed to improve both the robustness and efficiency of mean estimation in Poisson
regression contexts. The proposed approach seeks to offer a more reliable solution for count
data analysis, ultimately supporting enhanced policy decisions and outcomes across diverse
application areas.
1.3 MOTIVATION
In survey sampling and statistical estimation, achieving precise and reliable estimates of
population parameters is a central goal—particularly when analyzing count data. The Poisson
regression model has long been a standard tool for modelling such data due to its ability to
handle non-negative integer responses, which are common in fields like public health,
irregularities such as over-dispersion, outliers, and small sample sizes often limit the
When strata are properly constructed based on auxiliary information, stratified sampling leads
to reduced sampling variance and better representation of subgroups within the population.
Despite this advantage, there remains a gap in literature for robust estimators that effectively
combine Poisson regression with stratified sampling design, especially in the presence of
This study is driven by the need to address the inefficiencies and biases that arise in mean
estimation when standard Poisson regression methods are applied without accounting for the
under the Poisson model can be highly sensitive to outliers and dispersion, leading to
unreliable estimates.
sampling schemes.
methods and stratified sampling have each been studied individually, their integration
iv. Need for Efficient Estimation Tools in Applied Fields: Accurate mean estimation is
5
Hence, this research proposes a novel estimator that merges the strength of Poisson
incorporating robustness to make the estimator reliable under practical conditions. This
integrated approach aims to contribute both methodologically and empirically to the field
of statistical estimation.
The main objective of this study is to propose a robust unbiased Estimator for Poisson
ii. derive the mean square error (MSE) of the proposed estimator
iii. make algebraic comparisons of the MSE of the proposed estimator to that of existing
estimators
iv. check the performance of the proposed estimator using real data sets.
This research is highly relevant due to several key contributions it makes to the field of
statistical analysis. By introducing a novel estimator for mean estimation within the
framework of Poisson regression, the study adds to the ongoing development of advanced
statistical techniques. Specifically, it addresses the challenges associated with count data,
aiming to enhance estimation performance by minimizing both bias and mean square error.
Such improvements can have practical benefits, equipping researchers and analysts with more
6
Poisson regression plays a central role in various disciplines, including public health,
economics, social sciences, and finance. The outcomes of this study will be particularly
useful for professionals and researchers in these areas, as the improved estimator provides a
more accurate approach to calculating population means. In turn, this helps strengthen the
Moreover, the development and assessment of the proposed estimator will contribute to the
foundation for future research in statistical estimation. The comparative analysis conducted as
part of this study offers valuable insights into the performance differences among estimators,
shedding light on their advantages and shortcomings under different data conditions. These
findings can serve as a starting point for further refinement and exploration of estimation
This research extends the existing knowledge base surrounding Poisson regression and
relative to existing alternatives, offering a clearer understanding of when and how each
sensitivity to outliers and inefficiency in small samples—this study fills an important gap in
statistical methodology. The new estimator is designed to offer greater robustness and
The creation of a more effective estimator not only enhances the accuracy of analyses using
Poisson regression but also reinforces the overall reliability of empirical studies. This has
direct implications for domains where precise estimation is essential for developing policies,
planning interventions, and allocating resources. Although the study centers on Poisson
7
regression, the principles and techniques it introduces may be transferable to other forms of
regression modelling and estimation tasks, broadening its impact across disciplines.
In addition, the study provides empirical evidence on how different estimators perform under
varied conditions, helping practitioners choose the most appropriate method for their research
or operational needs. This empirical approach supports the adoption of best practices in
Finally, the methodological advancements proposed in this work have potential implications
for computational tools as well. The estimator and techniques developed may lead to the
more efficient and accessible statistical analysis across a range of professional fields.
1. Population: is the entire set of an object, individuals or units that have unique
2. Sample: is the fractional part of a population in which inferences related to the whole
“sample”.
5. Sampling frame: is a list of all sampling units from which a sample can be drawn.
population.
11. Estimator: is a rule that can be used to calculate an estimate of a given quantity based
12. Bias: is the mathematical difference an estimator’s expected value and the actual
value and the actual value of the parameter to be estimated. For a given population
parameter θ with statistic θ , which is an estimator based on any sample data x, the
Bias (θ ¿ = E (θ^ ) - θ
13. Mean Square Error (MSE): is the measure of the average squared difference
between the predicted value and the actual value of the quantity or variable being
estimated.
variability of a given dataset. Its expression is in form of the ratio of the standard
σ
CV = x 100%
μ
probability of a given number of events occurring in a fixed interval of time if these events
occur with a known constant mean rate and independently of the time since the last event.
9
CHAPTER TWO
LITERATURE REVIEW
2.1 INTRODUCTION
Sampling has been a very powerful technique playing a significant role across fields enabling
us to make informed decisions quickly and with less resources/efforts. Many researchers have
parameter estimation. For instance, Zaman (2019) and Zaman & Bulut (2020) introduced a
class of robust ratio-type estimators derived from a combination of existing ratio estimators.
Earlier, Raghav (2011) laid the groundwork for these developments by exploring robust
regression in constructing ratio estimators, replacing the conventional least squares approach.
Similarly, Ali et al. (2021) introduced a set of estimators utilizing robust regression strategies
in the context of sensitive surveys under simple random sampling. Tailor (2009) introduced a
improved efficiency over conventional estimators. Other studies have recommended down-
weighting data points with large residuals in order to reduce their influence and improve
estimator efficiency.
Comparative studies have also been conducted to evaluate the performance of different
population mean under simple random sampling without replacement (SRSWOR). Their
study derived expressions for the bias and mean square error (MSE) of these estimators to the
first order of approximation and presented theoretical comparisons with existing methods,
identifying conditions under which the proposed estimators outperformed traditional ones.
Abonazel (2019) investigated the use of Poisson regression in datasets containing outliers. A
Monte Carlo simulation study compared traditional maximum likelihood estimation with
robust alternatives like Mallows quasi-likelihood and weighted MLE. The results
demonstrated that robust estimators, particularly the weighted MLE, showed better
performance in the presence of outliers. More recently, Ahmed et al. (2024) evaluated three
robust methods—M, S, and MM estimators—on panel data with outliers using COVID-19
data from twelve European countries. Their analysis revealed that classical MLE was highly
sensitive to outliers, while MM estimators produced more stable and accurate predictions.
Ertan and Akay (2023) proposed a new class of biased estimators for Poisson regression,
developed based on the preliminary test estimator (PRE) framework. The authors evaluated
their efficiency using the asymptotic matrix mean square error criterion and validated the
findings with two Monte Carlo simulation studies. Their results showed that the proposed
estimators—have been proposed for estimating the population mean in Poisson regression
settings. Bhushan et al. (2023) introduced such classes under ranked set sampling (RSS) by
11
conventional approaches, including traditional ratio and regression estimators. Bias and MSE
expressions were derived, and their superiority was confirmed through studies on artificial
While MLE generally performs well in large samples, its efficiency tends to diminish in
small-sample scenarios. In contrast, Bayesian estimators often yield better results when prior
knowledge is available. Kadilar et al. (2007) presented a robust regression method based on
the Huber-M approach, while Oral and Kadilar (2011a, 2011b) refined estimation strategies
under SRS and enhanced MLE techniques. Abid et al. (2016) investigated estimators based
estimators for mean estimation in double sampling contexts. One key challenge with applying
linear regression to count data is that it may yield negative predicted values and does not
account for the discrete nature of the Poisson distribution. Cameron and Trivedi (1998)
emphasized this issue in their foundational work on count data modelling, where they
Simulation studies have proven essential in assessing the performance of estimators under
controlled settings. These simulations help identify the most efficient estimator across
varying sample sizes and parameter conditions. For example, Abonazel et al. (2023)
regression model. Using Monte Carlo simulations and real data applications, they compared
this new estimator with ridge, Liu, and modified one-parameter Liu estimators. The proposed
12
2.2 FAMILY OF EXISING ESTIMATOR
Kadilar and Cingi (2004) introduced a set of ratio estimators for estimating the population
mean Y in simple random sampling (SRS), assuming that the auxiliary variable A is known.
Their approach was motivated by earlier work from Sisodia and Dwivedi (1981) and
Upadhyaya et al. (1985), where Y represents the study variable and A denotes an auxiliary
variable.
ȳ +b ( Ā−ā)
Ȳ KC 1= Ā, (1)
ā
ȳ +b ( Ā−ā ) (2)
Ȳ KC 2= ( Ā +C a) ,
ā+C a
ȳ +b ( Ā−ā)
Ȳ KC 3= ( Ā+ϴ ( a ) ) , (3)
ā+ ϴ(a)
ȳ + b ( Ā−ā )
Ȳ KC 4 = ( Āϴ ( a )+ Ca ), (4)
āϴ ( a )+ Ca
ȳ +b ( Ā−ā)
Ȳ KC 5= ( Ā C a +ϴ ( a ) ) . (5)
ā Ca +ϴ ( a )
Where:
Ȳ, Ā: population means associated with study and auxiliary variables (Y, A), respectively.
13
𝝧(a): population coefficient of the kurtosis associated with A.
The mean squared error (MSE) formulas for the estimators were derived as presented by
1−f 2 2 2 2 2 2
MSE (Ȳ KCi )= [ R KCi S a +2 B R KCi S a+ B S a−2 R KCi Say −2 B S ay + S y ](6)
n
Where:
N = population size; n: sample size.
F = sample ratio, f = n/N.
S2y, S2a = population variance corresponding with Y and A, respectively.
Say = population covariance between X and Y, respectively.
𝜌xy = correlation coefficient between X and Y.
Ȳ Ȳ Ȳ Ȳ ϴ (a) Ȳ Ca
R KC 1= ,R = , R KC 3= , R KC 4= , R KC 5 = (7)
Ā KC 2 Ā+C a Ā+ϴ ( a ) Āϴ ( a )+C a Ā C a +ϴ ( a )
Usman Shahzad et al (2021) proposed a Poisson regression model under simple random sampling by
utilizing Poisson regression-based mean estimator and discovers its associated formula of the mean
y U = y+ b p ( X−x ) (8)
1−f 2 2 2 2 2
MSE ( y U )= [Y C y −2 B p Y X ρ xy C x C y +B p X C x ] (9)
n
14
15
2.2.3 Koҫ (2021) family of Estimator.
Recently, Koҫ (2021) proposed enhancements to the previously mentioned estimators by
introducing new ratio estimators derived from Poisson regression, as detailed below:
ȳ + b p ( Ā−ā)
ȳ H 1= Ā ,10
ā
ȳ + b p ( Ā−ā)
ȳ H 2=
ā+C a
[ Ā +C a ] , 11
ȳ +b p ( Ā−ā )
ȳ H 3= [ Ā +ϴ ( a ) ], 12
ā +ϴ ( a )
ȳ +b p ( Ā−ā)
ȳ H 4= [ Ā ϴ ( a )+C a ], 13
ā ϴ ( a ) +C a
ȳ +b p ( Ā−ā )
ȳ H 5= [ Ā C a+ϴ ( a ) ],14
ā C a +ϴ ( a )
Here, bp denotes the slope coefficients estimated through Poisson regression, with E(bp)=Bp.
1−f 2 2 2 2 2 2
MSE (Ȳ Hi)= [R KCi Sa +2 B p R KCi S a +B p S a−2 R KCi S ay −2 B p S ay +S y ]15
n
Furthermore, Koҫ (2021) showed that his estimators outperform those of Kadilar and Cingi
(2004) in efficiency, provided that at least one of the following conditions holds true:
B p−B<−2 R KCi
B p−B>−2 R KCi 16
estimator with simple random sampling and found its bias and mean square error formula as
presented below;
{( )}
δ
) (
−α G
a G+ β g
K pi = y α G+ β
+b p 1− (17)
a g+ β G
16
Bias ( K Pi )=
1−f
n [
Y ( w 2i + B p )
δ ( δ−1 ) 2
2
( wi + B p ) C 2g−δ C gy
] (18)
1−f 2 2 2 2
MSE ( K Pi )=
n
[
S y + R δ ( w i + B p ) S g−2 Rδ ( wi +B p ) S gy
2 2
] (19)
Y
They obtained the optimum value of δ by differentiating the equation wrt to δ, where δ= to
G
have
1−f 2 2
MSE ( K Pi )= S y (1−ρ ) (20)
n
Where Y is the research variable and G is an auxiliary variable, δ is a constant, α and β are
the coefficients that can take the values 0, 1, S g, C g, β 2(x) and the coefficient of correlation
where bp is the slope coefficient that can be obtained using a Poisson regression model and
E(bp)=Bp.
17
CHAPTER THREE
METHODOLOGY
3.0 INTRODUCTION
In this study, stratified random sampling was adopted, a new proposed Estimator was
derived, the MSE as well as the optimum value of the Estimator are obtained. Also, the newly
proposed estimators were applied to real life data sets and compare with existing estimators.
Stratified random sampling is a widely used probability sampling technique. Unlike the
traditional Simple Random Sampling (SRS), where some elements (or values) are selected
that a sample accurately represents the entire population by dividing it into smaller sub-
groups, known as strata, based on shared (similar) characteristics or attributes, such as age,
gender, income, education level, among others. After the strata are formed, random sample
are drawn from each sub-group. It is particularly useful in scenarios where the population is
divers, and simple random sampling might fail to capture all unique features (i.e., not a
representative of the population), reducing sampling bias and enhancing the precision of the
results.
Mainly, there are two (2) types of Stratified Random Sampling namely:
18
19
- Proportionated Stratified Random Sampling
It is a form of Stratified Random Sampling where the number of samples drawn from each
stratum is directly proportional to the size of that stratum relative to the entire population. In
other words, the proportion of the sample taken from each stratum reflects its actual share in
the population.
In disproportionate stratified random sampling, the sample size for each stratum is not
proportional to the stratum's size in the population. This means that a stratum that is
considered more important for the analysis may be over-sampled, while a stratum that is less
This type of stratified random sampling is most used when the strata are heterogeneous in
size or when some strata are considered more important than others. It can be more efficient
than proportionate stratified random sampling, but it may not be as representative of the
entire population.
events occurring within a specified time and follows a Poisson distribution defined by:
− µi yi
e µi
P ( y i ; µi ) = !
, µ>0 (21)
yi
Let A be the auxiliary variable matrix of size n x (k + 1). The relationship between Y i and ith
T
¿ ( µi )=ξ=ai W (23)
Here W = (W0, W1, …, Wk) denotes the vector of regression coefficients. This formulation is
commonly known as the Poisson regression model. The vector W represents the maximum
W.
∑ ¿¿ (24)
i=1
Iterative methods, such as the Fisher scoring and Newton-Raphson algorithms, are used to
solve these k equations (refer to Koç (2004) and Montgomery et al. (2006)).
The newly developed estimator for the population mean can be structured within the
frameworks of Zaman and Bulut (2020) and Ali et al. (2021). However, in this study, we
{( )}
δ
) (
−α X
a X +β x (25)
K st = y st α X+β
+b p 1− st
a x st + β X
y st −Y x st −X
ε ∘= , ε 1=
Y X
E (ε o) = E (ε 1) = 0
E ( ε o ) =∑ w h f 1 c y , E ( ε 1 ) =∑ w h f 1 c x , E ( ε o ε 1 )=∑ wh f 1 c y c x p y x
2 2 2 2
h h h h h h
21
Expressing equation (25) in terms of ε and neglecting the higher order, such that
αX
w=
α X+β
{( )}
δ
) (
−α X
α x st α X+β X −x st
K pst = y st +b p (26)
α X +β X
{( )}
δ
) (
−w
α x st + β+(α X−α X ) x st −X
¿ Y ( 1+ε o ) +b p
α X+β X
{( )}
δ
) (
−w
(α X+ β)+(α x st −α X ) x st −X
¿ Y ( 1+ε o ) −b p
α X +β X
{( }
δ
)
−w
α X + β α (x st − X )
¿ Y ( 1+ε o ) + −b p ε 1
α X +β α X +β
{[ ( )] }
−w δ
α (x st − X) X
¿ Y ( 1+ε o ) 1+ −b p ε 1
α X+ β X
{[ )] }
−w δ
)(
x st −X
¿ Y ( 1+ε o ) 1+
αX
α X +β ( X
−b p ε 1 (27)
¿ Y ( 1+ε o ) {( 1+w ε 1 ) −b p ε 1 }
−w δ
(28)
{( }
δ
K pst =Y ( 1+ ε o ) 1−w ε 1 +w
(w1 +1) 2
2
2
ε 1−b p ε 1 3
1 ) (29)
{ }
δ
2 3 (w+1) 2
K Pst =Y ( 1+ε o ) 1−w ε 1−b p ε 1 +w ε1
2
{ }
δ
2 3 ( w+1) 2
K Pst =Y ( 1+ε o ) 1−(w −b p )ε 1 +w ε1 (30)
2
22
[ ( ) ( )
]
3 2
w3 ( w+1 ) 2 δ ( δ−1 ) w ( w+1 ) ε 1
¿ 1+δ −( w +b p ) ε 1 + −( w +b p ) ε 1+
2 2
ε1 +
2 2 2
K pst =Y ( 1+ ε o )
[ ]
3
+ δ ( δ−1 ) ( δ−2 ) w 2 ( w+1 ) 2
− ( w + b p ) ε 1+
2
ε 1 + ...
6 2
[ ]
3 6 2
δ w ( w+1 ) 2 δ ( δ −1 ) 2 2 w ( w+ 1 ) 4
2
¿ Y ( 1+ε 0 ) ¿1−δ w ε i−δ b p ε 1 +
2
ε 1+
2
( w +b p ) ε 2i + ¿∧w3 ( w+1 ) ( w2 +b p) ε 31 +
2
ε 1 + ...
[ ]
3
δ w ( w+1 ) 2 δ ( δ−1 ) 2 2
¿ Y ( 1+ε 0 ) 1−δ ( w2 +b p ) ε 1 + ε 1+ ( w +b p ) ε 21
2 2
[
¿ Y ( 1+ε 0 ) 1−δ ( w2 +b p ) ε 1 +
δ ( δ−1 ) 2
2
(
2
w + b p ) ε 21
]
( )
δ ( δ−1 ) 2 2 2
¿ 1−δ ( w +b p ) ε 1 +
( w +b p ) ε 1 +ε 0
2
¿Y 2
δ ( δ−1 ) 2 2
¿−δ ( w +b p ) ε 0 ε i + ( w + b p ) ε0 . ε1
2 2
2
( )
δ ( δ−1 ) 2 2
¿ 1−δ ( w2 +b p ) ε 1 + ( w +b p ) ε 21 +ε 0
¿Y 2
¿−δ ( w +b p ) ε 0 ε i
2
(
K Pst =Y ¿ 1−δ ( w +b p) ε 1 +
2 δ ( δ−1 ) 2
2
(
2 2
w +b p ) ε 1 +ε o−¿∧δ ( w +b p ) ε 0 ε i
2
)
(31)
( )
δ ( δ−1 ) 2 2
¿−δ ( w2 +b p ) ε i + ( w +b p ) ε 21 +ε o
K Pst =Y −Y 2
¿−δ ( w + b p ) ε 0 ε i
2
( )
δ ( δ−1 ) 2 2
K Pst −Y =Y
¿
2
( w +b p ) ε 21 +ε 0 −δ ( w2 +b p ) ε i
(32)
¿−δ ( w2 +b p ) ε 0 ε i
E ( K Pst −Y ) =Bias
23
Bias ( K Pst )=Y
[ δ ( δ−1 ) 2
2
(
2
w + b p ) E ( ε 2i ) + E ( ε 0 )−δ ( w 2+ b p ) E ( ε i )−δ ( w 2+ b p ) E ( ε 0 ε i )
]
Y ( δ ( δ−1 ) 2
2
( w +b p ) ∑ W h f i c2x + 0−δ ( w 2+ b p ) ( 0 )−δ ( w 2+b p ) ∑ W h f i c y c x ρ y x
2
h h h h h )
¿Y ( δ ( δ−1
2
)
(w + b ) ∑ W f c −δ ( w +b ) ∑ W f c
2
p
2
h i
2
xh
2
p h i yh cx ρ y x
h h h )
Bias ( K Pst )=∑ W h f i Y ( w 2+b p ) [ δ ( δ−1 ) 2
2
( w + b p ) c 2x h−δ c y c x ρ y x h h h h ] (33)
To obtain the mean square error (MSE), we square both sides of equation (32) and take the
expectation.
2
( K Pst − y ) = Y [{ δ ( δ−1 )
2
(
2
}
w2 +b p ) ε 21 +ε 0−δ ( w2 +b p ) ε 1−δ ( w2 +b p ) ε 0 ε 1 ]
[
¿ ε o−δ ( w +b p ) ε 1 ε o +δ ( w +b p ) ε 1−δ ( w +b p ) ε 0 ε 1
2 2 2 2 2 2 2
] (34)
2 2 2 2
[
( K Pst − y ) =Y {δ ( w +b p ) ε 1−2 δ ( w +b p ) ε 0 ε 1 +ε o }
2 2 2 2
]
Taking expectation
[ {
E ( K Pst − y ) =E Y δ ( w + b p ) ε 1 −2 δ ( w +b p ) ε 0 ε 1 +ε o
2 2 2 2 2 2 2 2
}] (35)
2
MSE ( K Pst )=Y ¿ (36)
Recall that
2 2
2
sy 2
sx ρ s y h sx h ρ s y h x h
c =
yh
h
,c =xh
h
, c y x =ρ c y c x = = (37)
Y
2
X
2 h h h h
YX YX
{∑ }
2 2
Y Y
2∑ ∑ W h f i ρ s y sx
2
W f s + δ ( w +b p ) W h f i s x −2 δ ( w +b p )
2 2 2 2 2
MSE ( K Pst )= h i yh
X h
YX h h
24
¿ {∑ W h f i s y + R δ ( w +b p )
2
h
2 2 2 2
∑ W h f i s 2x −2 R δ ( w2 +b p )∑ W h f i ρ s y s x }
h h h
(38)
{
MSE ( K Pst )=∑ W h f i s 2y + R2 δ 2 ( w 2+ b p ) s 2x −2 R δ ( w 2+ b p ) ρ s y s x
h
2
h h h
} (39)
d ( MSE(K Pst ) ) d
dδ
=
dδ [∑ W f {s h i
2
yh
2
+ R2 δ 2 ( w2 +b p ) s 2x −2 R δ ( w2 +b p ) ρ s y s x h h h
}] (40)
{
= ∑ W h f i 0+ 2 Rδ ( w +b p ) s x −2 R ( w + b p ) ρ s y s x
2 2 2 2
h h h
}
¿>2 Rδ ( w +b p ) ∑ W h f i s 2x =2 R (w 2+ b p ) ∑ W h f i ρ s y s x
2 2
h h h
2 R ( w 2+ b p ) ∑ W h f i ρ s y s x ∑ W h f i ρ s y sx
δ opt = h h
= h h
2 R (w + b p )
2 2 2
∑W f s
2
h i xh
R(w 2+ b p ) ∑ W h f i s 2x h
(41)
Substitute δ opt into eqn (15) to get the optimum value of the MSE.
{ 2 R ∑ W h f i ρ s y sx
}
2
(W h f i ρ sy s x )
MSE ( K Pst )=∑ W h f i s + R
2 2
2 2
. ( w
2
+b ) s
h
−
h
. ( w +b p ) ρ s y s x
2 h h
R ( w +b p ) ∑ W h f i s x
y p x
R ( w + b p ) (∑ W h f i s x )
2 2 2 2 2 2 2 h h h
h h
{ (∑ W h f i ρ sy sx )
}
2
2 R ∑ W h f i ρ s y sx
∑ Wh f i sy+R ∑
2 2
¿
2
. ( w
2
+ b ) h
W f
h
s
2
− .( w + bp ) ρ s y sx
2 h h
R ( w + b p ) ∑ W h f i sx
2 p h i x
R ( w +b p ) (∑ W h f i s x )
2 2 2 2 2 2 h h h
h h
{ (∑ W h f i ρ s y s x )
}
2
2( ∑ W h f i ρ s y s x )
2
∑ Wh f i sy+
2
¿ −
h h h h
(∑ W h f i s 2x )
2
h
∑ W h f i sx 2
h
( ∑ W h f i ρ sy sx )
2
( ∑ W h f i ) ρ2 s2yh s 2xh
2
¿ ∑ W h f i s y− =∑ W h f i s y −
2 2 h h
(∑ W h f i s 2x )
2
∑ W h f i s 2x h
h
¿ ∑ W h f i s y (1−ρ )
2 2
(43)
25
CHAPTER FOUR
RESULT AND DISCUSSION
4.1 Introduction
The chapter comprises the computation and analysis of the data used in the validation of the
proposed Estimators which was obtained using Microsoft Excel and R statistical software.
Numerical analysis was conducted with the goal of evaluating the relative performance of the
proposed class of Estimators and family, as compared to the preliminary Estimators. The
evaluation was done empirically by utilizing two real population data sets and simulated data,
(¿ y )
PRE=var x 100(44)¿
MSE ( y )
4.3 Applications
In this study, We tried to validate our theoretical findings using two real count
dataset and a simulated dataset, the first dataset is the enrolment of the unified
26
4.4 Descriptive Statistics (Data 1: Schools from 3 LGAs)
Sample
Mean 49.3494
2.27889
Standard Error 5
Median 37
Mode 32
Standard 35.9603
Deviation 6
1293.14
Sample Variance 8
2.22253
Kurtosis 3
1.51261
Skewness 4
Range 199
Minimum 1
Maximum 200
Sum 12288
Count 249
27
Count of sample by States
100
90
80
70
60
50
40
30
20
10
0
State1 State2 State3
Figure 4.3.2 represents the distribution of the sample variable by states, showing the total
Distribution of LGAs
28% State1
36%
State2
State3
37%
Figure 4.3.3 illustrates the pie chart which represents the distribution of states and the
28
Population
Mean 1346.068
Standard Error 83.82328
Median 895
Mode 1756
Standard
Deviation 1322.709
Sample Variance 1749559
Kurtosis 8.378347
Skewness 2.291073
Range 9766
Minimum 118
Maximum 9884
Sum 335171
Count 249
Figure 4.3.4 is a histogram plot showing the distribution of the population variable
29
Sum of population by States
State1
28% 36% State2
State3
37%
Figure 4.3.5 illustrates the pie chart which represents the distribution of states and the
90
80
70
60
50
40
30
20
10
0
State1 State2 State3
Figure 4.3.6 represents the distribution of the population variable by states, showing the total
30
Figure 4.3.7 Box Plot of Population variable
31
250
200
150
100
50
0
0 2000 4000 6000 8000 10000 12000
Population
1304.49
Mean 8
90.2700
Standard Error 9
Median 890
Mode 400
Standard 1292.47
Deviation 1
Sample
Variance 1670481
10.6781
Kurtosis 7
2.54236
Skewness 2
Range 9766
Minimum 118
Maximum 9884
Sum 267422
Count 205
32
State3
State2
Total
State1
0 10 20 30 40 50 60 70 80
33
28%
36%
State1
State2
State3
37%
Sample
47.7073
Mean 2
2.41794
Standard Error 3
Median 36
Mode 26
Standard 34.6196
Deviation 7
Sample 1198.52
Variance 2
1.87502
Kurtosis 3
1.46606
Skewness 3
Range 170
Minimum 1
Maximum 171
Sum 9780
Count 205
34
State3
State2
State1
0 10 20 30 40 50 60 70 80
35
28%
36%
State1
State2
State3
37%
4.4.2 Computation of the MSE’s and PRE’s of the proposed Estimators using data I
Table # presents the MSE and PRE of the family of the proposed Estimator using
dataset I in comparison with the classical mean of stratified random sampling. Also,
Table # shows the performance of the proposed Estimators in comparison with the
registration for the years 2017 and 2018. The analysis shows the number of Applicants and
APP17
Mean 46538.32432
Median 42095
Mode #N/A
Kurtosis -1.184332828
Skewness 0.295468423
Range 95918
Minimum 6270
Maximum 102188
Sum 1721918
Count 37
Confidence Level
(95.0%) 8883.413532
Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“APP17” represents the number of JAMB Applicants in 2017. The result of the analysis
shows that:
37
1. The average number of students who applied for JAMB in 2017 are about 46,538
across all states (37, the Federal Capital inclusive) within the country (Nigeria).
2. Imo State appears to have the highest number of applicants for the year (2017) having
102,188 applicants, followed by Osun, Oyo, Ogun and Delta with 88,803, 87,827,
81,536 and 81,478 applicants respectively. This result shows that fairly the Ibos and
3. The top 5 states with the lowest number of applicants for 2017 are as follows;
Sokoto, Kebbi, Yobe, Zamfara and the FCT with 14,478, 13,927, 13,767, 10,522, and
6,270. This result indicates that the Northerns (or Hausas) are not interested in
education as the other 2 major ethnic groups in Nigeria as they fully dominant this
frame.
4. The total (sum) number of all applicants for the year 2017 is 1,721,918.
ADM17
Mean 15313.51351
Median 13665
Mode #N/A
Kurtosis -0.913318656
Skewness 0.212030526
Range 28415
Minimum 2120
Maximum 30535
38
Sum 566600
Count 37
Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“ADM17” represents the number of students that were admitted into tertiary institutions after
the examination (JAMB) in 2017. The result of the analysis shows that:
1. The average number of students who were admitted after the exam in 2017 are about
15,313 across all states (37, the Federal Capital inclusive) within the country
(Nigeria).
2. Imo State again appears to have the highest number of admitted students for the year
(2017) followed by Osun, Oyo, Ogun and Kano with 30535, 28209, 27,989, 26,589
3. The top 5 states with the lowest number of admitted students for 2017 are as follows;
Jigawa, Kebbi, Sokoto, Zamfara and the FCT with 5,776, 5,104, 4,004, 2,744, and
2,120.
4. The total (sum) number of admitted students for the year 2017 is 566,600.
APP18
Mean 44671.40541
Median 42560
Mode #N/A
Kurtosis -1.16341722
Skewness 0.306237264
Range 86610
Minimum 6438
Maximum 93048
Sum 1652842
Count 37
Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“APP18” represents the number of JAMB Applicants in 2018. The result of the analysis
shows that:
1. The average number of students who applied for JAMB in 2018 are about 44,671
across all states (37, the Federal Capital inclusive) within the country (Nigeria).
2. The same states from 2017 seems to appear as the top 5 States with the highest
number of applicants for 2018 (though their positions got changed). Imo, Oyo, Osun,
Ogun and Delta with 93,048, 86,687, 86,065 and 80,453 applicants respectively.
3. The same states from 2017 seems to appear as the top 5 States with the lowest number
of applicants for 2018 (though their positions got changed), they are; Yobe, Kebbi,
Sokoto, Zamfara and the FCT with 15,536, 15,341, 13,493, 10,090, and 6,438. This
result indicates that the Northerns (or Hausas) are not interested in education as the
other 2 major ethnic groups in Nigeria as they fully dominant this frame.
4. The total (sum) number of all applicants for the year 2018 is 1,652,842.
40
Table 5: Statistics of Admitted Students (2018)
ADM18
Mean 14855.43243
Median 14074
Mode #N/A
Kurtosis -0.749533705
Skewness 0.25991206
Range 27104
Minimum 2279
Maximum 29383
Sum 549651
Count 37
the examination (JAMB) in 2018. The result of the analysis shows that:
1. the average number of students who were admitted after the exam in 2018 are about
14855 across all states (37, the Federal Capital inclusive) within the country (Nigeria).
2. The states from 2017 seems to appear as the top 5 States with the highest number of
admitted students for 2018 expect Kano which was replaced by Anambra.
Imo, Oyo, Osun, Ogun and Anambra with 29,383, 28,743, 28,085, 25,804 and 25,369
applicants respectively.
41
3. The same states from 2017 seems to appear as the top 5 States with the lowest number
of applicants for 2018 (there positions are not changed), they are;
Jigawa, Kebbi, Sokoto, Zamfara and the FCT with 5,616, 4,494, 3,860, 2,596, and
2,279.
4. The total (sum) number of admitted students for the year 2018 is 549,651.
4.4.1 Charts of Number of Applicants and Admitted Students for the years 2017 &
2018
Figures 1-4 show the bar charts of Unified Tertiary Matriculation Examination registration
for the years 2017 and 2018. The charts show the number of Applicants and number of
Admitted Students in each state for the years 2017 and 2018.
ZAMFARA
TARABA
RIVERS
OYO
ONDO
NIGER
LAGOS
KOGI
KATSINA
KADUNA
IMO
FCT
EKITI
EBONYI
CROSS RIVER
BENUE
BAUCHI
AKWA IBOM
ABIA
0 20000 40000 60000 80000 100000 120000
42
Figure 1: No. of Applicants by State (2017)
ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 5000 10000 15000 20000 25000 30000 35000
43
ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 10000 20000 30000 40000 50000 60000 70000 80000 90000 100000
44
ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 5000 10000 15000 20000 25000 30000 35000
4.4.2 Computation of the MSE’s and PRE’s of the proposed Estimators using data II
Table # presents the MSE and PRE of the family of the proposed Estimator using dataset II in
comparison with the classical mean of stratified random sampling. Also, Table # shows the
performance of the proposed Estimators in comparison with the existing Estimators using
45
ESTIMATOR MSE BIAS PRE RANK
Classical Mean Estimator 7.18E+12 0
-
Yashpal Estimator 9.36154E+11 27.1224
KOC MSE 8.01E+06
Kadila & Cingi 3.76E+11
N1 N2 N3
W 1= =0.35 , W 2= =0.37 , W 3 = =0.28
N N N
[ ]
Y rat =∑ W h y h
Xh
xh
Y rat =1345.84
MSE (Y ¿¿ rat)=89841.145 ¿
………….
……………
47
CHAPTER FIVE
SUMMARY, CONCLUSION AND RECOMMENDATION
5.1 SUMMARY
5.2 CONCLUSION
5.3 RECOMMENDATION
…………
48
REFERENCES
T. Zaman and H. Bulut, “Modified regression estimators using robust regression methods and
covariance matrices in stratified random sampling,” Communications in Statistics- Theory
and Methods, vol. 49, no. 14, pp. 3407–3420, 2020.
N. Ali, I. Ahmad, M. Hanif, and U. Shahzad, “Robust-regression type estimators for
improving mean estimation of sensitive variables by using auxiliary information,”
Communications in Statistics: Theory and Methods, vol. 50, no. 4, pp. 979–992, 2021.
L. N. Upadhyaya, H. P. Singh, and J. W. E. Vos, “On the estimation of population means and
ratios using supplementary information,” Statistica Neerlandica, vol. 39, no. 3, pp. 309–318,
1985.
D. C. Montgomery, E. A. Peck, and G. G. Vining, Introduction to Linear Regression
Analysis, John Wiley and Sons, Hoboken, 4th edition, 2006.
S. Bhushan,and [Link], “Improved estimation of population mean in simple random
sampling using attribute”. Thailand Statistician 22 (2), 374–389, 2024.
S. Bhushan, A. Kumar, N. Alsadat, M.S. Mustafa, and M.M. Alsolmi, “Some optimal classes
of estimators based on , Axioms 12 (6), 515, 2023
H. Koç, “Ratio-type estimators for improving mean estimation using Poisson regression
method,” Communications in Statistics - Theory and Methods, vol. 50, no. 20, pp. 4685–
4691, 2021.
E.G. Koçyigit, and K.U.I. Rather. “The new sub-regression type estimator in ranked se˘ t “,
J. Statistic. Theor. Pract. 17 (2), 27, 2023.
C. Kadilar and H. Cingi, “Ratio estimators in simple random sampling,” Applied
Mathematics and Computation, vol. 151, no. 3, pp. 893–902, 2004.
F.A. Lukman, E. Adewuyi, K. Mansson and B.M. Golam, “A new estimator for the
multicollinear Poisson regression model: simulation and application,” Scientific Report 2021
Feb 12;11:3732. doi: 10.1038/s41598-021-82582-w
B. V. S. Sisodia and V. K. Dwivedi, “A modified ratio estimator using coefficient of
variation of auxiliary variable,” Journal of the Indian Society of Agricultural Statistics, vol.
33, no. 2, pp. 13–18, 1981
M. R. Abonazel and O. M. Sabar, “A comparative study of robust estimators for poisson
regression model with outliers”, Journal of Statistics Application & Probability, vol. 9, no. 2,
pp. 279-286, 2020
M.R, Abonazel, F.A Awwad, E. Tag Eldin, B.M.G Kibria and I.G. Khattab ”Developing a
two-parameter Liu estimator for the COM–Poisson regression model: Application and
simulation”, Front. Appl. Math. Stat. 9:956963 (2023). doi: 10.3389/fams.2023.956963.
E. Oral, and C. Kadilar. “Improved ratio estimators via modified maximum likelihood”.
Pakistan J. Statistic. 27 (3), 269–282, 2011a.
E. Oral and C. Kadilar. “Robust ratio-type estimators in simple random sampling”. J. Korean
Surg. Soc. 40 (4), 457–467. [Link] 2011b.
49
U. Shahzad, S. Shahzadi, N. Afshan, N.H. Al-Noor, D.A. Alilah, M. Hanif and M.M. Anas,
“Poisson regression-based mean estimator”, Mathematical Problems in Engineering Volume
2021, Article ID 9769029, 6 pages [Link]
M. Abid, N. Abbas, H.Z. Nazir, and Z. Lin. “Enhancing the mean ratio estimators for
estimating population mean using non-conventional location parameters”. Rev. Colomb.
Estadística 39 (1), 63–79. [Link] 2016
Z. H. Wani, S.E.H. Rizvi, M.I. Jeelani and S. Mushtaq, “Modified regression estimators for
improving mean estimation – Poisson regression approach,” Pakistan Journal of Statistics and
Operation Research, vol. 18, no. 4, pp 985-994, 2022
Y.S. Raghav, A.A.H. Ahmadini, A.M. Mahnashi and K.U.I Rather, “Enhancing estimation
efficiency with proposed estimator: A comparative analysis of Poisson regression-based
mean estimators,”, Kuwait Journal of Science, 52 (2025) 100282, pp 1-8
[Link]
K.U.I. Rather, E.G. Koçyigit˘ , R. Onyango, and C. Kadilar, Improved regression in ratio type
estimators based on robust M-estimation. PLoS One 17 (12), e0278868, 2022
==========
[Link]
[Link]
Ahmed, A., et al. (2024). Enhancing estimation efficiency with proposed estimator: A
comparative analysis of Poisson regression-based mean estimators. Kuwait Journal of
Science.
Tailor, R. (2009). A modified ratio-cum-product estimator of finite population mean in
stratified random sampling. Data Science Journal, 8, 182–189.
Zakari, Y., & Muhammad, I. (2023). Modified estimator of finite population variance under
stratified random sampling. Engineering Proceedings.
Khan, S., et al. (2024). An effective and economic estimation of population mean in stratified
random sampling using a linear cost function. Heliyon.
Singh, R., & Malik, S. (2014). A new estimator for population mean using two auxiliary
variables in stratified random sampling. arXiv preprint.
Verma, H. K., Sharma, P., & Singh, R. (2014). Improved estimator of finite population mean
using auxiliary attribute in stratified random sampling. arXiv preprint.
Oyeyemi, A. S. (2020). Robust estimation in stratification sampling. International Journal of
Engineering Research & Technology.
50
Iorlaha, P. I., Nwaosu, S. C., Uba, T., & Ikughur, A. J. (2024). Regression estimator of
population mean with random missing values in stratified two-stage sampling using auxiliary
information. Journal of Statistical Sciences and Computational Intelligence.
51