0% found this document useful (0 votes)
8 views52 pages

MRA Project L-1

This dissertation presents a robust unbiased estimator for the Poisson regression model using stratified random sampling, addressing challenges such as inefficiency and bias in traditional estimators. The study aims to enhance mean estimation accuracy in count data analysis, particularly in fields like healthcare and economics, by integrating auxiliary information and improving estimator robustness. The proposed methodology seeks to fill gaps in existing literature and contribute to better decision-making through more reliable statistical tools.

Uploaded by

pornjason1234567
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views52 pages

MRA Project L-1

This dissertation presents a robust unbiased estimator for the Poisson regression model using stratified random sampling, addressing challenges such as inefficiency and bias in traditional estimators. The study aims to enhance mean estimation accuracy in count data analysis, particularly in fields like healthcare and economics, by integrating auxiliary information and improving estimator robustness. The proposed methodology seeks to fill gaps in existing literature and contribute to better decision-making through more reliable statistical tools.

Uploaded by

pornjason1234567
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

A ROBUST UNBIASED ESTIMATOR FOR POISSON REGRESSION MODEL

USING STRATIFIED RANDOM SAMPLING

BY

ADEWOLE TIMOTHY BAMIDELE

PG/18/0340

[Link] (FUNAAB, ABEOKUTA)

A DISSERTATION SUBMITTED TO THE DEPARTMENT OF STATISTICS,


COLLEGE OF PHYSICAL SCIENCES,
FEDERAL UNIVERSITY OF AGRICULTURE, ABEOKUTA, NIGERIA
IN PARTIAL FULFILMENT OF THE REQUIREMENTS FOR THE AWARD OF
MASTERS OF SCIENCE IN STATISTICS
CHAPTER ONE

INTRODUCTION

1.1 Background of the Study

Poisson regression is a widely used method for analysing count data. It is typically applied to

explore the relationship between a dependent count variable and one or more independent

variables. Count variables are non-negative integers and often represent phenomena such as

the number of product defects, software bugs, road accidents, monthly equipment failures,

occurrences of infectious diseases, or environmental pollutant levels. The parameters of the

Poisson regression model are generally estimated through the Maximum Likelihood

Estimation technique (Lukman et al., 2021).

In numerous research disciplines, modelling count data has gained considerable importance.

While the Poisson regression model (PRM) is frequently used due to its simplicity,

alternative models are also employed, especially when the data exhibit over-dispersion or

under-dispersion. These include the negative binomial (NB) model, the bell model, and the

Conway–Maxwell–Poisson model. A major limitation of the PRM is its assumption that the

mean and variance of the response variable are equal, a condition not often met in practice.

When count data display greater variability than expected under the Poisson distribution, the

negative binomial model becomes a preferable alternative due to its flexibility (Abonazel et

al., 2023).

A defining feature of sample survey theory is its reliance on auxiliary information to enhance

the precision of estimators. Auxiliary variables are instrumental in several key survey design

components, including stratification, the determination of selection probabilities, and the

construction of estimators for population parameters. A central aim in most surveys is to

estimate the population mean or total. When there is a strong positive correlation between the
1
auxiliary and study variables, ratio and regression estimators prove to be effective tools. In

sampling methodologies, characteristics of the auxiliary variable's distribution—such as the

coefficient of variation or kurtosis—are often used to refine population mean estimates.

However, the presence of extreme values or outliers can reduce the effectiveness of these

traditional estimation methods (Shahzad et al., 2021).

Poisson regression is a fundamental statistical method used to model count data, which

frequently arises in disciplines such as healthcare, economics, and the social sciences. This

modelling approach is based on the assumption that the response variable follows a Poisson

distribution, characterized by equal mean and variance. However, real-world count data often

deviate from this assumption, displaying either overdispersion—where the variance exceeds

the mean—or under dispersion—where the variance is lower than the mean.

Accurately estimating the mean of the response variable is a central objective in Poisson

regression analysis. Traditionally, parameter estimation is carried out using the maximum

likelihood estimator (MLE), which has been widely adopted due to its theoretical properties.

Nevertheless, the MLE can be sensitive to certain data issues, such as the presence of extreme

values, over-dispersed data, or small sample sizes, which can compromise its reliability.

In response to these limitations, recent research has focused on developing alternative

estimation techniques aimed at improving the robustness and efficiency of parameter

estimates. Among these are robust regression methods and modified versions of MLE, which

are specifically designed to mitigate the influence of outliers and accommodate data

dispersion more effectively. These innovative approaches seek to enhance the overall

performance of Poisson regression models in applied settings.

When dealing with count data and large means, linear regression models become less

practical despite the fact that the Poisson distribution approximates normality. This is
2
primarily because linear regression can yield negative predictions, which are not meaningful

for count data. Moreover, hypothesis testing in linear models assumes constant variance of

the dependent variable—an assumption that does not hold for count data. Consequently,

Poisson regression is the most appropriate and frequently utilized method in the applied

sciences for analysing such data (Shahzad et al., 2021).

1.2 STATEMENT OF THE PROBLEM

Estimating the mean in Poisson regression models is a fundamental task in statistical analysis,

especially in domains such as healthcare, economics, and the social sciences. However,

standard mean estimators derived from Poisson regression often face significant challenges,

including inefficiency, susceptibility to bias, and lack of robustness—particularly when the

data are over-dispersed, contain outliers, or are based on small samples. These weaknesses

can compromise the reliability and accuracy of estimates, thereby affecting policy

formulation and practical decision-making.

In most survey-based research, estimating the population mean or total is a primary objective.

Numerous researchers have explored the incorporation of auxiliary information to enhance

the precision of mean estimators (e.g., Upadhyaya et al., 1985; Upadhyaya & Singh, 1999;

Shahzad, 2016; Shahzad et al., 2019, 2021; Rather et al., 2022; Bhushan et al., 2023;

Koçyigit & Rather, 2023; Bhushan & Kumar, 2024). When there is a strong correlation

between the auxiliary variable and the study variable, ratio-type and regression-based

estimators offer effective solutions. However, traditional methods tend to lose efficiency

when extreme values or outliers are present in the data.

To counter the influence of such anomalies, Kadilar et al. (2007) introduced a robust

regression technique using the Huber-M approach. Oral and Kadilar (2011a, 2011b) further

extended this by enhancing mean estimation using the basic simple random sampling (SRS)
3
method and refining the maximum likelihood approach. Abid et al. (2016) contributed by

applying non-traditional location parameters, while Koç et al. (2022) developed regression-

ratio estimators suitable for double sampling schemes.

Modelling count data using linear regression becomes problematic when the average count is

high, or the Poisson distribution approximates a normal one. Linear models, which link

predictions to auxiliary variables, can generate negative estimates—rendering them

unsuitable for non-negative count data. Cameron and Trivedi (1998) conducted a

comprehensive study on count data regression, reaffirming that regression models remain the

primary tools for such analyses.

Despite these various advancements, current Poisson regression-based estimators still face

persistent challenges—most notably inefficiency, sensitivity to bias, and vulnerability to data

irregularities. This study aims to address these issues by introducing a novel estimator

designed to improve both the robustness and efficiency of mean estimation in Poisson

regression contexts. The proposed approach seeks to offer a more reliable solution for count

data analysis, ultimately supporting enhanced policy decisions and outcomes across diverse

application areas.

1.3 MOTIVATION

In survey sampling and statistical estimation, achieving precise and reliable estimates of

population parameters is a central goal—particularly when analyzing count data. The Poisson

regression model has long been a standard tool for modelling such data due to its ability to

handle non-negative integer responses, which are common in fields like public health,

insurance, economics, and environmental sciences. However, in real-world applications, data

irregularities such as over-dispersion, outliers, and small sample sizes often limit the

performance of traditional estimators derived from Poisson regression.


4
At the same time, Stratified Random Sampling (SRS) is widely recognized as an effective

sampling technique that improves estimator precision by leveraging population heterogeneity.

When strata are properly constructed based on auxiliary information, stratified sampling leads

to reduced sampling variance and better representation of subgroups within the population.

Despite this advantage, there remains a gap in literature for robust estimators that effectively

combine Poisson regression with stratified sampling design, especially in the presence of

anomalies like outliers or unequal stratum sizes.

This study is driven by the need to address the inefficiencies and biases that arise in mean

estimation when standard Poisson regression methods are applied without accounting for the

structure of stratified populations or the influence of data abnormalities. The motivation

stems from the following:

i. Inadequacy of Traditional Estimators: Maximum Likelihood Estimators (MLEs)

under the Poisson model can be highly sensitive to outliers and dispersion, leading to

unreliable estimates.

ii. Importance of Auxiliary Information: Incorporating auxiliary variables into estimator

design has shown significant improvements in accuracy, especially under complex

sampling schemes.

iii. Lack of Robust Poisson-Based Estimators in Stratified Contexts: While robust

methods and stratified sampling have each been studied individually, their integration

within the context of Poisson regression is limited in current literature.

iv. Need for Efficient Estimation Tools in Applied Fields: Accurate mean estimation is

crucial for data-driven decision-making, particularly in disciplines relying on count

data where resource allocation, policy design, or intervention planning depend on

robust statistical outputs.

5
Hence, this research proposes a novel estimator that merges the strength of Poisson

regression with the precision-enhancing property of stratified sampling, while also

incorporating robustness to make the estimator reliable under practical conditions. This

integrated approach aims to contribute both methodologically and empirically to the field

of statistical estimation.

1.4 AIM AND OBJECTIVES

The main objective of this study is to propose a robust unbiased Estimator for Poisson

Regression Model using Stratified Random sampling.

The specific objectives areito:

i. derive new Poisson regression-based estimator.

ii. derive the mean square error (MSE) of the proposed estimator

iii. make algebraic comparisons of the MSE of the proposed estimator to that of existing

estimators

iv. check the performance of the proposed estimator using real data sets.

1.5 SIGNIFICANCE OF STUDY

This research is highly relevant due to several key contributions it makes to the field of

statistical analysis. By introducing a novel estimator for mean estimation within the

framework of Poisson regression, the study adds to the ongoing development of advanced

statistical techniques. Specifically, it addresses the challenges associated with count data,

aiming to enhance estimation performance by minimizing both bias and mean square error.

Such improvements can have practical benefits, equipping researchers and analysts with more

reliable tools for decision-making when dealing with count-based datasets.

6
Poisson regression plays a central role in various disciplines, including public health,

economics, social sciences, and finance. The outcomes of this study will be particularly

useful for professionals and researchers in these areas, as the improved estimator provides a

more accurate approach to calculating population means. In turn, this helps strengthen the

integrity of research findings and supports sound decision-making, particularly in contexts

where reliable statistical insights are critical.

Moreover, the development and assessment of the proposed estimator will contribute to the

foundation for future research in statistical estimation. The comparative analysis conducted as

part of this study offers valuable insights into the performance differences among estimators,

shedding light on their advantages and shortcomings under different data conditions. These

findings can serve as a starting point for further refinement and exploration of estimation

techniques for count data and beyond.

This research extends the existing knowledge base surrounding Poisson regression and

statistical estimation methods. It delivers a detailed evaluation of the proposed estimator

relative to existing alternatives, offering a clearer understanding of when and how each

method performs best. By addressing known weaknesses in current estimators—such as

sensitivity to outliers and inefficiency in small samples—this study fills an important gap in

statistical methodology. The new estimator is designed to offer greater robustness and

improved efficiency, making it a valuable addition to applied statistical tools.

The creation of a more effective estimator not only enhances the accuracy of analyses using

Poisson regression but also reinforces the overall reliability of empirical studies. This has

direct implications for domains where precise estimation is essential for developing policies,

planning interventions, and allocating resources. Although the study centers on Poisson

7
regression, the principles and techniques it introduces may be transferable to other forms of

regression modelling and estimation tasks, broadening its impact across disciplines.

In addition, the study provides empirical evidence on how different estimators perform under

varied conditions, helping practitioners choose the most appropriate method for their research

or operational needs. This empirical approach supports the adoption of best practices in

statistical work and improves the credibility of results in applied research.

Finally, the methodological advancements proposed in this work have potential implications

for computational tools as well. The estimator and techniques developed may lead to the

creation of new algorithms, software applications, or estimation frameworks, facilitating

more efficient and accessible statistical analysis across a range of professional fields.

1.6 Definition of related terms

1. Population: is the entire set of an object, individuals or units that have unique

characteristics and are of interest for analysis.

2. Sample: is the fractional part of a population in which inferences related to the whole

population are being made.

3. Sampling: is the procedure or process of selecting a part of a population known as the

“sample”.

4. Sampling unit: is the unit in which the population to be sampled is divided

5. Sampling frame: is a list of all sampling units from which a sample can be drawn.

6. Variable: can be a characteristic or attribute that is being measured or observed.

7. Population Parameter: can be referred to as the measurable characteristics of a

population.

8. Statistic: can be referred to as the measurable characteristics of the sample.

9. Correlation: is the statistical relationship between two or more variables.


8
10. Estimate: is a statistical value or range of values that can be gotten from a sample for

the purpose of making inferences about the population.

11. Estimator: is a rule that can be used to calculate an estimate of a given quantity based

on observed sample data.

12. Bias: is the mathematical difference an estimator’s expected value and the actual

value and the actual value of the parameter to be estimated. For a given population

parameter θ with statistic θ , which is an estimator based on any sample data x, the

bias of this estimator can be defined as

Bias (θ ¿ = E (θ^ ) - θ

13. Mean Square Error (MSE): is the measure of the average squared difference

between the predicted value and the actual value of the quantity or variable being

estimated.

MSE (θ^ ) = E (θ^ - θ )2 = Var (θ ) + (Bias (θ^ ¿)2

14. Coefficient of Variation (CV): is a measure of variation, it shows the extent of

variability of a given dataset. Its expression is in form of the ratio of the standard

deviation (σ ) to the mean ( μ ¿

σ
CV = x 100%
μ

15. Poisson Distribution: is a discrete probability distribution that expresses the

probability of a given number of events occurring in a fixed interval of time if these events

occur with a known constant mean rate and independently of the time since the last event.

9
CHAPTER TWO
LITERATURE REVIEW

2.1 INTRODUCTION

Sampling has been a very powerful technique playing a significant role across fields enabling

us to make informed decisions quickly and with less resources/efforts. Many researchers have

established and amended estimators in sampling techniques.

2.2 REVIEW OF PAST WORKS

Numerous researchers have explored the development and application of alternative

estimators—such as robust regression and modified maximum likelihood techniques—to

enhance estimation efficiency in Poisson regression models. These approaches are

particularly designed to reduce the adverse effects of outliers and overdispersion on

parameter estimation. For instance, Zaman (2019) and Zaman & Bulut (2020) introduced a

class of robust ratio-type estimators derived from a combination of existing ratio estimators.

Earlier, Raghav (2011) laid the groundwork for these developments by exploring robust

regression-based estimators. Shahzad et al. (2020) proposed a novel use of quantile

regression in constructing ratio estimators, replacing the conventional least squares approach.

Similarly, Ali et al. (2021) introduced a set of estimators utilizing robust regression strategies

in the context of sensitive surveys under simple random sampling. Tailor (2009) introduced a

modified ratio-cum-product estimator under stratified random sampling, demonstrating

improved efficiency over conventional estimators. Other studies have recommended down-

weighting data points with large residuals in order to reduce their influence and improve

estimator efficiency.

Comparative studies have also been conducted to evaluate the performance of different

estimation methods—such as the maximum likelihood estimator (MLE), moment-based


10
estimators, and Bayesian approaches—under various data conditions. Wani et al. (2022), for

example, proposed a Poisson-regression-based class of estimators for estimating the

population mean under simple random sampling without replacement (SRSWOR). Their

study derived expressions for the bias and mean square error (MSE) of these estimators to the

first order of approximation and presented theoretical comparisons with existing methods,

identifying conditions under which the proposed estimators outperformed traditional ones.

Abonazel (2019) investigated the use of Poisson regression in datasets containing outliers. A

Monte Carlo simulation study compared traditional maximum likelihood estimation with

robust alternatives like Mallows quasi-likelihood and weighted MLE. The results

demonstrated that robust estimators, particularly the weighted MLE, showed better

performance in the presence of outliers. More recently, Ahmed et al. (2024) evaluated three

robust methods—M, S, and MM estimators—on panel data with outliers using COVID-19

data from twelve European countries. Their analysis revealed that classical MLE was highly

sensitive to outliers, while MM estimators produced more stable and accurate predictions.

Ertan and Akay (2023) proposed a new class of biased estimators for Poisson regression,

developed based on the preliminary test estimator (PRE) framework. The authors evaluated

their efficiency using the asymptotic matrix mean square error criterion and validated the

findings with two Monte Carlo simulation studies. Their results showed that the proposed

estimators outperformed existing biased estimators, supported by real-data applications.

Further, optimal classes of estimators—such as ratio-type, product-type, and regression-type

estimators—have been proposed for estimating the population mean in Poisson regression

settings. Bhushan et al. (2023) introduced such classes under ranked set sampling (RSS) by

incorporating multiple auxiliary variables. They demonstrated, both theoretically and

computationally, that these estimators offered superior performance compared to

11
conventional approaches, including traditional ratio and regression estimators. Bias and MSE

expressions were derived, and their superiority was confirmed through studies on artificial

and real-world datasets.

While MLE generally performs well in large samples, its efficiency tends to diminish in

small-sample scenarios. In contrast, Bayesian estimators often yield better results when prior

knowledge is available. Kadilar et al. (2007) presented a robust regression method based on

the Huber-M approach, while Oral and Kadilar (2011a, 2011b) refined estimation strategies

under SRS and enhanced MLE techniques. Abid et al. (2016) investigated estimators based

on unconventional location parameters, and Koç et al. (2022) developed regression-ratio

estimators for mean estimation in double sampling contexts. One key challenge with applying

linear regression to count data is that it may yield negative predicted values and does not

account for the discrete nature of the Poisson distribution. Cameron and Trivedi (1998)

emphasized this issue in their foundational work on count data modelling, where they

discussed the limitations of linear regression in such scenarios.

Simulation studies have proven essential in assessing the performance of estimators under

controlled settings. These simulations help identify the most efficient estimator across

varying sample sizes and parameter conditions. For example, Abonazel et al. (2023)

introduced a two-parameter Liu estimator for the Conway–Maxwell–Poisson (COM-Poisson)

regression model. Using Monte Carlo simulations and real data applications, they compared

this new estimator with ridge, Liu, and modified one-parameter Liu estimators. The proposed

estimator incorporated two shrinkage parameters to address multicollinearity and

demonstrated superior performance in terms of MSE, confirming its effectiveness through

both theoretical derivation and empirical validation.

12
2.2 FAMILY OF EXISING ESTIMATOR

2.2.1 Kadilar and Cingi (2004) family of Estimators

Kadilar and Cingi (2004) introduced a set of ratio estimators for estimating the population

mean Y in simple random sampling (SRS), assuming that the auxiliary variable A is known.

Their approach was motivated by earlier work from Sisodia and Dwivedi (1981) and

Upadhyaya et al. (1985), where Y represents the study variable and A denotes an auxiliary

variable.

ȳ +b ( Ā−ā)
Ȳ KC 1= Ā, (1)
ā

ȳ +b ( Ā−ā ) (2)
Ȳ KC 2= ( Ā +C a) ,
ā+C a

ȳ +b ( Ā−ā)
Ȳ KC 3= ( Ā+ϴ ( a ) ) , (3)
ā+ ϴ(a)

ȳ + b ( Ā−ā )
Ȳ KC 4 = ( Āϴ ( a )+ Ca ), (4)
āϴ ( a )+ Ca

ȳ +b ( Ā−ā)
Ȳ KC 5= ( Ā C a +ϴ ( a ) ) . (5)
ā Ca +ϴ ( a )

Where:

Ȳ, Ā: population means associated with study and auxiliary variables (Y, A), respectively.

ȳ, ā: sample means associated with y and x, respectively.

Cy, Ca: population coefficients of variation associated with Y and A, respectively.

13
𝝧(a): population coefficient of the kurtosis associated with A.

The mean squared error (MSE) formulas for the estimators were derived as presented by

Kadilar and Cingi (2004).

1−f 2 2 2 2 2 2
MSE (Ȳ KCi )= [ R KCi S a +2 B R KCi S a+ B S a−2 R KCi Say −2 B S ay + S y ](6)
n

Where:
N = population size; n: sample size.
F = sample ratio, f = n/N.
S2y, S2a = population variance corresponding with Y and A, respectively.
Say = population covariance between X and Y, respectively.
𝜌xy = correlation coefficient between X and Y.

B = Slope coefficient estimated using the least squares method, B = Sxy/S2x.

The population ratio RKCi, i = 1, 2, …, 5 can be gained as follows:

Ȳ Ȳ Ȳ Ȳ ϴ (a) Ȳ Ca
R KC 1= ,R = , R KC 3= , R KC 4= , R KC 5 = (7)
Ā KC 2 Ā+C a Ā+ϴ ( a ) Āϴ ( a )+C a Ā C a +ϴ ( a )

2.2.2 Usman Shahzad et al (2021) family of Estimators

Usman Shahzad et al (2021) proposed a Poisson regression model under simple random sampling by

utilizing Poisson regression-based mean estimator and discovers its associated formula of the mean

square error (MSE). Their proposed estimator follows:

y U = y+ b p ( X−x ) (8)

And the MSE to be;

1−f 2 2 2 2 2
MSE ( y U )= [Y C y −2 B p Y X ρ xy C x C y +B p X C x ] (9)
n
14
15
2.2.3 Koҫ (2021) family of Estimator.
Recently, Koҫ (2021) proposed enhancements to the previously mentioned estimators by
introducing new ratio estimators derived from Poisson regression, as detailed below:
ȳ + b p ( Ā−ā)
ȳ H 1= Ā ,10
ā
ȳ + b p ( Ā−ā)
ȳ H 2=
ā+C a
[ Ā +C a ] , 11
ȳ +b p ( Ā−ā )
ȳ H 3= [ Ā +ϴ ( a ) ], 12
ā +ϴ ( a )
ȳ +b p ( Ā−ā)
ȳ H 4= [ Ā ϴ ( a )+C a ], 13
ā ϴ ( a ) +C a

ȳ +b p ( Ā−ā )
ȳ H 5= [ Ā C a+ϴ ( a ) ],14
ā C a +ϴ ( a )

Here, bp denotes the slope coefficients estimated through Poisson regression, with E(bp)=Bp.

The mean squared error (MSE) of ȳHi, Hi = 1, 2, …, 5, is expressed as follows:

1−f 2 2 2 2 2 2
MSE (Ȳ Hi)= [R KCi Sa +2 B p R KCi S a +B p S a−2 R KCi S ay −2 B p S ay +S y ]15
n

Furthermore, Koҫ (2021) showed that his estimators outperform those of Kadilar and Cingi
(2004) in efficiency, provided that at least one of the following conditions holds true:
B p−B<−2 R KCi
B p−B>−2 R KCi 16

2.2.4 Yashpal S. R. et al (2024) family of Estimators

Yashpal S. R. et al (2024) proposed a Poisson regression-based regression-type mean

estimator with simple random sampling and found its bias and mean square error formula as

presented below;

{( )}
δ

) (
−α G
a G+ β g
K pi = y α G+ β
+b p 1− (17)
a g+ β G
16
Bias ( K Pi )=
1−f
n [
Y ( w 2i + B p )
δ ( δ−1 ) 2
2
( wi + B p ) C 2g−δ C gy
] (18)

1−f 2 2 2 2
MSE ( K Pi )=
n
[
S y + R δ ( w i + B p ) S g−2 Rδ ( wi +B p ) S gy
2 2
] (19)

Y
They obtained the optimum value of δ by differentiating the equation wrt to δ, where δ= to
G

have

1−f 2 2
MSE ( K Pi )= S y (1−ρ ) (20)
n

Where Y is the research variable and G is an auxiliary variable, δ is a constant, α and β are

the coefficients that can take the values 0, 1, S g, C g, β 2(x) and the coefficient of correlation

between the study variables and supplementary.

where bp is the slope coefficient that can be obtained using a Poisson regression model and

E(bp)=Bp.

17
CHAPTER THREE

METHODOLOGY

3.0 INTRODUCTION

In this study, stratified random sampling was adopted, a new proposed Estimator was

derived, the MSE as well as the optimum value of the Estimator are obtained. Also, the newly

proposed estimators were applied to real life data sets and compare with existing estimators.

3.1 STRATIFIED RANDOM SAMPLING

Stratified random sampling is a widely used probability sampling technique. Unlike the

traditional Simple Random Sampling (SRS), where some elements (or values) are selected

randomly from a population without considering any factor (feature, attribute or

characteristic), Stratified Random Sampling (StRS) is a statistical technique used to ensure

that a sample accurately represents the entire population by dividing it into smaller sub-

groups, known as strata, based on shared (similar) characteristics or attributes, such as age,

gender, income, education level, among others. After the strata are formed, random sample

are drawn from each sub-group. It is particularly useful in scenarios where the population is

divers, and simple random sampling might fail to capture all unique features (i.e., not a

representative of the population), reducing sampling bias and enhancing the precision of the

results.

Mainly, there are two (2) types of Stratified Random Sampling namely:

i. Proportionated Stratified Random Sampling

ii. Disproportionated Stratified Random Sampling

18
19
- Proportionated Stratified Random Sampling

It is a form of Stratified Random Sampling where the number of samples drawn from each

stratum is directly proportional to the size of that stratum relative to the entire population. In

other words, the proportion of the sample taken from each stratum reflects its actual share in

the population.

- Disproportionated Stratified Random Sampling

In disproportionate stratified random sampling, the sample size for each stratum is not

proportional to the stratum's size in the population. This means that a stratum that is

considered more important for the analysis may be over-sampled, while a stratum that is less

important may be under-sampled.

This type of stratified random sampling is most used when the strata are heterogeneous in

size or when some strata are considered more important than others. It can be more efficient

than proportionate stratified random sampling, but it may not be as representative of the

entire population.

3.2 POISSON REGRESSION MODEL

In Poisson regression, the study variable yi (where yi = 0, 1, 2, . . .) represents the count of

events occurring within a specified time and follows a Poisson distribution defined by:

− µi yi
e µi
P ( y i ; µi ) = !
, µ>0 (21)
yi

and its mean and variance are both the same,


E(Yi) = var (Yi) = µi.

The natural logarithm of the likelihood function is given by:


20
n
l ( µ ; y )=∑ ( y i ∈ ( µ i )−µi−¿ ( y i! ) ) (22)
i =1

Let A be the auxiliary variable matrix of size n x (k + 1). The relationship between Y i and ith

row of matrix A, denoted as ai in association with a(µi), is given by:

T
¿ ( µi )=ξ=ai W (23)

Here W = (W0, W1, …, Wk) denotes the vector of regression coefficients. This formulation is

commonly known as the Poisson regression model. The vector W represents the maximum

likelihood estimator of W, which can be obtained by differentiating equation 3 with respect to

W.

∑ ¿¿ (24)
i=1

Iterative methods, such as the Fisher scoring and Newton-Raphson algorithms, are used to

solve these k equations (refer to Koç (2004) and Montgomery et al. (2006)).

3.3 PROPOSED ESTIMATOR AND ITS MEAN SQUARE ERROR (MSE)

The newly developed estimator for the population mean can be structured within the

frameworks of Zaman and Bulut (2020) and Ali et al. (2021). However, in this study, we

apply their frameworks using Poisson regression as follows:

{( )}
δ

) (
−α X
a X +β x (25)
K st = y st α X+β
+b p 1− st
a x st + β X

y st −Y x st −X
ε ∘= , ε 1=
Y X

E (ε o) = E (ε 1) = 0

E ( ε o ) =∑ w h f 1 c y , E ( ε 1 ) =∑ w h f 1 c x , E ( ε o ε 1 )=∑ wh f 1 c y c x p y x
2 2 2 2
h h h h h h

21
Expressing equation (25) in terms of ε and neglecting the higher order, such that

αX
w=
α X+β

{( )}
δ

) (
−α X
α x st α X+β X −x st
K pst = y st +b p (26)
α X +β X

{( )}
δ

) (
−w
α x st + β+(α X−α X ) x st −X
¿ Y ( 1+ε o ) +b p
α X+β X

{( )}
δ

) (
−w
(α X+ β)+(α x st −α X ) x st −X
¿ Y ( 1+ε o ) −b p
α X +β X

{( }
δ

)
−w
α X + β α (x st − X )
¿ Y ( 1+ε o ) + −b p ε 1
α X +β α X +β

{[ ( )] }
−w δ
α (x st − X) X
¿ Y ( 1+ε o ) 1+ −b p ε 1
α X+ β X

{[ )] }
−w δ

)(
x st −X
¿ Y ( 1+ε o ) 1+
αX
α X +β ( X
−b p ε 1 (27)

¿ Y ( 1+ε o ) {( 1+w ε 1 ) −b p ε 1 }
−w δ
(28)

{( }
δ

K pst =Y ( 1+ ε o ) 1−w ε 1 +w
(w1 +1) 2
2
2
ε 1−b p ε 1 3
1 ) (29)

Using Binomial Expansion Equation (29) becomes.

{ }
δ
2 3 (w+1) 2
K Pst =Y ( 1+ε o ) 1−w ε 1−b p ε 1 +w ε1
2

{ }
δ
2 3 ( w+1) 2
K Pst =Y ( 1+ε o ) 1−(w −b p )ε 1 +w ε1 (30)
2

22
[ ( ) ( )
]
3 2
w3 ( w+1 ) 2 δ ( δ−1 ) w ( w+1 ) ε 1
¿ 1+δ −( w +b p ) ε 1 + −( w +b p ) ε 1+
2 2
ε1 +
2 2 2
K pst =Y ( 1+ ε o )

[ ]
3
+ δ ( δ−1 ) ( δ−2 ) w 2 ( w+1 ) 2
− ( w + b p ) ε 1+
2
ε 1 + ...
6 2

[ ]
3 6 2
δ w ( w+1 ) 2 δ ( δ −1 ) 2 2 w ( w+ 1 ) 4
2
¿ Y ( 1+ε 0 ) ¿1−δ w ε i−δ b p ε 1 +
2
ε 1+
2
( w +b p ) ε 2i + ¿∧w3 ( w+1 ) ( w2 +b p) ε 31 +
2
ε 1 + ...

We neglect terms with order greater than 2

[ ]
3
δ w ( w+1 ) 2 δ ( δ−1 ) 2 2
¿ Y ( 1+ε 0 ) 1−δ ( w2 +b p ) ε 1 + ε 1+ ( w +b p ) ε 21
2 2

[
¿ Y ( 1+ε 0 ) 1−δ ( w2 +b p ) ε 1 +
δ ( δ−1 ) 2
2
(
2
w + b p ) ε 21
]

( )
δ ( δ−1 ) 2 2 2
¿ 1−δ ( w +b p ) ε 1 +
( w +b p ) ε 1 +ε 0
2

¿Y 2
δ ( δ−1 ) 2 2
¿−δ ( w +b p ) ε 0 ε i + ( w + b p ) ε0 . ε1
2 2
2

( )
δ ( δ−1 ) 2 2
¿ 1−δ ( w2 +b p ) ε 1 + ( w +b p ) ε 21 +ε 0
¿Y 2
¿−δ ( w +b p ) ε 0 ε i
2

(
K Pst =Y ¿ 1−δ ( w +b p) ε 1 +
2 δ ( δ−1 ) 2
2
(
2 2
w +b p ) ε 1 +ε o−¿∧δ ( w +b p ) ε 0 ε i
2
)
(31)

( )
δ ( δ−1 ) 2 2
¿−δ ( w2 +b p ) ε i + ( w +b p ) ε 21 +ε o
K Pst =Y −Y 2
¿−δ ( w + b p ) ε 0 ε i
2

( )
δ ( δ−1 ) 2 2

K Pst −Y =Y
¿
2
( w +b p ) ε 21 +ε 0 −δ ( w2 +b p ) ε i
(32)
¿−δ ( w2 +b p ) ε 0 ε i

Take the expression of equation (32) in order to obtain the bias

E ( K Pst −Y ) =Bias
23
Bias ( K Pst )=Y
[ δ ( δ−1 ) 2
2
(
2
w + b p ) E ( ε 2i ) + E ( ε 0 )−δ ( w 2+ b p ) E ( ε i )−δ ( w 2+ b p ) E ( ε 0 ε i )
]
Y ( δ ( δ−1 ) 2
2
( w +b p ) ∑ W h f i c2x + 0−δ ( w 2+ b p ) ( 0 )−δ ( w 2+b p ) ∑ W h f i c y c x ρ y x
2
h h h h h )
¿Y ( δ ( δ−1
2
)
(w + b ) ∑ W f c −δ ( w +b ) ∑ W f c
2
p
2
h i
2
xh
2
p h i yh cx ρ y x
h h h )
Bias ( K Pst )=∑ W h f i Y ( w 2+b p ) [ δ ( δ−1 ) 2
2
( w + b p ) c 2x h−δ c y c x ρ y x h h h h ] (33)

To obtain the mean square error (MSE), we square both sides of equation (32) and take the
expectation.

2
( K Pst − y ) = Y [{ δ ( δ−1 )
2
(
2
}
w2 +b p ) ε 21 +ε 0−δ ( w2 +b p ) ε 1−δ ( w2 +b p ) ε 0 ε 1 ]
[
¿ ε o−δ ( w +b p ) ε 1 ε o +δ ( w +b p ) ε 1−δ ( w +b p ) ε 0 ε 1
2 2 2 2 2 2 2
] (34)

2 2 2 2
[
( K Pst − y ) =Y {δ ( w +b p ) ε 1−2 δ ( w +b p ) ε 0 ε 1 +ε o }
2 2 2 2
]
Taking expectation

[ {
E ( K Pst − y ) =E Y δ ( w + b p ) ε 1 −2 δ ( w +b p ) ε 0 ε 1 +ε o
2 2 2 2 2 2 2 2
}] (35)

¿> MSE ( K Pst ) =Y 2 δ 2 ( w 2+ b p ) { 2


∑ W h f i c 2x −2 δ ( w2 +b p ) ∑ W h f i c y c x ρ y x +∑ W h f i c 2y }
h h h h h h

MSE ( K Pst )=Y 2 δ 2 ( w2 +b p ) { 2


∑ W h f i c 2x −2 δ ( w 2+ b p ) ∑ W h f i c y c x ρ y x +∑ W h f i c2y }
h h h h h h

2
MSE ( K Pst )=Y ¿ (36)

Recall that
2 2
2
sy 2
sx ρ s y h sx h ρ s y h x h
c =
yh
h
,c =xh
h
, c y x =ρ c y c x = = (37)
Y
2
X
2 h h h h
YX YX

Substitute (37) in (36)

{∑ }
2 2
Y Y
2∑ ∑ W h f i ρ s y sx
2
W f s + δ ( w +b p ) W h f i s x −2 δ ( w +b p )
2 2 2 2 2
MSE ( K Pst )= h i yh
X h
YX h h

24
¿ {∑ W h f i s y + R δ ( w +b p )
2
h
2 2 2 2
∑ W h f i s 2x −2 R δ ( w2 +b p )∑ W h f i ρ s y s x }
h h h

(38)

{
MSE ( K Pst )=∑ W h f i s 2y + R2 δ 2 ( w 2+ b p ) s 2x −2 R δ ( w 2+ b p ) ρ s y s x
h
2
h h h
} (39)

To obtain the optimum value of δ , we differentiate eqn (39) with respect to δ

d ( MSE(K Pst ) ) d

=
dδ [∑ W f {s h i
2
yh
2
+ R2 δ 2 ( w2 +b p ) s 2x −2 R δ ( w2 +b p ) ρ s y s x h h h
}] (40)

{
= ∑ W h f i 0+ 2 Rδ ( w +b p ) s x −2 R ( w + b p ) ρ s y s x
2 2 2 2
h h h
}
¿>2 Rδ ( w +b p ) ∑ W h f i s 2x =2 R (w 2+ b p ) ∑ W h f i ρ s y s x
2 2
h h h

2 R ( w 2+ b p ) ∑ W h f i ρ s y s x ∑ W h f i ρ s y sx
δ opt = h h
= h h

2 R (w + b p )
2 2 2
∑W f s
2
h i xh
R(w 2+ b p ) ∑ W h f i s 2x h

(41)

Substitute δ opt into eqn (15) to get the optimum value of the MSE.

{ 2 R ∑ W h f i ρ s y sx
}
2
(W h f i ρ sy s x )
MSE ( K Pst )=∑ W h f i s + R
2 2
2 2
. ( w
2
+b ) s
h

h
. ( w +b p ) ρ s y s x
2 h h

R ( w +b p ) ∑ W h f i s x
y p x
R ( w + b p ) (∑ W h f i s x )
2 2 2 2 2 2 2 h h h

h h

{ (∑ W h f i ρ sy sx )
}
2
2 R ∑ W h f i ρ s y sx
∑ Wh f i sy+R ∑
2 2
¿
2
. ( w
2
+ b ) h
W f
h
s
2
− .( w + bp ) ρ s y sx
2 h h

R ( w + b p ) ∑ W h f i sx
2 p h i x
R ( w +b p ) (∑ W h f i s x )
2 2 2 2 2 2 h h h

h h

{ (∑ W h f i ρ s y s x )
}
2
2( ∑ W h f i ρ s y s x )
2

∑ Wh f i sy+
2
¿ −
h h h h

(∑ W h f i s 2x )
2
h
∑ W h f i sx 2
h

( ∑ W h f i ρ sy sx )
2
( ∑ W h f i ) ρ2 s2yh s 2xh
2

¿ ∑ W h f i s y− =∑ W h f i s y −
2 2 h h

(∑ W h f i s 2x )
2
∑ W h f i s 2x h
h

MSE ( K Pst )=∑ W h f i s y −∑ W h f i ρ s yh


2 2 2
(42)

¿ ∑ W h f i s y (1−ρ )
2 2
(43)

25
CHAPTER FOUR
RESULT AND DISCUSSION

4.1 Introduction

The chapter comprises the computation and analysis of the data used in the validation of the

proposed Estimators which was obtained using Microsoft Excel and R statistical software.

4.2 Empirical Study

Numerical analysis was conducted with the goal of evaluating the relative performance of the

proposed class of Estimators and family, as compared to the preliminary Estimators. The

evaluation was done empirically by utilizing two real population data sets and simulated data,

the Mean Square Error (MSE) of the estimators was computed.

The percentage relative efficiency (PRE) can be computed using:

(¿ y )
PRE=var x 100(44)¿
MSE ( y )

4.3 Applications

In this study, We tried to validate our theoretical findings using two real count

dataset and a simulated dataset, the first dataset is the enrolment of the unified

tertiary matriculation examination for the year 2018(study variable) and

number of teachers (auxiliary variables from three local government areas in

ogun state). Also, Gender was used as the

26
4.4 Descriptive Statistics (Data 1: Schools from 3 LGAs)

Descriptive Statistics and Data Visualization of the Sample

Sample

Mean 49.3494
2.27889
Standard Error 5
Median 37
Mode 32
Standard 35.9603
Deviation 6
1293.14
Sample Variance 8
2.22253
Kurtosis 3
1.51261
Skewness 4
Range 199
Minimum 1
Maximum 200
Sum 12288
Count 249

Figure 4.3.1 shows the distribution of sample variable.

27
Count of sample by States
100
90
80
70
60
50
40
30
20
10
0
State1 State2 State3

Figure 4.3.2 represents the distribution of the sample variable by states, showing the total

count of each state.

Distribution of LGAs

28% State1
36%
State2
State3

37%

Figure 4.3.3 illustrates the pie chart which represents the distribution of states and the

percentage they cover within the sample variable.

28
Population

Mean 1346.068
Standard Error 83.82328
Median 895
Mode 1756
Standard
Deviation 1322.709
Sample Variance 1749559
Kurtosis 8.378347
Skewness 2.291073
Range 9766
Minimum 118
Maximum 9884
Sum 335171
Count 249

Figure 4.3.4 is a histogram plot showing the distribution of the population variable

29
Sum of population by States

State1
28% 36% State2
State3

37%

Figure 4.3.5 illustrates the pie chart which represents the distribution of states and the

percentage they cover within the population variable.

Count of population by States


100

90

80

70

60

50

40

30

20

10

0
State1 State2 State3

Figure 4.3.6 represents the distribution of the population variable by states, showing the total

count of each state.

30
Figure 4.3.7 Box Plot of Population variable

Figure 4.3.8 Box Plot of Sample variable

31
250

200

150

100

50

0
0 2000 4000 6000 8000 10000 12000

Figure 4.3.9 Scatter plot

Descriptive Statistics and Data Visualization of the Population

Population

1304.49
Mean 8
90.2700
Standard Error 9
Median 890
Mode 400
Standard 1292.47
Deviation 1
Sample
Variance 1670481
10.6781
Kurtosis 7
2.54236
Skewness 2
Range 9766
Minimum 118
Maximum 9884
Sum 267422
Count 205

32
State3

State2
Total

State1

0 10 20 30 40 50 60 70 80

33
28%
36%

State1
State2
State3

37%

Sample

47.7073
Mean 2
2.41794
Standard Error 3
Median 36
Mode 26
Standard 34.6196
Deviation 7
Sample 1198.52
Variance 2
1.87502
Kurtosis 3
1.46606
Skewness 3
Range 170
Minimum 1
Maximum 171
Sum 9780
Count 205

34
State3

State2

State1

0 10 20 30 40 50 60 70 80

35
28%
36%
State1
State2
State3

37%

4.4.2 Computation of the MSE’s and PRE’s of the proposed Estimators using data I
Table # presents the MSE and PRE of the family of the proposed Estimator using

dataset I in comparison with the classical mean of stratified random sampling. Also,

Table # shows the performance of the proposed Estimators in comparison with the

existing Estimators using data set I.

ESTIMATOR MSE BIAS PRE RANK


Classical Mean Estimator 2962170169 0 100
Yashpal Estimator 1885227102 0.52921
KOC MSE 1776809433
Kadila & Cingi 1429502562
Proposed Estimator (KPST) 173989.5 -131.796

ESTIMATOR MSE BIAS PRE RANK


Classical Mean Estimator Statified sampling 620468.1 0
-
Proposed Estimator (KPST) 173989.5 131.796

Descriptive Statistics (Data 2: UTME 2017 & 2018)


36
Tables ###2-5 show the descriptive analysis of Unified Tertiary Matriculation Examination

registration for the years 2017 and 2018. The analysis shows the number of Applicants and

number of Admitted Students for the two years.

Table 2: Statistics of Applicants (2017)

APP17

Mean 46538.32432

Standard Error 4380.178398

Median 42095

Mode #N/A

Standard Deviation 26643.58503

Sample Variance 709880623.5

Kurtosis -1.184332828

Skewness 0.295468423

Range 95918

Minimum 6270

Maximum 102188

Sum 1721918

Count 37

Confidence Level
(95.0%) 8883.413532

Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“APP17” represents the number of JAMB Applicants in 2017. The result of the analysis

shows that:

37
1. The average number of students who applied for JAMB in 2017 are about 46,538

across all states (37, the Federal Capital inclusive) within the country (Nigeria).

2. Imo State appears to have the highest number of applicants for the year (2017) having

102,188 applicants, followed by Osun, Oyo, Ogun and Delta with 88,803, 87,827,

81,536 and 81,478 applicants respectively. This result shows that fairly the Ibos and

Yoruba ethnic groups are more interested in tertiary education in Nigeria.

3. The top 5 states with the lowest number of applicants for 2017 are as follows;

Sokoto, Kebbi, Yobe, Zamfara and the FCT with 14,478, 13,927, 13,767, 10,522, and

6,270. This result indicates that the Northerns (or Hausas) are not interested in

education as the other 2 major ethnic groups in Nigeria as they fully dominant this

frame.

4. The total (sum) number of all applicants for the year 2017 is 1,721,918.

Table 3: Statistics of Admitted Students (2017)

ADM17

Mean 15313.51351

Standard Error 1276.810804

Median 13665

Mode #N/A

Standard Deviation 7766.536918

Sample Variance 60319095.7

Kurtosis -0.913318656

Skewness 0.212030526

Range 28415

Minimum 2120

Maximum 30535

38
Sum 566600

Count 37

Confidence Level (95.0%) 2589.492332

Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)

“ADM17” represents the number of students that were admitted into tertiary institutions after

the examination (JAMB) in 2017. The result of the analysis shows that:

1. The average number of students who were admitted after the exam in 2017 are about

15,313 across all states (37, the Federal Capital inclusive) within the country

(Nigeria).

2. Imo State again appears to have the highest number of admitted students for the year

(2017) followed by Osun, Oyo, Ogun and Kano with 30535, 28209, 27,989, 26,589

and 25,647 respectively.

3. The top 5 states with the lowest number of admitted students for 2017 are as follows;

Jigawa, Kebbi, Sokoto, Zamfara and the FCT with 5,776, 5,104, 4,004, 2,744, and

2,120.

4. The total (sum) number of admitted students for the year 2017 is 566,600.

Table 4: Statistics of Applicants (2018)

APP18

Mean 44671.40541

Standard Error 4091.690178

Median 42560

Mode #N/A

Standard Deviation 24888.7797


39
Sample Variance 619451355

Kurtosis -1.16341722

Skewness 0.306237264

Range 86610

Minimum 6438

Maximum 93048

Sum 1652842

Count 37

Confidence Level (95.0%) 8298.332304

Data 1: Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“APP18” represents the number of JAMB Applicants in 2018. The result of the analysis
shows that:
1. The average number of students who applied for JAMB in 2018 are about 44,671

across all states (37, the Federal Capital inclusive) within the country (Nigeria).

2. The same states from 2017 seems to appear as the top 5 States with the highest

number of applicants for 2018 (though their positions got changed). Imo, Oyo, Osun,

Ogun and Delta with 93,048, 86,687, 86,065 and 80,453 applicants respectively.

3. The same states from 2017 seems to appear as the top 5 States with the lowest number

of applicants for 2018 (though their positions got changed), they are; Yobe, Kebbi,

Sokoto, Zamfara and the FCT with 15,536, 15,341, 13,493, 10,090, and 6,438. This

result indicates that the Northerns (or Hausas) are not interested in education as the

other 2 major ethnic groups in Nigeria as they fully dominant this frame.

4. The total (sum) number of all applicants for the year 2018 is 1,652,842.

40
Table 5: Statistics of Admitted Students (2018)

ADM18

Mean 14855.43243

Standard Error 1229.864943

Median 14074

Mode #N/A

Standard Deviation 7480.976394

Sample Variance 55965007.81

Kurtosis -0.749533705

Skewness 0.25991206

Range 27104

Minimum 2279

Maximum 29383

Sum 549651

Count 37

Confidence Level (95.0%) 2494.281713


Data 1:
Source: Joint Admissions and Matriculation Board (JAMB) (2017 &2018)
“ADM18” represents the number of students that were admitted into tertiary institutions after

the examination (JAMB) in 2018. The result of the analysis shows that:

1. the average number of students who were admitted after the exam in 2018 are about

14855 across all states (37, the Federal Capital inclusive) within the country (Nigeria).

2. The states from 2017 seems to appear as the top 5 States with the highest number of

admitted students for 2018 expect Kano which was replaced by Anambra.

Imo, Oyo, Osun, Ogun and Anambra with 29,383, 28,743, 28,085, 25,804 and 25,369

applicants respectively.

41
3. The same states from 2017 seems to appear as the top 5 States with the lowest number

of applicants for 2018 (there positions are not changed), they are;

Jigawa, Kebbi, Sokoto, Zamfara and the FCT with 5,616, 4,494, 3,860, 2,596, and

2,279.

4. The total (sum) number of admitted students for the year 2018 is 549,651.

4.4.1 Charts of Number of Applicants and Admitted Students for the years 2017 &
2018
Figures 1-4 show the bar charts of Unified Tertiary Matriculation Examination registration

for the years 2017 and 2018. The charts show the number of Applicants and number of

Admitted Students in each state for the years 2017 and 2018.

ZAMFARA
TARABA
RIVERS
OYO
ONDO
NIGER
LAGOS
KOGI
KATSINA
KADUNA
IMO
FCT
EKITI
EBONYI
CROSS RIVER

BENUE
BAUCHI
AKWA IBOM

ABIA
0 20000 40000 60000 80000 100000 120000

42
Figure 1: No. of Applicants by State (2017)

ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 5000 10000 15000 20000 25000 30000 35000

Figure 2: No. of Admitted by State (2017)

43
ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 10000 20000 30000 40000 50000 60000 70000 80000 90000 100000

Figure 3: No. of Applicants by State (2018)

44
ZAMFARA
YOBE
TARABA
SOKOTO
RIVERS
PLATEAU
OYO
OSUN
ONDO
OGUN
NIGER
NASARAWA
LAGOS
KWARA
KOGI
KEBBI
KATSINA
KANO
KADUNA
JIGAWA
IMO
GOMBE
FCT
ENUGU
EKITI
EDO
EBONYI
DELTA
CROSS RIVER
BORNO
BENUE
BAYELSA
BAUCHI
ANAMBRA
AKWA IBOM
ADAMAWA
ABIA
0 5000 10000 15000 20000 25000 30000 35000

Figure 4: No. of Admitted by State (2018)

4.4.2 Computation of the MSE’s and PRE’s of the proposed Estimators using data II
Table # presents the MSE and PRE of the family of the proposed Estimator using dataset II in

comparison with the classical mean of stratified random sampling. Also, Table # shows the

performance of the proposed Estimators in comparison with the existing Estimators using

data set II.

45
ESTIMATOR MSE BIAS PRE RANK
Classical Mean Estimator 7.18E+12 0
-
Yashpal Estimator 9.36154E+11 27.1224
KOC MSE 8.01E+06
Kadila & Cingi 3.76E+11

ESTIMATOR MSE BIAS PRE RANK


Classical Mean Estimator Statified
sampling 8208847 0
455673. -
Proposed Estimator (KPST) 2 786.294

4.6 Estimation of Parameters


N 1=89 , N 2=91 , N 3 =69 , N =249 , N =249 , n=205

N1 N2 N3
W 1= =0.35 , W 2= =0.37 , W 3 = =0.28
N N N

n1=72, n2=75 ,n 3=57

[ ]
Y rat =∑ W h y h
Xh
xh

y 1=1546.137 , y 2=1472.373 , y 3=774.140

Y 1=1581.101 , Y 2=1514.670 , Y 3=820.55

x 1=65.18 , x 2=37.573 , x 3=38.67

X 1 =66.44 , X 2=38.54 , X 3=41.57

Y rat =1345.84

S x 1=40.11, S x 2=27.56 , S x 3 =25.98

S y 1=1220.48 , S y 2=1559.48 , S y 3=745.87


46
ρ1 xy=0.675 , ρ2 xy=0.920 , ρ3 xy =0.89

f 1=0.820 , f 2 =0.824 , f 3=0.826

B(Y rat )=1.1371

MSE (Y ¿¿ rat)=89841.145 ¿

………….

4.5 Proposed Estimator

……………

47
CHAPTER FIVE
SUMMARY, CONCLUSION AND RECOMMENDATION
5.1 SUMMARY

5.2 CONCLUSION

5.3 RECOMMENDATION

5.4 CONTRIBUTION TO KNOWLEDGE

5.5 AREA OF FURTHER STUDY

…………

48
REFERENCES

T. Zaman and H. Bulut, “Modified regression estimators using robust regression methods and
covariance matrices in stratified random sampling,” Communications in Statistics- Theory
and Methods, vol. 49, no. 14, pp. 3407–3420, 2020.
N. Ali, I. Ahmad, M. Hanif, and U. Shahzad, “Robust-regression type estimators for
improving mean estimation of sensitive variables by using auxiliary information,”
Communications in Statistics: Theory and Methods, vol. 50, no. 4, pp. 979–992, 2021.
L. N. Upadhyaya, H. P. Singh, and J. W. E. Vos, “On the estimation of population means and
ratios using supplementary information,” Statistica Neerlandica, vol. 39, no. 3, pp. 309–318,
1985.
D. C. Montgomery, E. A. Peck, and G. G. Vining, Introduction to Linear Regression
Analysis, John Wiley and Sons, Hoboken, 4th edition, 2006.
S. Bhushan,and [Link], “Improved estimation of population mean in simple random
sampling using attribute”. Thailand Statistician 22 (2), 374–389, 2024.
S. Bhushan, A. Kumar, N. Alsadat, M.S. Mustafa, and M.M. Alsolmi, “Some optimal classes
of estimators based on , Axioms 12 (6), 515, 2023
H. Koç, “Ratio-type estimators for improving mean estimation using Poisson regression
method,” Communications in Statistics - Theory and Methods, vol. 50, no. 20, pp. 4685–
4691, 2021.
E.G. Koçyigit, and K.U.I. Rather. “The new sub-regression type estimator in ranked se˘ t “,
J. Statistic. Theor. Pract. 17 (2), 27, 2023.
C. Kadilar and H. Cingi, “Ratio estimators in simple random sampling,” Applied
Mathematics and Computation, vol. 151, no. 3, pp. 893–902, 2004.
F.A. Lukman, E. Adewuyi, K. Mansson and B.M. Golam, “A new estimator for the
multicollinear Poisson regression model: simulation and application,” Scientific Report 2021
Feb 12;11:3732. doi: 10.1038/s41598-021-82582-w
B. V. S. Sisodia and V. K. Dwivedi, “A modified ratio estimator using coefficient of
variation of auxiliary variable,” Journal of the Indian Society of Agricultural Statistics, vol.
33, no. 2, pp. 13–18, 1981
M. R. Abonazel and O. M. Sabar, “A comparative study of robust estimators for poisson
regression model with outliers”, Journal of Statistics Application & Probability, vol. 9, no. 2,
pp. 279-286, 2020
M.R, Abonazel, F.A Awwad, E. Tag Eldin, B.M.G Kibria and I.G. Khattab ”Developing a
two-parameter Liu estimator for the COM–Poisson regression model: Application and
simulation”, Front. Appl. Math. Stat. 9:956963 (2023). doi: 10.3389/fams.2023.956963.
E. Oral, and C. Kadilar. “Improved ratio estimators via modified maximum likelihood”.
Pakistan J. Statistic. 27 (3), 269–282, 2011a.

E. Oral and C. Kadilar. “Robust ratio-type estimators in simple random sampling”. J. Korean
Surg. Soc. 40 (4), 457–467. [Link] 2011b.

49
U. Shahzad, S. Shahzadi, N. Afshan, N.H. Al-Noor, D.A. Alilah, M. Hanif and M.M. Anas,
“Poisson regression-based mean estimator”, Mathematical Problems in Engineering Volume
2021, Article ID 9769029, 6 pages [Link]
M. Abid, N. Abbas, H.Z. Nazir, and Z. Lin. “Enhancing the mean ratio estimators for
estimating population mean using non-conventional location parameters”. Rev. Colomb.
Estadística 39 (1), 63–79. [Link] 2016
Z. H. Wani, S.E.H. Rizvi, M.I. Jeelani and S. Mushtaq, “Modified regression estimators for
improving mean estimation – Poisson regression approach,” Pakistan Journal of Statistics and
Operation Research, vol. 18, no. 4, pp 985-994, 2022
Y.S. Raghav, A.A.H. Ahmadini, A.M. Mahnashi and K.U.I Rather, “Enhancing estimation
efficiency with proposed estimator: A comparative analysis of Poisson regression-based
mean estimators,”, Kuwait Journal of Science, 52 (2025) 100282, pp 1-8
[Link]
K.U.I. Rather, E.G. Koçyigit˘ , R. Onyango, and C. Kadilar, Improved regression in ratio type
estimators based on robust M-estimation. PLoS One 17 (12), e0278868, 2022

==========
[Link]
[Link]

Ahmed, A., et al. (2024). Enhancing estimation efficiency with proposed estimator: A
comparative analysis of Poisson regression-based mean estimators. Kuwait Journal of
Science.
Tailor, R. (2009). A modified ratio-cum-product estimator of finite population mean in
stratified random sampling. Data Science Journal, 8, 182–189.
Zakari, Y., & Muhammad, I. (2023). Modified estimator of finite population variance under
stratified random sampling. Engineering Proceedings.
Khan, S., et al. (2024). An effective and economic estimation of population mean in stratified
random sampling using a linear cost function. Heliyon.
Singh, R., & Malik, S. (2014). A new estimator for population mean using two auxiliary
variables in stratified random sampling. arXiv preprint.
Verma, H. K., Sharma, P., & Singh, R. (2014). Improved estimator of finite population mean
using auxiliary attribute in stratified random sampling. arXiv preprint.
Oyeyemi, A. S. (2020). Robust estimation in stratification sampling. International Journal of
Engineering Research & Technology.

50
Iorlaha, P. I., Nwaosu, S. C., Uba, T., & Ikughur, A. J. (2024). Regression estimator of
population mean with random missing values in stratified two-stage sampling using auxiliary
information. Journal of Statistical Sciences and Computational Intelligence.

51

You might also like