STA 202: Statistics II Course Overview
STA 202: Statistics II Course Overview
Page 2 of 174
STA202: Statistics II
Vice Chancellor’s Message
It is with great pleasure that I welcome you as learners to the Olabisi Onabanjo University
Open and Distance Learning Centre.
Massive and Democratisation of higher education via Open and Distance Learning as
advocated globally has since been one of the goals of Olabisi Onabanjo University
Management, hence, Open and Distance Learning constitutes one of the areas of focus
since my assumption of duty. Through the efforts of the University Governing Council
and Senate, the establishment of the Open and Distance Learning Centre was approved in
July, 2016.
Open and Distance Learning is a mode of study that affords tertiary education
opportunities to all and sundry regardless of age, gender, location, space and other limiting
factors.
Quite a large number of qualified applicants for tertiary education are denied admission
yearly, there are also several others who wish to advance educationally but could not,
because of their job which is their means of livelihood.
Olabisi Onabanjo University via its Open and Distance Learning Centre offers quality,
technology driven, flexible, self-directed and cost effective tertiary education. It is a viable
option for learners who wish to study online from their location and at desired time.
This course material provides learners with vital information relevant to our programme
and schedules. I advise learners to make judicious use of it. I congratulate our Open and
Distance Learning Centre Staff, Department and Faculty for their effort towards the
production of this handbook.
I hope your learning experience with the Olabisi Onabanjo University Open and Distance
Learning Centre is memorable and exciting.
Page 3 of 174
STA202: Statistics II
Course Study Guide
Introduction
STA 202 titled Statistics II is a 3-unit course for students studying towards acquiring a
Bachelor of Science in Accounting. The course is divided into 8 study sessions. The
course will introduce you to the basic statistics concept in solving practical problems.
The course study guide therefore gives you an overview of what STA 202 is all about, the
textbooks and other materials to be referenced, what you are expected to know in each
unit and how to work through the course materials. Define a set and understand sampling,
solve correlation and regression of any statistical data and applications of Times Series
Analysis.
Recommended Study Time
This course is a 3 unit course divided into 8 study sessions. You are enjoined to spend at
least 3 hours in studying the content of each study unit
What you are about to learn in this course
The overall aim of this course, STA 202 is to introduce you to sampling theory and
estimation techniques, Simple correlation analysis, Simple regression analysis, Test of
hypothesis, Index numbers and Time series analysis.
Course Aims
This course aims to introduce students to the basic statistical terms. It is expected that the
knowledge will help the reader to effectively use statistical principles to solve even life
problems.
Course Objectives
It is important to note that each unit has specific objectives. You should study them
carefully before proceeding to subsequent units. Therefore, it may be useful to refer to
these objectives in the course of your study of the unit to assess your progress. You should
always look at the unit objectives after completing a unit. In this way, you can be sure that
you have done what is required of you by the end of the unit.
Page 4 of 174
STA202: Statistics II
However, the overall objective of STA 202 is to enable students to analyze and interpret
data collected from a variety of types of research designs, within a linear model
framework.
Working through this course
In order to have a thorough understanding of the course units, you will need to read and
understand the contents, practice the steps by designing and implementing a mini
computer application system for your department and be committed to learning and
implementing your knowledge.
This course is designed to cover approximately fifteen weeks and it will require your
devoted attention. You should do the exercises in the Tutor-Marked Assignments and
submit to your tutors via the Learning Management System (LMS).
Course Materials
The major components of the course are;
1. Course Guide
2. Printed Lecture materials
3. Text Books
4. Interactive DVD
5. Electronic Lecture materials via LMS
6. Tutor Marked Assignments
Assessment
There are two aspects to the assessment of this course. First, there are tutor marked
assignments and second, the written examinations. Therefore, you are expected to take
note of the facts, information and problem solving gathered during the course. The tutor
marked assignments must be submitted to your tutor for formal assessment in accordance
to the deadline given. The work submitted will count for 30% of your total course mark.
At the end of the course, you will need to sit for a final written examination. This
examination will account for 70% of your total score. You will be required to submit some
Page 5 of 174
STA202: Statistics II
assignments by uploading them to STA 202 page on the Learning Management System
(LMS).
Tutor-Marked Assignment (TMA)
There are TMAs in this course. You need to submit all the TMAs. The best 10 will
therefore be counted. When you have completed each assignment, send them to your tutor
as soon as possible and make certain that it gets to your tutor on or before the stipulated
deadline. If for any reason you cannot complete your assignment on time, contact your
tutor before the assignment is due to discuss the possibility of extension. Extension will
not be granted after the deadline, unless on extraordinary cases.
Final Examination and Grading
The final examination for STA 202 will last for a period not more than 2hours and has a
value of 70% of the total course grade. The examination will consist of questions which
reflect the Self-Assessment Questions (SAQs), In-text Questions (ITQs), some applied
questions and tutor marked assignments that you have previously encountered.
Furthermore, all areas of the course will be examined. It would be better to use the time
between finishing the last unit and sitting for the examination to revise the entire course.
You might find it useful to review your TMAs and comment on them before the
examination. The final examination covers information from all parts of the course. Most
examinations will be conducted via Computer Based Testing (CBT)
Tutors and Tutorials
There are few hours of face-to-face tutorial provided in support of this course. You will be
notified of the dates, time and location together with the name and phone number of your
tutor as soon as you are allocated a tutorial group. Your tutor will mark and comment on
your assignments, keep a close watch on your progress and on any difficulties you might
encounter and provide assistance to you during the course. You must submit your tutor
marked assignment to your tutor well before the due date. At least two working days are
required for this purpose. They will be marked by your tutor and returned as soon as
possible via the same means of submission.
Page 6 of 174
STA202: Statistics II
Do not hesitate to contact your tutor by telephone, e-mail or discussion board if you need
help. The following might be circumstances in which you would find help necessary:
contact your tutor if:
You do not understand any part of the study units or the assigned readings.
You have difficulty with the self-test or exercise.
You have questions or problems with an assignment, with your tutor’s comments on an
assignment or with the grading of an assignment.
You should endeavour to attend the tutorials. This is the only opportunity to have face-to-
face contact with your tutor and ask questions which are answered instantly. You can raise
any problem encountered in the course of your study. To gain the maximum benefit from
the course tutorials, have some questions handy before attending them. You will learn a
lot from participating actively in discussions.
Good luck!
Recommended Texts
The following texts and Internet resource links will be of enormous benefit to you in
learning this course:
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 7 of 174
STA202: Statistics II
Table of Contents
Vice Chancellor’s Message .................................................................................................. 3
Course Study Guide .............................................................................................................. 4
Introduction....................................................................................................................... 4
Table of Contents .................................................................................................................. 8
Study Session 1: Sampling Theory and Estimation Techniques ........................................ 14
Introduction..................................................................................................................... 14
Learning Outcomes for Study Session 1 ........................................................................ 14
1.1 Definition of Sampling ........................................................................................ 15
1.1.1 Sample Survey .............................................................................................. 15
1.1.2 Advantages of Sample Survey...................................................................... 15
1.1.3 Disadvantages of Sample Survey ................................................................. 16
1.1.4 Sampling Frame ............................................................................................... 16
1.2 Type of Sampling Method ................................................................................... 17
1.2.1 Sampling with Replacement ......................................................................... 18
1.2.2 Sampling without Replacement .................................................................... 18
1.3 Types of Sampling Technique ............................................................................. 18
1.3.1 Non-Probability Sampling ............................................................................ 19
1.3.2 Judgment Sampling ...................................................................................... 20
1.3.3 Quota Sampling ............................................................................................ 20
1.3.4 Haphazard Sampling..................................................................................... 20
1.3.5 Probability Sampling .................................................................................... 20
1.3.6 Simple Random Sampling (SRS) ................................................................. 21
1.3.7 Systematic Sampling .................................................................................... 21
1.3.8 Stratified Sampling ....................................................................................... 22
1.3.9 Multi-Stage Sampling ................................................................................... 22
1.3.10 Cluster Sampling .......................................................................................... 23
1.3.11 Steps in Planning a Sample Survey .............................................................. 23
1.4 Sampling Distribution of the Sample Mean ........................................................ 24
Summary of Study Session 1 .......................................................................................... 26
Self-Assessment Questions (SAQs) for Study Session 1 ............................................... 27
Glossary of Terms........................................................................................................... 28
Page 8 of 174
STA202: Statistics II
References....................................................................................................................... 29
Study Session 2: Simple Correlation Analysis ................................................................... 30
Introduction..................................................................................................................... 30
Learning Outcomes for Study Session 2 ........................................................................ 30
2.1 Correlation Analysis ............................................................................................ 31
2.1.1 Karl Pearson’s’ Product Moment Correlation Coefficient ........................... 31
2.1.2 Spearman Rank Correlation Coefficient....................................................... 35
2.2 Tie in Ranks ......................................................................................................... 37
Summary of Study Session 2 .......................................................................................... 39
Self-Assessment Questions (SAQs) for Study Session 2 ............................................... 40
Glossary of Terms........................................................................................................... 42
References....................................................................................................................... 43
Study Session 3: Simple Regression Analysis ................................................................. 44
Introduction..................................................................................................................... 44
Learning Outcomes for Study Session 3 ........................................................................ 44
3.1 Simple Regression ............................................................................................... 45
3.2 The Least Squares Method .................................................................................. 45
Summary of Study Session 3 .......................................................................................... 51
Self-Assessment Questions (SAQs) for Study Session 3 ............................................... 52
Glossary of Terms........................................................................................................... 53
References....................................................................................................................... 54
Study Session 4: Test of Hypothesis ................................................................................ 55
Introduction..................................................................................................................... 55
Learning Outcomes for Study Session 4 ........................................................................ 55
4.1 Meaning of Test of Hypothesis ............................................................................ 56
4.2 Type I and Type II Errors .................................................................................... 57
4.2.1 One and Two Tailed Test ............................................................................. 57
4.2.2 Test Procedure and Steps .............................................................................. 58
4.2.3 Test Concerning the Mean (For Large Sample) ........................................... 58
4.2.4 Test Concerning Means (Small Samples) .................................................... 60
4.2.5 Test Concerning Two Population Means (Large Sample) ........................... 64
4.2.6 Test Statistics ................................................................................................ 66
Page 9 of 174
STA202: Statistics II
4..2.7 Test Concerning Two Population Means (Small Sample) ........................... 66
Summary of Study Session 4 .......................................................................................... 70
Self-Assessment Questions (SAQs) for Study Session 4 ............................................... 71
Glossary of Terms........................................................................................................... 72
References....................................................................................................................... 73
Study Session 5: Index Numbers ..................................................................................... 74
Introduction..................................................................................................................... 74
Learning Outcomes for Study Session 5 ........................................................................ 74
5.1 Meaning of Index Numbers ................................................................................. 75
5.1.1 Features of Index Numbers........................................................................... 75
5.1.2 Steps or Problems in the Construction of Price Index Numbers .................. 76
5.2 Construction of Price Index Numbers ................................................................. 78
5.2.1 Simple Aggregative Method......................................................................... 79
5.2.2 Simple Average of Price Relatives Method ................................................. 79
5.2.3 Weighted Aggregative Method .................................................................... 80
5.2.4 Weighted Average of Relatives Method ...................................................... 82
5.3 Difficulties in Measuring Changes in Value of Money ....................................... 84
5.3.1 Conceptual Difficulties ................................................................................. 84
Practical Difficulties ................................................................................................... 85
5.4 Types of Index Numbers...................................................................................... 87
5.4.1 Wholesale Price Index Numbers .................................................................. 88
5.4.2 Retail Price Index Numbers.......................................................................... 88
5.4.3 Cost-of-Living Index Numbers .................................................................... 88
5.4.4 Working Class Cost-of-Living Index Numbers ........................................... 89
5.4.5 Wage Index Numbers ................................................................................... 89
5.4.6 Industrial Index Numbers ............................................................................. 89
5.4.7 Uses of Index Number in the Economic Field ............................................. 89
5.4.8 Limitations of Index Numbers...................................................................... 90
Summary of Study Session 5 .......................................................................................... 91
Self-Assessment Questions (SAQs) for Study Session 5 ............................................... 92
Glossary of Terms........................................................................................................... 93
References...................................................................................................................... 94
Page 10 of 174
STA202: Statistics II
Study Session 6: Time Series Analysis ............................................................................ 95
Introduction..................................................................................................................... 95
Learning Outcomes for Study Session 6 ........................................................................ 95
6.1 Definition of Time Series .................................................................................... 96
6.1.1 Components of Time Series ......................................................................... 96
6.1.2 Time Series Model ....................................................................................... 96
6.2 Methods of Estimating Trend .............................................................................. 97
6.2.1 Moving Average Method ............................................................................. 98
6.2.2 Least Square Method .................................................................................. 101
6.3 Mathematical Representation ............................................................................ 101
Summary of Study Session 6 ........................................................................................ 105
Self-Assessment Questions (SAQs) for Study Session 6 ............................................. 106
Glossary of Terms......................................................................................................... 107
References..................................................................................................................... 108
Study Session 7: Multiple Regression .............................................................................. 109
Introduction................................................................................................................... 109
Learning Outcomes for Study Session 7 ...................................................................... 109
7.1 Introduction............................................................................................................. 110
7.1.1 First-Order Model with More than Two Predictor Variables ..................... 111
7.1.2 General Linear Regression Model .............................................................. 112
7.1.3 Qualitative Predictor Variables .................................................................. 113
7.1.4 Polynomial Regression ............................................................................... 113
7.1.5 Variables Transformation ........................................................................... 114
7.1.6 General Linear Regression Model in Matrix Terms................................... 115
7.2 Estimation of Regression Coefficients .............................................................. 117
7.2.1 Global Test: Testing Whether the Multiple Regression Model is Valid .... 120
7.2.2 Evaluating Individual Regression Coefficients .......................................... 121
7.2.3 Testing Individual Regression Coefficients ............................................... 122
7.2.4 Qualitative Independent Variables (Dummy Variables) ............................ 122
7.2.5 Dummy Variable ........................................................................................ 123
7.3 Polynomial Regression Model ............................................................................. 128
7.3.1 Uses of Polynomial Models........................................................................ 129
Page 11 of 174
STA202: Statistics II
7.3.2 One Predictor Variable-Second Order ........................................................ 129
7.3.3 One Predictor Variable-Third Order........................................................... 130
7.3.4 Two Predictor Variables-Second Order ..................................................... 130
7.3.5 Fitting Polynomial in One Variable ........................................................... 131
7.4 Analysis ............................................................................................................. 132
7.4.1 Test of Significance: ................................................................................... 135
Summary of Study Session 7 ........................................................................................ 138
Self-Assessment Questions (SAQs) for Study Session 7 ............................................. 139
Glossary of Terms......................................................................................................... 140
References..................................................................................................................... 141
Study Session 8: Partial Correlation .............................................................................. 142
Introduction................................................................................................................... 142
Learning Outcomes for Study Session 8 ...................................................................... 142
8.1 Simple Correlation Coefficient ............................................................................... 143
8.1.2 The Multiple Regression ............................................................................ 144
8.2 Partial Correlation Coefficient ........................................................................... 147
8.3 Tests for Association ......................................................................................... 150
8.3.1 Contingency Tables .................................................................................... 151
8.3.2 Expected Frequencies ................................................................................. 151
8.3.3 Expected Frequencies in Contingency Tables ............................................ 152
8.3.4 Test Statistic ............................................................................................... 153
8.3.5 𝝌𝟐 Test of Association .............................................................................. 153
8.3.6 Performing the Test ................................................................................... 154
8.3.7 Goodness-of-Fit Tests ................................................................................ 156
8.3.8 Observed and Expected Frequencies .......................................................... 157
8.3.9 Expected Frequencies in Goodness-of-Fit Tests ........................................ 157
8.3.10 The Goodness-of-Fit Test ........................................................................... 158
Summary of Study Session 8 ........................................................................................ 162
Self-Assessment Questions (SAQs) for Study Session 8 ............................................. 163
Glossary of Terms......................................................................................................... 164
References..................................................................................................................... 165
Notes on Self-Assessment Questions (SAQs) .................................................................. 166
Page 12 of 174
STA202: Statistics II
Notes on Self-Assessment Questions for Study Session 1 ........................................... 166
Notes on Self-Assessment Questions for Study Session 2 ........................................... 167
Notes on Self-Assessment Questions for Study Session 3 ........................................... 167
Notes on Self-Assessment Questions for Study Session 4 ........................................... 169
Notes on Self-Assessment Questions for Study Session 5 ........................................... 171
Notes on Self-Assessment Questions for Study Session 6 ........................................... 172
Page 13 of 174
STA202: Statistics II
Study Session 1: Sampling Theory and Estimation Techniques
Introduction
Each day, we observe the high, low, and close of stock market indexes from
around the world. Indexes such as Dangote Cement Index, Nestle Nigeria Index
and Stanbic IBTC Holdings Index are samples of stocks. Although Dangote
Cement, Nestle Nigeria and Stanbic IBTC Holdings do not represent the populations of
Nigerian stocks, we view them as valid indicators of the whole population’s behavior. As
analysts, we are accustomed to using this sample information to assess how various
markets from around the world are performing. Any statistics that we compute with
sample information, however, are only estimates of the underlying population parameters.
A sample, then, is a subset of the population—a subset studied to infer conclusions about
the population itself.
Page 14 of 174
STA202: Statistics II
1.1 Definition of Sampling
A sample survey can be defined as the collection and examination of data from a sample
in order to make inference about the whole or entire population. This is different from
census, which is a complete enumeration or survey of the whole units of enquiry. Hence,
sample survey theory deals with the process of sample selection, data collection,
estimation of the population characteristics using the sample data so collected and
determining the accuracy of the estimates.
i. Sample survey save time, labour and cost, especially when these resources are
limited than in census where a large number of these resources is needed.
Page 15 of 174
STA202: Statistics II
ii. Data can be collected and analysed quickly from a sample than from the entire
population. In a destructive investigation, such as determination of the
effectiveness or a newly produced drug by a pharmaceutical company, it is better
to use sample survey rather than complete enumeration.
iii. There is a greater and more efficient supervision of field staff in a sample survey
resulting in the collection of more reliable data. Sample survey makes for the use
of better-qualified staff and specialized equipment.
iv. It has greater subject coverage and less observational error than census. It is
possible to estimates the discrepancies between the sample and population values
in a sample survey.
v. Sample survey is used in conjunction with census in post enumeration checks in
order to determine the census coverage and accuracy of the census data.
i. Sample survey is not appropriate when information is required for every unit of
enquiry for example, compilation of voters list.
ii. It is less accurate in small area classification where data are needed for each
subdivision of the population.
iii. Sample survey cannot be used where the sampling frame is inadequate or
completely absent.
iv. It breeds sampling error which if not properly handled, may affect the accuracy of
the result. Finally, sample survey can be manipulated to suit the purpose of the
investigator.
A sampling frame contains the basic details of all members of the population from which
samples are to be drawn. It is generally believed by experts that without a complete
sampling frame, a truly random sample cannot be selected. Properties of a good frame are
completeness, no duplication, adequacy and clear identifiability of elements.
Page 16 of 174
STA202: Statistics II
In-Text Questions (ITQs)
There are two types of sampling method, these are sampling with replacement and
sampling without replacement.
sampling with
replacement
sampling
without
replacement
Sampling Method
Page 17 of 174
STA202: Statistics II
Fig 1.1: Types of Sampling Method
This is method of sampling where by a sample drawn is replaced before the subsequent.
This can occur where each member may be chosen more than once.
These can also occur when a sample element in a set of population is drawn for the first
time and it cannot be drawn again or cannot occur again.
There are two main types of sampling, namely; Probability (random) sampling and non-
probability (non-random) sampling. In probability sampling, every unit of enquiry in the
population has a known non-zero probability of being included in the sample. In non-
probability, the probability of selecting a unit in the sample is not known and cannot be
determined. Hence, statistical inference cannot be made objectively about the population
in non-random sampling. Hence we shall discuss briefly some examples of non-
Page 18 of 174
STA202: Statistics II
probability sampling and probability sampling respectively.
non-
probability
sampling
Sampling
Technique
Judgement
sampling
Stratified Quota
Sampling Sampling
Systematic Haphazard
Sampling Sampling
Non-
probability
sampling
Simple
Multi Stage
Random
Sampling
Sampling
Probability Cluster
Sampling Sampling
Page 19 of 174
STA202: Statistics II
1.3.2 Judgment Sampling
In this type of non- probability sampling experts pick units they judge to be representative
of the entire population. For example, a typical town may be picked to represent an urban
population or a village chosen to represent a rural population. Since experts differ in their
judgements, different experts may choose different units, which to them are more
preferable. Confidence is therefore not placed on the results from judgement sampling,
since there is no objective method for preferring one judgement to another. Finally,
accurate measure of sample variability cannot be obtained with judgement sampling.
In this type of non-random sampling, units are taken into the sample as they come along or
make themselves available. In a hospital survey on the characteristics of patients suffering
from a particular disease, the patients are taken into the sample in the order they report to
the hospital and volunteer to be part of the study. This sampling method lacks
representativeness of the population to be studied.
With probability sampling, each population’s element has a known and non-zero
probability of being selected. Samples of probability sampling are:
Page 20 of 174
STA202: Statistics II
1.3.6 Simple Random Sampling (SRS)
This is sampling procedure in which every member of the population has equal chance of
being selected as a member of the sample. It is mostly adopted in homogeneous (of same
kind) population. For example, if we have five coloured balls, red, orange, yellow, blue
and green, the chance of any of them being picked or selected is fraction of one and five,
that is, 1/5. Other examples are tossing of coin or rolling a die etc. All the above examples
are however, not perfectly objective. The most objective and widely used method is the
random number table. This is a table consisting of randomly allocated numbers that are
arranged in rows and columns of a standard statistical table.
Systematic sampling is another method of selecting a sample in such a way that every unit
in the population will have an equal chance of being selected in the sample. In this
selection procedure once the first unit is randomly selected the remaining units are
automatically selected.
Suppose a sample of n units is to be selected from N units in the population.
Let
N
K , k is an integer.
n
Select a random number between 1 and k inclusive. Suppose the random number selected
is r; add k to the random and select r successively until n numbers are obtained. The
sample data then consists of n units with serial numbers r, r+k, r+2k, …, r+(n-1)k. Thus
the sample consists of the first unit selected at random and every Kth unit thereafter.
This procedure of sample selection is called systematic sampling and the sample is called
1 n
systematic sample. K is called the sampling interval; is the sampling fraction.
K N
One of the advantages of systematic sampling is that it is easy and more convenient to
apply than simple random sampling without replacement (SRSWOR), especially in a large
scale sampling.
Page 21 of 174
STA202: Statistics II
In systematic sampling it is not necessary to serially number the units before selection or
even know the population size exactly in advance.
Application of intervals is easier to cross-check than the use of random numbers especially
if the selection is done in the field by the field enumerators. This helps to control the
field workers and minimize errors in sampling.
Systematic sample may be as precise as a simple random sample without replacement in a
population in which units are in random order. Systematic sampling yields an evenly
distributed sample.
This is a sampling procedure which consists of stratifying (dividing) the population into a
number of non-overlapping sub-populations (strata), then taking a sample from each
stratum. The items or sample from each stratum can then be selected by any suitable
random method.
Stratification does not guarantee good results, but if successful and properly executed, a
stratified sample will generally lead to a higher degree of precision, or reliability, then a
simple random sample of the same size drawn from the population.
This sampling procedure involves more than one stage. The first stage consists of breaking
down the population into sets of distinct groups, from these, a number of groups are
selected. Each group selected broken down into units from which a sample is taken. If we
stop at this point we have a two-stage sampling.
Further stages may be added and the number of stages involved is denoted in the name of
the sampling. For example, five stage sampling, means that five stages are involved. The
population is distributed into a number of first stage sampling units and a sample is taken
of these stage unit by some suitable method.
Page 22 of 174
STA202: Statistics II
1.3.10 Cluster Sampling
All sampling procedures so far discussed depend heavily on complete sampling frame.
Unfortunately, complete sampling frame is not always available. Cluster sampling was
developed to take care of this inadequacy. In this kind of sampling, the total population is
divided into a number of relatively small subdivisions, which are themselves clusters of
still smaller units, and then some of these sub-divisions or clusters, are randomly selected
for inclusion in the overall sample. If the clusters are geographical subdivisions this kind
of sampling is called area sampling and a cluster can be a household with members of the
household as elements.
Page 23 of 174
STA202: Statistics II
In-Text Questions (ITQs)
The sample mean is referred to as the point estimate of the population mean. The
arithmetic mean (𝜇𝑥̅ ) of the sampling distribution of mean values is equal to the
population mean (𝜇) regardless of the form of the population distribution that is, 𝜇𝑥̅ = 𝜇.
The sample mean is then said to be an unbiased estimate of the population mean.
0 + 2 + 4 + 6 12
𝜇= = =3
4 4
Page 24 of 174
STA202: Statistics II
The possible samples of size 2 and their mean from the population is
Sample Number Sample elements Sample mean
1 0,2 1
2 0,4 2
3 0,6 3
4 2,4 3
5 2,6 4
6 4,6 5
Table 1.1
The arithmetic mean of the sampling distribution of the mean value is
1+2+3+3+4+5
𝜇𝑥̅ = =3
6
Since 𝜇𝑥̅ = 𝜇, the sample mean is then an unbiased estimate of the population mean.
Page 25 of 174
STA202: Statistics II
Page 26 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 1
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
Page 27 of 174
STA202: Statistics II
Glossary of Terms
Page 28 of 174
STA202: Statistics II
References
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 29 of 174
STA202: Statistics II
Study Session 2: Simple Correlation Analysis
Introduction
We know that when the price of a product increases its demand will decrease. And
to the contrary quality supplied will increase with the increase of price. These are
nothing but correlation. Correlation quantifies the extent to which two
quantitative variables, X and Y, “go together.” When high values of X are associated
with high values of Y, a positive correlation exists. When high values of X are associated
with low values of Y, a negative correlation exists.
Page 30 of 174
STA202: Statistics II
2.1 Correlation Analysis
Spearman’s Karl
rank Pearson’s’
correlation product
coefficient moment
(R) correlation
coefficient
(r)
The Karl Pearson’s product moment correlation coefficient is devoted by r and given by:
Page 31 of 174
STA202: Statistics II
n xy x y
r
[n x 2 ( x) 2 ][n y 2 ( y ) 2 ]
It should be noted that the higher the magnitude of r, the stronger the association.
The table below is used to present Nigerian Government income (x) and Nigerian
Government expenditure (y) for a period of 12 months. This is given in million naira.
Table 2.1
Page 32 of 174
STA202: Statistics II
Calculate the coefficient of correlation
Solution
x y
xy - n
r
( y ) 2
[ x - ( x ) [ y -
2 2 2
138.00 x138.25
1608.12
r 12
2
(138.00) (138.25)
(1632.75 )(1607.81
12 12
Page 33 of 174
STA202: Statistics II
YZ 1.0 18.0 10.0 45.0 3.0 18.0 21.7 32.8 12.5 37.8 199.80
XZ 2.0 21.0 12.0 50.0 4.2 7.5 22.5 30.4 15.0 36.0 218.30
Table 2.2
Solution
n n
101.69
32.5(28.4)
10
32.52 28.4 2
[113.31 93.06
10 10
9.39
0.9620
9.76
(x(z )
xz n
rxz
2 x 2 .z 2
2
x z
n n
218.30
32.5(61.0)
10
32.5 2 612
113.31 457
10 10
20.05
0.7849
25.5436
Page 34 of 174
STA202: Statistics II
ryz 199.80
28.461.0
10 26.56
0.8185
28.4 2 612 32.45
93.06 457
10 10
Correlation coefficients are higher. It can be concluded that (i) Government income and
inflation are positively correlated. (ii) Government income and inflation are positively
correlated with GDP.
When variables do not follow normal distribution and one desires to assess the
relationship, correlation coefficient known as spearman rank correlation coefficient is
used. The variables are ranked based on the magnitude. The correlation between ranks of
variables x and y is obtained. The symbol used is R, the formula is:
6 d i2
R 1
n n 2 1
where d is the difference between ranks given to the variables of each pair and n is the
number of pairs studied. The procedure was developed by spearman. Hence, it is known as
spearman rank correlation coefficient. Its value also ranges from – 1 to 1.
Calculate the value correlation coefficient between the corresponding values of income
(X) and expenditure (Y) of a known company given below.
X 22 24 25 16 28 19
Y 48 42 40 38 47 45
Table 2.3
Solution
Page 35 of 174
STA202: Statistics II
The varying is in ascending order of magnitude
X Y RX RY d d2
22 48 3 6 -3 9
24 42 4 3 1 1
25 40 5 2 3 9
16 38 1 1 0 0
28 47 6 5 1 1
19 45 2 4 -2 4
24
Table 2.4
6 d 2
R = 1-
n(n 2 1)
6(24)
=1- = 1 - 0.6857 =0.3143
6(36 1)
R = 0.31
Page 36 of 174
STA202: Statistics II
In-Text Questions (ITQs)
Most times, two or more values of a variable might be equal. In such cases, we assign to
each of the tied observations the mean of the ranks which they jointly occupy. For
example if the 5th and 6th largest values of a variable are equal, we assign to each the rank
(5 6)
=5.5, and if the of fifth, smith and seventh largest values of a variable are the
2
(5 6 7 )
same we assign each the rank =6.6
3
Study 2.4
The table give below shows the respective yearly money deposit in two known
commercial banks over the period of a year in billons
Father ( ) 66 64 68 65 69 63 71 67 69 68 70 72
Sons ( ) 69 67 69 66 70 67 69 66 72 68 69 71
Table 2.5
Page 37 of 174
STA202: Statistics II
Calculate the coefficient of rank correlation and comment on the degree of correlation
between the two known commercial banks.
Solution
RX RX D= RX - RY d2
66 69 4 7.5 -3.5 12.25
64 67 2 3.5 -1.5 2.25
68 69 6.5 7.5 -1.0 1.00
65 66 3 1.5 1.5 2.25
69 70 8.5 10 -1.5 2.25
63 67 1 3.5 -2.5 6.25
71 69 11 7.5 3.5 12.25
67 66 5 1.5 3.5 12.25
69 72 8.5 1.5 -3.5 12.25
68 68 6.5 1.2 1.5 2.25
70 69 10 5 2.5 6.25
72 71 12 11 1.0 1.00
72.50
Table 2.6
6d 2
R=1-
n(n 2 1)
6(72.50)
=1-
12(144 1
72.50
=1-
2(143)
=1-0.2535
=0.7465
=0.75
Comment: There is a fairly high positive correlation between the money deposit of the tow
known commercial bank.
Page 38 of 174
STA202: Statistics II
1. The degree of relationship may be positive that is, an increase in one variable
accompanied by an increase in the other or negative when decrease in one variable
is accompanied by an increase in the other.
2. The patterns of correlation are perfect and positive correlation when r = 1, perfect
and negative correlation when r = - 1, positive correlation when r > 0, negative
correlation when r <0 and no correlation when r = 0.
3. Two types of the measures of correlation are:
i. Karl Pearson’s’ product moment correlation coefficient (r)
ii. Spearman’s rank correlation coefficient (R)
4. The Karl Pearson’s product moment correlation coefficient is devoted by r and
given by:
n xy x y
r
[n x 2 ( x) 2 ][n y 2 ( y ) 2 ]
6 d i2
R 1
5. Spearman rank correlation coefficient formula is
n n 2 1
6. Tie in ranks is settled by assigning to each of the tied observations the mean of the
ranks which they jointly occupy.
Page 39 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 2
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
Calculate the product moment correlation coefficient for the pair given below and
interpret your result.
1.
3. The table give below shows the respective income (X) and expenditure (Y) (in million
naira) of O.O.U, Ago-Iwoye for a year.
Page 40 of 174
STA202: Statistics II
Income ( ) 66 64 68 65 69 63 71 67 69 68 70 72
Expenditure ( ) 69 67 69 66 70 67 69 66 72 68 69 71
Calculate the coefficient of rank correlation and comment on the degree of correlation
between income (X) and expenditure (Y) (in million naira) of O.O.U, Ago-Iwoye.
Page 41 of 174
STA202: Statistics II
Glossary of Terms
Coefficient: a numerical or constant quantity placed before and multiplying the variable
in an algebraic expression
Linear: Arranged in or extending along a straight or nearly straight line
Magnitude: Size
Correlation analysis: statistical method that is used to discover if there is a relationship
between two variables/datasets, and how strong that relationship may be
Quantitative variables: any variables where the data represent amounts (e.g. height,
weight, or age)
Page 42 of 174
STA202: Statistics II
References
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 43 of 174
STA202: Statistics II
Study Session 3: Simple Regression Analysis
Introduction
Page 44 of 174
STA202: Statistics II
3.1 Simple Regression
When only two variables are involved, the regression is said to be simple. A simple linear
regression equation is therefore of the form = + X , once and are estimated,
we can substitute a given value of X into the equation and calculate the predicted value of
.
This is the most reliable of all the methods used to find regression lines. It leads to unique
regression line and regression coefficient. The least square method could be used to
estimate the parameters and from the model.
Yi X i i
as follow:
i2 (Yi x i ) 2 i 1, 2, ..., n
Minimized the function L with respect to and by taking the partial derivatives
L
2 (Yi X i )
Page 45 of 174
STA202: Statistics II
L
2 (Yi xi )
Set these partial derivatives equal to zero and solve for and , we obtain:
n xy x ( y)
=
n x 2 ( x ) 2
and
Y x
i i
where x
n
xi and Y Yi
n
Case Study 3.1
A study was made on the effect of income level on the standard of living. The following
data was obtained in coded form. Calculate the regression of standard of living on income
level.
Page 46 of 174
STA202: Statistics II
Solution
y x
xy
xy n
=
x x2 / n
2
Y x
x 0, y 102, x 2
110, xy 158
0 102
158
11 1.44
(0) 2
110
11
0
x 0
11
102
Y 9.27,
11
y 9.27 1.44 x
=9.27+1.44 (6)
Page 47 of 174
STA202: Statistics II
Case Study 3.2
The table below shows the Nigeria gross domestic product (X) and inflation rate (Y) over
a period of time.
X 65 63 67 64 68 62 70 66 68 67 69 71
Y 68 66 68 65 69 66 68 65 71 67 68 70
Table 3.2
Solution
2 Y2
Page 48 of 174
STA202: Statistics II
66 65 4356 4225 4290
Table 3.3
n xx x y
n x 2 ( x ) 2
12(54107) (800)(811)
12(53418) (800) 2
0.4764
yx
y x
n n
811 800
0.4764 x = 35.8233
12 12
x = Y
Page 49 of 174
STA202: Statistics II
n xy x y
n y 2 ( y ) 2
12(54107) (800)(811)
12(54849) (811) 2
1.036
3.38
Y = - 3.38 + 1.036Y
Page 50 of 174
STA202: Statistics II
Page 51 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 3
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
1. Explain the term Simple Regression analysis
2. The table give below shows the respective income (X) and expenditure (Y) (in million
naira) of O.O.U, Ago-Iwoye for a year.
Income ( ) 66 64 68 65 69 63 71 67 69 68 70 72
Expenditure ( ) 69 67 69 66 70 67 69 66 72 68 69 71
Page 52 of 174
STA202: Statistics II
Glossary of Terms
Parameter: a value that tells you something about a population and is the opposite from
a statistic,
Partial derivative: a function of several variables is its derivative with respect to one of
those variables, with the others held constant.
Regression Line: a single line that best fits the data (in terms of having the smallest
overall distance from the line to the points
Regression Coefficient: estimates of the unknown population parameters and describe the
relationship between a predictor variable and the response
Independent Variable: a variable that stands alone and isn't changed by the other
variables you are trying to measure
Linear Relationship: a straight-line relationship between two variables
Page 53 of 174
STA202: Statistics II
References
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 54 of 174
STA202: Statistics II
Study Session 4: Test of Hypothesis
Introduction
Page 55 of 174
STA202: Statistics II
4.1 Meaning of Test of Hypothesis
The most frequent application of statistics is to test some scientific hypotheses. Results of
experiments and investigations are usually not clear cut and, therefore, need statistical
tests to support decisions between alternative hypotheses. A statistical test examines a set
of sample data and on the basis of an expected distribution of the data, leads to a decision
on whether to accept the hypothesis or whether to reject that hypothesis and accept an
alternative one. The nature of the tests varies with the data and the hypothesis, but the
same general philosophy of hypothesis testing is common to all tests.
A statistical hypothesis is an assumption or statement which may or may not be true
concerning one or more population. A statistical hypothesis (or inference) is a statement
about the parameters or form of a population. A test of a statistical hypothesis is a criteria
which specifies for what sample results the hypothesis is to be accepted or rejected. The
hypothesis which is to be tested is generally called the Null hypothesis denoted by the H0
and hypothesis against which it is to be tested is called the alternative hypothesis and also
denoted by H1.
i. A statistical test is a test that examines a set of sample data and on the basis of an
expected distribution of the data, leads to a decision on whether to accept the
hypothesis or whether to reject that hypothesis and accept an alternative one.
ii. A statistical hypothesis is an assumption or statement which may or may not be
true concerning one or more population.
Page 56 of 174
STA202: Statistics II
4.2 Type I and Type II Errors
A type I error has been committed if we reject the null hypothesis when it is true and a
type II error has been committed if we accept the null hypothesis when it is false.
The following table summarizes the various situations that can arise when testing H0
against H1:
Accept H0 Accept H1
H0 is true No error Type I Error
H1 is true Type II error No error
Table 4.1
The probabilities of committing a type I and type II errors are called level of significance
of the tests and are written as and , respectively. is called the size of the test and (1-
) is called the power of the test, and (1-) is also the probability of rejecting null
hypothesis (H0) when it is false. The area such that if the sample point falls in it we reject
H0 is called the critical region. When the primary concern of a test is to see whether the
null hypothesis can be rejected, such a test is called a test of significance. In that case, the
quantity is called the level of significance at which the test is being conducted.
A test of any statistical hypothesis where the alternative is one sided such as:
H0: = 0 or H0: = 0
H1: > 0 H1: < 0
Is called a one-tailed test. The critical region for H1: > 0 lies entirely in the right tail
while the critical region for H1: < 0 lies entirely in the left tail.
A test of any statistical hypothesis where the alternative is two-sided such as:
H0: = 0
H1: 0
Page 57 of 174
STA202: Statistics II
Is called a two-tailed test, values in the both tails of the distribution constitute the critical
region.
The steps involved in general and in the utilization of any test of significance are:
We will assume that the sampling distribution of the sample estimates will be
approximately normal and that the variance is known. Hence, for large samples (n 30),
we can use the normal probability distribution for testing a hypothesized value of the
population mean.
X
Z
S .E. X
Page 58 of 174
STA202: Statistics II
is the population mean
S .E. X
n
where is the population standard deviation (usually known) and n is the sample size.
A bottling company which bottles a soft drink claims that the liquids content is 35cl with
standard deviation 0.75cl. A researcher randomly collects 50 bottles, measured their
contents and got mean of 34.2cl. Test at 0.01 level of significance that the bottling
company has been cheating their consumers.
Solution
= 35cl
= 0.75cl
n = 50
X = 34.2
= 0.01 (1%)
H0: = 35 that is, the company has not been cheating the consumers.
H1: < 35 that the company has been cheating the consumers.
Test statistics is
Z
X n
Page 59 of 174
STA202: Statistics II
34.2 35 50
0.75
0.8 7.0711
0.75
= -7.54
Decision: the Z calculated value 7.54 is greater than the Z tabulated value 2.33. we reject
H0 and accept H1.
Conclusion: There is significant difference between the population and sample mean.
Hence, the bottling company has been cheating their consumers.
There are situations in real life experiment, such as, testing the efficiency of a newly
produced drug, where it is impracticable to get a large sample and yet tests of significance
still have to be carried out. When we do not know the value of the population standard
deviation and the sample size is small (n < 30), we shall assume again that the population
we are sampling from has roughly the shape of a normal distribution. The test statistics is:
t
X
X n
S S
n
Whose sampling distribution is the t distribution with n-1 degree of freedom. S is the
sample standard deviation. As with large samples, we compare it with its value at a given
level of significance, and then draw our conclusions.
Page 60 of 174
STA202: Statistics II
Case Study 4.2
Suppose that we want to test on the basis of a random sample of size n = 5 whether or not
the fat content of a certain kind of ice cream exceeds 12 percent. What can we conclude
about the null hypothesis. = 12 percent at the 0.01 level of significance, if the sample
has the mean X as 12.7 percent and the standard deviation S is 0.38 percent.
Solution
H1: > 12
= 0.01
Test statistics
X
t
S
n
12.7 12
t
0.38
5
0. 7
t 4.12
0.1699
t0.01,4 = 4.12
Conclusion: Therefore, the content of the given kind of ice cream exceeds 12 percent.
Page 61 of 174
STA202: Statistics II
Case Study 4.3
The life time of telephone for a random sampling 10 from a large consignment give the
following data:
8 4.3 0 0
Page 62 of 174
STA202: Statistics II
Table 4.2
Can we accept the hypothesis that the average life time of telephone is 4,000hours at 5%
level of significance?
Solution
Hypothesis
H0: = 4,000hours
H1: 4,000hours
X i
4.2 4.0 ,..., 5.6 43.5
X i 1
n 10 10
X 4.3
X X
10
2
i
i 1
S2
n 1
S2
4.2 4.32 4.0 4.32 ... 4.4 4.3 5.6 4.3
2 2
10 1
Page 63 of 174
STA202: Statistics II
3.22
S2 = 0.358
9
Test Statistics
t
X n
S
t
4.3 4 10
0.598
where
S 0.358 = 0.598
t = 1.587
t0.025,9 = 2.262
Conclusion: Since tcal > ttab, then we accept H0 and conclude that the average life time is
4,000hours
The test statistics for large sample test concerning difference between two means is given
as:
X1 X 2
Z
12 22
n1 n2
Page 64 of 174
STA202: Statistics II
Case Study 4.5
In a study designed to test whether or not there is a difference between the average amount
used to buy food by families living in two different communities, random samples yield
the following results.
Table 4.3
The amount used to buy food are in Thousand Naira. Use the 0.05 level of significance to
test the null hypothesis that the corresponding population means are equal against the
alternative hypothesis that they are not equal.
Solution
H 0 1 2
H 1 1 2
Page 65 of 174
STA202: Statistics II
4.2.6 Test Statistics
X1 X 2
Z
12 22
n1 n2
62.7 61.8
Z
2.50 2 2.62 2
120 150
0.9
Z
0.0979
Z = 2.88
Conclusion: Since Z cal Z tab , the null hypothesis must be rejected and we conclude that
there is a difference between the true average heights of adult females in the two given
communities.
The test statistics for small sample test concerning difference between two means is given
as:
X1 X 2
t
Sp 2 1
n2 1
n2
where
Page 66 of 174
STA202: Statistics II
ii. The two populations have the same variance.
iii. The two samples are random ones.
Case Study 4.6
The following random samples are amount used by two states in Nigeria to provide health
facilities (in millions naira) for five months:
Table 4.4
Use 0.05 level of significance to test whether the difference between the means of these
two samples is significant.
Solution
X i
8400 8230 ... 7930
X1 i 1
5 5
X 1 8160
X i
7510 7690 ... 7660
X2 i 1
5 5
X 2 7730
X Xi
5
2
i
i 1
S 21
n1 1
Page 67 of 174
STA202: Statistics II
S 63450
1
2
X X2
5
2
i
i 1
S 22
n2 1
S 22 42650
H 0 1 2
H 1 1 2
Test statistics
X1 X 2
t
n1 1S12 n2 1S 22
n1 n2 2 1
n1 1
n2
8160 7730 430
t =t
4 63450 4 42650 1
552 5
1
5
21220
Conclusion: Since tcal ttab , the null hypothesis should be rejected then, we conclude that
the average amount spend on health facilities from the states are not the same.
Page 68 of 174
STA202: Statistics II
In-Text Answers (ITAs)
i. A type I error has been committed if we reject the null hypothesis when it is true
while a type II error has been committed if we accept the null hypothesis when it
is false
Page 69 of 174
STA202: Statistics II
Page 70 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 4
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
i. Statistical hypothesis
3. The following random samples are the spending power of families (in millions
naira)
Mine 1 84 82 83 78 79
Mine 2 75 76 77 80 76
Use 0.05 level of significance to test whether the difference between the means of these
two samples is significant.
Page 71 of 174
STA202: Statistics II
Glossary of Terms
Page 72 of 174
STA202: Statistics II
References
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 73 of 174
STA202: Statistics II
Study Session 5: Index Numbers
Introduction
The value of money does not remain constant over time. It rises or falls and is
inversely related to the changes in the price level. A rise in the price level
means a fall in the value of money and a fall in the price level means a rise in the
value of money. Thus, changes in the value of money are reflected by the changes in the
general level of prices over a period of time. Changes in the general level of prices can be
measured by a statistical device known as ‘index number.’
Page 74 of 174
STA202: Statistics II
5.1 Meaning of Index Numbers
Page 75 of 174
STA202: Statistics II
compared to that in 1960 taken as the base year) or the levels of a phenomenon at
different places on the same date (e.g., the price level in India in 1980 in
comparison with that in other countries in 1980).
The construction of the price index numbers involves the following steps or problems
1. Selection of Base Year: The first step or the problem in preparing the index
numbers is the selection of the base year. The base year is defined as that year with
reference to which the price changes in other years are compared and expressed as
percentages. The base year should be a normal year. In other words, it should be
free from abnormal conditions like wars, famines, floods, political instability, etc.
Base year can be selected in two ways- (a) through fixed base method in which the
base year remains fixed; and (b) through chain base method in which the base year
goes on changing, e.g., for 1980 the base year will be 1979, for 1979 it will be
1978, and so on.
2. Selection of Commodities: The second problem in the construction of index
numbers is the selection of the commodities. Since all commodities cannot be
included, only representative commodities should be selected keeping in view the
purpose and type of the index number.
In selecting items, the following points are to be kept in mind
a. The items should be representative of the tastes, habits and customs of the people.
b. Items should be recognizable,
c. Items should be stable in quality over two different periods and places.
d. The economic and social importance of various items should be considered
e. The items should be fairly large in number.
f. All those varieties of a commodity which are in common use and are stable in
character should be included.
3. Collection of Prices: After selecting the commodities, the next problem is
regarding the collection of their prices:
a. From where the prices to be collected;
Page 76 of 174
STA202: Statistics II
b. Whether to choose wholesale prices or retail prices;
c. Whether to include taxes in the prices or not etc.
While collecting prices, the following points are to be noted
a. Prices are to be collected from those places where a particular commodity is traded
in large quantities.
b. Published information regarding the prices should also be utilised,
c. In selecting individuals and institutions who would supply price quotations, care
should be taken that they are not biased.
d. Selection of wholesale or retail prices depends upon the type of index number to
be prepared. Wholesale prices are used in the construction of general price index
and retail prices are used in the construction of cost-of-living index number.
e. Prices collected from various places should be averaged.
4. Selection of Average: Since the index numbers are specialized average, the fourth
problem is to choose a suitable average. Theoretically, geometric mean is the
best for this purpose. But, in practice, arithmetic mean is used because it is easier
to follow.
5. Selection of Weights: Generally, all the commodities included in the construction’
of index numbers are not of equal importance. Therefore, if the index numbers are
to be representative, proper weights should be assigned to the commodities
according to their relative importance. For example, the prices of books will be
given more weightage while preparing the cost-of-living index for teachers than
while preparing the cost-of-living index for the workers. Weights should be
unbiased and be rationally and not arbitrarily selected.
6. Purpose of Index Number: The most important consideration in the construction of
the index numbers is the objective of the index numbers. All other problems or
steps are to be viewed in the light of the purpose for which a particular index
number is to be prepared. Since, different index numbers are prepared with
specific purposes and no single index number is ‘all purpose’ index number, it is
important to be clear about the purpose of the index number before its
construction.
Page 77 of 174
STA202: Statistics II
7. Selection of Method: The selection of a suitable method for the construction of
index numbers is the final step.
There are two methods of computing the index numbers:
a. Simple index number and
b. Weighted index number.
Simple index number again can be constructed either by:
i. Simple aggregate method or
ii. Simple average of price relative’s method.
Similarly, weighted index number can be constructed by:
i. Weighted aggregative method or
ii. Weighted average of price relative’s method.
The choice of method depends upon the availability of data, degree of accuracy required
and the purpose of the study.
Construction of price index numbers through various methods can be understood with the
help of the following examples:
Page 78 of 174
STA202: Statistics II
5.2.1 Simple Aggregative Method
In this method, the index number is equal to the sum of prices for the year for which index
number is to be found divided by the sum of actual prices for the base year.
The formula for finding the index number through this method is as follows in case study
5.1 below as follows
Case Study 5.1
∑ 𝑃1
𝑃01 = × 100
∑ 𝑃0
Where 𝑃01 stands for index number
∑ 𝑃1 stands for the sum of the prices for the year for which index number is to be found
∑ 𝑐 stands for the sum prices for the base year
Commodity Prices in Base Year 1980 Prices in Current Year 1988
(In Naira) 𝑃0 ( in Naira) 𝑃1
A 10 20
B 15 25
C 40 60
D 25 40
Total ∑ 𝑃0 = 90 ∑ 𝑃1 = 145
Table 5.1
∑𝑃 145
Index Number (𝑃01 ) = ∑ 𝑃1 × 100; 𝑃01 = × 100; 𝑃01 = 161.11
0 90
In this method, the index number is equal to the sum of price relatives divided by the
number of items and is calculated by using the following formula in case study 5.2 as
follows
Page 79 of 174
STA202: Statistics II
Case Study 5.2
Commodity Base Year Prices Current Year Price Relative
𝑃0 𝑃1 Prices (in Naira) 𝑃1
𝑅= × 100
𝑃0
A 10 20 20
× 100 = 200.0
10
B 15 25 25
× 100 = 166.7
15
C 40 60 60
× 100 = 150.0
40
D 25 40 40
× 100 = 160.0
25
N=4 ∑ 𝑅 = 676.7
Table 5.2
∑𝑅
Index Number (𝑃01 ) = 𝑁
676.7
(𝑃01 ) = = 169.2
4
In this method, different weights are assigned to the items according to their relative
importance. Weights used are the quantity weights. Many formulae have been developed
to estimate index numbers on the basis of quantity weights.
Some of them are explained below with some examples
i. Laspeyre’s Formula: In this formula, the quantities of base year are accepted
as weights.
∑ 𝑃1 𝑞0
𝑃01 = × 100
∑ 𝑃0 𝑞0
Where 𝑃1 is the price in the current year; 𝑃0 is the price in the base year; and 𝑞0 is the
quantity of the base year.
ii. Paasche’s Formula: In this formula, the quantities of the current year are
accepted as weights.
Page 80 of 174
STA202: Statistics II
∑ 𝑃1 𝑞1
𝑃01 = × 100
∑ 𝑃0 𝑞1
Where 𝑞1 is the quantity in the current year.
iii. Dorbish and Bowley’s Formula: Dorbish and Bowley’s Formula for
estimating weighted index number is as follows
∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
+
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1 𝐿+𝑃
𝑃01 = × 100 or 𝑃01 =
2 2
∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
𝑃01 = √ + × 100 𝑜𝑟 𝑃01 = √𝐿 × 𝑃 = 100
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1
∑ 𝑃0 𝑞0 ∑ 𝑃1 𝑞0 ∑ 𝑃0 𝑞1 ∑ 𝑃1 𝑞1
Table 5.3
i. Laspeyre’s Formula:
∑ 𝑃1 𝑞0
𝑃01 = × 100
∑ 𝑃0 𝑞0
440
𝑃01 = × 100 = 166.04
265
ii. Paasche’ Formula:
Page 81 of 174
STA202: Statistics II
∑ 𝑃1 𝑞1
𝑃01 = × 100
∑ 𝑃0 𝑞1
700
𝑃01 = × 100 = 158.3
480
∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
𝑃01 = √ + × 100
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1
440 760
𝑃01 = √ + × 100 = 162.1
265 480
In this method also different weights are used for the items according to their relative
importance. The price index number is found out with the help of the following formula
and with examples.
∑ 𝑅𝑊
𝑃01 =
∑𝑊
Where ∑ 𝑊 stands for the sum of weights of different commodities and
∑ 𝑅 stands for the sum of price relatives
Page 82 of 174
STA202: Statistics II
In-Text Answers (ITAs)
Simple aggregated method, simple average of price relatives method, weighted average of
relatives method.
A 5 10 20 20 1000.0
× 100 =
10
200.0
B 4 15 25 25 666.8
× 100
15
= 166.7
C 2 40 60 60 300.0
× 100
40
= 150.0
D 3 25 40 40 480.0
× 100
25
= 160.0
Total ∑ 𝑊 = 14 ∑ 𝑅𝑊 = 2446.8
Table 5.4
∑ 𝑅𝑊
Index Number (𝑃01 ) = ∑𝑊
2446.8
𝑃01 = = 174.8
14
Page 83 of 174
STA202: Statistics II
5.3 Difficulties in Measuring Changes in Value of Money
Measurement of changes in the value of money through price index number is not an easy
and reliable technique. There are a number of theoretical as well as practical difficulties
in the construction of price index numbers. Moreover, the index number technique itself
has many limitations.
The following are the conceptual difficulties during the construction of price index
numbers:
1. Vague Concept of Value of Money: The concept of money is vague, abstract and
cannot be clearly defined. The value of money is a relative concept which changes
from person to person depending upon the type of goods on which the money is
spent.
2. Inaccurate Measurement: Price index numbers do not measure the changes in the
value of money accurately and reliably. A rise or fall in the general level of prices
as indicated by the price index numbers does not mean that the price of every
commodity has risen or fallen to the same extent.
Page 84 of 174
STA202: Statistics II
3. Reflect General Changes: Price index numbers are averages and measure general
changes in the value of money on the average. Therefore, they are not of much
significance for the particular individuals who may be affected by the changes in
the actual prices quite differently from that indicated by the index numbers.
4. Limitations of Wholesale Price Index: The wholesale price index numbers, which
are generally used to measure changes in the value of money, suffer from certain
limitations:
i. They do not reflect the changes in the cost of living because retail prices are
generally higher than the wholesale prices.
ii. They ignore some of the important items concerning the urban population, such as,
expenditure on education, transport, house rent, etc.
iii. They do not take into consideration the changes in the consumers’ preferences.
Practical Difficulties
The practical difficulties in the way of constructing price index numbers, and therefore, in
measuring changes in the value of money are as follows:
1. Selection of Base Year: While preparing the index number, first difficulty arises
regarding the selection of base year. The base year should be a normal year. But, it
is very difficult to find out a fully normal year free from any unusual happening.
There is every possibility that the selected base year may be an abnormal year, or a
distant year, or may be selected by an immature or biased person.
2. Selection of Items: The selection of the representative commodities is the second
difficulty in the construction of index numbers:
[Link] the passage of time the quality of the product may change; if the quality of a
product changes in the year of enquiry from what it was in the base year, the product
becomes irrelevant
[Link] relative importance of certain commodities may change due to a change in the
consumption pattern of the people in the course of time; for example, Vanaspati Ghee
was not an important item of consumption in India in the pre-war period, but today it
Page 85 of 174
STA202: Statistics II
has become an item of necessity. Under such conditions, it is not easy to select the
appropriate commodities.
2 Collection of Prices: It is also difficult to obtain correct, adequate and representative
data regarding prices. It is not an easy job to select representative places from which
the information about prices to be collected and to select the experienced and unbiased
individuals or institutions who will supply price quotations. Moreover, there is the
problem of deciding which prices (wholesale or retail) are to be taken into
consideration. It is comparatively easy to get information about wholesale prices
which vary considerably.
3 Assigning Weights: Another important difficulty that arises in preparing the index
numbers is that of assigning proper weights to different items in order to arrive at
correct and unbiased conclusions. As there are no hard and fast rules to weights for the
commodities according to their relative importance, there is very likelihood that the
weights are decided arbitrarily on the basis of personal judgement and involve
biasness.
4 Selection of Averages: Another major problem is that which average should be
employed to find out the price relatives. There are many types of averages such as
arithmetic average, geometric average, mean, median, mode, etc. The use of different
averages gives different results. Therefore, it is essential to select the method with
great care. Dr. Marshall has advocated the use of chain index number to solve the
problem of averaging and weighing.
5 Problem of Dynamic Changes: In the dynamic world, the consumption pattern of the
individuals and the number and varieties of goods undergo continuous changes. They
create difficulties for preparing index numbers and making temporal comparisons:
i. Since, in the course of time, old commodities may disappear and many new ones
come into existence, the long-run comparison may become difficult
ii. The quantity and quality of commodities may also change over the period of time,
thus making the choice of commodities for constructing index numbers difficult
iii. A number of factors, like income, education, fashion, etc., bring changes in the
consumption pattern of the people which render the index numbers incomparable.
Page 86 of 174
STA202: Statistics II
In-Text Questions (ITQs)
Index numbers are of different types. Important types of index numbers are discussed
below
Whole sale
price
Cost-of-
Retail price
Living price
Index
Number
Industrial Wage
Working
class Cost-
of-Living
Page 87 of 174
STA202: Statistics II
5.4.1 Wholesale Price Index Numbers
Wholesale price index numbers are constructed on the basis of the wholesale prices of
certain important commodities. The commodities included in preparing these index
numbers are mainly raw-materials and semi-finished goods. Only the most important and
most price-sensitive and semi- finished goods which are bought and sold in the wholesale
market are selected and weights are assigned in accordance with their relative importance.
The wholesale price index numbers are generally used to measure changes in the value of
money. The main problem with these index numbers is that they include only the
wholesale prices of raw materials and semi-finished goods and do not take into
consideration the retail prices of goods and services generally consumed by the common
man. Hence, the wholesale price index numbers do not reflect true and accurate changes in
the value of money.
These index numbers are prepared to measure the changes in the value of money on the
basis of the retail prices of final consumption goods. The main difficulty with this index
number is that the retail price for the same goods and for continuous periods is not
available. The retail prices represent larger and more frequent fluctuations as compared to
the wholesale prices.
These index numbers are constructed with reference to the important goods and services
which are consumed by common people. Since the number of these goods and services is
very large, only representative items which form the consumption pattern of the people are
included. These index numbers are used to measure changes in the cost of living of the
general public.
Page 88 of 174
STA202: Statistics II
5.4.4 Working Class Cost-of-Living Index Numbers
The working class cost-of-living index numbers aim at measuring changes in the cost of
living of workers. These index numbers are consumed on the basis of only those goods
and services which are generally consumed by the working class. The prices of these
goods and index numbers are of great importance to the workers because their wages are
adjusted according to these indices.
The purpose of these index numbers is to measure time to time changes in money wages.
These index numbers, when compared with the working class cost-of-living index
numbers, provide information regarding the changes in the real wages of the workers.
Industrial index numbers are constructed with an objective of measuring changes in the
industrial production. The production data of various industries are included in preparing
these index numbers.
Some of the specific uses of index numbers in the economic field are:
i. They are useful in analysing markets for specific commodities.
ii. In the share market, the index numbers can provide data about the trends in the
share prices
iii. With the help of index numbers, the Railways can get information about the
changes in goods traffic.
iv. The bankers can get information about the changes in deposits by means of index
numbers.
Page 89 of 174
STA202: Statistics II
5.4.8 Limitations of Index Numbers
Index number technique itself has certain limitations which have greatly reduced its
usefulness:
Wholesale price index number, retail price index numbers, cost-of-living index numbers,
working class cost-of-living index numbers and industrial index numbers.
Page 90 of 174
STA202: Statistics II
Page 91 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 5
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
Page 92 of 174
STA202: Statistics II
Glossary of Terms
Limitations: a restriction
Page 93 of 174
STA202: Statistics II
References
1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Page 94 of 174
STA202: Statistics II
Study Session 6: Time Series Analysis
Introduction
You may have heard people saying that the price of a particular commodity has
increased or decreased with time. This commodity can be anything like gold, silver,
any eatables, petrol, diesel etc. Also, you may have heard that the rate of interest has
increased in banks. The rate of interest for home loans has decreased. What are all these? How
are they useful to us? These types of data are the time series of data.
Page 95 of 174
STA202: Statistics II
6.1 Definition of Time Series
A time series is defined as some quantity that is measured sequentially in time over some
interval. In its broadest form, time series analysis is about inferring what has happened to
a series of data points in the past and attempting to predict what will happen to it the
future.
i. Trend is the long term pattern of a time series: A trend can be positive or negative
depending on whether the time series exhibits an increasing long term pattern or a
decreasing long term pattern
ii. Cyclical movement: Any pattern showing an up and down movement around a
given trend is identified as a cyclical pattern. The duration of a cycle depends on
the type of business or industry being analyzed.
iii. Seasonal movement: Many time series contain seasonal variation. This is
particularly true in series representing business sales or climate levels. In
quantitative finance we often see seasonal variation in commodities, particularly
those related to growing seasons or annual temperature variation (such as natural
gas).
iv. Irregular movement: This component is unpredictable. Every time series has some
unpredictable component that makes it a random variable.
i. Additive model: This is a model in which the series value of y is the sum of all
four components, that is,
y = T + S + C+ I
Page 96 of 174
STA202: Statistics II
ii. Multiplicative model: It is a model in which y is the product of all the time series
components. This is denoted by
y = T * S * C* I
The methods used to estimate trend in time series analysis are many but here we will
discuss the followings
i. Moving Average
ii. Least Square method
Page 97 of 174
STA202: Statistics II
Moving Average
Method Least Square
Method
1
𝑦1 = (𝑥 + 𝑥2 +, … , +𝑥𝑛 )
𝑛 1
1
𝑦2 = (𝑥 + 𝑥3 +, … , +𝑥𝑛 + 𝑥𝑛+1 )
𝑛 2
1
𝑦3 = (𝑥 + 𝑥4 +, … , +𝑥𝑛 + 𝑥𝑛+1 + 𝑥𝑛+2 ) 𝑎𝑛𝑑 𝑠𝑜 𝑜𝑛.
𝑛 3
These averages are called-point moving averages. The n-points to show that n
observations are used in the averages. The averages are moving because they are averages
of successive n observations.
Page 98 of 174
STA202: Statistics II
In-Text Questions (ITQs)
The following table gives the quarterly indices of rental prices of certain foodstuffs.
Quarter
Year 1 2 3 4
2004 119 127 127 116
2005 123 142 133 127
2006 146 185 181 161
Table 6.1
(i) Calculate the four-quarterly (4-point) moving average and interpret your result.
Solution
Table for computation of 4-point moving average
Year/Quarter Indices of retail prices 4-point moving average
2004q1 119
2004q2 127
Page 99 of 174
STA202: Statistics II
¼(119+127+127+116)=122.25
2004q3 127
¼(127+127+116+123)=123.25
2004q4 116
¼(127+116+123+142)=127.00
2005q1 123
¼(116+123+142+133)=128.5
2005q2 142
¼(123+142+133+127)=131.25
2005q3 133
¼(142+133+127+146)=137.00
2005q4 127
¼(133+127+146+185)=147.75
2006q1 146
¼(127+146+185+181)=159.75
2006q2 185
¼(146+185+181+161)=168.25
2006q3 181
2006q4 161
Table 6.2
Note
Each average is placed between the two middle quarters of the four quarters used in
calculating the average.
Box 6.1
Each point on the fitted curve represents the relationship between a known independent
variable and an unknown dependent variable. In general, the least square method uses a
straight line in order to fit through the given points which are known as the method of linear
or ordinary least square. This line is termed as the line of best fit from which the sum of
squares of the distances from the points is minimized.
Equations with certain parameters usually represent the results in this method. The method of
least square actually defines the solution for the minimization of the sum of squares of
deviations or the errors in the result of each equation.
The least squares method is used mostly for data fitting. The best fit result minimizes the
sum of squared errors or residuals which are said to be the differences between the observed
or experimental value and corresponding fitted value given in the model. There are two basic
kinds of the least square methods – ordinary or linear least square and nonlinear least squares.
It is a mathematical method and with it gives a fitted trend line for the set of data in such a
manner that the following two conditions are satisfied.
1. The sum of the deviations of the actual values of Y and the computed values of Y is
zero.
2. The sum of the squares of the deviations of the actual values and the computed values
is least.
This method gives the line which is the line of best fit. This method is applicable to give
results either to fit a straight line trend or a parabolic trend.
The method of least square as studied in time series analysis is used to find the trend line of
best fit to a time series data.
Period
1996 1997 1998 1999 2000 2001 2002 2003 2004
(year)
Y 4 7 7 8 9 11 13 14 17
Table 6.3
Solution
PERIOD
Y X XY X2
YEAR)
1996 4 -4 -16 16
1997 7 -3 -21 9
1998 7 -2 -14 4
1999 8 -1 -8 1
2000 9 0 0 0
2001 11 1 11 1
2002 13 2 16 4
2003 14 3 42 9
2004 17 4 68 16
Table 6.4
From the table we find that value of n is 9, value of ΣY is 90, value of ΣX is 0, value
of ΣXY is 88 and value of ΣX2 is 60.
What are the conditions necessary for a fitted trend line for the set of data in mathematical
representation of times series satisfy?
i. The sum of the deviations of the actual values of Y and the computed values of Y is
zero.
ii. The sum of the squares of the deviations of the actual values and the computed
values is least.
1. Time series analysis is about inferring what has happened to a series of data points
in the past and attempting to predict what will happen to it the future.
2. Additive model is y = T + S + C+ I
3. Multiplicative model is y = T * S * C* I
4. Moving average technique is best suited for data that exhibit some form of regular
periodicity.
5. The least square method uses a straight line in order to fit through the given points
which are known as the method of linear.
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
Period
1996 1997 1998 1999 2000 2001 2002 2003 2004
(year)
Y 6 7 7 8 19 10 13 15 17
Glossary of Terms
References
1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
Introduction
For example, you could use multiple regression to understand whether exam performance
can be predicted based on revision time, test anxiety, lecture attendance and gender.
Alternately, you could use multiple regression to understand whether daily cigarette
consumption can be predicted based on smoking duration, age when started smoking,
smoker type, income and gender.
Multiple regression is a situation where there two or more predictors, and its analysis is
one of the most widely used of all statistical methods. Multiple regressions is a set of
techniques used to analyse the relationship between two or more independent variables
and a dependent variable. Variety of multiple regression models is discussed in this
section. Then we present the basic statistical results for multiple regression in matrix form.
Since the matrix expressions for multiple regression are the same as for simple linear
regression.
When there are two predictor variables 𝑋1 and 𝑋2 , the regression model:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖 (1.1)
is called a first-order model with two predictor variables. 𝑌𝑖 denotes as usual the response
in the ith trial, and 𝑋𝑖1 and 𝑋𝑖2 are the values of the two predictor vari;bles in the 𝑖 𝑡ℎ trial.
The parameters of the model are 𝛽0 , 𝛽1 , 𝑎𝑛𝑑 𝛽2 and the error term is 𝜀𝑖 .
Assuming that E[𝜀𝑖 ]= 0, the regression function for model (1.1) is:
𝐸(𝑌) = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 (1.2)
Similar to simple linear regression, where the regression function 𝐸(𝑌) = 𝛽0 + 𝛽1 𝑋is a
line, regression function (1.2) is a plane. Assuming the 𝛽0 = 12, 𝛽1 = 3, 𝑎𝑛𝑑 𝛽2 = 5,
we’ll have
𝐸(𝑌) = 12 + 3𝑋1 + 5𝑋2 (1.3)
The regression function in multiple regression is often called a regression surface or a
response surface.
The parameter 𝛽1 indicates the change in the mean response 𝐸[𝑌] per unit increase in
𝑋1 when 𝑋2 is held constant. Likewise, 𝛽2 indicates the change in the mean response per
unit increase in 𝑋2 when 𝑋1 is held constant. For example Suppose 𝑋2 is held at the level
𝑋2 = 3 . The regression function (1.3) now is;
𝐸(𝑌) = 12 + 3𝑋1 + 5(3) = 27 + 3𝑋1 (1.4)
Note that this response function is a straight line with slope 𝛽1 = 3. The same is true for
any other value of 𝑋2; only the intercept of the' response function will differ. Hence, 𝛽1 =
3 indicates that the mean response 𝐸 {𝑌} increases by 3 with a unit increase in 𝑋1 when
Page 110 of 174
STA202: Statistics II
𝑋2 is constant, no matter what the level of 𝑋2. We confirm therefore that 𝛽1 indicates the
change in 𝐸{𝑌} with a unit increase in 𝑋1 when 𝑋2is held constant.
The parameters 𝛽1 and 𝛽2 are sometimes called partial regression coefficients because
they reflect the partial effect of one predictor variable when the other predictor variable is
included in the model and is held constant.
We can readily establish the meaning of 𝛽1 and 𝛽2 by calculus, taking partial derivatives
of the response surface (1.2) with respect to 𝑋1 and 𝑋2 as follows:
𝜕𝐸{𝑌}
= 𝛽1
𝜕𝑥1
𝜕𝐸{𝑌}
= 𝛽2
𝜕𝑥2
The partial derivatives measure the rate of change in 𝐸{𝑌} with respect to one predictor
variable when the other is held constant.
We consider now the case where there are 𝑝 − 1 predictor variables 𝑋1 , … , 𝑋𝑝−1 The
regression model
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.5)
𝑝−1
𝑌𝑖 = 𝛽0 + ∑ 𝛽𝑘 𝑋𝑖𝑘 + 𝜀𝑖 (1.6)
𝑘=1
Assuming that E[𝜀𝑖 ]= 0, the response function for regression model 1.6) is:
For variables 𝑋1 , … , 𝑋𝑝−1 in a regression model, the general linear regression model, with
normal error terms, simply in terms of X variables is defined as
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.9)
Where:
𝛽0 , 𝛽1 , … , 𝛽𝑝−1 𝑎𝑟𝑒 𝑝𝑎𝑟𝑎𝑚𝑡𝑒𝑟𝑠
𝑋𝑖1 , … , 𝑋𝑖,𝑝−1 𝑎𝑟𝑒 𝑘𝑛𝑜𝑤𝑛 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡𝑠
𝜀𝑖 𝑎𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑒𝑛𝑑𝑒𝑛𝑡 𝑁(0, 𝜎 2 )
𝑖 = 1, … . 𝑛
If we let 𝑋𝑖0 = 1, we have
𝑌𝑖 = 𝛽0 𝑋𝑖0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.10)
𝑝−1
The response function for regression model (1.9) is, since E[𝜀𝑖 ]= 0:
𝐸(𝑌𝑖 ) = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 (1.12)
1 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑚𝑎𝑙𝑒
𝑋2={
0 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑓𝑒𝑚𝑎𝑙𝑒
1 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑚𝑎𝑙𝑒
𝑋𝑖2={
0 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑓𝑒𝑚𝑎𝑙𝑒
Polynomial regression models are special cases of the general linear regression model.
They contain squared and higher-order terms of the predictor variable (s), making the
response function curvilinear. The following is a polynomial regression model with one
predictor variable
Models with transformed variables involve complex, curvilinear response functions, yet
still are special cases of the general linear regression model. Consider the following
model with a transformed 𝑌 variable:
𝑙𝑜𝑔𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝛽3 𝑋𝑖3
+ 𝜀𝑖 (1.15)
If we let 𝑌𝑖′ = 𝑙𝑜𝑔𝑌𝑖 , we can write regression model in (1.15) as follows
𝑌𝑖′ = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝛽3 𝑋𝑖3 + 𝜀𝑖
which is in the form of general linear regression model (1.9). The response variable just
happens to be the logarithm of 𝑌. Many models can be transformed into the general linear
regression model. For instance, the model
1
𝑌𝑖 = (1.16)
𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖
We then have:
In order to use matrix form the important point is to know what the inverse of the matrix
represents. To express general linear regression model (1.9):
Y1 1 X 11 X 1, p 1
Y 1 X X 2, p 1
Y 2 X
21
n1 n p
Yn 1 X n1 X n , p 1
0 1
1 2
p1 n1
n n
With this compact notation, the linear regression model can be written in the form
y X
In linear algebra terms, the least-squares parameter estimates β are the vectors that
Minimize
𝑛
If 𝜀𝑖 = 0, then
̂ = 𝑿𝛽̂
𝒚
𝑿′ (𝒚 − 𝑿𝛽̂ ) = 0
𝑿′ 𝒚 = 𝑿′ 𝑿𝛽̂
𝒀 and 𝜺 vectors are the same as for simple linear regression. The 𝛽 vector contains
additional regression parameters, and the 𝑿 matrix contains a column of 𝐼𝑠 as well as a
column of the n observations for each of the 𝑝 − 1 𝑋 variables in the regression model.
The row subscript for each element 𝑋𝑖1in the 𝑿 matrix identifies the trial or case, and the
column subscript identifies the 𝑿 variable.
Y = X β+ ε
n×1 n×p p×1 p×1
(1.17)
where:
𝒀 is a vector of responses
𝜷 is a vector of parameters
𝑿 is a matrix of constants
𝜺 is a vector of independent normal random variables with expectation
𝑬{ 𝜺} = 0 and variance-covariance matrix:
2 0 0
0 2 0
(ε)
2
2I
nn
0 0 2
Therefore, the random vector 𝑌 has expectation
𝑬{𝒀} = 𝑿𝜷
(1.18)
n n
(1.19)
b0
b
b
1
(1.21)
p×1
b p 1
The least square normal equations for the general linear regression model (1.17) are:
X'Xb = X'Y (1.22)
b X'X X ' Y
1
∑ 𝜀 = ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 )2
2
𝑖=1 𝑖=1
𝑛
∑ 𝜀 2 𝑏𝑒 𝑟𝑒𝑝𝑟𝑒𝑠𝑒𝑛𝑡𝑒𝑑 𝑤𝑖𝑡ℎ 𝑆
𝑖=1
𝑆 = ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 )2
𝑖=1
𝜕𝑆 𝜕𝑆 𝜕𝑆
= 0; = 0; =0
𝜕𝛽0 𝜕𝛽1 𝜕𝛽2
𝜕𝑆
= −2 ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽0
𝜕𝑆
= −2 ∑ 𝑋1 (𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽1
𝜕𝑆
= −2 ∑ 𝑋2 (𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽2
Consequently, normal equation is obtained as follows
∑ 𝑌 = 𝑛𝛽̂0 + 𝛽̂1 ∑ 𝑋1+𝛽̂2 ∑ 𝑋2 (1.23)
ˆ X'X X ' Y
1
V ( ˆ ) X'X e2
1
∑ 𝑋2 2 ∑ 𝑋1 2
𝑉(𝛽̂1 ) = 𝜎𝑒2 ̂ 2
; 𝑉(𝛽2 ) = 𝜎𝑒
⎹𝑋 ′ 𝑋⎹ ⎹𝑋 ′ 𝑋⎹
2
𝑆𝑆𝑅 𝛽′𝑋′𝑌 𝑌 ′ 𝑌 − 𝑒′𝑒 𝑒 ′𝑒
𝑅 = = = = 1− ′ (1.27)
𝑆𝑆𝑇 𝑌′𝑌 𝑌′𝑌 𝑌𝑌
𝑆𝑆𝐸
= 1−
𝑆𝑆𝑇
𝑆𝑆𝐸 𝑛 − 1
𝑅̅ 2 = 1 − [( )( )] (1.28)
𝑆𝑆𝑇 𝑛 − 𝑘
𝑅̅ 2 is the adjusted 𝑅 2
The column headed MS refers to the mean square and is obtained by dividing the SS term
by the 𝑑𝑓 term. Thus, MSR, the mean square regression, is equal to 𝑆𝑆𝑅/𝑘, and MSE
equals 𝑆𝑆𝐸/ [𝑛 − (𝑘 + 1)]. The general format of the ANOVA table is:
Analysis of variance
Source df SS MS F
Regression 𝑘 SSR 𝑀𝑆𝑅 = 𝑆𝑆𝑅/𝑘 𝑀𝑆𝑅
/𝑀𝑆𝐸
Error 𝑛 − (𝑘 SSE 𝑀𝑆𝐸 = 𝑆𝑆𝐸/
+ 1) 𝑛 − (𝑘 + 1)
Total 𝑛−1 SST
7.2.1 Global Test: Testing Whether the Multiple Regression Model is Valid
The overall ability of the independent variables 𝑋1, 𝑋2,… 𝑋𝑘 , to explain the behaviour of
the dependent variable 𝑌 can be tested. Two tests of hypotheses are considered. The first
one is called the global test.
Global test: An overall test of the regression model. It investigates the possibility that all
the regression coefficients are equal to zero.
It tests the overall ability of the set of independent variables to explain differences in the
dependent variable. The null hypothesis is that all of the population regression
coefficients are zero. If accepted, it would imply that the set of coefficients is of no value
in explaining differences in the dependent variable. The alternate hypothesis is that at least
one of the coefficients is not zero. This test is written in symbolic form for three
independent variables as:
𝐻0 ∶ 𝛽1 = 𝛽2 = 𝛽3 = 0
𝐻1 ∶ 𝑁𝑜𝑡 𝑎𝑙𝑙 𝑡ℎ𝑒 𝛽 𝑠 = 0
Rejecting 𝐻0 and accepting 𝐻1 implies that one or more of the independent variables
𝐸𝑆𝑆/(𝑘 − 1)
𝐹= (1.29)
𝑅𝑒𝑔 𝑆𝑆/[𝑛 − 𝑘]
Where:
𝑆𝑆𝑅 is the sum of the squares “explained by” the regression.
𝑘 is the number of independent variables.
𝑆𝑆𝐸 is the sum of squares error.
𝑛 is the number of observations.
The second test of hypothesis identifies which of the set of independent variables are
significant predictors of the dependent variable. That is, it tests the independent variables
individually rather than as a unit. This test is useful because unimportant variables can be
eliminated from the regression model. The test statistic is the Student 𝑡 distribution with
𝑛 − (𝑘 + 1) degrees of freedom. For example, suppose we want to test whether the
hypothesis that the coefficient for the first independent variable in the model was equal to
zero versus the alternative hypothesis that it was not equal to zero. The null and alternate
hypotheses would be written as follows.
𝐻0 ∶ 𝛽1 = 0
𝐻1 ∶ 𝛽1 ≠ 0
Rejection of the null hypothesis and acceptance of the alternate hypothesis would
imply that variable number one is significant and that it has an inverse relationship
with the dependent variable.
𝛽̂1
𝑇= ~𝑡𝑛−2 (1.30)
𝑆. 𝐸(𝛽̂1 )
If the hypothesis test finds that the null hypothesis cannot be rejected, then the variable
should be dropped from the model. However, the above test only supports removing one
variable at a time from the model. After a variable is removed, a new regression model is
constructed using the remaining variables and a new t-test can be conducted for each of
the remaining variables.
The variables used in regression analysis have been mostly been quantitative variables. A
quantitative variable is a variable that is numerical in nature, such as, the number of hours
worked by employees, the number of traffic accidents ion 3rd mainland bridge in a week,
and the distance travelled to work.
A variable in which there are only two possible outcomes is called dummy variable. For
analysis, one of the outcomes is coded a 1 and the other a 0 as previously discussed in
section 1.3.
∑𝑌 𝒏 ∑ 𝑋1 ∑ 𝑋2 𝛽̂0
′ ′
𝑿 𝒀 = [∑ 𝑋1 𝑌]; 𝑿 𝑿 = [∑ 𝑋1 ∑ 𝑋1 2 ∑ 𝑋1 𝑋2 ]; 𝜷 = [𝛽̂1 ]
∑ 𝑋2 𝑌 ∑ 𝑋2 ∑ 𝑋1 𝑋2 ∑ 𝑋2 2 𝛽̂3
4 30 70
𝑿′ 𝑿 = [30 290 519 ]
70 519 1246
⎹𝑿′ 𝑿⎹ = 4(91,979) − 30((1050) + 70(−4730)
= 5,316
290 519
+𝑀11 = | | = (290 ∙ 1246) − 5192 = 91,979
519 1246
30 519
−𝑀12 = | | = −[(30 ∙ 1246) − (519 ∙ 70)] = −1,050
70 1246
It follows that:
𝛽̂0 −18.0380
𝛽̂ = [𝛽̂1 ] = [ −0.0903 ]
𝛽̂2 2.3552
𝑌 ′ 𝑌 = ∑ 𝑌 2 = 2,150
90
𝑅𝑒𝑔 𝑆𝑆=𝛽̂ ′ 𝑋 ′ 𝑌 = (−18.0380 −0.0903 )
2.3550 655 )
(
1625
𝑅𝑒𝑔 𝑆𝑆 = 2,144.63
𝐸𝑟𝑟𝑜𝑟 𝑆𝑆 = 𝑇𝑆𝑆 − 𝑅𝑒𝑔 𝑆𝑆
Source Df SS MS F
Regression 2 2,144.63 1072.315 199.68
Error 1 5.37 5.37
Total 3 2150
Table 7.2: Analysis of Variance
𝑆𝑆𝑅 2144.63
(b) 𝑅 2 = 𝑆𝑆𝑇 = = 0.9975
2150
5.37 4 − 1
𝑅̅ 2 = 1 − [( )( )]
2150 4 − 2
𝑆. 𝐸(𝛽̂1 ) = √𝑉(𝛽̂1)=0.2059
2.3550
𝑡= = 11.4376
0.2059
𝑡2 0.05 = 2.132
It is significant at 5%.
Regression Statistics
Adjusted R Square 0.636027664
𝐹 3.679636867
Observations 25
Table 7.3
Assume that the distribution of the least square estimates is approximately normal.
a) Test the null hypothesis that the coefficient 𝛽1 = 0 versus the alternative hypothesis of 𝛽1 ≠ 0 at
the 10% significance level.
b) Construct a 95% confidence interval for 𝛽2.
c) Test the overall significance of the regression model at the 5% level.
d) A researcher wonders whether the model would be improved by removing
𝑍1 and 𝑍2 . What are null and alternative hypotheses which are appropriate for this case?
Solution
(a)
𝛽̂1
𝑡= ~𝑡𝑛−2
𝐸. 𝑆. 𝐸(𝛽̂1 )
𝛽̂1 = 0.049, 𝐸. 𝑆. 𝐸(𝛽̂1 ) = 0.023
0.049
𝑡= = 2.13
0.023
Polynomial regression models may contain one, two, or more than two predictor variables.
Each predictor variable may be present in various powers. We start by considering a
polynomial regression model with one predictor variable raised to the first and second powers:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽2 𝑥𝑖2 + 𝜀𝑖 (7.4)
Where 𝑥𝑖 = 𝑋𝑖 − 𝑋̅
This polynomial model is called a second-order model with one predictor variable because the
single predictor variable is expressed in the model to the first and second powers. Note that the
predictor variable is centred-in other words, expressed as a deviation around its mean 𝑋̅ −and
that the ith centered observation is denoted by 𝑥𝑖 .
The regression coefficients in polynomial regression are frequently written in a slightly
different fashion, to reflect the pattern of the exponents:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽11 𝑥𝑖2 + 𝜀𝑖 (7.5)
The response function for regression model (7.5) is:
𝐸{𝑌}𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽11 𝑥𝑖2 (7.6)
This response function is a parabola and known as a quadratic response function.
The regression coefficient 𝛽0 represents the mean response of 𝑌 when 𝑥 = 0, i.e., wh
𝑋 = 𝑋̅. The regression coefficient 𝛽1 is often called the linear effect coefficient, and 𝛽11
is known as the quadratic effect coefficient.
The algebraic version of the least square normal equations:
Page 129 of 174
STA202: Statistics II
′
𝑿 𝑿𝒃 = 𝑿′𝒀
for the second-order polynomial regression model (7.5) can be readily obtained from. Since
∑ 𝑥𝑖 = 0, this yields the normal equations:
i. Order of the model: The order of the polynomial model is kept as low as possible. Some
transformations can be used to keep the model to be of first order. If this is not satisfactory,
then second order polynomial is tried. Arbitrary fitting of higher order polynomials can be a
serious abuse of regression analysis. A model which is consistent with the knowledge of data
and its environment should be taken into account. It is always possible for a polynomial of
order (𝑛 – 1) to pass through 𝑛 points so that a polynomial of sufficiently high degree can
always be found that provides a “good” fit to the data. Such models neither enhance the
understanding of the unknown function nor be a good predictor.
ii. Model building strategy: A good strategy should be used to choose the order of an
approximate polynomial. One possible approach is to successively fit the models in increasing
order and test the significance of regression coefficients at each step of model fitting. Keep the
order increasing until t-test for the highest order term is non-significant. This is called as
forward selection procedure. Another approach is to fit the appropriate highest order model and
then delete terms one at a time starting with highest order. This is continued until the highest
order remaining term has a significant t-statistic. This is called as backward elimination
procedure. The forward selection and backward elimination procedures does not necessarily
lead to same model. The first and second order polynomials are mostly used in practice.
iii. Extrapolation: One has to be very cautious in extrapolation with polynomial models. The
curvatures in the region of data and region of extrapolation can be different. So predicted
response will not be based on the true behaviour of the data.
iv. Ill-conditioning: A basic assumption in linear regression analysis is that 𝑋-matrix is of
full column rank. In polynomial regression models, as the order increases, the 𝑋’𝑋 matrix
becomes ill-conditioned. As a result, the matrix (𝑋’𝑋)−1 may not be accurate and the
parameters will be estimated with considerable error. If values of 𝑥 lie in a narrow range
then the degree of ill-conditioning increases and multicollinearity in the columns of
𝑦 = 𝛽0 + 𝛽1 𝑥 + 𝛽2 𝑥 2 + 𝛽4 𝑥 4 + 𝜀
7.4 Analysis
𝑝0 (𝑥𝑖 ) = 1
When we have the model as 𝑦 = 𝑋𝜷 + 𝜀, 𝑋 − matrix, in this case is given by
𝑝0 (𝑥1 ) ⋯ 𝑝𝑘 (𝑥1 )
𝑋=[ ⋮ ⋱ ⋮ ] (7.15)
𝑝0 (𝑥𝑛 ) ⋯ 𝑝𝑘 (𝑥𝑛 )
Since this 𝑋 -matrix has orthogonal columns, so 𝑋’𝑋 matrix becomes
𝑛
∑ 𝑃02 (𝑥𝑖 ) ⋯ 0
𝑖=1
𝑋’𝑋 = ⋮ ⋱ ⋮ (7.16)
𝑛
0 ⋯ ∑ 𝑃𝑘2 (𝑥𝑖 )
[ 𝑖=1 ]
The ordinary least squares estimator is 𝛼̂ = (𝑋′𝑋)−1 𝑋′𝑦 which is for 𝛼𝑗 is
∑𝑛𝑖=1 𝑃𝑗 (𝑥𝑖 )𝑦𝑖
𝛼̂𝑗 = 𝑛 , 𝑗 = 0,1,2, … , 𝑘 (7.17)
∑𝑖=1 𝑃𝑗2 (𝑥𝑖 )
This regression sum of squares does not depend on other parameters in the model.
The analysis of variance table, in this case, is given as follows
Source of variation Degree of freedom Sum of Squares Mean squares
𝛼̂0 1 𝑆𝑆(𝛼̂0 ) −
𝛼̂1 1 𝑆𝑆(𝛼̂1 ) 𝑆𝑆(𝛼̂1 )
𝛼̂2 1 𝑆𝑆(𝛼̂2 ) 𝑆𝑆(𝛼̂2 )
⋮ ⋮ ⋮ ⋮
𝛼̂𝑘 1 𝑆𝑆(𝛼̂𝑘 ) 𝑆𝑆(𝛼̂𝑘 )
Residual 𝑛−𝑘−1 𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘) (by subtraction) 𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘)
Total 𝑛 𝑆𝑆𝑇
Table 7.4
If we add another term 𝑃𝑘+1 (𝑥𝑖 )𝛼𝑘+1 in the model, then the model is
Notice that:
(i) We need not to bother for other terms in the model.
(ii) Simply concentrate on the newly added term only.
(iii) No re-computation of (𝑋′𝑋)−1or any other 𝛼̂𝑗 (𝑗 ≠ 𝑘 + 1) is necessary due to orthogonality of
polynomials.
(iv) Thus higher-order polynomials can be fitted with ease.
(v) Terminate the process when a suitably fitted model is obtained.
To test the significance of the highest order term, we test the null hypothesis
𝐻0 : 𝛼𝑘 = 0
This hypothesis is equivalent 𝑡𝑜 𝐻0 : 𝛽𝑘 = 0 in polynomial regression model.
The following apply
𝑆𝑆𝑟𝑒𝑔 (𝛼𝑘 )
𝐹0 = (2.23)
𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘)/(𝑛 − 𝑘 − 1)
~𝐹(1, 𝑛 − 𝑘 + 1) under 𝐻0
If the order of the model is changed to (𝑘 + 𝑟), we need to compute only r new coefficients.
The remaining coefficients 𝛼̂0 , 𝛼̂1 , … , 𝛼̂𝑘 do not change due to the orthogonality property of
polynomials. Thus the sequential fitting of the model is computationally easy.
When 𝑋𝑖 are equally spaced, the tables of orthogonal polynomials are available, and the
orthogonal polynomials can be easily constructed.
𝑥𝑖 − 𝑥̅ 3 𝑥𝑖 − 𝑥̅ 3𝑛3 − 7
𝑝3 (𝑥𝑖 ) = 𝜆3 [( ) −( )( )]
𝑠 𝑠 20
𝑥𝑖 − 𝑥̅ 4 𝑥𝑖 − 𝑥̅ 2 3𝑛3 − 13 3(𝑛2 − 1)(𝑛2 − 19)
𝑝4 (𝑥𝑖 ) = 𝜆4 [( ) −( ) ( )+ ]
𝑠 𝑠 14 560
An example of the table for n 𝑛 = 5 is as follows:
𝑥𝑖 𝑃1 𝑃2 𝑃3 𝑃4
𝜆 1 1 5⁄6 35⁄12
Table 7.5
The orthogonal polynomials can also be constructed when 𝑥’𝑠 are not equally spaced.
𝑦 = 71.102 + 6.3516𝑥
𝑦 = 64.897 + 6.3516𝑥 + 0.7521𝑡 2
𝑦 = 64.897 + 1.9492𝑥 + 0.7521𝑥 2 + 0.3005𝑥 3
𝑦 = 64.888 + 1.9492𝑥 + 0.7521𝑥 2 + 0.3005𝑥 3 − 0.0002𝑥 4
𝑦 = 64.888 − 0.5024𝑥 + 0.7521𝑥 2 + 0.8016𝑥 3 − 0.0002𝑥 4 − 0.0194𝑥 5
These lead to the following predictions for 1982.25
Degree 𝜇̂ 1982.25
1 113.98
2 142.04
3 204.74
4 204.50
5 70.260
1. Multiple regression is a situation where there two or more predictors, and its
analysis is one of the most widely used of all statistical methods.
2. Polynomial regression models contain squared and higher-order terms of the
predictor variable (s), making the response function curvilinear
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
1. 1. State whether the following statements are true or false. Give a brief
explanation.
2. (a) A negative chi-squared value shows that there is little association between the
variables tested.
3. (b) Similar observed and expected frequencies indicate strong evidence of an
association in an 𝑟 × 𝑐 contingency table.
Glossary of Terms
References
1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
7. [Link]
[Link]
Introduction
Before looking at partial correlation, we shall revise simple correlation. The simple
correlation coefficient ( rx1x2 ) between the two pair of variables ( 𝒙𝟏, 𝒙𝟐 ) is defined as
n x1 x2 x1 x2
rx1x2 (8.1)
n x 2 x 2 n x 2 x 2
1 1 2 2
n x1 y x1 y
rx1 y
n x 2 x 2 n y 2 y 2
1 1
We shall consider another approach to solving multiple regression problem. The least
square regression equation in three variables is given by
𝑦̂ = 𝛽1∙23 + 𝛽12∙3 𝑥1𝑖 + 𝛽13∙2
Let 𝑟𝑥1 𝑦 =𝑟12 = 0.4160,
𝑟𝑥2 𝑦 =𝑟13 = −0.693 and
𝑟𝑥1 𝑥2 =𝑟23 = −0.231
And
Partial correlation coefficient is the correlation coefficient between any two of a particular
three variables while the remaining one is kept constant. Thus, we can obtain the partial
correction coefficient between 𝑦 and 𝑥1 keeping 𝑥2 constant as follows:
𝑦 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 1
𝑥1 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 2
And
𝑥2 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 3
r12 r13r23
r123
1 r 1 r
∴ 2 2
12 23
Similarly,
ryx1 x2 , ryx2 x3 rx1x2 y , that is, r123 , r132 and 231r
Solution
ryx1 ryx2 rx1x2
ryx1 x2
1 r 1 r
2
yx2
2
x1 x2
0.4160 0.1601
0.5193 0.9466
0.2559
0.4912
0.2559
0.7008
ryx1 x2 0.365
0.693 0.0960
0.8269 0.9466
Table 8.4: SPSS output for partial Correlation coefficient between y and x2 taking
0.231 0.2883
0.8269 0.5198
0.0573
0.4298
0.0573
0.6556
rx1x2 y 0.087
Table 8.5: SPSS output for partial Correlation coefficient between x2 and x1 taking
y as constant or the controlling variable.
This type of test tests the null hypothesis that two factors (or attributes) are not associated (that
is, independent), against the alternative hypothesis that they are associated. Each data unit we
sample has one level (or `type' or `variety') of each factor.
Case Study 8.4
Suppose that we are sampling people, and that one factor of interest is hair colour (black,
blonde, brown etc.) while another factor of interest is eye colour (blue, brown, green etc.). In
this example, each sampled person has one level of each factor. We wish to test whether or not
these factors are associated. Hence:
𝐻0 : There is no association between hair colour and eye colour
𝐻1 : There is an association between hair colour and eye colour.
So, under 𝐻0 , the distribution of eye colour is the same for blondes as it is for brunettes etc.,
whereas if 𝐻1 is true it may be attributable to blonde-haired people having a (significantly)
higher proportion of blue eyes, say.
In a contingency table, also known as a cross-tabulation, the data are in the form of frequencies
(counts), where the observations are organised in cross-tabulated categories. We sample a
certain number of units (such as people) and classify them according to the two factors of
interest.
Case Study 8.5
In three areas of a city, a record has been kept of the numbers of Covid-19, Malaria and Lassa
fever which take place in a year. The total number of occurrences was 150, and they were
divided into the various categories as shown in the following contingency table:
Area Covid-19 Malaria Lassa fever Total
A 30 19 6 55
B 12 23 14 49
C 8 18 20 46
Total 50 60 40 150
Table 8.6
The cell frequencies are known as observed frequencies and show how the data are spread
across the different combinations of factor levels (known as contingencies). The first step
in any analysis is to complete the row and column totals (as already done in this table).
The expected frequency, 𝐸𝑖𝑗 , for the cell in row i and column 𝑗 of a contingency table with 𝑟
rows and 𝑐 columns, is:
𝑟𝑜𝑤 𝑖 𝑡𝑜𝑡𝑎𝑙 × 𝑐𝑜𝑙𝑢𝑚𝑛 𝑗 𝑡𝑜𝑡𝑎𝑙
𝐸𝑖𝑗 =
𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛
Where 𝑟 = 1, … , 𝑟 𝑎𝑛𝑑 𝑗 = 1, … , 𝑐
Area Covid-19 Malaria Lassa fever
A 𝐸11 𝐸12 𝐸13
B 𝐸21 𝐸22 𝐸23
C 𝐸31 𝐸32 𝐸33
Table 8.7
To motivate our choice of test statistic, if 𝐻0 is true then we would expect to observe small
differences between the observed and expected frequencies, while large differences
would suggest that 𝐻1 is true. This is because the expected frequencies have been
calculated conditional on the null hypothesis of independence hence, if 𝐻0 is actually
true, then what we actually observe (the observed frequencies) should be (approximately)
equal to what we expect to observe (the expected frequencies).
Let the contingency table have 𝑟 rows and 𝑐 columns, then formally the test statistic used
for tests of association is:
𝑟 𝑐 2
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 )
∑∑ ~𝜒 2 (𝑟−1)(𝑐−1) (6.1)
𝐸𝑖𝑗
𝑖=1 𝑗=1
Critical values are found using Statistical Tables. The `double summation' here just means
summing over all rows and columns. This test statistic follows an (approximate) chi-
squared distribution with (𝑟 − 1)(𝑐 − 1) degrees of freedom, where 𝑟 and 𝑐 denote the
number of rows and columns, respectively, in the contingency table.
𝑣 = 𝑟𝑐 − (𝑟 − 1) − (𝑐 − 1) − 1
= (𝑟 − 1)(𝑐 − 1)
As usual, we choose a significance level, 𝛼, at which to conduct the test. However, are we
performing a one-tailed test or a two-tailed test? To determine this, we need to consider
what sort of test statistic value would be considered extreme under 𝐻0 . If 𝐻0 is true, then
the observed and expected frequencies should be quite similar, since the expected
frequencies are computed conditional on the null hypothesis of independence. This
means that ⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ should be quite small for all cells. In contrast if 𝐻0 is not true,
then we would expect comparatively large values for⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ due to large differences
between the two sets of frequencies. Therefore, upon squaring ⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ suffciently
large test statistic values suggest rejection of 𝐻0 . Hence 𝜒 2 tests of association are always
upper-tailed tests.
Looking again at the contingency table, comparing observed and expected frequencies,
the interpretation of this association becomes clear Covid-19 is the main problem in area 𝐴
whereas Lassa Fever is a problem in area 𝐶. (We can deduce this by looking at the cells
However, for data involving more factors, and more factor levels, this type of analysis can
be very insightful. Cells which make a large contribution to the test statistic value (i.e.
2
which have large values of (𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) ⁄𝐸𝑖𝑗 should be studied carefully when determining
the nature of an association. This is because, in cases where 𝐻0 has been rejected, rather
than simply conclude that there is an association between two categorical variables, it is
helpful to describe the nature of the association.
In addition to tests of association, the chi-squared distribution is often used more generally
in so-called `goodness-of fit tests. We may, for example, wish to answer hypotheses such
as `Is it reasonable to assume the data follows a particular distribution? This justifies the
name `goodness-of-fit' tests, since we are testing whether or not a particular probability
distribution provides an adequate fit to the observed data. The null hypothesis will assert
that a specific hypothetical population distribution is the true one. The alternative
hypothesis is that this specific distribution is not the true one.
There is a special case when we are only dealing with one row or one column. This is
when we wish to test that the sample data are drawn from a (discrete) uniform
distribution, i.e. that each characteristic is equally likely.
As with tests of association, the goodness-of-fit test involves both observed and expected
frequencies. In all goodness-of-fit tests, the sample data must be expressed in the form of
observed frequencies associated with certain classifications of the data. Assume 𝑘
classifications, hence the observed frequencies can be denoted by 𝑂𝑖 , for 𝑖 = 1, … , 𝑘 .
For discrete uniform probability distributions, expected frequencies are computed as:
1
𝐸𝑖 = 𝑛 × 𝑓𝑜𝑟 𝑖 = 1,2 … , 𝑘
𝑘
where 𝑛 denotes the sample size and 11⁄𝑘 is the uniform (i.e. equal, same) probability for
each characteristic.
𝑘−1
𝐸𝑘 = 𝑛 − ∑ 𝐸𝑖 (6.2)
𝑖=1
This is because we have a constraint that the sum of the observed and expected
frequencies must be equal,5 that is:
𝑘 𝑘
∑ 𝑂𝑖 = ∑ 𝐸𝑖 (6.3)
𝑖=1 𝑖=1
𝑘
(𝑂𝑖 − 𝐸𝑖 )2
∑ ~𝜒 2 (𝑘−1) 𝑎𝑝𝑝𝑟𝑜𝑥𝑖𝑚𝑎𝑡𝑙𝑦 𝑢𝑛𝑑𝑒𝑟 𝐻0 (6.4)
𝐸𝑖
𝑗=1
Note that this test statistic does not have a true 𝜒 2 (𝑘−1) distribution under 𝐻0 , rather it is
only an approximating distribution. An important point to note is that this approximation
is only good enough provided all the expected frequencies are at least there are 𝑘 − 1
degrees of freedom when testing a discrete uniform distribution. 𝑘 is the number of
categories, and we lose one degree of freedom due to the constraint that:
∑ 𝑂𝑖 = ∑ 𝐸𝑖
As with the test of association, goodness-of-fit t tests are upper-tailed tests as, under H0,
we would expect to see small differences between the observed and expected
frequencies, as the expected frequencies are computed conditional on 𝐻0 . Hence large
test statistic values are considered extreme under 𝐻0 , since these arise due to large
differences between the observed and expected frequencies.
33
𝐸𝑖 = = 11
3
Starting 1 2 3 4 5 6 7 8
Position
Number of 29 19 18 25 17 10 15 11
wins
Table 8.10
Solution
We test whether the data follow a discrete uniform distribution of 8 categories. Let
𝑝𝑖 = 𝑃(𝑋 = 𝑖), 𝑓𝑜𝑟 𝑖 = 1, . . . , 8.
Starting 1 2 3 4 5 6 7 8
Positio
n
𝑂𝑖 29 19 18 25 17 10 15 11
𝐸𝑖 18 18 18 18 18 18 18 18
𝑂𝑖 − 𝐸𝑖 11 1 0 7 −1 −8 −3 −7
(𝑂𝑖 6.72 0.06 0 2.72 0.06 3.66 0.50 2.72
− 𝐸𝑖 )2
/𝐸𝑖
Table 8.11
(𝑂𝑖 −𝐸𝑖 )2
Under 𝐻0 , ~𝜒 2 7 . At 5% significant level, the critical value is 14.067. Since
𝐸𝑖
14.067<16.34 we reject the null hypothesis. Turning to 1% significance level, the critical
value is 18.475. Since 16.34 < 18.475, we cannot reject the null hypothesis; hence we
conclude that the test is moderately significant. There is moderate, but not strong,
evidence to support the claim that the chances of winning in the different starting positions
are not all the same.
The simple correlation coefficient ( rx1x2 ) between the two pair of variables ( 𝒙𝟏, 𝒙𝟐) is
n x1 x2 x1 x2
defined as rx1x2
n x 2 x 2 n x 2 x 2
1 1 2 2
1. The expected frequency, 𝐸𝑖𝑗 , for the cell in row i and column 𝑗 of a contingency table
with 𝑟 rows and 𝑐 columns, is:
𝑟𝑜𝑤 𝑖 𝑡𝑜𝑡𝑎𝑙 × 𝑐𝑜𝑙𝑢𝑚𝑛 𝑗 𝑡𝑜𝑡𝑎𝑙
𝐸𝑖𝑗 =
𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛
Where 𝑟 = 1, … , 𝑟 𝑎𝑛𝑑 𝑗 = 1, … , 𝑐
2. Let the contingency table have 𝑟 rows and 𝑐 columns, then formally the test
statistic used for tests of association is:
𝒓 𝒄 𝟐
(𝑶𝒊𝒋 − 𝑬𝒊𝒋 )
∑∑ ~𝝌𝟐 (𝒓−𝟏)(𝒄−𝟏)
𝑬𝒊𝒋
𝒊=𝟏 𝒋=𝟏
Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
(a) Based on the data in the table, and without conducting any significance test, would you
say there is an association between age and watch preference?
Provide a brief justification for your answer.
(b) Calculate the 𝜒 2 statistic for the hypothesis of independence between age and watch
preference, and test that hypothesis. What do you conclude?
Glossary of Terms
Contingency table: a table showing the distribution of one variable in rows and another
in columns, used to study the correlation between the two variables
1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
7. [Link]
[Link]
X = a + by
Where
𝑛 ∑ 𝑥𝑦 − ∑ 𝑥 ∑ 𝑦
𝑏=
𝑛 ∑ 𝑦 2 − (∑ 𝑦)2
𝑎 = 𝑥̅ − 𝑏𝑦̅
X Y Xy X2 Y2
66 69 4554 4356 4761
64 67 4288 4096 4489
68 69 4692 4624 4761
65 66 4290 4225 4356
69 70 4830 4761 4900
63 67 4221 3969 4489
71 69 4899 5041 4761
1.
i. Statistical hypothesis is a statement about the parameters or form of a
population. A test of a statistical hypothesis is a criteria which specifies for
what sample results the hypothesis is to be accepted or rejected.
ii. A type I error has been committed if we reject the null hypothesis when it
is true and a type II error has been committed if we accept the null
hypothesis when it is false.
iii. A test of any statistical hypothesis where the alternative is one sided such
as:
H0: = 0 or H0: = 0
H1: > 0 H1: < 0
is called a one-tailed test. The critical region for H1: > 0 lies entirely in the right tail
while the critical region for H1: < 0 lies entirely in the left tail. A test of any statistical
hypothesis where the alternative is two-sided such as:
H0: = 0 vs H1: 0
is called a two-tailed test, values in the both tails of the distribution constitute the
critical region.
2. The steps are:
𝑋1 (𝑋1 − 𝑥̅ )2 𝑋2 (𝑋2 − 𝑥̅ )2
84 9 75 4
82 1 76 1
83 4 77 0
78 9 80 9
79 4 76 1
406 27 384 15
∑ 𝑋1 406
𝑥̅1 = = = 81.2 ~ 81
𝑛1 5
∑(𝑋1 − 𝑋̅2 ) 27
𝑆12 = = = 5.4
𝑛 5
𝑋̅1 − 𝑋̅2
𝑡=
𝑆1 𝑆
+ 2
√𝑛1 √𝑛2
81 − 77 4
= = = 2.197
2.32 1.73 1.82
+
√5 √5
𝑡𝑡𝑎𝑏𝑢𝑙𝑎𝑡𝑒𝑑 = 𝑡0.05 (𝑛1 + 𝑛2−2 ) = 𝑡0.05 (8) = 1.894
Interpretation
Null hypothesis (H0): P = 40% = 0.4
Alternate hypothesis (H1) = P ≠ 40% ≠ 0.4
Hence: P = 0.40, q = 0.60
Observed sample population
180
𝑝̂ = = 0.36
500
The observed value of Z is -0.579 which is the acceptance region and such H0 is
accepted.
Null hypothesis (H0): 𝑃̂1 = 𝑃̂2
Alternative Hypothesis (H1): 𝑃̂1 ≠ 𝑃̂2
450
𝑃̂1 = = 0.45 𝑞̂1 = 1 − 𝑃1 = 1 − 0.45 = 0.55, 𝑛1 = 1000
1000
400
𝑃̂2 = = 0.5 𝑞̂1 = 1 − 𝑞2 = 1 − 0.5 = 0.5, 𝑛2 = 400
800
The test statistic
𝑃̂1 − 𝑃̂2 0.45 − 0.50 −0.05
𝑍= = =
√0.00025 + 0.00031
̂ ̂ √0.45(0.55) + 0.5(0.5)
√𝑃1 𝑞̂1 + 𝑃2 𝑞̂2 1000 800
𝑛1 𝑛2
−0.05 −0.005
𝑍= = = −0.213
√0.000563 0.0237
𝑍𝑡𝑎𝑏𝑙𝑒 𝑎𝑡 1% = 1.64
𝑍𝑡𝑎𝑏𝑙𝑒 𝑎𝑡 5% = 1.96
The observed value of Z is -0.213 which is acceptance region at 1% and 5% level and
such H0 is accepted.
1. Time series is defined as some quantity that is measured sequentially in time over
some interval.
2. The components of time series are Trend, Cyclical movement, Seasonal movement
and Irregular movement.
3.
i. Additive model: This is a model in which the series value of y is the sum of
all four components, that is,
y = T + S + C+ I
ii. Multiplicative model: It is a del in which y is the product of all the time
series components. This is denoted by
PERIOD
Y t Yt t2
YEAR)
1996 6 -4 -24 16
1997 7 -3 -21 9
1998 7 -2 -14 4
1999 8 -1 -8 1
2000 19 0 0 0
2001 10 1 10 1
2002 13 2 16 4
2003 15 3 45 9
2004 17 4 68 16
The fundamental steps in planning a sample survey include defining the survey's objective, establishing the scope, determining subject coverage, choosing a method of data collection, organizing fieldwork, conducting pretests and pilot surveys, analyzing the survey data, and reporting the findings. Each step is crucial for ensuring that the data collected is reliable, valid, and applicable. Clearly defining objectives ensures the survey addresses specific questions; establishing scope ensures the correct population is studied; subject coverage ensures comprehensive data collection; method selection impacts data accuracy; organization ensures operational efficiency; pretests improve survey design; and data analysis and reporting ensure findings are communicated effectively .
Multiple regression analysis is used to predict the value of a dependent variable by modeling its relationship with two or more independent variables. The technique estimates the parameters of the regression equation, which represent the average change in the dependent variable for a one-unit change in each predictor while holding others constant. By analyzing the coefficients, multicollinearity, and interaction effects, it can predict and provide insights into the relationships between variables, thus aiding in informed decision-making and forecasting .
Index numbers have several practical limitations. They are never cent-percent accurate due to the inherent difficulties in computation. Selecting a representative basket of goods can be problematic due to changes in consumption patterns over time. Different countries may use disparate base years, limiting international comparability. Index numbers also measure only average changes, which can hide specific sectoral performances, and can be influenced by quality changes in goods rather than solely price changes. These factors complicate their use in accurately reflecting true economic changes .
Systematic sampling involves selecting every nth item in a list after a random start, unlike simple random sampling where each unit is selected entirely by random chance from the population. Its advantages include simplicity, ease of implementation, and reduced costs. It often yields a more evenly distributed sample across the population, which can improve representativeness, especially when population elements are naturally ordered. However, it assumes an absence of periodicity within the list that could bias results .
Type I error occurs when a true null hypothesis is incorrectly rejected, whereas Type II error occurs when a false null hypothesis is not rejected. Understanding their trade-off is important because minimizing one often increases the other. The level of significance (α) indicates the probability of a Type I error, while the power of the test (1-β) is the probability of correctly rejecting a false null hypothesis. Balancing these errors is crucial for designing tests that are both accurate and reliable, with implications in research validation and decision accuracy .
Stratified sampling increases the precision of estimates by dividing the population into non-overlapping sub-populations known as strata. Each stratum is more homogeneous compared to the entire population, which reduces variance within each subgroup. Sampling from each stratum ensures that each sub-population is adequately represented, leading to overall reduced variability in the estimate of the population mean. This typically results in higher accuracy and precision compared to a simple random sample of the same size, as it captures important variations within the population .
Cluster sampling is beneficial when a complete sampling frame is not available, making it impractical to conduct a simple random or stratified sampling. It is particularly useful in geographically dispersed populations. By dividing the population into clusters, which can be naturally occurring groups like neighborhoods or schools, and then randomly selecting clusters to assess, it reduces fieldwork and administrative costs. It is preferable when the population is large and spread out, and when budget and time constraints are significant .
Probability sampling methods, such as simple random sampling, ensure that every unit in the population has a known, non-zero probability of being selected, allowing for objective statistical inference about the population. Non-probability sampling methods, including judgement sampling and quota sampling, do not provide this assurance, making statistical inference less reliable. In non-probability sampling, selection is based on subjective judgment rather than randomization, which introduces potential biases and limits generalizability .
The sample mean is considered an unbiased estimator of the population mean because the expected value of the sample mean equals the population mean, irrespective of the population distribution's form. This means when samples of the same size are repeatedly drawn from the population, the average of their means will converge to the true population mean, reflecting no systematic overestimation or underestimation .
Constructing price index numbers involves challenges such as selecting appropriate base years, identifying representative items, collecting accurate price data, and assigning suitable weights. These factors introduce potential biases and discrepancies that can impact the index's reliability. Inaccurate or non-representative index numbers can lead to misinterpretations in economic policies, affecting inflation estimates, cost of living adjustments, and monetary policy decisions. Moreover, changes in consumption patterns and product quality further complicate this construction, potentially overstating or understating economic changes .