0% found this document useful (0 votes)
60 views174 pages

STA 202: Statistics II Course Overview

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
60 views174 pages

STA 202: Statistics II Course Overview

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

OLABISI ONABANJO UNIVERSITY

OPEN AND DISTANCE LEARNING CENTRE


AGO-IWOYE

STA 202: Statistics II


STA202: Statistics II

COURSE DEVELOPMENT TEAM

Prof. T.O. Olatayo – Subject Expert

Prof. D.A. Agunbiade – Course Reviewer

Dr. Oluwakemi Olayemi – Language Editor

Prof. Oyesoji Aremu – ODL Expert

Mr. Moyosola Ayodele – Instructional Designer

Page 2 of 174
STA202: Statistics II
Vice Chancellor’s Message
It is with great pleasure that I welcome you as learners to the Olabisi Onabanjo University
Open and Distance Learning Centre.
Massive and Democratisation of higher education via Open and Distance Learning as
advocated globally has since been one of the goals of Olabisi Onabanjo University
Management, hence, Open and Distance Learning constitutes one of the areas of focus
since my assumption of duty. Through the efforts of the University Governing Council
and Senate, the establishment of the Open and Distance Learning Centre was approved in
July, 2016.

Open and Distance Learning is a mode of study that affords tertiary education
opportunities to all and sundry regardless of age, gender, location, space and other limiting
factors.

Quite a large number of qualified applicants for tertiary education are denied admission
yearly, there are also several others who wish to advance educationally but could not,
because of their job which is their means of livelihood.

Olabisi Onabanjo University via its Open and Distance Learning Centre offers quality,
technology driven, flexible, self-directed and cost effective tertiary education. It is a viable
option for learners who wish to study online from their location and at desired time.

This course material provides learners with vital information relevant to our programme
and schedules. I advise learners to make judicious use of it. I congratulate our Open and
Distance Learning Centre Staff, Department and Faculty for their effort towards the
production of this handbook.

I hope your learning experience with the Olabisi Onabanjo University Open and Distance
Learning Centre is memorable and exciting.

Prof Ganiyu Olatunji Olatunde


Vice Chancellor OOU

Page 3 of 174
STA202: Statistics II
Course Study Guide
Introduction

STA 202 titled Statistics II is a 3-unit course for students studying towards acquiring a
Bachelor of Science in Accounting. The course is divided into 8 study sessions. The
course will introduce you to the basic statistics concept in solving practical problems.
The course study guide therefore gives you an overview of what STA 202 is all about, the
textbooks and other materials to be referenced, what you are expected to know in each
unit and how to work through the course materials. Define a set and understand sampling,
solve correlation and regression of any statistical data and applications of Times Series
Analysis.
Recommended Study Time
This course is a 3 unit course divided into 8 study sessions. You are enjoined to spend at
least 3 hours in studying the content of each study unit
What you are about to learn in this course
The overall aim of this course, STA 202 is to introduce you to sampling theory and
estimation techniques, Simple correlation analysis, Simple regression analysis, Test of
hypothesis, Index numbers and Time series analysis.
Course Aims
This course aims to introduce students to the basic statistical terms. It is expected that the
knowledge will help the reader to effectively use statistical principles to solve even life
problems.
Course Objectives
It is important to note that each unit has specific objectives. You should study them
carefully before proceeding to subsequent units. Therefore, it may be useful to refer to
these objectives in the course of your study of the unit to assess your progress. You should
always look at the unit objectives after completing a unit. In this way, you can be sure that
you have done what is required of you by the end of the unit.

Page 4 of 174
STA202: Statistics II
However, the overall objective of STA 202 is to enable students to analyze and interpret
data collected from a variety of types of research designs, within a linear model
framework.
Working through this course
In order to have a thorough understanding of the course units, you will need to read and
understand the contents, practice the steps by designing and implementing a mini
computer application system for your department and be committed to learning and
implementing your knowledge.
This course is designed to cover approximately fifteen weeks and it will require your
devoted attention. You should do the exercises in the Tutor-Marked Assignments and
submit to your tutors via the Learning Management System (LMS).

Course Materials
The major components of the course are;
1. Course Guide
2. Printed Lecture materials
3. Text Books
4. Interactive DVD
5. Electronic Lecture materials via LMS
6. Tutor Marked Assignments

Assessment
There are two aspects to the assessment of this course. First, there are tutor marked
assignments and second, the written examinations. Therefore, you are expected to take
note of the facts, information and problem solving gathered during the course. The tutor
marked assignments must be submitted to your tutor for formal assessment in accordance
to the deadline given. The work submitted will count for 30% of your total course mark.
At the end of the course, you will need to sit for a final written examination. This
examination will account for 70% of your total score. You will be required to submit some

Page 5 of 174
STA202: Statistics II
assignments by uploading them to STA 202 page on the Learning Management System
(LMS).
Tutor-Marked Assignment (TMA)
There are TMAs in this course. You need to submit all the TMAs. The best 10 will
therefore be counted. When you have completed each assignment, send them to your tutor
as soon as possible and make certain that it gets to your tutor on or before the stipulated
deadline. If for any reason you cannot complete your assignment on time, contact your
tutor before the assignment is due to discuss the possibility of extension. Extension will
not be granted after the deadline, unless on extraordinary cases.
Final Examination and Grading
The final examination for STA 202 will last for a period not more than 2hours and has a
value of 70% of the total course grade. The examination will consist of questions which
reflect the Self-Assessment Questions (SAQs), In-text Questions (ITQs), some applied
questions and tutor marked assignments that you have previously encountered.
Furthermore, all areas of the course will be examined. It would be better to use the time
between finishing the last unit and sitting for the examination to revise the entire course.
You might find it useful to review your TMAs and comment on them before the
examination. The final examination covers information from all parts of the course. Most
examinations will be conducted via Computer Based Testing (CBT)
Tutors and Tutorials
There are few hours of face-to-face tutorial provided in support of this course. You will be
notified of the dates, time and location together with the name and phone number of your
tutor as soon as you are allocated a tutorial group. Your tutor will mark and comment on
your assignments, keep a close watch on your progress and on any difficulties you might
encounter and provide assistance to you during the course. You must submit your tutor
marked assignment to your tutor well before the due date. At least two working days are
required for this purpose. They will be marked by your tutor and returned as soon as
possible via the same means of submission.

Page 6 of 174
STA202: Statistics II
Do not hesitate to contact your tutor by telephone, e-mail or discussion board if you need
help. The following might be circumstances in which you would find help necessary:
contact your tutor if:
 You do not understand any part of the study units or the assigned readings.
 You have difficulty with the self-test or exercise.
 You have questions or problems with an assignment, with your tutor’s comments on an
assignment or with the grading of an assignment.
You should endeavour to attend the tutorials. This is the only opportunity to have face-to-
face contact with your tutor and ask questions which are answered instantly. You can raise
any problem encountered in the course of your study. To gain the maximum benefit from
the course tutorials, have some questions handy before attending them. You will learn a
lot from participating actively in discussions.
Good luck!

Recommended Texts
The following texts and Internet resource links will be of enormous benefit to you in
learning this course:

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Page 7 of 174
STA202: Statistics II
Table of Contents
Vice Chancellor’s Message .................................................................................................. 3
Course Study Guide .............................................................................................................. 4
Introduction....................................................................................................................... 4
Table of Contents .................................................................................................................. 8
Study Session 1: Sampling Theory and Estimation Techniques ........................................ 14
Introduction..................................................................................................................... 14
Learning Outcomes for Study Session 1 ........................................................................ 14
1.1 Definition of Sampling ........................................................................................ 15
1.1.1 Sample Survey .............................................................................................. 15
1.1.2 Advantages of Sample Survey...................................................................... 15
1.1.3 Disadvantages of Sample Survey ................................................................. 16
1.1.4 Sampling Frame ............................................................................................... 16
1.2 Type of Sampling Method ................................................................................... 17
1.2.1 Sampling with Replacement ......................................................................... 18
1.2.2 Sampling without Replacement .................................................................... 18
1.3 Types of Sampling Technique ............................................................................. 18
1.3.1 Non-Probability Sampling ............................................................................ 19
1.3.2 Judgment Sampling ...................................................................................... 20
1.3.3 Quota Sampling ............................................................................................ 20
1.3.4 Haphazard Sampling..................................................................................... 20
1.3.5 Probability Sampling .................................................................................... 20
1.3.6 Simple Random Sampling (SRS) ................................................................. 21
1.3.7 Systematic Sampling .................................................................................... 21
1.3.8 Stratified Sampling ....................................................................................... 22
1.3.9 Multi-Stage Sampling ................................................................................... 22
1.3.10 Cluster Sampling .......................................................................................... 23
1.3.11 Steps in Planning a Sample Survey .............................................................. 23
1.4 Sampling Distribution of the Sample Mean ........................................................ 24
Summary of Study Session 1 .......................................................................................... 26
Self-Assessment Questions (SAQs) for Study Session 1 ............................................... 27
Glossary of Terms........................................................................................................... 28

Page 8 of 174
STA202: Statistics II
References....................................................................................................................... 29
Study Session 2: Simple Correlation Analysis ................................................................... 30
Introduction..................................................................................................................... 30
Learning Outcomes for Study Session 2 ........................................................................ 30
2.1 Correlation Analysis ............................................................................................ 31
2.1.1 Karl Pearson’s’ Product Moment Correlation Coefficient ........................... 31
2.1.2 Spearman Rank Correlation Coefficient....................................................... 35
2.2 Tie in Ranks ......................................................................................................... 37
Summary of Study Session 2 .......................................................................................... 39
Self-Assessment Questions (SAQs) for Study Session 2 ............................................... 40
Glossary of Terms........................................................................................................... 42
References....................................................................................................................... 43
Study Session 3: Simple Regression Analysis ................................................................. 44
Introduction..................................................................................................................... 44
Learning Outcomes for Study Session 3 ........................................................................ 44
3.1 Simple Regression ............................................................................................... 45
3.2 The Least Squares Method .................................................................................. 45
Summary of Study Session 3 .......................................................................................... 51
Self-Assessment Questions (SAQs) for Study Session 3 ............................................... 52
Glossary of Terms........................................................................................................... 53
References....................................................................................................................... 54
Study Session 4: Test of Hypothesis ................................................................................ 55
Introduction..................................................................................................................... 55
Learning Outcomes for Study Session 4 ........................................................................ 55
4.1 Meaning of Test of Hypothesis ............................................................................ 56
4.2 Type I and Type II Errors .................................................................................... 57
4.2.1 One and Two Tailed Test ............................................................................. 57
4.2.2 Test Procedure and Steps .............................................................................. 58
4.2.3 Test Concerning the Mean (For Large Sample) ........................................... 58
4.2.4 Test Concerning Means (Small Samples) .................................................... 60
4.2.5 Test Concerning Two Population Means (Large Sample) ........................... 64
4.2.6 Test Statistics ................................................................................................ 66

Page 9 of 174
STA202: Statistics II
4..2.7 Test Concerning Two Population Means (Small Sample) ........................... 66
Summary of Study Session 4 .......................................................................................... 70
Self-Assessment Questions (SAQs) for Study Session 4 ............................................... 71
Glossary of Terms........................................................................................................... 72
References....................................................................................................................... 73
Study Session 5: Index Numbers ..................................................................................... 74
Introduction..................................................................................................................... 74
Learning Outcomes for Study Session 5 ........................................................................ 74
5.1 Meaning of Index Numbers ................................................................................. 75
5.1.1 Features of Index Numbers........................................................................... 75
5.1.2 Steps or Problems in the Construction of Price Index Numbers .................. 76
5.2 Construction of Price Index Numbers ................................................................. 78
5.2.1 Simple Aggregative Method......................................................................... 79
5.2.2 Simple Average of Price Relatives Method ................................................. 79
5.2.3 Weighted Aggregative Method .................................................................... 80
5.2.4 Weighted Average of Relatives Method ...................................................... 82
5.3 Difficulties in Measuring Changes in Value of Money ....................................... 84
5.3.1 Conceptual Difficulties ................................................................................. 84
Practical Difficulties ................................................................................................... 85
5.4 Types of Index Numbers...................................................................................... 87
5.4.1 Wholesale Price Index Numbers .................................................................. 88
5.4.2 Retail Price Index Numbers.......................................................................... 88
5.4.3 Cost-of-Living Index Numbers .................................................................... 88
5.4.4 Working Class Cost-of-Living Index Numbers ........................................... 89
5.4.5 Wage Index Numbers ................................................................................... 89
5.4.6 Industrial Index Numbers ............................................................................. 89
5.4.7 Uses of Index Number in the Economic Field ............................................. 89
5.4.8 Limitations of Index Numbers...................................................................... 90
Summary of Study Session 5 .......................................................................................... 91
Self-Assessment Questions (SAQs) for Study Session 5 ............................................... 92
Glossary of Terms........................................................................................................... 93
References...................................................................................................................... 94

Page 10 of 174
STA202: Statistics II
Study Session 6: Time Series Analysis ............................................................................ 95
Introduction..................................................................................................................... 95
Learning Outcomes for Study Session 6 ........................................................................ 95
6.1 Definition of Time Series .................................................................................... 96
6.1.1 Components of Time Series ......................................................................... 96
6.1.2 Time Series Model ....................................................................................... 96
6.2 Methods of Estimating Trend .............................................................................. 97
6.2.1 Moving Average Method ............................................................................. 98
6.2.2 Least Square Method .................................................................................. 101
6.3 Mathematical Representation ............................................................................ 101
Summary of Study Session 6 ........................................................................................ 105
Self-Assessment Questions (SAQs) for Study Session 6 ............................................. 106
Glossary of Terms......................................................................................................... 107
References..................................................................................................................... 108
Study Session 7: Multiple Regression .............................................................................. 109
Introduction................................................................................................................... 109
Learning Outcomes for Study Session 7 ...................................................................... 109
7.1 Introduction............................................................................................................. 110
7.1.1 First-Order Model with More than Two Predictor Variables ..................... 111
7.1.2 General Linear Regression Model .............................................................. 112
7.1.3 Qualitative Predictor Variables .................................................................. 113
7.1.4 Polynomial Regression ............................................................................... 113
7.1.5 Variables Transformation ........................................................................... 114
7.1.6 General Linear Regression Model in Matrix Terms................................... 115
7.2 Estimation of Regression Coefficients .............................................................. 117
7.2.1 Global Test: Testing Whether the Multiple Regression Model is Valid .... 120
7.2.2 Evaluating Individual Regression Coefficients .......................................... 121
7.2.3 Testing Individual Regression Coefficients ............................................... 122
7.2.4 Qualitative Independent Variables (Dummy Variables) ............................ 122
7.2.5 Dummy Variable ........................................................................................ 123
7.3 Polynomial Regression Model ............................................................................. 128
7.3.1 Uses of Polynomial Models........................................................................ 129

Page 11 of 174
STA202: Statistics II
7.3.2 One Predictor Variable-Second Order ........................................................ 129
7.3.3 One Predictor Variable-Third Order........................................................... 130
7.3.4 Two Predictor Variables-Second Order ..................................................... 130
7.3.5 Fitting Polynomial in One Variable ........................................................... 131
7.4 Analysis ............................................................................................................. 132
7.4.1 Test of Significance: ................................................................................... 135
Summary of Study Session 7 ........................................................................................ 138
Self-Assessment Questions (SAQs) for Study Session 7 ............................................. 139
Glossary of Terms......................................................................................................... 140
References..................................................................................................................... 141
Study Session 8: Partial Correlation .............................................................................. 142
Introduction................................................................................................................... 142
Learning Outcomes for Study Session 8 ...................................................................... 142
8.1 Simple Correlation Coefficient ............................................................................... 143
8.1.2 The Multiple Regression ............................................................................ 144
8.2 Partial Correlation Coefficient ........................................................................... 147
8.3 Tests for Association ......................................................................................... 150
8.3.1 Contingency Tables .................................................................................... 151
8.3.2 Expected Frequencies ................................................................................. 151
8.3.3 Expected Frequencies in Contingency Tables ............................................ 152
8.3.4 Test Statistic ............................................................................................... 153
8.3.5 𝝌𝟐 Test of Association .............................................................................. 153
8.3.6 Performing the Test ................................................................................... 154
8.3.7 Goodness-of-Fit Tests ................................................................................ 156
8.3.8 Observed and Expected Frequencies .......................................................... 157
8.3.9 Expected Frequencies in Goodness-of-Fit Tests ........................................ 157
8.3.10 The Goodness-of-Fit Test ........................................................................... 158
Summary of Study Session 8 ........................................................................................ 162
Self-Assessment Questions (SAQs) for Study Session 8 ............................................. 163
Glossary of Terms......................................................................................................... 164
References..................................................................................................................... 165
Notes on Self-Assessment Questions (SAQs) .................................................................. 166

Page 12 of 174
STA202: Statistics II
Notes on Self-Assessment Questions for Study Session 1 ........................................... 166
Notes on Self-Assessment Questions for Study Session 2 ........................................... 167
Notes on Self-Assessment Questions for Study Session 3 ........................................... 167
Notes on Self-Assessment Questions for Study Session 4 ........................................... 169
Notes on Self-Assessment Questions for Study Session 5 ........................................... 171
Notes on Self-Assessment Questions for Study Session 6 ........................................... 172

Page 13 of 174
STA202: Statistics II
Study Session 1: Sampling Theory and Estimation Techniques

Introduction

Each day, we observe the high, low, and close of stock market indexes from
around the world. Indexes such as Dangote Cement Index, Nestle Nigeria Index
and Stanbic IBTC Holdings Index are samples of stocks. Although Dangote
Cement, Nestle Nigeria and Stanbic IBTC Holdings do not represent the populations of
Nigerian stocks, we view them as valid indicators of the whole population’s behavior. As
analysts, we are accustomed to using this sample information to assess how various
markets from around the world are performing. Any statistics that we compute with
sample information, however, are only estimates of the underlying population parameters.
A sample, then, is a subset of the population—a subset studied to infer conclusions about
the population itself.

Learning Outcomes for Study Session 1

On completion of this study session, you should be able to:

1.1 Understand what sampling is


1.2 Differentiate between the types of sampling method
1.3 Explain the types of sampling techniques
1.4 Define sampling distribution of the sample mean

Page 14 of 174
STA202: Statistics II
1.1 Definition of Sampling

Sampling is a scientific method of selecting and using a representative part (sample) of a


whole to seek the truth about the whole. Sampling is used extensively, consciously or
unconsciously, in everyday life to obtain the required information or to carry out a course
of action. For example, a housewife uses a small quality of soup from a pot of soup to
ascertain whether there is adequate salt in the soup. A medical doctor or a laboratory
scientist, by using a few milliliters of a patient’s blood sample obtains the quantity of
malaria parasites in the blood system. Also an agriculturist uses little quantity of soil
sample to measure the amount of nutrient in a farmland. A market researcher makes use of
a fraction of consumer of a product to gauge the acceptability of the product.
In all of the above illustrations an alternative method of obtaining the necessary
information is to inspect or survey the whole items called population.
Therefore, population consists of the entire objects in a defined area of interest and sample
is that part of the population in which further analysis is sought, to make inference about
the population. It may be defined as a selected part or a subset of a population obtained
with the objective of investigating population parameters.

1.1.1 Sample Survey

A sample survey can be defined as the collection and examination of data from a sample
in order to make inference about the whole or entire population. This is different from
census, which is a complete enumeration or survey of the whole units of enquiry. Hence,
sample survey theory deals with the process of sample selection, data collection,
estimation of the population characteristics using the sample data so collected and
determining the accuracy of the estimates.

1.1.2 Advantages of Sample Survey

i. Sample survey save time, labour and cost, especially when these resources are
limited than in census where a large number of these resources is needed.

Page 15 of 174
STA202: Statistics II
ii. Data can be collected and analysed quickly from a sample than from the entire
population. In a destructive investigation, such as determination of the
effectiveness or a newly produced drug by a pharmaceutical company, it is better
to use sample survey rather than complete enumeration.
iii. There is a greater and more efficient supervision of field staff in a sample survey
resulting in the collection of more reliable data. Sample survey makes for the use
of better-qualified staff and specialized equipment.
iv. It has greater subject coverage and less observational error than census. It is
possible to estimates the discrepancies between the sample and population values
in a sample survey.
v. Sample survey is used in conjunction with census in post enumeration checks in
order to determine the census coverage and accuracy of the census data.

1.1.3 Disadvantages of Sample Survey

i. Sample survey is not appropriate when information is required for every unit of
enquiry for example, compilation of voters list.
ii. It is less accurate in small area classification where data are needed for each
subdivision of the population.
iii. Sample survey cannot be used where the sampling frame is inadequate or
completely absent.
iv. It breeds sampling error which if not properly handled, may affect the accuracy of
the result. Finally, sample survey can be manipulated to suit the purpose of the
investigator.

1.1.4 Sampling Frame

A sampling frame contains the basic details of all members of the population from which
samples are to be drawn. It is generally believed by experts that without a complete
sampling frame, a truly random sample cannot be selected. Properties of a good frame are
completeness, no duplication, adequacy and clear identifiability of elements.

Page 16 of 174
STA202: Statistics II
In-Text Questions (ITQs)

Define the following


i. Sampling
ii. Sample survey
iii. Sampling frame

In-Text Answers (ITAs)

i. Sampling is a scientific method of selecting and using a representative part


(sample) of a whole to seek the truth about the whole
ii. A sample survey can define as the collection and examination of data from a
sample in order to make inference about the whole or entire population
iii. A sampling frame contains the basic details of all members of the population from
which samples are to be drawn.

1.2 Type of Sampling Method

There are two types of sampling method, these are sampling with replacement and
sampling without replacement.

sampling with
replacement

sampling
without
replacement

Sampling Method

Page 17 of 174
STA202: Statistics II
Fig 1.1: Types of Sampling Method

1.2.1 Sampling with Replacement

This is method of sampling where by a sample drawn is replaced before the subsequent.
This can occur where each member may be chosen more than once.

1.2.2 Sampling without Replacement

These can also occur when a sample element in a set of population is drawn for the first
time and it cannot be drawn again or cannot occur again.

In-Text Questions (ITQs)

What are the types of sampling method?

In-Text Answers (ITAs)

Sampling with replacement and sampling without replacement

1.3 Types of Sampling Technique

There are two main types of sampling, namely; Probability (random) sampling and non-
probability (non-random) sampling. In probability sampling, every unit of enquiry in the
population has a known non-zero probability of being included in the sample. In non-
probability, the probability of selecting a unit in the sample is not known and cannot be
determined. Hence, statistical inference cannot be made objectively about the population
in non-random sampling. Hence we shall discuss briefly some examples of non-

Page 18 of 174
STA202: Statistics II
probability sampling and probability sampling respectively.

non-
probability
sampling

Sampling
Technique

Fig 1.2: Types of Sampling Technique

1.3.1 Non-Probability Sampling

Types of non-probability sampling techniques are discussed below

Judgement
sampling
Stratified Quota
Sampling Sampling

Systematic Haphazard
Sampling Sampling
Non-
probability
sampling

Simple
Multi Stage
Random
Sampling
Sampling

Probability Cluster
Sampling Sampling

Fig 1.3: Types of Non-Probability Sampling

Page 19 of 174
STA202: Statistics II
1.3.2 Judgment Sampling

In this type of non- probability sampling experts pick units they judge to be representative
of the entire population. For example, a typical town may be picked to represent an urban
population or a village chosen to represent a rural population. Since experts differ in their
judgements, different experts may choose different units, which to them are more
preferable. Confidence is therefore not placed on the results from judgement sampling,
since there is no objective method for preferring one judgement to another. Finally,
accurate measure of sample variability cannot be obtained with judgement sampling.

1.3.3 Quota Sampling

In this type of non-random sampling, sampling is set up according to some specified


characteristics such as, gender, age, social class and many more is assigned to an
enumerator. The enumerator then selects the required representative number of
individuals to be in the sample within each quota (group). This is in a way similar to
judgement sampling.
It is useful in obtaining quickly the people’s reaction to current issues. Quota sampling can
be used even where there is no suitable list of population units.

1.3.4 Haphazard Sampling

In this type of non-random sampling, units are taken into the sample as they come along or
make themselves available. In a hospital survey on the characteristics of patients suffering
from a particular disease, the patients are taken into the sample in the order they report to
the hospital and volunteer to be part of the study. This sampling method lacks
representativeness of the population to be studied.

1.3.5 Probability Sampling

With probability sampling, each population’s element has a known and non-zero
probability of being selected. Samples of probability sampling are:

Page 20 of 174
STA202: Statistics II
1.3.6 Simple Random Sampling (SRS)

This is sampling procedure in which every member of the population has equal chance of
being selected as a member of the sample. It is mostly adopted in homogeneous (of same
kind) population. For example, if we have five coloured balls, red, orange, yellow, blue
and green, the chance of any of them being picked or selected is fraction of one and five,
that is, 1/5. Other examples are tossing of coin or rolling a die etc. All the above examples
are however, not perfectly objective. The most objective and widely used method is the
random number table. This is a table consisting of randomly allocated numbers that are
arranged in rows and columns of a standard statistical table.

1.3.7 Systematic Sampling

Systematic sampling is another method of selecting a sample in such a way that every unit
in the population will have an equal chance of being selected in the sample. In this
selection procedure once the first unit is randomly selected the remaining units are
automatically selected.
Suppose a sample of n units is to be selected from N units in the population.
Let
N
K  , k is an integer.
n
Select a random number between 1 and k inclusive. Suppose the random number selected
is r; add k to the random and select r successively until n numbers are obtained. The
sample data then consists of n units with serial numbers r, r+k, r+2k, …, r+(n-1)k. Thus
the sample consists of the first unit selected at random and every Kth unit thereafter.
This procedure of sample selection is called systematic sampling and the sample is called
1 n
systematic sample. K is called the sampling interval;  is the sampling fraction.
K N
One of the advantages of systematic sampling is that it is easy and more convenient to
apply than simple random sampling without replacement (SRSWOR), especially in a large
scale sampling.

Page 21 of 174
STA202: Statistics II
In systematic sampling it is not necessary to serially number the units before selection or
even know the population size exactly in advance.
Application of intervals is easier to cross-check than the use of random numbers especially
if the selection is done in the field by the field enumerators. This helps to control the
field workers and minimize errors in sampling.
Systematic sample may be as precise as a simple random sample without replacement in a
population in which units are in random order. Systematic sampling yields an evenly
distributed sample.

1.3.8 Stratified Sampling

This is a sampling procedure which consists of stratifying (dividing) the population into a
number of non-overlapping sub-populations (strata), then taking a sample from each
stratum. The items or sample from each stratum can then be selected by any suitable
random method.
Stratification does not guarantee good results, but if successful and properly executed, a
stratified sample will generally lead to a higher degree of precision, or reliability, then a
simple random sample of the same size drawn from the population.

1.3.9 Multi-Stage Sampling

This sampling procedure involves more than one stage. The first stage consists of breaking
down the population into sets of distinct groups, from these, a number of groups are
selected. Each group selected broken down into units from which a sample is taken. If we
stop at this point we have a two-stage sampling.
Further stages may be added and the number of stages involved is denoted in the name of
the sampling. For example, five stage sampling, means that five stages are involved. The
population is distributed into a number of first stage sampling units and a sample is taken
of these stage unit by some suitable method.

Page 22 of 174
STA202: Statistics II
1.3.10 Cluster Sampling

All sampling procedures so far discussed depend heavily on complete sampling frame.
Unfortunately, complete sampling frame is not always available. Cluster sampling was
developed to take care of this inadequacy. In this kind of sampling, the total population is
divided into a number of relatively small subdivisions, which are themselves clusters of
still smaller units, and then some of these sub-divisions or clusters, are randomly selected
for inclusion in the overall sample. If the clusters are geographical subdivisions this kind
of sampling is called area sampling and a cluster can be a household with members of the
household as elements.

1.3.11 Steps in Planning a Sample Survey

The basic steps in planning and execution a sample survey are:


i. Objective of the survey: the objective of the survey must be stated in clear,
concrete and concise terms.
ii. Scope: the population to be covered must be defined in terms of content,
units, extent, and time.
iii. Subject coverage: a detailed description of the items of information
collected should be given.
iv. Method of data collection
v. Organization of field works.
vi. Pretest and pilot survey
vii. Analysis of survey data.
viii. Report writing.

Page 23 of 174
STA202: Statistics II
In-Text Questions (ITQs)

i. What are the types of sampling method?


ii. What are the types of non-probability sampling?

In-Text Answers (ITAs)

i. Probability and non-probability sampling


ii. Judgement sampling, Quota sampling, Haphazard Sampling, Probability
Sampling, Simple Sampling, Systematic Sampling, Stratified Sampling, Multi-Stage
Sampling and Cluster Sampling

1.4 Sampling Distribution of the Sample Mean

The sample mean is referred to as the point estimate of the population mean. The
arithmetic mean (𝜇𝑥̅ ) of the sampling distribution of mean values is equal to the
population mean (𝜇) regardless of the form of the population distribution that is, 𝜇𝑥̅ = 𝜇.
The sample mean is then said to be an unbiased estimate of the population mean.

Case Study 1.1


A population consists of the numbers 0, 2, 4 and 6. The population mean, 𝜇 = 3. Consider
all the possible samples of size 2 without replacement from the population and show that
the sample mean is an unbiased estimator.
Solution
The population mean is

0 + 2 + 4 + 6 12
𝜇= = =3
4 4

Page 24 of 174
STA202: Statistics II
The possible samples of size 2 and their mean from the population is
Sample Number Sample elements Sample mean
1 0,2 1
2 0,4 2
3 0,6 3
4 2,4 3
5 2,6 4
6 4,6 5
Table 1.1
The arithmetic mean of the sampling distribution of the mean value is
1+2+3+3+4+5
𝜇𝑥̅ = =3
6
Since 𝜇𝑥̅ = 𝜇, the sample mean is then an unbiased estimate of the population mean.

Page 25 of 174
STA202: Statistics II

Summary of Study Session 1

In study session 1, you have learnt that:

1. Sampling is used extensively, consciously or unconsciously, in everyday life to


obtain the required information or to carry out a course of action.
2. The two types of sampling method are sampling with replacement and sampling
without replacement.
3. The types of probability sampling techniques are: Judgement sampling, Quota
sampling, Haphazard Sampling, Probability Sampling, Simple Sampling,
Systematic Sampling, Stratified Sampling, Multi-Stage Sampling and Cluster
Sampling.
4. The arithmetic mean (𝜇𝑥̅ ) of the sampling distribution of mean values is equal to
the population mean (𝜇) regardless of the form of the population distribution that
is, 𝜇𝑥̅ = 𝜇

Page 26 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 1

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. Why would you prefer sample to population?


2. Explain the following terms:
i. Simple random sampling
ii. Cluster sampling
iii. Systematic sampling
3. What is different between probability and non-probability sampling?
4. A population consists of the numbers 0, 2, 4, 6 and 8. Consider all the possible
samples of size 2 without replacement from the population. Hence, show that the
sample mean is an unbiased estimator.
5. What do you understand by Sampling frame?
6. Define Sampling survey.

Page 27 of 174
STA202: Statistics II

Glossary of Terms

Theory: a supposition or a system of ideas intended to explain something, especially one


based on general principles independent of the thing to be explained.
Estimation: the process by which one makes inferences about a population, based on
information obtained from a sample
Census: study of every unit, everyone or everything, in a population. It is known as a
complete enumeration, which means a complete count.
Enumerator: survey personnel charged with carrying out that part of
an enumeration consisting of the counting and listing of people or assisting respondents in
answering the questions and in completing the questionnaire.
Precision: how close estimates from different samples are to each other

Page 28 of 174
STA202: Statistics II

References

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 29 of 174
STA202: Statistics II
Study Session 2: Simple Correlation Analysis

Introduction

We know that when the price of a product increases its demand will decrease. And
to the contrary quality supplied will increase with the increase of price. These are
nothing but correlation. Correlation quantifies the extent to which two
quantitative variables, X and Y, “go together.” When high values of X are associated
with high values of Y, a positive correlation exists. When high values of X are associated
with low values of Y, a negative correlation exists.

Learning Outcomes for Study Session 2

On completion of this study session, you should be able to:

2.1 Explain correlation analysis

2.2 Break tie in ranks

Page 30 of 174
STA202: Statistics II
2.1 Correlation Analysis

Correlation analysis is a technique for estimating the closeness or degree of relationship


between two or more variables. Correlation is the degree of association between two or
more variables. The degree of relationship may be positive that is, an increase in one
variable accompanied by an increase in the other or negative when decrease in one
variable is accompanied by an increase in the other. The patterns of correlation are perfect
and positive correlation when r = 1, perfect and negative correlation when r = - 1, positive
correlation when r > 0, negative correlation when r <0 and no correlation when r = 0.
The correlation coefficient or coefficient of correlation denoted by r, is a measure of the
strength of the linear relationship between two variables. Two types of the measures of
correlation are:
i. Karl Pearson’s’ product moment correlation coefficient (r)
ii. Spearman’s rank correlation coefficient (R)

Spearman’s Karl
rank Pearson’s’
correlation product
coefficient moment
(R) correlation
coefficient
(r)

Fig 2.1: Types of Correlation analysis

2.1.1 Karl Pearson’s’ Product Moment Correlation Coefficient

The Karl Pearson’s product moment correlation coefficient is devoted by r and given by:

Page 31 of 174
STA202: Statistics II
n xy  x  y
r
[n x 2  ( x) 2 ][n y 2  ( y ) 2 ]

where – 1 < r <1

It should be noted that the higher the magnitude of r, the stronger the association.

Case Study 2.1

The table below is used to present Nigerian Government income (x) and Nigerian
Government expenditure (y) for a period of 12 months. This is given in million naira.

Month(s) Nigerian Government Nigerian Government expenditure


income
1 11.50 11.25
2 9.50 11.75
3 13.00 11.75
4 15.50 12.50
5 12.50 12.50
6 11.50 12.75
7 9.00 9.50
8 11.50 10.75
9 9.25 11.00
10 9.75 9.50
11 14.25 13.00
12 10 12.00

Table 2.1

Page 32 of 174
STA202: Statistics II
Calculate the coefficient of correlation

Solution

x y
 xy - n
r
( y ) 2
[ x - ( x ) [ y -
2 2 2

 x  138.00, y  138.25,  x  1608.12 ,  x 2


 1632.75,  y 2 1602.81

138.00 x138.25
1608.12 
r 12
2
(138.00) (138.25)
(1632.75  )(1607.81 
12 12

r = 0.70 (to 2 decimal places)

There is a significant relationship between Nigerian Government expenditure and


government income.

Case Study 2.2


Calculate the correlation coefficients between the following pairs of variables: Nigerian
Government income and inflation and gross domestic product over 10years given in the
table below:
Child number
1 2 3 4 5 6 7 8 9 10 Total
Govt. Income 1.0 3.0 2.5 4.5 1.5 2.0 3.1 4.1 2.5 4.2 28.4
(y)
Inflation (x) 2.0 3.5 3.0 5.0 2.1 2.5 3.6 3.8 3.0 4.0 32.5
GDP (Z) 1.0 6.0 4.0 10.0 2.0 9.0 7.0 8.0 5.0 9.0 61.0
XY 2.0 10.5 7.5 22.5 3.15 50 11.16 15.58 7.5 16.8 101.69

Page 33 of 174
STA202: Statistics II
YZ 1.0 18.0 10.0 45.0 3.0 18.0 21.7 32.8 12.5 37.8 199.80
XZ 2.0 21.0 12.0 50.0 4.2 7.5 22.5 30.4 15.0 36.0 218.30
Table 2.2
Solution

 x  113.31, y  93.06,  z  457,  xy  101.69,


2 2 2

 xz  218.3, yz  199.80,  x  32.5,  y  28.4,

 z  61.0, X  3.25, Y  2.84, Z  6.10


 xy  ( 
x y)
rxy  n
 x
 x 2  
 2

  y 
2
 y 
2


 n   n 

101.69 
32.5(28.4)
 10
32.52  28.4 2 
[113.31  93.06  
10  10 

9.39
  0.9620
9.76
(x(z )
 xz  n
rxz 
 2  x   2  .z 2
2

 x     z  

  n   n 

218.30 
32.5(61.0)
 10
 32.5 2   612 
113.31   457  
 10   10 

20.05
  0.7849
25.5436

Page 34 of 174
STA202: Statistics II

ryz  199.80 
28.461.0
10 26.56
  0.8185
 28.4 2   612  32.45
93.06   457  
 10   10 

Correlation coefficients are higher. It can be concluded that (i) Government income and
inflation are positively correlated. (ii) Government income and inflation are positively
correlated with GDP.

2.1.2 Spearman Rank Correlation Coefficient

When variables do not follow normal distribution and one desires to assess the
relationship, correlation coefficient known as spearman rank correlation coefficient is
used. The variables are ranked based on the magnitude. The correlation between ranks of
variables x and y is obtained. The symbol used is R, the formula is:

6 d i2
R  1

n n 2 1 
where d is the difference between ranks given to the variables of each pair and n is the
number of pairs studied. The procedure was developed by spearman. Hence, it is known as
spearman rank correlation coefficient. Its value also ranges from – 1 to 1.

Case Study 2.3

Calculate the value correlation coefficient between the corresponding values of income
(X) and expenditure (Y) of a known company given below.

X 22 24 25 16 28 19

Y 48 42 40 38 47 45

Table 2.3
Solution

Page 35 of 174
STA202: Statistics II
The varying is in ascending order of magnitude

X Y RX RY d d2

22 48 3 6 -3 9

24 42 4 3 1 1

25 40 5 2 3 9

16 38 1 1 0 0

28 47 6 5 1 1

19 45 2 4 -2 4

24

Table 2.4

6 d 2
R = 1-
n(n 2  1)

6(24)
=1- = 1 - 0.6857 =0.3143
6(36  1)

R = 0.31

There is a low or weak positive correlation between the two variables.

Page 36 of 174
STA202: Statistics II
In-Text Questions (ITQs)

i. Define Correlation analysis


ii. What are the two types of the measure of correlation

In-Text Answers (ITAs)

i. Correlation analysis is a technique for estimating the closeness or degree of


relationship between two or more variables
ii. Karl Pearson’s’ product moment correlation coefficient (r) and Spearman’s rank
correlation coefficient (R)

2.2 Tie in Ranks

Most times, two or more values of a variable might be equal. In such cases, we assign to
each of the tied observations the mean of the ranks which they jointly occupy. For
example if the 5th and 6th largest values of a variable are equal, we assign to each the rank
(5  6)
=5.5, and if the of fifth, smith and seventh largest values of a variable are the
2
(5  6  7 )
same we assign each the rank =6.6
3
Study 2.4
The table give below shows the respective yearly money deposit in two known
commercial banks over the period of a year in billons

Father (  ) 66 64 68 65 69 63 71 67 69 68 70 72

Sons (  ) 69 67 69 66 70 67 69 66 72 68 69 71
Table 2.5

Page 37 of 174
STA202: Statistics II
Calculate the coefficient of rank correlation and comment on the degree of correlation
between the two known commercial banks.

Solution
  RX RX D= RX - RY d2
66 69 4 7.5 -3.5 12.25
64 67 2 3.5 -1.5 2.25
68 69 6.5 7.5 -1.0 1.00
65 66 3 1.5 1.5 2.25
69 70 8.5 10 -1.5 2.25
63 67 1 3.5 -2.5 6.25
71 69 11 7.5 3.5 12.25
67 66 5 1.5 3.5 12.25
69 72 8.5 1.5 -3.5 12.25
68 68 6.5 1.2 1.5 2.25
70 69 10 5 2.5 6.25
72 71 12 11 1.0 1.00
72.50
Table 2.6
6d 2
R=1-
n(n 2  1)

6(72.50)
=1-
12(144  1
72.50
=1-
2(143)
=1-0.2535
=0.7465
=0.75
Comment: There is a fairly high positive correlation between the money deposit of the tow
known commercial bank.

Page 38 of 174
STA202: Statistics II

Summary of Study Session 2

In study session 2, you have learnt that:

1. The degree of relationship may be positive that is, an increase in one variable
accompanied by an increase in the other or negative when decrease in one variable
is accompanied by an increase in the other.
2. The patterns of correlation are perfect and positive correlation when r = 1, perfect
and negative correlation when r = - 1, positive correlation when r > 0, negative
correlation when r <0 and no correlation when r = 0.
3. Two types of the measures of correlation are:
i. Karl Pearson’s’ product moment correlation coefficient (r)
ii. Spearman’s rank correlation coefficient (R)
4. The Karl Pearson’s product moment correlation coefficient is devoted by r and
given by:

n xy  x  y
r
[n x 2  ( x) 2 ][n y 2  ( y ) 2 ]

where – 1 < r <1

6 d i2
R  1
 
5. Spearman rank correlation coefficient formula is
n n 2 1
6. Tie in ranks is settled by assigning to each of the tied observations the mean of the
ranks which they jointly occupy.

Page 39 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 2

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

Calculate the product moment correlation coefficient for the pair given below and
interpret your result.

1.

Month(s) Inflation rate Interest rate


1 11.25 11.50
2 11.75 9.50
3 11.75 13.00
4 12.50 15.50
5 12.50 12.50
6 12.75 11.50
7 9.50 9.00
8 10.75 11.50
9 11.00 9.25
10 9.50 9.75
11 13.00 14.25
12 12.00 10.0

2. Explain the term Correlation analysis

3. The table give below shows the respective income (X) and expenditure (Y) (in million
naira) of O.O.U, Ago-Iwoye for a year.

Page 40 of 174
STA202: Statistics II

Income (  ) 66 64 68 65 69 63 71 67 69 68 70 72

Expenditure (  ) 69 67 69 66 70 67 69 66 72 68 69 71

Calculate the coefficient of rank correlation and comment on the degree of correlation
between income (X) and expenditure (Y) (in million naira) of O.O.U, Ago-Iwoye.

Page 41 of 174
STA202: Statistics II

Glossary of Terms

Coefficient: a numerical or constant quantity placed before and multiplying the variable
in an algebraic expression
Linear: Arranged in or extending along a straight or nearly straight line
Magnitude: Size
Correlation analysis: statistical method that is used to discover if there is a relationship
between two variables/datasets, and how strong that relationship may be
Quantitative variables: any variables where the data represent amounts (e.g. height,
weight, or age)

Coefficient of Correlation: a statistical measure of the strength of the relationship


between the relative movements of two variables

Page 42 of 174
STA202: Statistics II

References

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 43 of 174
STA202: Statistics II
Study Session 3: Simple Regression Analysis

Introduction

Regression models describe the relationship between variables by fitting a line


to the observed data. Linear regression models use a straight line, while logistic
and nonlinear regression models use a curved line. Regression allows you to
estimate how a dependent variable changes as the independent variable(s) change.
Simple linear regression is used to estimate the relationship between two quantitative
variables. It is also used to test the statistical significance which can be used to test whether
the observed linear relationship could have emerged. if it fits a linear equation to observed
data called a regression equation.

Learning Outcomes for Study Session 3

On completion of this study session, you should be able to:

3.1 Explain simple regression


3.2 Understand the least square method

Page 44 of 174
STA202: Statistics II
3.1 Simple Regression

When only two variables are involved, the regression is said to be simple. A simple linear
regression equation is therefore of the form  =  + X , once  and  are estimated,
we can substitute a given value of X into the equation and calculate the predicted value of
.

In-Text Question (ITQ)

When is regression said to be simple?

In-Text Answer (ITA)

When only two variables are involved

3.2 The Least Squares Method

This is the most reliable of all the methods used to find regression lines. It leads to unique
regression line and regression coefficient. The least square method could be used to
estimate the parameters  and  from the model.

Yi    X i   i
as follow:
 i2  (Yi    x i ) 2 i  1, 2, ..., n

where the least square function is


n n
L= 2 1
 i2   (Y2    X i )2
i 1

Minimized the function L with respect to  and  by taking the partial derivatives
L
 2 (Yi     X i )


Page 45 of 174
STA202: Statistics II
L
 2 (Yi    xi )


Set these partial derivatives equal to zero and solve for  and , we obtain:

n xy   x ( y)
=
n x 2  ( x ) 2
and
 Y x

i i
where x 
n
 xi and Y   Yi
n
Case Study 3.1

A study was made on the effect of income level on the standard of living. The following
data was obtained in coded form. Calculate the regression of standard of living on income
level.

Income level Standard of Living


-5 1
-4 5
-3 4
-2 7
-1 10
0 8
1 9
2 13
3 14
4 13
5 18
Table 3.1

Page 46 of 174
STA202: Statistics II
Solution

The regression equation is

y    x

xy
 xy  n
=
 x   x2 / n
2

 Y x

From the table we compute

 x  0,  y  102,  x 2
 110,  xy  158

0  102
158 
 11  1.44
(0) 2
110 
11

0
x 0
11

102
Y  9.27,
11

  9.27  0(0)  9.27

y  9.27  1.44 x

Suppose we are to estimate or predict the value of Y when X= 6 we obtain

 =9.27+1.44 (6)

The predicted value of  =17.91 respectively

Page 47 of 174
STA202: Statistics II
Case Study 3.2

The table below shows the Nigeria gross domestic product (X) and inflation rate (Y) over
a period of time.

i) Find least square regression line of Y on X

ii) Find least square regression line of X on Y is considering X as dependent and Y as


independent variable respectively

X 65 63 67 64 68 62 70 66 68 67 69 71

Y 68 66 68 65 69 66 68 65 71 67 68 70

Table 3.2
Solution

i) The regression line of  on  is given by Y    x

  2 Y2 

65 68 4225 4624 4420

63 66 3969 4356 4158

67 68 4489 4624 4556

64 65 1096 4225 4160

68 69 4624 4769 4692

62 66 3844 4356 4092

70 68 4900 4624 4760

Page 48 of 174
STA202: Statistics II
66 65 4356 4225 4290

68 71 4624 5041 4290

67 67 4489 4889 4889

69 68 4761 4624 4692

71 70 5041 4900 4970

800 811 53418 54849 54107

Table 3.3

i. The regression line of y on x is given by

n xx   x y

n x 2  ( x ) 2

12(54107)  (800)(811)

12(53418)  (800) 2

  0.4764

      yx
y x
n n

811 800
  0.4764 x = 35.8233
12 12

The regression equation of Y on X is given as Y = 35.823 + 0.476X

(ii) The regression line of x on y is given by

x =   Y

Page 49 of 174
STA202: Statistics II
  n xy   x y
n y 2  ( y ) 2

12(54107)  (800)(811)

12(54849)  (811) 2

  1.036

  x    y  800  1.036 x 811


n n 12 12

 3.38

The regression equation of x on y is given as

Y = - 3.38 + 1.036Y

Page 50 of 174
STA202: Statistics II

Summary of Study Session 3

In study session 3, you have learnt that:

1. A simple linear regression equation is therefore of the form  =  + X


2. The least square methods is the most reliable of all the methods used to find
regression lines.
3. The least square method could be used to estimate the parameters  and  from
the model.

Page 51 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 3

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.
1. Explain the term Simple Regression analysis
2. The table give below shows the respective income (X) and expenditure (Y) (in million
naira) of O.O.U, Ago-Iwoye for a year.

Income (  ) 66 64 68 65 69 63 71 67 69 68 70 72

Expenditure (  ) 69 67 69 66 70 67 69 66 72 68 69 71

Find the regression of lines of X on Y and Y on X.

3. Given that ∑ 𝑥 = 628, ∑ 𝑦 = 1682, ∑ 𝑥𝑖 𝑦𝑖 = 1775.32, ∑ 𝑥𝑖2 = 1545.68,

∑ 𝑦𝑖2 = 2553.68 𝑎𝑛𝑑 ∑(𝑦̂ − 𝑦̅)2 = 1979.7792

(i) Fit a regression line of x on y.


(ii) Fit a regression line of y on x.

Page 52 of 174
STA202: Statistics II

Glossary of Terms

Parameter: a value that tells you something about a population and is the opposite from
a statistic,
Partial derivative: a function of several variables is its derivative with respect to one of
those variables, with the others held constant.
Regression Line: a single line that best fits the data (in terms of having the smallest
overall distance from the line to the points
Regression Coefficient: estimates of the unknown population parameters and describe the
relationship between a predictor variable and the response
Independent Variable: a variable that stands alone and isn't changed by the other
variables you are trying to measure
Linear Relationship: a straight-line relationship between two variables

Page 53 of 174
STA202: Statistics II

References

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 54 of 174
STA202: Statistics II
Study Session 4: Test of Hypothesis

Introduction

Data must be interpreted in order to add meaning. We can interpret data by


assuming a specific structure our outcome and use statistical methods to
confirm or reject the assumption. The assumption is called a hypothesis and
the statistical tests used for this purpose are called statistical hypothesis tests.
Whenever we want to make claims about the distribution of data or whether one set of
results are different from another set of results in applied machine learning, we must
rely on statistical hypothesis tests.

Learning Outcomes for Study Session 4

On completion of this study session, you should be able to:


4.1 Test for hypothesis
4.2 Differentiate between the types of errors

Page 55 of 174
STA202: Statistics II
4.1 Meaning of Test of Hypothesis

The most frequent application of statistics is to test some scientific hypotheses. Results of
experiments and investigations are usually not clear cut and, therefore, need statistical
tests to support decisions between alternative hypotheses. A statistical test examines a set
of sample data and on the basis of an expected distribution of the data, leads to a decision
on whether to accept the hypothesis or whether to reject that hypothesis and accept an
alternative one. The nature of the tests varies with the data and the hypothesis, but the
same general philosophy of hypothesis testing is common to all tests.
A statistical hypothesis is an assumption or statement which may or may not be true
concerning one or more population. A statistical hypothesis (or inference) is a statement
about the parameters or form of a population. A test of a statistical hypothesis is a criteria
which specifies for what sample results the hypothesis is to be accepted or rejected. The
hypothesis which is to be tested is generally called the Null hypothesis denoted by the H0
and hypothesis against which it is to be tested is called the alternative hypothesis and also
denoted by H1.

In-Text Questions (ITQs)

i. What is a statistical test?


ii. Define statistical hypothesis

In-Text Answers (ITAs)

i. A statistical test is a test that examines a set of sample data and on the basis of an
expected distribution of the data, leads to a decision on whether to accept the
hypothesis or whether to reject that hypothesis and accept an alternative one.
ii. A statistical hypothesis is an assumption or statement which may or may not be
true concerning one or more population.

Page 56 of 174
STA202: Statistics II
4.2 Type I and Type II Errors

A type I error has been committed if we reject the null hypothesis when it is true and a
type II error has been committed if we accept the null hypothesis when it is false.

The following table summarizes the various situations that can arise when testing H0
against H1:
Accept H0 Accept H1
H0 is true No error Type I Error
H1 is true Type II error No error
Table 4.1
The probabilities of committing a type I and type II errors are called level of significance
of the tests and are written as  and , respectively.  is called the size of the test and (1-
) is called the power of the test, and (1-) is also the probability of rejecting null
hypothesis (H0) when it is false. The area such that if the sample point falls in it we reject
H0 is called the critical region. When the primary concern of a test is to see whether the
null hypothesis can be rejected, such a test is called a test of significance. In that case, the
quantity  is called the level of significance at which the test is being conducted.

4.2.1 One and Two Tailed Test

A test of any statistical hypothesis where the alternative is one sided such as:
H0:  = 0 or H0:  = 0
H1:  > 0 H1:  < 0
Is called a one-tailed test. The critical region for H1:  > 0 lies entirely in the right tail
while the critical region for H1:  < 0 lies entirely in the left tail.
A test of any statistical hypothesis where the alternative is two-sided such as:
H0:  = 0
H1:   0

Page 57 of 174
STA202: Statistics II
Is called a two-tailed test, values in the both tails of the distribution constitute the critical
region.

4.2.2 Test Procedure and Steps

The steps involved in general and in the utilization of any test of significance are:

i. Find the type of problem and the question to be answered.


ii. To state the null hypothesis (H0) and the appropriate alternative (H1)
hypothesis
iii. Selection of the appropriate test to be utilized and calculation of the test
criterion based on the type of test.
iv. Fixation of the level of significance 
v. Decision making on test criterion value, whether to reject or accept the
hypothesis.
vi. Drawing of the conclusion (or inference) on the basis of level of significance is
deciding whether the difference observed is due to chance or due to some other
known factors.

4.2.3 Test Concerning the Mean (For Large Sample)

We will assume that the sampling distribution of the sample estimates will be
approximately normal and that the variance is known. Hence, for large samples (n  30),
we can use the normal probability distribution for testing a hypothesized value of the
population mean.

The test statistics

X  
Z 
S .E. X 

where X is the sample mean

Page 58 of 174
STA202: Statistics II
 is the population mean

S.E. ( X ) is the standard error of the sample mean.


S .E. X  
n

where  is the population standard deviation (usually known) and n is the sample size.

Case Study 4.1

A bottling company which bottles a soft drink claims that the liquids content is 35cl with
standard deviation 0.75cl. A researcher randomly collects 50 bottles, measured their
contents and got mean of 34.2cl. Test at 0.01 level of significance that the bottling
company has been cheating their consumers.

Solution

 = 35cl

 = 0.75cl

n = 50

X = 34.2

 = 0.01 (1%)

H0:  = 35 that is, the company has not been cheating the consumers.

H1:  < 35 that the company has been cheating the consumers.

Test statistics is

Z 
X   n

Page 59 of 174
STA202: Statistics II


34.2  35 50
0.75

 0.8  7.0711

0.75

= -7.54

Thus, |Z| = |-7.541| = 7.54

At 0.01 level of significance the Z tabulated value (one tailed) is 2.33

Decision: the Z calculated value 7.54 is greater than the Z tabulated value 2.33. we reject
H0 and accept H1.

Conclusion: There is significant difference between the population and sample mean.
Hence, the bottling company has been cheating their consumers.

4.2.4 Test Concerning Means (Small Samples)

There are situations in real life experiment, such as, testing the efficiency of a newly
produced drug, where it is impracticable to get a large sample and yet tests of significance
still have to be carried out. When we do not know the value of the population standard
deviation and the sample size is small (n < 30), we shall assume again that the population
we are sampling from has roughly the shape of a normal distribution. The test statistics is:

t
X 

X   n
S S
n

Whose sampling distribution is the t distribution with n-1 degree of freedom. S is the
sample standard deviation. As with large samples, we compare it with its value at a given
level of significance, and then draw our conclusions.

Page 60 of 174
STA202: Statistics II
Case Study 4.2

Suppose that we want to test on the basis of a random sample of size n = 5 whether or not
the fat content of a certain kind of ice cream exceeds 12 percent. What can we conclude
about the null hypothesis.  = 12 percent at the 0.01 level of significance, if the sample
has the mean X as 12.7 percent and the standard deviation S is 0.38 percent.

Solution

Hypothesis: H0:  = 12%

H1:  > 12

 = 0.01

n = 5, d.f. = n – 1 = t0.01,4 degree of freedom

Test statistics

X 
t
S
n

12.7  12
t
0.38
5

0. 7
t  4.12
0.1699

t0.01,4 = 4.12

Decision: Since tcal > ttab, we reject H0

Conclusion: Therefore, the content of the given kind of ice cream exceeds 12 percent.

Page 61 of 174
STA202: Statistics II
Case Study 4.3

The life time of telephone for a random sampling 10 from a large consignment give the
following data:

Item Life in 1,000hrs x- X (X - X )2

1 4.2 -0.1 0.01

2 4.0 -0.3 0.09

3 3.9 -0.4 0.16

4 4.1 -0.2 0.04

5 5.2 0.9 0.81

6 3.8 -0.5 0.25

7 3.9 -0.5 0.16

8 4.3 0 0

9 4.4 0.1 0.01

10 5.6 1.3 1.69

Page 62 of 174
STA202: Statistics II
Table 4.2

Case Study 4.4

Can we accept the hypothesis that the average life time of telephone is 4,000hours at 5%
level of significance?

Solution

Hypothesis

H0:  = 4,000hours

H1:   4,000hours

 = 0.05 leel of significance

Since, n = 10 d.f. = n-1 = 10 – 1, 9

t  2  n 1  n 1  t 0.25,9

X i
4.2  4.0  ,..., 5.6 43.5
X  i 1
 
n 10 10

X  4.3

 X  X
10
2
i
i 1
S2 
n 1

S2 
4.2  4.32  4.0  4.32  ... 4.4  4.3  5.6  4.3
2 2

10 1

0.01  0.09  , ...,  0.01  1.69



9

Page 63 of 174
STA202: Statistics II
3.22
S2  = 0.358
9

Test Statistics

t
X   n
S

t
4.3  4 10
0.598

where

S  0.358 = 0.598

t = 1.587

t0.025,9 = 2.262

Decision: Reject H0 if tcal > ttab

Conclusion: Since tcal > ttab, then we accept H0 and conclude that the average life time is
4,000hours

4.2.5 Test Concerning Two Population Means (Large Sample)

The test statistics for large sample test concerning difference between two means is given
as:

X1  X 2
Z 
 12  22
n1  n2

Page 64 of 174
STA202: Statistics II
Case Study 4.5

In a study designed to test whether or not there is a difference between the average amount
used to buy food by families living in two different communities, random samples yield
the following results.

n1 = 120 x1  62.7 1 = 2.50

n2 = 150 x2  61.8 2 = 2.62

Table 4.3

The amount used to buy food are in Thousand Naira. Use the 0.05 level of significance to
test the null hypothesis that the corresponding population means are equal against the
alternative hypothesis that they are not equal.

Solution

H 0  1   2

H 1  1   2

n1  120, x1  62.7,  1  2.50

n2  150, x2  61.8,  2  2.62

 = 0.05 level of significance

Page 65 of 174
STA202: Statistics II
4.2.6 Test Statistics

X1  X 2
Z 
 12  22
n1  n2

62.7  61.8
Z 
 2.50 2  2.62 2
120  150

0.9
Z 
0.0979

Z = 2.88

Conclusion: Since Z cal  Z tab , the null hypothesis must be rejected and we conclude that
there is a difference between the true average heights of adult females in the two given
communities.

4..2.7 Test Concerning Two Population Means (Small Sample)

The test statistics for small sample test concerning difference between two means is given
as:

X1  X 2
t
Sp 2  1
n2  1
n2

where

n1 1 S12  n2 1S 22


Sp 2 
n1  n2  2

where t-distribution with n1 + n2 – 2 is known as the pooled variance. Assumption when


using t-distribution.

i. The two populations are normal

Page 66 of 174
STA202: Statistics II
ii. The two populations have the same variance.
iii. The two samples are random ones.
Case Study 4.6

The following random samples are amount used by two states in Nigeria to provide health
facilities (in millions naira) for five months:

State 1 8400 8230 8380 7860 7930

State 2 7510 7690 7720 8070 7660

Table 4.4

Use 0.05 level of significance to test whether the difference between the means of these
two samples is significant.

Solution

X i
8400  8230  ...  7930
X1  i 1

5 5

X 1  8160

X i
7510  7690  ...  7660
X2  i 1

5 5

X 2  7730

 X  Xi 
5
2
i
i 1
S 21 
n1  1

Page 67 of 174
STA202: Statistics II
S  63450
1
2

 X  X2
5
2
i
i 1
S 22 
n2  1

S 22  42650

H 0  1   2

H 1  1   2

Test statistics

X1  X 2
t
n1  1S12   n2  1S 22
n1  n2  2  1
n1  1
n2 
8160  7730 430
t =t
4 63450   4  42650  1
552 5
  1
5
 21220

t = 2.9 = ttab = 2.306

Conclusion: Since tcal  ttab , the null hypothesis should be rejected then, we conclude that
the average amount spend on health facilities from the states are not the same.

In-Text Questions (ITQs)

i. Differentiate between type I error and type II error

Page 68 of 174
STA202: Statistics II
In-Text Answers (ITAs)

i. A type I error has been committed if we reject the null hypothesis when it is true
while a type II error has been committed if we accept the null hypothesis when it
is false

Page 69 of 174
STA202: Statistics II

Summary of Study Session 4

In study session 4, you have learnt that:

1. A test of a statistical hypothesis is a criterion which specifies for what sample


results the hypothesis is to be accepted or rejected
2. The probabilities of committing a type I and type II errors are called level of
significance of the tests and are written as  and , respectively.  is called the
size of the test and (1-) is called the power of the test.
3. The probability of rejecting null hypothesis (H0) when it is false is (1-).

Page 70 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 4

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. Explain the following terms

i. Statistical hypothesis

ii. Type I and Type II Errors

iii. One tailed and two tailed test

2. What are the procedure in testing statistical hypothesis?

3. The following random samples are the spending power of families (in millions
naira)

for two states in Nigeria:

Mine 1 84 82 83 78 79
Mine 2 75 76 77 80 76

Use 0.05 level of significance to test whether the difference between the means of these
two samples is significant.

Page 71 of 174
STA202: Statistics II

Glossary of Terms

Hypothesis: the process that an analyst uses to test a statistical hypothesis


Null: no value
Scientific hypotheses: an idea that proposes a tentative explanation about a phenomenon
or a narrow set of phenomena observed in the natural world
Null hypothesis: a typical statistical theory which suggests that no statistical relationship
and significance exists in a set of given single observed variable, between two sets of
observed data and measured phenomena
Tests of significance: a formal procedure for comparing observed data with a claim (also
called a hypothesis), the truth of which is being assessed
Sample test: the studies to be performed by each Party using the applicable Samples

Page 72 of 174
STA202: Statistics II

References

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 73 of 174
STA202: Statistics II
Study Session 5: Index Numbers
Introduction

The value of money does not remain constant over time. It rises or falls and is
inversely related to the changes in the price level. A rise in the price level
means a fall in the value of money and a fall in the price level means a rise in the
value of money. Thus, changes in the value of money are reflected by the changes in the
general level of prices over a period of time. Changes in the general level of prices can be
measured by a statistical device known as ‘index number.’

Learning Outcomes for Study Session 5

On completion of this study session, you should be able to:


5.1 Explain the meaning of index number
5.2 Construct price index numbers
5.3 Understand the difficulties in measuring changes in value of money
5.4 Differentiate between the various types of index numbers

Page 74 of 174
STA202: Statistics II
5.1 Meaning of Index Numbers

Index number is a technique of measuring changes in a variable or group of variables with


respect to time, geographical location or other characteristics. There can be various types
of index numbers, but, in the present context, we are concerned with price index numbers,
which measures changes in the general price level (or in the value of money) over a period
of time.
Price index number indicates the average of changes in the prices of representative
commodities at one time in comparison with that at some other time taken as the base
period. An index number of prices is a figure showing the height of average prices at one
time relative to their height at some other time which is taken as the base period.”

5.1.1 Features of Index Numbers

The following are the main features of index numbers:


i. Index numbers are a special type of average. Whereas mean, median and mode
measure the absolute changes and are used to compare only those series which are
expressed in the same units, the technique of index numbers is used to measure the
relative changes in the level of a phenomenon where the measurement of absolute
change is not possible and the series are expressed in different types of items.
ii. Index numbers are meant to study the changes in the effects of such factors which
cannot be measured directly. For example, the general price level is an imaginary
concept and is not capable of direct measurement. But, through the technique of
index numbers, it is possible to have an idea of relative changes in the general
level of prices by measuring relative changes in the price level of different
commodities.
iii. The technique of index numbers measures changes in one variable or group of
related variables. For example, one variable can be the price of wheat, and group
of variables can be the price of sugar, the price of milk and the price of rice.
iv. The technique of index numbers is used to compare the levels of a phenomenon on
a certain date with its level on some previous date (e.g., the price level in 1980 as

Page 75 of 174
STA202: Statistics II
compared to that in 1960 taken as the base year) or the levels of a phenomenon at
different places on the same date (e.g., the price level in India in 1980 in
comparison with that in other countries in 1980).

5.1.2 Steps or Problems in the Construction of Price Index Numbers

The construction of the price index numbers involves the following steps or problems
1. Selection of Base Year: The first step or the problem in preparing the index
numbers is the selection of the base year. The base year is defined as that year with
reference to which the price changes in other years are compared and expressed as
percentages. The base year should be a normal year. In other words, it should be
free from abnormal conditions like wars, famines, floods, political instability, etc.
Base year can be selected in two ways- (a) through fixed base method in which the
base year remains fixed; and (b) through chain base method in which the base year
goes on changing, e.g., for 1980 the base year will be 1979, for 1979 it will be
1978, and so on.
2. Selection of Commodities: The second problem in the construction of index
numbers is the selection of the commodities. Since all commodities cannot be
included, only representative commodities should be selected keeping in view the
purpose and type of the index number.
In selecting items, the following points are to be kept in mind
a. The items should be representative of the tastes, habits and customs of the people.
b. Items should be recognizable,
c. Items should be stable in quality over two different periods and places.
d. The economic and social importance of various items should be considered
e. The items should be fairly large in number.
f. All those varieties of a commodity which are in common use and are stable in
character should be included.
3. Collection of Prices: After selecting the commodities, the next problem is
regarding the collection of their prices:
a. From where the prices to be collected;

Page 76 of 174
STA202: Statistics II
b. Whether to choose wholesale prices or retail prices;
c. Whether to include taxes in the prices or not etc.
While collecting prices, the following points are to be noted
a. Prices are to be collected from those places where a particular commodity is traded
in large quantities.
b. Published information regarding the prices should also be utilised,
c. In selecting individuals and institutions who would supply price quotations, care
should be taken that they are not biased.
d. Selection of wholesale or retail prices depends upon the type of index number to
be prepared. Wholesale prices are used in the construction of general price index
and retail prices are used in the construction of cost-of-living index number.
e. Prices collected from various places should be averaged.
4. Selection of Average: Since the index numbers are specialized average, the fourth
problem is to choose a suitable average. Theoretically, geometric mean is the
best for this purpose. But, in practice, arithmetic mean is used because it is easier
to follow.
5. Selection of Weights: Generally, all the commodities included in the construction’
of index numbers are not of equal importance. Therefore, if the index numbers are
to be representative, proper weights should be assigned to the commodities
according to their relative importance. For example, the prices of books will be
given more weightage while preparing the cost-of-living index for teachers than
while preparing the cost-of-living index for the workers. Weights should be
unbiased and be rationally and not arbitrarily selected.
6. Purpose of Index Number: The most important consideration in the construction of
the index numbers is the objective of the index numbers. All other problems or
steps are to be viewed in the light of the purpose for which a particular index
number is to be prepared. Since, different index numbers are prepared with
specific purposes and no single index number is ‘all purpose’ index number, it is
important to be clear about the purpose of the index number before its
construction.

Page 77 of 174
STA202: Statistics II
7. Selection of Method: The selection of a suitable method for the construction of
index numbers is the final step.
There are two methods of computing the index numbers:
a. Simple index number and
b. Weighted index number.
Simple index number again can be constructed either by:
i. Simple aggregate method or
ii. Simple average of price relative’s method.
Similarly, weighted index number can be constructed by:
i. Weighted aggregative method or
ii. Weighted average of price relative’s method.
The choice of method depends upon the availability of data, degree of accuracy required
and the purpose of the study.

In-Text Questions (ITQs)

i. What is index number?


ii. What are the steps or problems in the construction of price index numbers?

In-Text Answers (ITAs)

i. Index number is a technique of measuring changes in a variable or group of


variables with respect to time, geographical location or other characteristics
ii. Selection of base year, selection of commodities, selection of prices, selection of
average, selection of weights, purpose of index number, selection of method.

5.2 Construction of Price Index Numbers

Construction of price index numbers through various methods can be understood with the
help of the following examples:

Page 78 of 174
STA202: Statistics II
5.2.1 Simple Aggregative Method

In this method, the index number is equal to the sum of prices for the year for which index
number is to be found divided by the sum of actual prices for the base year.
The formula for finding the index number through this method is as follows in case study
5.1 below as follows
Case Study 5.1
∑ 𝑃1
𝑃01 = × 100
∑ 𝑃0
Where 𝑃01 stands for index number
∑ 𝑃1 stands for the sum of the prices for the year for which index number is to be found
∑ 𝑐 stands for the sum prices for the base year
Commodity Prices in Base Year 1980 Prices in Current Year 1988
(In Naira) 𝑃0 ( in Naira) 𝑃1

A 10 20
B 15 25
C 40 60
D 25 40
Total ∑ 𝑃0 = 90 ∑ 𝑃1 = 145

Table 5.1
∑𝑃 145
Index Number (𝑃01 ) = ∑ 𝑃1 × 100; 𝑃01 = × 100; 𝑃01 = 161.11
0 90

5.2.2 Simple Average of Price Relatives Method

In this method, the index number is equal to the sum of price relatives divided by the
number of items and is calculated by using the following formula in case study 5.2 as
follows

Page 79 of 174
STA202: Statistics II
Case Study 5.2
Commodity Base Year Prices Current Year Price Relative
𝑃0 𝑃1 Prices (in Naira) 𝑃1
𝑅= × 100
𝑃0
A 10 20 20
× 100 = 200.0
10
B 15 25 25
× 100 = 166.7
15
C 40 60 60
× 100 = 150.0
40
D 25 40 40
× 100 = 160.0
25
N=4 ∑ 𝑅 = 676.7

Table 5.2
∑𝑅
Index Number (𝑃01 ) = 𝑁

676.7
(𝑃01 ) = = 169.2
4

5.2.3 Weighted Aggregative Method

In this method, different weights are assigned to the items according to their relative
importance. Weights used are the quantity weights. Many formulae have been developed
to estimate index numbers on the basis of quantity weights.
Some of them are explained below with some examples
i. Laspeyre’s Formula: In this formula, the quantities of base year are accepted
as weights.
∑ 𝑃1 𝑞0
𝑃01 = × 100
∑ 𝑃0 𝑞0
Where 𝑃1 is the price in the current year; 𝑃0 is the price in the base year; and 𝑞0 is the
quantity of the base year.
ii. Paasche’s Formula: In this formula, the quantities of the current year are
accepted as weights.
Page 80 of 174
STA202: Statistics II
∑ 𝑃1 𝑞1
𝑃01 = × 100
∑ 𝑃0 𝑞1
Where 𝑞1 is the quantity in the current year.
iii. Dorbish and Bowley’s Formula: Dorbish and Bowley’s Formula for
estimating weighted index number is as follows
∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
+
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1 𝐿+𝑃
𝑃01 = × 100 or 𝑃01 =
2 2

Where L is Laspeyre’s index and P is Paasche’s index.


iv. Fisher’s Ideal Formula: In this formula, the geometric mean of two indices
(i.e. Laspeyre’s index and P is Paasche’s index) is taken.

∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
𝑃01 = √ + × 100 𝑜𝑟 𝑃01 = √𝐿 × 𝑃 = 100
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1

Where L is Laspeyre’s index and P is Paasche’s index.


Case Study 5.3
Commodity Base Year Current Year 𝑃0 𝑞0 𝑃1 𝑞0 𝑃0 𝑞1 𝑃1 𝑞1
𝑃0 𝑞0 𝑃1 𝑞1
A 10 5 20 2 50 100 20 40
B 15 4 25 8 60 100 120 200
C 40 2 60 6 80 120 140 360
D 25 3 40 4 75 120 100 160
Total 265 440 480 760

∑ 𝑃0 𝑞0 ∑ 𝑃1 𝑞0 ∑ 𝑃0 𝑞1 ∑ 𝑃1 𝑞1

Table 5.3
i. Laspeyre’s Formula:
∑ 𝑃1 𝑞0
𝑃01 = × 100
∑ 𝑃0 𝑞0
440
𝑃01 = × 100 = 166.04
265
ii. Paasche’ Formula:

Page 81 of 174
STA202: Statistics II
∑ 𝑃1 𝑞1
𝑃01 = × 100
∑ 𝑃0 𝑞1
700
𝑃01 = × 100 = 158.3
480

iii. Dorbish and Bowley’s Formula:


∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
+
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1
𝑃01 = × 100 = 162.2
2
iv. Fisher’s Ideal Formula:

∑ 𝑃1 𝑞0 ∑ 𝑃1 𝑞1
𝑃01 = √ + × 100
∑ 𝑃0 𝑞0 ∑ 𝑃0 𝑞1

440 760
𝑃01 = √ + × 100 = 162.1
265 480

5.2.4 Weighted Average of Relatives Method

In this method also different weights are used for the items according to their relative
importance. The price index number is found out with the help of the following formula
and with examples.
∑ 𝑅𝑊
𝑃01 =
∑𝑊
Where ∑ 𝑊 stands for the sum of weights of different commodities and
∑ 𝑅 stands for the sum of price relatives

In-Text Questions (ITQs)

i. What are the methods of constructing Price index numbers?


ii. What are the steps or problems in the construction of price index numbers

Page 82 of 174
STA202: Statistics II
In-Text Answers (ITAs)

Simple aggregated method, simple average of price relatives method, weighted average of
relatives method.

Case Study 5.4

Commodity Weights Base Current Price RW


W Prices Year Relatives
Year Prices 𝑃1
× 100
𝑃0 𝑃1 𝑃0

A 5 10 20 20 1000.0
× 100 =
10

200.0
B 4 15 25 25 666.8
× 100
15
= 166.7
C 2 40 60 60 300.0
× 100
40
= 150.0
D 3 25 40 40 480.0
× 100
25
= 160.0
Total ∑ 𝑊 = 14 ∑ 𝑅𝑊 = 2446.8

Table 5.4
∑ 𝑅𝑊
Index Number (𝑃01 ) = ∑𝑊

2446.8
𝑃01 = = 174.8
14

Page 83 of 174
STA202: Statistics II
5.3 Difficulties in Measuring Changes in Value of Money

Measurement of changes in the value of money through price index number is not an easy
and reliable technique. There are a number of theoretical as well as practical difficulties
in the construction of price index numbers. Moreover, the index number technique itself
has many limitations.

Fig 5.1: Difficulties in Measuring Changes in Value of Money

5.3.1 Conceptual Difficulties

The following are the conceptual difficulties during the construction of price index
numbers:
1. Vague Concept of Value of Money: The concept of money is vague, abstract and
cannot be clearly defined. The value of money is a relative concept which changes
from person to person depending upon the type of goods on which the money is
spent.
2. Inaccurate Measurement: Price index numbers do not measure the changes in the
value of money accurately and reliably. A rise or fall in the general level of prices
as indicated by the price index numbers does not mean that the price of every
commodity has risen or fallen to the same extent.

Page 84 of 174
STA202: Statistics II
3. Reflect General Changes: Price index numbers are averages and measure general
changes in the value of money on the average. Therefore, they are not of much
significance for the particular individuals who may be affected by the changes in
the actual prices quite differently from that indicated by the index numbers.
4. Limitations of Wholesale Price Index: The wholesale price index numbers, which
are generally used to measure changes in the value of money, suffer from certain
limitations:
i. They do not reflect the changes in the cost of living because retail prices are
generally higher than the wholesale prices.
ii. They ignore some of the important items concerning the urban population, such as,
expenditure on education, transport, house rent, etc.
iii. They do not take into consideration the changes in the consumers’ preferences.

Practical Difficulties

The practical difficulties in the way of constructing price index numbers, and therefore, in
measuring changes in the value of money are as follows:
1. Selection of Base Year: While preparing the index number, first difficulty arises
regarding the selection of base year. The base year should be a normal year. But, it
is very difficult to find out a fully normal year free from any unusual happening.
There is every possibility that the selected base year may be an abnormal year, or a
distant year, or may be selected by an immature or biased person.
2. Selection of Items: The selection of the representative commodities is the second
difficulty in the construction of index numbers:
[Link] the passage of time the quality of the product may change; if the quality of a
product changes in the year of enquiry from what it was in the base year, the product
becomes irrelevant
[Link] relative importance of certain commodities may change due to a change in the
consumption pattern of the people in the course of time; for example, Vanaspati Ghee
was not an important item of consumption in India in the pre-war period, but today it

Page 85 of 174
STA202: Statistics II
has become an item of necessity. Under such conditions, it is not easy to select the
appropriate commodities.
2 Collection of Prices: It is also difficult to obtain correct, adequate and representative
data regarding prices. It is not an easy job to select representative places from which
the information about prices to be collected and to select the experienced and unbiased
individuals or institutions who will supply price quotations. Moreover, there is the
problem of deciding which prices (wholesale or retail) are to be taken into
consideration. It is comparatively easy to get information about wholesale prices
which vary considerably.
3 Assigning Weights: Another important difficulty that arises in preparing the index
numbers is that of assigning proper weights to different items in order to arrive at
correct and unbiased conclusions. As there are no hard and fast rules to weights for the
commodities according to their relative importance, there is very likelihood that the
weights are decided arbitrarily on the basis of personal judgement and involve
biasness.
4 Selection of Averages: Another major problem is that which average should be
employed to find out the price relatives. There are many types of averages such as
arithmetic average, geometric average, mean, median, mode, etc. The use of different
averages gives different results. Therefore, it is essential to select the method with
great care. Dr. Marshall has advocated the use of chain index number to solve the
problem of averaging and weighing.
5 Problem of Dynamic Changes: In the dynamic world, the consumption pattern of the
individuals and the number and varieties of goods undergo continuous changes. They
create difficulties for preparing index numbers and making temporal comparisons:
i. Since, in the course of time, old commodities may disappear and many new ones
come into existence, the long-run comparison may become difficult
ii. The quantity and quality of commodities may also change over the period of time,
thus making the choice of commodities for constructing index numbers difficult
iii. A number of factors, like income, education, fashion, etc., bring changes in the
consumption pattern of the people which render the index numbers incomparable.

Page 86 of 174
STA202: Statistics II
In-Text Questions (ITQs)

What are the difficulties in measuring changes in value of money?

In-Text Answers (ITAs)

i. Conceptual difficulties: vague concept of value of money, inaccurate


measurement, reflect general changes, limitations of wholesale price index
ii. Practical difficulties: selection of base year, selection of items, collection of
prices, assigning weights, selection of averages, selection of averages and problem
of dynamic changes

5.4 Types of Index Numbers

Index numbers are of different types. Important types of index numbers are discussed
below

Whole sale
price

Cost-of-
Retail price
Living price

Index
Number

Industrial Wage

Working
class Cost-
of-Living

Fig 5.2: Types of Index Number

Page 87 of 174
STA202: Statistics II
5.4.1 Wholesale Price Index Numbers

Wholesale price index numbers are constructed on the basis of the wholesale prices of
certain important commodities. The commodities included in preparing these index
numbers are mainly raw-materials and semi-finished goods. Only the most important and
most price-sensitive and semi- finished goods which are bought and sold in the wholesale
market are selected and weights are assigned in accordance with their relative importance.
The wholesale price index numbers are generally used to measure changes in the value of
money. The main problem with these index numbers is that they include only the
wholesale prices of raw materials and semi-finished goods and do not take into
consideration the retail prices of goods and services generally consumed by the common
man. Hence, the wholesale price index numbers do not reflect true and accurate changes in
the value of money.

5.4.2 Retail Price Index Numbers

These index numbers are prepared to measure the changes in the value of money on the
basis of the retail prices of final consumption goods. The main difficulty with this index
number is that the retail price for the same goods and for continuous periods is not
available. The retail prices represent larger and more frequent fluctuations as compared to
the wholesale prices.

5.4.3 Cost-of-Living Index Numbers

These index numbers are constructed with reference to the important goods and services
which are consumed by common people. Since the number of these goods and services is
very large, only representative items which form the consumption pattern of the people are
included. These index numbers are used to measure changes in the cost of living of the
general public.

Page 88 of 174
STA202: Statistics II
5.4.4 Working Class Cost-of-Living Index Numbers

The working class cost-of-living index numbers aim at measuring changes in the cost of
living of workers. These index numbers are consumed on the basis of only those goods
and services which are generally consumed by the working class. The prices of these
goods and index numbers are of great importance to the workers because their wages are
adjusted according to these indices.

5.4.5 Wage Index Numbers

The purpose of these index numbers is to measure time to time changes in money wages.
These index numbers, when compared with the working class cost-of-living index
numbers, provide information regarding the changes in the real wages of the workers.

5.4.6 Industrial Index Numbers

Industrial index numbers are constructed with an objective of measuring changes in the
industrial production. The production data of various industries are included in preparing
these index numbers.

5.4.7 Uses of Index Number in the Economic Field

Some of the specific uses of index numbers in the economic field are:
i. They are useful in analysing markets for specific commodities.
ii. In the share market, the index numbers can provide data about the trends in the
share prices
iii. With the help of index numbers, the Railways can get information about the
changes in goods traffic.
iv. The bankers can get information about the changes in deposits by means of index
numbers.

Page 89 of 174
STA202: Statistics II
5.4.8 Limitations of Index Numbers

Index number technique itself has certain limitations which have greatly reduced its
usefulness:

i. Because of the various practical difficulties involved in their computation, the


index numbers are never cent per cent correct.
ii. There are no all-purpose index numbers. The index numbers prepared for one
purpose cannot be used for another purpose. For example, the cost-of-living index
numbers of factory workers cannot be used to measure changes in the value of
money of the middle income group.
iii. Index numbers cannot be reliably used to make international comparisons.
Different countries include different items with different qualities and use different
base years in constructing index numbers.
iv. Index numbers measure only average change and indicate only broad trends. They
do not provide accurate information.
v. While preparing index numbers, quality of items is not considered. It may be
possible that a general rise in the index is due to an improvement in the quality of a
product and not because of a rise in its price.

In-Text Questions (ITQs)

What are the types of index number?

In-Text Answers (ITAs)

Wholesale price index number, retail price index numbers, cost-of-living index numbers,
working class cost-of-living index numbers and industrial index numbers.

Page 90 of 174
STA202: Statistics II

Summary of Study Session 5

In study session 5, you have learnt that:

1. Index number is a technique of measuring changes in a variable or group of


variables with respect to time, geographical location or other characteristics.
2. Price index number indicates the average of changes in the prices of representative
commodities at one time in comparison with that at some other time taken as the
base period
3. The two methods of computing the index numbers are Simple index number and
Weighted index number.
4. Simple index number again can be constructed either by Simple aggregate method
Simple average of price relative’s method.
5. Weighted index number can be constructed by Weighted aggregative method or
Weighted average of price relative’s method.

Page 91 of 174
STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 5

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. Define index number.


2. Outline the features of index number
3. Outline the problems in constructing index number.

Page 92 of 174
STA202: Statistics II

Glossary of Terms

Theoretical: concerned with or involving the theory of a subject


Arbitrary: subject to individual will or judgment without restriction
Vague: of uncertain
Statistical device: Statistical methods involved in carrying out a study include planning,
designing, collecting data, analysing, drawing meaningful interpretation and reporting of
the research findings
Absolute changes: the simple difference in the indicator over two periods in time
Geometric mean: used for a set of numbers whose values are meant to be multiplied
together or are exponential in nature

Limitations: a restriction

Page 93 of 174
STA202: Statistics II

References

1. Probability and statistics for engineers & scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 94 of 174
STA202: Statistics II
Study Session 6: Time Series Analysis

Introduction

You may have heard people saying that the price of a particular commodity has
increased or decreased with time. This commodity can be anything like gold, silver,
any eatables, petrol, diesel etc. Also, you may have heard that the rate of interest has
increased in banks. The rate of interest for home loans has decreased. What are all these? How
are they useful to us? These types of data are the time series of data.

Learning Outcomes for Study Session 6

On completion of this study session, you should be able to:


6.1 Define time series
6.2 Explain the methods of estimating trend
6.3 Represent time series mathematically

Page 95 of 174
STA202: Statistics II
6.1 Definition of Time Series

A time series is defined as some quantity that is measured sequentially in time over some
interval. In its broadest form, time series analysis is about inferring what has happened to
a series of data points in the past and attempting to predict what will happen to it the
future.

6.1.1 Components of Time Series

The components of a time series are:

i. Trend is the long term pattern of a time series: A trend can be positive or negative
depending on whether the time series exhibits an increasing long term pattern or a
decreasing long term pattern
ii. Cyclical movement: Any pattern showing an up and down movement around a
given trend is identified as a cyclical pattern. The duration of a cycle depends on
the type of business or industry being analyzed.
iii. Seasonal movement: Many time series contain seasonal variation. This is
particularly true in series representing business sales or climate levels. In
quantitative finance we often see seasonal variation in commodities, particularly
those related to growing seasons or annual temperature variation (such as natural
gas).
iv. Irregular movement: This component is unpredictable. Every time series has some
unpredictable component that makes it a random variable.

6.1.2 Time Series Model

Time series model are:

i. Additive model: This is a model in which the series value of y is the sum of all
four components, that is,

y = T + S + C+ I

Page 96 of 174
STA202: Statistics II
ii. Multiplicative model: It is a model in which y is the product of all the time series
components. This is denoted by

y = T * S * C* I

In-Text Questions (ITQs)

i. Define time series


ii. What are the components of time series?
iii. What are the time series models?

In-Text Answers (ITAs)

i. A time series is defined as some quantity that is measured sequentially in time


over some interval
ii. Trend is the long term pattern of a time series, cyclical movement, seasonal,
movement, and irregular movement.
iii. Additive model and multiplicative model

6.2 Methods of Estimating Trend

The methods used to estimate trend in time series analysis are many but here we will
discuss the followings

i. Moving Average
ii. Least Square method

Page 97 of 174
STA202: Statistics II

Moving Average
Method Least Square
Method

Fig 6.1: Methods of Estimating Trend

6.2.1 Moving Average Method

A moving average (rolling average or running average) is a calculation to analyze data


points by creating a series of averages of different subsets of the full data set. Moving
average technique is best suited for data that exhibit some form of regular periodicity. In
general, if a set of T observation is arranged chronologically as 𝑥1 , 𝑥2 , 𝑥3 , … , 𝑥𝑡 , 𝑥𝑡+1 , … , 𝑥𝑇.

We obtain a set of averages thus

1
𝑦1 = (𝑥 + 𝑥2 +, … , +𝑥𝑛 )
𝑛 1
1
𝑦2 = (𝑥 + 𝑥3 +, … , +𝑥𝑛 + 𝑥𝑛+1 )
𝑛 2
1
𝑦3 = (𝑥 + 𝑥4 +, … , +𝑥𝑛 + 𝑥𝑛+1 + 𝑥𝑛+2 ) 𝑎𝑛𝑑 𝑠𝑜 𝑜𝑛.
𝑛 3

These averages are called-point moving averages. The n-points to show that n
observations are used in the averages. The averages are moving because they are averages
of successive n observations.

Page 98 of 174
STA202: Statistics II
In-Text Questions (ITQs)

What are the methods of estimating trend?

In-Text Answers (ITAs)

Moving Average and Least Square methods

Case Study 6.1

The following table gives the quarterly indices of rental prices of certain foodstuffs.

Quarterly indices of retail prices of foodstuff

Quarter
Year 1 2 3 4
2004 119 127 127 116
2005 123 142 133 127
2006 146 185 181 161

Table 6.1

(i) Calculate the four-quarterly (4-point) moving average and interpret your result.

Solution
Table for computation of 4-point moving average
Year/Quarter Indices of retail prices 4-point moving average

2004q1 119

2004q2 127

Page 99 of 174
STA202: Statistics II
¼(119+127+127+116)=122.25

2004q3 127

¼(127+127+116+123)=123.25

2004q4 116

¼(127+116+123+142)=127.00

2005q1 123

¼(116+123+142+133)=128.5

2005q2 142

¼(123+142+133+127)=131.25

2005q3 133

¼(142+133+127+146)=137.00

2005q4 127

¼(133+127+146+185)=147.75

2006q1 146

¼(127+146+185+181)=159.75

2006q2 185

¼(146+185+181+161)=168.25

2006q3 181

2006q4 161

Table 6.2
Note
Each average is placed between the two middle quarters of the four quarters used in
calculating the average.
Box 6.1

Page 100 of 174


STA202: Statistics II
The 4-point moving average showed quite clearly an upward trend in the series. If one
were to predict the future from this result, one would say that if the trend continues, then
there is likely to be a rise in retail price index in the coming years.

6.2.2 Least Square Method

Each point on the fitted curve represents the relationship between a known independent
variable and an unknown dependent variable. In general, the least square method uses a
straight line in order to fit through the given points which are known as the method of linear
or ordinary least square. This line is termed as the line of best fit from which the sum of
squares of the distances from the points is minimized.
Equations with certain parameters usually represent the results in this method. The method of
least square actually defines the solution for the minimization of the sum of squares of
deviations or the errors in the result of each equation.
The least squares method is used mostly for data fitting. The best fit result minimizes the
sum of squared errors or residuals which are said to be the differences between the observed
or experimental value and corresponding fitted value given in the model. There are two basic
kinds of the least square methods – ordinary or linear least square and nonlinear least squares.

6.3 Mathematical Representation

It is a mathematical method and with it gives a fitted trend line for the set of data in such a
manner that the following two conditions are satisfied.
1. The sum of the deviations of the actual values of Y and the computed values of Y is
zero.
2. The sum of the squares of the deviations of the actual values and the computed values
is least.
This method gives the line which is the line of best fit. This method is applicable to give
results either to fit a straight line trend or a parabolic trend.
The method of least square as studied in time series analysis is used to find the trend line of
best fit to a time series data.

Page 101 of 174


STA202: Statistics II
The trend line (Y) is defined by the following equation:
Y=a+bX
Where Y = predicted value of the dependent variable
a = intercept i.e. the height of the line above origin (when X = 0, Y = a)
b = slope of the line (the rate of change in Y for a given change in X)
When b is positive the slope is upwards, when b is negative, the slope is downward.
X = independent variable (in this case it is time)
To estimate the constants a and b, the following two equations have to be solved
simultaneously:
ΣY = na + b ΣX
ΣXY = aΣX + bΣX2
To simplify the calculations, if the midpoint of the time series is taken as origin, then the
negative values in the first half of the series balance out the positive values in the second half
so that ΣΣX = 0. In this case, the above two normal equations will be as follows:
ΣY = na
ΣXY = bΣX2
In such a case the values of a and b can be calculated as under:
Since ΣY = na
a = ∑Yn∑Yn
Since, ΣXY = bΣX2
Case Study 6.2
Fit a straight line trend on the following data using the Least Squares Method.

Period
1996 1997 1998 1999 2000 2001 2002 2003 2004
(year)

Y 4 7 7 8 9 11 13 14 17

Table 6.3
Solution

Page 102 of 174


STA202: Statistics II
Total of 9 observations are there. So, the origin is taken at the Year 2000 for which X is
assumed to be 0.

PERIOD
Y X XY X2
YEAR)

1996 4 -4 -16 16

1997 7 -3 -21 9

1998 7 -2 -14 4

1999 8 -1 -8 1

2000 9 0 0 0

2001 11 1 11 1

2002 13 2 16 4

2003 14 3 42 9

2004 17 4 68 16

Total ΣY = 90 ΣX = 0 ΣXY = 88 ΣX2 =60

Table 6.4
From the table we find that value of n is 9, value of ΣY is 90, value of ΣX is 0, value
of ΣXY is 88 and value of ΣX2 is 60.

Page 103 of 174


STA202: Statistics II
Substituting these values in the two given equations,
That is,
90 = 9𝑎
𝑎 = 1088 = 60𝑏
𝑏 = 1.47
The Trend equation is : Y = 10 + 1.47 X

In-Text Questions (ITQs)

What are the conditions necessary for a fitted trend line for the set of data in mathematical
representation of times series satisfy?

In-Text Answers (ITAs)

i. The sum of the deviations of the actual values of Y and the computed values of Y is
zero.
ii. The sum of the squares of the deviations of the actual values and the computed
values is least.

Page 104 of 174


STA202: Statistics II

Summary of Study Session 6

In study session 6, you have learnt that:

1. Time series analysis is about inferring what has happened to a series of data points
in the past and attempting to predict what will happen to it the future.
2. Additive model is y = T + S + C+ I
3. Multiplicative model is y = T * S * C* I
4. Moving average technique is best suited for data that exhibit some form of regular
periodicity.
5. The least square method uses a straight line in order to fit through the given points
which are known as the method of linear.

Page 105 of 174


STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 6

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. Define time series


2. Outline the components of time series
3. Explain the followings
(i) Additive time series model
(ii) Multiplicative time series model
4. State two types of method used for trend estimation methods
5. Fit a straight line trend on the following data using the Least Squares Method.

Period
1996 1997 1998 1999 2000 2001 2002 2003 2004
(year)

Y 6 7 7 8 19 10 13 15 17

Page 106 of 174


STA202: Statistics II

Glossary of Terms

Sequential: used in sequence


Data: individual pieces of factual information recorded and used for the purpose of
analysis
Measured sequentially: a special kind of realization of a joint measurement
Variation: a change or slight difference in condition, amount, or level, typically within
certain limits
Successive: following one another

Page 107 of 174


STA202: Statistics II

References

1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 108 of 174


STA202: Statistics II
Study Session 7: Multiple Regression

Introduction

Multiple regression is an extension of simple linear regression. It is used when


we want to predict the value of a variable based on the value of two or more
other variables. The variable we want to predict is called the dependent variable
(or sometimes, the outcome, target or criterion variable). The variables we are using to
predict the value of the dependent variable are called the independent variables (or
sometimes, the predictor, explanatory or regressor variables).

For example, you could use multiple regression to understand whether exam performance
can be predicted based on revision time, test anxiety, lecture attendance and gender.
Alternately, you could use multiple regression to understand whether daily cigarette
consumption can be predicted based on smoking duration, age when started smoking,
smoker type, income and gender.

Learning Outcomes for Study Session 7

On completion of this study session, you should be able to:

7.1 Explain multiple regression

7.2 Estimate regression coefficient

7.3 Formulate polynomial regression model

7.4 Analyse polynomial model

Page 109 of 174


STA202: Statistics II
7.1 Introduction

Multiple regression is a situation where there two or more predictors, and its analysis is
one of the most widely used of all statistical methods. Multiple regressions is a set of
techniques used to analyse the relationship between two or more independent variables
and a dependent variable. Variety of multiple regression models is discussed in this
section. Then we present the basic statistical results for multiple regression in matrix form.
Since the matrix expressions for multiple regression are the same as for simple linear
regression.
When there are two predictor variables 𝑋1 and 𝑋2 , the regression model:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖 (1.1)
is called a first-order model with two predictor variables. 𝑌𝑖 denotes as usual the response
in the ith trial, and 𝑋𝑖1 and 𝑋𝑖2 are the values of the two predictor vari;bles in the 𝑖 𝑡ℎ trial.
The parameters of the model are 𝛽0 , 𝛽1 , 𝑎𝑛𝑑 𝛽2 and the error term is 𝜀𝑖 .
Assuming that E[𝜀𝑖 ]= 0, the regression function for model (1.1) is:
𝐸(𝑌) = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 (1.2)
Similar to simple linear regression, where the regression function 𝐸(𝑌) = 𝛽0 + 𝛽1 𝑋is a
line, regression function (1.2) is a plane. Assuming the 𝛽0 = 12, 𝛽1 = 3, 𝑎𝑛𝑑 𝛽2 = 5,
we’ll have
𝐸(𝑌) = 12 + 3𝑋1 + 5𝑋2 (1.3)
The regression function in multiple regression is often called a regression surface or a
response surface.
The parameter 𝛽1 indicates the change in the mean response 𝐸[𝑌] per unit increase in
𝑋1 when 𝑋2 is held constant. Likewise, 𝛽2 indicates the change in the mean response per
unit increase in 𝑋2 when 𝑋1 is held constant. For example Suppose 𝑋2 is held at the level
𝑋2 = 3 . The regression function (1.3) now is;
𝐸(𝑌) = 12 + 3𝑋1 + 5(3) = 27 + 3𝑋1 (1.4)
Note that this response function is a straight line with slope 𝛽1 = 3. The same is true for
any other value of 𝑋2; only the intercept of the' response function will differ. Hence, 𝛽1 =
3 indicates that the mean response 𝐸 {𝑌} increases by 3 with a unit increase in 𝑋1 when
Page 110 of 174
STA202: Statistics II
𝑋2 is constant, no matter what the level of 𝑋2. We confirm therefore that 𝛽1 indicates the
change in 𝐸{𝑌} with a unit increase in 𝑋1 when 𝑋2is held constant.

The parameters 𝛽1 and 𝛽2 are sometimes called partial regression coefficients because
they reflect the partial effect of one predictor variable when the other predictor variable is
included in the model and is held constant.

We can readily establish the meaning of 𝛽1 and 𝛽2 by calculus, taking partial derivatives
of the response surface (1.2) with respect to 𝑋1 and 𝑋2 as follows:
𝜕𝐸{𝑌}
= 𝛽1
𝜕𝑥1
𝜕𝐸{𝑌}
= 𝛽2
𝜕𝑥2
The partial derivatives measure the rate of change in 𝐸{𝑌} with respect to one predictor
variable when the other is held constant.

7.1.1 First-Order Model with More than Two Predictor Variables

We consider now the case where there are 𝑝 − 1 predictor variables 𝑋1 , … , 𝑋𝑝−1 The
regression model
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.5)

is called a first-order model with 𝑝 − 1 predictor variables. It can also be written:

𝑝−1

𝑌𝑖 = 𝛽0 + ∑ 𝛽𝑘 𝑋𝑖𝑘 + 𝜀𝑖 (1.6)
𝑘=1

or, if 𝑋𝑖0 = 1 We let, it can be written as:

Page 111 of 174


STA202: Statistics II
𝑝−1

𝑌𝑖 = ∑ 𝛽𝑘 𝑋𝑖𝑘 + 𝜀𝑖 𝑤ℎ𝑒𝑟𝑒 𝑋𝑖0 = 1 (1.7)


𝑘=1

Assuming that E[𝜀𝑖 ]= 0, the response function for regression model 1.6) is:

𝑌𝑖 = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 + ⋯ + 𝛽𝑝−1 𝑋𝑝−1 (1.8)

When 𝑝 − 1 = 1, regression model (1.6) reduces to:


𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝜀𝑖
which is the simple linear regression model

7.1.2 General Linear Regression Model

For variables 𝑋1 , … , 𝑋𝑝−1 in a regression model, the general linear regression model, with
normal error terms, simply in terms of X variables is defined as
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.9)
Where:
𝛽0 , 𝛽1 , … , 𝛽𝑝−1 𝑎𝑟𝑒 𝑝𝑎𝑟𝑎𝑚𝑡𝑒𝑟𝑠
𝑋𝑖1 , … , 𝑋𝑖,𝑝−1 𝑎𝑟𝑒 𝑘𝑛𝑜𝑤𝑛 𝑐𝑜𝑛𝑠𝑡𝑎𝑛𝑡𝑠
𝜀𝑖 𝑎𝑒 𝑖𝑛𝑑𝑒𝑝𝑒𝑒𝑛𝑑𝑒𝑛𝑡 𝑁(0, 𝜎 2 )
𝑖 = 1, … . 𝑛
If we let 𝑋𝑖0 = 1, we have
𝑌𝑖 = 𝛽0 𝑋𝑖0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖 (1.10)
𝑝−1

𝑌𝑖 = 𝛽0 + ∑ 𝛽𝑘 𝑋𝑖𝑘 + 𝜀𝑖 𝑋𝑖0 = 1 (1.11)


𝑘=1

The response function for regression model (1.9) is, since E[𝜀𝑖 ]= 0:
𝐸(𝑌𝑖 ) = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 (1.12)

Page 112 of 174


STA202: Statistics II
Thus, the general linear regression model with normal error terms implies that the
observations 𝑌𝑖 are independent normal variables, with mean (𝑌𝑖 ) as given by (1.12) and
with constant variance 𝜎 2 .

7.1.3 Qualitative Predictor Variables

Regression model takes qualitative factors as well as quantitative factors, Consider a


regression analysis to predict the length of hospital stay (𝑌) based on the age (𝑋1 ) and
gender (𝑋2 ) of the patient. We define 𝑋2 as follows:

1 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑚𝑎𝑙𝑒
𝑋2={
0 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑓𝑒𝑚𝑎𝑙𝑒

The first-order regression model then is as follows:


𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖 (1.13)
Where:
𝑋𝑖1 = 𝑝𝑎𝑡𝑖𝑒𝑛𝑡′𝑠 age

1 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑚𝑎𝑙𝑒
𝑋𝑖2={
0 𝑖𝑓 𝑝𝑎𝑡𝑖𝑒𝑛𝑡 𝑖𝑠 𝑓𝑒𝑚𝑎𝑙𝑒

7.1.4 Polynomial Regression

Polynomial regression models are special cases of the general linear regression model.
They contain squared and higher-order terms of the predictor variable (s), making the
response function curvilinear. The following is a polynomial regression model with one
predictor variable

Page 113 of 174


STA202: Statistics II
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖 + 𝛽2 𝑋𝑖2 + 𝜀𝑖 (1.14)
it is a special case of general linear regression model (1.9). If we let 𝑋𝑖1 = 𝑋𝑖 and 𝑋𝑖2 =
𝑋𝑖2 , we can write (1.14) as follows:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖
which is in the form of general linear regression model (1.9). While (1.14) illustrates a
curvilinear regression model where the response function is quadratic, models with
higher-degree polynomial response functions are also particular cases of the general
linear regression model. Polynomial regression will be treated in details later.

7.1.5 Variables Transformation

Models with transformed variables involve complex, curvilinear response functions, yet
still are special cases of the general linear regression model. Consider the following
model with a transformed 𝑌 variable:
𝑙𝑜𝑔𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝛽3 𝑋𝑖3
+ 𝜀𝑖 (1.15)
If we let 𝑌𝑖′ = 𝑙𝑜𝑔𝑌𝑖 , we can write regression model in (1.15) as follows
𝑌𝑖′ = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝛽3 𝑋𝑖3 + 𝜀𝑖
which is in the form of general linear regression model (1.9). The response variable just
happens to be the logarithm of 𝑌. Many models can be transformed into the general linear
regression model. For instance, the model

1
𝑌𝑖 = (1.16)
𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖

can be transformed to the general linear regression model by letting


𝑌𝑖′ = 1⁄𝑌
𝑖

We then have:

Page 114 of 174


STA202: Statistics II
𝑌𝑖′ = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + 𝜀𝑖

7.1.6 General Linear Regression Model in Matrix Terms

In order to use matrix form the important point is to know what the inverse of the matrix
represents. To express general linear regression model (1.9):

𝑌𝑖 = 𝛽0 + 𝛽1 𝑋𝑖1 + 𝛽2 𝑋𝑖2 + ⋯ + 𝛽𝑝−1 𝑋𝑖,𝑝−1 + 𝜀𝑖


in matrix terms, we need to define the following matrices:

 Y1  1 X 11 X 1, p 1 
Y  1 X X 2, p 1 
Y   2 X 
21
n1   n p  
   
Yn  1 X n1 X n , p 1 

 0   1 
   
   1    2
p1   n1  
   
n   n 

With this compact notation, the linear regression model can be written in the form
y  X  
In linear algebra terms, the least-squares parameter estimates β are the vectors that
Minimize
𝑛

∑ 𝜀𝑖2 = 𝜀′𝜀 = (𝒚 − 𝑿𝛽)′ (𝒚 − 𝑿𝛽)


𝑖=1

If 𝜀𝑖 = 0, then
̂ = 𝑿𝛽̂
𝒚
𝑿′ (𝒚 − 𝑿𝛽̂ ) = 0

Page 115 of 174


STA202: Statistics II
𝑿 𝒚 − 𝑿 𝑿𝛽̂ = 0
′ ′

𝑿′ 𝒚 = 𝑿′ 𝑿𝛽̂
𝒀 and 𝜺 vectors are the same as for simple linear regression. The 𝛽 vector contains
additional regression parameters, and the 𝑿 matrix contains a column of 𝐼𝑠 as well as a
column of the n observations for each of the 𝑝 − 1 𝑋 variables in the regression model.
The row subscript for each element 𝑋𝑖1in the 𝑿 matrix identifies the trial or case, and the
column subscript identifies the 𝑿 variable.

In matrix terms, the general linear regression model (1.9) is:

Y = X β+ ε
n×1 n×p p×1 p×1

(1.17)
where:
𝒀 is a vector of responses
𝜷 is a vector of parameters
𝑿 is a matrix of constants
𝜺 is a vector of independent normal random variables with expectation
𝑬{ 𝜺} = 0 and variance-covariance matrix:

 2 0 0
 
0 2 0
 (ε)  
2
  2I
nn  
 
0 0 2
Therefore, the random vector 𝑌 has expectation
𝑬{𝒀} = 𝑿𝜷
(1.18)

and the variance-covariance matrix of Y is the same as that of 𝜀:

Page 116 of 174


STA202: Statistics II
 (Y)   I
2 2

n n

(1.19)

7.2 Estimation of Regression Coefficients

From equation (1.19) using least square approach, we have


𝑛

𝑆 = ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋𝑖1 − 𝛽2 𝑋𝑖2 − ⋯


𝑖=1
2
− 𝛽𝑝−1 𝑋𝑖,𝑝−1 ) (1.20)
The least square estimators are those values of 𝛽0 , 𝛽1 , … 𝛽𝑝−1 that minimize 𝑆. Let us
denote the vector of the least squares estimated regression coefficients 𝑏0 , 𝑏1 , … 𝑏𝑝−1 as 𝒃:

 b0 
 b 
b 
1 
(1.21)
p×1  
 
b p 1 
The least square normal equations for the general linear regression model (1.17) are:
X'Xb = X'Y (1.22)

and the least square estimators are:

b   X'X  X ' Y
1

For a 3-variable model, 𝑌, 𝑋1 𝑎𝑛𝑑 𝑋2


𝑛 𝑛

∑ 𝜀 = ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 )2
2

𝑖=1 𝑖=1
𝑛

∑ 𝜀 2 𝑏𝑒 𝑟𝑒𝑝𝑟𝑒𝑠𝑒𝑛𝑡𝑒𝑑 𝑤𝑖𝑡ℎ 𝑆
𝑖=1

Page 117 of 174


STA202: Statistics II
𝑛

𝑆 = ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 )2
𝑖=1

𝜕𝑆 𝜕𝑆 𝜕𝑆
= 0; = 0; =0
𝜕𝛽0 𝜕𝛽1 𝜕𝛽2
𝜕𝑆
= −2 ∑(𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽0
𝜕𝑆
= −2 ∑ 𝑋1 (𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽1
𝜕𝑆
= −2 ∑ 𝑋2 (𝑌𝑖 − 𝛽0 − 𝛽1 𝑋1 − 𝛽2 𝑋2 ) = 0
𝜕𝛽2
Consequently, normal equation is obtained as follows
∑ 𝑌 = 𝑛𝛽̂0 + 𝛽̂1 ∑ 𝑋1+𝛽̂2 ∑ 𝑋2 (1.23)

∑ 𝑋1 𝑌 = 𝛽̂0 ∑ 𝑋1 + 𝛽̂1 ∑ 𝑋1 2 +𝛽̂2 ∑ 𝑋1 𝑋2 (1.24)

∑ 𝑋2 𝑌 = 𝛽̂0 ∑ 𝑋2 + 𝛽̂1 ∑ 𝑋1 𝑋2+𝛽̂2 ∑ 𝑋2 2 (1.25)

Putting the normal equation in matrix form we have


∑𝑌 𝒏 ∑ 𝑋1 ∑ 𝑋2 𝛽̂0
′ ′
𝑿 𝒀 = [∑ 𝑋1 𝑌]; 𝑿 𝑿 = [∑ 𝑋1 ∑ 𝑋1 2 ∑ 𝑋1 𝑋2 ]; 𝜷 = [𝛽̂1 ]
∑ 𝑋2 𝑌 ∑ 𝑋2 ∑ 𝑋1 𝑋2 ∑ 𝑋2 2 𝛽̂3

ˆ   X'X  X ' Y
1

V ( ˆ )   X'X   e2
1

𝑆𝑆𝑅 =Sum of square regression


𝑆𝑆𝑇 =Sum of square Total

𝑅 2 is coefficient of determination. It is used to used to test for goodness of fit of a model


so it measures the total variation in 𝑌 which has been explained by the variation of 𝑋.

Page 118 of 174


STA202: Statistics II

𝑆𝑆𝐸 𝑒′𝑒 𝑌 ′ 𝑌 − 𝛽′𝑋′𝑌


𝜎𝑒2 = = = (1.26)
𝑛−𝑘 𝑛−𝑘 𝑛−𝑘

Were 𝑘 is the number of parameters in the model

𝑉(𝛽̂1 ) 𝐶𝑜𝑣(𝛽̂1 𝛽̂2 )


𝑉(𝛽̂ ) = (𝑋′𝑌)−1 𝜎𝑒2 = [ ]
𝐶𝑜𝑣(𝛽̂1 𝛽̂2 ) 𝑉(𝛽̂2 )

∑ 𝑋2 2 ∑ 𝑋1 2
𝑉(𝛽̂1 ) = 𝜎𝑒2 ̂ 2
; 𝑉(𝛽2 ) = 𝜎𝑒
⎹𝑋 ′ 𝑋⎹ ⎹𝑋 ′ 𝑋⎹

2
𝑆𝑆𝑅 𝛽′𝑋′𝑌 𝑌 ′ 𝑌 − 𝑒′𝑒 𝑒 ′𝑒
𝑅 = = = = 1− ′ (1.27)
𝑆𝑆𝑇 𝑌′𝑌 𝑌′𝑌 𝑌𝑌
𝑆𝑆𝐸
= 1−
𝑆𝑆𝑇
𝑆𝑆𝐸 𝑛 − 1
𝑅̅ 2 = 1 − [( )( )] (1.28)
𝑆𝑆𝑇 𝑛 − 𝑘
𝑅̅ 2 is the adjusted 𝑅 2

Another way of computing sum of squares are as follows

𝑆𝑆𝑇 = ∑(𝑌 − 𝑌̅)2


2
𝑆𝑆𝐸 = ∑(𝑌 − 𝑌̂)
2
𝑆𝑆𝑅 = ∑(𝑌̂ − 𝑌̅) = (𝑆𝑆𝑇 − 𝑆𝑆𝐸)

The column headed MS refers to the mean square and is obtained by dividing the SS term
by the 𝑑𝑓 term. Thus, MSR, the mean square regression, is equal to 𝑆𝑆𝑅/𝑘, and MSE
equals 𝑆𝑆𝐸/ [𝑛 − (𝑘 + 1)]. The general format of the ANOVA table is:

Page 119 of 174


STA202: Statistics II

Analysis of variance
Source df SS MS F
Regression 𝑘 SSR 𝑀𝑆𝑅 = 𝑆𝑆𝑅/𝑘 𝑀𝑆𝑅
/𝑀𝑆𝐸
Error 𝑛 − (𝑘 SSE 𝑀𝑆𝐸 = 𝑆𝑆𝐸/
+ 1) 𝑛 − (𝑘 + 1)
Total 𝑛−1 SST

Coefficient of multiple determination: The proportion of the variation in the


dependent variable, 𝑌, that is explained by the set of independent variables 𝑋1,
𝑋2,… 𝑋𝑘 .

7.2.1 Global Test: Testing Whether the Multiple Regression Model is Valid

The overall ability of the independent variables 𝑋1, 𝑋2,… 𝑋𝑘 , to explain the behaviour of
the dependent variable 𝑌 can be tested. Two tests of hypotheses are considered. The first
one is called the global test.
Global test: An overall test of the regression model. It investigates the possibility that all
the regression coefficients are equal to zero.
It tests the overall ability of the set of independent variables to explain differences in the
dependent variable. The null hypothesis is that all of the population regression
coefficients are zero. If accepted, it would imply that the set of coefficients is of no value
in explaining differences in the dependent variable. The alternate hypothesis is that at least
one of the coefficients is not zero. This test is written in symbolic form for three
independent variables as:
𝐻0 ∶ 𝛽1 = 𝛽2 = 𝛽3 = 0
𝐻1 ∶ 𝑁𝑜𝑡 𝑎𝑙𝑙 𝑡ℎ𝑒 𝛽 𝑠 = 0
Rejecting 𝐻0 and accepting 𝐻1 implies that one or more of the independent variables

Page 120 of 174


STA202: Statistics II
is useful in explaining differences in the dependent variable. However, a word of
caution, it does not suggest how many or identify which regression coefficients are
not zero. Note also that 𝛽𝑖 denotes the population value of the slope, whereas 𝑏𝑗 , a
point estimate of 𝛽𝑗 , is computed from sample data.

𝐸𝑆𝑆/(𝑘 − 1)
𝐹= (1.29)
𝑅𝑒𝑔 𝑆𝑆/[𝑛 − 𝑘]
Where:
𝑆𝑆𝑅 is the sum of the squares “explained by” the regression.
𝑘 is the number of independent variables.
𝑆𝑆𝐸 is the sum of squares error.
𝑛 is the number of observations.

7.2.2 Evaluating Individual Regression Coefficients

The second test of hypothesis identifies which of the set of independent variables are
significant predictors of the dependent variable. That is, it tests the independent variables
individually rather than as a unit. This test is useful because unimportant variables can be
eliminated from the regression model. The test statistic is the Student 𝑡 distribution with
𝑛 − (𝑘 + 1) degrees of freedom. For example, suppose we want to test whether the
hypothesis that the coefficient for the first independent variable in the model was equal to
zero versus the alternative hypothesis that it was not equal to zero. The null and alternate
hypotheses would be written as follows.

𝐻0 ∶ 𝛽1 = 0
𝐻1 ∶ 𝛽1 ≠ 0

Rejection of the null hypothesis and acceptance of the alternate hypothesis would
imply that variable number one is significant and that it has an inverse relationship
with the dependent variable.

Page 121 of 174


STA202: Statistics II

7.2.3 Testing Individual Regression Coefficients

𝛽̂1
𝑇= ~𝑡𝑛−2 (1.30)
𝑆. 𝐸(𝛽̂1 )

If the hypothesis test finds that the null hypothesis cannot be rejected, then the variable
should be dropped from the model. However, the above test only supports removing one
variable at a time from the model. After a variable is removed, a new regression model is
constructed using the remaining variables and a new t-test can be conducted for each of
the remaining variables.

7.2.4 Qualitative Independent Variables (Dummy Variables)

The variables used in regression analysis have been mostly been quantitative variables. A
quantitative variable is a variable that is numerical in nature, such as, the number of hours
worked by employees, the number of traffic accidents ion 3rd mainland bridge in a week,
and the distance travelled to work.

However, frequently we want to use nominal-scale variables such as gender, whether a


home has a swimming pool, or whether an answer is yes or no. These are called
qualitative variables. To use a qualitative variable in a regression model we use a scheme
of dummy variables in which one of the two possible conditions is coded 0 and the other
1.

Page 122 of 174


STA202: Statistics II
7.2.5 Dummy Variable

A variable in which there are only two possible outcomes is called dummy variable. For
analysis, one of the outcomes is coded a 1 and the other a 0 as previously discussed in
section 1.3.

Case Study 7.1


Given that
 25 1 18 
30  9 21
y   and X  
15  8 15 
   
 20  12 16 
By carrying out an analysis of the model 𝑌𝑖 = 𝛽0 + 𝛽1 𝑋1 + 𝛽2 𝑋2 + 𝜀𝑖

(a) Obtain the linear regression model of the form y  X   ,


(b) Deduce 𝑅 2 and adjusted 𝑅 2 and interpret your result
(c) Compute
(d) Test 𝐻0 : 𝛽1 = 𝛽2 = 0
(e) Test the significance of 𝑋2
Solution
𝑌 𝑋1 𝑋2 𝑋12 𝑋22 𝑌2 𝑋1 𝑋2 𝑋1 𝑌 𝑋2 𝑌
2 1 1 1 32 62 18 25 450
5 8 4 5
3 9 2 81 44 90 189 270 630
0 1 1 0
1 8 1 64 22 22 120 120 225
5 5 5 5
2 1 1 14 25 40 192 340 320
0 2 6 4 6 0
Table 7.1

Page 123 of 174


STA202: Statistics II

∑𝑌 𝒏 ∑ 𝑋1 ∑ 𝑋2 𝛽̂0
′ ′
𝑿 𝒀 = [∑ 𝑋1 𝑌]; 𝑿 𝑿 = [∑ 𝑋1 ∑ 𝑋1 2 ∑ 𝑋1 𝑋2 ]; 𝜷 = [𝛽̂1 ]
∑ 𝑋2 𝑌 ∑ 𝑋2 ∑ 𝑋1 𝑋2 ∑ 𝑋2 2 𝛽̂3

4 30 70
𝑿′ 𝑿 = [30 290 519 ]
70 519 1246
⎹𝑿′ 𝑿⎹ = 4(91,979) − 30((1050) + 70(−4730)
= 5,316
290 519
+𝑀11 = | | = (290 ∙ 1246) − 5192 = 91,979
519 1246

30 519
−𝑀12 = | | = −[(30 ∙ 1246) − (519 ∙ 70)] = −1,050
70 1246
It follows that:

+𝑀11 −𝑀12 +𝑀13 91,979 −1,050 −4730


𝑪=[ −𝑀21 +𝑀22 −𝑀23 ]=[−1,050 84 +24 ]
+𝑀31 −𝑀32 +𝑀33 −4,730 24 260
91,979 −1,050 −4730
′ 𝑇
𝐴𝑑𝑗(𝑋 𝑋) = 𝐶 = [−1,050 84 +24 ]
−4,730 24 260
1
(𝑋′𝑋)−1 = 𝑎𝑑𝑗(𝑋 ′ 𝑋)
𝑑𝑒𝑡(𝑋 ′ 𝑋)

1 91,979 −1,050 −4730


(𝑋′𝑋)−1 = [−1,050 84 +24 ]
5,316 −4,730 24 260
∑𝑌 90

𝑿 𝒀 = [∑ 𝑋1 𝑌]=[ 655 ]
∑ 𝑋2 𝑌 1625
𝛽̂ = (𝑋′𝑋)−1 𝑋′𝑌

1 91,979 −1,050 −4730 90


[−1,050 84 +24 ] [ 655 ]
5,316 −4,730 24 260 1625

Page 124 of 174


STA202: Statistics II
1 −95,890
[ −480 ]
5,316 12,520

𝛽̂0 −18.0380
𝛽̂ = [𝛽̂1 ] = [ −0.0903 ]
𝛽̂2 2.3552

(a) 𝑌̂ = −18.0380 − 0.0903𝑋1 + 2.3552𝑋2

𝑌 ′ 𝑌 = ∑ 𝑌 2 = 2,150

90
𝑅𝑒𝑔 𝑆𝑆=𝛽̂ ′ 𝑋 ′ 𝑌 = (−18.0380 −0.0903 )
2.3550 655 )
(
1625

𝑅𝑒𝑔 𝑆𝑆 = 2,144.63
𝐸𝑟𝑟𝑜𝑟 𝑆𝑆 = 𝑇𝑆𝑆 − 𝑅𝑒𝑔 𝑆𝑆

Source Df SS MS F
Regression 2 2,144.63 1072.315 199.68
Error 1 5.37 5.37
Total 3 2150
Table 7.2: Analysis of Variance
𝑆𝑆𝑅 2144.63
(b) 𝑅 2 = 𝑆𝑆𝑇 = = 0.9975
2150

5.37 4 − 1
𝑅̅ 2 = 1 − [( )( )]
2150 4 − 2

= 1 − (0.0024977 × 1.5) = 0.9962


Adjusted 𝑅 2 = 0.9962
𝑆𝑆𝐸 𝑒′𝑒 𝑌 ′ 𝑌−𝛽′𝑋′𝑌
(c) 𝜎𝑒2 = 𝑛−𝑘 = 𝑛−𝑘 = 𝑛−𝑘

Page 125 of 174


STA202: Statistics II
5.37
𝜎𝑒2 = = 2.685
4−2
∑ 𝑋2 2 1246
𝑉(𝛽̂1 ) = 𝜎𝑒2 ′
= 2.685 ∙ = 0.6291
⎹𝑋 𝑋⎹ 5316
∑ 𝑋1 2 84
𝑉(𝛽̂2 ) = 𝜎𝑒2 = 2.685 ∙ = 0.04242
⎹𝑋 ′ 𝑋⎹ 5316
𝐸𝑆𝑆/(𝑘 − 1)
𝐹=
𝑅𝑒𝑔 𝑆𝑆/[𝑛 − 𝑘]
(d)
5.37/1
𝐹= = 0.0035008
2144.63/2
𝐹0.05 (1,3) =10.13

The sample fall short of the 5% critical value

(e) To test the significant of 𝑋2 , we test 𝐻0 : 𝛽2 = 0, so we employ t-test


𝛽̂1
𝑡= ~𝑡𝑛−2
𝑆. 𝐸(𝛽̂1 )

𝑆. 𝐸(𝛽̂1 ) = √𝑉(𝛽̂1)=0.2059

2.3550
𝑡= = 11.4376
0.2059

𝑡2 0.05 = 2.132
It is significant at 5%.

Case Study 7.2


A regression model takes the following form:
𝑌𝑖 = 𝛼 + 𝛽1 𝑋1𝑖 + 𝛽2 𝑋2𝑖 + 𝛾1 𝑍1𝑖 + 𝛾2 𝑍2𝑖 +𝜀𝑖
where the errors 𝜀𝑖 are normally distributed. The least square estimates based on a dataset of 25
observations, together with the associated standard errors, are given below.

Page 126 of 174


STA202: Statistics II

Coefficient Standard Error

Intercept 0.315 0.338


𝑋1 0.049 0.023
𝑋2 -0.036 0.022
𝑍1 0.206 0.183
𝑍2 0.048 0.209

Regression Statistics
Adjusted R Square 0.636027664
𝐹 3.679636867
Observations 25
Table 7.3
Assume that the distribution of the least square estimates is approximately normal.
a) Test the null hypothesis that the coefficient 𝛽1 = 0 versus the alternative hypothesis of 𝛽1 ≠ 0 at
the 10% significance level.
b) Construct a 95% confidence interval for 𝛽2.
c) Test the overall significance of the regression model at the 5% level.
d) A researcher wonders whether the model would be improved by removing
𝑍1 and 𝑍2 . What are null and alternative hypotheses which are appropriate for this case?

Solution
(a)
𝛽̂1
𝑡= ~𝑡𝑛−2
𝐸. 𝑆. 𝐸(𝛽̂1 )
𝛽̂1 = 0.049, 𝐸. 𝑆. 𝐸(𝛽̂1 ) = 0.023
0.049
𝑡= = 2.13
0.023

Page 127 of 174


STA202: Statistics II
𝑡0.10,23 = 1.319, from T-Statistical Table, Since 2.13>1.319, we reject the null hypothesis that 𝛽1 =
0, in favour of alternative hypothesis that 𝛽1 ≠ 0.
(𝑏)
𝛽̂2 − 𝑡𝛼⁄2,𝑛−2 × 𝐸. 𝑆. 𝐸(𝛽̂2 )

(𝛽̂2 − 𝑡𝛼⁄2,𝑛−2 × 𝑆. 𝐸(𝛽̂2 ), 𝛽̂2 − 𝑡𝛼⁄2,𝑛−2 × 𝑆. 𝐸(𝛽̂2 ))

(𝛽̂2 − 𝑡0.025 × 𝐸. 𝑆. 𝐸(𝛽̂2 ), 𝛽̂2 + 𝑡0.025 × 𝐸. 𝑆. 𝐸(𝛽̂2 ))


𝛽̂2 = −0.036, 𝑔𝑖𝑣𝑒𝑛
𝐸. 𝑆. 𝐸(𝛽̂1 ) = 0.022, 𝑔𝑖𝑣𝑒𝑛
𝑡0.025 = 2.069 , from the statistical Table.
(−0.036 − 2.069 × 0.022, −0.036 + 2.069 × 0.022)
(−0.08152, 0.09518)
(c) 𝛼 = 0.1, the null hypothesis is single tail
𝐹2⁄23 , 0.05 = 3.422, and 𝐹 calculated is 3.67963 > 3.422
The 𝐹 calculated is greater than the 𝐹 critical, we reject the null hypothesis that the overall model is
significant.
(d) 𝐻0 : 𝛽1 − 𝛽2 = 0 Versus 𝐻𝐴 : 𝛽1 − 𝛽2 ≠ 0

7.3 Polynomial Regression Model

A model is said to be linear when it is linear in parameters. So the model


𝑦 = 𝛽0 + 𝛽1 𝑥 + 𝛽2 𝑥 2 + 𝜀 (7.1)
And
𝑦 = 𝛽0 + 𝛽1 𝑥1 + 𝛽2 𝑥2 + 𝛽11 𝑥12 + 𝛽22 𝑥22 + 𝜀 (7.2)
are also the linear model. And they are the second order polynomials in one and two variables
respectively.
The polynomial models can be used in those situations where the relationship between study and
explanatory variables is curvilinear. Sometimes a nonlinear relationship in a small range of
explanatory variable can also be modelled by polynomials.
The 𝑘 𝑡ℎ order polynomial model in one variable is given by

Page 128 of 174


STA202: Statistics II
2 𝑘
𝑦 = 𝛽0 + 𝛽1 𝑥 + 𝛽2 𝑥 +. . +𝛽𝑘 𝑥 + 𝜀 (7.3)

7.3.1 Uses of Polynomial Models

Polynomial regression models have two basic types of uses:


l. When the true curvilinear response function is indeed a polynomial function.
2. When the true curvilinear response function is unknown (or complex) but a polynomial
function is a good approximation to the true function.

7.3.2 One Predictor Variable-Second Order

Polynomial regression models may contain one, two, or more than two predictor variables.
Each predictor variable may be present in various powers. We start by considering a
polynomial regression model with one predictor variable raised to the first and second powers:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽2 𝑥𝑖2 + 𝜀𝑖 (7.4)
Where 𝑥𝑖 = 𝑋𝑖 − 𝑋̅
This polynomial model is called a second-order model with one predictor variable because the
single predictor variable is expressed in the model to the first and second powers. Note that the
predictor variable is centred-in other words, expressed as a deviation around its mean 𝑋̅ −and
that the ith centered observation is denoted by 𝑥𝑖 .
The regression coefficients in polynomial regression are frequently written in a slightly
different fashion, to reflect the pattern of the exponents:
𝑌𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽11 𝑥𝑖2 + 𝜀𝑖 (7.5)
The response function for regression model (7.5) is:
𝐸{𝑌}𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽11 𝑥𝑖2 (7.6)
This response function is a parabola and known as a quadratic response function.
The regression coefficient 𝛽0 represents the mean response of 𝑌 when 𝑥 = 0, i.e., wh
𝑋 = 𝑋̅. The regression coefficient 𝛽1 is often called the linear effect coefficient, and 𝛽11
is known as the quadratic effect coefficient.
The algebraic version of the least square normal equations:
Page 129 of 174
STA202: Statistics II

𝑿 𝑿𝒃 = 𝑿′𝒀
for the second-order polynomial regression model (7.5) can be readily obtained from. Since
∑ 𝑥𝑖 = 0, this yields the normal equations:

∑ 𝑌𝑖 = 𝑛𝑏0 + 𝑏11 ∑ 𝑥𝑖2

∑ 𝑥𝑖 𝑌𝑖 = 𝑏1 ∑ 𝑥𝑖2 + 𝑏11 ∑ 𝑥𝑖3 (7.7)

∑ 𝑥𝑖2 𝑌𝑖 = 𝑏0 ∑ 𝑥𝑖2 + 𝑏1 ∑ 𝑥𝑖3 + 𝑏11 ∑ 𝑥𝑖4

7.3.3 One Predictor Variable-Third Order

The regression model:


𝑌 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽11 𝑥𝑖2 + 𝛽111 𝑥𝑖3 + 𝜀 (7.8)
is a third-order model with one predictor variable. The response function for regression
model (7.8) is:
𝐸{𝑌} = 𝛽0 + 𝛽1 𝑥 + 𝛽11 𝑥 2 + 𝛽111 𝑥 3 (7.9)

7.3.4 Two Predictor Variables-Second Order

The regression model:


𝑌𝑖
2 2
= 𝛽0 + 𝛽1 𝑥𝑖1 + 𝛽2 𝑥𝑖2 + 𝛽11 𝑥𝑖1 + 𝛽22 𝑥𝑖2 + 𝛽12 𝑥𝑖1 𝑥𝑖2 + 𝜀𝑖 (7.10)
𝑤ℎ𝑒𝑟𝑒:
𝑥𝑖1 = 𝑋𝑖1 − 𝑋̅1
𝑥𝑖2 = 𝑋𝑖2 − 𝑋̅2
is a second-order model with two predictor variables. The response function is:
𝐸{ 𝑌𝑖 } = 𝛽0 + 𝛽1 𝑥1 + 𝛽2 𝑥2 + 𝛽11 𝑥12 + 𝛽22 𝑥22 + 𝛽12 𝑥1 𝑥2 (7.11)

Page 130 of 174


STA202: Statistics II
which is the equation of a conic section. Note that regression model (7.11) contains separate
linear and quadratic components for each of the two predictor variables and a cross-product
term. The coefficient 𝛽12 is often called the interaction effect coefficient.

7.3.5 Fitting Polynomial in One Variable

i. Order of the model: The order of the polynomial model is kept as low as possible. Some
transformations can be used to keep the model to be of first order. If this is not satisfactory,
then second order polynomial is tried. Arbitrary fitting of higher order polynomials can be a
serious abuse of regression analysis. A model which is consistent with the knowledge of data
and its environment should be taken into account. It is always possible for a polynomial of
order (𝑛 – 1) to pass through 𝑛 points so that a polynomial of sufficiently high degree can
always be found that provides a “good” fit to the data. Such models neither enhance the
understanding of the unknown function nor be a good predictor.
ii. Model building strategy: A good strategy should be used to choose the order of an
approximate polynomial. One possible approach is to successively fit the models in increasing
order and test the significance of regression coefficients at each step of model fitting. Keep the
order increasing until t-test for the highest order term is non-significant. This is called as
forward selection procedure. Another approach is to fit the appropriate highest order model and
then delete terms one at a time starting with highest order. This is continued until the highest
order remaining term has a significant t-statistic. This is called as backward elimination
procedure. The forward selection and backward elimination procedures does not necessarily
lead to same model. The first and second order polynomials are mostly used in practice.
iii. Extrapolation: One has to be very cautious in extrapolation with polynomial models. The
curvatures in the region of data and region of extrapolation can be different. So predicted
response will not be based on the true behaviour of the data.
iv. Ill-conditioning: A basic assumption in linear regression analysis is that 𝑋-matrix is of
full column rank. In polynomial regression models, as the order increases, the 𝑋’𝑋 matrix
becomes ill-conditioned. As a result, the matrix (𝑋’𝑋)−1 may not be accurate and the
parameters will be estimated with considerable error. If values of 𝑥 lie in a narrow range
then the degree of ill-conditioning increases and multicollinearity in the columns of

Page 131 of 174


STA202: Statistics II
2
𝑋 matrix enters. For example, if 𝑥 varies between 2 and 3 , then 𝑥 varies between 4 and 9.
This introduces strong multicollinearity between 𝑥 and 𝑥 2 .
v. Hierachy: A model is said to be hierarchical if it contains the terms 𝑥, 𝑥 2 , … , 𝑥 𝑛 . in a
hierarchy. For example, the model
𝑦 = 𝛽0 + 𝛽1 𝑥 + 𝛽2 𝑥 2 + 𝛽3 𝑥 3 + 𝛽4 𝑥 4 + 𝜀
is hierarchical as it contains all the terms up to order four. The model

𝑦 = 𝛽0 + 𝛽1 𝑥 + 𝛽2 𝑥 2 + 𝛽4 𝑥 4 + 𝜀

is not hierarchical if does not contain the term of 𝑥 3 .


It is expected that all polynomial models should have this property because only
hierarchical models are invariant under linear transformation. This requirement is more
attractive from mathematics point of view. In many situations, the need of model may be
different. For example, the model
𝑦 = 𝛽0 + 𝛽1 𝑥1 + 𝛽12 𝑥1 𝑥2 + 𝜀
needs a two factor interaction which is provided by the cross-product term. A hierarchical
model would need inclusion of 𝑥2 which is not needed from the point of view of statistical
significance perspective.

7.4 Analysis

Consider the polynomial model of order 𝑘 is one variable as


𝑦𝑖 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽2 𝑥𝑖2 +. . +𝛽𝑘 𝑥𝑖𝑘 + 𝜀𝑖 𝑖 = 1, 2, … , 𝑛 (7.12)
If this model is written as
𝑦 = 𝑋𝜷 + 𝜀
the columns of 𝑋 will not be orthogonal. If we add another term 𝛽𝑘+1 𝑥𝑖𝑘+1 then the matrix
(𝑋′𝑋)−1 has to be recomputed and consequently, the lower order parameters 𝛽̂0 , 𝛽̂1 , … , 𝛽̂𝑘 will also
change.
Consider the fitting of the following model:
𝑦𝑖 = 𝛼0 𝑃0 (𝑥𝑖 ) + 𝛼1 𝑃1 (𝑥𝑖 ) + 𝛼2 𝑃2 (𝑥𝑖 ) … + 𝛼𝑘 𝑃𝑘 (𝑥𝑖 )+ 𝜀𝑖 𝑖 = 1,2, … , 𝑛 (7.13)

Page 132 of 174


STA202: Statistics II
𝑡ℎ
Where 𝑝𝑢 (𝑥𝑖 ) is the 𝑢 order orthogonal defines as
𝑛

∑ 𝑝𝑟 (𝑥𝑖 )𝑝𝑠 (𝑥𝑖 ) = 0, 𝑟 ≠ 𝑠 𝑟, 𝑠 = 0, 1, 2, … , 𝑘 (7.14)


𝑖=1

𝑝0 (𝑥𝑖 ) = 1
When we have the model as 𝑦 = 𝑋𝜷 + 𝜀, 𝑋 − matrix, in this case is given by
𝑝0 (𝑥1 ) ⋯ 𝑝𝑘 (𝑥1 )
𝑋=[ ⋮ ⋱ ⋮ ] (7.15)
𝑝0 (𝑥𝑛 ) ⋯ 𝑝𝑘 (𝑥𝑛 )
Since this 𝑋 -matrix has orthogonal columns, so 𝑋’𝑋 matrix becomes
𝑛

∑ 𝑃02 (𝑥𝑖 ) ⋯ 0
𝑖=1
𝑋’𝑋 = ⋮ ⋱ ⋮ (7.16)
𝑛

0 ⋯ ∑ 𝑃𝑘2 (𝑥𝑖 )
[ 𝑖=1 ]
The ordinary least squares estimator is 𝛼̂ = (𝑋′𝑋)−1 𝑋′𝑦 which is for 𝛼𝑗 is
∑𝑛𝑖=1 𝑃𝑗 (𝑥𝑖 )𝑦𝑖
𝛼̂𝑗 = 𝑛 , 𝑗 = 0,1,2, … , 𝑘 (7.17)
∑𝑖=1 𝑃𝑗2 (𝑥𝑖 )

and its variance is obtained from 𝑉𝑎𝑟(𝛼̂) = 𝜎 2 (𝑋′𝑋)−1 as


𝜎2
𝑉𝑎𝑟(𝛼̂𝑗 ) = 𝑛 (7.18)
∑𝑖=1[𝑃𝑘2 (𝑥𝑖 )]2
When is 𝜎 2 unknown, it can be estimated from the analysis of variance table.
Since 𝑝0 (𝑥1 ) is a polynomial of order zero, set it as 𝑝0 (𝑥𝑖 ) = 1 and consequently
𝛼̂0 = 𝑦̂ − 𝑦̅
The residual sum of squares is
𝑘 𝑛 𝑛

𝑆𝑆𝑟𝑒𝑠𝑑 = 𝑆𝑆𝑇 − ∑ 𝛼̂𝑗 [∑ ∑ 𝑃𝑗 (𝑥𝑖 )𝑦𝑖 ] (7.19)


𝑗=1 𝑖=1 𝑖=1

The regression sum of squares is


𝑛 𝑛

𝑆𝑆𝑟𝑒𝑔𝑟 (𝛼̂𝑗 ) = 𝛼̂𝑗 ∑ ∑ 𝑃𝑗 (𝑥𝑖 )𝑦𝑖


𝑖=1 𝑖=1

Page 133 of 174


STA202: Statistics II
2
[∑𝑛𝑖=1 𝑃𝑗 (𝑥𝑖 )𝑦𝑖 ]
= (7.20)
∑𝑛𝑖=1 𝑃𝑗2 (𝑥𝑖 )

This regression sum of squares does not depend on other parameters in the model.
The analysis of variance table, in this case, is given as follows
Source of variation Degree of freedom Sum of Squares Mean squares
𝛼̂0 1 𝑆𝑆(𝛼̂0 ) −
𝛼̂1 1 𝑆𝑆(𝛼̂1 ) 𝑆𝑆(𝛼̂1 )
𝛼̂2 1 𝑆𝑆(𝛼̂2 ) 𝑆𝑆(𝛼̂2 )
⋮ ⋮ ⋮ ⋮
𝛼̂𝑘 1 𝑆𝑆(𝛼̂𝑘 ) 𝑆𝑆(𝛼̂𝑘 )
Residual 𝑛−𝑘−1 𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘) (by subtraction) 𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘)
Total 𝑛 𝑆𝑆𝑇
Table 7.4
If we add another term 𝑃𝑘+1 (𝑥𝑖 )𝛼𝑘+1 in the model, then the model is

𝑦𝑖 = 𝛼0 𝑃0 (𝑥𝑖 ) + 𝛼1 𝑃1 (𝑥𝑖 ) + 𝛼2 𝑃2 (𝑥𝑖 ) … + 𝛼𝑘+1 𝑃𝑘+1 (𝑥𝑖 )+ 𝜀𝑖 𝑖 = 1,2, … , 𝑛 (2.21)


then we just need 𝛼̂𝑘+1 which can be obtained as
∑𝑛𝑖=1 𝑃𝑘+1 (𝑥𝑖 )𝑦𝑖
𝛼̂𝑘+1 = (2.22)
∑𝑛𝑖=1[𝑃𝑘+1
2
(𝑥𝑖 )]2

Notice that:
(i) We need not to bother for other terms in the model.
(ii) Simply concentrate on the newly added term only.
(iii) No re-computation of (𝑋′𝑋)−1or any other 𝛼̂𝑗 (𝑗 ≠ 𝑘 + 1) is necessary due to orthogonality of
polynomials.
(iv) Thus higher-order polynomials can be fitted with ease.
(v) Terminate the process when a suitably fitted model is obtained.

Page 134 of 174


STA202: Statistics II
7.4.1 Test of Significance:

To test the significance of the highest order term, we test the null hypothesis
𝐻0 : 𝛼𝑘 = 0
This hypothesis is equivalent 𝑡𝑜 𝐻0 : 𝛽𝑘 = 0 in polynomial regression model.
The following apply
𝑆𝑆𝑟𝑒𝑔 (𝛼𝑘 )
𝐹0 = (2.23)
𝑆𝑆𝑟𝑒𝑠𝑑 (𝑘)/(𝑛 − 𝑘 − 1)
~𝐹(1, 𝑛 − 𝑘 + 1) under 𝐻0
If the order of the model is changed to (𝑘 + 𝑟), we need to compute only r new coefficients.
The remaining coefficients 𝛼̂0 , 𝛼̂1 , … , 𝛼̂𝑘 do not change due to the orthogonality property of
polynomials. Thus the sequential fitting of the model is computationally easy.
When 𝑋𝑖 are equally spaced, the tables of orthogonal polynomials are available, and the
orthogonal polynomials can be easily constructed.

First 5 orthogonal polynomials are as follows:


Let 𝑠 be the spacing between levels of 𝑥 and {𝜆𝑗 } be the constants chosen so that polynomials
will have integer values. The tables are available
𝑝0 (𝑥𝑖 ) = 1
𝑥𝑖 − 𝑥̅
𝑝1 (𝑥𝑖 ) = 𝜆1 [ ]
𝑠
𝑥𝑖 − 𝑥̅ 2 𝑛2 − 1
𝑝2 (𝑥𝑖 ) = 𝜆2 [( ) −( )]
𝑠 12

𝑥𝑖 − 𝑥̅ 3 𝑥𝑖 − 𝑥̅ 3𝑛3 − 7
𝑝3 (𝑥𝑖 ) = 𝜆3 [( ) −( )( )]
𝑠 𝑠 20
𝑥𝑖 − 𝑥̅ 4 𝑥𝑖 − 𝑥̅ 2 3𝑛3 − 13 3(𝑛2 − 1)(𝑛2 − 19)
𝑝4 (𝑥𝑖 ) = 𝜆4 [( ) −( ) ( )+ ]
𝑠 𝑠 14 560
An example of the table for n 𝑛 = 5 is as follows:
𝑥𝑖 𝑃1 𝑃2 𝑃3 𝑃4

Page 135 of 174


STA202: Statistics II
1 -2
2 -1
⋮ ⋮ ⋮ ⋮ ⋮
5
𝑛
2
10 14 10 70
∑{𝑃𝑗 (𝑥𝑖 )}
𝑖=1

𝜆 1 1 5⁄6 35⁄12
Table 7.5
The orthogonal polynomials can also be constructed when 𝑥’𝑠 are not equally spaced.

Case Study 7.3


Data: average claims paid per policy for automobile insurance in
New Brunswick in the years 1971-1980:
Year 1971 1972 1973 1974 1975 1976 1977 1978 1979 1980
Cost 45.13 51.71 60.17 64.83 65.24 65.17 67.65 79.80 96.13 115.19
One goal of analysis is to extrapolate the Costs for 2.25 years beyond the end of the data; this
should help the insurance company set premiums.
i. We fit polynomials of degrees from 1 to 5, plot the fits, compute error sums of squares
and examine the 5 resulting extrapolations to the year 1982.25.
ii. The model equation for a 𝑝𝑡ℎ degree polynomial is
𝑦 = 𝛽0 + 𝛽1 𝑥𝑖 + 𝛽2 𝑥𝑖2 +. . +𝛽𝑝 𝑥𝑖𝑝 + 𝜀
where the 𝑥𝑖 are the covariate values (the dates in the example).

𝑝 + 1 parameters (sometimes there will be p parameters in total and sometimes a total


of 𝑝 + 1 – the intercept plus 𝑝 others). And 𝛽0 is the intercept
The design matrix is given by
1 𝑥1 ⋯ 𝑥1𝑝
𝑋 = [⋮ 𝑥2 ⋱ ⋮]
1 𝑥𝑛 ⋯ 𝑥𝑛𝑝
We estimate βs by select good value of p. This presents a trade-off:
Page 136 of 174
STA202: Statistics II
i. large p fits data better but
ii. small p is easier to interpret
Using SAS proc glm 5 times, once for each model. The fitted models are

𝑦 = 71.102 + 6.3516𝑥
𝑦 = 64.897 + 6.3516𝑥 + 0.7521𝑡 2
𝑦 = 64.897 + 1.9492𝑥 + 0.7521𝑥 2 + 0.3005𝑥 3
𝑦 = 64.888 + 1.9492𝑥 + 0.7521𝑥 2 + 0.3005𝑥 3 − 0.0002𝑥 4
𝑦 = 64.888 − 0.5024𝑥 + 0.7521𝑥 2 + 0.8016𝑥 3 − 0.0002𝑥 4 − 0.0194𝑥 5
These lead to the following predictions for 1982.25
Degree 𝜇̂ 1982.25
1 113.98
2 142.04
3 204.74
4 204.50
5 70.260

Page 137 of 174


STA202: Statistics II

Summary of Study Session 7

In study session 7, you have learnt that:

1. Multiple regression is a situation where there two or more predictors, and its
analysis is one of the most widely used of all statistical methods.
2. Polynomial regression models contain squared and higher-order terms of the
predictor variable (s), making the response function curvilinear

Page 138 of 174


STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 7

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. 1. State whether the following statements are true or false. Give a brief
explanation.
2. (a) A negative chi-squared value shows that there is little association between the
variables tested.
3. (b) Similar observed and expected frequencies indicate strong evidence of an
association in an 𝑟 × 𝑐 contingency table.

Page 139 of 174


STA202: Statistics II

Glossary of Terms

MSR: the mean square regression


Linear: Having the property that the output is proportional to the input
Curvilinear: contained by or consisting of a curved line or lines
Criterion variable: the dependent variable in a variety of statistical modeling contexts,
including multiple regression, discriminant analysis, and canonical correlation
Trials: a test of the performance, qualities, or suitability
Parameters: a numerical or other measurable factor forming one of a set that defines a
system or sets the conditions of its operation
Extrapolation: the extension of a graph, curve, or range of values by inferring unknown
values from trends in the known data

Page 140 of 174


STA202: Statistics II

References

1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
7. [Link]
[Link]

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 141 of 174


STA202: Statistics II
Study Session 8: Partial Correlation

Introduction

Partial correlation is a measure of the strength and direction of a linear


relationship between two continuous variables whilst controlling for the effect
of one or more other continuous variables (also known as 'covariates' or 'control'
variables). Although partial correlation does not make the distinction between independent
and dependent variables, the two variables are often considered in such a manner (i.e., you
have one continuous dependent variable and one continuous independent variable, as well
as one or more continuous control variables).

Learning Outcomes for Study Session 8

On completion of this study session, you should be able to:

8.1 Explain simple correlation coefficient

8.2 Explain partial correlation coefficient

8.3 Test for association

Page 142 of 174


STA202: Statistics II
8.1 Simple Correlation Coefficient

Before looking at partial correlation, we shall revise simple correlation. The simple

correlation coefficient ( rx1x2 ) between the two pair of variables ( 𝒙𝟏, 𝒙𝟐 ) is defined as

n x1 x2   x1  x2
rx1x2  (8.1)
 n x 2   x 2   n x 2   x 2 
  1  1    2  2 

Case Study 8.1


Use table 1to calculate 𝑟𝑥1 𝑦 , and 𝑟𝑥2 𝑦 , 𝑟𝑥1 𝑥2
Table 1
Corn yield amount of amount of
per acre water applied fertilizer
y (‘000) per acre x1 applied per

(‘000) Liter acre x2


(’000) kg
3 1 4
4 3 2
2 2 3
6 2 1
5 4 3
Table 8.1
Solution
𝑦 𝑥1 𝑥2 𝑦2 𝑥12 𝑥22 𝑥1 𝑦 𝑥2 𝑦 𝑥1 𝑥2
3 1 4 9 1 16 3 12 4
4 3 2 16 9 4 12 8 6
2 2 3 4 4 9 4 6 6

Page 143 of 174


STA202: Statistics II
6 2 1 36 4 1 12 6 2
5 4 3 25 16 9 20 15 12
20 12 13 90 34 39 51 47 30
Table 8.2
𝑛=5

n x1 y   x1  y
rx1 y 
 n x 2   x 2    n y 2   y 2 
  1  1     

5 × 51 − (12)(20) 255 − 240 15


𝑟𝑥1 𝑦 = = =
√5 × 34 − 122 × √5 × 90 − 202 √26 × 50 √1300
15
= 0.4160
36.05551
n x2 y   x2  y
rx2 y 
 n x 2   x 2    n y 2   y 2 
  2  2     
n x1 x2   x1  x2
rx1x2 
 n x 2   x 2    n x 2   x 2 
  1  1    2  2 

Similarly, 𝑟𝑥2 𝑦 =−0.693 and 𝑟𝑥1 𝑥2 = −0.231 respectively

8.1.2 The Multiple Regression

We shall consider another approach to solving multiple regression problem. The least
square regression equation in three variables is given by
𝑦̂ = 𝛽1∙23 + 𝛽12∙3 𝑥1𝑖 + 𝛽13∙2
Let 𝑟𝑥1 𝑦 =𝑟12 = 0.4160,
𝑟𝑥2 𝑦 =𝑟13 = −0.693 and
𝑟𝑥1 𝑥2 =𝑟23 = −0.231
And

Page 144 of 174


STA202: Statistics II
1
𝛽1∙23 = [∑ 𝑦𝑖 − 𝛽12∙3 ∑ 𝑥1𝑖 − 𝛽13∙2 ∑ 𝑥2𝑖 ]
𝑛
𝑟12 − 𝑟13 𝑟23 𝑆1
𝛽12∙3 = ×
1 − 𝑟23 2 𝑆2
𝑟13 − 𝑟12 𝑟23 𝑆1
𝛽13∙2 = ×
1 − 𝑟23 2 𝑆3
1
𝑆12 = ∑(𝑦𝑖 − 𝑦̅)2
𝑛
1 2
𝑆22 = ∑(𝑥1 𝑖 − 𝑥̅1𝑖 )
𝑛
1 2
𝑆32 = ∑(𝑥2 𝑖 − 𝑥̅2𝑖 )
𝑛

Case Study 8.2


Use the data given in example 1 and results to obtain the least squares regression in three
variables.
Solution
∑ 𝑦 20 12 13
𝑦̅ = = = 4, 𝑥̅1 = = 2.4, 𝑥̅2 = =2
𝑛 5 5 5
𝑦 𝑦 − 𝑦̅ (𝑦 − 𝑦̅)2 𝑥1 𝑥1 − 𝑥̅1 (𝑥1 − 𝑥̅1 )2 𝑥2 𝑥2 − 𝑥̅ 2 (𝑥2 − 𝑥̅2 )2
3 -1 1 1 -1.4 1.96 4 1.4 1.96
4 0 0 3 0.6 0.36 2 -0.6 0.36
2 -2 4 2 -0.4 0.16 3 0.4 0.16
6 2 4 2 -0.4 0.16 1 -1.6 2.56
5 1 1 4 1.6 2.56 3 0.4 0.16
10 5.2 5.2
Table 8.3
1 10
𝑆12 = ∑(𝑦𝑖 − 𝑦̅)2 = = 2 ⇒ 𝑆1 = 1.414
𝑛 5
1 2 5.2
𝑆22 = ∑(𝑥1 𝑖 − 𝑥̅1𝑖 ) = = 1.04 ⇒ 𝑆2 = 1.02
𝑛 5
1 2 5.2
𝑆32 = ∑(𝑥2 𝑖 − 𝑥̅2𝑖 ) = = 1.04 ⇒ 𝑆3 = 1.02
𝑛 5
Page 145 of 174
STA202: Statistics II
𝑟12 − 𝑟13 𝑟23 𝑆1
𝛽12∙3 = ×
1 − 𝑟23 2 𝑆2
0.4160 − (−0.693)(−0.231) 1.414
= ×
1 − (−0.231)2 1.02
0.4160 − 0.160 1.414
= ×
1 − 0.0534 1.02
0.256 1.414 0.3619
= × =
0.9947 1.02 1.0145
= 0.3567
𝑟13 − 𝑟12 𝑟23 𝑆1
𝛽13∙2 = ×
1 − 𝑟23 2 𝑆3
−0.693 − (0.4160)(−0.231) 1.414
= ×
1 − (−0.231)2 1.02
−0.693 + 0.0960 1.414
= ×
1 − 0.0534 1.02
−0.597 1.414 −0.8442
= × =
0.9947 1.02 1.0145
= −0.8320
1
𝛽1∙23 = [∑ 𝑦𝑖 − 𝛽12∙3 ∑ 𝑥1𝑖 − 𝛽13∙2 ∑ 𝑥2𝑖 ]
𝑛
1
= [20 − 0.3567 × 12 − (−0.8320 × 13)]
5
1
[20 − 4.2804 + 10.816]
5
26.5356
= = 5.307
5
∴ The least square regression line is
𝑦̂ = 5.307 + 0.3567𝑥1 − 0.8320𝑥2

Page 146 of 174


STA202: Statistics II
8.2 Partial Correlation Coefficient

Partial correlation coefficient is the correlation coefficient between any two of a particular
three variables while the remaining one is kept constant. Thus, we can obtain the partial
correction coefficient between 𝑦 and 𝑥1 keeping 𝑥2 constant as follows:

ryx1  ryx2 rx1x2


ryx1 x2 
1  r 1  r 
2
yx2
2
x1 x2

Taking ryx1 x2  r123 that is

𝑦 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 1
𝑥1 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 2
And
𝑥2 𝑡𝑎𝑘𝑒𝑠 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛 𝑜𝑓 3

r12  r13r23
r123 
1  r 1  r 
∴ 2 2
12 23

Similarly,

ryx2  ryx1 rx1x2


ryx2 x1 
1  r 1  r 
2
yx1
2
x1 x2

r13  r12 r23


r132 
1  r 1  r 
2
12
2
23

Equation (10.15) is the partial correlation coefficient between 𝑦 and 𝑥2 keeping 𝑥1


constant
Also,

rx1x2  ryx1 ryx2


rx1x2  y 
1  r 1  r 
2
yx1
2
yx2

Page 147 of 174


STA202: Statistics II
r23  r12 r13
r231 
1  r 1  r 
2
12
2
13

Case Study 8.3


Use the given data and previous results to compute the partial correlation coefficient

ryx1 x2 , ryx2  x3 rx1x2  y , that is, r123 , r132 and 231r

Solution
ryx1  ryx2 rx1x2
ryx1 x2 
1  r 1  r 
2
yx2
2
x1 x2

0.4160   0.693  (0.231)


ryx1  x2 
1   0.693  1   0.231 
2 2

0.4160  0.1601

0.5193  0.9466
0.2559

0.4912
0.2559

0.7008
ryx1  x2  0.365

0.693   0.4160  (0.231) 


ryx2  x1 
1   0.4160  1   0.231 
2 2

0.693  0.0960

0.8269  0.9466

Page 148 of 174


STA202: Statistics II
0.597

0.7827
0.597

0.8847
ryx2  x1  0.6748

Table 8.4: SPSS output for partial Correlation coefficient between y and x2 taking

x1 as constant or the controlling variable


0.231   0.4160  (0.693) 
rx1x2  y 
1   0.4160  1   0.693 
2 2

0.231  0.2883

0.8269  0.5198
0.0573

0.4298
0.0573

0.6556
rx1x2  y  0.087

Page 149 of 174


STA202: Statistics II

Table 8.5: SPSS output for partial Correlation coefficient between x2 and x1 taking
y as constant or the controlling variable.

8.3 Tests for Association

This type of test tests the null hypothesis that two factors (or attributes) are not associated (that
is, independent), against the alternative hypothesis that they are associated. Each data unit we
sample has one level (or `type' or `variety') of each factor.
Case Study 8.4
Suppose that we are sampling people, and that one factor of interest is hair colour (black,
blonde, brown etc.) while another factor of interest is eye colour (blue, brown, green etc.). In
this example, each sampled person has one level of each factor. We wish to test whether or not
these factors are associated. Hence:
𝐻0 : There is no association between hair colour and eye colour
𝐻1 : There is an association between hair colour and eye colour.
So, under 𝐻0 , the distribution of eye colour is the same for blondes as it is for brunettes etc.,
whereas if 𝐻1 is true it may be attributable to blonde-haired people having a (significantly)
higher proportion of blue eyes, say.

Page 150 of 174


STA202: Statistics II
The association might also depend on the sex of the person, and that would be a third factor
which was associated with (i.e. interacted with) both of the others. The main way of analysing
these questions is by using a contingency table.

8.3.1 Contingency Tables

In a contingency table, also known as a cross-tabulation, the data are in the form of frequencies
(counts), where the observations are organised in cross-tabulated categories. We sample a
certain number of units (such as people) and classify them according to the two factors of
interest.
Case Study 8.5
In three areas of a city, a record has been kept of the numbers of Covid-19, Malaria and Lassa
fever which take place in a year. The total number of occurrences was 150, and they were
divided into the various categories as shown in the following contingency table:
Area Covid-19 Malaria Lassa fever Total
A 30 19 6 55
B 12 23 14 49
C 8 18 20 46

Total 50 60 40 150
Table 8.6
The cell frequencies are known as observed frequencies and show how the data are spread
across the different combinations of factor levels (known as contingencies). The first step
in any analysis is to complete the row and column totals (as already done in this table).

8.3.2 Expected Frequencies

We proceed by computing a corresponding set of expected frequencies, conditional on the


null hypothesis of no association between the factors, i.e. that the factors are independent.

Page 151 of 174


STA202: Statistics II
Now suppose that you are only given the row and column totals for the frequencies. If the
factors were assumed to be independent, consider how you would calculate the expected
frequencies. If A and B are two independent events, then 𝑃(𝐴 ∩ 𝐵) = 𝑃(𝐴) 𝑃(𝐵). We now
apply this idea.
Case Study 8.6
For the data in example above, if a record was selected at random from the 150 records:
P(a patient is Covid-19 from area 𝐴) = 50/150
P( a patient is in area 𝐴) = 55/150.
Hence, under 𝐻0 , we have:
P(a patient being covid-19 in area 𝐴)
50 55
= ×
150 150
and so the expected number of patients in area 𝐴 is:
50 55
= 150 × ×
150 150
So the expected frequency is obtained by multiplying the product of the `marginal' probabilities
by n, the total number of observations. This can be generalised as follows.

8.3.3 Expected Frequencies in Contingency Tables

The expected frequency, 𝐸𝑖𝑗 , for the cell in row i and column 𝑗 of a contingency table with 𝑟
rows and 𝑐 columns, is:
𝑟𝑜𝑤 𝑖 𝑡𝑜𝑡𝑎𝑙 × 𝑐𝑜𝑙𝑢𝑚𝑛 𝑗 𝑡𝑜𝑡𝑎𝑙
𝐸𝑖𝑗 =
𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛
Where 𝑟 = 1, … , 𝑟 𝑎𝑛𝑑 𝑗 = 1, … , 𝑐
Area Covid-19 Malaria Lassa fever
A 𝐸11 𝐸12 𝐸13
B 𝐸21 𝐸22 𝐸23
C 𝐸31 𝐸32 𝐸33
Table 8.7

Page 152 of 174


STA202: Statistics II
50 × 55
𝐸11 = = 18.33
150
Completing the remaining expectations we have
Area Covid-19 Malaria Lassa fever Total
A 18.33 22.00 14.67 55
B 16.33 19.60 13.07 49
C 15.33 18.40 12.27 46
Total 50 60 40 150
Table 8.8

8.3.4 Test Statistic

To motivate our choice of test statistic, if 𝐻0 is true then we would expect to observe small
differences between the observed and expected frequencies, while large differences
would suggest that 𝐻1 is true. This is because the expected frequencies have been
calculated conditional on the null hypothesis of independence hence, if 𝐻0 is actually
true, then what we actually observe (the observed frequencies) should be (approximately)
equal to what we expect to observe (the expected frequencies).

8.3.5 𝝌𝟐 Test of Association

Let the contingency table have 𝑟 rows and 𝑐 columns, then formally the test statistic used
for tests of association is:

𝑟 𝑐 2
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 )
∑∑ ~𝜒 2 (𝑟−1)(𝑐−1) (6.1)
𝐸𝑖𝑗
𝑖=1 𝑗=1

Critical values are found using Statistical Tables. The `double summation' here just means
summing over all rows and columns. This test statistic follows an (approximate) chi-
squared distribution with (𝑟 − 1)(𝑐 − 1) degrees of freedom, where 𝑟 and 𝑐 denote the
number of rows and columns, respectively, in the contingency table.

Page 153 of 174


STA202: Statistics II
For a 𝑟 × 𝑐 contingency table, we begin with 𝑟𝑐 cells. We lose one degree of freedom for
needing to use the total number of observations to compute the expected frequencies.
However, we also use the row and column totals in these calculations, but we only need
𝑟 − 1 row totals and 𝑐 − 1 column totals, as the final one in each case can be deduced
using the total number of observations. Hence we only lose 𝑟 − 1 degrees of freedom for
the row totals, similarly we only lose 𝑐 − 1 degrees of freedom for the column totals.
Hence the overall degrees of freedom are:

𝑣 = 𝑟𝑐 − (𝑟 − 1) − (𝑐 − 1) − 1
= (𝑟 − 1)(𝑐 − 1)

8.3.6 Performing the Test

As usual, we choose a significance level, 𝛼, at which to conduct the test. However, are we
performing a one-tailed test or a two-tailed test? To determine this, we need to consider
what sort of test statistic value would be considered extreme under 𝐻0 . If 𝐻0 is true, then
the observed and expected frequencies should be quite similar, since the expected
frequencies are computed conditional on the null hypothesis of independence. This
means that ⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ should be quite small for all cells. In contrast if 𝐻0 is not true,
then we would expect comparatively large values for⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ due to large differences
between the two sets of frequencies. Therefore, upon squaring ⎹𝑂𝑖𝑗 − 𝐸𝑖𝑗 ⎹ suffciently
large test statistic values suggest rejection of 𝐻0 . Hence 𝜒 2 tests of association are always
upper-tailed tests.

Case Study 8.7


Using above data, we proceed with the hypothesis test. Note it is advisable to present your
calculations as an extended contingency table as shown below, where the three rows in
each cell correspond to the observed frequencies, the expected frequencies and the test
statistic contributor.

Page 154 of 174


STA202: Statistics II
2
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) (𝑂11 − 𝐸11 )2 (30 − 18.33)2 11.67
= = = = 7.43
𝐸𝑖𝑗 𝐸11 18.33 18.33

Area Covid-19 Malaria Lassa fever Total


𝐴 𝑂1∙ 30 19 6 55
𝐸1∙ 18.33 22.00 14.67 55
2 7.43 0.41 5.15
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) ⁄𝐸𝑖𝑗
𝐵 𝑂1∙ 12 23 14 49
𝐸1∙ 16.33 19.60 13.07 49
2 1.13 0.59 0.06
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) ⁄𝐸𝑖𝑗
𝐶 𝑂1∙ 8 18 20 46
𝐸1∙ 15.33 18.40 12.27 46
2 3.48 0.01 4.82
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) ⁄𝐸𝑖𝑗
Total 50 60 40 150
Table 8.8
using
𝑟 𝑐 2
(𝑂𝑖𝑗 − 𝐸𝑖𝑗 )
∑∑ = 7.43 + 0.41 + ⋯ + 4.82 = 23.13
𝐸𝑖𝑗
𝑖=1 𝑗=1

Since 𝑟 = 𝑐 = 3, we have (𝑟 − 1)(𝑐 − 1) = (3 − 1)(3 − 1) = 4 degrees of freedom.


For α = 0.05, using Table 8 of the New Cambridge Statistical Tables, we obtain an upper-
tail critical value of 9.488. Hence we reject 𝐻0 since 9.488 < 23.13. Moving to the 1%
significance level, the critical value is now 13.28 so, again, we reject 𝐻0 since 13.28 <
23.13. Therefore, the test is highly significant and we conclude that there is strong
evidence of an association between the factors.

Looking again at the contingency table, comparing observed and expected frequencies,
the interpretation of this association becomes clear Covid-19 is the main problem in area 𝐴
whereas Lassa Fever is a problem in area 𝐶. (We can deduce this by looking at the cells

Page 155 of 174


STA202: Statistics II
with large test statistic contributors, which are a consequence of large differences between
observed and expected frequencies.)

Incidentally, the p-value for this upper-tailed test is (using a computer)


𝑃(𝑋 > 23.13) = 0.000119, where 𝑋~𝜒 2 4 , emphasising the extreme significance of
this test statistic value.

However, for data involving more factors, and more factor levels, this type of analysis can
be very insightful. Cells which make a large contribution to the test statistic value (i.e.
2
which have large values of (𝑂𝑖𝑗 − 𝐸𝑖𝑗 ) ⁄𝐸𝑖𝑗 should be studied carefully when determining
the nature of an association. This is because, in cases where 𝐻0 has been rejected, rather
than simply conclude that there is an association between two categorical variables, it is
helpful to describe the nature of the association.

8.3.7 Goodness-of-Fit Tests

In addition to tests of association, the chi-squared distribution is often used more generally
in so-called `goodness-of fit tests. We may, for example, wish to answer hypotheses such
as `Is it reasonable to assume the data follows a particular distribution? This justifies the
name `goodness-of-fit' tests, since we are testing whether or not a particular probability
distribution provides an adequate fit to the observed data. The null hypothesis will assert
that a specific hypothetical population distribution is the true one. The alternative
hypothesis is that this specific distribution is not the true one.

There is a special case when we are only dealing with one row or one column. This is
when we wish to test that the sample data are drawn from a (discrete) uniform
distribution, i.e. that each characteristic is equally likely.

Page 156 of 174


STA202: Statistics II
Case Study 8.9
Is a given die fair? If the die is fair, then the values of the faces (1, 2, 3, 4, 5 and 6) are all
equally likely so we have the following hypotheses:
𝐻0 : Score is uniformly distributed vs. 𝐻1 : Score is not uniformly distributed.

8.3.8 Observed and Expected Frequencies

As with tests of association, the goodness-of-fit test involves both observed and expected
frequencies. In all goodness-of-fit tests, the sample data must be expressed in the form of
observed frequencies associated with certain classifications of the data. Assume 𝑘
classifications, hence the observed frequencies can be denoted by 𝑂𝑖 , for 𝑖 = 1, … , 𝑘 .

Case Study 8.10


Extending the previous example, for a die the obvious classifications would be the six
faces. If the die is thrown n times, then our observed frequency data would be the number
of times each face appeared. Here 𝑘 = 6.
Recall that in hypothesis testing we always assume that the null hypothesis, 𝐻0 , is true. In
order to conduct a goodness-of-fit test, expected frequencies are computed conditional
on the probability distribution expressed in 𝐻0 . The test statistic will then involve a
comparison of the observed and expected frequencies. If 𝐻0 is true, then we would
expect small differences between these two sets of frequencies, while large differences
would indicate support for 𝐻1 . We now consider how to compute expected frequencies
for discrete uniform probability distributions.

8.3.9 Expected Frequencies in Goodness-of-Fit Tests

For discrete uniform probability distributions, expected frequencies are computed as:
1
𝐸𝑖 = 𝑛 × 𝑓𝑜𝑟 𝑖 = 1,2 … , 𝑘
𝑘
where 𝑛 denotes the sample size and 11⁄𝑘 is the uniform (i.e. equal, same) probability for
each characteristic.

Page 157 of 174


STA202: Statistics II
Expected frequencies should not be rounded, just as we do not round sample means, say.
Note that the final expected frequency (for the 𝑘𝑡ℎ category) can easily be computed
using the formula:

𝑘−1

𝐸𝑘 = 𝑛 − ∑ 𝐸𝑖 (6.2)
𝑖=1

This is because we have a constraint that the sum of the observed and expected
frequencies must be equal,5 that is:

𝑘 𝑘

∑ 𝑂𝑖 = ∑ 𝐸𝑖 (6.3)
𝑖=1 𝑖=1

which results in a loss of one degree of freedom (discussed below).

8.3.10 The Goodness-of-Fit Test

For a discrete uniform distribution with 𝑘 categories, observed frequencies 𝑂𝑖 and


expected frequencies 𝐸𝑖 , the test statistic is:

𝑘
(𝑂𝑖 − 𝐸𝑖 )2
∑ ~𝜒 2 (𝑘−1) 𝑎𝑝𝑝𝑟𝑜𝑥𝑖𝑚𝑎𝑡𝑙𝑦 𝑢𝑛𝑑𝑒𝑟 𝐻0 (6.4)
𝐸𝑖
𝑗=1

Note that this test statistic does not have a true 𝜒 2 (𝑘−1) distribution under 𝐻0 , rather it is
only an approximating distribution. An important point to note is that this approximation
is only good enough provided all the expected frequencies are at least there are 𝑘 − 1
degrees of freedom when testing a discrete uniform distribution. 𝑘 is the number of
categories, and we lose one degree of freedom due to the constraint that:

∑ 𝑂𝑖 = ∑ 𝐸𝑖

Page 158 of 174


STA202: Statistics II

As with the test of association, goodness-of-fit t tests are upper-tailed tests as, under H0,
we would expect to see small differences between the observed and expected
frequencies, as the expected frequencies are computed conditional on 𝐻0 . Hence large
test statistic values are considered extreme under 𝐻0 , since these arise due to large
differences between the observed and expected frequencies.

Case Study 8.11


A confectionery company is trying out different wrappers for a chocolate bar- its original,
𝐴, and two new ones, 𝐵 and 𝐶. It puts the bars on display in a supermarket and looks to
see how many of each wrapper type have been sold in the first hour, with the following
results.
Wrapper Type A B C Total
Observed 8 10 15 33
frequency
Table 8.9
Is there a difference between wrapper types in the consumer choices made?
To answer this we need to test:

𝐻0 : There is no difference in preference for the wrapper types.


vs.
𝐻1 : There is a difference in preference for the wrapper types.

Expected is the mean of the entries

33
𝐸𝑖 = = 11
3

our test statistic value is:

Page 159 of 174


STA202: Statistics II
𝑘
(𝑂𝑖 − 𝐸𝑖 )2 (8 − 11)2 (10 − 11)2 (15 − 11)2
∑ = + + = 2.364
𝐸𝑖 11 11 11
𝑗=1

The degrees of freedom will be 𝑘 − 1 = 3 − 1 = 2. At the 5% significance level, the


upper-tail critical value is 5.991, using Table 8 of the New Cambridge Statistical
Tables, so we do not reject 𝐻0 since 2.364 < 5.991. If we now consider the 10%
significance level6 the critical value is 4.605, so again we do not reject 𝐻0 since 2.364 <
4605. Hence the test is not significant. Therefore, there is no evidence of a strict
preference for a particular wrapper type based on the choices observed in the supermarket
during the first hour.

Case Study 8.12


Many people believe that when horse races, it has a better chance of winning if its starting
line-up position is closer to the rail on the inside of the track. The starting position of 1 is
closest to the inside rail, followed by position 2, and so on. The table below lists the
numbers of wins of horses in the different starting positions. Do the data support the claim
that the probabilities of winning in the different starting positions are not all the same?

Starting 1 2 3 4 5 6 7 8
Position
Number of 29 19 18 25 17 10 15 11
wins
Table 8.10
Solution
We test whether the data follow a discrete uniform distribution of 8 categories. Let
𝑝𝑖 = 𝑃(𝑋 = 𝑖), 𝑓𝑜𝑟 𝑖 = 1, . . . , 8.

We test the null hypothesis 𝐻0 : pi = 1=8, for 𝑖 = 1, . . . , 8.

Page 160 of 174


STA202: Statistics II
Note 𝑛 = 29 + 19 + 18 + 25 + 17 + 10 + 15 + 11 = 144. The expected
frequencies are 𝐸𝑖 = 144⁄8 = 18, for all 𝑖 = 1, . . . , 8.

Starting 1 2 3 4 5 6 7 8
Positio
n
𝑂𝑖 29 19 18 25 17 10 15 11
𝐸𝑖 18 18 18 18 18 18 18 18
𝑂𝑖 − 𝐸𝑖 11 1 0 7 −1 −8 −3 −7
(𝑂𝑖 6.72 0.06 0 2.72 0.06 3.66 0.50 2.72
− 𝐸𝑖 )2
/𝐸𝑖
Table 8.11
(𝑂𝑖 −𝐸𝑖 )2
Under 𝐻0 , ~𝜒 2 7 . At 5% significant level, the critical value is 14.067. Since
𝐸𝑖

14.067<16.34 we reject the null hypothesis. Turning to 1% significance level, the critical
value is 18.475. Since 16.34 < 18.475, we cannot reject the null hypothesis; hence we
conclude that the test is moderately significant. There is moderate, but not strong,
evidence to support the claim that the chances of winning in the different starting positions
are not all the same.

Page 161 of 174


STA202: Statistics II

Summary of Study Session 8

In study session 8, you have learnt that:

The simple correlation coefficient ( rx1x2 ) between the two pair of variables ( 𝒙𝟏, 𝒙𝟐) is

n x1 x2   x1  x2
defined as rx1x2 
 n x 2   x 2   n x 2   x 2 
  1  1    2  2 
1. The expected frequency, 𝐸𝑖𝑗 , for the cell in row i and column 𝑗 of a contingency table
with 𝑟 rows and 𝑐 columns, is:
𝑟𝑜𝑤 𝑖 𝑡𝑜𝑡𝑎𝑙 × 𝑐𝑜𝑙𝑢𝑚𝑛 𝑗 𝑡𝑜𝑡𝑎𝑙
𝐸𝑖𝑗 =
𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑜𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛
Where 𝑟 = 1, … , 𝑟 𝑎𝑛𝑑 𝑗 = 1, … , 𝑐
2. Let the contingency table have 𝑟 rows and 𝑐 columns, then formally the test
statistic used for tests of association is:
𝒓 𝒄 𝟐
(𝑶𝒊𝒋 − 𝑬𝒊𝒋 )
∑∑ ~𝝌𝟐 (𝒓−𝟏)(𝒄−𝟏)
𝑬𝒊𝒋
𝒊=𝟏 𝒋=𝟏

Page 162 of 174


STA202: Statistics II
Self-Assessment Questions (SAQs) for Study Session 8

Now that you have completed this study session, you can assess how well you have
achieved its learning outcomes by answering these questions. Write your answers in your
study diary and discuss them with your tutor at the next study support meeting. You can
check your answers with the notes on the Self-Assessment Questions at the end of this
session.

1. An experiment was conducted to examine whether age, in particular being over


30 or not, has any effect on preferences for a digital or an analogue watch.
Specifically, 129 randomly-selected people were asked what watch they prefer
and their responses are summarised in the table below

Analogue Undecided Digital watch


watch
30 year old or 10 17 37
younger
Over 30 years old 31 22 12

(a) Based on the data in the table, and without conducting any significance test, would you
say there is an association between age and watch preference?
Provide a brief justification for your answer.
(b) Calculate the 𝜒 2 statistic for the hypothesis of independence between age and watch
preference, and test that hypothesis. What do you conclude?

Page 163 of 174


STA202: Statistics II

Glossary of Terms

Null: having no elements, or only zeros as elements

Discrete: individually separate and distinct

Contingency table: a table showing the distribution of one variable in rows and another
in columns, used to study the correlation between the two variables

Controlling variable: anything that is held constant or limited in a research study

Expected frequencies: computed by multiplying the probability that an event occurs by


the total number of possible times that the event could occur

Discrete uniform distribution: have a finite number of outcomes

Page 164 of 174


STA202: Statistics II
References

1. Probability and Statistics for Engineers & Scientists by Walpole and Myers.
2. Introduction to Statistics. Jedidiah Publishers by Sojobi O.A.
3. Fundamentals of Statistics. Rasmed Publications by Shangodoyin & Agunbiade
4. Schaum’s Outline Series Theory and Problems of Probability (S.I. Metric) Edition
McGraw Hill Book Company, New York by Symour L.
5. An Introduction to Statistical Methods. Vikas Publishing House. Delhi by GUPTA
C. B.
6. Introductory Statistics (A learner’s Motivated Approach). Evan Brothers (Nigeria
Publishers) Limited by Afonja, B, Olubusoye O. E., Ossai E. and Arinola J. B.
7. [Link]
[Link]

Should you require more explanations on this study session? Please


do not hesitate to contact your e-tutor via the LMS.

Page 165 of 174


STA202: Statistics II
Notes on Self-Assessment Questions (SAQs)
Notes on Self-Assessment Questions for Study Session 1

1 Sample is preferred to population because is less costly, saves time, efficient


and is better especially when these resources are limited than population where
a large number of these resources is needed.
2
i. Simple random sampling is a sampling procedure in which every member
of the population has equal chance of being selected as a member of the
sample. It is mostly adopted in homogeneous (of same kind) population.
ii. Cluster sampling is a sampling method where the total population is
divided into a number of relatively small subdivisions, which are
themselves clusters of still smaller units, and then some of these sub-
divisions or clusters, are randomly selected for inclusion in the overall
sample.
iii. Systematic sampling is a sampling method used to select a sample in such a
way that every unit in the population will have an equal chance of being
selected in the sample.
3. In probability sampling, every unit of enquiry in the population has a known non-
zero probability of being included in the sample while in non-probability, the
probability of selecting a unit in the sample is not known and cannot be
determined.
4. Since the population mean is 𝜇 = 4 and the arithmetic mean of the sampling
distribution of the mean value is 𝜇𝑥̅ = 4. Therefore, the sample mean is an
unbiased estimator.
5. A sampling frame contains the basic details of all members of the population from
which samples are to be drawn
6. Sampling survey is a scientific method of selecting and using a representative part
(sample) of a whole to seek the truth about the whole

Page 166 of 174


STA202: Statistics II
Notes on Self-Assessment Questions for Study Session 2

1. The Pearson’s correlation coefficient, r = 0.70. Therefore, there is a fairly high


positive correlation between inflation rate and interest rate.
2. Correlation analysis Correlation is the degree of association between two or more
variables.
3. The rank correlation coefficient is R= 0.75 and comment there is a fairly high
positive correlation between income (X) and expenditure (Y) (in million naira) of
O.O.U, Ago-Iwoye.

Notes on Self-Assessment Questions for Study Session 3


1. Simple Regression Analysis is a statistical tool which helps to study the trend and

pattern of movement in one variable in response to changes in another variable on the

basis of an assumed relationship existing between them

2. To find the regression of line of X on Y

X = a + by
Where
𝑛 ∑ 𝑥𝑦 − ∑ 𝑥 ∑ 𝑦
𝑏=
𝑛 ∑ 𝑦 2 − (∑ 𝑦)2
𝑎 = 𝑥̅ − 𝑏𝑦̅
X Y Xy X2 Y2
66 69 4554 4356 4761
64 67 4288 4096 4489
68 69 4692 4624 4761
65 66 4290 4225 4356
69 70 4830 4761 4900
63 67 4221 3969 4489
71 69 4899 5041 4761

Page 167 of 174


STA202: Statistics II
67 66 4422 4489 4356
69 72 4968 4761 5184
68 68 4420 4624 4624
70 69 4830 4900 4761
72 71 5112 5184 5041
812 823 55,526 55,030 56,483

12(55,526) − 812 (823) 1,964


𝑏= = − = −4.206
12(56,483) − (823)2 467
𝑎 = 67.667— 4.206(68.583)
= 67.667 + 288.460
= 356.127
Regression equation of X on Y becomes
𝑋 = 𝑎 + 𝑏𝑦
𝑋 = 356.127 − 4.206𝑌
To find the regression analysis of line Y on X
𝑌 = 𝑎 + 𝑏𝑥 𝑤ℎ𝑒𝑟𝑒;
𝑛 ∑ 𝑥𝑦 − ∑ 𝑥𝑦
𝑏=
𝑛 ∑ 𝑥 2 − (∑ 𝑥)2
𝑎 = 𝑌̅ − 𝑏𝑥̅
12(55,526) − 812 (823)
𝑏=
12(56,030) − (812)2
1964
= − = −1.933
1016
𝑎 = 68.583— 1.933(67.667)
= 130.80
Regression equation of line Y on X becomes
𝑌 = 𝑎 + 𝑏𝑥
= 130.80 − 1.933𝑥

Page 168 of 174


STA202: Statistics II

Notes on Self-Assessment Questions for Study Session 4

1.
i. Statistical hypothesis is a statement about the parameters or form of a
population. A test of a statistical hypothesis is a criteria which specifies for
what sample results the hypothesis is to be accepted or rejected.
ii. A type I error has been committed if we reject the null hypothesis when it
is true and a type II error has been committed if we accept the null
hypothesis when it is false.
iii. A test of any statistical hypothesis where the alternative is one sided such
as:
H0:  = 0 or H0:  = 0
H1:  > 0 H1:  < 0
is called a one-tailed test. The critical region for H1:  > 0 lies entirely in the right tail
while the critical region for H1:  < 0 lies entirely in the left tail. A test of any statistical
hypothesis where the alternative is two-sided such as:
H0:  = 0 vs H1:   0
is called a two-tailed test, values in the both tails of the distribution constitute the
critical region.
2. The steps are:

i. Find the type of problem and the question to be answered.


ii. To state the null hypothesis (H0) and the appropriate alternative (H1)
hypothesis
iii. Selection of the appropriate test to be utilized and calculation of the test
criterion based on the type of test.
iv. Fixation of the level of significance 
v. Decision making on test criterion value, whether to reject or accept the
hypothesis.

Page 169 of 174


STA202: Statistics II
vi. Drawing of the conclusion (or inference) on the basis of level of significance is
deciding whether the difference observed is due to chance or due to some other
known factors.

3. Mine 1 (X1) Mine 2 (X2)

𝑋1 (𝑋1 − 𝑥̅ )2 𝑋2 (𝑋2 − 𝑥̅ )2
84 9 75 4
82 1 76 1
83 4 77 0
78 9 80 9
79 4 76 1
406 27 384 15

∑ 𝑋1 406
𝑥̅1 = = = 81.2 ~ 81
𝑛1 5

∑(𝑋1 − 𝑋̅2 ) 27
𝑆12 = = = 5.4
𝑛 5
𝑋̅1 − 𝑋̅2
𝑡=
𝑆1 𝑆
+ 2
√𝑛1 √𝑛2
81 − 77 4
= = = 2.197
2.32 1.73 1.82
+
√5 √5
𝑡𝑡𝑎𝑏𝑢𝑙𝑎𝑡𝑒𝑑 = 𝑡0.05 (𝑛1 + 𝑛2−2 ) = 𝑡0.05 (8) = 1.894
Interpretation
Null hypothesis (H0): P = 40% = 0.4
Alternate hypothesis (H1) = P ≠ 40% ≠ 0.4
Hence: P = 0.40, q = 0.60
Observed sample population
180
𝑝̂ = = 0.36
500

Page 170 of 174


STA202: Statistics II
Test statistic
𝑝̂ − 𝑝 0.36 − 0.40 −0.40 −0.04
𝑧= = = = = −0.579
𝑝𝑞 √0.00048 0.069
√ √0.24
𝑛 500
As (H0) is two-sided, we shall determine the rejection regions applying two-failed test
at 5% level of significance
𝑍5%
= 1.96
2

The observed value of Z is -0.579 which is the acceptance region and such H0 is
accepted.
Null hypothesis (H0): 𝑃̂1 = 𝑃̂2
Alternative Hypothesis (H1): 𝑃̂1 ≠ 𝑃̂2
450
𝑃̂1 = = 0.45 𝑞̂1 = 1 − 𝑃1 = 1 − 0.45 = 0.55, 𝑛1 = 1000
1000
400
𝑃̂2 = = 0.5 𝑞̂1 = 1 − 𝑞2 = 1 − 0.5 = 0.5, 𝑛2 = 400
800
The test statistic
𝑃̂1 − 𝑃̂2 0.45 − 0.50 −0.05
𝑍= = =
√0.00025 + 0.00031
̂ ̂ √0.45(0.55) + 0.5(0.5)
√𝑃1 𝑞̂1 + 𝑃2 𝑞̂2 1000 800
𝑛1 𝑛2

−0.05 −0.005
𝑍= = = −0.213
√0.000563 0.0237
𝑍𝑡𝑎𝑏𝑙𝑒 𝑎𝑡 1% = 1.64
𝑍𝑡𝑎𝑏𝑙𝑒 𝑎𝑡 5% = 1.96
The observed value of Z is -0.213 which is acceptance region at 1% and 5% level and
such H0 is accepted.

Notes on Self-Assessment Questions for Study Session 5


1. Index number is a technique of measuring changes in a variable or group of
variables with respect to time, geographical location or other characteristics
2.

Page 171 of 174


STA202: Statistics II
i. Index numbers are a special type of average.
ii. Index numbers are meant to study the changes in the effects of such
factors which cannot be measured directly.
iii. The technique of index numbers measures changes in one variable or
group of related variables.
iv. The technique of index numbers is used to compare the levels of a
phenomenon on a certain date with its level on some previous date
3. Problems of constructing index number are
i. Selection of Base Year
ii. Selection of Commodities
iii. Collection of Prices
iv. Selection of Average
v. Selection of Weights
vi. Purpose of Index Numbers
vii. Selection of Method

Notes on Self-Assessment Questions for Study Session 6

1. Time series is defined as some quantity that is measured sequentially in time over
some interval.
2. The components of time series are Trend, Cyclical movement, Seasonal movement
and Irregular movement.
3.
i. Additive model: This is a model in which the series value of y is the sum of
all four components, that is,

y = T + S + C+ I

ii. Multiplicative model: It is a del in which y is the product of all the time
series components. This is denoted by

Page 172 of 174


STA202: Statistics II
y = T * S * C* I

4. Moving average and Least Squares method


5.

PERIOD
Y t Yt t2
YEAR)

1996 6 -4 -24 16

1997 7 -3 -21 9

1998 7 -2 -14 4

1999 8 -1 -8 1

2000 19 0 0 0

2001 10 1 10 1

2002 13 2 16 4

2003 15 3 45 9

2004 17 4 68 16

Total ΣY = 102 Σt = 0 ΣYt = 72 Σt2 =60

Page 173 of 174


STA202: Statistics II
Substituting these values in the two given equations,
That is,
102 = 9a
a = 11.3
72 = 60b
b = 1.20
The Trend equation is : Y = 11.3 + 1.20 t

Page 174 of 174

Common questions

Powered by AI

The fundamental steps in planning a sample survey include defining the survey's objective, establishing the scope, determining subject coverage, choosing a method of data collection, organizing fieldwork, conducting pretests and pilot surveys, analyzing the survey data, and reporting the findings. Each step is crucial for ensuring that the data collected is reliable, valid, and applicable. Clearly defining objectives ensures the survey addresses specific questions; establishing scope ensures the correct population is studied; subject coverage ensures comprehensive data collection; method selection impacts data accuracy; organization ensures operational efficiency; pretests improve survey design; and data analysis and reporting ensure findings are communicated effectively .

Multiple regression analysis is used to predict the value of a dependent variable by modeling its relationship with two or more independent variables. The technique estimates the parameters of the regression equation, which represent the average change in the dependent variable for a one-unit change in each predictor while holding others constant. By analyzing the coefficients, multicollinearity, and interaction effects, it can predict and provide insights into the relationships between variables, thus aiding in informed decision-making and forecasting .

Index numbers have several practical limitations. They are never cent-percent accurate due to the inherent difficulties in computation. Selecting a representative basket of goods can be problematic due to changes in consumption patterns over time. Different countries may use disparate base years, limiting international comparability. Index numbers also measure only average changes, which can hide specific sectoral performances, and can be influenced by quality changes in goods rather than solely price changes. These factors complicate their use in accurately reflecting true economic changes .

Systematic sampling involves selecting every nth item in a list after a random start, unlike simple random sampling where each unit is selected entirely by random chance from the population. Its advantages include simplicity, ease of implementation, and reduced costs. It often yields a more evenly distributed sample across the population, which can improve representativeness, especially when population elements are naturally ordered. However, it assumes an absence of periodicity within the list that could bias results .

Type I error occurs when a true null hypothesis is incorrectly rejected, whereas Type II error occurs when a false null hypothesis is not rejected. Understanding their trade-off is important because minimizing one often increases the other. The level of significance (α) indicates the probability of a Type I error, while the power of the test (1-β) is the probability of correctly rejecting a false null hypothesis. Balancing these errors is crucial for designing tests that are both accurate and reliable, with implications in research validation and decision accuracy .

Stratified sampling increases the precision of estimates by dividing the population into non-overlapping sub-populations known as strata. Each stratum is more homogeneous compared to the entire population, which reduces variance within each subgroup. Sampling from each stratum ensures that each sub-population is adequately represented, leading to overall reduced variability in the estimate of the population mean. This typically results in higher accuracy and precision compared to a simple random sample of the same size, as it captures important variations within the population .

Cluster sampling is beneficial when a complete sampling frame is not available, making it impractical to conduct a simple random or stratified sampling. It is particularly useful in geographically dispersed populations. By dividing the population into clusters, which can be naturally occurring groups like neighborhoods or schools, and then randomly selecting clusters to assess, it reduces fieldwork and administrative costs. It is preferable when the population is large and spread out, and when budget and time constraints are significant .

Probability sampling methods, such as simple random sampling, ensure that every unit in the population has a known, non-zero probability of being selected, allowing for objective statistical inference about the population. Non-probability sampling methods, including judgement sampling and quota sampling, do not provide this assurance, making statistical inference less reliable. In non-probability sampling, selection is based on subjective judgment rather than randomization, which introduces potential biases and limits generalizability .

The sample mean is considered an unbiased estimator of the population mean because the expected value of the sample mean equals the population mean, irrespective of the population distribution's form. This means when samples of the same size are repeatedly drawn from the population, the average of their means will converge to the true population mean, reflecting no systematic overestimation or underestimation .

Constructing price index numbers involves challenges such as selecting appropriate base years, identifying representative items, collecting accurate price data, and assigning suitable weights. These factors introduce potential biases and discrepancies that can impact the index's reliability. Inaccurate or non-representative index numbers can lead to misinterpretations in economic policies, affecting inflation estimates, cost of living adjustments, and monetary policy decisions. Moreover, changes in consumption patterns and product quality further complicate this construction, potentially overstating or understating economic changes .

You might also like