0% found this document useful (0 votes)
5 views35 pages

Probability Theory Fundamentals

This document provides an overview of probability theory, including definitions of theoretical and empirical probability, and the laws governing probability calculations. It explains concepts such as mutually exclusive events, independent and dependent events, and includes examples to illustrate these principles. Additionally, it covers joint, marginal, and conditional probabilities, along with the Bayes theorem.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views35 pages

Probability Theory Fundamentals

This document provides an overview of probability theory, including definitions of theoretical and empirical probability, and the laws governing probability calculations. It explains concepts such as mutually exclusive events, independent and dependent events, and includes examples to illustrate these principles. Additionally, it covers joint, marginal, and conditional probabilities, along with the Bayes theorem.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MODULE ONE

PROBABILITY THEORY AND APPLICATIONS


Definitions of Probability
Probability is a concept that most people understand naturally, since such words as a chance,
likelihood, possibilities and proportion are used as part of everyday speech. It is a term used
in making decisions involving uncertainty. Though the concept is often viewed as very
abstract and difficult to relate to real world activities, it remains the best tool for solving
uncertainties problems.
There are basically two separate ways of calculating probability. 1. Calculation based on
theoretical probability. This is the name given to probability that is calculated without an
experiment that is, using only information that is known about the physical situation. 2.
Calculation based on empirical probability. This is probability calculated using the results
of an experiment that has been performed a number of times. Empirical probability is often
referred to as relative frequency or Subjective probability.
Definition of Theoretical Probability Let E represent an event of an experiment that has an
equally likely outcome set, U, then the theoretical probability of event E occurring when the
experiment is written as Pr (E) and given by:

𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑤𝑎𝑦𝑠 𝑡ℎ𝑟𝑜𝑢𝑔ℎ 𝑤ℎ𝑖𝑐ℎ 𝑒𝑣𝑒𝑛𝑡 𝑐𝑎𝑛 𝑜𝑐𝑐𝑢𝑟


Pr (E) = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑡 𝑝𝑜𝑠𝑠𝑖𝑏𝑙𝑒 𝑜𝑢𝑡𝑐𝑜𝑚𝑒𝑠

𝑛(𝐸)
= 𝑛(𝑈) where n (E) = the number of outcomes in event set E and n (U) = total possible number
of outcomes in outcome set, U. If, for example, an ordinary six-sided die is to be rolled, the
equally likely outcome set, U, is {1,2,3,4,5,6} and the event ―even number has event set
{2,4,6}. It follows that the theoretical probability of obtaining an even number can be
calculated as:
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑒𝑣𝑒𝑛 𝑛𝑢𝑚𝑏𝑒𝑟𝑠 3
Pr (even numbers) = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑝𝑜𝑠𝑠𝑖𝑏𝑙𝑒 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑡 𝑜𝑢𝑡𝑐𝑜𝑚𝑒 𝑛(𝑈) = 6 = 0.5

Other Examples
A wholesaler stocks heavy (2B), medium (HB), fine (2H) and extra fine (3H) pencils which
come in packs of 10. Currently in stock are 2 packs of 3H, 14 packs of 2H, 35 packs of HB
and 8 packs of 2B. If a pack of pencil is chosen at random for inspection, what is the
probability that they are: (a) medium (b) heavy (c ) not very fine (d) neither heavy nor
medium?
Solutions
Since the pencil pack is chosen at random, each separate pack of pencils can be regarded as
a single equally likely outcome. The total number of outcomes is the number of pencil packs,
that is, 2+14+35+8 = 59.
Thus, n (U) = 59
(c) Pr (not very fine). Note that the number of pencil packs that are not very fine is 14+35+8
= 57.
𝑛(𝑛𝑜𝑡 𝑣𝑒𝑟𝑦 𝑓𝑖𝑛𝑒) 57
Therefore, Pr(not very fine) = = 59 = 0.966
𝑛(𝑈)

(d) “Neither heavy nor medium” is equivalent to “fine” or “very fine” in the problem. There
is 2+14 = 16 of these pencil packs.
𝑛(𝑛𝑒𝑖𝑡ℎ𝑒𝑟 ℎ𝑒𝑎𝑣𝑦 𝑛𝑜𝑟 𝑚𝑒𝑑𝑖𝑢𝑚) 16
Thus, Pr (neither heavy nor medium) = = 59 = 0.271
𝑛(𝑈)
Definition of Empirical (Relative Frequency) Probability
If E is some event of an experiment that has been performed a number of times, yielding a
frequency distribution of events or outcomes, then the empirical probability of event E
occurring when the experiment is performed one more time is given by:
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡𝑖𝑚𝑒𝑠 𝑡ℎ𝑎𝑡 𝑒𝑣𝑒𝑛𝑡 𝑜𝑐𝑐𝑢𝑟𝑒𝑑 𝑓(𝐸)
Pr(E) = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡𝑖𝑚𝑒𝑠 𝑡ℎ𝑎𝑡 𝑒𝑥𝑝𝑒𝑟𝑖𝑚𝑒𝑛𝑡 𝑤𝑎𝑠 𝑝𝑒𝑟𝑓𝑜𝑟𝑚𝑒𝑑 = ∑𝑓

Where f(E) = the frequency of event E , Σf = total frequency of the experiment.


Put differently, the empirical probability of an event E occurring is simply the proportion of
times that event E actually occurred when the experiment was performed. For example, if,
out of 60 orders received so far this financial year, 12 were not completely satisfied, the
proportion, 12/60 = 0.2 is the empirical probability that the next order received will not be
completely satisfied.

Other Examples
A number of families of a particular type were measured by the number of children they have,
given the following frequency distribution:
Number of children 0 1 2 3 4 5
Number of families 12 28 22 8 2 2

Use this information to calculate the (relative frequency) probability that another family of
this type chosen at random will have: (a) 2 children (b) 3 or more children (c) less than 2
children.
Solutions
Here, Σf = total number of families = 74
𝑓(2 𝑐ℎ𝑖𝑙𝑑𝑟𝑒𝑛) 22
(a)Pr (2 children) = = = 0.297
∑𝑓 74
(b) f (3 or more children) = 8+2+2 = 12
Thus, Pr(3 or more children = 1274 = 0.162
𝑓(𝑙𝑒𝑠𝑠 𝑡ℎ𝑎𝑛 2 𝑐ℎ𝑖𝑙𝑑𝑟𝑒𝑛) 40
(c) Pr(less than 2 children) = = 74 =0.541
∑𝑓

Laws of Probability
There are four basic laws of probability.
1. Addition Law for mutually exclusive events
2. Addition Law for events that are not mutually exclusive
3. Multiplication Law for Independent events
4. Multiplication Law for Dependent events.
Addition Law for Mutually Exclusive Events

Two events are said to be mutually exclusive events if they cannot occur at the same time.
The addition law states that if events A and B are mutually exclusive events, then: Pr (A or
B) = Pr (A) + Pr (B)
Examples
The purchasing department of a big company has analysed the number of orders placed by
each of the 5 departments in the company by type as follows:
Table 9.1: Departmental Orders
Type of Order Sales Purchasing Production Accounts Maintenance Total
Consumables 10 12 4 8 4 38
Equipment 1 3 9 1 1 15
Special 0 0 4 1 2 7
Total 11 15 17 10 7 60

An error has been found in one of the orders. What is the probability that the incorrect order
came from
(a) Came from maintenance?
(b) Came from Production?
(c) Came from Maintenance or Production?
(d) Came from neither maintenance nor production?

Solutions:
a) since there are 7 maintenance orders out of the 60,
Pr (maintenance) = 7/60 = 0.117
b) Similarly, Pr (Production) = 17/60 = 0.283
c) Maintenance and production departments are two mutually exclusive events so that, Pr
(maintenance or production) = Pr (maintenance) + Pr(production = 0.117 + 0.283 = 0.40
d) Pr (neither maintenance nor production) = 1-Pr (maintenance or production) = 1-0.4 = 0.6
Addition Law for Events that are Not Mutually Exclusive Events
If events A and B are not mutually exclusive, that is, they can either occur together or occur
separately, then according to the Law:
Pr (A or B or Both) = Pr (A) +Pr (B) – Pr (A).Pr (B)
Example
Consider the following contingency table for the salary range of 94 employees:
Table 9.2: Contingency Table for the Salary range of 94 Employees
Salary /Month Men Women Total
N10,000 and Above 20 37 57
Below N10,000 15 22 37
Total 35 59 94
What is the probability of selecting an employee who is a man or earns below N10,000 per
month?
Solution
The two events of being a man and earning below N10, 000 is not mutually exclusive. It
follows that:
Pr(employee man OR earning below N10,000) = Pr(been a Man) + Pr(Earning below
N10,000) – Pr(been a Man).Pr(Employee earning below N10,000)
35 37 35 37
= 94 + 94 − 94 . 94
= 0.372 +0.394 – (0.372).(0.395)
= 0.766 – 0.147
= 0.619 0r 61.9%

Multiplication Law for Independent Events


This law states that if A and B are independent events, then:
Pr (A and B) = Pr (A). Pr (B)
As an example, suppose, in any given week, the probability of an assembly line failing is
0.03 and the probability of a raw material shortage is 0.1. If these two events are independent
of each other, then the probability of an assembly line failing, and a raw material shortage is
given by:
Pr (Assembly line failing and Material shortage) = (0.03)(0.1) = 0.003
Multiplication Law for Dependent Events
This Law states that if A and B are dependent events, then:
Pr (A and B) = Pr(A).Pr(B/A)
Note that Pr (B/A) in interpreted as probability of B given that event A has occurred.
Example
A display of 15 T-shirts in a Sports shop contains three different sizes: small,
medium and large. Of the 15 T-shirts:
3 are small
6 are medium
6 are large.
If two T-shirts are randomly selected from the T-shirts, what is the probability of selecting
both a small T-shirt and a large T-shirt, the first not being replaced before the second is
selected?
Solution
Since the first selected T-shirt is not replaced before the second T-shirt is selected, the two
events are said to be dependent events. It follows that:
Pr (Small T-shirt and Large T-shirt) = Pr(Small).Pr(Large/Small)
= (3/15)(4/14)
= (0.2)(0.429)
= 0.086
Computational Formula for Multiple Occurrence of an Event
The probability of an event, E occurring X times in n number of trials are given by the
formula:
Pr (En,x) = Cn,xPxq(n - x)

𝑛!
Where Cn,x = 𝑥!(𝑛−𝑥)!
p = probability of success
q = probability of failure
p+q=1

Example Assume there is a drug store with 10 antibiotic capsules of which 6 capsules are
effective and 4 are defective. What is the probability of purchasing the effective capsules
from the drug store?
Solution
From the given information: The probability of purchasing an effective capsule is:
P = 6/10 = 0.60
Since p + q = 1; q = 1 – 0.60 = 0.40; n = 10; x = 6
Pr (6EC) = probability of purchasing the 6 effective capsules
= C10,6(0.6)6(0.4)4
10!
= (0.047)(0.026)
(6!(10−6)!)
[Link].6!
= (0.0012)
6!.4!
[Link]
= (0.0012)
[Link]
= = 210(0.0012) = 0.252
Hence, the probability of purchasing the 6 effective capsules out of the 10 capsules is 25.2
percent.
Joint, Marginal, Conditional Probabilities, and the Bayes Theorem
Joint Probabilities
A joint probability implies the probability of joint events. Joint probabilities can be
conveniently analysed with the aid of joint probability tables.
The Joint Probability Table
A joint probability table is a contingency table in which all possible events for a variable are
recorded in a row and those of other variables are recorded in a column, with the values listed
in corresponding cells as in the following example.

Example: Consider a research activity with the following observations on the number of
customers that visit XYZ supermarket per day. The observations (or events) are recorded in
a joint probability table as follows:
Table 9.3: Joint Probability Table
Age(Years) Male (M) Female (F) Total
Below 30 (B) 60 70 130
30 and Above (A) 60 20 80
Total 120 90 210

We can observe four joint events from the above table:


Below 30 and Male (B∩M) = 60
Below 30 and Female (B∩F) = 70
30 and Above and Male (A∩M) = 60
30 and Above and Female (A∩F) = 20
Total events or sample space = 210
The joint probabilities associated with the above joint events are
60
Pr( B ∩ M ) = 210
Pr (B ∩ F) 70/120 = 0 3333
Pr (A ∩ M) 60/210 = 0.2857
Pr (A ∩ F) 20/210 = 0.0952

Marginal Probabilities
The Marginal Probability of an event is its simple probability of occurrence, given the sample
space. In the present discussion, the results of adding the joint probabilities in rows and
columns are known as marginal probabilities.
The marginal probability of each of the above events:
Male (M), Female (F), Below 30 (B), and Above 30 (A) are as follows:
Pr (M) = Pr(B∩M) + Pr(A∩M) = 0.2857 + 0.2857 = 0.57
Pr (F) = Pr(B∩F)+Pr(A∩F) = 0..3333 + 0.0952 = 0.43
Pr (B) = Pr(B∩M)+Pr(B∩F) = 0.2857+0.3333 = 0.62
Pr (A) = {r(A∩M)+Pr(A∩F) = 0.2857+0.0952 = 0.38
The joint and marginal probabilities above can be summarised in a contingency table as
follows:
Table 9.4: Joint and Marginal Probability Table.
Age(Years) Male (M) Female (F) Marginal Probability
Below 30 (B) 0.2857 0.3333 0.62
30 and Above (A) 0.2857 0.0952 0.38
Marginal Probability 0.57 0.43 1.00

Conditional Probability
Assuming two events, A and B, the probability of event A, given that event B has occurred is
referred to as the conditional probability of event A.
In symbolic term:
Pr( A∩ B) Pr(A).Pr(B)
Pr (A/B) = = = Pr(A)
Pr(B) Pr(B)
Where Pr (A/B) = conditional probability of event A
Pr (A∩B) = joint probability of events A and B
Pr (B) = marginal probability of event B
𝐽𝑜𝑖𝑛𝑡 𝑃𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑒𝑣𝑒𝑛𝑡𝑠 𝐴 𝑎𝑛𝑑 𝐵
In general, Pr(A/ B ) = 𝑀𝑎𝑟𝑔𝑖𝑛𝑎𝑙 𝑃𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑒𝑣𝑒𝑛𝑡 𝐵

The Bayes Theorem


Bayes theorem is a formula which can be thought of as ―reversing‖ conditional probability.
That is, it finds a conditional probability, A/B given, among other things, its inverse, B/A.
According to the theorem, given events A and B,

Pr(A).Pr(B/A)
Pr(A/ B ) = Pr(B)

As an example in the use of Bayes theorem, if the probability of meeting a business contract
date is 0.8, the probability of good weather is 0.5 and the probability of meeting the date given
good weather is 0.9, we can calculate the probability that there was good weather given that
the contract date was met.
Let G = good weather, and m = contract date was met
Given that: Pr (m) = 0.8; Pr (G) = 0.5; Pr (m/G) = 0.9, we need to find Pr (G/m):
From the Bayes theorem:
𝑃𝑟( 𝐺 )⋅ 𝑃𝑟( 𝑚 𝐺) (0.5)(0.9)
Pr( Gm ) = =
𝑃𝑟(𝑚) 0.8
= 0.5625 or 56.25%

Probability and Expected Values


The expected value of a set of values, with associated probabilities, is the arithmetic mean of
the set of values. If some variable, X, has its values specified with associated probabilities, P,
then:
Expected value of X = E (X) = Σ PX
Example
An ice-cream salesman divides his days into ‗Sunny‘ ‗Medium‘ or ‗Cold‘. He estimates that
the probability of a sunny day is 0.2 and that 30% of his days are cold. He has also calculated
that his average revenue on the three types of days is N220, N130, and N40 respectively. If
his average total cost per day is N80, calculate his expected profit per day.
Solution
We first calculate the different values of profit that are possible since we are required to
calculate expected profit per day, as well as their respective probabilities.
Given that Pr (sunny day) = 0.2; Pr (cold day) = 0.3
Since in theory, Pr (sunny day) + Pr (cold day) +Pr (medium day) = 1
It follows that:
Pr (medium day) = 1 - 0.2 - 0.3 = 0.5
The total costs are the same for any day (#80), so that the profits that the salesman makes on
each day of the three types of day are;
Sunnyday:N(220-80)=N140
Medium day: N (130-80) = N50
Cold day: N(40-80)= -N40 (loss)
We can summarise the probability table as follows:

BINOMIAL DISTRIBUTION
This is a discrete probability distribution use for a repeated random experiment that has only
two possible outcomes. These outcomes are usually called success and failure with
probability p and q respectively.
There are many experiments that conform either exactly or approximately to the following
list of requirements:
1. The experiment consists of a sequence of n smaller experiments called trials, where n is
fixed in advance of the experiment.
2. Each trial can result in one of the same two possible outcomes (dichotomous trials), which
we generically denote by success (S) and failure (F).
3. The trials are independent, so that the outcome on any particular trial does not influence
the outcome on any other trial.
4. The probability of success P(S) is constant from trial to trial; we denote this probability by
p.
Definition
An experiment for which Conditions 1–4 are satisfied is called a binomial experiment.
Examples of Binomial distribution are ;
(i) Experiment of tossing a fair coin repeatedly; the two possible outcomes are “a head” and
“not a head i.e. the tail ”.
(ii) The conduct of election in a ward; the outcomes are “winning in the election” and “losing
in the election”
PROPERTIES OF BINOMIAL DISTRIBUTION
If μ and σ respectively represents (denotes) the mean and standard deviation of a binomial
distribution, then:
(i) The mean , μ =np
(ii) The variance, σ2 =npq
(iii) The standard deviation, σ =√𝑛𝑝𝑞
𝑞−𝑝
(iv) Moment of coefficient of skewness =
√𝑛𝑝𝑞
1−6𝑝𝑞
(v) Moment of coefficient of kurtosis = 3 +
𝑛𝑝𝑞
Example:
A coin is tossed successively and independently n times.
We arbitrarily use S to denote the outcome H (heads) and F to denote the outcome T (tails).
Then this experiment satisfies Conditions 1–4.
Tossing a thumbtack n times, with S = point up and F = point down, also results in a binomial
experiment.
The Binomial Random Variable and Distribution
In most binomial experiments, it is the total number of S’s, rather than knowledge of exactly
which trials yielded S’s, that is of interest.
Definition
The binomial random variable X associated with a binomial experiment consisting of n
trials is defined as
X = the number of S’s (Success) among the n trials
Suppose, for example, that n= 3. Then there are eight possible outcomes for the experiment:
SSS SSF SFS SFF FSS FSF FFS FFF
From the definition of X, X(SSF) = 2, X(SFF) = 1, and so on. Possible values for X in an n-
trial experiment are x = 0, 1, 2, . . . , n. We will often write X ~ Bin(n, p) to indicate that X is
a binomial rv based on n trials with success probability p.
Notation
Because the probability mass function (pmf) of a binomial random variable X depends on the
two parameters n and p, we denote the pmf by b(x; n, p).

Example 1
Each of six randomly selected cola drinkers is given a glass containing cola S and one
containing cola F. The glasses are identical in appearance except for a code on the bottom to
identify the cola. Suppose there is actually no tendency among cola drinkers to prefer one
cola to the other.
Then p = P(a selected individual prefers S) = .5, so with X= the number among the six who
prefer S, X ~ Bin(6,.5).
Thus

The probability that at least three prefer S is


P(3 ≤ X) = ∑6𝑥=3 𝑏(𝑥; 6,5)
= ∑6𝑥=3(𝑥6) (0.5)𝑥 (0.5)6−𝑥
= .656
and the probability that at most one prefers S is
P(X ≤ 1) = ∑1𝑥=0 𝑏(𝑥; 6,5)
= .109
Example 2
If 75% of all purchases at a certain store are made with a credit card and X is the number
among ten randomly selected purchases made with a credit card, then X ~ Bin(10, .75).
Thus E(X) = np = (10)(.75) = 7.5,
V(X) = npq = 10(.75)(.25)
= 1.875,
and σ = √1.875
= 1.37.

Example 3:
In an examination 70% of the candidates pass. Use the binomial distribution to calculate the
probability that in a random sample of 10 candidates contains: (i) exactly 2 (ii) between 1 and
3 inclusive and (iii) at most 2 places.
Solution:
Let P be the probability that a candidate passes in the examination, then
70
P= = 0.7, q=1-p = 0.3, n=10
100
(i) P(exactly 2 passes) = P(x=2)
Using binomial distribution
P(x=r) = nCr Pr qn-r; x=0,1,2,3,……
10!
=2!(10−2)!(0.7)2 (0.3)8
= 45 x 0.49 x 0.0000656
=0.00145

(ii) P(x=1) =10C1 (0.7)1 (0.3)9


= 10 x 0.7 x 0.000019683
= 0.000138

P(x=2) =10C2 (0.7)2 (0.3)8


= 45 x 0.49 x 0.00006561
= 0.001447

P(x=3) =10C3 (0.7)3 (0.3)7


= 120 x 0.343 x 0.0002187
= 0.009001

P(1 and 3 inclusive) = P(1 ≤ x ≤ 3)


=P(1) + P(2) + P(3)
0.000138+ 0.001447+ 0.009001
0.010586 ≈ 0.01059

(iii) P(at most 2 passes) = P(x≤ 2)


= P(x=0) + P(x=1) + P(x=2)
P(x=0) = 10C0 (0.7)0 (0.3)10
= 1 x 1 x 0.0000059
= 0.000006

P(at most 2 passes) =P(x≤2)


= P(x=0) + P(x=1) + P(x=2)
= 0.000006 + 0.000138 + 0.001447
= 0.001591

Example 2:
If 40% of cocoa seed bought by a produce buyer are defective, find the probability that out of
5 cocoa seeds selected at random (i) at least 4 (ii) between 1 and 3 (iii) exactly 3 (iv) at most
1 are non-defective.
Solution:
Let P be the probability that a seed bought is non-defective. Then:
60 3 3 2
P = 60% = 100 = 5 ; =1-5 = 5 ; n =5

Let P(x= r) be the probability that in n trials, there are r non-defective


P(x=4) = 5C4 (35)4 (25)1
5!
= (5−4)!4! x 0.1296 x 0.4000
= 5 x 0.1296 x 0.4000
= 0.2592
= 25.92%
P(x=5) = 5C5 (35)5 (25)0
5!
= (5−5)!5! x 0.07776 x 1
= 1 x 0.07776 x 1
= 0. 0.07776
= 7.776 %

(i) P(at least 4) = P(x ≥ 4) = P(x =4 )+ P(x =5)


= 0.2592 + 0. 0.07776
= 0.33696 ≈ 0.337
= 33.7%
(ii) Between 1 and 3 implies that 1 < x < 3. Since x can only assume integer values it means
in effect that x=2
P(1 < x < 3 ) = P(x =2)
3 2
= 5C2 (5)2 (5)3
= 10 x 0.3600 x 0.0640
= 0.2304
= 23.04%
(iii) P(exactly 3) = P(x =3)
3 2
=5C3 (5)3 (5)2
= 10 x 0.02160 x 0.1600
= 0.3456
= 34.56%
(iv) P( x= at most 1 ) = P(x < 1)
= P(x =0) + P(x=1)
3 2
P(x =0) =5C0 (5)0 (5)5
= 1 x 1 x 0.01024
= 0.0102
= 1.02%
3 2
P(x =1) =5C1 (5)1 (5)4
= 5 x 0.1296 x 0.4
= 0.2592
= 25.92%
P( x = at most 1) = 0.0102 + 0.2592
= 0.2694
= 26.94%

Example 3:
The probability that a diabetic patient survives when injected with a newly discovered drug
is 0.75. Find the probability that exactly 8 out of 10 diabetic patients survives on being
injected with the new drug.
Solution:
Let P be the probability that a patient survives when injected.
Then P=0.75, q= 1-P = 1- 0.75 =0.25, n=10
Let P(x= r) be the probability that in n trials, r patient would survive
Using Binomial distribution
P(x= exactly 8) = P(x = 8 ) = 10C8 (0.75)8 (0.25)2
= 45 x 0.1001 x 0.0625
= 0.2815
= 28.15%
Example 5:
Ten aspirants contest in 10 wards of local government area for a particular post. The
1
probability that a political aspirant wins in a ward is5. Find the probability that four of the
aspirants win their elections.
Solution: Let P be the probability that an aspirant wins a ward election.
1 4
Then P= 15 , q = 1 - 5 = 5 , n=10
Using binomial distribution, P (x= four aspirants win) = P(x=4)
1 4
= 10C4 (5)4 (5)6
10! 1 4
= (10−4)!4! (5)4 (5)6
10! 1 4
= (10−4)!4! (5)4 (5)6

[Link].6! 1 4 4 6
= ( ) ( )
6!4! 5 5
[Link] 1 4 4 6
= [Link] (5) (5)
10.3.7
= x 0.0016 x 0.2621
1
= 10 x 3 x 7 x 0.0016 x 0.2621

= 210 x 0.0016 x 0.2621

= 0.0880656

≈ = 8.81%
MODULE TWO
INFERENTIAL PROCEDURES
INFERENTIAL PROCEDURES
There are two types of inferential procedures: (1) Estimation, (2) Hypothesis testing

Estimation
In estimation a sample is drawn and studied, and inference is made about the population
characteristics on the basis of what is discovered about the sample. There may be sampling
variations because of chance fluctuations, variations in sampling techniques, and other
sampling errors. We, therefore, do not expect our estimate of the population characteristics to
be exactly correct. We do, however, expect it to be close. The real question in estimation is
not whether our estimate is correct or not but how close is it to be the true value. Our first
interest is in using the sample mean (X̅) to estimate the population mean (μ).

Characteristics of X as an estimate of (μ).


The sample mean (𝑋̅) often is used to estimate a population mean (μ). For example, the sample
mean of 45.0 from the Academic Anxiety Test may be used to estimate the mean Academic
Anxiety of population of college students. Using this sample would lead to an estimate of
45.0 for the population mean. Thus, sample mean is an unbiased and consistent estimator of
population mean.

Unbiased Estimator
An unbiased estimator is one which, if we were to obtain an infinite number of random
samples of a certain size, the mean of the statistic would be equal to the parameter. The sample
mean, (X̅) is an unbiased estimate of (μ) because if we look at possible random samples of
size N from a population, mean of the sample would be equal to μ. Consistent

Estimator
A consistent estimator is one that as the sample size increases, the probability that estimate
has a value close to the parameter also increase. Because it is a consistent estimator, a sample
mean based on 20 scores has a greater probability of being closer to (μ) than does a sample
mean based upon only 5 scores. Better estimates of a population mean should be more
probable from large samples.

Accuracy of Estimation
The sample mean is an unbiased and consistent estimator of (μ) . But we should not overlook
the fact that an estimate is just a rough or approximate calculation. It is unlikely in any
estimate that (X̅) will be exactly equal to (μ). Whether or not X̅ is a good estimate of (μ)
depends upon the representativeness of sample, the sample size, and the variability of scores
in the population.

HYPOTHESIS TESTING
In addition to estimating the accuracy of parameter estimates, the sampling distribution of the
mean serves a very important function in hypothesis testing. Imagine that someone told you
that they had a magic die that was loaded to show 6 when it was thrown.
Would you simply believe them and purchase the die for N100? You would surely want to
test the die before purchasing it? Specifically, you would want to test the hypothesis that the
die is in fact loaded.
A hypothesis is a tentative statement of a relationship between two variables, or as Neuman
(1997, p. 108) puts it, hypotheses are educated ‘guesses about how the social world works’.
Hypothesis testing is a logical and empirical procedure whereby hypotheses are formally set
up and subjected to empirical test. In the first stage of hypothesis testing, the researcher states
a research question, and poses two hypotheses that refer to the possible outcomes of the
empirical investigation. The research question is the question that the researcher wants to
answer by doing the research. In our loaded die problem, we would want to test whether the
die shows 6 more often than a fair die. This would tell us whether the die was loaded or not.
The research question for this investigation would be: ‘Does the “magic” die show 6 more
often than a fair die?’ Answering this question would be the whole point of the research.

Examples of research questions


All of the following research questions can be investigated with a hypothesis testing
approach:
a) Are individuals less intelligent in crowd situations?
b) Do women and men perform similarly at facial recognition tasks?
c) Have a group of children who were involved in a bus accident suffered mental impairment?
d) Are schizophrenics violent?
You will note that all the research questions above presuppose two conditions or groups, and
a comparison between them. We are comparing a fair die with a loaded die; individuals in
crowds with the same individuals when they are not in crowds; women and men; and children
involved in a bus accident with similar children who were not involved in a bus accident. A
comparison group is also implied by the research question ‘Are schizophrenics violent?’
What we want to know here is whether schizophrenics are more violent than non-
schizophrenic people.
The research question in a hypothesis-testing situation typically seeks to determine whether
groups are the same or not. Before conducting an empirical investigation to determine this,
the research question is first translated into two hypotheses, known as the null and alternative
hypotheses. The null hypothesis is a statement that maintains that there is no difference
between the groups or conditions. It is represented by the symbol H0. From the loaded die
research question, we would derive the following null hypothesis:

H0: The loaded die shows 6 with the same probability as a fair die.

Examples of null hypotheses


1. H0: There is no difference between the intelligence of individuals when they are in crowds
and when they are not in crowds.
2. H0: Men and women perform similarly at facial recognition tasks.
3. H0: There is no difference in mental functioning between the children who were involved
in the bus accident and similar children who have not experienced trauma.
4. H0: Schizophrenic and non-schizophrenic people display similar levels of violence.

In contrast to the null hypothesis, the alternative hypothesis is a statement that maintains that
there are differences between the groups or conditions. This hypothesis makes a conjecture
that is diametrically opposed to the null hypothesis. The alternative hypothesis is represented
by the symbol H1. The alternative hypothesis can take two forms, depending on the nature of
the research question: it can be either directional or non-directional. A directional alternative
hypothesis anticipates the direction of difference. It states the researcher’s expectation
regarding whether one group is going to score higher or lower than the other group. A non-
directional hypothesis merely states that a difference is expected, without anticipating the
direction of the difference.
The ‘loaded die’ research question involves a directional alternative hypothesis because we
want to determine whether the loaded die shows heads more often than a fair die:
H1: The loaded die shows 6 more often than a fair die.

Examples of alternative hypotheses


1. H1: Individuals in crowds are less intelligent than when they are not in crowds.
2. H1: Men and women perform differently at facial recognition tasks.
3. H1: The children who were involved in the bus accident show impaired mental functioning
in comparison with similar children who have not experienced trauma.
4. H1: Schizophrenics are more violent than non-schizophrenics.
Can you identify which of the research questions in “Examples of null hypotheses” require
directional or non-directional alternative hypotheses? Hypotheses 1, 3, and 4 are all
directional, whereas 2 is non-directional (see Examples of alternative hypotheses). Can you
see why?

Thus far the null hypothesis and alternative hypothesis have been written out in words.
However, they are usually written in symbolic format. At the outset of a research project,
before engaging in any empirical testing, the researcher should state the research question and
hypotheses closely analogous to the following:

1. Research question: Are individuals less intelligent in crowd situations?


H0: μ1 = μ2
H1: μ1 < μ2

2. Research question: Do women and men perform the same at facial recognition tasks?
H0: μ1 = μ2
H1: μ1 ≠ μ2

The research question in the first example implies a directional alternative hypothesis. The
words ‘less intelligent’ in the research question indicate that a ‘less than’ sign (i.e.< ) should
be used in H1 to show the researcher’s expectation. The research question in the second
example is non-directional, and a ≠ sign is used in H1 to indicate the absence of direction.
In hypothesis testing, we are not really interested in whether or not our sample means differ.
They may differ because of random variation introduced by the sampling process (i.e. error
variance). We are interested in whether or not the population means differ, therefore the
hypotheses are stated in terms of the population parameter (μ) not the sample statistic (X̄).
The mean of the first population (e.g. individuals in crowds; women) is represented by μ1,
and the mean of the second population (e.g. individuals not in crowds; men) is represented
by μ2. Once the research question and hypotheses have been stated, the researcher may
proceed to test the hypotheses empirically. The results of the empirical investigation will
indicate whether the null hypothesis or the alternative hypothesis should be rejected.
A Type I error is made by rejecting the null hypothesis when it in fact is true.
A Type II error is made by not rejecting the null hypothesis when it is false.

Example A
We have a medicine that is being manufactured and each pill is supposed to have 14
milligrams of the active ingredient. What are our null and alternative hypotheses?
Solution
H0 : μ = 14
Ha : μ ≠ 14
Our null hypothesis states that the population has a mean equal to 14 milligrams. Our
alternative hypothesis states that the population has a mean that is different than 14
milligrams.

Example B
The school principal wants to test if it is true what teachers say – that high school juniors use
the computer an average 3.2 hours a day. What are our null and alternative hypotheses?
Solution
H0 : μ = 3:2
Ha : μ ≠ 3:2
Our null hypothesis states that the population has a mean equal to 3.2 hours. Our alternative
hypothesis states that the population has a mean that differs from 3.2 hours.

Deciding Whether to Reject the Null Hypothesis: One and Two-Tailed Hypothesis Tests
The alternative hypothesis can be supported only by rejecting the null hypothesis. To reject
the null hypothesis means to find a large enough difference between your sample mean and
the hypothesized (null) mean that it raises real doubt that the true population mean is 20. If
the difference between the hypothesized mean and the sample mean is very large, we reject
the null hypothesis. If the difference is very small, we do not. In each hypothesis test, we
have to decide in advance what the magnitude of that difference must be to allow us to reject
the null hypothesis.
Below is an overview of this process. Notice that if we fail to find a large enough difference
to reject, we fail to reject the null hypothesis. Those are your only two alternatives.

When a hypothesis is tested, a statistician must decide on how much of a difference between
means is necessary in order to reject the null hypothesis.
Statisticians first choose a level of significance or alpha (a) level for their hypothesis test.
Similar, to the significance level you used in constructing confidence intervals, this alpha
level tells us how improbable a sample mean must be for it to be deemed "significantly
different" from the hypothesized mean. The most frequently used levels of significance are
0:05 and 0:01: An alpha level of 0.05 means that we will consider our sample mean to be
significantly different from the hypothesized mean if the chances of observing that sample
mean are less than 5%. Similarly, an alpha level of 0.01 means that we will consider our
sample mean to be significantly different from the hypothesized mean if the chances of
observing that sample mean are less than 1%.

Type I and Type II Errors


Remember that there will be some sample means that are extremes – that is going to happen
about 5% of the time, since 95% of all sample means fall within about two standard deviations
of the mean. What happens if we run a hypothesis test and we get an extreme sample mean?
It won’t look like our hypothesized mean, even if it comes from that distribution. We would
be likely to reject the null hypothesis. But we would be wrong.
When we decide to reject or not reject the null hypothesis, we have four possible scenarios:
a. A true hypothesis is rejected.
b. A true hypothesis is not rejected.
c. A false hypothesis is not rejected.
d. A false hypothesis is rejected.
If a hypothesis is true and we do not reject it (Option 2) or if a false hypothesis is rejected
(Option 4), we have made the correct decision. But if we reject a true hypothesis (Option 1)
or a false hypothesis is not rejected (Option 3) we have made an error. Overall, one type of
error is not necessarily more serious than the other. Which type is more serious depends on
the specific research situation, but ideally both types of errors should be minimized during
the analysis.

TABLE 12.1: The Four Possible Outcomes in Hypothesis Testing

CHI SQUARE DISTRIBUTION


The square of a standard normal variable is called a Chi-square variate with 1 degree of
freedom, abbreviated as d.f. Thus, if x is a random variable following normal distribution with
mean μ and standard deviation σ, then (X- μ)/σ is a standard normal variate.
𝑋̅− 𝜇 2
Therefore, Z =( ) is a chi-square (abbreviated by the letter χ2 of the Greek alphabet)
𝜎
variate with 1 d.f.
If X1, X2, X3, ...........................Xv are v independent random variables following normal
distribution with means μ1, μ2, μ3,................... μv, and standard deviations σ1, σ2, σ3,..... σv
respectively then the variate

𝑥1−𝜇1 2 𝑥2−𝜇2 2 𝑥3− 𝜇3 2 𝑥𝑣− 𝜇𝑣 2


X2 = ( ) + ( ) +( ) + …… … . ( )
𝜎1 𝜎2 𝜎3 𝜎𝑣
𝑋̅1−𝜇1 2
= ∑𝑣𝑖=1 ( )
𝜎1

which is the sum of the squares of v independent standard normal variates, follow Chi-square
distribution with v d.f.
Applications of the χ2-Distribution
Chi-square distribution has a number of applications, some of which are enumerated below:
(i) Chi-square test of goodness of fit.
(ii) χ2-test for independence of attributes
(iii) To test if the population has a specified value of variance σ2.
(iv) To test the equality of several population proportions

Observed and Theoretical Frequencies


Suppose that in a particular sample a set of possible events E1, E2, E3,..................Ek are
observed to occur with frequencies O1, O2, O3, ..........Ok, called observed frequencies, and
that according to probability rules they are expected to occur with frequencies e1, e2,
e3,.....ek, called expected or theoretical frequencies. Often, we wish to know whether the
observed frequencies differ significantly from expected frequencies.
Definition of χ2
A measure of discrepancy existing between the observed and expected frequencies is supplied
by the statistics χ2 given by

Chi-Square test of goodness of fit


A Chi-Square Test is used to examine whether the observed results are in order with the
expected values. When the data to be analysed is from a random sample, and when the
variable is the question is a categorical variable, then Chi-Square proves the most appropriate
test for the same. A categorical variable consists of selections such as breeds of dogs, types
of cars, genres of movies, educational attainment, male v/s female etc. Survey responses and
questionnaires are the primary sources of these types of data. The Chi-square test is most
commonly used for analysing this kind of data. This type of analysis is helpful for researchers
who are studying survey response data. The research can range from customer and marketing
research to political sciences and economics.

The chi-square test can be used to determine how well theoretical distributions such as the
normal and binomial distributions) fit empirical distributions (i.e. those obtained from sample
data). Suppose we are given a set of observed frequencies obtained under some experiment
and we want to test if the experimental results support a particular hypothesis or theory. Karl
Pearson in 1900, developed a test for testing the significance of the discrepancy between
experimental values and the theoretical values obtained under some theory or hypothesis.
This test is known as χ2-test of goodness of fit and is used to test if the deviation between
observation (experiment) and theory may be attributed to chance (fluctuations of sampling)
or if it is really due to the inadequacy of the theory to fit the observed data.
Under the null hypothesis that there is no significant difference between the observed
(experimental and the theoretical or hypothetical values i.e. there is good compatibility
between theory and experiment.
Karl Pearson proved that the statistic.
Follows χ2-distribution with v = n-1, d.f. where O1, O2,..................On are the observed
frequencies and E1, E2,..................En are the corresponding expected or theoretical
frequencies obtained under some theory or hypothesis.

Chi-square distributions (X2) are a type of continuous probability distribution. They're


commonly utilized in hypothesis testing, such as the chi-square goodness of fit and
independence tests. The parameter k, which represents the degrees of freedom, determines
the shape of a chi-square distribution.

A chi-square distribution is followed by very few real-world observations. The objective of


chi-square distributions is to test hypotheses, not to describe real-world distributions. In
contrast, most other commonly used distributions, such as normal and Poisson distributions,
may explain important things like baby birth weights or illness cases per year.
Because of its close resemblance to the conventional normal distribution, chi-square
distributions are excellent for hypothesis testing. Many essential statistical tests rely on the
conventional normal distribution.

In statistical analysis, the Chi-Square distribution is used in many hypothesis tests and is
determined by the parameter k degree of freedoms. It belongs to the family of continuous
probability distributions. The Sum of the squares of the k independent standard random
variables is called the Chi-Squared distribution. Pearson’s Chi-Square Test formula is

Where X2 is the Chi-Square test symbol


Σ is the summation of observations
O is the observed results
E is the expected results
The shape of the distribution graph changes with the increase in the value of k, i.e. degree of
freedoms.
When k is 1 or 2, the Chi-square distribution curve is shaped like a backwards ‘J’. It means
there is a high chance that X2 becomes close to zero.

Courtesy: Scribbr
When k is greater than 2, the shape of the distribution curve looks like a hump and has a low
probability that X2 is very near to 0 or very far from 0. The distribution occurs much longer
on the right-hand side and shorter on the left-hand side. The probable value of X2 is (X2 - 2).

Courtesy: Scribbr

When k is greater than ninety, a normal distribution is seen, approximating the Chi-square
distribution.
Chi-Square P-Values
Here P denotes the probability; hence for the calculation of p-values, the Chi-Square test
comes into the picture. The different p-values indicate different types of hypothesis
interpretations.
1. P <= 0.05 (Hypothesis interpretations are rejected)
2. P>= 0.05 (Hypothesis interpretations are accepted)

The concepts of probability and statistics are entangled with Chi-Square Test. Probability is
the estimation of something that is most likely to happen. Simply put, it is the possibility of
an event or outcome of the sample. Probability can understandably represent bulky or
complicated data. And statistics involves collecting and organising, analysing, interpreting
and presenting the data.

Finding P-Value
When you run all of the Chi-square tests, you'll get a test statistic called X2. You have two
options for determining whether this test statistic is statistically significant at some alpha
level:
1. Compare the test statistic X2 to a critical value from the Chi-square distribution table.
2. Compare the p-value of the test statistic X2 to a chosen alpha level.

Test statistics are calculated by taking into account the sampling distribution of the test
statistic under the null hypothesis, the sample data, and the approach which is chosen for
performing the test.
The p-value will be as mentioned in the following cases.
• A lower-tailed test is specified by: P(TS ts | H0 is true) p-value = cdf (ts)
• Lower-tailed tests have the following definition: P(TS ts | H0 is true) p-value = cdf (ts)
• A two-sided test is defined as follows, if we assume that the test static distribution of H0 is
symmetric about 0. 2 * P(TS |ts| | H0 is true) = 2 * (1 - cdf(|ts|))

Where:
P: probability Event
TS: Test statistic is computed observed value of the test statistic from your sample cdf():
Cumulative distribution function of the test statistic's distribution (TS)
Types of Chi-square Tests
Pearson's chi-square tests are classified into two types:
1. Chi-square goodness-of-fit analysis
2. Chi-square independence test

These are, mathematically, the same exam. However, because they are utilized for distinct
goals, we generally conceive of them as separate tests.
Properties
The chi-square test has the following significant properties:
1. If you multiply the number of degrees of freedom by two, you will receive an answer that
is equal to the variance.
2. The chi-square distribution curve approaches the data is normally distributed as the degree
of freedom increases.
3. The mean distribution is equal to the number of degrees of freedom.

Properties of Chi-Square Test


1. Variance is double the times the number of degrees of freedom.
2. Mean distribution is equal to the number of degrees of freedom.
3. When the degree of freedom increases, the Chi-Square distribution curve becomes normal.
Limitations of Chi-Square Test
There are two limitations to using the chi-square test that you should be aware of.
• The chi-square test, for starters, is extremely sensitive to sample size. Even insignificant
relationships can appear statistically significant when a large enough sample is used. Keep in
mind that "statistically significant" does not always imply "meaningful" when using the chi-
square test.
• Be mindful that the chi-square can only determine whether two variables are related. It does
not necessarily follow that one variable has a causal relationship with the other. It would
require a more detailed analysis to establish causality.

Steps for computing χ2 and drawing conclusions


(i) Compute the expected frequencies E1, E2, .....................En corresponding to the observed
frequencies O1, O2, ...................On under some theory or hypothesis.
(ii) Compute the deviations (O-E) for each frequency and then square them to obtain (O-E)2.
(iii) Divide the square of the deviations (O-E)2 by the corresponding expected frequency to
obtain (O-E)2/E.
(𝑂𝑖−𝐸𝑖)2
(iv) Add values obtained in step (iii) to compute χ2 = ∑𝑛𝑖=1 𝐸
(v) Under the null hypothesis that the theory fits the data well, the statistic follows χ2-
distribution with v = n-1 d.f.
(vi) Look for the tabulated (critical) values of χ2 for (n-1) d.f. at certain level of significance,
usually 5% or 1%, from any Chi-square distribution table.
If calculated value of χ2 obtained in step (iv) is less than the corresponding tabulated value
obtained in step (vi), then it is said to be non-significant at the required level of significance.
This implies that the discrepancy between observed values (experiment) and the expected
values (theory) may be attributed to chance, i.e. fluctuations of sampling. In other words, data
do not provide us any evidence against the null hypothesis [given in step (v)] which may,
therefore, be accepted at the required level of significance and we may conclude that there is
good correspondence (fit) between theory and experiment.

(vii) On the other hand, if calculated value of χ2 is greater than the tabulated value, it is said
to be significant. In other words, discrepancy between observed and expected frequencies
cannot be attributed to chance and we reject the null hypothesis. Thus, we conclude that the
experiment does not support the theory.

Example 1:A pair of dice is rolled 500 times with the sums in the table below
Sum(x) Observed frequency
2 15
3 35
4 49
5 58
6 65
7 76
8 72
9 60
10 35
11 29
12 6

Take α = 5%
It should be noted that the expected sums if the dice are fair, are determined from the
distribution of x as in the table below:
Sum(x) P(x)
2 1/36
3 2/36
4 3/36
5 4/36
6 5/36
7 6/36
8 5/36
9 4/36
10 3/36
11 2/36
12 1/36

To obtain the expected frequencies, the P(x) is multiplied by the total number of trials.
Sum(x) Observed Frequency P(x) Expected Frequency
(O) (P(x).500)
2 15 1/36 13.9
3 35 2/36 27.8
4 49 3/36 41.7
5 58 4/36 55.6
6 65 5/36 69.5
7 76 6/36 83.4
8 72 5/36 69.5
9 60 4/36 55.6
10 35 3/36 41.7
11 29 2/36 27.8
12 6 1/36 13.9

Recall that χi2 = (Oi – Ei)2/Ei


Therefore χ12 = (O1 – E1)2/E1 = (15 – 13.9)2/13.9 = 0.09
χ22 = (O2 – E2)2/E2 = (35 – 27.8)2/27.8 = 1.86
χ32 = (O3 – E3)2/E3 = (49 – 41.7)2/41.7 = 1.28
χ42 = (O4 – E4)2/E4 = (58 – 55.6)2/55.6 = 0.10
χ52 = (O5 – E5)2/E5 = (65 – 69.5)2/69.5 = 0.29
χ62 = (O6 – E6)2/E6 = (76 – 83.4)2/83.4 = 0.66
χ72 = (O7 – E7)2/E7 = (72 – 69.5)2/69.5 = 0.09
χ82 = (O8 – E8)2/E8 = (60 – 55.6)2/55.6 = 0.35
χ92 = (O9 – E9)2/E9 = (35 – 41.7)2/41.7 = 1.08
χ102 = (O10 – E10)2/E10 = (29 – 27.8)2/27.8 = 0.05
χ112 = (O11 – E11)2/E11 = (6 – 13.9)2/13.9 = 4.49
(𝑂𝑖−𝐸𝑖)2
To calculate the overall Chi-squared value, recall that χ2 = ∑𝑛𝑖=1 i.e. we add the
𝐸
individual χ2 value.
Therefore, χ2 = 0.09 + 1.86 + 1.28+ 0.10 + 0.29 + 0.66 + 0.09 + 0.35 + 1.08 + 0.05 + 4.49
χ2 = 10.34
For the critical value, since n=11, d.f. = 10
Therefore, table value = 18.3
Decision: since the calculated value which is 10.34 is less than table (critical) value, the null
hypothesis is accepted.
Conclusion: There is no significant difference between observed and expected frequencies.
The slight observed differences occurred due to chance.
Example 2
Over a period of 2 years a psychiatrist has classified by socioeconomic class the women aged
20-64 admitted to her unit suffering from self-poisoning sample A. At the same time, she has
likewise classified the women of similar age admitted to a gastroenterological unit in the same
hospital sample B. She has employed the Registrar General’s five socioeconomic classes, and
generally classified the women by reference to their father’s or husband’s occupation. The
results are set out in table 8.1.
Table 8.1 Distribution by socioeconomic class of patients admitted to
self-poisoning (sample A) and gastroenterological (sample B) units
Socio economic Samples Total Proportion in group A
class A B
a b n=a+b P=a/n
I 17 5 22 0.77
II 25 21 46 0.54
III 39 34 73 0.53
IV 42 49 91 0.46
V 32 25 57 0.56
Total 155 134 289

The psychiatrist wants to investigate whether the distribution of the patients by social class
differed in these two units.
She therefore erects the null hypothesis that there is no difference between the two
distributions. This is what is tested by the chi squared (χ²) test. By default, all χ² tests are
two sided.
It is important to emphasise here that χ² tests may be carried out for this purpose only on the
actual numbers of occurrences, not on percentages, proportions, means of observations, or
other derived statistics. Note, we distinguish here the Greek (χ²) for the test and the
distribution and the Roman (x²) for the calculated statistic, which is what is obtained from the
test.
Ensure you complete the exercise and submit as part of your C.A.
Example3
Let's say you want to know if gender has anything to do with political party preference. You
poll 440 voters in a simple random sample to find out which political party they prefer. The
results of the survey are shown in the table below:
Republican Democrat Independent Total
Male 100 70 30 200
Female 140 60 40 240
Total 240 130 70 440

To see if gender is linked to political party preference, we perform a Chi-Square test of


independence using the steps below.
Step 1: Define the Hypothesis
H0: There is no link between gender and political party preference.
H1: There is a link between gender and political party preference.
Step 2: Calculate the Expected Values
Now you will calculate the expected frequency.
(𝑅𝑜𝑤 𝑇𝑜𝑡𝑎𝑙)∗(𝐶𝑜𝑙𝑢𝑚𝑛 𝑇𝑜𝑡𝑎𝑙)
Expected value = 𝑇𝑜𝑡𝑎𝑙 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑂𝑏𝑠𝑒𝑟𝑣𝑎𝑡𝑖𝑜𝑛𝑠

For example, the expected value for Male Republicans is:


(240)∗(200)
= = 109
440

Similarly, you can calculate the expected value for each of the cells.
Expected values.

Republican Democrat Independent Total


Male 109 59 22.72 200
Female 120 65 25 240
Total 240 130 70 440

Step 3: Calculate (O-E)2 / E for Each Cell in the Table


Now you will calculate the (O - E)2 / E for each cell in the table.
Where, O = Observed Value E = Expected Value
(𝑂 − 𝐸)2
𝐸
Republican Democrat Independent Total
Male 0.74311927 2.050847 2.332676056 200
Female 3.33333333 0.384615 1 240
Total 240 130 70 440

Step 4: Calculate the Test Statistic X2


X2 is the sum of all the values in the last table
= 0.743 + 2.05 + 2.33 + 3.33 + 0.384 + 1
X2 calculated= 9.837
Before you can conclude, you must first determine the critical statistic, which requires
determining our degrees of freedom. The degrees of freedom in this case are equal to the
table's number of columns minus one multiplied by the table's number of rows minus one, or
(r-1) (c-1). We have (3-1)(2-1) = 2.
Finally, you compare our obtained statistic to the critical statistic found in the chi-square
table. As you can see, for an alpha level of 0.05 and two degrees of freedom, the critical
statistic is 5.991(X2table from the statistical table), which is less than our obtained statistic of
9.83. You can reject our null hypothesis because the critical statistic is higher than your
obtained statistic.
This means you have sufficient evidence to say that there is an association between gender
and political party preference.
Below is a sample chi-square table (it is also available in a mathematical /statistical
table) Critical values of the Chi-Square distribution with d degrees of freedom
Probability of exceeding the critical value
d 0.05 0.01 0.001 d 0.05 0.01 0.001
1 3.841 6.635 10.828 11 19.675 24.725 31.264
2 5.991 9.210 13.816 12 21.026 26.217 32.910
3 7.815 11.345 16.266 13 22.362 27.688 34.528
4 9.488 13.277 18.467 14 23.685 29.141 36.123
5 11.070 15.086 20.515 15 24.996 30.578 37.697
6 12.592 16.812 22.458 16 26.286 30.578 37.697
7 14.067 18.475 24.322 17 27.587 33.409 40.790
8 15.507 20.090 26.125 18 28.869 34.805 42.312
9 16.919 21.666 27.877 19 30.144 36.191 43.820
10 18.307 23.209 29.588 20 31.410 37.566 45.315

The student’s t-test


𝒙̅−µ
Z= 𝝈 ≈ N(0,1), asymptotically
√𝒏
If the population variance is unknown then for the large samples, its estimates provided by
sample variance S2 is used, and normal test is applied. For small samples an unbiased estimate
of population variance σ2 is given by:
It is quite conventional to replace σ2 byS2 (for small samples) and then apply the normal test
even for small samples. W.S Goset, who wrote under the pen name of Student, obtained the
𝒙̅−µ
sampling distribution 𝑺 of the statistic for small samples and showed that it is far from
√𝒏
normality.

This discovery started a new field, viz ‘Exact Sample Test’ in the history of statistical
inference.
Note: If x1, x2...............xn is a random sample of size n from a normal population with mean
μ and variance σ2 then the Student’s t statistic is defined as:

𝛴𝑥
Where 𝑋 = 𝑛is the sample mean and is an unbiased estimate of the population variance
𝑛
σ2 .
Applications of t-distribution
(i) t-test for the significance of single mean, population variance being unknown
(ii) t-test for the significance of the difference between two sample means, the population
variances being equal but unknown
(iii) t-test for the significance of an observed sample correlation coefficient.
Test for Single Mean
Sometimes, we may be interested in testing if:
(i) The given normal population has a specified value of the population mean, say μo.
(ii) The sample mean differs significantly from specified value of population mean.
(iii) A given random sample x1, x2...............xn of size n has been drawn from a normal
population with specified mean μo.
Basically, all the three problems are the same. We set up the corresponding null hypothesis
thus:
(a) Ho: μ = μo i.e. the population mean is μo
(b) Ho: There is no significant difference between the sample mean and the population mean.
In order words, the difference between and μ is due to fluctuations of sampling.
(c) Ho: The given random sample has been drawn from the normal population with mean μo.
Under Ho the test-statistic is:
And it follows Student’s t-distribution with (n-1) degrees of freedom.
We compute the test-statistic using the formula above under Ho and compare it with the
tabulated value of t for (n-1) [Link] the given level of significance. If the absolute value of the
calculated t is greater than tabulated t, we say it is significant and the null hypothesis is
rejected. But if the calculated t is less than tabulated t, Ho may be accepted at the level of
significance adopted.

Assumptions for Student’s test


(i) The parent population from which the sample is drawn is normal
(ii) The sample observations are independent i.e. the given sample is random.
(iii) The population standard deviation σ is unknown
Example: Ten cartons are taken at random from an automatic filling machine. The mean net
weight of the 10 cartons is 11.8kg and standard deviation is 0.15kg. Does the sample mean
differ significantly from the intended weight of 12kg, α=0.05
Hint: You are given that for d.f. =9, t0.05 = 2.26
Solution: n= 10, 𝑋̅ = 11.8kg, s = 0.15kg
Null hypothesis, Ho: μ = 12 kg (i.e. the sample mean of 𝑋̅=11.8 kg does not differ significantly
from the population mean μ = 12 kg
Alternative Hypothesis H1. Ho: ≠ 12kg (Two tailed)

The tabulated value of t for 9 d.f. at 5% level of significance is 2.26. Since the calculated t is
much greater than the tabulated t, it is highly significant. Hence, null hypothesis is rejected at
5% level of significance, and we conclude that the sample mean differ significantly.

t-Test for difference of means


Assume we are interested in testing if two independent samples have been drawn from two
normal populations having the same means, the population variances being equal.
Let x1, x2...............xn and y1, y2...............yn be two independent random samples from the given
normal populations.
Ho: μx = μy i.e the two samples have been drawn from the normal populations with the same
means. Under the hypothesis that the σ12 = σ22 = σ2 i.e. population variances are equal but
unknown, the test statistic under Ho is:

This is an unbiased estimate of the common population variance σ2 based on both the samples.
By comparing the computed value of t with the tabulated value of t for n1 + n2 -2 d.f. and at
desired level of significance, usually 5% or 1%, we reject the null hypothesis.
Example: The nicotine content in milligram of two samples of tobacco were found to be as
follows:
Sample A: 24 27 26 21 25
Sample B: 27 30 28 31 22 36
Can it be said that the two samples come from the same normal population having the same
mean?
Solution Hints: Applying the above formula and calculating the variance as appropriate, the
calculated t-value is -1.92. the tabulated value for 9 d.f. at 5% level of significance for two-
tailed test is 2.262. Since calculated t is less than the tabulated t, it is not significant, and the
null hypothesis is accepted.

T-test has very wide applications. It can be applied in the tests of single mean, in the
comparison of two different means and in the test of significance of other parameter estimates.

ANALYSIS OF VARIANCE (ANOVA)


In day-to-day business management and in sciences, instances may arise where we need to
compare means. If there are only two means e.g., average recharge card expenditure between
male and female students in a faculty of a university, the typical t-test for the difference of
two means becomes handy to solve this type of problem. However, in real life situation man
is always confronted with situation where we need to compare more than two means at the
same time. The typical t-test for the difference of two means is not capable of handling this
type of problem; otherwise, the obvious method is to compare two means at a time by using
the t-test earlier treated. This process is very time consuming, since as few as 4 sample means
would require 4C2 = 6, different tests to compare 6 possible pairs of sample means. Therefore,
there must be a procedure that can compare all means simultaneously. One such procedure is
the analysis of variance (ANOVA). For instance, we may be interested in the mean telephone
recharge expenditures of various groups of students in the university such as student in the
faculty of Science, Arts, Social Sciences, Medicine, and Engineering. We may be interested
in testing if the average monthly expenditure of students in the five faculties are equal or not
or whether they are drawn from the same normal population. The answer to this problem is
provided by the technique of analysis of variance. It should be noted that the basic purpose
of the analysis of variance is to test the homogeneity of several means.

The term Analysis of Variance was introduced by Prof. R.A Fisher in 1920s to deal with
problem s in the analysis of agronomical data. Variation is inherent in nature. The total
variation in any set of numerical data is due to a number of causes which may be classified
as:
(i) Assignable causes and (ii) chance causes
The variation due to assignable causes can be detected and measured whereas the variation
due to chances is beyond the control of human and cannot be traced separately.

Assumption for ANOVA test


ANOVA test is based on the test statistic F (or variance ratio). For the validity of the F-test
in ANOVA, the following assumptions are made:
(i) The observations are independent.
(ii) Parent population from which observation are taken are normal.
(iii) Various treatment and environmental effects are additive in nature.
ANOVA as a tool has different dimensions and complexities. ANOVA can be (a) One-way
classification or (b) two-way classification. However, the one-way ANOVA we will deal with
in this course material.
Note
(i) ANOVA technique enables us to compare several populations means simultaneously and
thus results in lot of saving in terms of time and money as compared to several experiments
required for comparing two populations means at a time.
(ii) The origin of the ANOVA technique lies in agricultural experiments and as such its
language is loaded with such terms as treatments, blocks, plots etc. However, ANOVA
technique is so versatile that it finds applications in almost all types of design of experiments
in various diverse fields such as industry, education, psychology, business, economics etc.
(iii) It should be clearly understood that ANOVA technique is not designed to test equality of
several population variances. Rather, its objective is to test the equality of several population
means or the homogeneity of several independent sample means.
(iv) In addition to testing the homogeneity of several sample means, the ANOVA technique
is now frequently applied in testing the linearity of the fitted regression line or the significance
of the correlation ratio.
The one-way classification
Assuming n sample observations of random variable X are divided into k classes on the basis
of some criterion or factor of classification. Let the ith class consist of ni observations and
let:
Xij = jth member of the ith class; {j=1,2,......ni; i= 1,2, ........k}
n = n1 +n2 +...........................+ nk =
The n sample observations can be expressed as in the table below:
Such scheme of classification according to a single criterion is called one-way classification
and its analysis of variance is known as one-way analysis of variance.
The total variation in the observations Xij can be split into the following two components:
(i) The variation between the classes or the variation due to different bases of classification
(commonly known as treatments in pure sciences, medicine and agriculture). This type of
variation is due to assignable causes which can be detected and controlled by human
endeavour.
(ii) The variation within the classes, i.e. the inherent variation of the random variable within
the observations of a class. This type of variation is due to chance causes which are beyond
the control of man.
The main objective of the analysis of variance technique is to examine if there is significant
difference between the class means in view of the inherent variability within the separate
classes.
Steps for testing hypothesis for more than two means (ANOVA): Here, we adopt the
rejection region method, and the steps are as follows:
Step1: Set up the hypothesis:
Null Hypothesis: Ho: μ1 = μ2 = μ3 = ..............= μk i.e, all means are equal
Alternative hypothesis: H1 : At least two means are different.
Step 2: Compute the means and standard deviations for each of the by the formular:

Also, compute the mean 𝑋̅of all the data observations in the k-classes by the formula:

Step 3: Obtain the Between Classes Sum of Squares (BSS) by the formula:

Step 4: Obtain the Between Classes Mean Sum of Squares (MBSS)


𝐵𝑒𝑡𝑤𝑒𝑒𝑛 𝑐𝑎𝑠𝑠𝑒𝑠𝑠 𝑆𝑢𝑚 𝑜𝑓 𝑆𝑞𝑢𝑎𝑟𝑒 𝐵𝑆𝑆
MBSS = = 𝐾−1
𝐷𝑒𝑔𝑟𝑒𝑒𝑠 𝑜𝑓 𝑓𝑟𝑒𝑒𝑑𝑜𝑚

Step 5: Obtain the Within Classes Sum of Squares (WSS) by the formula:

Step 6: Obtain the Within Classes Mean Sum of Squares (MWSS)

Step 7: Obtain the test statistic F or Variance Ratio (V.R)

Which follows F-distribution with (v1 = k-1, v2 = n-k) d.f (This implies that the degrees of
freedom are two in number. The first one is the number of classes (treatment) less one, while
the second d.f is number of observations less number of classes)

Step 8: Find the critical value of the test statistic F for the degree of freedom and at desired
level of significance in any standard statistical table.
If computed value of test-statistic F is greater than the critical (tabulated) value, reject (Ho,
otherwise Ho may be regarded as true.
Step 9: Write the conclusion in simple language.

Example 1: To test the hypothesis that the average number of days a patient is kept in the
three local hospitals A, B and C is the same, a random check on the number of days that seven
patients stayed in each hospital reveals the following:
Hospital A 8 5 9 2 7 8 2
Hospital B 4 3 8 7 7 1 5
Hospital C 1 4 9 8 7 2 3

Solution: Let X1j, X2j, X3j denote the number of days the jth patient stays in the hospitals A,
B and C respectively
Calculations for various Sum of Squares
Within Sample Sum of Square: To find the variation within the sample, we compute the
sum of the square of the deviations of the observations in each sample from the mean values
of the respective samples (see the table above)
Sum of Squares within Samples =

= 50.8572 + 38 + 58.8572 = 147.7144 ~ 147.71

To obtain the variation between samples, we compute the sum of the squares of the deviations
of the various sample means from the overall (grand) mean.

Sum of square Between Samples (hospitals):

= 7(0.3844) + 7(0.0576) + 7(0.1444)


= 2.6908 + 0.4032 + 1.0108 = 4.1048 = 4.10

The total variation in the sample data is obtained on calculating the sum of the squares of the
deviations of each sample observation from the grand mean, for all the samples as in the table
below:
= 53.5232 + 38.4032 + 59.8832 = 151.81
Note: Sum of Squares Within Samples + S.S Between Samples = 147.71 + 4.10 =151.81
= Total Sum of Squares

Ordinarily, there is no need to find the sum of squares within the samples (i.e, the error sum
of squares), the calculations of which are quite tedious and time consuming. In practice, we
find the total sum of squares and between samples sum of squares which are relatively simple
to calculate. Finally, within samples sum of squares is obtained by subtracting Between
Samples Sum of Squares from the Total Sum of Squares:

W.S.S.S = T.S.S – B.S.S.S

Therefore, Within Sample (Error) Sum of Square = 151.8096 – 4.1048 = 147.7044

Degrees of freedom for:


Between classes (hospitals) Sum of Squares = k-1 = 3-1=2
Total Sum of Squares = n-1 = 21-1 = 20
Within Classes (or Error) Sum of Squares = n-k = 21 – 3= 18

ANOVA TABLE
Critical Value: The tabulated (critical) value of F for d.f (v1=2, v2=18) d.f at 5% level of
significance is 3.55
Since the calculated F = 0.25 is less than the critical value 3.55, it is not significant. Hence,
we fail to accept Ho.
However, in cases like this when MSS between classes is less than the MSS within classes,
we need not calculate F and we may conclude that the means,𝑋̅1, 𝑋̅2 ,𝑋̅3 and do not differ
significantly. Hence, Ho may be regarded as true.
Conclusion: Ho : μ1 = μ2 = μ3, may be regarded as true and we may conclude that there is no
significant difference in the average stay at each of the three hospitals.

Critical Difference: If the classes (called treatments in pure sciences) show significant effect
then we would be interested to find out which pair(s) of treatment differ significantly. Instead
of calculating Student’s t for different pairs of classes (treatments) means, we calculate the
Least Significant Difference (LSD) at the given level of significance. This LSD is also known
as Critical Difference (CD).
The LSD between any two classes (treatments) means, say 𝑋̅i and 𝑋̅j at level of significance
‘α’ is given by:

LSD (𝑋̅i - 𝑋̅j) = [The critical value of t at level of significance α and error d.f] X [S.E (𝑋̅i -
𝑋̅j)]
Note: S.E means Standard Error. Therefore, the S.E ( -) above mean the standard error of the
difference between the two means being considered.

MSSE means sum of squares due to Error.


If the difference |𝑋̅i - 𝑋̅j| between any two classes (treatments) means is greater than the LSD
or CD, it is said to be significant.
Note: There is another Method for the computation of various sums of squares.
LIKELY EXAMINATIONS QUESTIONS
1. What is the major difference between the crude mode and the interpolated mode?
2. If the first quartile is 104 and quartile deviation is 18. Find the third quartile.
3. 1. What do we mean by the terms ―mutually exclusive events and―independent
events?
4. State, with simple examples, the four laws of probability.
5. A firm has tendered for two independent contracts. It estimates that it has probability
0.4 of obtaining contract A and probability 0.1 of obtaining contract B.
6. Find the probability that the firm:
(a) obtains both contracts
(b) obtains neither of the contracts
(c) obtains exactly one contract
7. The mean weekly sale of the chocolate bar in candy stores was 146.3 bars per store.
After advertising campaign, the mean weekly sales in 22 stores for typical week
increased to 153.7 and showed a standard deviation of 17.2. Was the advertising
campaign successful?
8. Prices of shares of a company on the different days in a month were found to be: 66,
65, 69, 70, 69, 71, 70, 63, 64 and 68. Discuss whether the mean price of the price of
the shares in the month is 65.
9. he probability that a diabetic patient survives when injected with a newly discovered
drug is 0.72, find the probability that exactly 8 out of 10 diabetic patients survive on
being injected with the new drug.
10. If the birth of a male child and that of a female child are equiprobable, find the
probability that a family of five children, exactly 3 will be males.
11. Two salesmen A and B are working in certain district. From a Sample Survey
Conducted by the Head Office the following results were obtained. State whether
there is any significant difference in the average sales between the two salesmen.

A B
No of Sales 20 18
Average sales in (N ‘000) 170 205
Average sales in (N ‘000 ) 20 25

[Link] following figures show the distribution of digits in numbers chosen at random from a
telephone directory:
Digit 0 1 2 3 4 5 6 7 8 9 Total
Freq 1026 1107 997 966 1,075 933 1,107 972 964 853 10000
Test whether the digits may be taken to occur equally frequently in the directory. The table
value of χ2 for d.f at 5% level of significance is 16.92.
Hint: Set up the null hypothesis that the digits 0, 1, 2, 3, ..........9 in the numbers in the
telephone directory are uniformly distributed, i.e. all digits occur equally frequently in the
directory. Then, under the null hypothesis, the expected frequency for each of the digits 0, 1,
2, 3,.............9 is 10,000/10 = 1,000

13. The table below gives the retail prices of a commodity in some shops selected at random
in four cities of Lagos, Calabar, Kano and Abuja. Carry out the Analysis of Variance
(ANOVA) to test the significance of the differences between the mean prices of the
commodity in the four cities.
City Price per unit of the commodity in different shops
Lagos 9 7 10 8
Calabar 5 4 5 6
Kano 10 8 9 9
Abuja 7 8 9 8

If significant difference is established, calculate the Least Significant Difference (LSD) and
use it to compare all the possible combinations of two means (α=0.05).

14. Concord Bus Company just bought four different Brands of tyres and wishes to determine
if the average lives of the brands of tyres are the same or otherwise in order to make an
important management decision. The Company uses all the brands of tyres on randomly
selected buses. The table below shows the lives (in ‘000Km) of the tyres:
Brand 1: 10, 12, 9, 9
Brand 2: 9, 8, 11, 8, 10
Brand 3: 11, 10, 10, 8, 7
Brand 4: 8, 9, 13, 9
Test the hypothesis that the average life for each of brand of tyres is the same. Take α = 0.01

You might also like