Probability Theory Fundamentals
Probability Theory Fundamentals
𝑛(𝐸)
= 𝑛(𝑈) where n (E) = the number of outcomes in event set E and n (U) = total possible number
of outcomes in outcome set, U. If, for example, an ordinary six-sided die is to be rolled, the
equally likely outcome set, U, is {1,2,3,4,5,6} and the event ―even number has event set
{2,4,6}. It follows that the theoretical probability of obtaining an even number can be
calculated as:
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑒𝑣𝑒𝑛 𝑛𝑢𝑚𝑏𝑒𝑟𝑠 3
Pr (even numbers) = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑝𝑜𝑠𝑠𝑖𝑏𝑙𝑒 𝑑𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑡 𝑜𝑢𝑡𝑐𝑜𝑚𝑒 𝑛(𝑈) = 6 = 0.5
Other Examples
A wholesaler stocks heavy (2B), medium (HB), fine (2H) and extra fine (3H) pencils which
come in packs of 10. Currently in stock are 2 packs of 3H, 14 packs of 2H, 35 packs of HB
and 8 packs of 2B. If a pack of pencil is chosen at random for inspection, what is the
probability that they are: (a) medium (b) heavy (c ) not very fine (d) neither heavy nor
medium?
Solutions
Since the pencil pack is chosen at random, each separate pack of pencils can be regarded as
a single equally likely outcome. The total number of outcomes is the number of pencil packs,
that is, 2+14+35+8 = 59.
Thus, n (U) = 59
(c) Pr (not very fine). Note that the number of pencil packs that are not very fine is 14+35+8
= 57.
𝑛(𝑛𝑜𝑡 𝑣𝑒𝑟𝑦 𝑓𝑖𝑛𝑒) 57
Therefore, Pr(not very fine) = = 59 = 0.966
𝑛(𝑈)
(d) “Neither heavy nor medium” is equivalent to “fine” or “very fine” in the problem. There
is 2+14 = 16 of these pencil packs.
𝑛(𝑛𝑒𝑖𝑡ℎ𝑒𝑟 ℎ𝑒𝑎𝑣𝑦 𝑛𝑜𝑟 𝑚𝑒𝑑𝑖𝑢𝑚) 16
Thus, Pr (neither heavy nor medium) = = 59 = 0.271
𝑛(𝑈)
Definition of Empirical (Relative Frequency) Probability
If E is some event of an experiment that has been performed a number of times, yielding a
frequency distribution of events or outcomes, then the empirical probability of event E
occurring when the experiment is performed one more time is given by:
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡𝑖𝑚𝑒𝑠 𝑡ℎ𝑎𝑡 𝑒𝑣𝑒𝑛𝑡 𝑜𝑐𝑐𝑢𝑟𝑒𝑑 𝑓(𝐸)
Pr(E) = 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑡𝑖𝑚𝑒𝑠 𝑡ℎ𝑎𝑡 𝑒𝑥𝑝𝑒𝑟𝑖𝑚𝑒𝑛𝑡 𝑤𝑎𝑠 𝑝𝑒𝑟𝑓𝑜𝑟𝑚𝑒𝑑 = ∑𝑓
Other Examples
A number of families of a particular type were measured by the number of children they have,
given the following frequency distribution:
Number of children 0 1 2 3 4 5
Number of families 12 28 22 8 2 2
Use this information to calculate the (relative frequency) probability that another family of
this type chosen at random will have: (a) 2 children (b) 3 or more children (c) less than 2
children.
Solutions
Here, Σf = total number of families = 74
𝑓(2 𝑐ℎ𝑖𝑙𝑑𝑟𝑒𝑛) 22
(a)Pr (2 children) = = = 0.297
∑𝑓 74
(b) f (3 or more children) = 8+2+2 = 12
Thus, Pr(3 or more children = 1274 = 0.162
𝑓(𝑙𝑒𝑠𝑠 𝑡ℎ𝑎𝑛 2 𝑐ℎ𝑖𝑙𝑑𝑟𝑒𝑛) 40
(c) Pr(less than 2 children) = = 74 =0.541
∑𝑓
Laws of Probability
There are four basic laws of probability.
1. Addition Law for mutually exclusive events
2. Addition Law for events that are not mutually exclusive
3. Multiplication Law for Independent events
4. Multiplication Law for Dependent events.
Addition Law for Mutually Exclusive Events
Two events are said to be mutually exclusive events if they cannot occur at the same time.
The addition law states that if events A and B are mutually exclusive events, then: Pr (A or
B) = Pr (A) + Pr (B)
Examples
The purchasing department of a big company has analysed the number of orders placed by
each of the 5 departments in the company by type as follows:
Table 9.1: Departmental Orders
Type of Order Sales Purchasing Production Accounts Maintenance Total
Consumables 10 12 4 8 4 38
Equipment 1 3 9 1 1 15
Special 0 0 4 1 2 7
Total 11 15 17 10 7 60
An error has been found in one of the orders. What is the probability that the incorrect order
came from
(a) Came from maintenance?
(b) Came from Production?
(c) Came from Maintenance or Production?
(d) Came from neither maintenance nor production?
Solutions:
a) since there are 7 maintenance orders out of the 60,
Pr (maintenance) = 7/60 = 0.117
b) Similarly, Pr (Production) = 17/60 = 0.283
c) Maintenance and production departments are two mutually exclusive events so that, Pr
(maintenance or production) = Pr (maintenance) + Pr(production = 0.117 + 0.283 = 0.40
d) Pr (neither maintenance nor production) = 1-Pr (maintenance or production) = 1-0.4 = 0.6
Addition Law for Events that are Not Mutually Exclusive Events
If events A and B are not mutually exclusive, that is, they can either occur together or occur
separately, then according to the Law:
Pr (A or B or Both) = Pr (A) +Pr (B) – Pr (A).Pr (B)
Example
Consider the following contingency table for the salary range of 94 employees:
Table 9.2: Contingency Table for the Salary range of 94 Employees
Salary /Month Men Women Total
N10,000 and Above 20 37 57
Below N10,000 15 22 37
Total 35 59 94
What is the probability of selecting an employee who is a man or earns below N10,000 per
month?
Solution
The two events of being a man and earning below N10, 000 is not mutually exclusive. It
follows that:
Pr(employee man OR earning below N10,000) = Pr(been a Man) + Pr(Earning below
N10,000) – Pr(been a Man).Pr(Employee earning below N10,000)
35 37 35 37
= 94 + 94 − 94 . 94
= 0.372 +0.394 – (0.372).(0.395)
= 0.766 – 0.147
= 0.619 0r 61.9%
𝑛!
Where Cn,x = 𝑥!(𝑛−𝑥)!
p = probability of success
q = probability of failure
p+q=1
Example Assume there is a drug store with 10 antibiotic capsules of which 6 capsules are
effective and 4 are defective. What is the probability of purchasing the effective capsules
from the drug store?
Solution
From the given information: The probability of purchasing an effective capsule is:
P = 6/10 = 0.60
Since p + q = 1; q = 1 – 0.60 = 0.40; n = 10; x = 6
Pr (6EC) = probability of purchasing the 6 effective capsules
= C10,6(0.6)6(0.4)4
10!
= (0.047)(0.026)
(6!(10−6)!)
[Link].6!
= (0.0012)
6!.4!
[Link]
= (0.0012)
[Link]
= = 210(0.0012) = 0.252
Hence, the probability of purchasing the 6 effective capsules out of the 10 capsules is 25.2
percent.
Joint, Marginal, Conditional Probabilities, and the Bayes Theorem
Joint Probabilities
A joint probability implies the probability of joint events. Joint probabilities can be
conveniently analysed with the aid of joint probability tables.
The Joint Probability Table
A joint probability table is a contingency table in which all possible events for a variable are
recorded in a row and those of other variables are recorded in a column, with the values listed
in corresponding cells as in the following example.
Example: Consider a research activity with the following observations on the number of
customers that visit XYZ supermarket per day. The observations (or events) are recorded in
a joint probability table as follows:
Table 9.3: Joint Probability Table
Age(Years) Male (M) Female (F) Total
Below 30 (B) 60 70 130
30 and Above (A) 60 20 80
Total 120 90 210
Marginal Probabilities
The Marginal Probability of an event is its simple probability of occurrence, given the sample
space. In the present discussion, the results of adding the joint probabilities in rows and
columns are known as marginal probabilities.
The marginal probability of each of the above events:
Male (M), Female (F), Below 30 (B), and Above 30 (A) are as follows:
Pr (M) = Pr(B∩M) + Pr(A∩M) = 0.2857 + 0.2857 = 0.57
Pr (F) = Pr(B∩F)+Pr(A∩F) = 0..3333 + 0.0952 = 0.43
Pr (B) = Pr(B∩M)+Pr(B∩F) = 0.2857+0.3333 = 0.62
Pr (A) = {r(A∩M)+Pr(A∩F) = 0.2857+0.0952 = 0.38
The joint and marginal probabilities above can be summarised in a contingency table as
follows:
Table 9.4: Joint and Marginal Probability Table.
Age(Years) Male (M) Female (F) Marginal Probability
Below 30 (B) 0.2857 0.3333 0.62
30 and Above (A) 0.2857 0.0952 0.38
Marginal Probability 0.57 0.43 1.00
Conditional Probability
Assuming two events, A and B, the probability of event A, given that event B has occurred is
referred to as the conditional probability of event A.
In symbolic term:
Pr( A∩ B) Pr(A).Pr(B)
Pr (A/B) = = = Pr(A)
Pr(B) Pr(B)
Where Pr (A/B) = conditional probability of event A
Pr (A∩B) = joint probability of events A and B
Pr (B) = marginal probability of event B
𝐽𝑜𝑖𝑛𝑡 𝑃𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑒𝑣𝑒𝑛𝑡𝑠 𝐴 𝑎𝑛𝑑 𝐵
In general, Pr(A/ B ) = 𝑀𝑎𝑟𝑔𝑖𝑛𝑎𝑙 𝑃𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑒𝑣𝑒𝑛𝑡 𝐵
Pr(A).Pr(B/A)
Pr(A/ B ) = Pr(B)
As an example in the use of Bayes theorem, if the probability of meeting a business contract
date is 0.8, the probability of good weather is 0.5 and the probability of meeting the date given
good weather is 0.9, we can calculate the probability that there was good weather given that
the contract date was met.
Let G = good weather, and m = contract date was met
Given that: Pr (m) = 0.8; Pr (G) = 0.5; Pr (m/G) = 0.9, we need to find Pr (G/m):
From the Bayes theorem:
𝑃𝑟( 𝐺 )⋅ 𝑃𝑟( 𝑚 𝐺) (0.5)(0.9)
Pr( Gm ) = =
𝑃𝑟(𝑚) 0.8
= 0.5625 or 56.25%
BINOMIAL DISTRIBUTION
This is a discrete probability distribution use for a repeated random experiment that has only
two possible outcomes. These outcomes are usually called success and failure with
probability p and q respectively.
There are many experiments that conform either exactly or approximately to the following
list of requirements:
1. The experiment consists of a sequence of n smaller experiments called trials, where n is
fixed in advance of the experiment.
2. Each trial can result in one of the same two possible outcomes (dichotomous trials), which
we generically denote by success (S) and failure (F).
3. The trials are independent, so that the outcome on any particular trial does not influence
the outcome on any other trial.
4. The probability of success P(S) is constant from trial to trial; we denote this probability by
p.
Definition
An experiment for which Conditions 1–4 are satisfied is called a binomial experiment.
Examples of Binomial distribution are ;
(i) Experiment of tossing a fair coin repeatedly; the two possible outcomes are “a head” and
“not a head i.e. the tail ”.
(ii) The conduct of election in a ward; the outcomes are “winning in the election” and “losing
in the election”
PROPERTIES OF BINOMIAL DISTRIBUTION
If μ and σ respectively represents (denotes) the mean and standard deviation of a binomial
distribution, then:
(i) The mean , μ =np
(ii) The variance, σ2 =npq
(iii) The standard deviation, σ =√𝑛𝑝𝑞
𝑞−𝑝
(iv) Moment of coefficient of skewness =
√𝑛𝑝𝑞
1−6𝑝𝑞
(v) Moment of coefficient of kurtosis = 3 +
𝑛𝑝𝑞
Example:
A coin is tossed successively and independently n times.
We arbitrarily use S to denote the outcome H (heads) and F to denote the outcome T (tails).
Then this experiment satisfies Conditions 1–4.
Tossing a thumbtack n times, with S = point up and F = point down, also results in a binomial
experiment.
The Binomial Random Variable and Distribution
In most binomial experiments, it is the total number of S’s, rather than knowledge of exactly
which trials yielded S’s, that is of interest.
Definition
The binomial random variable X associated with a binomial experiment consisting of n
trials is defined as
X = the number of S’s (Success) among the n trials
Suppose, for example, that n= 3. Then there are eight possible outcomes for the experiment:
SSS SSF SFS SFF FSS FSF FFS FFF
From the definition of X, X(SSF) = 2, X(SFF) = 1, and so on. Possible values for X in an n-
trial experiment are x = 0, 1, 2, . . . , n. We will often write X ~ Bin(n, p) to indicate that X is
a binomial rv based on n trials with success probability p.
Notation
Because the probability mass function (pmf) of a binomial random variable X depends on the
two parameters n and p, we denote the pmf by b(x; n, p).
Example 1
Each of six randomly selected cola drinkers is given a glass containing cola S and one
containing cola F. The glasses are identical in appearance except for a code on the bottom to
identify the cola. Suppose there is actually no tendency among cola drinkers to prefer one
cola to the other.
Then p = P(a selected individual prefers S) = .5, so with X= the number among the six who
prefer S, X ~ Bin(6,.5).
Thus
Example 3:
In an examination 70% of the candidates pass. Use the binomial distribution to calculate the
probability that in a random sample of 10 candidates contains: (i) exactly 2 (ii) between 1 and
3 inclusive and (iii) at most 2 places.
Solution:
Let P be the probability that a candidate passes in the examination, then
70
P= = 0.7, q=1-p = 0.3, n=10
100
(i) P(exactly 2 passes) = P(x=2)
Using binomial distribution
P(x=r) = nCr Pr qn-r; x=0,1,2,3,……
10!
=2!(10−2)!(0.7)2 (0.3)8
= 45 x 0.49 x 0.0000656
=0.00145
Example 2:
If 40% of cocoa seed bought by a produce buyer are defective, find the probability that out of
5 cocoa seeds selected at random (i) at least 4 (ii) between 1 and 3 (iii) exactly 3 (iv) at most
1 are non-defective.
Solution:
Let P be the probability that a seed bought is non-defective. Then:
60 3 3 2
P = 60% = 100 = 5 ; =1-5 = 5 ; n =5
Example 3:
The probability that a diabetic patient survives when injected with a newly discovered drug
is 0.75. Find the probability that exactly 8 out of 10 diabetic patients survives on being
injected with the new drug.
Solution:
Let P be the probability that a patient survives when injected.
Then P=0.75, q= 1-P = 1- 0.75 =0.25, n=10
Let P(x= r) be the probability that in n trials, r patient would survive
Using Binomial distribution
P(x= exactly 8) = P(x = 8 ) = 10C8 (0.75)8 (0.25)2
= 45 x 0.1001 x 0.0625
= 0.2815
= 28.15%
Example 5:
Ten aspirants contest in 10 wards of local government area for a particular post. The
1
probability that a political aspirant wins in a ward is5. Find the probability that four of the
aspirants win their elections.
Solution: Let P be the probability that an aspirant wins a ward election.
1 4
Then P= 15 , q = 1 - 5 = 5 , n=10
Using binomial distribution, P (x= four aspirants win) = P(x=4)
1 4
= 10C4 (5)4 (5)6
10! 1 4
= (10−4)!4! (5)4 (5)6
10! 1 4
= (10−4)!4! (5)4 (5)6
[Link].6! 1 4 4 6
= ( ) ( )
6!4! 5 5
[Link] 1 4 4 6
= [Link] (5) (5)
10.3.7
= x 0.0016 x 0.2621
1
= 10 x 3 x 7 x 0.0016 x 0.2621
= 0.0880656
≈ = 8.81%
MODULE TWO
INFERENTIAL PROCEDURES
INFERENTIAL PROCEDURES
There are two types of inferential procedures: (1) Estimation, (2) Hypothesis testing
Estimation
In estimation a sample is drawn and studied, and inference is made about the population
characteristics on the basis of what is discovered about the sample. There may be sampling
variations because of chance fluctuations, variations in sampling techniques, and other
sampling errors. We, therefore, do not expect our estimate of the population characteristics to
be exactly correct. We do, however, expect it to be close. The real question in estimation is
not whether our estimate is correct or not but how close is it to be the true value. Our first
interest is in using the sample mean (X̅) to estimate the population mean (μ).
Unbiased Estimator
An unbiased estimator is one which, if we were to obtain an infinite number of random
samples of a certain size, the mean of the statistic would be equal to the parameter. The sample
mean, (X̅) is an unbiased estimate of (μ) because if we look at possible random samples of
size N from a population, mean of the sample would be equal to μ. Consistent
Estimator
A consistent estimator is one that as the sample size increases, the probability that estimate
has a value close to the parameter also increase. Because it is a consistent estimator, a sample
mean based on 20 scores has a greater probability of being closer to (μ) than does a sample
mean based upon only 5 scores. Better estimates of a population mean should be more
probable from large samples.
Accuracy of Estimation
The sample mean is an unbiased and consistent estimator of (μ) . But we should not overlook
the fact that an estimate is just a rough or approximate calculation. It is unlikely in any
estimate that (X̅) will be exactly equal to (μ). Whether or not X̅ is a good estimate of (μ)
depends upon the representativeness of sample, the sample size, and the variability of scores
in the population.
HYPOTHESIS TESTING
In addition to estimating the accuracy of parameter estimates, the sampling distribution of the
mean serves a very important function in hypothesis testing. Imagine that someone told you
that they had a magic die that was loaded to show 6 when it was thrown.
Would you simply believe them and purchase the die for N100? You would surely want to
test the die before purchasing it? Specifically, you would want to test the hypothesis that the
die is in fact loaded.
A hypothesis is a tentative statement of a relationship between two variables, or as Neuman
(1997, p. 108) puts it, hypotheses are educated ‘guesses about how the social world works’.
Hypothesis testing is a logical and empirical procedure whereby hypotheses are formally set
up and subjected to empirical test. In the first stage of hypothesis testing, the researcher states
a research question, and poses two hypotheses that refer to the possible outcomes of the
empirical investigation. The research question is the question that the researcher wants to
answer by doing the research. In our loaded die problem, we would want to test whether the
die shows 6 more often than a fair die. This would tell us whether the die was loaded or not.
The research question for this investigation would be: ‘Does the “magic” die show 6 more
often than a fair die?’ Answering this question would be the whole point of the research.
H0: The loaded die shows 6 with the same probability as a fair die.
In contrast to the null hypothesis, the alternative hypothesis is a statement that maintains that
there are differences between the groups or conditions. This hypothesis makes a conjecture
that is diametrically opposed to the null hypothesis. The alternative hypothesis is represented
by the symbol H1. The alternative hypothesis can take two forms, depending on the nature of
the research question: it can be either directional or non-directional. A directional alternative
hypothesis anticipates the direction of difference. It states the researcher’s expectation
regarding whether one group is going to score higher or lower than the other group. A non-
directional hypothesis merely states that a difference is expected, without anticipating the
direction of the difference.
The ‘loaded die’ research question involves a directional alternative hypothesis because we
want to determine whether the loaded die shows heads more often than a fair die:
H1: The loaded die shows 6 more often than a fair die.
Thus far the null hypothesis and alternative hypothesis have been written out in words.
However, they are usually written in symbolic format. At the outset of a research project,
before engaging in any empirical testing, the researcher should state the research question and
hypotheses closely analogous to the following:
2. Research question: Do women and men perform the same at facial recognition tasks?
H0: μ1 = μ2
H1: μ1 ≠ μ2
The research question in the first example implies a directional alternative hypothesis. The
words ‘less intelligent’ in the research question indicate that a ‘less than’ sign (i.e.< ) should
be used in H1 to show the researcher’s expectation. The research question in the second
example is non-directional, and a ≠ sign is used in H1 to indicate the absence of direction.
In hypothesis testing, we are not really interested in whether or not our sample means differ.
They may differ because of random variation introduced by the sampling process (i.e. error
variance). We are interested in whether or not the population means differ, therefore the
hypotheses are stated in terms of the population parameter (μ) not the sample statistic (X̄).
The mean of the first population (e.g. individuals in crowds; women) is represented by μ1,
and the mean of the second population (e.g. individuals not in crowds; men) is represented
by μ2. Once the research question and hypotheses have been stated, the researcher may
proceed to test the hypotheses empirically. The results of the empirical investigation will
indicate whether the null hypothesis or the alternative hypothesis should be rejected.
A Type I error is made by rejecting the null hypothesis when it in fact is true.
A Type II error is made by not rejecting the null hypothesis when it is false.
Example A
We have a medicine that is being manufactured and each pill is supposed to have 14
milligrams of the active ingredient. What are our null and alternative hypotheses?
Solution
H0 : μ = 14
Ha : μ ≠ 14
Our null hypothesis states that the population has a mean equal to 14 milligrams. Our
alternative hypothesis states that the population has a mean that is different than 14
milligrams.
Example B
The school principal wants to test if it is true what teachers say – that high school juniors use
the computer an average 3.2 hours a day. What are our null and alternative hypotheses?
Solution
H0 : μ = 3:2
Ha : μ ≠ 3:2
Our null hypothesis states that the population has a mean equal to 3.2 hours. Our alternative
hypothesis states that the population has a mean that differs from 3.2 hours.
Deciding Whether to Reject the Null Hypothesis: One and Two-Tailed Hypothesis Tests
The alternative hypothesis can be supported only by rejecting the null hypothesis. To reject
the null hypothesis means to find a large enough difference between your sample mean and
the hypothesized (null) mean that it raises real doubt that the true population mean is 20. If
the difference between the hypothesized mean and the sample mean is very large, we reject
the null hypothesis. If the difference is very small, we do not. In each hypothesis test, we
have to decide in advance what the magnitude of that difference must be to allow us to reject
the null hypothesis.
Below is an overview of this process. Notice that if we fail to find a large enough difference
to reject, we fail to reject the null hypothesis. Those are your only two alternatives.
When a hypothesis is tested, a statistician must decide on how much of a difference between
means is necessary in order to reject the null hypothesis.
Statisticians first choose a level of significance or alpha (a) level for their hypothesis test.
Similar, to the significance level you used in constructing confidence intervals, this alpha
level tells us how improbable a sample mean must be for it to be deemed "significantly
different" from the hypothesized mean. The most frequently used levels of significance are
0:05 and 0:01: An alpha level of 0.05 means that we will consider our sample mean to be
significantly different from the hypothesized mean if the chances of observing that sample
mean are less than 5%. Similarly, an alpha level of 0.01 means that we will consider our
sample mean to be significantly different from the hypothesized mean if the chances of
observing that sample mean are less than 1%.
which is the sum of the squares of v independent standard normal variates, follow Chi-square
distribution with v d.f.
Applications of the χ2-Distribution
Chi-square distribution has a number of applications, some of which are enumerated below:
(i) Chi-square test of goodness of fit.
(ii) χ2-test for independence of attributes
(iii) To test if the population has a specified value of variance σ2.
(iv) To test the equality of several population proportions
The chi-square test can be used to determine how well theoretical distributions such as the
normal and binomial distributions) fit empirical distributions (i.e. those obtained from sample
data). Suppose we are given a set of observed frequencies obtained under some experiment
and we want to test if the experimental results support a particular hypothesis or theory. Karl
Pearson in 1900, developed a test for testing the significance of the discrepancy between
experimental values and the theoretical values obtained under some theory or hypothesis.
This test is known as χ2-test of goodness of fit and is used to test if the deviation between
observation (experiment) and theory may be attributed to chance (fluctuations of sampling)
or if it is really due to the inadequacy of the theory to fit the observed data.
Under the null hypothesis that there is no significant difference between the observed
(experimental and the theoretical or hypothetical values i.e. there is good compatibility
between theory and experiment.
Karl Pearson proved that the statistic.
Follows χ2-distribution with v = n-1, d.f. where O1, O2,..................On are the observed
frequencies and E1, E2,..................En are the corresponding expected or theoretical
frequencies obtained under some theory or hypothesis.
In statistical analysis, the Chi-Square distribution is used in many hypothesis tests and is
determined by the parameter k degree of freedoms. It belongs to the family of continuous
probability distributions. The Sum of the squares of the k independent standard random
variables is called the Chi-Squared distribution. Pearson’s Chi-Square Test formula is
Courtesy: Scribbr
When k is greater than 2, the shape of the distribution curve looks like a hump and has a low
probability that X2 is very near to 0 or very far from 0. The distribution occurs much longer
on the right-hand side and shorter on the left-hand side. The probable value of X2 is (X2 - 2).
Courtesy: Scribbr
When k is greater than ninety, a normal distribution is seen, approximating the Chi-square
distribution.
Chi-Square P-Values
Here P denotes the probability; hence for the calculation of p-values, the Chi-Square test
comes into the picture. The different p-values indicate different types of hypothesis
interpretations.
1. P <= 0.05 (Hypothesis interpretations are rejected)
2. P>= 0.05 (Hypothesis interpretations are accepted)
The concepts of probability and statistics are entangled with Chi-Square Test. Probability is
the estimation of something that is most likely to happen. Simply put, it is the possibility of
an event or outcome of the sample. Probability can understandably represent bulky or
complicated data. And statistics involves collecting and organising, analysing, interpreting
and presenting the data.
Finding P-Value
When you run all of the Chi-square tests, you'll get a test statistic called X2. You have two
options for determining whether this test statistic is statistically significant at some alpha
level:
1. Compare the test statistic X2 to a critical value from the Chi-square distribution table.
2. Compare the p-value of the test statistic X2 to a chosen alpha level.
Test statistics are calculated by taking into account the sampling distribution of the test
statistic under the null hypothesis, the sample data, and the approach which is chosen for
performing the test.
The p-value will be as mentioned in the following cases.
• A lower-tailed test is specified by: P(TS ts | H0 is true) p-value = cdf (ts)
• Lower-tailed tests have the following definition: P(TS ts | H0 is true) p-value = cdf (ts)
• A two-sided test is defined as follows, if we assume that the test static distribution of H0 is
symmetric about 0. 2 * P(TS |ts| | H0 is true) = 2 * (1 - cdf(|ts|))
Where:
P: probability Event
TS: Test statistic is computed observed value of the test statistic from your sample cdf():
Cumulative distribution function of the test statistic's distribution (TS)
Types of Chi-square Tests
Pearson's chi-square tests are classified into two types:
1. Chi-square goodness-of-fit analysis
2. Chi-square independence test
These are, mathematically, the same exam. However, because they are utilized for distinct
goals, we generally conceive of them as separate tests.
Properties
The chi-square test has the following significant properties:
1. If you multiply the number of degrees of freedom by two, you will receive an answer that
is equal to the variance.
2. The chi-square distribution curve approaches the data is normally distributed as the degree
of freedom increases.
3. The mean distribution is equal to the number of degrees of freedom.
(vii) On the other hand, if calculated value of χ2 is greater than the tabulated value, it is said
to be significant. In other words, discrepancy between observed and expected frequencies
cannot be attributed to chance and we reject the null hypothesis. Thus, we conclude that the
experiment does not support the theory.
Example 1:A pair of dice is rolled 500 times with the sums in the table below
Sum(x) Observed frequency
2 15
3 35
4 49
5 58
6 65
7 76
8 72
9 60
10 35
11 29
12 6
Take α = 5%
It should be noted that the expected sums if the dice are fair, are determined from the
distribution of x as in the table below:
Sum(x) P(x)
2 1/36
3 2/36
4 3/36
5 4/36
6 5/36
7 6/36
8 5/36
9 4/36
10 3/36
11 2/36
12 1/36
To obtain the expected frequencies, the P(x) is multiplied by the total number of trials.
Sum(x) Observed Frequency P(x) Expected Frequency
(O) (P(x).500)
2 15 1/36 13.9
3 35 2/36 27.8
4 49 3/36 41.7
5 58 4/36 55.6
6 65 5/36 69.5
7 76 6/36 83.4
8 72 5/36 69.5
9 60 4/36 55.6
10 35 3/36 41.7
11 29 2/36 27.8
12 6 1/36 13.9
The psychiatrist wants to investigate whether the distribution of the patients by social class
differed in these two units.
She therefore erects the null hypothesis that there is no difference between the two
distributions. This is what is tested by the chi squared (χ²) test. By default, all χ² tests are
two sided.
It is important to emphasise here that χ² tests may be carried out for this purpose only on the
actual numbers of occurrences, not on percentages, proportions, means of observations, or
other derived statistics. Note, we distinguish here the Greek (χ²) for the test and the
distribution and the Roman (x²) for the calculated statistic, which is what is obtained from the
test.
Ensure you complete the exercise and submit as part of your C.A.
Example3
Let's say you want to know if gender has anything to do with political party preference. You
poll 440 voters in a simple random sample to find out which political party they prefer. The
results of the survey are shown in the table below:
Republican Democrat Independent Total
Male 100 70 30 200
Female 140 60 40 240
Total 240 130 70 440
Similarly, you can calculate the expected value for each of the cells.
Expected values.
This discovery started a new field, viz ‘Exact Sample Test’ in the history of statistical
inference.
Note: If x1, x2...............xn is a random sample of size n from a normal population with mean
μ and variance σ2 then the Student’s t statistic is defined as:
𝛴𝑥
Where 𝑋 = 𝑛is the sample mean and is an unbiased estimate of the population variance
𝑛
σ2 .
Applications of t-distribution
(i) t-test for the significance of single mean, population variance being unknown
(ii) t-test for the significance of the difference between two sample means, the population
variances being equal but unknown
(iii) t-test for the significance of an observed sample correlation coefficient.
Test for Single Mean
Sometimes, we may be interested in testing if:
(i) The given normal population has a specified value of the population mean, say μo.
(ii) The sample mean differs significantly from specified value of population mean.
(iii) A given random sample x1, x2...............xn of size n has been drawn from a normal
population with specified mean μo.
Basically, all the three problems are the same. We set up the corresponding null hypothesis
thus:
(a) Ho: μ = μo i.e. the population mean is μo
(b) Ho: There is no significant difference between the sample mean and the population mean.
In order words, the difference between and μ is due to fluctuations of sampling.
(c) Ho: The given random sample has been drawn from the normal population with mean μo.
Under Ho the test-statistic is:
And it follows Student’s t-distribution with (n-1) degrees of freedom.
We compute the test-statistic using the formula above under Ho and compare it with the
tabulated value of t for (n-1) [Link] the given level of significance. If the absolute value of the
calculated t is greater than tabulated t, we say it is significant and the null hypothesis is
rejected. But if the calculated t is less than tabulated t, Ho may be accepted at the level of
significance adopted.
The tabulated value of t for 9 d.f. at 5% level of significance is 2.26. Since the calculated t is
much greater than the tabulated t, it is highly significant. Hence, null hypothesis is rejected at
5% level of significance, and we conclude that the sample mean differ significantly.
This is an unbiased estimate of the common population variance σ2 based on both the samples.
By comparing the computed value of t with the tabulated value of t for n1 + n2 -2 d.f. and at
desired level of significance, usually 5% or 1%, we reject the null hypothesis.
Example: The nicotine content in milligram of two samples of tobacco were found to be as
follows:
Sample A: 24 27 26 21 25
Sample B: 27 30 28 31 22 36
Can it be said that the two samples come from the same normal population having the same
mean?
Solution Hints: Applying the above formula and calculating the variance as appropriate, the
calculated t-value is -1.92. the tabulated value for 9 d.f. at 5% level of significance for two-
tailed test is 2.262. Since calculated t is less than the tabulated t, it is not significant, and the
null hypothesis is accepted.
T-test has very wide applications. It can be applied in the tests of single mean, in the
comparison of two different means and in the test of significance of other parameter estimates.
The term Analysis of Variance was introduced by Prof. R.A Fisher in 1920s to deal with
problem s in the analysis of agronomical data. Variation is inherent in nature. The total
variation in any set of numerical data is due to a number of causes which may be classified
as:
(i) Assignable causes and (ii) chance causes
The variation due to assignable causes can be detected and measured whereas the variation
due to chances is beyond the control of human and cannot be traced separately.
Also, compute the mean 𝑋̅of all the data observations in the k-classes by the formula:
Step 3: Obtain the Between Classes Sum of Squares (BSS) by the formula:
Step 5: Obtain the Within Classes Sum of Squares (WSS) by the formula:
Which follows F-distribution with (v1 = k-1, v2 = n-k) d.f (This implies that the degrees of
freedom are two in number. The first one is the number of classes (treatment) less one, while
the second d.f is number of observations less number of classes)
Step 8: Find the critical value of the test statistic F for the degree of freedom and at desired
level of significance in any standard statistical table.
If computed value of test-statistic F is greater than the critical (tabulated) value, reject (Ho,
otherwise Ho may be regarded as true.
Step 9: Write the conclusion in simple language.
Example 1: To test the hypothesis that the average number of days a patient is kept in the
three local hospitals A, B and C is the same, a random check on the number of days that seven
patients stayed in each hospital reveals the following:
Hospital A 8 5 9 2 7 8 2
Hospital B 4 3 8 7 7 1 5
Hospital C 1 4 9 8 7 2 3
Solution: Let X1j, X2j, X3j denote the number of days the jth patient stays in the hospitals A,
B and C respectively
Calculations for various Sum of Squares
Within Sample Sum of Square: To find the variation within the sample, we compute the
sum of the square of the deviations of the observations in each sample from the mean values
of the respective samples (see the table above)
Sum of Squares within Samples =
To obtain the variation between samples, we compute the sum of the squares of the deviations
of the various sample means from the overall (grand) mean.
The total variation in the sample data is obtained on calculating the sum of the squares of the
deviations of each sample observation from the grand mean, for all the samples as in the table
below:
= 53.5232 + 38.4032 + 59.8832 = 151.81
Note: Sum of Squares Within Samples + S.S Between Samples = 147.71 + 4.10 =151.81
= Total Sum of Squares
Ordinarily, there is no need to find the sum of squares within the samples (i.e, the error sum
of squares), the calculations of which are quite tedious and time consuming. In practice, we
find the total sum of squares and between samples sum of squares which are relatively simple
to calculate. Finally, within samples sum of squares is obtained by subtracting Between
Samples Sum of Squares from the Total Sum of Squares:
ANOVA TABLE
Critical Value: The tabulated (critical) value of F for d.f (v1=2, v2=18) d.f at 5% level of
significance is 3.55
Since the calculated F = 0.25 is less than the critical value 3.55, it is not significant. Hence,
we fail to accept Ho.
However, in cases like this when MSS between classes is less than the MSS within classes,
we need not calculate F and we may conclude that the means,𝑋̅1, 𝑋̅2 ,𝑋̅3 and do not differ
significantly. Hence, Ho may be regarded as true.
Conclusion: Ho : μ1 = μ2 = μ3, may be regarded as true and we may conclude that there is no
significant difference in the average stay at each of the three hospitals.
Critical Difference: If the classes (called treatments in pure sciences) show significant effect
then we would be interested to find out which pair(s) of treatment differ significantly. Instead
of calculating Student’s t for different pairs of classes (treatments) means, we calculate the
Least Significant Difference (LSD) at the given level of significance. This LSD is also known
as Critical Difference (CD).
The LSD between any two classes (treatments) means, say 𝑋̅i and 𝑋̅j at level of significance
‘α’ is given by:
LSD (𝑋̅i - 𝑋̅j) = [The critical value of t at level of significance α and error d.f] X [S.E (𝑋̅i -
𝑋̅j)]
Note: S.E means Standard Error. Therefore, the S.E ( -) above mean the standard error of the
difference between the two means being considered.
A B
No of Sales 20 18
Average sales in (N ‘000) 170 205
Average sales in (N ‘000 ) 20 25
[Link] following figures show the distribution of digits in numbers chosen at random from a
telephone directory:
Digit 0 1 2 3 4 5 6 7 8 9 Total
Freq 1026 1107 997 966 1,075 933 1,107 972 964 853 10000
Test whether the digits may be taken to occur equally frequently in the directory. The table
value of χ2 for d.f at 5% level of significance is 16.92.
Hint: Set up the null hypothesis that the digits 0, 1, 2, 3, ..........9 in the numbers in the
telephone directory are uniformly distributed, i.e. all digits occur equally frequently in the
directory. Then, under the null hypothesis, the expected frequency for each of the digits 0, 1,
2, 3,.............9 is 10,000/10 = 1,000
13. The table below gives the retail prices of a commodity in some shops selected at random
in four cities of Lagos, Calabar, Kano and Abuja. Carry out the Analysis of Variance
(ANOVA) to test the significance of the differences between the mean prices of the
commodity in the four cities.
City Price per unit of the commodity in different shops
Lagos 9 7 10 8
Calabar 5 4 5 6
Kano 10 8 9 9
Abuja 7 8 9 8
If significant difference is established, calculate the Least Significant Difference (LSD) and
use it to compare all the possible combinations of two means (α=0.05).
14. Concord Bus Company just bought four different Brands of tyres and wishes to determine
if the average lives of the brands of tyres are the same or otherwise in order to make an
important management decision. The Company uses all the brands of tyres on randomly
selected buses. The table below shows the lives (in ‘000Km) of the tyres:
Brand 1: 10, 12, 9, 9
Brand 2: 9, 8, 11, 8, 10
Brand 3: 11, 10, 10, 8, 7
Brand 4: 8, 9, 13, 9
Test the hypothesis that the average life for each of brand of tyres is the same. Take α = 0.01