PROBABILITY AND THE
NORMAL DISTRIBUTION
Dr. Omosivie Maduka
Consultant Public Health Physician
Senior Lecturer
Preventive and Social Medicine
What is probability
• Measure of the likelihood of the occurrence of an event.
• The bridge between descriptive and inferential statistics
• The basis of all statistical inference
• Probability values always run between 0 and 1 – 0 means the event
cannot occur while 1 means the event must occur
Probability – the science of uncertainty
• The use of probability theory allows the decision maker with only
limited information to analyze risks and minimize gamble.
• It enables the decision maker to choose between different courses of
action when the consequence of each is subject to uncertainty.
• Since there is considerable uncertainty in decision-making, it is
important that all the known risks involved be scientifically evaluated.
• Helpful in this evaluation is probability theory, which has often been
referred to as the science of uncertainty.
Approaches to probability
Approaches
to probability
Subjective Objective
approach approach
Relative
Classical frequency
approach approach
Classical definition of probability
Also called theoretical or apriori probability. This probability can be
calculated without carrying out the experiment because all the possible
outcomes are known.
Pr (E) = Number of outcomes favourable to event
Total number of possible outcomes
Relative frequency definition of probability
Also called empirical probability or aposteriori. Calculation is only
based on data collection.
Pr (E) = Number of times event occurred in past
Total number of observations
Group work
List and discuss all the possible outcomes of tossing:
1. one coin
2. one die
Group work
The national centre for health statistics reported that of every 1,000
deaths, 81 results from road traffic accident, 55 from cancer, 156 from
heart disease, 367 from infectious disease, 22 from childbirth and 319
from other causes.
Estimate the probability that a particular death is caused by heart
disease
Probability distributions
1. Discrete
2. Continuous
Probability – the mathematics underlying
statistics
• Empirical or relative frequency or aposteriori probability is used when
summarizing the characteristics of the sample
• Classical or Theoretical or apriori probability is useful in making
inference from the sample to the population
Probability Distributions
• A probability distribution is a distribution that shows the probabilities
of all possible values of the random variable.
• A random variable is a quantity that can take one of a set of mutually
exclusive values with a given probability
• A probability distribution is a theoretical distribution that is
expressed in a tabular form, graphically or mathematically.
Characteristics of Probability Distribution
The two important characteristics of a probability distribution are:
• The probability of a particular outcome must always be between 0
and 1 inclusive.
• The sum of the probabilities of all mutually exclusive outcomes is 1.
Types of Probability Distribution - Discrete
Discrete Probability Distribution:
This is the probability distribution in which we can derive probabilities
corresponding to every possible value of the random variable.
Examples include:
1. Geometric distribution
2. Binomial distribution - proportions
3. Poisson distribution - rates
Types of Probability distributions - Continuous
Continuous Probability Distribution:
• This is the probability distribution in which we can only derive
probability of the random variable, x, taking values in certain ranges
(because there are infinitely many values of x).
• It can be represented by a curve of the frequency on the value
(probability density function).
• The total area under curve is one.
• This area represents the probability of all possible events.
• The probability that x lies between two limits is equal to the area
under the curve between these values.
• Examples are Normal, Chi-squared, t and F distributions.
Normal Distribution, Standard
Normal Distribution and Central
Limit Theorem
Normal (Gaussian) distribution
This is perhaps the most important distribution in statistics. It is in fact the
basis of hypothesis testing. Its probability density function is:
• completely described by two parameters, the mean (μ) and the variance
(σ2)
• bell shaped (unimodal)
• symmetrical about the mean
• shifted to the right if the mean is increased and to the left if the mean is
decreased (assuming constant variance)
• flattened as the variance is increased but becomes more peaked as the
variance is decreased (assuming fixed mean)
• the mean, mode and median are equal
Normal curve
The Standard Normal distribution
• There are infinite numbers of Normal distributions depending on the
values of μ and σ.
• The Standard Normal distribution is one Normal distribution with
mean zero and variance one (and also standard deviation one).
• Done using the z score or the Standard Normal Deviate (z).
z = x-μ
σ
Standard Normal Curve (μ = 0, σ = 1)
Normal and Standard Normal Curves
Percentages of the Area under the Standard
Normal Curve
Important to Note!
• 68% of the distribution lies between -1 and +1 standard deviation.
• 95.5% of the distribution lies between -2 and +2 standard deviation.
• 99.7% of the distribution lies between -3 and +3 standard deviation.
• 95% of the distribution lies between -1.96 and +1.96 standard
deviation.
• 99% of the distribution lies between -2.58 and +2.58 standard
deviation.
Sampling Distributions
A distribution of sample statistics obtained from samples repeatedly
drawn from one or more populations
Illustration of standard error
Central Limit Theorem
This states that sampling distribution of certain statistics will approach
normality as sample size (n) increases regardless of the shape of the
sampled population
INFERENTIAL STATISTICS;
Hypothesis Testing and
Confidence Interval Estimation
INFERENTIAL STATISTICS
• Inferential statistics is made up of various techniques used to provide
information about parameter values based on observations made on the
values of statistics.
• There are two broad areas of statistical inference, which are hypothesis
testing and estimation.
• In hypothesis testing, a specific statement or hypothesis is generated about
a population parameter and sample statistics are used to assess the
likelihood that the hypothesis is true.
• In estimation, sample statistics are used to generate estimates about
unknown population parameters
• This includes confidence interval estimation and sample size estimation.
INFERENTIAL STATISTICS
• This is to say that inferential statistics can be employed
through:
- Hypothesis Testing
- Confidence Interval Estimation
- Sample Size Estimation
• These processes are based on Probability Theory and
Central Limit Theory.
HYPOTHESIS TESTING
• Hypothesis testing is essentially a method for decision making.
• The decision relates to the choice between two competing, mutually
exclusive, statements regarding one or more population parameters.
• The competing statements of fact are referred to respectively as the
null and alternative hypothesis.
• Null hypothesis (H0) is a precise statement concerning the
parameters of interest
• Alternate (research) hypothesis (Ha or H1) is a less precise competing
statement.
HYPOTHESIS TESTING
• The term null is used because the null hypothesis is a statement of no
difference.
• Null hypothesis states that there is no difference between population
mean and the sample mean or between the means of two groups.
• The alternate hypothesis asserts that there is a difference between
the population mean and the sample mean or between means of two
groups.
• The alternate hypothesis is the investigator’s belief
HYPOTHESIS TESTING
Null Hypothesis (H0)
• There is no difference between the mean weight of Ikwerre men (μ1)
and the mean weight of Kalabari men (μ2)
• H0: μ1=μ2 OR H0: μ1-μ2=0
HYPOTHESIS TESTING
Alternate (Research) Hypothesis (H1)
• There is true difference between the mean weight of Ikwerre men
(μ1) and the mean weight of Kalabari men (μ2)
• H0: μ1≠μ2 (Two tailed)
• The mean weight of Ikwerre men (μ1) is greater than the mean
weight of Kalabari men (μ2)
• H1: μ1>μ2 (Upper one-tailed)
• The mean weight of Ikwerre men (μ1) is less than the mean weight of
Kalabari men (μ2)
• H1: μ1<μ2 (Lower one-tailed)
HYPOTHESIS TESTING
• The decision as to whether the assertion of null hypothesis is to be
abandoned and the alternative assertion established as the true
condition is made by using a model of the sampling distribution of the
involved statistic to establish a decision criterion.
• If the criterion is met the null hypothesis is rejected.
• If the criterion is not met, one fails to reject the null hypothesis.
• Because the sampling distribution tends to be approximately normal
(with a large sample size), the normal curve is chosen as the test
model.
Steps in Hypothesis Testing
The main steps in hypothesis testing:
1. State the null hypothesis (H0) and alternate hypothesis (H1)
2. Set the level of significance
3. Select the appropriate test statistic
4. Set up the decision rule
5. Compute the test statistic
6. Draw your conclusion
7. State your assumptions
STATISTICAL ERRORS
RULES ERRORS
• Do not reject null hypothesis Type 1 Error (α Error)
(H0) when it is true - Rejecting null hypothesis (H0)
when it is true
• Reject null hypothesis (H0) when Type 2 Error (β Error)
it is false (Statistical Power). - Not rejecting null hypothesis
(H0) when it is false
p-value, level of significance, and hypothesis
testing
• p-value: The probability of making type I error
• Level of Significance (α): The value at which the difference is
considered to be statistically significant is the significance level (α).
• Simply put, p-value is like examination score, while α is like the pass
mark.
• Hypothesis testing is a method of deciding whether the data are
consistent with the null hypothesis.
• The calculation of p-value is an important part of the procedure.
STATISTICAL POWER AND TYPE II ERROR
• Statistical power is the probability of correctly rejecting null
hypothesis (H0) when a particular alternative value of the parameter
is true.
• In other words, power is rejecting null hypothesis when it is false.
• Power is 1 – β.
• In other word, β + power = 1.
• Power can be increased by increasing the sample size.
Test Statistic (Hypothesis tests)
Sample or group to be compared Parametric test Non-parametric test
Two independent samples with • Student’s t-test (n<30) • Wilcoxon rank sum test
continuous outcome • Z-test for means (n≥30) • Mann-Whitney U test
• Kendall’s S test
Two dependent samples with • Paired t-test • Wilcoxon signed rank test
continuous outcome
More than two independent • One-way ANOVA (F-test) • Kruskal-Wallis one-way ANOVA
samples with continuous outcome • Two-way ANOVA • Friedman two-way ANOVA
Two independent samples with • Z-test for proportion • Chi-Square test
dichotomous outcome • Fisher’s exact test
Two dependent samples with - • McNemar’s Chi-Square test
dichotomous outcome
Two or more independent samples - • Chi-Square test
with categorical (nominal or • Chi-Square test for trend (for
ordinal) outcome ordinal outcome)
Tests of Association
Purpose Parametric test Non-parametric test
To determine the degree of • Pearson’s • Spearman’s rank
association between two correlation correlation
quantitative variable • Kendall’s rank
correlation
To determine the relationship • Linear regression -
between quantitative variables
To determine the relationship - • Logistic regression
between a dependent categorical
variables and other variables
CONFIDENCE INTERVAL ESTIMATION
• In estimation, sample statistics are used to generate estimates about
unknown population parameters
• There are two types of estimates that can be produced for any
population parameter;
- Point Estimate
- Interval Estimate
• Point Estimate for a population parameter is a single-valued estimate
of that parameter
• Interval Estimate is a range of values for a population parameter with
a level of confidence attached.
CONFIDENCE INTERVAL ESTIMATION
• Confidence Interval is therefore a range of values which will
likely contain the population parameter. The lower boundary
of the confidence interval is called the lower limit, while the
upper boundary is called the upper limit.
• Confidence Interval (CI) is given thus:
• CI=Point Estimate ± Margin of Error
CONFIDENCE INTERVAL ESTIMATION
• The point estimate for continuous variable is mean (X) and that of a
categorical variable is proportion (p).
• Margin of Error= Test Statistic x Standard Error of point estimate
• CI=Point Estimate ± Test Statistic x Standard Error of point estimate
• Just as we have different formula for standard error of mean and
standard error of proportion, so do we have different formula for
confidence interval for mean and confidence interval for proportion
CONFIDENCE INTERVAL ESTIMATION
• The level of confidence set determines the probability level, and this
can be used to draw conclusion whether the result is statistically
significant or not.
• At 95% confidence interval, the probability level is 5% (i.e p=0.05).
Thus, all results with p<0.05 are concluded to be statistically
significant while those with p≥0.05 are concluded not statistically
significant.
CONFIDENCE INTERVAL ESTIMATION
• The sample mean, proportion or rate is the best estimate we have of
the true population mean, proportion or rate. The distribution of
these parameter estimates from many samples of the same size will
roughly be Normal.
Such an interval for the population mean µ is defined by:
- 1.96 x SE ( ) to 1.96 x SE ( ) for 95% Confidence Interval
Or Mean ± 1.96 x SE
For a proportion or rate the interval for the population estimate
Estimate ± 1.96 x SE of estimate
Graphical representation of confidence
interval and confidence limits
Adopted from Sylvia Wassertheil-Smoller, Biostatistics and Epidemiology: A Primer for Health Professionals, Springer-Verlag, New York
Hypothesis Test, p-value, Statistical
Significance and Confidence Interval
• Hypothesis test tells whether there is a difference in the parameter of
concern and gives significant and non-significant result;
• Hypothesis test can also be used to obtain the p-value which also tells
whether the result is statistically significant or not-significant;
• But does not tell what the difference is and how large the difference
is.
• Confidence interval gives a range of values in which we are confident
(sure) that the population difference would be contained in, tells
whether there is a difference in the parameter, and also tells what the
difference is and how large the difference is.
Hypothesis Test, p-value, Statistical
Significance and Confidence Interval
• Also note that statistical significance does not necessarily mean the
results is clinically or practically significant. The p-value does not
relate to the clinical importance of a finding, as it depends on a large
extent on the size of the study.
• Thus, a large study may find small, unimportant differences that are
highly significant and a small study may fail to find important
differences.
Hypothesis Test, p-value, Statistical
Significance and Confidence Interval
• Supplementing the hypothesis test with a confidence interval will
indicate the magnitude of the result and this will aid the investigator
to decide whether the difference is of interest clinically.
• The confidence interval gives an estimate of the precision with which
a statistic estimates a population value.
• If the 95% confidence interval does not include zero (i.e. the value
specified in the null hypothesis), then the hypothesis test will return a
statistically significant result.
Hypothesis Test, p-value, Statistical
Significance and Confidence Interval
• If the 95% confidence interval include zero then the hypothesis test will
return a non-significant result.
• The confidence interval shows the magnitude of the difference and the
uncertainty or lack of precision in the estimate of interest.
• Therefore, the confidence interval conveys more useful information than
the p-value.
• The presentation of both the p-value and the confidence interval is
desirable. But if only one of them is to be presented, it should be the
confidence interval. This is because presenting a 95% CI indicates whether
the result is statistically significant at the 5% level apart from indicating the
magnitude of parameter of interest together with the level of uncertainty.
VERY IMPORTANT POINTS TO NOTE
1. The size of the p-values does not indicate the
importance of the result
2. Result may be statistically significant but practically
unimportant
3. Differences that are not statistically significant are
not necessarily unimportant