Normal distribution
Normal distribution
The normal distribution was discovered in 1735 by French
mathematician De-moiver. He was the first person to discover
the normal distribution. He got this distribution as a limiting
case of Binomial distribution. While studying the binomial
distribution he showed that when the number of trials is large,
it can be approximated by a continuous curve. And he
applied it to problem associating in the game of chance.
Later it was known due to Laplace. Laplace expanded De -
Moiver’s ideas and connected them to the Central limit
theorem. He showed that the sum of many small independent
random variables tends to follow a normal distribution. This
made the normal distribution very important in probability
theory.
The mathematician Gauss used the normal
distribution to analyze errors in astronomical
measurements. He showed that these errors are often
normally distributed, later the curve known as the
Gaussian distribution. And popularized by Gauss, and
applied broadly in the 19th century to human and
natural measurements.
Normal distribution has become the most important
probability model in statistical analysis.
Normal distribution
Properties of a normal distribution
Normal distribution is bell – shaped and is symmetrical around a vertical line.
Mean, median, Mode of the distribution coinside,
Mean = Median = Mode
skewness =0, kurtosis =3
4
Mean deviation about mean = 5 σ
Q1 and Q2 are equidistant from Q2.
Approximately 65% of the observations may lie in the interval mean+ SD.
Approximately 95% of the observations may lie in the interval mean+ 2SD.
Approximately 99.7% of the observations may lie in the interval mean+3 SD.
Total area under the curve =1.( Represents total probability)
Asymptotic to the X – axis. The curve approaches the horizontal axis but never
touches it.
Unimodal (the curve has only one peak)
Defined by two parameters.
mean (µ) – defines the location (center)
Standard deviation (σ) – defines the spread(width)
The curve is smooth and continuous.
Tails are infinite.
The distribution extends to infinitely on both sides.
The shape of 3 normal distribution with same Sd but different mean is shown below
Constant SD
Graph of 3 Normal distribution with same mean but different SD is given below.
Constant mean
The larger the value of SD the wider and flatter is the Normal curve, thus showing more
variability in the data.
Standard normal distribution.
A standard normal distribution is a special type of normal distribution that has,
Mean µ = o, Sd = 1
It is also called Z distribution.
Probability density function of standard normal distribution is
𝑧2
1 −2
f(Z) = 𝑒
2π
Z – score
Any normal variable X can be converted into a standard
normal variate ‘Z’ using
𝑥−µ
z= , this helps in finding probabilities from Z – tables.
σ
Why we use standard normal distribution
Because instead of having many normal tables for different means
and Sd. We convert all normal problems to Z and use Z – table.
68% of the data between -1<Z<1.
95% of the data between -2<Z<2.
99.7% of the data between -3<Z<3.
Application
To calculate the probabilities.
we can find the probability that a value lie above a point, below a
point.
To convert any normal distribution into Z – scores.
𝑥−µ
z=
σ
Application
The normal distribution is widely used because many
natural and human processes follow a Bell shaped
pattern.
It helps in prediction, decision-making, quality control and
statistical inference.
1. In natural and biological measurement.
most biological characteristics follow a normal pattern,
such as height, weight, BP, body temperature, etc.
2. In quality control
Industries assume measurements (like length, weight,
diameter) follows a normal curve.
It is used In setting control limits, checking defective items, monitoring
production consistency.
3. In Ayurvedic drug manufacturing (e.g. tablet weight,
churnam particle size), the variation
around the mean is assumed normal to maintain quality.
4. In psychological and educational testing
scores of large group in tests usually follow a normal
distribution. Used in – grading, cutoff
setting, analyzing performance.
[Link] medical diagnosis and research.
Normal distribution helps in
• Identifying abnormal values.
• Setting reference ranges(normal BP, Hb level)
6. In sampling distribution and statistical inference used to
construct
(i)confidence interval
(ii) perform hypothesis testing.
(iii) Estimate population parameters.
Note:
Normal distribution is used when data
[Link] around an average.
[Link] symmetric variation.
[Link] fever extreme values.
Real life examples
1. Human heights
Most people have height near the average.
very tall and very short people are rare.
the distribution forms a bell shaped curve.
2. Human body weight
3. BP
4. Exam marks
Many students score around the average. Very high or
low scores are fever.
marks follows normal distribution.
5. Measurement errors follows normal distribution.
Assumptions of Normal distribution
A variable X is said to follow a Normal distribution if,
The data is continuous
Normal distribution applies only to continuous variables. E.g.
Height, weight, BP etc.
The distribution is Symmetrical.
The curve should be perfectly symmetric around the mean.
The left side and right side of the curve are mirror images.
Mean=Median=Mode.
All the measures of central tendency are equal and located at
the centre.
Bell shaped curve.
The curve should be unimodal(one peak) and shaped
like a bell.
Data is independent.
Each observation is independent of others. Ie, one
variable should not influence another.
No extreme outliers
Data should not have very large or very small values,
because they distort normality.
Large sample size.
If the sample size is large (n ≥30), the mean tends to be
normal, even if population is not
perfectly normal (CLT).
Constant variance ( Homo Scedasticity)
Constant variance means that the spread ( variability)of
data remains the same across all levels of the variable.
The data should not become more scattered or
less scattered as the values increase. The width of the
distribution is uniform every where.
How to test the assumptions of Normality
Graphical methods
(i) Histogram
Draw a histogram of the data and check if it looks like
bell-shaped and symmetric.
(ii) Q-Q plot (quantile- quantile plot)
If points lie on a straight line, data is normal. Small
deviations at extreme ends are
acceptable.
(iii) Box – plot
Summary statistics
(i) skewness (ii) Kurtosis
Box plot
A box plot is a graphical method used to display the spread and distribution of a dataset
based on five key summary values.
1. Minimum (lowest value)
2. Q1 (first quartile) – 25% of the data below this.
3. Median (Q2) – 50% of the data is below this.
4. Q3( third quartile) – 75% of the data is below this.
5. Maximum (largest value)
These values create the ‘box’ and ‘whiskers’
It is very useful in understanding
The center of the data.
The spread of the data.
The presence of outliers.
The summary or skewness of the data.
Positively skewed
If the distance from the median to the maximum is greater than the distance from
the median to minimum, then the data is positively skewed.
Negatively skewed
If the distance from the median to minimum is greater than the distance from the
median to the maximum, then the box plot is negatively skewed.
Skewness
Skewness is a measure that tells us how the data are
distributed around the mean, whether the distribution is
symmetric or tilted to one side. If the data are distributed
exactly in the same way on either side of the central value,
the distribution is said to be symmetrical otherwise it is said
to be skewed.
The concentration of observation is measured by
Skewness.
Normal curve or symmetric curve
The shape of the frequency curve gives an indication of
skewness of the distribution. If the Skewness is zero, the
distribution is symmetric and in this case the mean, median
and mode will co inside . The curve is bell shaped and it is
called a Normal curve.
𝑿 = Median = Mode
Positive skewness (Right skewed)
If the observations are distributed more to the right of the mode ie more than half the area
under the curve is to right of the mode, then the distribution is said to be positively skewed.
𝑿 > Median>mode or Mode< Median<Mean
Negatively skewed (left skewed)
If more than half of the area is to the left of the mode, the skewness is said to be negatively
skewed. The tail is longer on the left side.
Mean<Median<Mode or Mode>Median>Mean
𝑀𝑒𝑎𝑛 −𝑀𝑜𝑑𝑒 3 𝑀𝑒𝑎𝑛 −𝑀𝑒𝑑𝑖𝑎𝑛
Coefficient of skewness = or
𝑆𝐷 𝑆𝐷
kurtosis
The measurement of peakedness of the frequency curve is
called kurtosis.
If the frequency curve is more
Peaked than the normal curve
It is called leptokurtic. Heavy
Tail and sharp peak.
And it is flattered than the normal
Curve it is said to be platykurtic
(Flat and wide peak, data are
More spread out).
The normal curve usually called
Mesokurtic.
The degree of peakedness or kurtosis is measured
mathematically by computing the coefficients of kurtosis.
Coeffi mesociant of kurtosis denoted by β2 .
If β2 = 3 Mesokurtic.
β2 > 3 Leptokurtic.
β2 < 3 Platykurtic.
Empirical rule
Approximately 65% of the observations may lie in the interval mean+ SD.
Approximately 95% of the observations may lie in the interval mean+ 2SD.
Approximately 99.7% of the observations may lie in the interval mean+3 SD.
Calculate the percentage of data and compare.
Normal curve
Summary test
Mean = Median = Mode
If very different not normal.
Statistical test for normality
These tests mathematically check if the data deviates
from a normal distribution.
(i) Shapiro Wilk test
(ii) Kolmogorov smirnov test
Shapiro-Wilk test
The Shapiro –wilk test is a statistical test used to check whether a
given data set comes from a normal distribution.
It is one of the most powerful and widely used tests for normality.
Null hypothesis (H0) – The data follows a normal distribution.
Alternative hypothesis (H1) – The data does not follow a normal
distribution.
At 5% significance level the calculated p value ≤ 0.05,
We reject the null hypothesis and
P> 0.05 we accept the nullhypothesis.
W=
Approximation of binomial to Normal
Binomial → Normal
No of trials ‘n’ is large (n≥30 is a common rule of thumb)
Probability of success P, is not too close to ‘0’ or ‘1’.
Both np ≥ 5 and nq ≥ 5, q = 1-p
When ‘n’ is large , the binomial curve becomes bell-shaped, similar to normal
distribution.
Also for large ‘n’ poisson curve becomes bell- shaped, symmetric and similar to
Normal distribution.