NORMAL
DISTRIBUTION
[Link]
BRINDA
NORMAL PROBABILITY CURVE:
The literal meaning of the term normal is average. Most of the things like
intelligence, wealth, beauty, height etc. are quite equally distributed. There are
quite a few persons who deviate noticeably from average, either above or
below it. If we plot such a distribution on a graph paper, we get a bell-shaped
curve, referred to as Normal Curve.
Normal curve was derived by Laplace and Gauss (1777-1855)
independently. They also named it 'curve of error', where 'error' is used in the
sense of a deviation from the normal, true value. In the honour of Gauss, it is also
known as Gaussian Curve’
The normal curve takes into account the law which states that the
greater the deviation from the mean or an average, the less
frequently it occurs. For e.g. in terms of Intelligence, it is rare to find
people with very low or very high intelligence. It's normally distributed
in the population.
Normal Probability Curve or most popularly known as NPC.
CONCEPT OF NORMAL CURVE:
Now, suppose if we draw a frequency polygone with the help of above distribution, we will have a curve as shown in the
fig.
The shape of the curve in Fig. is just like a 'Bell' and is symmetrical on both the sides.
If you compute the values of Mean, Median and Mode, you will find that these three are approximately the same (M52;
Md52 and Mo = 52).
This Bell shaped curve technically known as Normal Probability Curve
or simply Normal Curve and the corresponding frequency distribution of
scores, having just the same values of all three measures of central tendency
(Mean, Median and Mode) is known as Normal Distribution.
Many variables in the physical (e.g. height, weight, temperature etc.)
biological (e.g. age, longevity, blood sugar level) and behavioural (e.g.
Intelligence; Achievement; Adjustment, Anxiety, Socio-Economic-Status etc.)
sciences are normally distributed in the nature. This normal curve has a great
significance in mental measurement. Hence to measure such behavioural
aspects, the Normal Probability Curve in simple terms Normal Curve worked
as reference curve and the unit of measurement is described as o (Sigma).
THEORETICAL BASE OF THE NORMAL PROBABILITY CURVE
The normal probability curve is based upon the law of Probability (the
various games of chance) discovered by French Mathematician Abraham
Demoiver (1667-1754). In the eighteenth century, he developed its
mathematical equation and graphical representation also.
CHARACTERISTICS OF NORMAL PROBABILITY CURVE:
Some of the major characteristics of normal probability curve are as follows:
1. The normal curve is symmetrical:- The Normal Probability Curve (N.P.C.) is symmetrical
about the ordinate of the central point of the curve. It implies that the size, shape and slope of
the curve on one side of the curve is identical to that of the other. That is, the normal curve
has a bilateral symmetry
2. The normal curve is unimodal:- Since there is only one point in the curve which has
maximum frequency, the normal probability curve is unimodal, i.e. it has only one mode.
3. Mean, median and mode coincide:- The mean, median and mode of the normal distribution
are the same and they lie at the centre. [Mean =Median =Mode]
4. The maximum ordinate occurs at the centre:- The maximum height of the
ordinate always occurs at the central point of the curve that is, at the mid-
point.
5. The normal curve is asymptotic to the X-axis:- The Normal Probability
Curve approaches the horizontal axis asymptotically i.e., the curve
continues to decrease in height on both ends away from the middle point
(the maximum ordinate point); but it never touches the horizontal axis. It
extends infinitely in both directions i.e. from minus infinity (-) to plus infinity
(+) as shown in Figure. As the distance from the mean increases the curve
approaches to the base line more and more closely.
5) The Height of the Curve declines Symmetrically: In the normal probability
curve the height declines symmetrically in either direction from the maximum
point.
6) The Points of Influx occur at point ±1 Standard Deviation (plus/minus 1 *
sigma) The normal curve changes its direction from convex to concave at a point
recognised as point of influx. If we draw the perpendiculars from these two points
of influx of the curve to the horizontal X axis; touch at a distance one standard
deviation unit from above and below the mean (the central point).
7) The Total Percentage of Area of the Normal Curve within Two Points
of Influxation is fixed: Approximately 68.26% area of the curve lies within
the limits of plus/minus 1 standard deviation (plus/minus 1 * sigma) unit from
the mean.
8) The Total Area under Normal Curve may be also considered 100
Percent Probability: The total area under the normal curve may be
considered to approach 100 percent probability; interpreted in terms of
standard deviations. The specified area under each unit of standard deviation
are shown in this figure.
9) The Normal Curve is Bilateral: The 50% area of the curve lies to the
left side of the maximum central ordinate and 50% of the area lies to the
right side. Hence the curve is bilateral.
10) The Normal Curve is a mathematical model in behavioural Sciences
Specially in Mental Measurement: This curve is used as a measurement
scale. The measurement unit of this scale is± 10 (the unit standard
deviation).
Skewness (Sk) which displays the lack of equality between tails if one
tail is extended longer than other tail. If one tail gets stretched toward the left,
there is a negatively skewed distribution. If one tail is pulled toward the right,
there will be a positively skewed distribution (Johnson & Christensen, 2008).
A distribution is said to be "Skewed when the mean and median fall at
different points in the distribution and the balance i.e. the point of center of
gravity is shifted to one side or the other to left or right. In a normal
distribution the mean equals the median exactly and there is no skewness.
Skewness:-
Distribution can be characterized by their skewness: the nature and extent to which symmetry is absent.
It is an indication of how the measurements in a distribution are distributed.
There are two types of skewness which appear in the Normal Curve.
a) Positive Skewness
b) Negative Skewness
a) Positive Skewness:- Distributions are skewed positively or to the right, when scores are massed at the
low, i.e. the left end of the scale, and are spread out gradually toward the right end.
b) Negative Skewness:- Distribution is said to be skewed negatively or to the left, when scores are massed at
the high end of the scale, i.e. the right side of the curve, and are spread out gradually towards the low end
i.e. the left side of the curve.
IMPORTANCE OF NORMAL DISTRIBUTION
The Normal distribution is by far the most used distribution in inferential
statistics because of the following reasons:
1) Number of evidences are accumulated to show that normal distribution
provides a good fit or describe the frequencies of occurrence of many
variable facts in biological statistics, eg, sex ratio in births, in a country
over a number of years. The anthropometrical data, eg height, weight, etc.
The social and economic data e.g. rate of births, marriages and deaths. In
psychological measurements
.
2) The normal distribution is of great value in educational evaluation and
educational research, when we make use of mental measurement. It may be
noted that normal distribution is not an actual distribution of scores on
any test of ability or academic achievement, but is, instead, a mathematical
model. The distributions of test scores approach the theoretical normal
distribution as a limit, but the fit is rarely ideal and perfect.
APPLICATIONS/USES OF NORMAL DISTRIBUTION CURVE
There are number of applications of normal curve in the field of psychology as well as
educational measurement and evaluation. These are:
1) To determine the percentage of cases (in a normal distribution) within given limits or
scores.
2) To determine the percentage of cases that are above or below a given score or reference
point.
3) To determine the limits of scores which include a given percentage of cases to determine
the percentile rank of an individual or a student in his own group.
4) To find out the percentile value of an individual on the basis of his percentile
rank.
5) Dividing a group into sub-groups according to certain ability and assigning
the grades.
6) To compare the two distributions in terms of overlapping.
USES:
1. It is used to compare two distributions in terms of-overlapping:
If scores of two groups on a particular variable are normally distributed.
What we know about the group is the mean and standard deviation of both the
groups. And we want to know how much the first group over-laps the second
group or vice-versa at that time we can determine this by using the table area
under NPC.
2. NPC helps us in dividing a group into sub-groups according to certain
ability and assigning the grades:
When we want to divide a large group in to certain sub-groups according
to some specified ability at that time we use the standard deviation unite of a
NPC unite of scale.
3. NPC helps to determine the relative difficulty of test items or problems:
When it is known that what percentage of students successfully solved a
problem we can determine the difficulty level of the item or problem by using
table area under NPC.
4. NPC is useful to normalize a frequency distribution:
In order to normalize a frequency distribution we use Normal Probability
Curve. For the process of standardizing a psychological test this process is very
much necessary.
5. To test the significance of observations of experiments we use NPC:
In an experiment we test the relationship among variables whether these
are due to chance fluctuations or errors of sampling procedure or it is real
relationship. This is done with the help of table area under NPC.
6. NPC is used to generalize about population from the sample:
We compute standard error of mean, standard error of standard deviation
and other statistics to generalize about the population from which the sample
are drawn. For this computation we use the table area under NPC.
POINTS TO BE KEPT IN MIND WHILE CONSULTING TABLE OF AREA UNDER
NORMAL PROBABILITY CURVE
The following points are to be kept in mind to avoid errors, while consulting the N.P.C. Table.
1) Every given score or observation must be converted into standard measure i.e. Z score, by
using the following formula:
Z=X-M \σ
2) The mean of the curve is always the reference point, and all the values of areas are given in
terms of distances from mean which is zero.
3) The area in terms of proportion can be converted into percentage.
4) While consulting the table, absolute values of z should be taken. However, a negative value
of z shows that the scores and the area lie below the mean and this fact should be kept in mind
while doing further calculation on the area. A positive value of z shows that the score lies
above the mean ie. right side.
The Normal Probability Curve helps us to determine:
i. What percent of cases fall between two scores of a distribution?
ii. What percent of scores lie above a particular score of a distribution?
iii. What percent of scores lie below a particular score of a distribution?
1) To determine the percentage of cases in a normal distribution within
given limits of scores.
Often a psychometrician or psychology teacher is interested to know the number of
cases or individuals that lie in between two points or two limits. For example, a teacher
may be interested as to how many students of his class got marks in between 60% and 70%
in the annual examination, or he may be interested in how many students of his got marks
above 80%.
Example 1
An adjustment test was administered on a sample of 500 students of class VIII. The mean
of the adjustment scores of the total sample obtained was 40 and standard deviation obtained was
8, what percentage of cases lie between the score 36 and 48, if the distribution of adjustment
scores is normal in the universe.
Solution: In the problem it is given that :
N = 500
M=40
σ= 8
We have to find out the total % of the students who obtained score in between 36 and 48
on the adjustment test.
To find the required percentage of cases, first we have to find out the 2 scores for the raw
scores (X) 36 and 48, by using the formula Z = (X - M)/ σ
According to table of area under Normal Probability curve (N.P.C.) i.e. Table
No.
the total area of the curve lie in between M to +16 is 34.13 and in between M
to -0.56 is 19.15. .. The total area of the curve in between -0.5 0 to +1 6 is
19.15+34.13=53.28
Thus the total percentage of students who got scores in between 36 and 48 on
the adjustment test is 53.28 (Ans.)
SAMPLING ERRORS
Sampling is an act of extracting a representative part of
population for determining characteristics of the whole
population.
Sampling error is the deviation of the selected sample
from the true characteristics, traits, behaviours,
qualities, or figures of the entire population
Reasons of the Sampling Errors:
Sampling process error occurs because researchers draw different subjects from the
same population, but the subjects have individual differences.
The most frequent cause of the said error is a biased sampling procedure.
Another possible cause of this error is chance. The process of random sample
selection and probability sampling is done to minimize sampling process error, but it
is still possible that all the randomly chosen subjects are not representative of the
population.
The most common result of sampling error is systematic error wherein the results
from the sample differ significantly from the results from the entire population.
Two basic reasons of sampling error are:
Chance error: The error occurs by chance. For example, someone did a comparative study
on malnutrition in under five year children in two cities A and B. Unfortunately; city B had a
large number of slum dwellers. So it comprised a large number of malnourished children,
thus skewing the result.
Sampling bias: Sampling bias is a tendency to favour a selection of sample units that
possess particular characteristics. It may occur in the form of over-representation bias. For
example, a study is done on nursing students' satisfaction with staying at a hostel or as a
paying guest. This study is biased towards those students who are hostellers or paying guests,
but it excluded the students who come from their own homes or who are day scholars.
Another way to understand is if we do a study on the time spent by the students on Internet.
The students who participate will be those who have an email address. This study is biased
towards those who have an email address, but the ones who do not have an address are
excluded.
TYPES OF SAMPLING BIAS:
Self-selection bias: This type of bias happens in a situation when the participants
in the study have some kind of control over the study to participate or not. For
example, if a study is conducted on the number of people who can carry a load of
10 kg for 20 min, then only well-built people will have a preference to participate.
Exclusion bias: This type of bias happens when some people of the group are
eliminated from the study as the instance of day scholars mentioned above.
Healthy user bias: This type of bias occurs when the sample selected has more
likelihood to be healthier as compared to general population. For example, if the
sample is selected from students who take the food from the mess and not the
canteen where junk food is not served, thus there will be chances of healthy user
bias.
MINIMIZE SAMPLING ERRORS/BIAS
There is only one way to eliminate this error. This solution is to eliminate the concept of sample, and to test the
entire population. In most cases, this is not possible; consequently, what a researcher must to do is to minimize sampling
process error.
Sampling
sample error Sample
Sampling
error
population population
Larger the sample , lesser the chances of sampling errors
This can be achieved by a proper and unbiased probability sampling and by using a
large sample size. However, to minimize sampling bias, following things should be
done:
Avoid convenience or judgement sampling.
To ensure that the target population is well defined and the sample frame
should match it as much as possible.
When complete population cannot be sampled then care should be taken that
the population that is excluded one is not taking away the desired features,
which were supposed to be measured from the population.