0% found this document useful (0 votes)
5 views45 pages

Understanding Averages in Statistics

Chapter 3 discusses the importance of averages in statistical analysis, defining them as central values that summarize data characteristics. It covers various types of averages, including arithmetic mean, median, and mode, along with methods for calculating them in both ungrouped and grouped data. Additionally, the chapter addresses variability and its measures, emphasizing the significance of understanding dispersion in biological data.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views45 pages

Understanding Averages in Statistics

Chapter 3 discusses the importance of averages in statistical analysis, defining them as central values that summarize data characteristics. It covers various types of averages, including arithmetic mean, median, and mode, along with methods for calculating them in both ungrouped and grouped data. Additionally, the chapter addresses variability and its measures, emphasizing the significance of understanding dispersion in biological data.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Chapter 3

One of the most important objectives of statistical analysis is to get one single
valve that describe the characteristic of the entire mean of unwieldy data. Such a
valve is called the central valve or an average or the expected valve of the
variable.
“Average is an attempt to find one single figure to describe whole of figure”.-
Clark
Objective of Average:- There are two main objectives of average:-
(a). To get single valve that describe the characteristic of the entire group.
(b). Measure of central valve, by reducing the mass of data to one single figure,
enable comparison to be made. Comparison can be made either at a time or over
a period of time.

Requisition of a Good Average:- An average is a single valve representing a group


of valves, it is desired that such a valve satisfies the following proportions:-
(1). Easy to understand.
(2). Simple to compute.
(3). Based on all the items.
(4). Not be unduly affected by extreme observations.
(5). Capable of further Algebraic treatment.
(6). Sampling stability

Type of Average:- The following are the important types of Average—

(a). Arithmatic Mean or Mean:- The most popular and widely used measure of
representing the entire data by one valve is most layman call an “Average” and
the statistician call the arithmetic mean.
- It is the simplest form of measurement of central tendency.
- It is calculated by dividing the sum of all the observations by the number of
observations.
- Mean is denoted by X.
Calculation of Mean:-
Mean = Total or sum of the observations
Numbers of observations
X = ∑(X1+X2+……..+Xn)
N
The mean is calculated by different methods in two types of series, ungrouped
and grouped.

Ungrouped Series:- In such series the number of observations in small and there
are two methods for calculating the means. The choice depends upon the size of
observations in the series.
(a). When the observations are small in size

Example:- Hb of 10 ANC women is as follows: 9.5, 11.5, 12.5, 10, 11.7, 13, 10.8,
12.7, 13.2 and 14.2 mg/dl.
Mean X=
9.5+11.5+12.5+10+11.7+13+10.8+12.7+13.2+14.2__________________________
___________
10
Means Hb = 11.91 mg/dl

(b). When individual observations are large in size:- The arithmetic means can be
calculated by using as an arbitrary origin. When deviations are taken from an
arbitrary origin, the formula for mean is
X = A + _∑ x
N
Where A is assumed means and x is the deviations of items from assumed means
ie.,
x= (X – A)

Example:- Height in centimeters for 7 school children are given below:


148,143,160,152,157,150,155. Find the mean.
150 may be taken as the working origin to calculate the mean height
X X–A x
148 148-150 -2
143 143-150 -7
160 160-150 10
152 152-150 2
157 157-150 7
150 150-150 0
155 155-150 5

-------------------------------------------------
1065 15
X =A + _∑x_
N
=150 + 15/7
=150 + 2.1= 152.1

Grouped Series:- When the number of observations is large, the data are
arranged in groups and frequency distribution.
Mean X = ∑ fx / ∑f

Example:- Calculate the average period duration of menstrual cycle.


Days of period(X) No of women(f) Fx
3 22 66
4 34 136
5 25 125
6 18 108
7 12 84
-------- -----------
111 519
Mean = ∑ fX / ∑ f
= 519 / 111 = 4.67 days of period

Example:- Calculate average daily protein intake of formula from the following
data of 200 female. On the basis of their protein intake we can summarize data in
frequency distribution table as follows:-
class interval no of female (f) midpoint of class interval x fx
15-25 15 20 300
25-35 20 30 600
35-45 50 40 2000
45-55 55 50 2750
55-65 40 60 2400
65-75 15 70 1050
75-85 05 80 400
________ _______
200 9500
Mean= 9500/200 = 47.5

Merits and Demerits of Arithmetic Mean


Merits:- It is easy to calculate and to understand.
It is based on all observation and least affected by sampling fluctuations.
Demerits:- It is not possible to calculate mean for qualitative data.
It is very sensitive to extreme observation.
It cannot be determined graphically.
It cannot be calculated in open end series.

Median:-The data is first arranged in an ascending or descending of magnitude


and the valve of the middle observations is located, which is called the median.

Mode:- The mode is the commonly occurring valve in a distribution of data. It is


the most frequent item or the most “fashionable” valve in a series of observation.

For example:- the diastolic blood pressure of 20 individuals was:-


85,81,79,71,75,95,75,90,75,77,79,71,75,95,95,77,84,75,75,81.
The mode is the most frequently occurring valve is 75. The advantages of
mode are that it is easy to understand, and is not affected by the extreme items.
It can be used to describe qualitative phenomenon. The disadvantage are that the
exact location is often uncertain and is often not clearly defined. Therefore, mode
is not often in biological or medical statistics.

Variability and its Measures


Dispersion is the measure of the variation of the items.
Biological data, quantitative or qualitative, collected by measurement or
counting are very variable. No two measurements in man are absolutely equal,
not even the means or proportions of two series in health or in disease are equal
although we compare like with the like. Apart from Physical characters, the
mental qualities such as intelligence quotient, behavior and tendency to do wrong
or right vary from man to man and group to group. Variability is essentially a
normal character. In other words, occurrence of variability is an biological
phenomenon.
Type of Variability:- There are three type of variability are as follows:-
(1). Biological Variability:- Individuals, in similar environments differ when
compared as regards sex, class and other attributes but the difference noted may
be small and is send to occure by chance. Such a difference or variability is called
biological variability. Biological variability may be classified under following
heads:-
(a). Individual Variability:- One student height is 150 cm and that of another of the
same age is 175 cm. To find whether their heights are normal or not we find the
mean and standard deviation of heights in the population, by taking a large
representative sample.
(b). Periodical Variability:- The same individuals show variations in temperature,
pulse rate, blood pressure, WBC count, blood sugar, urea, chalosteral etc. at
different times of the day. The variations can be studied or analyzed by applying
suitable statistical techniques.
( c). Class, group of Category variability:- Height, weight, blood pressure etc. very
from class to class depending on age, sex, caste, social states or nature of work,
whether the mean height of males is more than female is determined by standard
error of difference between two means or by applying unpaired t- test if the
sample is small.
(d). Sampling Variability:- We do not examine each and every individual in the
population but a sample is taken which is much smaller. The valve of any sample,
therefore, will differ from those of the populations and further those will be
variability from one sample to another. This is a biological variability from one
sample to another sample. This is a biological variability of sample and is called,
sampling variability or sampling error or statistical error which is natural or
inevitable.
(e). Real Variability:- When the difference between two readings, observations or
valves of classes or samples is more than the defined limits in universe, it is said to
be real. The heights and weights in England are more than those in India perhaps
due to better socio-economic conditions. Attack rate in the group, vaccinated
against measles is very much less as compared with that in the unvaccinated
group because of protection given by vaccinated in theformer and not due to
chance.
Difference in the incidence of cancer among smokers and nonsmokers may
be due to excessive smoking and not due to chance. Higher coronary incidence
may be associated with high cholesterol in blood, tension, obesity and such other
factors. Standard error of mean and proportion help to verify and solve such
problems.
(f). Experimental Variability:- Error or difference or variation may be due to
materials, methods, procedures employed in the study or defects in the
techniques involved in the experiment. They are of three type:- Observer,
Instrumental and sampling.
Observer error:- Observer error may be subjected or objected
(a). Subjective:- An interviewer may even ask embarrassing questions which the
person may not like to answer such as menstrual history, pregnancy, use of family
planning methods, sexually transmitted diseases etc. some subjects are very keen
while others do not wish to give any information.
(b). Objective:- Objective errors may be added by an untrained observer while
recordings the measurements such as BP, pulse rate, in recording height, weight
etc. Biochemical tests may give different results in different hands and so on.
Instrumental Error:- This may be negligible or gross. Defects in weighing machine,
height measures, chemical apparatus and other tools may cause undesirable or
error in observation.
Observer and experimental errors are sometimes called as non sampling errors.
Measures of Variability:-Measures of variability of observations help to find how
individual observations are dispersed around the mean of a large series. They may
also be called measures of dispersion, variation or scatter as against the averages
which are measures of central tendency. The following are the important
methods of studying variations:-
(1). Measures of Variability of Individual Observation
- Range
- Interquartile range
- Mean Deviation
- Standard Deviation
- Coefficient of Variation
(2). Measures of Variability of Samples
- Standard error of mean
- Standard error of difference between two means
- Standard error of proportion
- Standard error of difference between two proportions
- Standard error of correlation coefficient
- Standard deviation of regression coefficient
Range:= Range is the simplest method of studying dispersion. It is defined the
difference between the valve of the smallest items and the valves of the largest
items included in the distribution.
Range = L – S, Where S – smallest items, L – largest items
A range defines the normal limits of biological characteristics. Some ranges are
given below:-
Systolic Blood pressure 100—140 mm Hg
Diastolic Blood sugar 80---90 mm Hg
Fasting blood sugar 80---120 mm Hg
Cholesterol 120---150 mg/Dl
Uric Acid 2---4 mg/Dl
Menstrual cycle 21—34 days.
Ordinarily observations falling within a particular range are considered normal
and those falling outside the normal range are considered as abnormal. Reasons
for an observation being lower than the lowest and higher than the highest of the
range.
Merits:- Easy to calculate and easy to understand.
Demerits:- It is not based on all observation.
-It is dependent upon only the two extreme observations.
- Range cannot be computed in case of open end distribution.
Interquartile Range:- One quarter of the observation at the lower and another
quarter of the observation at the upper end of the distribution are enclosed in
computing the interquartile range. In other words, interquartile range represents
the difference between the third quartile and the first quartile.
Interquartile range = Q3 – Q1
Quartile deviation or QD = (Q3 – Q1)/2
Mean Deviation:-- Mean deviation (MD) of observations or measurements from
the mean, ignore the sign, add the difference from the mean and divide by the
number of observations.
MD = ∑ I X – mean I / n
Example :- Weight of LBW baby has given below. Calculate mean deviation for the
valve:- 2.6,2.9,3.0,2.5,3.1.
X X – Mean I X – Mean I

2.6 -0.22 0.22


2.9 +0.08 0.08
3.0 +0.18 0.18
2.5 -0.32 0.32
3.1 -0.28 0.28
14.1 1.08
MD = 1.08/5 =0.216
Standard Deviation:- Standard deviation is also known as root mean square
deviation for the reasons that it is the square root of the mean of the squared
deviation from the arithmatic mean, SD is denoted by the small Greek letter
sigma.
The standard deviation measures the absolute dispersion ( or variability or
distribution; the greater the amount of dispersion or variability), the greater the
standard deviations, the greater will be the magnitude of the deviation of the
valves from their means. A small standard deviation means a high degree of
uniformity of the observation as well as homogeneity of a series.
SD = root ∑ (x – mean)square / n-1
Uses of Standard deviation:- 1. It summarizes the deviation of a large distribution
from mean in one figure used as a unit of variations.
2. Indicates whether the variation of difference of an individual from the mean is
by chance, ie natural or real due to some special reasons.
3. Helps in finding the standard error which determines whether the difference
between means of two similar samples is by chance or real.
4. It also helps in finding the suitable size of sample for valid conclusions.
Merits:- 1. It is very rigid as based on all observations.
2. It is treated further mathematically to have better idea about the type of
distribution and the level of significance.
3. It is not affected much due to sampling fluctuation.
Demerits:- 1. Time consuming if to calculate manually,
2. It is not possible to calculate SD for qualitative and open end classes type
frequency distribution data.
3. Like means, SD also unduly gets affected due to the extreme deviation.
Example:- Find the means respiratory rate per minute and its SD when in 0 cases
the rate was found to be 23,22,20,24,16,17,18,19, & 21.
X – 23,22,20,24,16,17,18,19,21. ∑ X= 180, Mean= 180/9= 20
(X—Mean)=x —3,2,0,4,-4,-3,-2,-1,1
X square – 9,4,0,16,16,9,4,1,1 = 60
SD = Root ∑(X – Mean)square/ n-1 =Root(60)/9-1 =2.74
Example:- Find SD of the erythrocyte sedimentation rate (ESR) found to be
3,4,5,4,2,4,5,3, in 8 normal individuals.
Coefficient of Variations:- CV is used to compare the variability of one character in
two different groups having different magnitude of valves or two characters in the
same group by expressing in percentage. The CV is calculated from SD and mean
of the characteristic. The ratio of SD and means is found in percentage. This SD
expressed as percentage of mean is the coefficient of variation
CV – SD *100/Mean
Example: - Mean blood sugar among housewives was 105 mg/dl & SD was 10. In
the same group of housewives mean body weight was 75 kg and SD was 5 kg. Find
the level of variation (a). CV of blood sugar = 10*100/105 = 9.52%
(b). CV of blood weight =5*100/75 =6.66%
Now the ratio = 9.52/6.66= 1.4
It means blood sugar is 1.4 times more variable character than body weight.
Normal Distribution: - One of the most common distributions in statistics that has
the widest application is normal distribution. This is due to its proportion that is
most desirable making the distribution versatile for further use.
The normal distribution is one of the most important continuous
theoretical distributions in statistics. Unlike experimental distribution, it can be
used for variables that include negative valves. The probability density function
for a variable “x” following normal distribution is given as:
f(x)=
x = valves of the continuous random variables

Graph of Normal Distribution:-


1. The normal distribution can have different shapes depending on different
valves of mean and sigma but there is one and only one normal distribution for
any given pair of valves for mean and sigma.
2. Normal distribution is a limiting case of binomial distribution when n- and
neither p nor q is very small.
3. The mean of a normally distributed population lies at the centre of its normal
curve..
4. Normal distribution is a limiting case of Poisson distribution when its mean is
large.
5. The two tails of the normal probability distribution extend infinitely and never
touch the horizontal axis.
Importance of Normal Distribution:- 1. The normal distribution has the
remarkable property stated in the so called central limit theorem.
2. As n become large normal distribution serves as a good approximation of many
discrete distributions.
3. In theoretical statistics many problems can be solved only under the
assumption of a normal distribution.
4. It is easy to manipulate.
5. It is used extensively in statistical quality control in industry in setting up of
central limits.
Properties of the Normal Distribution:- The normal curve is bell shaped and
symmetrical.
2. The height of the normal curve is at its maximum at the mean.
3. In normal distribution mean, median and mode are all equals.
4. No portion of curve lies between the X-axis the probability can be never
negative.
5. There is only one max. Point, the normal curve is unimodal, ie. It has only one
mode.
6. The point of inflexion, ie., the point where the change in curvature occurs are
mean plus minus sigma.
7. The first and third quartile are equidistant from the median.
8. The area under the normal curve distributed as follows:
(a). Means plus minus sigma covers 68.27% area ; 34.135%, area will lie on
either side of the mean.
(b). Mean plus minus 2 sigma covers 95.45%
( c). Means plus minus 3 sigma covers 99.73%.

Standard Normal Distribution:- A normal cure with zero mean and unit SD is
known as the Standard Normal Curve. Any variable following a normal
distribution can be standardized to the standard normal distribution by
subtracting its mean from every valve and dividing this difference by its standard
deviation. In the following expression X N ( ) IS TRANSFORMED TO z n( 0, 1)
As the proportion of normal distribution suggests, the normal distribution is
completely characterized by its mean and standard deviation. The transformation
from normal to standard normal distribution is often known as Z- transformation.

Sampling Methods
When a large proportion of individuals or items or units have to be studied, we
take a sample. Sampling is the process of selecting observations (a sample) to
provide an adequate description and inference of the population. The following
point are essentials of sampling:-
A sample should be selected that it truly represents the universe otherwise
results obtained may be misleading.
The size of sample should be adequate.
Independence – All items of the universe should have the same chance of being
selected in the sample.
Homogeneity – There is no basic difference in the nature of units of the universe.
Sample design process
Define Population
Determine sample frame
Determine sample procedure
Probability Sampling Non Probability Sampling
Simple random sample Convenient
Stratified sampling Judgement
Cluster sampling Quota
Systematic sampling Snowball
Multistage sampling Self selecting sampling
Determine appropriate sample size
Execute sampling design
Methods of Sampling
(a). Non-Probability Sampling:- The probability of each case being selected from
the total population is not known.
Units of sample are chosen on the basic of personal judgment or convenience.
There are no statistical techniques for measuring random sampling error in a non-
probability sample.
The major problem with these methods is the difficulty in generalizing the study
sample to the whole population. Researchers should always try to avoid these
sample methods.
(1). Convenience Sampling:- The researcher selects a sample by choosing these
who are easy to select.
The subjects are either the easiest to select or they are most likely to respond in
the study.
It is common method for selecting participants in a focus group discussion.
Advantage:- Very low costs, extensively used.
Disadvantages:- Variability and bias cannot be measured or controlled.
Projecting data beyond sample not justified.
Restriction of generalization.
(2). Quota Sampling:- The population is first segmented into mutually exclusive
subgroups, just as in stratified sampling.
Then convenience or judgment sampling is used to select the required numbers of
subjects from each stratum.
Advantages:- Used when research budget is limited.
Very extensively used.
No need for list of population elements.
Disadvantages:- Variability and bias cannot be measured/ controlled.
Time consuming.
Projecting data beyond sample not justified.
(3). Judgement Sampling:- The researcher selects a sample deliberately or
purposely on the basis of his/ her own judgement. For example, representative
block, even though the population includes several blocks in a city.
Advantages:- There is a assurance of quality response. Meet the specific objective.
Disadvantages:- Bias selection of sample may occur. Time consuming process.
(4). Snowball Sampling:- The research starts with a key person and introduce the
next one to become a chain.
Who meet the criteria for inclusion
It is purely based on referrals and that is how a researcher is able to generate
sample.
It is extensively used where a population is unknown and rare.
Advantages:- Cost effective.
Its quicker to find sample.
Useful in specific circumstance for looking rare population.
Disadvantages:- Sampling bias and margin of error.
Lack of corporation.
Projecting data beyoud sample not justified.
(5). Self Selection Sampling:- It occurs when you allow each case usually
individuals to identify their desire to take part in the research.
Advantage:- More accurate.
Useful in specific circumstance to serve the purpose.
Disadvantages:- More costly due to advertizing.
(b). Probability Sampling:- A probability sampling is one in which each members of
the population has an equal chance of being selected.
(1). Simple Random Sampling:- The every units of the population has an equal
chance of being selected. It is used in experimental medicine or clinical trial like
testing the efficacy of a particular drug. The first step is to draw up a list of all the
individual in the population. This is called sampling from. The second step to
decide on the size of of the sample. Thirdly, the serial no in the list are written on
small pieces of paper and placed in the box. Numbers are then picked up from the
box until the required total is reached. This is called lottery method. A more
convenient way to do the random selection is to use the random number,
Advantages:- If the population is homogeneous, it is likely to be more
representative.
Easy to analyses data.
Disadvantages:- Low frequency of use.
Larger risk of random errors.
(2). Systematic random sampling:- The individuals are random choosen at regular
interval from the population systematically rather than randomly, the starting
point being chosen at random.
Order all units in the sample fram.
Then every nth no on the list is selected
K=N/n
K= sampling interval, N= Universe size, n= sample size.
Example:- Suppose 100 student are these. It is desire to take sample of 10 student
K= N/n = 100/10= 10
Suppose the first student come out to be 4th. The sample should be
4,14,24,34,44,54,64,74,84,94.
Advantages:- Simple to draw sample. Easy to verify.
Disadvantages:- Periodic ordering required.
(3). Stratified random sampling:- The entire population is divided into certain
homogeneous sub-group ( or strata) depending upon the characteristics to be
studied (the basis for status stratification being age-group, sex-group, area-wise,
socio-economic status etc.)
The sample is drawn from each stratum at simple random sampling in population
of its size.
Advantages:- It gives great accuracy.
Assure representation of all groups in sample population.
No sub sample is less than 30 in size.
Characteristic of each stratum can be estimated and comparison made.
Disadvantages:- If the stratification is not done properly, the sample may be
biased.
A list of all the units within each stratum is required
Stratified list costly to prepare.
(4). Cluster sampling:- The population is divided into sub-group (cluster) such as
families in village, villages in a district, schools and wards of a city etc. A sample of
cluster proportionate to their size is randomly drawn.
All houses with population are numbered
A simple random sample is taken from each cluster.
Calculate cumulative population and divide the same by 30. This give the sample
interval.
Advantages:- Only a lit of units is needed in the selected cluster.
It is cost effective.
Disadvantages:- It is less precise than simple random sample
Each stage in cluster sampling introduce sampling error, the more stages there
are the more error there.
(5). Multistage sampling:- With large population, it is necessary to carry out the
sampling in several stages. In multistage sampling the no of subjects/ units gets
reduced in every succeeding phase, thereby reducing the magnitude of the
complicated and costly procedure reserved for the last phase. This multistage
sampling procedure makes the studies less expensive, less time consuming, less
laborious and more purposeful.
Size of the sample:- The estimation of the sample size involves the following
factors:-
The approximate idea of the estimate of the characteristic under observation is
required and that is obtained either from previous studies or from Pilot study.
The maximum permissible error that can be allowed should be decided in
advance. If the error is large, than a small sample will serve the purpose and vice
versa.
Higher the probability, bigger the sample size.
Availability of the resources such as men, money and material also determine the
size of the sample.
Qualitative Data:- In a field survey to estimate the prevalence of a particular
disease the sample size is calculated by the formula,
N= 4pq / E square
Where, n= required sample size
P= Approximate prevalence rate of the disease obtain from previous
studies or from pilot study
Q = 1- p
E= Permissible error in the estimate of “p”
Example:- To estimate the prevalence rate of ascariasis in a community, where it
is approximately known to be 40 percent, then the required sample size to
estimate the morbidity with 5 percent error with a probability of 0.05 is
calculated as follows:
N= 4pq /E square, where p= 40%, q = 1-p-= 60%
E = 5% of 40 = 5*40/100= 2
N = 4*40*60 / 2*2 =2400
2400 person are to be examined to estimate the prevalence rate of ascariasis with
5 percent error.
Example:- Hookwarm prevalence rate was 30%. Calculate the size of sample if
allowable error is .03 and .06.
P= 30% = 0.3, q= 1 – p= 1 - 0.3=0.7
N= 4*0.3*0.7/ .03*.03 = 933.3 (934 size)
If E=0.06, n=4*0.3*0.7 /0.06*0.06 = 233.3 (234 size)
Quantitative Data:- If the SD in a population is known
E = 2 Sigma / root n or n = 4 sigma square/ E square
Example:- If mean pulse rate of a population I 70 per minute, SD= 8 beats, size of
sample=
If allowable error E=plus minus beat at 5% risk.
N = 4*8*8 / 1*1 = 256
If E= plus minus2 beat with 5% risk.
N = 4*8*8 / 2*2 = 64.
Example:- In a community survey to estimate the hemoglobin level, from the data
already available if it is known that the mean Hb percent level is about 12g
percent with a standard deviation of 1.5g percent than the sample size required
to estimate the Hb level with a permissible error of 0.5 g percent on either side is
obtained as follows
N = 2*2*1.5*1.5 / 0.5*0.5 = 36 person
Sampling Errors:- If we take repeated samples from the same population or
universe, the results obtained from one sample will differ to some extent from
the results of another sample. This type of variation from one sample to another
is called sampling error. Sampling errors are of two types: biased and unbiased.
(1). Biased Error:- These errors arise from any bias in selection, estimations etc.
For example, if in the place of simple random sampling deliberate sampling has
been used in a particular case some bias is introduced in the result and hence
such errors are called biased sampling errors.
(2). Unbiased Error:- These error arise due to chance difference between the
members of population included in the sample and those not included.
Causes of Bias:- Bias may arise due to:-
Faulty process of selection.
Faulty work during the collection.
Faulty method of analysis.
Method of Reducing Sampling Error:-
Specific problem solutions
Systematic documentation of related research.
Effective enumeration survey in case of a sample survey
Effective pretesting
Controlling methodological bias.
Selection of appropriate sampling techniques.
Non- sampling Error:-Non sampling errors can occure at every stage of planning
and execution of the survey. Such error can arise due to no of causes such as
defective methods of data collection and tabulation, faulty definition, incomplete
coverage of the population or sample [Link] sampling error may arise from one
or more of the following factors:-
Data specification being inadequate and inconsistent with respect to the objective
of the censes or survey.
Inappropriate statistical unit.
In accurate or inappropriate methods of interview observation or measurement.
Lack of trained and experienced investigations.
Lack of adequate inspiration and supervision of primary staff.
Error due to non-response
Error in data processing such as coding, punching, verification etc. Non sampling
error may arise due to defective frame and faulty selection of sampling units.
SIGNIFICANCE OF DIFFERENCE OF MEAN
After making experiment in medical problem, certain results like means and
populations are obtained which vary from sample to sample and ample to
universe.
The difference observed is expressed in terms of significance or probability or
relative frequency of its occurrence by chance and it stated on the basis of
sampling distribution. There are two basic methods of drawing the conclusion or
knowing the significance of the result obtained
(1). The estimation of a population parameter from a sample statistic such as
Mean.
(2). The testing of hypothesis about the population parameter (u).
We then set up certain limits on both sides of the population mean (u) on the
basis of the fact that mean (X) of sample size 30 or more one normally distributed
around the population mean (u). These limits are called the confidence limits and
the range between the two is called the confidence interval.
The limit of the region at which we no longer regard the chance to be
operating is called the level of significance.
If the chance limit is set at mean plus minus 1.96 SE, it implies 5% or 0.05 level
of significance, also called the critical level of significance. A valve lying beyond
this area is said to be significantly different from the population valve such a
different valve will be found by chance only 5 times in 100 results (P, 0.05).
If the linc is drawn at a distance of 2.58 SE from the mean, its level of
significance is said to be 1%. A valve lying beyond this area is said to be highly
significant because such a different valve will be found by chance only one in 100
results (P, 0.01).
In most the statistical studies, the level of significance are set at 5%, (P, 0.05),
1% (p, 0.01) and 0.5% (p, 0.005). Significant or insignificant indicate whether a
valve is likely to occur by chance or it is unlikely to occur by chance.
Statistically Hypothesis:-- A Hypothesis is a supposition made as a basis for
reasoning. According to Prof. Moris Hambarg, “A Hypothesis in statistics is simply
a quantitative statement about a population”. The two hypothesis in a statistical
test are normally referred to is (1). Null hypothesis, (2). Alternative hypothesis
(1). A null hypothesis or hypothesis of no difference (Ho) between statistic of a
sample and parameter of population or between statistic of two samples. This
hypothesis nullifies the claim that the experimental result is different from or
better than the one observed already. For example, If we want to find out
whether a particular drug is effective in curing malaria. We will take te null
hypothesis that the drug is not effective in curing malaria. The rejection of null
hypothesis indicate that the difference have statistical significance and the
acceptance of null hypothesis indicate that the difference are due to chance.
(2). The Alternative hypothesis of significant difference (Ho) stating that the
sample result in different greater or smaller than the hypothetical valve of
population e.g., weight gain or less due to new fading regimen.
To make minimum error in rejection or acceptance of Ho. We divide ±±
±sampling distribution or the area under the normal curve into two region or
zones (1). A sum of acceptance, (2). A zone of rejection
(Diagram)

(1). Zone of acceptance: -- If the result of a sample falls in the plain area, i.e.,
within the mean± 1.96 SE the null hypothesis is accepted, hence this area is called
the zone of acceptance for null hypothesis.
(2). Zone of Rejection:-- I the result of a sample falls in the shaded area, i.e.,
beyond mean±1.96 SE it is significantly different from the universe valve. Hence,
the Ho of no difference is rejected and alternative Ha is accepted. This shaded
area is called the zone of rejection of null hypothesis.
Type- 1 Error (α- Error):-- Consider a clinical trail comparing a new drug against a
standard drug in recovery from a disease. The alternative hypothesis is that the
new drug is more effective in comparision to the standard drug, whereas the null
hypothesis is that the two drugs are equally effective. It is possible to obtain
better recovery in the sample receiving the new drug in comparison to the sample
receiving the standard drug even when the drugs are indeed equally effective in
the population. This may give rise to a false- positive finding. This erroneous
conclusion of falsely rejecting the null hypothesis is called the type-1 error or α-
error.
Type-II Error (β- Error):-- Considering the same example, if the new drug is
actually more effective than the standard drug, it is still possible to obtain
samples that show no evidence of difference between the two groups due to
more sampling variation. This may lead to a false- negative finding. This erroneous
conclusion of not rejecting a false null hypothesis is called type-II error or β-error.
When a statistical hypothesis is tested these are four possibilities:-
Hypothesis Ho is true and test accepts it because the result falls within the zone
of acceptance at 5% level.
Hypothesis Ho is false and our test rejects in because the estimate falls in shaded
area of rejection.
Hypothesis Ho is true still it is rejected, the estimate falls in acceptance zone at
5% level in plain area (Type I Error).
Hypothesis Ho is false but it is accepted, the estimate falls in the zone of rejection
(Type II Error).
Z – Test:-- Z-test is applied to the sampling variability, the difference observed
between a sample estimate and that of population is expressed in terms of SE
instead of SD. The score of valve of the ratio between the observed difference
and SE is called “Z-test”.
If the Z score falls within mean±1.96 SE i.e. in the zone of acceptance (95%
confidence limits) the Ho is accepted. If Ho is rejected is called the level of
significance.
The Z=test for mean has two application:
(1). To test the significance of difference between a sample (X) mean and a know
valve of population (µ)
Observed difference between sample Z = Mean(X) – Population mean(µ)
SE of sample mean
(2). To test the significance of difference between two sample mean or between
experiment sample mean and a control sample mean.
Z= observed difference between two sample mean
SE of difference between two sample means
The four prerequisites to apply Z- test for mean are:-
Sample must be randomly selected.
The data must be quantitative.
The variable is assumed to follow normal distribution in the population.
The sample size must be larger than 30.
If SD of population is known, Z-test can still be applied even if the sample is
smaller than 30.
The significance of Z valve, the probability (p) valve is found from the following
table, constructed on the basis of normal distribution.
Z : 1.6(1.65) , 2.0(1.96) , 2.3(2.2) , 2.6(2.58)
P : 0.10 , 0.05 , 0.02 , 0.01
If the Z valve increases, the P valve or probability of an event happening by
chance decreases and alternative happening due to same external factor is to be
considered.
Two tailed test: -- A two tailed test of hypothesis will reject the null hypothesis, of
the sample statistic is significantly higher than or lower than the hypothesized
population parameter. In two tail test rejection region is located in both the tails.
If we are testing a hypothesis at 5% level of significance, the size of the
acceptance region on each sides of the mean would be 0.475 and the size of
rejection region is 0.025. If the sample mean falls into this area, the hypothesis is
accepted. If the sample mean falls in the area beyond 1.96 SE, the hypothesis is
rejected because it falls into te rejection region. For example, when you want to
know the action of a particular drug is different from that of another,it will be two
tailed test.
One tailed test: -- In case of one tail test the rejection region will be located in
only one tail which may be located in only one tail which may be either left or
right depending upon the alternative hypothesis. For example, if we know
whether one particular drug is better than the other, it will be one tailed test.

Standard Error of the Mean (SE X):--The sample estimate of statistics (X, s or p)
will differ from population parameter (µ,σ or p) because of chance or biological
variability. Such a difference between sample and population valves is measured
by statistic know as sampling error or standard error.
Standard error is thus a measure of chance variation and it does not mean
error or mistake.
SE X = σ/ √n , σ= standard deviation of population
n= no of observation in the sample
Uses of SE X
TO find the confidence limits of population mean. If standard deviation of the
sample is known.
Whether the sample is drawn from a known population or not when (µ) its mean
is known. If the sample mean is larger than the known population mean 1.96 SE,
95% chance are that the sample is not drawn from the same population or it is
under the influence of same other factor.
(a). Z = (X - µ)/σ/√n , when µ and σ or known
(b). Z = (X -µ)/σ/√n , when µ is known but s is not known.
To find the SE of difference between two mean to know if the observed difference
between the means of two samples is real and statistically significant or it is
apparent and insignificant due to chance.
(Example page 170 & 171 mahajan)
Example :-- Mean & SD for the height of 50 boys were 150 and 7cm respectively.
Find SE of mean and 95 percent confidence limits of heights in nature. Could this
sample be from the universe with a population mean (m) of 154 cm.
Mean height= 150 cm , SD = 7 cm , n = 50 boys
SE X = 7/ √50 = 7/7.07 = 0.98
95% confidence limits in nature will be : mean height± 2 SE
150 + 2*0.98 =151.96 AND 150 – 2*0.98 = 148.04
The sample of 50 boys with a mean height of 150 cm is not drawn from the
universe with a population mean of 154 cm, because 154 cm is much more than
95% of confidence limit of 151.96 cm
Z= (observed – mean) / SE X = (154 – 150)/ 0.98 =4.08
4.08 is more than 4 times of the SE.
Since the ratio is more than 3 ( i.e. beyond 3 SD), it is considered to be highly
significant i.e., beyond 99% confident limit.
Standard Error of Difference between two mean of large samples SE (X1- X2):---
If independent, large and random samples are drawn in pairs, repeatedly from
the same population and each time difference between the two mean of each
pair is calculated, the mean of the differences in the paired samples would be
zero or nearly so.
Moreover, the differences in means of the different sets of twin samples will
follow normal frequency distribution curve. The SD of such a distribution of
difference is known as standard error of difference between two means. In
practice it is not possible to find difference of large no of samples and than find SE
of these difference. The test is applied to one pair directly if SD of two mean are
known.
(a). Two independent sample from the same population of SD
SE of the difference between mean = √SD² (1/n₁ + 1/n₂)
= SD √ (1/n₁ + 1/n₂)
The action of two different drug or two different doses of the drug can be
compared and as this test based on normal distribution sample should be large.
(b). If two random sample from different population
SE (X₁ - X₂) = √ SD₁²/n₁ + SD₂²/n₂
Z = (X₁ - X₂)/ SE (X₁-X₂)
EXAMPLE PAGE 175 MAHAJAN, 718.
Example:-- Pregnant women attaining an anganwadi and receiving nutritional
supplementation numbering 49 were match with 64 pregnant women not
attaining anganwadi and not getting the supplementation. Both groups were
followed up. The mean birth-weight of babies born to the farmer was 3.5 kg with
SD 1.4 kg and that of those born to the better, 3.0 kg with SD 1.6 KG. Is the higher
birth weight in the farmer due to the nutrition supplementation.
Where SD₁²= 1.4 kg, SD₂² = 1.6 kg , n₁= 49 , n₂= 64, X₁= 3.5 kg, X₂= 3.0 kg
SEDM = √ SD₁²/n₁ + SD₂²/n₂ = √ (1.4)²/49 + (1.6)²/64
= √(1.96/49) + (2.56/64) = √.04 + .04 =0.28
Z = (X₁-X₂)/ SE(X₁-X₂) = (3.5-3.0) / 0.28 = 1.78
Since the observed difference is less than 2 times the SE (within 95% confidence
limits) it is not significant. That means the difference in mean both weight is due
to chance and not due to nutritional supplementation.
Example:- A simple sample of the height of 6400 Englishman has a mean of 67.85
inches, SD of 2.56 inches. While a simple sample of height of 1600 Austrian has a
mean of 68.55 & SD of 2.52 inches. Do the data indicate that the Austrian are on
the average taller than the Englishmen. Give reasons?
Let us take the hypothesis that there is no significance difference in the mean
of Englishman and Austrian.
SE(X₁-X₂) = √σ₁²/n₁ +σ₂²/n₂ = √(2.56)²/6400 + (2.52)²/1600 = 0.0707
t = I 67.85 – 68.55I/ 0.0707 =9.9
Since the difference is more than SE (1% Level of significance), the hypothesis is
rejected. Hence the data indicate that the Austrian are on average taller than the
Englishman. Significance of Difference Between Mean of small sample by
Student’s t-test:
The ratio of observed difference between two means of small sample to the
SE of difference in the same is divided by letter ‘t’.
The ‘t’- statistics is defined as: t= (X-µ)*√n / s
Where s = √ ∑(X- Mean)²/ n-1
X = the mean of the sample
µ = the actual or hypothetical mean of the population
n = sample size
s = SD of the sample
If the calculated ‘t’ valve exceeds the valve given under p=0.05 in the table, it is
said to be significant at 5% level and null hypothesis Ho is rejected and alternative
hypothesis (Ha) is accepted.
Criteria for applying t-test: --
Random samples
Quantitative data
Variable normally distributed
Population SD is known
The variable t-distribution ranges from minus infinity to plus infinity
T-distribution is symmetrical and has a mean zero, like the standard normal
distribution.
Unpaired t-test (Independent sample):-- This test is applied to unpaired data of
independent observation made on two different group or separate group or
sample drawn from two population, like control group and treated (experimental)
group and their means are compared for their significant difference, it is known as
‘Unpaired comparisons.
t = (X₁ - X₂)/ SE , SE= SD√(1/n₁ + 1/n₂)
SD calculated by the following formula:
SD= √∑(X₁-X₁)²+∑(X₂-X₂)² /(n₁+n₂-2)
Where X₁= mean of the first sample
X₂= mean of the second sample
n₁= no of observation in the first sample
n₂= no of observation in the second sample
SD= combined standard deviation
EXAMPLE PAGE 180 MAHAJAN
Example:-- Two type of drug used on 5 and 7 patients for reducing their weight:
Drug A was imported and drug B indigenous. The decrease in the weight after
using the drugs for six months was as follows.
Drug A : 10,12,13,11,14
Drug B : 8,9,12,14,15,10,9
Is there is significant difference in the efficacy of the two drugs? If not, which drug
should you buy? (v=10,t₀.₀₅=2.223)
Let us take the hypothesis that there is no significant difference in efficacy of
the two drug. Applying t-test
t =(X₁-X₂)/s √(n₁n₂)/n₁+n₂
X₁ (X₁-X₁) (X₁-X₁)² X₂ (X₂-X₂) (X₂-X₂)²
10 -2 4 8 -3 9
12 0 0 9 -2 4
13 1 1 12 1 1
11 1 1 14 3 9
14 2 4 15 4 4
10 -1 1
9 -2 4
∑X₁= 60 ∑(X₁-X₁)=10 ∑X₂=77 ∑(X₂-
X₂)=44
X₁= ∑X₁/n = 60/5 = 12 X₂ = ∑X₂/n = 77/7 = 11 =
SD = √∑(X₁-X ₁)²+∑(X₂-X₂)² /(n₁+n₂-2) = √(10+44)/5+7-2 = √54/10 =2.324
t = (12-11)/2.324 √5*7/5+7 = 0.735
v =n₁+n₂-2 = 5+7-2 =10 t₀.₀₅=2.223
Calculated valve is less than the table valve, the hypothesis is accepted. Hence
there is no significance in the efficacy of two drug. Since Drug B is indigenous and
there is no difference in the efficacy of imported and indigenous drug, we should
buy indigenous drug B.
Example :-- For a random sample of 10 person, fed on diet it, the increased weight
in pounds in a certain period were: 10,6,16,17,13,12,8,14,15,9.
For another random sample of 12 persons, fed on diet B, the increase in the same
period were: 7,13,22,15,12,14,18,8,21,23,10,17.
Test whether the diets A and B differ significantly as regards their effect an
increase in weight. D.f is 20 valve of t at 5% level 2.09.
Let us take the null hypothesis that A and B do not differ significantly weight
regards to their effect on increase in weight. Applying t-test
t =(X₁-X₂)/s √(n₁n₂)/n₁+n₂ SD=
√∑(X₁-X₁)²+∑(X₂-X₂)² /(n₁+n₂-2)
X₁ (X₁-X₁) (X₁-X₁)² X₂ (X₂-X₂) (X₂-X₂)²
10 -2 4 7 -8 64
6 -6 36 13 -2 4
16 4 16 22 7 49
17 5 25 15 0 0
13 1 1 12 -3 9
12 0 0 14 -1 1
8 -4 16 18 3 9
14 2 4 8 -7 49
15 3 9 21 6 36
9 -3 9 23 8 64
10 -5 25
17 2 4
∑X₁= 120 ∑(X₁-X₁) 120 ∑X₂=180 ∑(X₂-X₂)=314
Mean increase in weight of 10 person fed on diet A, X₁= 120/10= 12 person
Mean increase in weight of 12 person fed on diet B, X₂= 180/12= 15 person
SD =√(120+314)/10+12-2 = 4.46
t= (12-15)/4.66 √10*12/10+12 =3*2.34/4.66 = 1.51
For v=20, the table valve of t at 5% level is 2.09. The calculated valve is less than
the table valve and hence the experiment provides no evidence against the
hypothesis. We, therefore, conclude that diets A and B do not differ significantly
as regards their effect on increase in weight is concerned.
Paired t-test:-- In the test it was assumed that the two sample are said to be
dependent when the in one sample are related to there in the other in any
significant or meaningful manner. In fact, the two sample may consist of pair of
observation made on the same object, individuals or, more generally, on the same
population elements. Such type of situation are commonly faced in medical
sciences, as in example below:-
To study the role of a factor or cause when the observations are made before and
after its eg. Of exertion on pulse rate, effect of a drug on BP etc.
To compare the effect of two drugs, given to same individuals in the sample on
two different occasions,eg. Adrenaline and noradrenaline on pulse rate.
To study the comparative accuracy of two different instruments eg. Two types of
sphygmomanometers.
To compare results of two different laboratory techniques, eg. Microfilaria
infection rate by thick smear and concentration techniques; examination of stools
for hookworm over by zinc floatation and concentration methods and so on.
To compare observations made at two different sites in the same body, eg.
Compare temperature between toes and between fingers or in axilla and mouth,
or in mouth and rectum of the same individuals.
The t-test based on paired observation is defined by the following formula:
t = (X-0)*√n /s = X√n /s ,
Where, X = the mean of the difference, s= SD of the difference
S = √∑(X-X)²/n-1 or √∑X² - n(X)²/ n-1, D.f = n-1
EXAMPLE PAGE 187 MAHAJAN
Example:- A drug is given to 10 patients, and the increments in their blood
pressure where recorded to be 3,6,-2,4,-3,4,6,0,0,2. Is it reasonable to belive that
the drug has no effect on change of BP? (5% valve of t for 9 d.f = 2.26).
Let us take the hypothesis that the drug has no effect on change of BP.
Applying the difference test:
t = X√n/s or X/SE
X : 3, 6, -2, 4, -3, 4, 6, 0, 0, 2 = 20 X= ∑X/n =20/10= 2
(X-X): 1, 4, -4, 2, -5, 2, 4, -2, -2, 0 SD= √∑(X-X)²/n-1 = √90/10-1=3.16
(X-X)²: 1, 16, 16, 4, 25, 4, 16, 4, 4, 0 =90 t= 2√10/3.16 =2
V=n-1= 10-1 =9, t₀.₀₅=2.26
The CV of t is less than table valve. The hypothesis is accepted. Hence it is
reasonable to believe that the drug has no effect on change of BP.

F- Test
The object of the F- test is to find out whether the two independent estimates of
population variance differ significantly, or whether the two samples may be
regarded as drawn from the normal populations having the same variance. For
carrying out the test of significance, we calculate the ratio F. F is defined as:
F= S₁²/S₂² Where, S₁²= ∑(X₁ - X₁)²/n₁-1, S₂²= ∑(X₂ - X₂)²/n₂-1
It should ne noted that S₁² is always the larger estimate of variance ie S₁²>S₂²,
v₁=n₁-1, v₂=n₂-1. Where
v₁= degree of freedom for sample having larger variance
v₂= degree of freedom for sample having smaller variance
The calculated valve of F is greater than the table valve than the F ratio is
significant and null hypothesis is rejected. On the other hand, if calculated of F is
less than the table valve the null hypothesis is accepted and it is inferred that
both the samples have come from the population having same variance. F-test is
based on the ratio of two variance, it is known as the Variance ratio test.
Assumption in F-test:
Normality i.e. the valve in each group are normally distributed.
Homogeneity i.e. the variance within each group should be equal for all groups.

Independence of error, it states that the error should be independent for each
valve.
Example:- Two random samples were drawn from two normal population and
their valves are:-
A: 66, 67, 75, 76, 82, 84, 88, 90, 92
B: 64, 66, 74, 78, 82, 85, 87, 92, 93, 95, 97
Test whether the two population have the same variance at the 5% level of
significance (F= 3.36) at 5% level of v₁=10 and v₂=8.
Let us take the hypothesis that the two population have the same variance.
Applying F-test
X₁ (X₁ - X₁) (X₁ - X₁)² X₂ (X₂ - X₂) (X₂ - X₂)²
66 -14 196 64 -19 361
67 -13 169 66 -17 289
75 -5 25 74 -9 81
76 -4 16 78 -5 25
82 2 4 82 -1 1
84 4 16 85 2 4
88 8 64 87 4 16
90 10 100 92 9 81
92 12 144 93 10 100
95 12 144
97 14 196
720 734 913 1298
X₁= 720/9 =80, X₂= 913/11 =83
S₁²= ∑(X₁-X₁)²/n-1 =734/9-1 =91.75, S₂² =∑(X₂-X₂)²/n-1= 1298/11-1=129.8
F=S₁²/S₂² =91.75/129.8 =0.707
The calculated valve of F is less than the table alve. The hypothesis is accepted.
Hence the two population have the same variance.
ANOVA
An analysis of variance could be used to test whether the mean cholesterol level
between three treatment group is equal or different. The aim is to compare
means by several groups using a one-way analysis of variance. One way ANOVA is
appropride when the subgroups to be compared by just one factor, eg. In the
comparison of mean cholesterol by three treatment group. In general, ANOVA is
a techniques used to test how cholesterol level is being influenced by more than
one factor, e.g. the combined effect of age, sex, and treatment.
One-way ANOVA:- In two sample independent t-test, the variability (or difference)
between groups is compared to the variability within groups. In one-way , the
same principle is extended to measure the variability between the groups versus
variability within the groups.
The ANOVA calculation is Shown:=
Total sum of square (SST) = ∑Xᵢ² - (∑Xᵢ)²/n , d.f = n-1
Sum of square between the sample (SSB)=(∑X₁)²/N+(∑X₂)²/N +….+(∑X)²/N –
(∑Xᵢ)²/N
Sum of square within the samples = Total sum of square – Sum of square between
sample
Analysis of Variance ANOVA Table (one-way classification)
Sources of Sum of square V (d.f) Mean square Variance ratio
variance of F
Between SSB v₁=c-1 MSB = SSB / c-
sample 1
Within sample SSW v₂= n-c MSW =
SSW/n-c
Total SST n-1 MSB/MSW

SSB = Sum of square between sample


SSW = Sum of square within sample
SST = Total sum of square of variations
MSB = Mean sum of square between sample
MSW = Mean sum of square within sample
If calculated valve is greater than the table valve, null hypothesis rejected and
alternative hypothesis of significant difference between the mean is accepted. If
calculated valve is smaller than the table valve, null hypothesis is accepted and
the inference would be that the sample are drawn from the same population.
EXAMPLE PAGE 191 MAHAJAN
Example :-- Fasting serum cholesterol level (mg/Dl) in a sample of 15 young
adults in each of normal, IGT (Impaired glucose tolerance) and T2DM (Type2
Diabetes) groups:
Normal: 157.7, 99.5, 194.7, 132.8, 184.2, 133.8, 141, 162.5, 167.2, 147.7, 114.8,
97.7, 162.4, 154.7, 161.7.
IGT : 182, 160.6, 148, 129.5, 163.9, 174.6, 145.3, 145.4, 138.6, 133.4, 192, 141.2,
187.4, 135.7, 144.8.
T2DM : 220.3, 184, 201.1, 173.4, 201.8, 210.6, 195.7, 181.8, 161.4, 169.3, 175,
181.2, 158.5, 201.6, 174.7.
ANS:- Assumption:- The distribution of cholesterol level is normally distributed for
each of the three groups and the variance are equal. The three groups constitute
independent random samples.
Hypothesis: Null hypothesis:- The three group means are equal, ie. (µ₁=µ₂=µ₃)
Alternative hypothesis: - The three group means are not equal.
Computation:-
Total Mean SD
Normal 2212.40 147.49 28.11
IGT 2322.45 154.83 20.60
T2DM 2790.45 186.03 18.31
Grand total 7325.30 162.78

TSS= (157.7)²+(99.5)²+……………..+(201.6)²+(174.7)² - (7325.30)²/45 = 34254.88


SSB = (2212.40)²/15 +(2322.45)²/15 +(2790.45)²/15 – (7325.30)²/45 = 12563.80
Within groups SS = 34254.88 – 12563.80 = 21691.08
Source of Sum of Degree of Mean sum F P
variation square freedom of square
Between 12563.80 2 6281.90 6281.90/516.45= <0.001
groups 12.16
Within 21691.08 42 516.45
groups
Total 34254.88 44
Since the observed valve of F is 12.16 which is greater than the expected valve of
8.25 at 0.001 level of significance of the F distribution with 2,42 d.f, we reject the
null hypothesis and conclude that the mean cholesterol level were significantly
different between three groups (p<0.001).
Two-way ANOVA:- In many situation in which the response variable of interest
may be affected by more than one factor. When it is believe that two
independent factors have an effect on the response variable of interest, it is
possible to design the test so that an analysis of variable can be used to test for
the effect of the two factors simultaneously. Such a test is called a two factor
analysis of variance. In that case, we can test two sets of hypothesis with the
same data at the same time.
In a two-way classification the data are classified according to two different
criteria or factor. In a two-way classification the analysis of variance table takes
the following form:
Source of Sum of square Degree of Mean sum of Ratio of F
variation freedom square
Between SSC c-1 MSC= SSC/c-1 MSC/MSE
Samples
Between SSR r-1 MSR= SSB/r-1 MSR/MSE
Rows
Residual or SSE (c-1) (r-1) MSE= SSE/ (r-
error 1) (c-1)
Total SST n-1

SSC = Sum of square between colums


SSR = Sum of square between rows
SSE = Sum of square due to error
SST = Total sum of squares
SSE = SST –(SSC+SSR)
Total no of degree of freedom= n-1 or cr-1
Where c reforms to no. of columns and r reforms to no. of rows.
If calculated valve of F is greater than the table valve at reassigned level of
significance, the null hypothesis is rejected, otherwise accepted.
Chi-Square test
Categorical data are frequently collected in medical investigations. The categorical
variables assumed might be, for example sex, blood groups, treatment groups,
smoking (yes or no) and diseased (yes or no) and patients survival (dead or alive).
Categorical variables includes the grouped quantitative variables such as material
age groups, preterm or term babies birth weight groups and hypertensive or
normotensive. Categorical Variables usually presented in frequencies or counts.
To test the association between two categorical variables using a chi-square
[Link] chi-square test is an important test amongst the several test of
significance.
It is non-parametric test in which no constant of a population is used. Data do not
follow any specific distribution of any variable and no assumption are made in the
test.
In general the test we use to measure the difference between what is observed
and what is expected according to an assumed hypothesis is called the chi-square
test.
Used to evaluate unpaired/ unrelated samples and proportion.
This test can also be applied to a complex contingency table with several classes
and as such is a very useful test in research work
Condition for the application of χ²-test:-
All items in the sample must be independent.
Total no of items should be large, say at least 50.
No group should contain very few items says less than 10.
Chi-square distribution is not symmetrical.
As the d.f increase, the chi-square curve approaches a normal distribution
It is based on frequencies.
It is profound in analyzing complex and contingency tables.
Contingency table:- When the table is prepared by enumeration of qualitative
data by entering the actual frequencies and if that table represent accurance of
two set of event, that table is called contingency table. It is also called as an
association table.
χ²- Test for Association:-- The χ²- test is used to compare the difference between
the observed counts and the expected counts in a contingency tables. If the
difference are large, than the two variables under the investigation are likely to be
associated with each other. If the null hypothesis of no association is true, the
expected counts, for example, in call ‘a’ is expressed as

Exposure Outcomes
Yes No Total
Yes A b r₁
No C d r₂
Total C₁ C₂ N

E(a) = (CT*RT)/N = C₁*r₁/n,


Where RT= Row total for the row containing the cell
CT= Colum total for the Colum containing the cell
In the same way, the expected counts in the remaining cells are calculated.
The test statistics given by χ² = ∑ (Oᵢ - Eᵢ)²/Eᵢ
Where Oᵢ= Observed events, Eᵢ= Expected events, d.f= (r-1) (c-1).
An alternative method to calculate χ² in a 2*2 table is given below which would be
easier and shortcut are χ² = n(ad –bc)²/r₁*r₂*c₁*c₂
If calculated valve of χ² is greater than table valve of χ², the difference between
theory and observation is considered to be significant on the other hand, the
calculated valve is less than table valve, the difference between theory and
observation is not considered as significant.
Uses of χ²- Test:- It can be used in the following situation:
(1). Test of Independence or Association:- The test of association between
two events in binomial or multinomial samples. It measures the probability of
association between two discrete or qualitative variables. For example, to test the
association between an treatment and outcome of disease, smoking and cancer,
vaccination and immunity, cholesterol and coronary disease, weight and diabetes,
alcohol and gastric ulcer and so on. These are two possibilities, either they are
independent of each other or they are dependent on each others, i.e. associated.
The chi-square test has an add advantages. It can be applied to find association
or relationship between two discrete attributes when there are more than two
classes or groups as happen multinomial sample ie to test the association
between no of cigarettes, equal to 10, more than 10, 11-20, 21-30 and more than
30 smoked per day and the incidence of lung cancer, state of nutrition less than
60%, 61-80% and 81-100% and intelligence quotient and so on.
Example:- In a random sample of 500 mothers who had singleton birth from 2011
to 2015 in a large tertiary hospital, one of our aims is to examine whether
maternal hypertension is a risk factor for preterm delivery. The summery of the
data is presented as follows:
Maternal Gestational Age
hypertension
Preterm Term Total
Yes 7 (25.9%) 20 (74.1%) 27
No 41 (8.7%) 432 (91.3%) 473
Total 48 452 500

Test whether there is an increased risk of preterm birth among pregnant


mothers with hypertension.
ANS:- Data are displayed with a=7, b=20, c=41, d=432.
Is there an association between maternal hypertension and preterm birth? Or
Is gestational hypertension a risk factor for preterm birth?
Hypothesis: H₀ : There is no association between maternal hypertension and
preterm birth.
Hₔ : There is an association between maternal hypertension and
preterm birth.
χ² = ∑(O – E)²/ E, d.f= 1
Expected valve of cell ‘a’ ie E(a)= E(7)= 48*27/500 = 2.6
E(b)= E(20)=452*27/500 =24.4, E(c)=E(41)=48*473/500= 45.4
E(d)= E(432)=452*473/500 = 427.6
Now, (a) – O=7, E=2.6, (b) – O=20, E=24.4, (c) – O=41, E=45.4, (d) –O=432, E=427.6
χ²= (7-2.6)²/2.6 + (20-24.4)²/24.4 + (41-45.4)²/45.4 + (432-427.6)²/427.
= 7.45+0.79+0.43+0.05 = 8.72
The table valve of chi-square distribution is 6.63 with 1 degree of freedom at
0.01 level of significance. The observed chi-square valve of 8.72 is greater than
the table valve of 6.63, at 0.01 level of significance and hence we reject the null
hypothesis. We associate that matters with gestational hypertension have more
frequently delivered a preterm baby than mothers without hypertension.
(2). Test of Homogeneity or Proportions:- The χ²-test of homogeneity is an
extension of the chi-square test of independence. The test are designed to
determine whether two or more independent random sample are drawn from the
same population or from different population. For example, do men or women
have the same proportion of overweight? Or do male or female newborns have
the same prevalence of very low birth weight (LBW)? To compare the frequencies
of two multinomial samples such as no of diabetes and nondiabetes in groups
weighting 40-50 kg, 60-70 kg and more than 70 kg.
Example:-- I n a study on cardiovascular risk factors among young adults from
Japan birth cohort study, the proportion of overweight in men and women are
shown in the table below. The overweight in defined as body index (BMI)≥ 23 and
normal weight as BMI<23
Weight Men Women Total
Overweight 285(24.6%) 299(28.4%) 584
Normal 873(75.4%) 753(71.6%) 1626
Total 1158 1052 2210
Do the proportion of overweight in men and women differ significantly?
ANS:--H₀: The proportion of overweight is equal in men or women.
Hₔ: The proportion of overweight is not equal in men and women.
E(a) = (1158*584) / 2210 = 306, E(b) = (1052*584) / 2210 = 278
E( c) = (1158*1626) / 2210 = 852, E(D)= (1052*1626)/2210 = 774
Weight Men Women Total
O E O E
Overweight 285 306 299 278 584
Normal 873 852 753 774 1626
Total 1158 1052 2210

χ² = (285-306)²/306 + (299-278)²/278 + (873-852)²/852 + (753-774)²/774


= 1.44+1.59+0.52+0.57 = 4.12
Table valve of χ² is 3.84 at 0.05 level of significance with 1 degree of freedom.
The observed valve of chi-square is 4.12 is greater than the table valve 3.84 at
0.05 level of significance with 1 degree of freedom and hence we reject the null
hypothesis. The proportion of overweight is higher in women compared to men.
(3). Test of Goodness of Fit:- We can find whether the observed frequency
distribution fits in a hypothetical or theoretical distribution of a qualitative data.
The χ²- test determines whether the observed frequency distribution differs from
the theoretical distribution by chance or the sample is drawn from a different
population. If the calculated valve of χ²-test is higher than the table valve, it is
significant (null hypothesis rejected) and vice-versa.
Significance of Difference in Proportion
In qualitative data, the characters remain the same while the frequency variations
will be studies. If a sample consists of characters of two attributes (positive or
negative) only such as male and female, rich and poor, vaccinated and
nonvaccinated, died and survived, successes and failures etc, it is a sample of
binomial classification. If the sample is divided into more than two classes such as
blood groups A, B, AB, O or WBCs into polymorphic, lymphocytes, eosinophils etc,
it is said to have a multinomial or polynomial classification.
The proportion of individuals, having a specific character or attributes, in a
binomial distribution, is expressed as ‘p’ either as a fraction of 1 or percentage.
P = No of individuals having a specific character/ Total no in the sample
The remaining proportion of individual having the other (negative) character, is
represented as ‘q’.
P= 1-q in term of fraction of 1 or 100-p in percentage.
Thus, ‘p’ is the probability of occurrence of a positive attribute and ‘q’ is the
probability of occurrence of the negative attribute.
Standard error of proportion may be defined as a measure of variation occurring
by chance between the sample proportion (p) and the population proportions(P)
in a qualitative data. This test of significance is employed to find the efficacy of a
drug or vaccine or line of treatment or a surgical procedure etc.
Standard error of proportion (SEP) = √(p*q)/n
Where, p= Percentage of positive character
q= Percentage of negative character
n= No of observation (size of sample)
The significance of difference is found by relative deviate (Z) test.
Z = Observed difference/ Sep
The binomial confidence limits are follows:-
68% of sample population will lie within the range P±1SE of population
95% of sample proportion will lie within the range P±2SE of population
99% of sample proportion will lie within the range P±3SE of population.
Example: - Find the population proportion of in a sample of 50 children 25 give
history of whooping cough.
Ans :- p =25*100/50 = 50, and q = 100 – 50= 50
SEp = √pq/n = √50*50/ 50 = 7.08
The 95% confidence limits of population proportion will be 50±2*7 i.e. 36 and
64%.

Example:- Polymorph count was 350 out of 500 WBCs. At 95% confidence level,
within limits the population will lie?
Ans:- P= 350*100/500 = 70, q= 100-70 = 30
Sep = √70*30/500 = √4.2 = 2.05
The 95% confidence limits for the population proportion will be 70±2*2 =
66% and 74%.

Example:- The proportion of Example blood group A among Indians is 35%. In a


batch of 100 individuals if it is observed as 30%, what is your conclusion about the
group?
Ans:- SEp = √30*70/100 = √21 = 4.58
Z = p –P/SEp = 35 -30 /4.58 =1.09
The difference is in significance at 95% confidence limits, because 35% is
less than the Indian limit of p+2SE, IE. 30+2*4.58 =39.16.

Example:- Constipation was considered to be a common features as observed in


55% of typhoid cases. In a study of 500 typhoid cases, 25% had constipation. Can
you consider constipation as a common feature of typhoid on this observation?

Ans :- SEp = √25*75 /500 =√3.75 = 1.93%

Z = (55 – 25) /1.93 =15.54

Difference is highly significant as the proportion (55%) lies much beyond


the 99% confidence limits, ie. 25+3*1.93 = 30.79. In other words, constipation as a
common symptom in 55% is a wrong presumption which might have been from a
biased or a small sample.
Example:- What should be the size of the sample of assessing prevalence rate of
diabetics in an urban population where the prevalence was given as 3% in age
above 15 years. Allowable error, E is 10% of positive character with 5% risk.

Ans:- E =10*3/100 = 0.3


With 5% risk (95% confidence limits)
N = 4pq /E² = 4*3*97/0.3*0.3 = 12933 persons

Standard Error of Difference Between two Population:-


SE of difference between two proportions measures the
chance between the sample proportion of paired samples drawn from the same
population or universe.
SE (p₁- p₂)= √(p₁q₁/n₁) + (p₂q₂/n₂) or √PQ (1/n₁ +1/n₂)
Where,
p₁ & p₂ are the estimates of the proportions of the two samples.
q₁ & q₂= (1- p₁) and (1- p₂) respectively.
n₁ & n₂ = no of observation in the two samples.

Observed difference between two sample proportion


Z= -------------------------------------------------------------------------
SE of difference between two proportion

P = p₁ + p₂ /n₁ + n₂, Q= q₁+q₂ / n₁+n₂, Z= p₁ - p₂ /SE


If the p valve is less than 0.05 (95%), null hypothesis is rejected and concluded
that the difference between two sample estimate is significant and vice-versa ie. if
the p valve is more than 0.05, the null hypothesis is accepted and concluded that
the difference between the sample estimate is insignificant.

Example :- In an epidemic of gastroenteritis in an area the no of cases reported in


two population. Consuming water from two different sources were as follow:

Sources of water No of people consuming No of cases


Water from the sources
Tap water 800 35
Hand pump water 2400 120
Total 3200 155
Find out whether the difference in the proportion of cases in the two group is
significant.
Ans:- The null hypothesis in this case is that the difference is significant.
P₁ = 35/800 =0.044, q₁ = 1-P₁ = 1- 0.044 = 0.956, n₁= 800
SEp₁ = √P₁q₁/n₁ = √0.044*0.956 / 800 = √0.0000525 =0.0072

P₂ = 120/2400 = 0.05, q₂= 1- P₂= 1-0.05= 0.95, n₂=2400


SEp₂ = √P₂q₂/n₂ = √0.05*0.95 /2400 =√0.0000197 = 0.0044
Difference between two proportion = P₁ - P₂ =0.044 -0.05 = 0.006
SE of Difference = √ P₁q₁/n₁ + P₂q₂/n₂
= √.0000525 + .0000197 = 0.0085
Z = 0.006 / 0.0085 = 0.706
Calculated valve is less than table valve, the null hypothesis is not rejected
And it is concluded that the difference is insignificant.

Example :- In school A, tonsillectomy had been done in 23 student out of 50 while


in the other school B it was done in 77 out of 350. Find if the difference observed
in two schools is by chance or due to some influence.

In school A, n₁ = 50
p₁ with tonsillectomy = 23 out of 50 = 23*100 /50 = 46%,
q₁ without tonsillectomy = 100 – 46 = 54%
In school B, n₂ = 350, p₂ = 77*100/ 350= 22%, q₂= 100 – 22 = 78%
SE(p₁-p₂) = √(p₁q₁/n₁) + ( p₂q₂/n₂)
= √(46*54/50) + (22*78/350) = 7.39

By another formula
P = 23+77, out of 400 =25%
Q = 27+273, out of 400 = 75%
SE(p₁-p₂) = √PQ (1/n₁ +1/n₂) = √25*75 (1/50 + 1/350) = √600/14 = 6.55
Observed difference 46-22
Z=----------------------------- = --------- = 3.7
SE of difference 6.55

Difference is highly significant because at 99% confidence limits it is more than 3


times the SE. It can happen less than 1 in 100 times by chance. So tonsillectomy is
more common in school A. The factor playing part may be investigated.
Example:- If typhoid mortality in one sample of 100 is 20% and in another sample
of 100 it is 30%, find the standard error of difference between two proportions by
both the formulae. Is the difference in mortality rates significant?

As per formula, p₁ = 20, q₁ = 80, n₁= 100


p₂= 30, q₂ = 70, n₂= 100
SE(p₁ - p₂) = √(20*80/100) + (30*70/100) = √16+21 =6.08

As per another formula, P = (p₁+p₂)*100 /(n₁+n₂)= (20+30)*100 /(100+100)= 25%


Q = (q₁+q₂)*100 /(n₁+n₂)= (80+70)*100 /(100+100)= 75%
SE(p₁-p₂) =√PQ(1/n₁ +1/n₂) = √25*75(1/100 +1/100)
=√25*75*2 /100 =√37.5 =6.12
A and B are very close to each other
Observed difference 30 - 20
Z = ----------------------------- = ----------- = 1.64
SE of difference 6.08

It is less than 1.96, the critical level of significance, hence insignificant at 95%
confidence limits.
CORRELATION ANALYSIS

When two variable or attributes are simultaneously recorded (i.e.,bivariate


data) to know whether are influences the other, i.e, age and blood pressure,
temperature and pulse, height and weight, weight of pregnant mother and weight
of the newborn, etc. The purpose of such study is to know whether change in one
variable is associated with the change in the other variable, then it is assumed
that there is a ‘correlation’ between two variables and the degree of this
correlation is measured in terms of ‘Correlation Coefficient’.
Correlation is the linear relationship or mutuality of association of two
variables; it is independent of the units of measurement. It help us to determine
the degree of relationship between two or more variables, it does not tell us
anything about cause and effect relationship. The correlation varies between -1≤ r
≤ 1.

TYPES OF Correlation:- There are five type of correlation depending on its extent
and direction are as follows:-

(1). Perfect Positive Correlation:- The two variables X and Y are directly
proportional and fully correlated with each other. The correlation coefficient ® =
+1 i.e., both variables rise or fall in the same proportion. Example- height and
weight, age and height, age and weight, temperature and pulse etc. X varies
directly and proportionally to Y.

(2). Perfect Negative Correlation :- The two variable X and Y are inversely
proportional to each other, i.e., when one rises the other falls in the same
proportion and vice-versa. The correlation coefficient (r) = -1. Example-
socioeconomic status and incidence of tuberculosis, income and malnutrition.

(3). Partial Positive Correlation:- In this type, the regression lines incline on one
another and ascent toward right. The nonzero valves of coefficient ( r) lie between
0 and +1, i.e., 0<r<1 such as temperature and pulse rate, age of husband and age
of wife, infant mortality rate and over crowding etc.

(4). Partial Negative Correlation :- In this type, the regression lines incline on
another and ascent toward left. The nonzero valves of coefficient ( r) lie between -
1 and 0, i.e., -1<r<0 such as income and infant mortality rate, age and vital
capacity among adults.

(5). Absolutely No Correlation:- The valve of correlation coefficient is zero,


indicating that no linear relationship exists between two variables X is completely
independent of Y such as height and pulse.

Karl Pearson’s Coefficient of Correlation:- When associated variables are normally


distributed such as height and weight the correlation coefficient is called
Pearson’s correlation coefficient. It is denoted by symbol ‘r’. It is one of the very
few symbols that one used universally for describing the degree of correlation
between two series. The formula for computing Pearsonion r is :

∑xy x=(X-X), y= (Y-Y) , σₓ= SD of series X , σ = SD of series Y


r =---------- , N = Number of pairs of observations
N σₓσᵧ r = Correlation coefficient

r = ∑xy / √∑x²*∑y² , σₓ= √∑x² /N , σᵧ = √∑y² /N

REGRESSION ANALYSIS
The meaning of the term ‘regression’ is the act of returning or going back.
The term regression was first used by Sir Francis Galton in 1877.
(1). Regression is the measure of the average relationship between two or more
variables in terms of the original units of the data
(2). “Regression analysis attempts to establish the nature of the relation-ship
between variables – that is, to study the functional relationship between the
variables and thereby provide a mechanism for prediction, or forecasting.”

It is clear from the above definition that regression analysis is a statistical


device with the help of which we are in a position to estimate (or predict) the
unknown valves of one variables from known valve of another variable. The
variable which is used to predict is called independent variable and the variable
we are trying to predict is called dependent variable. The independent variable is
denoted by X and the dependent variable by Y. The analysis used is called the
simple linear regression analysis – simple because there is only one predictor or
independent variable, and linear because of the assumed linear relationship
between the dependent and the independent variable. The term linear means
that an equation of a straight line of the form Y=a+Bx, where a and b one
constants, is used to describe the average relationship that exists between the
two variables.

Difference between Correlation and Regression Analysis


We cannot say one variable is the cause It possible to study the cause and effect
and other the effect. relationship.
rₓᵧ and rᵧₓ are symmetrical rₓᵧ= rᵧₓ bₓᵧ and bᵧₓ are not symmetrical bₓᵧ≠bᵧₓ
There may be nonsense correlation There is nothing like nonsense
between two variables. correlation.
Independent of change of scale & Independent of change of origin but not
origin. scale.

Regression lines:-- I f we take two variables X and Y, we shall have two regression
Lines as the regression of X on Y and the regression of X on Y.
The regression line of Y on X gives the most probable valves of Y for given valves
of X and the regression line of X on Y gives the most probable valves of X for given
valves of Y.
When there is perfect correlation between the two variables (r=±1), the
regression lines will coincide, i.e., we will have only one line. On the other hand,
when correlation is partial, the lines will be separate and diverge forming an acute
angle at the meeting point of perpendiculars drawn from the means of variables.
Linear the correlation, greater will be the divergence of angle. If the variables are
independent, r is zero and the lines of regression are at right angles i.e., parallel to
OX and OY.

Regression Equation:-- (a).The Regression equation of Y on X is expressed as


follows
:--
Y= a+Bx
In the equation ‘Y’ is a dependent variables i.e., its valve depends on X. ‘X’ is
independent variables, i.e., we can take a given valves of X and compute the valve
of Y.
∑Y = Na + b∑X
∑XY = a∑X + b∑X²
(b). Regression equation of X on Y is expressed as follows:-
X = a + By
∑X = Na + b∑Y
∑XY = a∑Y + b∑Y²

Regression Coefficient :- Regression lines X on Y defined as


X – X= bₓᵧ(Y - Ῡ)
Where bₓᵧ= regression coefficient of X on Y.
Regression line Y on X defined as
Y -- Ῡ= bᵧₓ(X –X)
Where bᵧₓ = regression coefficient of Y on X.

(a). If Correlation coefficient (r ) is already calculated the regression coefficient is


derived as: r* SD of Y series ∑xy ∑ (X – X)(Y - Ῡ)
bᵧₓ=---------------------- = --------- = --------------------
SD of X series ∑x² ∑ (X –X)²

r* SD of X series ∑xy ∑ (X –X)(Y - Ῡ)


bₓᵧ= ---------------------- = --------- = ---------------------
SD of Y series ∑y² ∑(Y - Ῡ)²
(b). If means are not to be calculated
∑XY - ∑X∑Y /n
bᵧₓ= ----------------------
∑X² - (∑X)² /n
∑XY - ∑X∑Y /n
bₓᵧ= -----------------------
∑Y² - (∑Y)² /n

Karl Person Correlation Coefficient


r = √ bᵧₓ* bₓᵧ

Epidemiological Study Design and Analysis

Epi – among, upon ; Demos – people, population ; Logos – science study


Epidemiologic principles and methods are used increasingly in medical
research to investigate the causes of disease and intervention to prevent or
control disease. It is used to set up effective policies and programmes to improve
the public health. Over the years, epidemiology has contributed to change and
increased the understanding of major classes of disease like cancer, stroke and
coronary heart diseases and presented with greater opportunities for prevention
and control.
Epidemiology is defined as a ‘study of the distribution and determinants of
health related states or events in specified populations and application of this
study to control health problems.’ ([Link], 1981)

The meaning of key words need to be explained.

Distribution :- This refers to the pattern of occurrence of disease in the


community with reference to time, place, and person. The study is known as ‘
Descriptive epidemiology’. This helps to study the trend of the disease over the
years (decades) geographical areas and over different population groups. This
study also helps to know the magnitude of the problem, gives a clue about the
etiology, mode of transmission of the disease and also helps to formulate
etiological hypothesis.

Determinants: - The study of determinants of the disease is called analytical


epidemiology. Hence, the focus is an the factor that influence the increase or
decrease of the risk of occurrence of health-related events. These factor could be
classified as demographic, genetic, behavioral, environmental and life style
factors.

You might also like