Understanding Averages in Statistics
Understanding Averages in Statistics
One of the most important objectives of statistical analysis is to get one single
valve that describe the characteristic of the entire mean of unwieldy data. Such a
valve is called the central valve or an average or the expected valve of the
variable.
“Average is an attempt to find one single figure to describe whole of figure”.-
Clark
Objective of Average:- There are two main objectives of average:-
(a). To get single valve that describe the characteristic of the entire group.
(b). Measure of central valve, by reducing the mass of data to one single figure,
enable comparison to be made. Comparison can be made either at a time or over
a period of time.
(a). Arithmatic Mean or Mean:- The most popular and widely used measure of
representing the entire data by one valve is most layman call an “Average” and
the statistician call the arithmetic mean.
- It is the simplest form of measurement of central tendency.
- It is calculated by dividing the sum of all the observations by the number of
observations.
- Mean is denoted by X.
Calculation of Mean:-
Mean = Total or sum of the observations
Numbers of observations
X = ∑(X1+X2+……..+Xn)
N
The mean is calculated by different methods in two types of series, ungrouped
and grouped.
Ungrouped Series:- In such series the number of observations in small and there
are two methods for calculating the means. The choice depends upon the size of
observations in the series.
(a). When the observations are small in size
Example:- Hb of 10 ANC women is as follows: 9.5, 11.5, 12.5, 10, 11.7, 13, 10.8,
12.7, 13.2 and 14.2 mg/dl.
Mean X=
9.5+11.5+12.5+10+11.7+13+10.8+12.7+13.2+14.2__________________________
___________
10
Means Hb = 11.91 mg/dl
(b). When individual observations are large in size:- The arithmetic means can be
calculated by using as an arbitrary origin. When deviations are taken from an
arbitrary origin, the formula for mean is
X = A + _∑ x
N
Where A is assumed means and x is the deviations of items from assumed means
ie.,
x= (X – A)
-------------------------------------------------
1065 15
X =A + _∑x_
N
=150 + 15/7
=150 + 2.1= 152.1
Grouped Series:- When the number of observations is large, the data are
arranged in groups and frequency distribution.
Mean X = ∑ fx / ∑f
Example:- Calculate average daily protein intake of formula from the following
data of 200 female. On the basis of their protein intake we can summarize data in
frequency distribution table as follows:-
class interval no of female (f) midpoint of class interval x fx
15-25 15 20 300
25-35 20 30 600
35-45 50 40 2000
45-55 55 50 2750
55-65 40 60 2400
65-75 15 70 1050
75-85 05 80 400
________ _______
200 9500
Mean= 9500/200 = 47.5
Standard Normal Distribution:- A normal cure with zero mean and unit SD is
known as the Standard Normal Curve. Any variable following a normal
distribution can be standardized to the standard normal distribution by
subtracting its mean from every valve and dividing this difference by its standard
deviation. In the following expression X N ( ) IS TRANSFORMED TO z n( 0, 1)
As the proportion of normal distribution suggests, the normal distribution is
completely characterized by its mean and standard deviation. The transformation
from normal to standard normal distribution is often known as Z- transformation.
Sampling Methods
When a large proportion of individuals or items or units have to be studied, we
take a sample. Sampling is the process of selecting observations (a sample) to
provide an adequate description and inference of the population. The following
point are essentials of sampling:-
A sample should be selected that it truly represents the universe otherwise
results obtained may be misleading.
The size of sample should be adequate.
Independence – All items of the universe should have the same chance of being
selected in the sample.
Homogeneity – There is no basic difference in the nature of units of the universe.
Sample design process
Define Population
Determine sample frame
Determine sample procedure
Probability Sampling Non Probability Sampling
Simple random sample Convenient
Stratified sampling Judgement
Cluster sampling Quota
Systematic sampling Snowball
Multistage sampling Self selecting sampling
Determine appropriate sample size
Execute sampling design
Methods of Sampling
(a). Non-Probability Sampling:- The probability of each case being selected from
the total population is not known.
Units of sample are chosen on the basic of personal judgment or convenience.
There are no statistical techniques for measuring random sampling error in a non-
probability sample.
The major problem with these methods is the difficulty in generalizing the study
sample to the whole population. Researchers should always try to avoid these
sample methods.
(1). Convenience Sampling:- The researcher selects a sample by choosing these
who are easy to select.
The subjects are either the easiest to select or they are most likely to respond in
the study.
It is common method for selecting participants in a focus group discussion.
Advantage:- Very low costs, extensively used.
Disadvantages:- Variability and bias cannot be measured or controlled.
Projecting data beyond sample not justified.
Restriction of generalization.
(2). Quota Sampling:- The population is first segmented into mutually exclusive
subgroups, just as in stratified sampling.
Then convenience or judgment sampling is used to select the required numbers of
subjects from each stratum.
Advantages:- Used when research budget is limited.
Very extensively used.
No need for list of population elements.
Disadvantages:- Variability and bias cannot be measured/ controlled.
Time consuming.
Projecting data beyond sample not justified.
(3). Judgement Sampling:- The researcher selects a sample deliberately or
purposely on the basis of his/ her own judgement. For example, representative
block, even though the population includes several blocks in a city.
Advantages:- There is a assurance of quality response. Meet the specific objective.
Disadvantages:- Bias selection of sample may occur. Time consuming process.
(4). Snowball Sampling:- The research starts with a key person and introduce the
next one to become a chain.
Who meet the criteria for inclusion
It is purely based on referrals and that is how a researcher is able to generate
sample.
It is extensively used where a population is unknown and rare.
Advantages:- Cost effective.
Its quicker to find sample.
Useful in specific circumstance for looking rare population.
Disadvantages:- Sampling bias and margin of error.
Lack of corporation.
Projecting data beyoud sample not justified.
(5). Self Selection Sampling:- It occurs when you allow each case usually
individuals to identify their desire to take part in the research.
Advantage:- More accurate.
Useful in specific circumstance to serve the purpose.
Disadvantages:- More costly due to advertizing.
(b). Probability Sampling:- A probability sampling is one in which each members of
the population has an equal chance of being selected.
(1). Simple Random Sampling:- The every units of the population has an equal
chance of being selected. It is used in experimental medicine or clinical trial like
testing the efficacy of a particular drug. The first step is to draw up a list of all the
individual in the population. This is called sampling from. The second step to
decide on the size of of the sample. Thirdly, the serial no in the list are written on
small pieces of paper and placed in the box. Numbers are then picked up from the
box until the required total is reached. This is called lottery method. A more
convenient way to do the random selection is to use the random number,
Advantages:- If the population is homogeneous, it is likely to be more
representative.
Easy to analyses data.
Disadvantages:- Low frequency of use.
Larger risk of random errors.
(2). Systematic random sampling:- The individuals are random choosen at regular
interval from the population systematically rather than randomly, the starting
point being chosen at random.
Order all units in the sample fram.
Then every nth no on the list is selected
K=N/n
K= sampling interval, N= Universe size, n= sample size.
Example:- Suppose 100 student are these. It is desire to take sample of 10 student
K= N/n = 100/10= 10
Suppose the first student come out to be 4th. The sample should be
4,14,24,34,44,54,64,74,84,94.
Advantages:- Simple to draw sample. Easy to verify.
Disadvantages:- Periodic ordering required.
(3). Stratified random sampling:- The entire population is divided into certain
homogeneous sub-group ( or strata) depending upon the characteristics to be
studied (the basis for status stratification being age-group, sex-group, area-wise,
socio-economic status etc.)
The sample is drawn from each stratum at simple random sampling in population
of its size.
Advantages:- It gives great accuracy.
Assure representation of all groups in sample population.
No sub sample is less than 30 in size.
Characteristic of each stratum can be estimated and comparison made.
Disadvantages:- If the stratification is not done properly, the sample may be
biased.
A list of all the units within each stratum is required
Stratified list costly to prepare.
(4). Cluster sampling:- The population is divided into sub-group (cluster) such as
families in village, villages in a district, schools and wards of a city etc. A sample of
cluster proportionate to their size is randomly drawn.
All houses with population are numbered
A simple random sample is taken from each cluster.
Calculate cumulative population and divide the same by 30. This give the sample
interval.
Advantages:- Only a lit of units is needed in the selected cluster.
It is cost effective.
Disadvantages:- It is less precise than simple random sample
Each stage in cluster sampling introduce sampling error, the more stages there
are the more error there.
(5). Multistage sampling:- With large population, it is necessary to carry out the
sampling in several stages. In multistage sampling the no of subjects/ units gets
reduced in every succeeding phase, thereby reducing the magnitude of the
complicated and costly procedure reserved for the last phase. This multistage
sampling procedure makes the studies less expensive, less time consuming, less
laborious and more purposeful.
Size of the sample:- The estimation of the sample size involves the following
factors:-
The approximate idea of the estimate of the characteristic under observation is
required and that is obtained either from previous studies or from Pilot study.
The maximum permissible error that can be allowed should be decided in
advance. If the error is large, than a small sample will serve the purpose and vice
versa.
Higher the probability, bigger the sample size.
Availability of the resources such as men, money and material also determine the
size of the sample.
Qualitative Data:- In a field survey to estimate the prevalence of a particular
disease the sample size is calculated by the formula,
N= 4pq / E square
Where, n= required sample size
P= Approximate prevalence rate of the disease obtain from previous
studies or from pilot study
Q = 1- p
E= Permissible error in the estimate of “p”
Example:- To estimate the prevalence rate of ascariasis in a community, where it
is approximately known to be 40 percent, then the required sample size to
estimate the morbidity with 5 percent error with a probability of 0.05 is
calculated as follows:
N= 4pq /E square, where p= 40%, q = 1-p-= 60%
E = 5% of 40 = 5*40/100= 2
N = 4*40*60 / 2*2 =2400
2400 person are to be examined to estimate the prevalence rate of ascariasis with
5 percent error.
Example:- Hookwarm prevalence rate was 30%. Calculate the size of sample if
allowable error is .03 and .06.
P= 30% = 0.3, q= 1 – p= 1 - 0.3=0.7
N= 4*0.3*0.7/ .03*.03 = 933.3 (934 size)
If E=0.06, n=4*0.3*0.7 /0.06*0.06 = 233.3 (234 size)
Quantitative Data:- If the SD in a population is known
E = 2 Sigma / root n or n = 4 sigma square/ E square
Example:- If mean pulse rate of a population I 70 per minute, SD= 8 beats, size of
sample=
If allowable error E=plus minus beat at 5% risk.
N = 4*8*8 / 1*1 = 256
If E= plus minus2 beat with 5% risk.
N = 4*8*8 / 2*2 = 64.
Example:- In a community survey to estimate the hemoglobin level, from the data
already available if it is known that the mean Hb percent level is about 12g
percent with a standard deviation of 1.5g percent than the sample size required
to estimate the Hb level with a permissible error of 0.5 g percent on either side is
obtained as follows
N = 2*2*1.5*1.5 / 0.5*0.5 = 36 person
Sampling Errors:- If we take repeated samples from the same population or
universe, the results obtained from one sample will differ to some extent from
the results of another sample. This type of variation from one sample to another
is called sampling error. Sampling errors are of two types: biased and unbiased.
(1). Biased Error:- These errors arise from any bias in selection, estimations etc.
For example, if in the place of simple random sampling deliberate sampling has
been used in a particular case some bias is introduced in the result and hence
such errors are called biased sampling errors.
(2). Unbiased Error:- These error arise due to chance difference between the
members of population included in the sample and those not included.
Causes of Bias:- Bias may arise due to:-
Faulty process of selection.
Faulty work during the collection.
Faulty method of analysis.
Method of Reducing Sampling Error:-
Specific problem solutions
Systematic documentation of related research.
Effective enumeration survey in case of a sample survey
Effective pretesting
Controlling methodological bias.
Selection of appropriate sampling techniques.
Non- sampling Error:-Non sampling errors can occure at every stage of planning
and execution of the survey. Such error can arise due to no of causes such as
defective methods of data collection and tabulation, faulty definition, incomplete
coverage of the population or sample [Link] sampling error may arise from one
or more of the following factors:-
Data specification being inadequate and inconsistent with respect to the objective
of the censes or survey.
Inappropriate statistical unit.
In accurate or inappropriate methods of interview observation or measurement.
Lack of trained and experienced investigations.
Lack of adequate inspiration and supervision of primary staff.
Error due to non-response
Error in data processing such as coding, punching, verification etc. Non sampling
error may arise due to defective frame and faulty selection of sampling units.
SIGNIFICANCE OF DIFFERENCE OF MEAN
After making experiment in medical problem, certain results like means and
populations are obtained which vary from sample to sample and ample to
universe.
The difference observed is expressed in terms of significance or probability or
relative frequency of its occurrence by chance and it stated on the basis of
sampling distribution. There are two basic methods of drawing the conclusion or
knowing the significance of the result obtained
(1). The estimation of a population parameter from a sample statistic such as
Mean.
(2). The testing of hypothesis about the population parameter (u).
We then set up certain limits on both sides of the population mean (u) on the
basis of the fact that mean (X) of sample size 30 or more one normally distributed
around the population mean (u). These limits are called the confidence limits and
the range between the two is called the confidence interval.
The limit of the region at which we no longer regard the chance to be
operating is called the level of significance.
If the chance limit is set at mean plus minus 1.96 SE, it implies 5% or 0.05 level
of significance, also called the critical level of significance. A valve lying beyond
this area is said to be significantly different from the population valve such a
different valve will be found by chance only 5 times in 100 results (P, 0.05).
If the linc is drawn at a distance of 2.58 SE from the mean, its level of
significance is said to be 1%. A valve lying beyond this area is said to be highly
significant because such a different valve will be found by chance only one in 100
results (P, 0.01).
In most the statistical studies, the level of significance are set at 5%, (P, 0.05),
1% (p, 0.01) and 0.5% (p, 0.005). Significant or insignificant indicate whether a
valve is likely to occur by chance or it is unlikely to occur by chance.
Statistically Hypothesis:-- A Hypothesis is a supposition made as a basis for
reasoning. According to Prof. Moris Hambarg, “A Hypothesis in statistics is simply
a quantitative statement about a population”. The two hypothesis in a statistical
test are normally referred to is (1). Null hypothesis, (2). Alternative hypothesis
(1). A null hypothesis or hypothesis of no difference (Ho) between statistic of a
sample and parameter of population or between statistic of two samples. This
hypothesis nullifies the claim that the experimental result is different from or
better than the one observed already. For example, If we want to find out
whether a particular drug is effective in curing malaria. We will take te null
hypothesis that the drug is not effective in curing malaria. The rejection of null
hypothesis indicate that the difference have statistical significance and the
acceptance of null hypothesis indicate that the difference are due to chance.
(2). The Alternative hypothesis of significant difference (Ho) stating that the
sample result in different greater or smaller than the hypothetical valve of
population e.g., weight gain or less due to new fading regimen.
To make minimum error in rejection or acceptance of Ho. We divide ±±
±sampling distribution or the area under the normal curve into two region or
zones (1). A sum of acceptance, (2). A zone of rejection
(Diagram)
(1). Zone of acceptance: -- If the result of a sample falls in the plain area, i.e.,
within the mean± 1.96 SE the null hypothesis is accepted, hence this area is called
the zone of acceptance for null hypothesis.
(2). Zone of Rejection:-- I the result of a sample falls in the shaded area, i.e.,
beyond mean±1.96 SE it is significantly different from the universe valve. Hence,
the Ho of no difference is rejected and alternative Ha is accepted. This shaded
area is called the zone of rejection of null hypothesis.
Type- 1 Error (α- Error):-- Consider a clinical trail comparing a new drug against a
standard drug in recovery from a disease. The alternative hypothesis is that the
new drug is more effective in comparision to the standard drug, whereas the null
hypothesis is that the two drugs are equally effective. It is possible to obtain
better recovery in the sample receiving the new drug in comparison to the sample
receiving the standard drug even when the drugs are indeed equally effective in
the population. This may give rise to a false- positive finding. This erroneous
conclusion of falsely rejecting the null hypothesis is called the type-1 error or α-
error.
Type-II Error (β- Error):-- Considering the same example, if the new drug is
actually more effective than the standard drug, it is still possible to obtain
samples that show no evidence of difference between the two groups due to
more sampling variation. This may lead to a false- negative finding. This erroneous
conclusion of not rejecting a false null hypothesis is called type-II error or β-error.
When a statistical hypothesis is tested these are four possibilities:-
Hypothesis Ho is true and test accepts it because the result falls within the zone
of acceptance at 5% level.
Hypothesis Ho is false and our test rejects in because the estimate falls in shaded
area of rejection.
Hypothesis Ho is true still it is rejected, the estimate falls in acceptance zone at
5% level in plain area (Type I Error).
Hypothesis Ho is false but it is accepted, the estimate falls in the zone of rejection
(Type II Error).
Z – Test:-- Z-test is applied to the sampling variability, the difference observed
between a sample estimate and that of population is expressed in terms of SE
instead of SD. The score of valve of the ratio between the observed difference
and SE is called “Z-test”.
If the Z score falls within mean±1.96 SE i.e. in the zone of acceptance (95%
confidence limits) the Ho is accepted. If Ho is rejected is called the level of
significance.
The Z=test for mean has two application:
(1). To test the significance of difference between a sample (X) mean and a know
valve of population (µ)
Observed difference between sample Z = Mean(X) – Population mean(µ)
SE of sample mean
(2). To test the significance of difference between two sample mean or between
experiment sample mean and a control sample mean.
Z= observed difference between two sample mean
SE of difference between two sample means
The four prerequisites to apply Z- test for mean are:-
Sample must be randomly selected.
The data must be quantitative.
The variable is assumed to follow normal distribution in the population.
The sample size must be larger than 30.
If SD of population is known, Z-test can still be applied even if the sample is
smaller than 30.
The significance of Z valve, the probability (p) valve is found from the following
table, constructed on the basis of normal distribution.
Z : 1.6(1.65) , 2.0(1.96) , 2.3(2.2) , 2.6(2.58)
P : 0.10 , 0.05 , 0.02 , 0.01
If the Z valve increases, the P valve or probability of an event happening by
chance decreases and alternative happening due to same external factor is to be
considered.
Two tailed test: -- A two tailed test of hypothesis will reject the null hypothesis, of
the sample statistic is significantly higher than or lower than the hypothesized
population parameter. In two tail test rejection region is located in both the tails.
If we are testing a hypothesis at 5% level of significance, the size of the
acceptance region on each sides of the mean would be 0.475 and the size of
rejection region is 0.025. If the sample mean falls into this area, the hypothesis is
accepted. If the sample mean falls in the area beyond 1.96 SE, the hypothesis is
rejected because it falls into te rejection region. For example, when you want to
know the action of a particular drug is different from that of another,it will be two
tailed test.
One tailed test: -- In case of one tail test the rejection region will be located in
only one tail which may be located in only one tail which may be either left or
right depending upon the alternative hypothesis. For example, if we know
whether one particular drug is better than the other, it will be one tailed test.
Standard Error of the Mean (SE X):--The sample estimate of statistics (X, s or p)
will differ from population parameter (µ,σ or p) because of chance or biological
variability. Such a difference between sample and population valves is measured
by statistic know as sampling error or standard error.
Standard error is thus a measure of chance variation and it does not mean
error or mistake.
SE X = σ/ √n , σ= standard deviation of population
n= no of observation in the sample
Uses of SE X
TO find the confidence limits of population mean. If standard deviation of the
sample is known.
Whether the sample is drawn from a known population or not when (µ) its mean
is known. If the sample mean is larger than the known population mean 1.96 SE,
95% chance are that the sample is not drawn from the same population or it is
under the influence of same other factor.
(a). Z = (X - µ)/σ/√n , when µ and σ or known
(b). Z = (X -µ)/σ/√n , when µ is known but s is not known.
To find the SE of difference between two mean to know if the observed difference
between the means of two samples is real and statistically significant or it is
apparent and insignificant due to chance.
(Example page 170 & 171 mahajan)
Example :-- Mean & SD for the height of 50 boys were 150 and 7cm respectively.
Find SE of mean and 95 percent confidence limits of heights in nature. Could this
sample be from the universe with a population mean (m) of 154 cm.
Mean height= 150 cm , SD = 7 cm , n = 50 boys
SE X = 7/ √50 = 7/7.07 = 0.98
95% confidence limits in nature will be : mean height± 2 SE
150 + 2*0.98 =151.96 AND 150 – 2*0.98 = 148.04
The sample of 50 boys with a mean height of 150 cm is not drawn from the
universe with a population mean of 154 cm, because 154 cm is much more than
95% of confidence limit of 151.96 cm
Z= (observed – mean) / SE X = (154 – 150)/ 0.98 =4.08
4.08 is more than 4 times of the SE.
Since the ratio is more than 3 ( i.e. beyond 3 SD), it is considered to be highly
significant i.e., beyond 99% confident limit.
Standard Error of Difference between two mean of large samples SE (X1- X2):---
If independent, large and random samples are drawn in pairs, repeatedly from
the same population and each time difference between the two mean of each
pair is calculated, the mean of the differences in the paired samples would be
zero or nearly so.
Moreover, the differences in means of the different sets of twin samples will
follow normal frequency distribution curve. The SD of such a distribution of
difference is known as standard error of difference between two means. In
practice it is not possible to find difference of large no of samples and than find SE
of these difference. The test is applied to one pair directly if SD of two mean are
known.
(a). Two independent sample from the same population of SD
SE of the difference between mean = √SD² (1/n₁ + 1/n₂)
= SD √ (1/n₁ + 1/n₂)
The action of two different drug or two different doses of the drug can be
compared and as this test based on normal distribution sample should be large.
(b). If two random sample from different population
SE (X₁ - X₂) = √ SD₁²/n₁ + SD₂²/n₂
Z = (X₁ - X₂)/ SE (X₁-X₂)
EXAMPLE PAGE 175 MAHAJAN, 718.
Example:-- Pregnant women attaining an anganwadi and receiving nutritional
supplementation numbering 49 were match with 64 pregnant women not
attaining anganwadi and not getting the supplementation. Both groups were
followed up. The mean birth-weight of babies born to the farmer was 3.5 kg with
SD 1.4 kg and that of those born to the better, 3.0 kg with SD 1.6 KG. Is the higher
birth weight in the farmer due to the nutrition supplementation.
Where SD₁²= 1.4 kg, SD₂² = 1.6 kg , n₁= 49 , n₂= 64, X₁= 3.5 kg, X₂= 3.0 kg
SEDM = √ SD₁²/n₁ + SD₂²/n₂ = √ (1.4)²/49 + (1.6)²/64
= √(1.96/49) + (2.56/64) = √.04 + .04 =0.28
Z = (X₁-X₂)/ SE(X₁-X₂) = (3.5-3.0) / 0.28 = 1.78
Since the observed difference is less than 2 times the SE (within 95% confidence
limits) it is not significant. That means the difference in mean both weight is due
to chance and not due to nutritional supplementation.
Example:- A simple sample of the height of 6400 Englishman has a mean of 67.85
inches, SD of 2.56 inches. While a simple sample of height of 1600 Austrian has a
mean of 68.55 & SD of 2.52 inches. Do the data indicate that the Austrian are on
the average taller than the Englishmen. Give reasons?
Let us take the hypothesis that there is no significance difference in the mean
of Englishman and Austrian.
SE(X₁-X₂) = √σ₁²/n₁ +σ₂²/n₂ = √(2.56)²/6400 + (2.52)²/1600 = 0.0707
t = I 67.85 – 68.55I/ 0.0707 =9.9
Since the difference is more than SE (1% Level of significance), the hypothesis is
rejected. Hence the data indicate that the Austrian are on average taller than the
Englishman. Significance of Difference Between Mean of small sample by
Student’s t-test:
The ratio of observed difference between two means of small sample to the
SE of difference in the same is divided by letter ‘t’.
The ‘t’- statistics is defined as: t= (X-µ)*√n / s
Where s = √ ∑(X- Mean)²/ n-1
X = the mean of the sample
µ = the actual or hypothetical mean of the population
n = sample size
s = SD of the sample
If the calculated ‘t’ valve exceeds the valve given under p=0.05 in the table, it is
said to be significant at 5% level and null hypothesis Ho is rejected and alternative
hypothesis (Ha) is accepted.
Criteria for applying t-test: --
Random samples
Quantitative data
Variable normally distributed
Population SD is known
The variable t-distribution ranges from minus infinity to plus infinity
T-distribution is symmetrical and has a mean zero, like the standard normal
distribution.
Unpaired t-test (Independent sample):-- This test is applied to unpaired data of
independent observation made on two different group or separate group or
sample drawn from two population, like control group and treated (experimental)
group and their means are compared for their significant difference, it is known as
‘Unpaired comparisons.
t = (X₁ - X₂)/ SE , SE= SD√(1/n₁ + 1/n₂)
SD calculated by the following formula:
SD= √∑(X₁-X₁)²+∑(X₂-X₂)² /(n₁+n₂-2)
Where X₁= mean of the first sample
X₂= mean of the second sample
n₁= no of observation in the first sample
n₂= no of observation in the second sample
SD= combined standard deviation
EXAMPLE PAGE 180 MAHAJAN
Example:-- Two type of drug used on 5 and 7 patients for reducing their weight:
Drug A was imported and drug B indigenous. The decrease in the weight after
using the drugs for six months was as follows.
Drug A : 10,12,13,11,14
Drug B : 8,9,12,14,15,10,9
Is there is significant difference in the efficacy of the two drugs? If not, which drug
should you buy? (v=10,t₀.₀₅=2.223)
Let us take the hypothesis that there is no significant difference in efficacy of
the two drug. Applying t-test
t =(X₁-X₂)/s √(n₁n₂)/n₁+n₂
X₁ (X₁-X₁) (X₁-X₁)² X₂ (X₂-X₂) (X₂-X₂)²
10 -2 4 8 -3 9
12 0 0 9 -2 4
13 1 1 12 1 1
11 1 1 14 3 9
14 2 4 15 4 4
10 -1 1
9 -2 4
∑X₁= 60 ∑(X₁-X₁)=10 ∑X₂=77 ∑(X₂-
X₂)=44
X₁= ∑X₁/n = 60/5 = 12 X₂ = ∑X₂/n = 77/7 = 11 =
SD = √∑(X₁-X ₁)²+∑(X₂-X₂)² /(n₁+n₂-2) = √(10+44)/5+7-2 = √54/10 =2.324
t = (12-11)/2.324 √5*7/5+7 = 0.735
v =n₁+n₂-2 = 5+7-2 =10 t₀.₀₅=2.223
Calculated valve is less than the table valve, the hypothesis is accepted. Hence
there is no significance in the efficacy of two drug. Since Drug B is indigenous and
there is no difference in the efficacy of imported and indigenous drug, we should
buy indigenous drug B.
Example :-- For a random sample of 10 person, fed on diet it, the increased weight
in pounds in a certain period were: 10,6,16,17,13,12,8,14,15,9.
For another random sample of 12 persons, fed on diet B, the increase in the same
period were: 7,13,22,15,12,14,18,8,21,23,10,17.
Test whether the diets A and B differ significantly as regards their effect an
increase in weight. D.f is 20 valve of t at 5% level 2.09.
Let us take the null hypothesis that A and B do not differ significantly weight
regards to their effect on increase in weight. Applying t-test
t =(X₁-X₂)/s √(n₁n₂)/n₁+n₂ SD=
√∑(X₁-X₁)²+∑(X₂-X₂)² /(n₁+n₂-2)
X₁ (X₁-X₁) (X₁-X₁)² X₂ (X₂-X₂) (X₂-X₂)²
10 -2 4 7 -8 64
6 -6 36 13 -2 4
16 4 16 22 7 49
17 5 25 15 0 0
13 1 1 12 -3 9
12 0 0 14 -1 1
8 -4 16 18 3 9
14 2 4 8 -7 49
15 3 9 21 6 36
9 -3 9 23 8 64
10 -5 25
17 2 4
∑X₁= 120 ∑(X₁-X₁) 120 ∑X₂=180 ∑(X₂-X₂)=314
Mean increase in weight of 10 person fed on diet A, X₁= 120/10= 12 person
Mean increase in weight of 12 person fed on diet B, X₂= 180/12= 15 person
SD =√(120+314)/10+12-2 = 4.46
t= (12-15)/4.66 √10*12/10+12 =3*2.34/4.66 = 1.51
For v=20, the table valve of t at 5% level is 2.09. The calculated valve is less than
the table valve and hence the experiment provides no evidence against the
hypothesis. We, therefore, conclude that diets A and B do not differ significantly
as regards their effect on increase in weight is concerned.
Paired t-test:-- In the test it was assumed that the two sample are said to be
dependent when the in one sample are related to there in the other in any
significant or meaningful manner. In fact, the two sample may consist of pair of
observation made on the same object, individuals or, more generally, on the same
population elements. Such type of situation are commonly faced in medical
sciences, as in example below:-
To study the role of a factor or cause when the observations are made before and
after its eg. Of exertion on pulse rate, effect of a drug on BP etc.
To compare the effect of two drugs, given to same individuals in the sample on
two different occasions,eg. Adrenaline and noradrenaline on pulse rate.
To study the comparative accuracy of two different instruments eg. Two types of
sphygmomanometers.
To compare results of two different laboratory techniques, eg. Microfilaria
infection rate by thick smear and concentration techniques; examination of stools
for hookworm over by zinc floatation and concentration methods and so on.
To compare observations made at two different sites in the same body, eg.
Compare temperature between toes and between fingers or in axilla and mouth,
or in mouth and rectum of the same individuals.
The t-test based on paired observation is defined by the following formula:
t = (X-0)*√n /s = X√n /s ,
Where, X = the mean of the difference, s= SD of the difference
S = √∑(X-X)²/n-1 or √∑X² - n(X)²/ n-1, D.f = n-1
EXAMPLE PAGE 187 MAHAJAN
Example:- A drug is given to 10 patients, and the increments in their blood
pressure where recorded to be 3,6,-2,4,-3,4,6,0,0,2. Is it reasonable to belive that
the drug has no effect on change of BP? (5% valve of t for 9 d.f = 2.26).
Let us take the hypothesis that the drug has no effect on change of BP.
Applying the difference test:
t = X√n/s or X/SE
X : 3, 6, -2, 4, -3, 4, 6, 0, 0, 2 = 20 X= ∑X/n =20/10= 2
(X-X): 1, 4, -4, 2, -5, 2, 4, -2, -2, 0 SD= √∑(X-X)²/n-1 = √90/10-1=3.16
(X-X)²: 1, 16, 16, 4, 25, 4, 16, 4, 4, 0 =90 t= 2√10/3.16 =2
V=n-1= 10-1 =9, t₀.₀₅=2.26
The CV of t is less than table valve. The hypothesis is accepted. Hence it is
reasonable to believe that the drug has no effect on change of BP.
F- Test
The object of the F- test is to find out whether the two independent estimates of
population variance differ significantly, or whether the two samples may be
regarded as drawn from the normal populations having the same variance. For
carrying out the test of significance, we calculate the ratio F. F is defined as:
F= S₁²/S₂² Where, S₁²= ∑(X₁ - X₁)²/n₁-1, S₂²= ∑(X₂ - X₂)²/n₂-1
It should ne noted that S₁² is always the larger estimate of variance ie S₁²>S₂²,
v₁=n₁-1, v₂=n₂-1. Where
v₁= degree of freedom for sample having larger variance
v₂= degree of freedom for sample having smaller variance
The calculated valve of F is greater than the table valve than the F ratio is
significant and null hypothesis is rejected. On the other hand, if calculated of F is
less than the table valve the null hypothesis is accepted and it is inferred that
both the samples have come from the population having same variance. F-test is
based on the ratio of two variance, it is known as the Variance ratio test.
Assumption in F-test:
Normality i.e. the valve in each group are normally distributed.
Homogeneity i.e. the variance within each group should be equal for all groups.
Independence of error, it states that the error should be independent for each
valve.
Example:- Two random samples were drawn from two normal population and
their valves are:-
A: 66, 67, 75, 76, 82, 84, 88, 90, 92
B: 64, 66, 74, 78, 82, 85, 87, 92, 93, 95, 97
Test whether the two population have the same variance at the 5% level of
significance (F= 3.36) at 5% level of v₁=10 and v₂=8.
Let us take the hypothesis that the two population have the same variance.
Applying F-test
X₁ (X₁ - X₁) (X₁ - X₁)² X₂ (X₂ - X₂) (X₂ - X₂)²
66 -14 196 64 -19 361
67 -13 169 66 -17 289
75 -5 25 74 -9 81
76 -4 16 78 -5 25
82 2 4 82 -1 1
84 4 16 85 2 4
88 8 64 87 4 16
90 10 100 92 9 81
92 12 144 93 10 100
95 12 144
97 14 196
720 734 913 1298
X₁= 720/9 =80, X₂= 913/11 =83
S₁²= ∑(X₁-X₁)²/n-1 =734/9-1 =91.75, S₂² =∑(X₂-X₂)²/n-1= 1298/11-1=129.8
F=S₁²/S₂² =91.75/129.8 =0.707
The calculated valve of F is less than the table alve. The hypothesis is accepted.
Hence the two population have the same variance.
ANOVA
An analysis of variance could be used to test whether the mean cholesterol level
between three treatment group is equal or different. The aim is to compare
means by several groups using a one-way analysis of variance. One way ANOVA is
appropride when the subgroups to be compared by just one factor, eg. In the
comparison of mean cholesterol by three treatment group. In general, ANOVA is
a techniques used to test how cholesterol level is being influenced by more than
one factor, e.g. the combined effect of age, sex, and treatment.
One-way ANOVA:- In two sample independent t-test, the variability (or difference)
between groups is compared to the variability within groups. In one-way , the
same principle is extended to measure the variability between the groups versus
variability within the groups.
The ANOVA calculation is Shown:=
Total sum of square (SST) = ∑Xᵢ² - (∑Xᵢ)²/n , d.f = n-1
Sum of square between the sample (SSB)=(∑X₁)²/N+(∑X₂)²/N +….+(∑X)²/N –
(∑Xᵢ)²/N
Sum of square within the samples = Total sum of square – Sum of square between
sample
Analysis of Variance ANOVA Table (one-way classification)
Sources of Sum of square V (d.f) Mean square Variance ratio
variance of F
Between SSB v₁=c-1 MSB = SSB / c-
sample 1
Within sample SSW v₂= n-c MSW =
SSW/n-c
Total SST n-1 MSB/MSW
Exposure Outcomes
Yes No Total
Yes A b r₁
No C d r₂
Total C₁ C₂ N
Example:- Polymorph count was 350 out of 500 WBCs. At 95% confidence level,
within limits the population will lie?
Ans:- P= 350*100/500 = 70, q= 100-70 = 30
Sep = √70*30/500 = √4.2 = 2.05
The 95% confidence limits for the population proportion will be 70±2*2 =
66% and 74%.
In school A, n₁ = 50
p₁ with tonsillectomy = 23 out of 50 = 23*100 /50 = 46%,
q₁ without tonsillectomy = 100 – 46 = 54%
In school B, n₂ = 350, p₂ = 77*100/ 350= 22%, q₂= 100 – 22 = 78%
SE(p₁-p₂) = √(p₁q₁/n₁) + ( p₂q₂/n₂)
= √(46*54/50) + (22*78/350) = 7.39
By another formula
P = 23+77, out of 400 =25%
Q = 27+273, out of 400 = 75%
SE(p₁-p₂) = √PQ (1/n₁ +1/n₂) = √25*75 (1/50 + 1/350) = √600/14 = 6.55
Observed difference 46-22
Z=----------------------------- = --------- = 3.7
SE of difference 6.55
It is less than 1.96, the critical level of significance, hence insignificant at 95%
confidence limits.
CORRELATION ANALYSIS
TYPES OF Correlation:- There are five type of correlation depending on its extent
and direction are as follows:-
(1). Perfect Positive Correlation:- The two variables X and Y are directly
proportional and fully correlated with each other. The correlation coefficient ® =
+1 i.e., both variables rise or fall in the same proportion. Example- height and
weight, age and height, age and weight, temperature and pulse etc. X varies
directly and proportionally to Y.
(2). Perfect Negative Correlation :- The two variable X and Y are inversely
proportional to each other, i.e., when one rises the other falls in the same
proportion and vice-versa. The correlation coefficient (r) = -1. Example-
socioeconomic status and incidence of tuberculosis, income and malnutrition.
(3). Partial Positive Correlation:- In this type, the regression lines incline on one
another and ascent toward right. The nonzero valves of coefficient ( r) lie between
0 and +1, i.e., 0<r<1 such as temperature and pulse rate, age of husband and age
of wife, infant mortality rate and over crowding etc.
(4). Partial Negative Correlation :- In this type, the regression lines incline on
another and ascent toward left. The nonzero valves of coefficient ( r) lie between -
1 and 0, i.e., -1<r<0 such as income and infant mortality rate, age and vital
capacity among adults.
REGRESSION ANALYSIS
The meaning of the term ‘regression’ is the act of returning or going back.
The term regression was first used by Sir Francis Galton in 1877.
(1). Regression is the measure of the average relationship between two or more
variables in terms of the original units of the data
(2). “Regression analysis attempts to establish the nature of the relation-ship
between variables – that is, to study the functional relationship between the
variables and thereby provide a mechanism for prediction, or forecasting.”
Regression lines:-- I f we take two variables X and Y, we shall have two regression
Lines as the regression of X on Y and the regression of X on Y.
The regression line of Y on X gives the most probable valves of Y for given valves
of X and the regression line of X on Y gives the most probable valves of X for given
valves of Y.
When there is perfect correlation between the two variables (r=±1), the
regression lines will coincide, i.e., we will have only one line. On the other hand,
when correlation is partial, the lines will be separate and diverge forming an acute
angle at the meeting point of perpendiculars drawn from the means of variables.
Linear the correlation, greater will be the divergence of angle. If the variables are
independent, r is zero and the lines of regression are at right angles i.e., parallel to
OX and OY.