DR.
NTR UNIVERSITY OFHEALTH SCIENCES-EXAMINATIONS
DEPARTMENT OF COMMUNITY MEDICINE
RAJIV GANDHI INSTITUTE OF MEDICAL SCIENCES; ONGOLE
STATISTICAL EXERCISES
1. Calculate mean, median, mode for the following data
10,7,15,13,10,9,12,15,8,6,10.
Solution: Mean (X̅) = sum of observations/no. of observations
=∑X/n
=115/11
=10.45
Median: Arrange the data in increasing or decreasing order.
6,7,8,9,10,10,10,12,13,15,15
Median=middle value in the order i.e 10
Mode: Most frequently appeared no. i.e 10
[Link] range, mean deviation, standard deviation of the
following data.
7,8,10,3,15,8,6,5,6,2,8
Solution: Range=Highest value-Lowest value
=15-2
Range =13
Mean=∑X/n
=78/11
=7.09
X X̅ (X-X̅) (X-X̅) 2
7 7.09 -0.09 0.0081
8 7.09 0.91 0.8281
10 7.09 2.91 8.4681
3 7.09 -4.09 16.7281
15 7.09 6.94 48.16
8 7.09 0.91 0.8281
6 7.09 -1.09 1.18
5 7.09 -2.09 4.36
6 7.09 -1.09 1.18
2 7.09 -5.09 25.9
8 7.09 0.91 0.8281
2
∑(X-X̅) =26.12 ; ∑(X-X̅) =108.428
Mean deviation=∑(X-X̅)/n
=26.12/11
=2.37
2
Standard deviation=√∑(X-X̅) /n-1
=√108.428/10
=√10.8428
=3.29
[Link] two series adults aged 21 years and children 3 months old
following values are obtained for the height. Find which series shows
greater deviation.
Persons Mean height SD
Adults 160cms 10cms
Children 60cms 5cms
Solution: Coefficient of variation(CV)=SD/Mean×100
CV1=10/160×100
=6.25
CV2=5/60×100
=8.33
Children show greater variation than adults
[Link] standard deviation and variance from the following data
Respiratory rate/min:23,22,20,24,16,17,18,19,21.
Solution: Mean X̅ =∑X/n
=180/9
=20
Respiratory rate Mean Deviation from (X−X̅) 2
mean (X−X̅)
23 20 3 9
22 20 2 4
20 20 0 0
24 20 4 16
16 20 −4 16
17 20 −3 9
18 20 −2 4
19 20 −1 1
21 20 1 1
Total 20 60
Standard deviation =√{∑(X−X)2/n−1}
=√60/8
=2.73
2
Variance =∑(X−X̅) /n−1
=60/8
=7.5
5. A group of 15 normal children had a mean bilirubin with infantile
level of 1.05mg% and SD of 0.34. Another group of 15 children with
infantile cirrhosis of liver had mean bilirubin level of 4.89mg% and SD
of 2.52. What is the difference between two means.
Observation Mean(X̅) SD
n1 1.05mg% 0.34
n2 4.89mg% 2.52
Standard error of difference between two means
SEM =√(SD21/n1 +SD22/n2)
=√(0.34)2/15+(2.52)2/15
=√0.1156/15+6.3504/15
=√6.466/15
=√0.4310
=0.656
6.A group of 15 normal children had a mean bilirubin with infantile
level of 1.05mg% and SD of [Link] group of 15 children with
infantile cirrhosis of lever had mean bilirubin level of 4.89% and SD of
[Link] the difference between the two means statistically significant
?
Sex Number Mean SD
Normal 15 1.05mg% 0.34
Cirrhosis 15 4.89mg% 2.52
Standard error of difference between two mean
SEM=√SD21/n1+SD22/n2
=√(0.34)2/15+(2.52)2/15
=√0.1156+6.3504/15
=√6.466/15
=√0.4301
=0.656
Z=Observed difference/SEM
=4.89−1.05/0.656
=3.84/0.656
=5.39
As the Z value is more than 2 standard error of difference between
two means is significant
[Link] of7 subjects aged <5 years are 5,3,4,6,4,7,9 and that of adult
aged>25 years are 10,13,16,18,21,25,28,.Find out which group ESR
show greater variation
Age <5 years: ESR 5,3,4,6,4,7,9
Mean X̅ =∑X/n
=38/7=5.4
2
Observation Mean X Deviation from (X−X̅)
mean(X−X̅)
5 5.4 −0.4 0.64
3 5.4 −2.4 5.76
4 5.4 −1.4 1.96
6 5.4 0.6 0.36
4 5.4 −1.4 1.96
7 5.4 1.6 2.56
9 5.4 3.6 12.96
Total 25.72
2
Standard deviation=√∑(X−X̅) /n−1
=√25.72/6
=√4.28
=2.06
Adult aged >25 year :
ESR-10,13,16,18,21,25,28.
Mean=∑X/n
=131/7
=18.71
2
ESR values mean (X−X̅) (X−X̅)
10 18.71 −8.71 75.86
13 18.71 −5.71 32.60
16 18.71 −2.71 7.34
18 18.71 0.71 0.54
21 18.71 2.29 5.24
25 18.71 6.29 39.56
28 18.17 9.29 86.30
Total 247.4
2
Standard deviation=√∑(X−X̅) /n−1
=√247.4/6
=√41.23
=6.42
Coefficient of variation CV1=SD/Mean×100
= 2.06×100/5.42
=38
CV2=SD/Mean×100
=6.42×100/18.71
=34.31
ESR of 7 subjects aged <5 years show greater variation than aged>25
years
[Link] mean and standard deviation with working mean method
of following blood serum cholesterol levels of 10 subjects. 240,260,
290,245,255,288,272,263,277,250.
Solution: Mean(X̅)= ∑X/n
=2640/10
=264
X X̅ (X-X̅) (X-X̅) 2
240 264 -24 576
260 264 -4 16
290 264 26 676
245 264 -19 361
255 264 -9 81
288 264 24 576
272 264 8 64
263 264 -1 1
277 264 13 169
250 264 -14 196
TOTAL 142 2716
2
SD=√∑(X-X̅) /n-1
=√2716/9
=√301.7
=17.36
9. Find the significance of difference in the mean index of brightness
of 50 boys and 50 girls with the following values.
Solution:
SEX MEAN SD
Boys 91.2 37.0
Girls 90.1 37.2
SEM= √SD21/n1+SD22/n2
=√(37)2/50+(31.2)2/50
=√1369+973.44/50
=√2342.44/50
=√46.84
=6.84
Z=Observed value of difference between 2 means /SEM
=1.1/6.34
=0.16
As the Z value <2 there is no statistical significant difference
between the brightness of boys and girls
10. In School A, Tonsillectomy has been done in 23 students out of 50
, while in another school it was done 77 out of 350. Find whether the
difference observed in two schools by chance or influence.
Solution:
School Tonsillectomy Tonsillectomy Total
done not done
A 23 27 50
B 77 273 350
100 300 400
Null hypothesis: No difference between the Tonsillectomy done in 2
schools
Chi-square X2 = ∑(O−E)/E
Expected value=Column total× row total/grand total
Expected value of 1 cell E1=100×50/400
=12.5
Expected value of 2 cell E2=300×50/400
=37.5
Expected value of 3 cell E3=100×350/400
=
87.5
Expected value of 4 cell E4=300×350/400
=262.5
X21 = (23−12.5)2/12.5
= (10.5)2/12.5
=110.25/12.5
8.82
X22= (27−37.5)2/37.5
=110.25/37.5
=2.94
X23= (77−87.5)2/87.5
=110.25/87.5
=1.26
X24= (273−262.5)2/262.5
=110.25/262.5
=0.42
Total X2= X21+X22+ X23+X24
=8.82+2.94+1.26+0.42
=13.4
Degree of freedom = (no. of rows− 1)(no. of columns−1)
= (2−1)(2−1)
=1
Calculated value of X is higher than that of table X value at 1 df at 0.
05. So, null hypothesis is not accepted.
There is a statistical significant difference between the Tonsillectomy
done in school A and B.
[Link] schools referred in example above, there was history of
whooping cough in 25 students in A and 215 in B. Determine
whether the difference is by chance or real?
Solution:
School Whooping cough No whooping Total
cough
A 25 25 50
B 215 85 350
240 110 400
Null hypothesis: No difference in incidence of whooping cough in
school A and B
X2=∑(O−E)/E
Expected value E= row total× column total/ grand total
E1=50×240/400
=30
E2=50×110/400
=13.75
E3=350×240/400
=210
E4=350×110/400
=96.25
X2= X21+X22+X32+X42
=(25−30)2/30+(25−13.75)2/13.75+(215−210)2/210+(85−96.25)2/96.
25
= (−5)2/30+(11.5)2/13.75+52/210+(11.25)2/96.25
=25/30+126.56/13.75+25/210+126.56/96.25
=0.83+9.2+0.119+1.31
=11.459
Degree of freedom = (r−1)(c−1)
= (2−1)(2−1)
=1
Calculated X2 value is higher than the X2 table value (i.e3.84) 1 df at 0
.05. So, null hypothesis is not accepted. There is statistical significant
difference between the whooping cough in School A and B.
[Link] chi-square test to find out the efficacy of drug from data
given below
group died survived Total
Control on 10 20 30
placebo
Experiment on 5 60 65
drug
Total 15 80 95
Chi-square X2=∑(O−E)2/E
Expected value=row total× column total/grand total
E1=30×15/95
=4.736
E2=30×80/95
=25.26
E3=65×15/95
=10.26
E4=65×80/95
=54.74
Total X=X12+X22+X23+X24
= (10−4.74)2/4.74+(20−25.26)2/25.26+(5−10.26)2/10.26+(
60−54.74)2/54.74
=27.66/4.74+27.66/25.26+27.66/10.26+27.66/54.74
=5.83+1.095+2.695+0.505
=10.125
Degree of freedom= (r−1)(c−1)
= (2−1)(2−1)
=1
Calculated X value is higher than X table value 1 df at 0.05. So, null
hypothesis is not accepted. There is statistical significant difference
in mortality. So, the drug is efficacious.
[Link] scores in 2 groups of new born, one born to high risk
mothers and other to normal mothers are given below. Comment
whether there is significant difference in the Apgar scores of these 2
groups.
New born of high risk mother-5,3,2,4,7,6,3,2
New born of normal mothers-1,1,1,2,1,13,5,2.
Solution: Apgar scores of new born of high risk group
5,3,2,4,7,6,3,2
Mean X̅=∑X/n
=32/8
=4
2
Apgar scores X−X̅ (X−X̅)
X̅
5 4 1 1
3 4 −1 1
2 4 −2 4
4 4 0 0
7 4 3 9
6 4 2 4
3 4 −1 1
2 4 −2 4
∑(X−X̅)2=24
2
Standard deviation=√∑(X−X̅) /n−1
=√24/7
=√3.428
=1.85
Apgar scores of new born of normal mothers
1,1,1,2,1,13,5,2
Mean=∑X/n
=26/8
=3.25
2
Apgar scores X−X̅ (X−X̅)
X̅
1 3.25 −2.25 5.06
1 3.25 −2.25 5.06
1 3.25 −2.25 5.06
2 3.25 −1.25 1.56
1 3.25 −2.25 5.06
13 3.25 9.75 95.06
5 3.25 1.75 3.06
2 3.25 −1.25 1.56
∑(X−X̅) 2=121.48
Standard deviation=√∑(X−X)2/n−1
=√121.48/7
=√17.35
=4.16
Apgar scores of new mean SD
born
High risk mothers 4 1.85
Normal mothers 3.25 4.16
SEM=√SD21/n1+SD22/n2
=√(1.85)2/8+(4.16)2/8
=√3.42/8+17.3/8
=√20.72/8
=√2.59
=1.6
Z=Observed value/SEM
=4−3.25/1.6
=0.75/1.6
=0.046
Z value is less than 2 so there is no statistical significant difference
between 2 means.
[Link] an epidemiological study of diabetes in rural and urban
population of Ahmedabad district, the following data was obtained.
Compute the prevalence in the areas and determine if the results
differ statistically.
Area Diabetes No diabetes Total
Rural 45 3450 3495
Urban 107 3409 3516
52 6859 7011
Null hypothesis: No significant difference between prevalence in 2
areas.
Chi-square X2=∑(O−E)2/E
E1=3495×52/7011
=25.92
E2=3495×6859/7011
=3419.22
E3=3516×52/7011
=26.077
E4=3516×6859/7011
=3439.77
Total X=X21+X22+X23+X24
= (45−25.92)2/25.92+(3450−3419.22)2/3419.22+(107−26.
077)2/26.077+(3409−3439.77)2/3439.77
=364.04/25.92+947.408/3419.22+6548.04/26.077+946.79/3439.77
=14.04+0.277+251.104+0.275
=265.696
Degree of freedom= (r−1)(c−1)
= (2−1)(2−1)
=1
Calculated X value is more than X table value at 1 df at 0.05. there is
significant difference between prevalence of diabetes in urban and
rural area.
[Link] rate of 12 individuals are given below
58,66,70,74,80,86,90,100,79,96,88,97. Calculate measures of
dispersion.
Solution: Range=Highest value− Lowest value
=100−58
=42
Mean=∑X/n
=984/12
=82
Pulse rate (X−X̅) (X−X)2
X̅
58 82 −24 576
66 82 −16 256
70 82 −12 144
74 82 −8 64
80 82 −2 4
86 82 4 16
80 82 8 64
79 82 −3 9
100 82 18 364
96 82 14 196
88 82 6 36
97 82 15 225
∑(X−X)2=1914
Mead deviation=∑(X−X)2/n−1
= 1914/11
=174
Standard deviation=√∑(X−X)2/n−1
=√1914/11
=√174
=13.1
16. Incidence rate in the last influenza epidemic was found to be 50
per 100 of the population exposed . what should be the size of
sample to find incidence rate in the current epidemic if allowable
error is 5%?
Solution: Sample size =4pq/l2
P=prevalence
P=50/1000
=0.05
l=5% of p
=5×0.05/100
= 0.25
Q=100−p
=100−0.05
=99.95
Sample size=4×0.05×99.95/0.25×0.25
=319.84
17. what is maternal mortality rate? In the year 1992 in the district A,
30 women died due to puerperal sepsis and there were 8000 live
births. Calculate the maternal mortality rate of the district.
Maternal mortality rate=[Link]
maternal deaths/[Link] live births×100000
No. of maternal deaths =30
No. of live births=8000
MMR=30×1000000/8000
=375
18. Calculate the crude death rate and crude birth rate, infant
mortality rate of a town whose midyear population is 20,00,000, live
births in the year 60,000 no. of deaths 1,600, no. of infant deaths
were 2400.
Solution: Midyear population-20,00,000
No. of live births-60,000
No. of deaths-16,000
No. of infant deaths-2400
Crude death rate=
no. of deaths during the year/midyear population×1000
=16,000×1000/20,00,000
=8
Crude birth rate=
No. of live births during the year/midyear population×1000
=60,000×1000/20,00,000
=30
Infant mortality rate
No. of infants deaths/no. of live births×1000
=2400×1000/60000
=4
[Link] a town with mid-year population 2 lakhs there were 6050 live
births, 110 still births and 750 deaths in the first year of life in
particular year of which 250 deaths were within one week after birth
and 480 deaths in the first month of life. Calculate and comment on
perinatal, neonatal, and infant mortality rates for the following data.
Solution: live births=6050
Still births=110
Deaths in the first year=750
Early neonatal death =250
Late neonatal death=480
Perinatal mortality rate=still births+ early neonatal deaths/total
Live births×1000
=110+250/6050×1000
=360/6050×1000
=59.5
Neonatal mortality rate=early+ late neonatal deaths/no. of live
births×1000
=250+480/6050×1000
=730/6050×1000
=120.66
Infant mortality rate=no. of deaths of infants/no of live
births×1000
=750/6050×1000
=123.97
20. In a town with midyear population of 1,50,000 following vital
events occurred
Total live births=3200
Total deaths=1400
Infants deaths=270
Maternal deaths=10
Calculate and comment on crude birth rate, crude death rate, infant
mortality rate, maternal mortality rate.
Solution:
Crude death rate=no. of deaths during the year in the town/midyear
Population of the town×1000
=1400/1,50,000×1000
=9.33
Crude birth rate=no. of live births during the year/midyear
population×1000
=3200/1,50,000×1000
=21.3
Infant mortality rate=no. of infant deaths/no. of live births×1000
=270/3200×1000
=84.37
Maternal mortality rate=no. of maternal deaths/no. of live
births×10000
=10/3200×10000
=31.25
21. Data computed in a primary health centre is given to you,
calculate and comment on all possible rates
Total live births=4500
Total still births=44
Deaths under 7 days=100
Deaths between 7 to 28 days=75
Deaths between 28 to 1 year=165
Solution:
Perinatal mortality rate=still births+ early neonatal deaths/total
Live births×1000
=44+100/4500×1000
=32
Neonatal mortality rate=early+ late neonatal deaths/total live
births×1000
=100+75/4500×1000
=38.8
Infant mortality rate=total no. of deaths of infants/total live
births×1000
=100+75+165/4500×1000
=75.5
22. Midyear population of a city in 2001 was 1,020,000. The
following events occurred in the year 2001
Total live births=30,000
Maternal deaths=120
Calculate and comment crude birth rate and maternal mortality rate.
Solution:
Crude birth rate=no. of live births during the year/midyear
population×1000
=30,000/1,020,000×1000
=29.41
Maternal mortality rate=no. of maternal deaths/no. of live
births×10000
=120/30,000×10000
=400
23. A new screening test for a certain disease administered to 490
persons , 60 of whom are known to have the disease . the test was
positive in 50 of the persons with the disease, as well as in 20
persons without the disease .
Calculate the following
Sensitivity of test Specificity of the test
Predictive value of a positive test Predictive value of a negative
test.
Solution:
Screening test Disease positive Disease Total
negative
Positive 50 (a) 20 (b) 70
Negative 10 (c) 410 (d) 420
Total 60 430 490
Sensitivity – To detect the true positives
Specificity – To detect the true negatives
Sensitivity = a/a+c×100
= 50/50+10×100
= 83.33%
Specificity = d/b+d×100
= 410/20+410×100
= 4100/43
= 95.34%
Predictive value of a positive test = a/a+b×100
= 50/50+20×100
= 71.42%
Predictive value of a negative test = d/c+d×100
= 410/420×100
= 97.6%
24.A screening test was done for cervical cancer Results are given
below.
Screening test Disease positive Disease negative Total
PAP smear
Positive 400 (a) 150 (b) 550
Negative 100 (c) 4350 (d) 4450
Total 500 4500 5000
Calculate the following
Sensitivity of test Specificity of the test
Predictive valve of a positive test Predictive value of a negative test
Sensitivity – to detect the true positives
Specificity – to detect the true negatives
Sensitivity of test = a/a+c×100
= 400/400+100×100
= 80%
Specificity of test = d/b+d×100
= 4350/4500×100
= 96.6%
Predictive value of a positive test = a/a+b×100
= 400/550×100
= 72.72%
Predictive value of a negative test = d/c+d×100
=4350/4450×100
= 97.7%