Introduction to Biostatistics Concepts
Introduction to Biostatistics Concepts
CHAPTER 1
1. INTRODUCTION
1.1. Definition of biostatistics and Classification of statistics
Definition
Statistics: is defined as the science of collecting, organizing, presenting, analyzing
and interpreting numerical data for the purpose of assisting in making a more
effective decision.
Biostatistics: is the branch of applied statistics directed toward applications in the health
sciences and biology. It is sometimes distinguished from the field of biometry based upon
whether applications are in the health sciences (biostatistics) or in broader biology
(biometry; e.g., agriculture, ecology, wildlife biology).
Classification of statistics
Depending on how data can be used, statistics is sometimes divided in to two main areas
or branches.
1) Descriptive Statistics: is concerned with summary calculations, graphs, charts and
tables.
2) Inferential Statistics: is a method used to generalize from a sample to a population.
For example:
The average income of all families (the population) in Ethiopia can be estimated from
figures obtained from a few hundred (the sample) families.
It is important because statistical data usually arises from sample.
Statistical techniques based on probability theory are required.
Page 1 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
1.2. Data Types
Data are observations of random variables made on the elements of a population or
sample. Data are the quantities (numbers) or qualities (attributes) measured or observed
that are to be collected and/or analyzed. The word data is plural but datum is singular. A
collection of data is often called a data set (singular)
Scales of measurement
Proper knowledge about the nature and type of data to be dealt with is essential in order
to specify and apply the proper statistical method for their analysis and inferences.
Measurement scale refers to the property of value assigned to the data based on the
properties of order, distance and fixed zero.
In mathematical terms measurement is a functional mapping from the set of objects {Oi}
to the set of real numbers {M(Oi)}.
The goal of measurement systems is to structure the rule for assigning numbers to objects
in such a way that the relationship between the objects is preserved in the numbers
assigned to the objects. The different kinds of relationships preserved are called
properties of the measurement system.
Order
The property of order exists when an object that has more of the attribute than another
object, is given a bigger number by the rule system. This relationship must hold for all
objects in the "real world".
The property of ORDER exists
When for all i, j if Oi > Oj, then M(Oi) > M(Oj).
Page 2 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Distance
The property of distance is concerned with the relationship of differences between
objects. If a measurement system possesses the property of distance it means that the unit
of measurement means the same thing throughout the scale of numbers. That is, an inch
is an inch, no matters were it falls - immediately ahead or a mile downs the road.
More precisely, an equal difference between two numbers reflects an equal difference in
the "real world" between the objects that were assigned the numbers. In order to define
the property of distance in the mathematical notation, four objects are required: Oi, Oj,
Ok, and Ol . The difference between objects is represented by the "-" sign; Oi - Oj refers to
the actual "real world" difference between object i and object j, while M(Oi) - M(Oj)
refers to differences between numbers.
Fixed Zero
A measurement system possesses a rational zero (fixed zero) if an object that has none of
the attribute in question is assigned the number zero by the system of rules. The object
does not need to really exist in the "real world", as it is somewhat difficult to visualize a
"man with no height". The requirement for a rational zero is this: if objects with none of
the attribute did exist would they be given the value zero. Defining O0 as the object with
none of the attribute in question, the definition of a rational zero becomes:
The property of fixed zero is necessary for ratios between numbers to be meaningful.
SCALE TYPES
Measurement is the assignment of numbers to objects or events in a systematic fashion.
Four levels of measurement scales are commonly distinguished: nominal, ordinal,
interval, and ratio and each possessed different properties of measurement systems.
Nominal Scales
Nominal scales are measurement systems that possess none of the three properties stated
above.
Level of measurement which classifies data into mutually exclusive, all inclusive
categories in which no order or ranking can be imposed on the data.
No arithmetic and relational operation can be applied.
Examples:
o Political party preference (Republican, Democrat, or Other,)
o Sex (Male or Female.)
o Marital status(married, single, widow, divorce)
o Country code
o Regional differentiation of Ethiopia.
Page 3 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Ordinal Scales
Ordinal Scales are measurement systems that possess the property of order, but not the
property of distance. The property of fixed zero is not important if the property of
distance is not satisfied.
Level of measurement which classifies data into categories that can be ranked.
Differences between the ranks do not exist.
Arithmetic operations are not applicable but relational operations are applicable.
Ordering is the sole property of ordinal scale.
Examples:
o Letter grades (A, B, C, D, F).
o Rating scales (Excellent, Very good, Good, Fair, poor).
o Military status.
Interval Scales
Interval scales are measurement systems that possess the properties of Order and
distance, but not the property of fixed zero.
Level of measurement which classifies data that can be ranked and differences are
meaningful. However, there is no meaningful zero, so ratios are meaningless.
All arithmetic operations except division are applicable.
Relational operations are also possible.
Examples:
o IQ
o Temperature in oF.
Ratio Scales
Ratio scales are measurement systems that possess all three properties: order, distance,
and fixed zero. The added power of a fixed zero allows ratios of numbers to be
meaningfully interpreted; i.e. the ratio of Bekele's height to Martha's height is 1.32,
whereas this is not possible with interval scales.
Level of measurement which classifies data that can be ranked, differences are
meaningful, and there is a true zero. True ratios exist between the different units
of measure.
All arithmetic and relational operations are applicable.
Examples:
o Weight
o Height
o Number of students
o Age
Operations that make sense for variables of different scales:
No. Scale Operation that make sense
Counting Ranking Addition/ Multiplication/
Subtraction Division
1. Nominal
2. Ordinal
3. Interval
4. Ratio
Page 4 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Generally, variable types can be distinguished based on their scale_ Typically_ different
statistical methods are appropriate for variables of different scales.
No. Scale Characteristic Question Examples
1. Nominal Is A different than B? Marital status
Eye color
Gender
Religious affiliation
Race
2. Ordinal Is A bigger than B? Stage of disease
Severity of pain
Level of satisfaction
3. Interval By how many units do A and B differ? Temperature
SAT score
4. Ratio How many times bigger than B is A? Distance
Length
Time until death
Weight
The following present a list of different attributes and rules for assigning numbers to
objects. Try to classify the different measurement systems into one of the four types of
scales. (Exercise)
Page 5 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Page 6 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.1.2. Methods of data presentation
Having collected and edited the data, the next important step is to organize it. That is to
present it in a readily comprehensible condensed form that aids in order to draw
inferences from it. It is also necessary that the like be separated from the unlike ones.
Classification is a preliminary and it prepares the ground for proper presentation of data.
Definitions:
Raw data: recorded information in its original collected form, whether it be
counts or measurements, is referred to as raw data.
Frequency: is the number of values in a specific class of the distribution.
Frequency distribution: is the organization of raw data in table form using classes
and frequencies.
Used for data that can be place in specific categories such as nominal, or ordinal. e.g. marital
status.
Example:
Page 7 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Step 1: Make a table as shown.
Class Tally Frequency Percent
(1) (2) (3) (4)
M
S
D
W
Step 2: Tally the data and place the result in column (2).
Step 3: Count the tally and place the result in column (3).
Step 4: Find the percentages of values in each class by using;
f
% *100 Where f= frequency of the class, n=total number of value.
n
Percentages are not normally a part of frequency distribution but they can be added since
they are used in certain types diagrammatic such as pie charts.
Step 5: Find the total for column (3) and (4).
Combing the entire steps one can construct the following frequency distribution.
Class Tally Frequency Percent
(1) (2) (3) (4)
M 5 20
////
S //// // 7 28
D //// // 7 28
W //// 6 24
2) Ungrouped frequency Distribution:
It is a table of all the potential raw score values that could possible occur in the data along
with the number of times each actually occurred. It is often constructed for small set or data
on discrete variable.
Solution:
Step 1: Find the range, Range=Max-Min=90-60=30.
Step 2: Make a table as shown
Step 3: Tally the data.
Step 4: Compute the frequency.
Page 8 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Mark Tally Frequency
60 // 2
62 / 1
63 / 1
65 / 1
70 //// 4
74 / 1
75 // 2
76 / 1
80 /// 3
85 /// 3
90 / 1
Each individual value is presented separately, that is why it is named ungrouped
frequency distribution.
3) Grouped frequency Distribution:
When the range of the data is large, the data must be grouped in to classes that are more than
one unit in width.
Definitions:
Grouped Frequency Distribution: a frequency distribution when several numbers are
grouped in one class.
Class limits: Separates one class in a grouped frequency distribution from another. The limits
could actually appear in the data and have gaps between the upper limits of one class and
lower limit of the next.
Units of measurement (U): the distance between two possible consecutive measures. It is
usually taken as 1, 0.1, 0.01, 0.001, -----.
Class boundaries: Separates one class in a grouped frequency distribution from another. The
boundaries have one more decimal places than the row data and therefore do not appear in the
data. There is no gap between the upper boundary of one class and lower boundary of the
next class. The lower class boundary is found by subtracting U/2 from the corresponding
lower class limit and the upper class boundary is found by adding U/2 to the corresponding
upper class limit.
Class width: the difference between the upper and lower class boundaries of any class. It is
also the difference between the lower limits of any two consecutive classes or the difference
between any two consecutive class marks.
Class mark (Mid points): it is the average of the lower and upper class limits or the average
of upper and lower class boundary.
Cumulative frequency: is the number of observations less than/more than or equal to a
specific value.
Cumulative frequency above: it is the total frequency of all values greater than or equal to
the lower class boundary of a given class.
Cumulative frequency blow: it is the total frequency of all values less than or equal to the
upper class boundary of a given class.
Cumulative Frequency Distribution (CFD): it is the tabular arrangement of class interval
together with their corresponding cumulative frequencies. It can be more than or less than
type, depending on the type of cumulative frequency used.
Relative frequency (rf): it is the frequency divided by the total frequency.
Relative cumulative frequency (rcf): it is the cumulative frequency divided by the total
frequency.
Page 9 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Guidelines for classes
Example*:
11 29 6 33 14 31 22 27 19 20
18 17 22 38 23 21 26 34 39 27
Page 10 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
Step 1: Find the highest and the lowest value H=39, L=6
Step 6: Find the upper class limit; e.g. the first upper class=12-U=12-1=11
11, 17, 23, 29, 35, 41 are the upper class limits.
So combining step 5 and step 6, one can construct the following classes.
Class limits
6 – 11
12 – 17
18 – 23
24 – 29
30 – 35
36 – 41
Class boundary
5.5 – 11.5
11.5 – 17.5
17.5 – 23.5
23.5 – 29.5
29.5 – 35.5
35.5 – 41.5
Page 11 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Step 8: tally the data.
Step 9: Write the numeric values for the tallies in the frequency column.
Class Class boundary Class Tally Freq. Cf (less Cf (more rf. rcf (less
limit Mark than than type) than type
type)
6 – 11 5.5 – 11.5 8.5 // 2 2 20 0.10 0.10
12 – 17 11.5 – 17.5 14.5 // 2 4 18 0.10 0.20
18 – 23 17.5 – 23.5 20.5 7 11 16 0.35 0.55
//////
24 – 29 23.5 – 29.5 26.5 //// 4 15 9 0.20 0.75
30 – 35 29.5 – 35.5 32.5 /// 3 18 5 0.15 0.90
36 – 41 35.5 – 41.5 38.5 // 2 20 2 0.10 1.00
These are techniques for presenting data in visual displays using geometric and pictures.
Importance:
-The three most commonly used diagrammatic presentation for discrete as well as qualitative
data are:
Pie charts
Pictogram
Bar charts
Page 12 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Pie chart
- A pie chart is a circle that is divided in to sections or wedges according to the percentage of
frequencies in each category of the distribution. The angle of the sector is obtained using:
Solutions:
Step 3: Using a protractor and compass, graph each section and write its name corresponding
percentage.
Page 13 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Pictogram
-In these diagram, we represent data by means of some picture symbols. We decide
abut a suitable picture to represent a definite number of units in which the variable is
measured.
Solution
Page 14 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Bar Charts:
- A set of bars (thick lines or narrow rectangles) representing some magnitude over time
space.
- They are useful for comparing aggregate over time space.
- Bars can be drawn either vertically or horizontally.
- There are different types of bar charts. The most common being :
Simple bar chart
Deviation o0r two way bar chart
Broken bar chart
Component or sub divided bar chart.
Multiple bar charts.
Solutions:
30
25
Sales in $
20
15
10
5
0
A B C
product
Page 15 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Component Bar chart
-When there is a desire to show how a total (or aggregate) is divided in to its component
parts, we use component bar chart.
-The bars represent total value of a variable with each total broken in to its component parts
and different colours or designs are used for identifications
Example:
Draw a component bar chart to represent the sales by product from 1957 to 1959.
Solutions:
100
80
Sales in $
Product C
60
Product B
40
Product A
20
0
1957 1958 1959
Year of production
60
50
Sales in $
40 Product A
30 Product B
20 Product C
10
0
1957 1958 1959
Year of production
Page 16 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Graphical Presentation of data
- The histogram, frequency polygon and cumulative frequency graph or ogive are most
commonly applied graphical representation for continuous data.
Procedures for constructing statistical graphs:
Draw and label the X and Y axes.
Choose a suitable scale for the frequencies or cumulative frequencies and label it on the Y
axes.
Represent the class boundaries for the histogram or ogive or the mid points for the
frequency polygon on the X axes.
Plot the points.
Draw the bars or lines to connect the points.
Histogram
A graph which displays the data by using vertical bars of various height to represent
frequencies. Class boundaries are placed along the horizontal axes. Class marks and class limits
are some times used as quantity on the X axes.
Solution
Page 17 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Frequency Polygon:
- A line graph. The frequency is placed along the vertical axis and classes mid points are placed
along the horizontal axis. It is customer to the next higher and lower class interval with
corresponding frequency of zero, this is to make it a complete polygon.
Example:
Draw a frequency polygon for the above data (example *).
Solutions:
8
4
Value Frequency
0
2. 5 8. 5 14.5 20.5 26.5 32.5 38.5 44.5
Example:
Draw an Ogive curve(less than type) for the above data. (Example *)
Page 18 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The expression is read, "the sum of X sub i from i equals 1 to N." It means "add up all the
numbers."
Example: Suppose the following were scores made on the first homework assignment for
five students in the class: 5, 7, 7, 6, and 8. In this example set of five numbers, where
N=5, the summation could be written:
The "i=1" in the bottom of the summation notation tells where to begin the sequence of
summation. If the expression were written with "i=3", the summation would start with the
third number in the set. For example:
In the example set of numbers, this would give the following result:
Page 19 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The "N" in the upper part of the summation notation tells where to end the sequence of
summation. If there were only three scores then the summation and example would be:
Sometimes if the summation notation is used in an expression and the expression must be
written a number of times, as in a proof, then a shorthand notation for the shorthand
notation is employed. When the summation sign "" is used without additional notation,
then "i=1" and "N" are assumed.
For example:
PROPERTIES OF SUMMATION
n
1. k nk
i 1
where k is any constant
n n
2. kX i k X i where k is any constant
i 1 i 1
n n
3. (a bX
i 1
i ) na b X i
i 1
where a and b are any constant
n n n
4. (X
i 1
i Yi ) X i Yi
i 1 i 1
Xi ( X i Yi ) X
2
a) d) g) i
i 1 i 1 i 1
5 5 5 5
b) Yi
i 1
e) ( X i Yi )
i 1
h) ( X i )( Yi )
i 1 i 1
5 5
c) 10
i 1
f) X Y
i 1
i i
Solutions:
Page 20 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
5
a) X
i 1
i 5 7 7 6 8 33
5
b) Y
i 1
i 6 7 8 7 8 36
5
c) 10 5 *10 50
i 1
5
d) (X
i 1
i Yi ) (5 6) (7 7) (7 8) (6 7) (8 8) 69 33 36
5
e) (X
i 1
i Yi ) (5 6) (7 7) (7 8) (6 7) (8 8) 3 33 36
5
f) X Y
i 1
i i 5 * 6 7 * 7 7 * 8 6 * 7 8 * 8 241
5
X 5 2 7 2 7 2 6 2 8 2 223
2
g) i
i 1
5 5
h) ( X i )( Yi ) 33 * 36 1188
i 1 i 1
𝑋1 + 𝑋2 + … + 𝑋𝑛
𝑋 =
𝑛
∑𝑛𝑖=1 𝑋𝑖
𝑋 =
𝑛
If X1 occurs f1 times
If X2occurs f2 times
.
.
If Xn occurs fn times
Page 21 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
k
f X , i i
X i 1
k
f
i 1
i
k
where k is the number of classes and f
i 1
i n
Example:
Obtain the mean of the following number
2, 7, 8, 2, 7, 3, 7
Solution:
Xi fi Xifi
2 2 4
3 1 3
7 3 21
8 1 8
Total 7 36
4
f i Xi
36
X i 1
4
5.15
f
7
i
i 1
If data are given in the shape of a continuous frequency distribution, then the mean is
obtained as follows:
∑𝑘𝑖=1 𝑓𝑖 𝑋𝑖
𝑋̅ =
∑𝑘𝑖=1 𝑓𝑖
where Xi =the class mark of the ith class and fi = the frequency of the ith class
Example:
calculate the mean for the following age distribution.
Class Class boundaries frequency
6- 10 5.5 - 10.5 35
11- 15 10.5 - 15.5 23
16- 20 15.5 - 20.5 15
21- 25 20.5 - 25.5 12
26- 30 25.5 - 30.5 9
31- 35 30.5 - 35.5 6
Solutions:
First find the class marks
Find the product of frequency and class marks
Find mean using the formula.
Page 22 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Class Class boundaries fi Xi Xifi
6- 10 5.5 - 10.5 35 8 280
11- 15 10.5 - 15.5 23 13 299
16- 20 15.5 - 20.5 15 18 270
21- 25 20.5 - 25.5 12 23 276
26- 30 25.5 - 30.5 9 28 252
31- 35 30.5 - 35.5 6 33 198
Total 10 1575
0
f i Xi
1575
X i 1
6
15.75
f
100
i
i 1
Exercises:
1. Marks of 75 students are summarized in the following frequency distribution:
Marks No. of students
40-44 7
45-49 10
50-54 22
55-59 f4
60-64 f5
65-69 6
70-74 3
If 20% of the students have marks between 55 and 59
i. Find the missing frequencies f4 and f5.
ii. Find the mean.
If the values in a series or mid values of a class are large enough, coding of values is a
good device
to simplify the calculations.
For raw data suppose we have used the following coding system.
di X i A
X i di A
n n
Xi (d i A)
X i 1
i 1
n n
n
d i
X A i 1
n
X Ad
Where A is an assumed mean and d is the mean of the coded data.
Page 23 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
If the data are expressed in terms of ungrouped frequency distribution
di X i A
X i di A
k k
fX i i f (d i i A)
X i 1
i 1
n n
k
fd i i
X A i 1
n
X Ad
In both cases the true mean is the assumed mean plus the average of the deviations
from the assumed mean.
Suppose the data is given in the shape of continuous frequency distribution with a
constant class size of w then the following coding is appropriate.
X A
d i
i w
X wd A
i i
k k
f X f ( wd A)
i i i i
X i 1 i 1
n n
k
f wd
i i
X A i 1
n
X A wd
Where: Xi is the original class mark for the ith class.
di is the transformed class mark for the ith class.
A is an assumed mean usually the mean of the class marks.
(i =1, 2… k)
Example:
1. Suppose the deviations of the observations from an assumed mean of 7 are:
1, -1, -2, -2, 0, -3, -2, 2, 0, -3.
a) Find the true mean b) Find the original observation
Solutions:
10
A 7, d i 10
i 1
a) d 10 1
10
X A d 7 1 6
The true mean is 6.
b) Using Xi=A+di we obtain the following original observations:
8, 6, 5, 5, 7, 4, 5, 9, 7, 4.
Page 24 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Special properties of Arithmetic mean
1. The sum of the deviations of a set of items from their mean is always zero. i.e.
n
( X X ) 0.
i 1
i
2. The sum of the squared deviations of a set of items from their mean is the
n n
minimum. i.e. ( Xi X ) 2 ( X A) 2 , A X
i
i 1 i 1
X1n1 X 2 n 2 .... X k n k
Xini
Xc i1k
n1 n 2 ...n k
n
i 1
i
Example:
In a class there are 30 females and 70 males. If females averaged 60 in an examination
and boys averaged 72, find the mean for the entire class.
Solutions:
Females
Males
X 1 60 X 2 72
n1 30 n2 70
X n
X 1 n1 X 2 n2 i 1 i i
Xc 2
n1 n2 nii 1
4. If a wrong figure has been used when calculating the mean the correct mean can be
obtained with out repeating the whole process using:
(CorrectValue WrongValue)
CorrectMean WrongMean
n
Where n is total number of observations.
Example:
An average weight of 10 students was calculated to be [Link] it was discovered
that one weight was misread as 40 instead of 80 k.g. Calculate the correct average
weight.
Page 25 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
(CorrectValue WrongValue)
CorrectMean WrongMean
n
(80 40)
CorrectMean 65 65 4 69k.g.
10
Weighted Mean
When a proper importance is desired to be given to different data a weighted mean
is appropriate.
Weights are assigned to each item in proportion to its relative importance.
Let X1, X2, …Xn be the value of items of a series and W1, W2, …Wn their
corresponding weights , then the weighted mean denoted X w is defined as:
n
X W i i
Xw i 1
n
W
i 1
i
Example:
A student obtained the following percentage in an examination:
English 60, Biology 75, Mathematics 63, Physics 59, and chemistry [Link] the
students weighted arithmetic mean if weights 1, 2, 1, 3, 3 respectively are allotted
to the subjects.
Page 26 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
5
X W i i
60 * 1 75 * 2 63 * 1 59 * 3 55 * 3 615
Xw i 1
61.5
1 2 1 3 3
5
10
W
i 1
i
The geometric mean of a set of n observation is the nth root of their product.
The geometric mean of X1, X2 ,X3 …Xn is denoted by G.M and given by:
𝐺. 𝑀 = 𝑛√𝑋1 ∗ 𝑋2 ∗ … ∗ 𝑋𝑛
Page 27 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Harmonic Mean
The harmonic mean of X1, X2 , X3 …Xn is denoted by H.M and given by:
n
H.M n , This is called simple harmonic mean.
1
i 1 X i
n k
H.M , n fi
k
fi
i 1 X i
i 1
If observations X1, X2, …Xn have weights W1, W2, …Wn respectively, then their
harmonic mean is given by
W i
H.M n
i 1
, This is called Weighted Harmonic Mean.
W
i 1
i Xi
Remark: The Harmonic Mean is useful and appropriate in finding average speeds and
average rates.
Example: A cyclist pedals from his house to his college at speed of 10 km/hr and back
from the college to his house at 15 km/hr. Find the average speed.
Solution: Here the distance is constant
The simple H.M is appropriate for this problem.
X1= 10km/hr X2=15km/hr
2
H.M 12km /hr
1 1
10 15
The Mode
- Mode is a value which occurs most frequently in a set of values
- The mode may not exist and even if it does exist, it may not be unique.
- In case of discrete distribution the value having the maximum frequency is the model
value.
Examples:
1. Find the mode of 5, 3, 5, 8, 9
Mode =5
2. Find the mode of 8, 9, 9, 7, 8, 2, and 5.
It is a bimodal Data: 8 and 9
3. Find the mode of 4, 12, 3, 6, and 7.
No mode for this data.
- The mode of a set of numbers X1, X2, …Xn is usually denoted by X̂ .
Page 28 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
If data are given in the shape of continuous frequency distribution, the mode is defined
as:
1
X̂ L mo w
1 2
Where:
Xˆ the mod e of the distribution
w the sizeof the mod al class
1 f mo f 1
2 f mo f 2
f mo frequencyof the mod al class
f 1 frequencyof the classpreceedingthe mod al class
f 2 frequencyof the class followingthe mod al class
Example: Following is the distribution of the size of certain farms selected at random
from a district. Calculate the mode of the distribution.
Lmo 45 ˆ 45 10 2
X
2 26
w 10 45.71
1 f mo f 1 2
2 f mo f 2 26
f mo 31
f 1 29
f2 5
Page 29 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Merits and Demerits of Mode
Merits:
It is not affected by extreme observations.
Easy to calculate and simple to understand.
It can be calculated for distribution with open end class
Demerits:
It is not rigidly defined.
It is not based on all observations
It is not suitable for further mathematical treatment.
It is not stable average, i.e. it is affected by fluctuations of sampling to
some extent.
Often its value is not unique.
Note: being the point of maximum density, mode is especially useful in finding the most
popular size in studies relating to marketing, trade, business, and industry. It is the
appropriate average to be used to find the ideal size.
The Median
- In a distribution, median is the value of the variable which divides it in to two
equal halves.
- In an ordered series of data median is an observation lying exactly in the middle
of the series. It is the middle most value in the sense that the number of values
less than the median is equal to the number of values greater than it.
- If X1, X2, …Xn be the observations, then the numbers arranged in ascending
order will be X[1], X[2], …X[n], where X[i] is ith smallest value.
X[1]< X[2]< …<X[n]
-Median is denoted by 𝑋̃.
Median for ungrouped data
𝑿 𝒏+𝟏 , 𝒊𝒇 𝒏 𝒊𝒔 𝒐𝒅𝒅
[ ]
𝟐
̃=
𝑿
𝟏
(𝑿 𝒏 + 𝑿 𝒏+𝟐 ) , 𝒊𝒇 𝒏 𝒊𝒔 𝒆𝒗𝒆𝒏
{𝟐 [𝟐] [
𝟐
]
Example: Find the median of the following numbers.
a) 6, 5, 2, 8, 9, 4.
b) 2, 1, 8, 3, 5.
Solutions: ~ 1
X (X n X n )
a) First order the data: 2, 4, 5, 6, 8, 9 2 [2] [ 1]
2
Here n=6 1
(X [3] X [ 4 ] )
2
1
(5 6) 5.5
2
b) Order the data :1, 2, 3, 5, 8 ~ X
X n 1
[ ]
Here n=5 2
X[3]
3
Page 30 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Median for grouped data
If data are given in the shape of continuous frequency distribution, the median is defined
~ w n
X L med ( c)
f med 2
Where :
L med lower class boundaryof the medianclass.
as: w the size of the medianclass
n total numberof observations.
c the cumulativefrequency(less than type) preceedingthe medianclass.
f med thefrequency of the medianclass.
Remark:
The median class is the class with the smallest cumulative frequency (less than type) greater
n
than or equal to .
2
Example: Find the median of the following distribution.
Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3
Solutions:
First find the less than cumulative frequency.
Identify the median class.
Find median using formula.
Page 31 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
n 75
37.5
2 2
39 is the first cumulativefrequencyto be greaterthanor equalto 37.5
50 54 is the median class.
L 49.5, w 5
med
n 75, c 17, f 22
med
~
X L w ( n c)
med f 2
med
49.5 5 (37.5 17)
22
54.16
Demerits:
It is not a good representative of data if the number of items is small.
It is not amenable to further algebraic treatment.
It is susceptible to sampling fluctuations.
Quantiles
When a distribution is arranged in order of magnitude of items, the median is the value of the
middle term. Their measures that depend up on their positions in distribution quartiles, deciles,
and percentiles are collectively called quantiles.
Quartiles:
- Quartiles are measures that divide the frequency distribution in to four equal parts.
- The value of the variables corresponding to these divisions are denoted Q1, Q2, and
Q3 often called the first, the second and the third quartile respectively.
- Q1 is a value which has 25% items which are less than or equal to it. Similarly Q2 has
50%items with value less than or equal to it and Q3 has 75% items whose values are
less than or equal to it.
N 1
- To find Qi (i=1, 2, 3) we count i of the classes beginning from the lowest class.
4
Page 32 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
- For grouped data: we have the following formula
w iN
Q
i LQ i f ( 4 c) , i 1,2,3
Qi
Where :
L lower class boundary of the quartile class.
Qi
w the size of the quartile class
N total number of observations.
c the cumulative frequency (less than type) preceeding the quartile class.
f thefrequency of the quartile class.
Qi
Remark:
The quartile class (class containing Qi ) is the class with the smallest cumulative frequency
iN
(less than type) greater than or equal to .
4
Deciles:
- Deciles are measures that divide the frequency distribution in to ten equal parts.
- The values of the variables corresponding to these divisions are denoted D1, D2,.. D9
often called the first, the second,…, the ninth decile respectively.
N 1
- To find Di (i=1, 2,..9) we count i of the classes beginning from the lowest class.
10
- For grouped data: we have the following formula
w iN
Di L Di ( c) , i 1,2,...,9
f Di 10
Where :
L Di lower class boundaryof the decile class.
w the size of the decileclass
N total numberof observations.
c the cumulativefrequency( less than type) preceedingthe decile class.
f Di thefrequency of the decileclass.
Remark: The decile class (class containing Di )is the class with the smallest cumulative
iN
frequency (less than type) greater than or equal to .
10
Page 33 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Percentiles:
- Percentiles are measures that divide the frequency distribution in to hundred equal
parts.
- The values of the variables corresponding to these divisions are denoted P1, P2,.. P99
often called the first, the second,…, the ninety-ninth percentile respectively.
N 1
- To find Pi (i=1, 2,..99) we count i of the classes beginning from the lowest
100
class.
Remark:
The percentile class (class containing Pi )is the class with the smallest cumulative
iN
frequency (less than type) greater than or equal to .
100
Example1: Considering the following data
a) 64,76, 77, 81, 62,64, 63, 70, 81, 72
b) 29, 40, 42, 25, 27, 26, 30, 41, 28
Calculate:
i. All quartiles.
ii. The 5th and 7th deciles
iii. The 50th and 90th percentiles
Solution:
a) Order: 62, 63, 64, 64, 70, 72, 76, 77, 81, 81
here: N = 10
i. All quartiles
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄1 = 1 ( ) 𝑣𝑎𝑙𝑢𝑒 = ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
(2.75) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒 + 0.75(3𝑟𝑑 𝑣𝑎𝑙𝑢𝑒 − 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒)
= 63 + 0.75(64 − 63) = 63 + 0.75(1)
= 63.75
Page 34 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄2 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
(5.5) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄3 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (8.25)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.25(9𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 77 + 0.25(81 − 77)
= 77 + 0.25(4)
= 77 + 1
= 78
ii. The 5th and 7th deciles
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝐷5 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (5.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝐷7 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (7.7)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.7(8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 7𝑡ℎ 𝑉𝑎𝑙𝑢𝑒 )
= 76 + 0.7(77 − 76)
= 76 + 0.7(1)
= 76.7
𝑁 + 1 𝑡ℎ 11 𝑡ℎ 11 𝑡ℎ
𝑃50 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
= (5.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71
Page 35 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
𝑁 + 1 𝑡ℎ 11 𝑡ℎ 11 𝑡ℎ
𝑃90 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 9 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
(9.9) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 9𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.9(10𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 9𝑡ℎ 𝑉𝑎𝑙𝑢𝑒 )
= 81 + 0.9(81 − 81)
= 81 + 0.9(0)
= 81
𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄1 = 1 ( ) 𝑣𝑎𝑙𝑢𝑒 = ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (2.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒 + 0.5(3𝑟𝑑 𝑣𝑎𝑙𝑢𝑒 − 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒)
= 26 + 0.5(27 − 26) = 26 + 0.5(1)
= 26.5
𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄2 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29
𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄3 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (7.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 40 + 0.5(41 − 40)
= 40 + 0.5(1)
= 40.5
𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝐷5 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29
𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝐷7 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
(7) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 40
Page 36 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
th th
iii. The 50 and 90 percentiles
𝑁 + 1 𝑡ℎ 10 𝑡ℎ 10 𝑡ℎ
𝑃50 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29
𝑁 + 1 𝑡ℎ 10 𝑡ℎ 10 𝑡ℎ
𝑃90 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 9 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
(9) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 42
Values Frequency
140- 150 17
150- 160 29
160- 170 42
170- 180 72
180- 190 84
190- 200 107
200- 210 49
210- 220 34
220- 230 31
230- 240 16
240- 250 12
Solutions:
First find the less than cumulative frequency.
Use the formula to calculate the required quantile.
Values Frequency [Link](less than type)
140- 150 17 17
150- 160 29 46
160- 170 42 88
170- 180 72 160
180- 190 84 244
190- 200 107 351
200- 210 49 400
210- 220 34 434
220- 230 31 465
230- 240 16 481
240- 250 12 493
Page 37 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
a) Quartiles:
i. Q1
- determine the class containing the first quartile.
N
123.25
4
170 180 is the classcontainingthe first quartile.
LQ 170 , w 10 w N
1 Q1 LQ 1 ( c)
N 493 , c 88 , f Q 72 f Q1 4
1
10
170 (123.25 88)
72
174.90
ii. Q2
- determine the class containing the second quartile.
2* N
246.5
4
190 200 is the class containing the sec ond quartile.
LQ 190 , w 10 w 2* N
2
Q2 LQ2 ( c)
N 493 , c 244 , f Q 107
2
f Q2 4
10
190 (246.5 244)
107
190.23
iii. Q3
- determine the class containing the third quartile.
3* N
369.75
4
200 210 is the class containing the third quartile.
LQ 200 ,
3
w 10 Q3 LQ 3
w 3* N
( c)
f Q3 4
N 493 , c 351 , f Q 49
3
10
200 (369.75 351)
49
203.83
Page 38 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
b) D7
- determine the class containing the 7th decile.
7* N
345.1
10
190 200 is the class containing the seventh decile.
LD 190 , w 10 w 7* N
7
D7 LD ( c)
N 493 , c 244 , f D 107 f D 10
7
7
7
10
190 (345.1 244)
107
199.45
c) P90
- determine the class containing the 90th percentile.
90 * N
443.7
100
220 230 is the class containing the 90th percentile.
L P90 220 , w 10 w 90 * N
P90 LP ( c)
N 493 , c 434 , f P90 31 f P 100
90
90
10
220 (443.7 434)
31
223.13
Page 39 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.2.2. Measures of Dispersion (Variation)
The measures of dispersion which are expressed in terms of the original unit of a series
are termed as absolute measures. Such measures are not suitable for comparing the
variability of two distributions which are expressed in different units of measurement and
different average size. Relative measures of dispersions are a ratio or percentage of a
measure of absolute dispersion to an appropriate measure of central tendency and are thus
pure numbers independent of the units of measurement. For comparing the variability of
two distributions (even if they are measured in the same unit), we compute the relative
measure of dispersion instead of absolute measures of dispersion.
Various measures of dispersions are in use. The most commonly used measures of
dispersions are:
1) Range and relative range
2) Quartile deviation and coefficient of Quartile deviation
3) Mean deviation and coefficient of Mean deviation
4) Standard deviation and coefficient of variation.
Page 40 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Range for grouped data:
If data are given in the shape of continuous frequency distribution, the range is computed
as:
Merits:
It is rigidly defined.
It is easy to calculate and simple to understand.
Demerits:
It is not based on all observation.
It is highly affected by extreme observations.
It is affected by fluctuation in sampling.
It is not liable to further algebraic treatment.
It can not be computed in the case of open end distribution.
It is very sensitive to the size of the sample.
Relative Range (RR)
-it is also some times called coefficient of range and given by:
LS R
RR
LS LS
Example:
1. Find the relative range of the above two distribution.(exercise!)
2. If the range and relative range of a series are 4 and 0.25 respectively. Then what is the
value of:
a) Smallest observation
b) Largest observation
Solutions :( 2)
R 4 L S 4 _________________(1)
RR 0.25 L S 16 _____________(2)
Solving (1) and (2) at the same time , one can obtain the following value
L 10 and S 6
Page 41 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Quartile Deviation (Semi-inter quartile range), Q.D
The inter quartile range is the difference between the third and the first quartiles of a set of
items and semi-inter quartile range is half of the inter quartile range.
Q3 Q1
Q.D
2
Coefficient of Quartile Deviation (C.Q.D)
(Q3 Q1 2 2 * Q.D Q3 Q1
C. Q.D
(Q3 Q1 ) 2 Q3 Q1 Q3 Q1
It gives the average amount by which the two quartiles differ from the median.
Example: Compute Q.D and its coefficient for the following distribution.
Values Frequency
140- 150 17
150- 160 29
160- 170 42
170- 180 72
180- 190 84
190- 200 107
200- 210 49
210- 220 34
220- 230 31
230- 240 16
240- 250 12
Solutions:
In the previous chapter we have obtained the values of all quartiles as:
Q1= 174.90, Q2= 190.23, Q3=203.83
Q3 Q1 203.83 174.90
Q.D 14.47
2 2
2 * Q.D 2 *14.47
C.Q.D 0.076
Q3 Q1 203.83 174.90
Remark: Q.D or C.Q.D includes only the middle 50% of the observation.
Page 42 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Mean Deviation (M.D):
The mean deviation of a set of items is defined as the arithmetic mean of the values of the
absolute deviations from a given average. Depending up on the type of averages used we
have different mean deviations.
a) Mean Deviation about the mean
Denoted by M.D( X ) and given by
n
Xi X
M .D( X ) i 1
n
For the case of frequency distribution it is given as:
k
fi X i X
M .D ( X ) i 1
n
n ~
~
Xi X
M .D ( X ) i 1
n
For the case of frequency distribution it is given as:
k ~
~
fi X i X
M .D( X ) i 1
n
~
Steps to calculate M.D ( X ):
~
1. Find the median, X
~
2. Find the deviations of each reading from X .
3. Find the arithmetic mean of the deviations, ignoring sign.
Page 43 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
X i
ˆ
X
ˆ)
M.D( X i 1
n
k
f i X i Xˆ
M .D ( Xˆ ) i 1
n
Examples:
1. The following are the number of visit made by ten mothers to the local doctor’s surgery.
8, 6, 5, 5, 7, 4, 5, 9, 7, 4
Find mean deviation about mean, median and mode.
Solutions:
First calculate the three averages
~
X 6, X 5.5, Xˆ 5
Then take the deviations of each observation from these averages.
Xi 4 4 5 5 5 6 7 7 8 9 total
X 6
i
2 2 1 1 1 0 1 1 2 3 14
X i 5.5 1.5 1.5 0.5 0.5 0.5 0.5 1.5 1.5 2.5 3.5 14
Xi 5 1 1 0 0 0 1 2 2 3 4 14
10
X i 6)
14
M .D( X ) i 1
1.4
10 10
10
~
X i 5.5
14
M .D( X ) i 1
1.4
10 10
10
X i 5)
14
M .D( Xˆ ) i 1
1.4
10 10
Page 44 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
2. Find mean deviation about mean, median and mode for the following
distributions.(exercise)
Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3
M .D
C.M .D
Average about which deviations are taken
M .D( X )
C.M .D( X )
X
~
~ M .D( X )
C.M .D( X ) ~
X
M .D( Xˆ )
C.M .D( Xˆ )
Xˆ
Example:
Calculate the C.M.D about the mean, median and mode for the data in example 1 above.
Solutions:
M .D
C.M .D
Average about which deviations are taken
M .D( X ) 1.4
C.M .D( X ) 0.233
X 6
~
~ M .D( X ) 1.4
C.M .D( X ) ~ 0.255
X 5.5
M .D( Xˆ ) 1.4
C.M .D( Xˆ ) 0.28
Xˆ 5
Exercise:
Identify the merits and demerits of Mean Deviation
Page 45 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Variance
Population Variance
Page 46 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
1. Find the arithmetic mean.
2. Find the difference between each observation and the mean.
3. Square these differences.
4. Sum the squared differences.
5. Since the data is a sample, divide the number (from step 4 above) by the number of
observations minus one, i.e., n-1 (where n is equal to the number of observations in the
data set).
Examples: Find the variance and standard deviation of the following sample data
1. 5, 17, 12, 10.
2. The data is given in the form of frequency distribution.
No. Class Frequency
1. 40-44 7
2. 45-49 10
3. 50-54 22
4. 55-59 15
5. 60-64 12
6. 65-69 6
7. 70-74 3
Solutions:
1. X 11
Xi 5 10 12 17 Total
(Xi- X) 2 36 1 1 36 74
n
( X i X )2 74
S 2 i 1 24.67.
n 1 3
S S 2 24.67 4.97.
2. X 55
Xi(C.M) 42 47 52 57 62 67 72 Total
fi(Xi- X) 2 1183 640 198 60 588 864 867 4400
n
fi ( X i X )2 4400
S 2 i 1 59.46.
n 1 74
S S 2 59.46 7.71.
Special properties of Standard deviations
Page 47 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
1.
( X i X )2 ( X i A) 2 , A X
n 1 n 1
2. For normal (symmetric distribution the following holds.
Approximately 68.27% of the data values fall within one standard deviation of the
mean. i.e. with in ( X S , X S )
Approximately 95.45% of the data values fall within two standard deviations of the
mean. i.e. with in ( X 2S , X 2S )
Approximately 99.73% of the data values fall within three standard deviations of the
mean. i.e. with in ( X 3S , X 3S )
3. Chebyshev's Theorem
For any data set ,no matter what the pattern of variation, the proportion of the values that
fall with in k standard deviations of the mean or ( X kS, X kS) will be at least
1
1 2 , where k is an number greater than 1. i.e. the proportion of items falling beyond k
k
standard deviations of the mean is at most 1
k2
Example: Suppose a distribution has mean 50 and standard deviation
[Link] percent of the numbers are:
a) Between 38 and 62
b) Between 32 and 68
c) Less than 38 or more than 62.
d) Less than 32 or more than 68.
Solutions:
a) 38 and 62 are at equal distance from the mean,50 and this distance is 12
ks 12
12 12
k 2
S 6
1
Applying the above theorem at least (1 ) *100% 75% of the numbers lie
k2
between 38 and 62.
b) Similarly done.
1
c) It is just the complement of a) i.e. at most *100% 25% of the numbers lie
k2
less than 32 or more than 62.
d) Similarly done.
Example 2:
Page 48 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The average score of a special test of knowledge of wood refinishing has a mean of 53
and standard deviation of 6. Find the range of values in which at least 75% the scores will
lie. (Exercise)
b)
kX1 , kX 2 , .....kX n would be k S
i.e., if Var ( X ) S 2 , then
Var (kX ) k 2 Var ( X ) k 2 S 2
s tan dard deviation k S
c)
a kX1 , a kX 2 , .....a kX n would be k S
i.e., if Var ( X ) S 2 , then
Var (a kX ) Var (a) Var (kX ) 0 k 2 Var ( X ) k 2 S 2
s tan dard deviation k S
Examples:
1. The mean and standard deviation of n Tetracycline Capsules X 1 , X 2 , ..... X n are
known to be 12 gm and 3 gm respectively. New set of capsules of another drug are
obtained by the linear transformation Yi = 2Xi – 0.5 ( i = 1, 2, …, n ) then what will
be the standard deviation of the new set of capsules
2. The mean and the standard deviation of a set of numbers are respectively 500 and 10.
a. If 10 is added to each of the numbers in the set, then what will
be the variance and standard deviation of the new set?
b. If each of the numbers in the set are multiplied by -5, then what
will be the variance and standard deviation of the new set?
Solutions:
1. Using c) above the new standard deviation = |k|S = 2*3 = 6
2. a. They will remain the same. i.e., S = 10 and S 2 10 *10 100
b. New standard deviation = |k|S = |-5|*10 = 5*10 = 50
Page 49 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Coefficient of Variation (C.V)
Is defined as the ratio of standard deviation to the mean usually expressed as percents.
S
C.V *100
X
The distribution having less C.V is said to be less variable or more consistent.
Examples:
1. An analysis of the monthly wages paid (in Birr) to workers in two firms A and B belonging to
the same industry gives the following results
Solutions:
Calculate coefficient of variation for both firms.
SA 10
[Link] *100 *100 19.05%
XA 52.5
S 11
[Link] B *100 *100 23.16%
XB 47.5
Since [Link] < [Link], in firm B there is greater variability in individual wages.
2. A meteorologist interested in the consistency of temperatures in three cities during a given
week collected the following data. The temperatures for the five days of the week in the three
cities were
City 1 25 24 23 26 17
City2 22 21 24 22 20
City3 32 27 35 24 28
Which city have the most consistent temperature, based on these data?
(Exercise)
Page 50 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Z gives the deviations from the mean in units of standard deviation
Z gives the number of standard deviation a particular observation lie above or
below the mean.
It is used to compare two observations coming from different groups.
Examples:
1. Two sections were given introduction to statistics examinations. The following
information was given.
Student A from section 1 scored 90 and student B from section 2 scored [Link]
speaking who performed better?
Solutions:
Calculate the standard score of both students.
X A X 1 90 78
ZA 2
S1 6
X B X 2 95 90
ZB 1
S2 5
Student A performed better relative to his section because the score of student A is
two standard deviation above the mean score of his section while, the score of student B
is only one standard deviation above the mean score of his section.
2. Two groups of people were trained to perform a certain task and tested to find out
which group is faster to learn the task. For the two groups the following information
was given:
Relatively speaking:
a) Which group is more consistent in its performance
b) Suppose a person A from group one take 9.2 minutes while
person B from Group two take 9.3 minutes, who was faster in
performing the task? Why?
Page 51 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
a) Use coefficient of variation.
S1 1.2
C.V1 *100 *100 11.54%
X1 10.4
S 1.3
C.V2 2 *100 *100 10.92%
X2 11.9
Since C.V2 < C.V1, group 2 is more consistent.
b) Calculate the standard score of A and B
X A X 1 9.2 10.4
ZA 1
S1 1.2
X B X 2 9.3 11.9
ZB 2
S2 1.3
Child B is faster because the time taken by child B is two standard deviation shorter
than the average time taken by group 2 while, the time taken by child A is only one
standard deviation shorter than the average time taken by group 1.
Moments
- If X is a variable that assume the values X1, X2,…..,Xn then
1. The rth moment is defined as:
X X 2 ... X n
r r r
X 1
r
n
n
Xi
r
i 1
n
- For the case of frequency distribution this is expressed as:
k
fi X i
r
X r i 1
n
- If r 1,it is the simple arithmetic mean, this is called the first moment.
2. The rth moment about the mean ( the rth central moment)
n n
- Denoted by Mr and defined as:
( X i X )r (n 1) (X i X )r
Mr i 1
i 1
n n n 1
Page 52 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
- For the case of frequency distribution this is expressed as:
k
fi ( X i X )r
M r i 1
n
- If r 2 , it is population variance, this is called the second central moment. If we
assume n 1 n ,it is also the sample variance.
3. The rth moment about any number A is defined as:
'
- Denoted by M r and
n n
(X i A) r
(n 1) (X i A) r
Mr i 1
i 1
'
n n n 1
- For the case of frequency distribution this is expressed as:
k
f i ( X i A) r
M r i 1
'
n
Example:
1. Find the first two moments for the following set of numbers 2, 3, 7
2. Find the first three central moments of the numbers in problem 1
3. Find the third moment about the number 3 of the numbers in problem 1.
Solutions:
1. Use the rth moment formula.
n
Xi
r
X r i 1
n
237
X1 4 X
3
2 2 32 7 2
X
2
20.67
3
2. Use the rth central moment formula.
Page 53 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
n
( X i X )r
M r i 1
n
(2 4) (3 4) (7 4)
M1 0
3
(2 4) 2 (3 4) 2 (7 4) 2
M2 4.67
3
(2 4)3 (3 4)3 (7 4)3
M3 6
th
3
3. Use the r moment about A.
n
( X i A) r
M r i 1
n
(2 3)3 (3 3)3 (7 3)3
M3 21
'
Page 54 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.2.3. Measures of location
[Link]. Skewness
M3 M3 M3
3 , Where is the population s tan dard deviation.
M2
32
( ) 2 32
3
Page 55 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Examples:
1. Suppose the mean, the mode, and the standard deviation of a certain distribution
are 32, 30.5 and 10 respectively. What is the shape of the curve representing the
distribution?
Solutions:
Use the Pearsonian coefficient of skewness
Mean Mode 32 30.5
3 0.15
S tan dard deviation 10
3 0 The distributi on is positively skewed.
2. In a frequency distribution, the coefficient of skewness based on the quartiles is
given to be 0.5. If the sum of the upper and lower quartile is 28 and the median is
11, find the values of the upper and lower quartiles.
Solutions:
~
Given: 3 0.5, X Q2 11 Required: Q1 ,Q3
Q1 Q3 28...........................(*)
Solutions: (exercise)
Page 56 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4. For a moderately skewed frequency distribution, the mean is 10 and the median is
8.5. If the coefficient of variation is 20%, find the Pearsonian coefficient of
skewness and the probable mode of the distribution. (exercise)
5. The sum of fifteen observations, whose mode is 8, was found to be 150 with
coefficient of variation of 20%
(a) Calculate the pearsonian coefficient of skewness and give appropriate
conclusion.
(b) Are smaller values more or less frequent than bigger values for this
distribution?
(c) If a constant k was added on each observation, what will be the new
pearsonian coefficient of skewness? Show your steps. What do you conclude
from this?
Solutions: (exercise)
[Link]. Kurtosis
Page 57 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
M3 60
3 32
0.94 0
a) M2 163 2
The distribution is negatively skewed .
M 4 162
4 2
2 0.6 3
b) M2 16
The curve is platykurtic.
2. The median and the mode of a mesokurtic distribution are 32 and 34 respectively. The
4th moment about the mean is 243. Compute the Pearsonian coefficient of skewness and
identify the type of skewness. Assume (n-1 = n).
3. If the standard deviation of a symmetric distribution is 10, what should be the value of
the fourth moment so that the distribution is mesokurtic?
Solutions (exercise).
Page 58 of 58