0% found this document useful (0 votes)
2 views58 pages

Introduction to Biostatistics Concepts

The document provides an introduction to biostatistics, defining its purpose and distinguishing it from general statistics. It covers key concepts such as statistical populations, samples, data types, and measurement scales, along with methods for data collection and presentation. The chapter emphasizes the importance of understanding data types and appropriate statistical methods for analysis in health sciences.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views58 pages

Introduction to Biostatistics Concepts

The document provides an introduction to biostatistics, defining its purpose and distinguishing it from general statistics. It covers key concepts such as statistical populations, samples, data types, and measurement scales, along with methods for data collection and presentation. The chapter emphasizes the importance of understanding data types and appropriate statistical methods for analysis in health sciences.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

CHAPTER 1
1. INTRODUCTION
1.1. Definition of biostatistics and Classification of statistics
Definition
Statistics: is defined as the science of collecting, organizing, presenting, analyzing
and interpreting numerical data for the purpose of assisting in making a more
effective decision.

Biostatistics: is the branch of applied statistics directed toward applications in the health
sciences and biology. It is sometimes distinguished from the field of biometry based upon
whether applications are in the health sciences (biostatistics) or in broader biology
(biometry; e.g., agriculture, ecology, wildlife biology).

Why biostatistics? What is the difference?


 Because some statistical methods are more heavily used in health applications
than elsewhere (e.g., survival analysis, longitudinal data analysis).
 Because examples are drawn from health sciences.
 Makes subject more appealing to those interested in health.

Definitions of some terms


 Statistical Population: It is the collection of all possible observations of a specified
characteristic of interest (possessing certain common property) and being under
study. An example is all of the students in AAU 3101 course in this term.
 Sample: It is a subset of the population, selected using some sampling technique in
such a way that they represent the population.
 Sampling: The process or method of sample selection from the population.
 Sample size: The number of elements or observation to be included in the sample.
 Census: Complete enumeration or observation of the elements of the population. Or
it is the collection of data from every element in a population
 Parameter: Characteristic or measure obtained from a population.
 Statistic: Characteristic or measure obtained from a sample.
 Variable: It is an item of interest that can take on many different numerical values.

Classification of statistics
Depending on how data can be used, statistics is sometimes divided in to two main areas
or branches.
1) Descriptive Statistics: is concerned with summary calculations, graphs, charts and
tables.
2) Inferential Statistics: is a method used to generalize from a sample to a population.
For example:
The average income of all families (the population) in Ethiopia can be estimated from
figures obtained from a few hundred (the sample) families.
 It is important because statistical data usually arises from sample.
 Statistical techniques based on probability theory are required.

Page 1 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
1.2. Data Types
Data are observations of random variables made on the elements of a population or
sample. Data are the quantities (numbers) or qualities (attributes) measured or observed
that are to be collected and/or analyzed. The word data is plural but datum is singular. A
collection of data is often called a data set (singular)

1.3. Types of Variables


 Qualitative Variables are nonnumeric variables and can't be measured. Examples
include gender, religious affiliation, and state of birth.
 Quantitative Variables are numerical variables and can be measured. Examples include
balance in checking account, number of children in family. Note that quantitative
variables are either discrete (which can assume only certain values, and there are usually
"gaps" between the values, such as the number of bedrooms in your house) or continuous
(which can assume any value within a specific range, such as the air pressure in a tire.)

Scales of measurement
Proper knowledge about the nature and type of data to be dealt with is essential in order
to specify and apply the proper statistical method for their analysis and inferences.
Measurement scale refers to the property of value assigned to the data based on the
properties of order, distance and fixed zero.
In mathematical terms measurement is a functional mapping from the set of objects {Oi}
to the set of real numbers {M(Oi)}.

The goal of measurement systems is to structure the rule for assigning numbers to objects
in such a way that the relationship between the objects is preserved in the numbers
assigned to the objects. The different kinds of relationships preserved are called
properties of the measurement system.

Order
The property of order exists when an object that has more of the attribute than another
object, is given a bigger number by the rule system. This relationship must hold for all
objects in the "real world".
The property of ORDER exists
When for all i, j if Oi > Oj, then M(Oi) > M(Oj).

Page 2 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Distance
The property of distance is concerned with the relationship of differences between
objects. If a measurement system possesses the property of distance it means that the unit
of measurement means the same thing throughout the scale of numbers. That is, an inch
is an inch, no matters were it falls - immediately ahead or a mile downs the road.

More precisely, an equal difference between two numbers reflects an equal difference in
the "real world" between the objects that were assigned the numbers. In order to define
the property of distance in the mathematical notation, four objects are required: Oi, Oj,
Ok, and Ol . The difference between objects is represented by the "-" sign; Oi - Oj refers to
the actual "real world" difference between object i and object j, while M(Oi) - M(Oj)
refers to differences between numbers.

The property of DISTANCE exists, for all i, j, k, l

If Oi-Oj ≥ Ok- Ol then M(Oi)-M(Oj) ≥ M(Ok)-M( Ol ).

Fixed Zero
A measurement system possesses a rational zero (fixed zero) if an object that has none of
the attribute in question is assigned the number zero by the system of rules. The object
does not need to really exist in the "real world", as it is somewhat difficult to visualize a
"man with no height". The requirement for a rational zero is this: if objects with none of
the attribute did exist would they be given the value zero. Defining O0 as the object with
none of the attribute in question, the definition of a rational zero becomes:

The property of FIXED ZERO exists if M(O0) = 0.

The property of fixed zero is necessary for ratios between numbers to be meaningful.

SCALE TYPES
Measurement is the assignment of numbers to objects or events in a systematic fashion.
Four levels of measurement scales are commonly distinguished: nominal, ordinal,
interval, and ratio and each possessed different properties of measurement systems.
Nominal Scales
Nominal scales are measurement systems that possess none of the three properties stated
above.
 Level of measurement which classifies data into mutually exclusive, all inclusive
categories in which no order or ranking can be imposed on the data.
 No arithmetic and relational operation can be applied.
Examples:
o Political party preference (Republican, Democrat, or Other,)
o Sex (Male or Female.)
o Marital status(married, single, widow, divorce)
o Country code
o Regional differentiation of Ethiopia.

Page 3 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Ordinal Scales
Ordinal Scales are measurement systems that possess the property of order, but not the
property of distance. The property of fixed zero is not important if the property of
distance is not satisfied.
 Level of measurement which classifies data into categories that can be ranked.
Differences between the ranks do not exist.
 Arithmetic operations are not applicable but relational operations are applicable.
 Ordering is the sole property of ordinal scale.
Examples:
o Letter grades (A, B, C, D, F).
o Rating scales (Excellent, Very good, Good, Fair, poor).
o Military status.

Interval Scales
Interval scales are measurement systems that possess the properties of Order and
distance, but not the property of fixed zero.
 Level of measurement which classifies data that can be ranked and differences are
meaningful. However, there is no meaningful zero, so ratios are meaningless.
 All arithmetic operations except division are applicable.
 Relational operations are also possible.
Examples:
o IQ
o Temperature in oF.

Ratio Scales
Ratio scales are measurement systems that possess all three properties: order, distance,
and fixed zero. The added power of a fixed zero allows ratios of numbers to be
meaningfully interpreted; i.e. the ratio of Bekele's height to Martha's height is 1.32,
whereas this is not possible with interval scales.
 Level of measurement which classifies data that can be ranked, differences are
meaningful, and there is a true zero. True ratios exist between the different units
of measure.
 All arithmetic and relational operations are applicable.
Examples:
o Weight
o Height
o Number of students
o Age
Operations that make sense for variables of different scales:
No. Scale Operation that make sense
Counting Ranking Addition/ Multiplication/
Subtraction Division
1. Nominal 
2. Ordinal  
3. Interval   
4. Ratio    

Page 4 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Generally, variable types can be distinguished based on their scale_ Typically_ different
statistical methods are appropriate for variables of different scales.
No. Scale Characteristic Question Examples
1. Nominal Is A different than B? Marital status
Eye color
Gender
Religious affiliation
Race
2. Ordinal Is A bigger than B? Stage of disease
Severity of pain
Level of satisfaction
3. Interval By how many units do A and B differ? Temperature
SAT score
4. Ratio How many times bigger than B is A? Distance
Length
Time until death
Weight

The following present a list of different attributes and rules for assigning numbers to
objects. Try to classify the different measurement systems into one of the four types of
scales. (Exercise)

1. Your checking account number as a name for your account.


2. Your checking account balance as a measure of the amount of money you have in
that account.
3. The order in which you were eliminated in a spelling bee as a measure of your
spelling ability.
4. Your score on the first statistics test as a measure of your knowledge of statistics.
5. Your score on an individual intelligence test as a measure of your intelligence.
6. The distance around your forehead measured with a tape measure as a measure of
your intelligence.
7. A response to the statement "Abortion is a woman's right" where "Strongly
Disagree" = 1, "Disagree" = 2, "No Opinion" = 3, "Agree" = 4, and "Strongly
Agree" = 5, as a measure of attitude toward abortion.
8. Times for swimmers to complete a 50-meter race
9. Months of the year Meskerm, Tikimit…
10. Socioeconomic status of a family when classified as low, middle and upper
classes.
11. Blood type of individuals, A, B, AB and O.
12. Pollen counts provided as numbers between 1 and 10 where 1 implies there is
almost no pollen and 10 that it is rampant, but for which the values do not
represent an actual counts of grains of pollen.
13. Regions numbers of Ethiopia (1, 2, 3 etc.)
14. The number of students in a college;
15. the net wages of a group of workers;
16. the height of the men in the same town;

Page 5 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

4.1. Methods of Data Collections and Data Presentations

4.1.1. Introduction to methods of data collection


There are two sources of data:
1. Primary Data
 Data measured or collect by the investigator or the user directly from the
source.
 Two activities involved: planning and measuring.
a) Planning:
 Identify source and elements of the data.
 Decide whether to consider sample or census.
 If sampling is preferred, decide on sample size, selection
method,… etc
 Decide measurement procedure.
 Set up the necessary organizational structure.
b) Measuring: there are different options.
 Focus Group
 Telephone Interview
 Mail Questionnaires
 Door-to-Door Survey
 Mall Intercept
 New Product Registration
 Personal Interview and
 Experiments are some of the sources for collecting the
primary data.
2. Secondary Data
 Data gathered or compiled from published and unpublished
sources or files.
 When our source is secondary data check that:
 The type and objective of the situations.
 The purpose for which the data are collected and
compatible with the present problem.
 The nature and classification of data is appropriate to our
problem.
 There are no biases and misreporting in the published data.
Note:
Data which are primary for one may be secondary for the other.

Page 6 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.1.2. Methods of data presentation
Having collected and edited the data, the next important step is to organize it. That is to
present it in a readily comprehensible condensed form that aids in order to draw
inferences from it. It is also necessary that the like be separated from the unlike ones.

The presentation of data is broadly classified in to the following two categories:


 Tabular presentation
 Diagrammatic and Graphic presentation.
The process of arranging data in to classes or categories according to similarities
technically is called classification.

Classification is a preliminary and it prepares the ground for proper presentation of data.

Definitions:
 Raw data: recorded information in its original collected form, whether it be
counts or measurements, is referred to as raw data.
 Frequency: is the number of values in a specific class of the distribution.
 Frequency distribution: is the organization of raw data in table form using classes
and frequencies.

There are three basic types of frequency distributions


 Categorical frequency distribution
 Ungrouped frequency distribution
 Grouped frequency distribution

There are specific procedures for constructing each type.

1) Categorical frequency Distribution:

Used for data that can be place in specific categories such as nominal, or ordinal. e.g. marital
status.

Example:

A social worker collected the following data on marital status for 25


persons.(M=married, S=single, W=widowed, D=divorced)
M S D W D
S S M M M
W D S M M
W D D S S
S W W D D
Solution:
Since the data are categorical, discrete classes can be used. There are four types of marital
status M, S, D, and W. These types will be used as class for the distribution. We follow
procedure to construct the frequency distribution.

Page 7 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Step 1: Make a table as shown.
Class Tally Frequency Percent
(1) (2) (3) (4)
M
S
D
W
Step 2: Tally the data and place the result in column (2).
Step 3: Count the tally and place the result in column (3).
Step 4: Find the percentages of values in each class by using;
f
%  *100 Where f= frequency of the class, n=total number of value.
n
Percentages are not normally a part of frequency distribution but they can be added since
they are used in certain types diagrammatic such as pie charts.
Step 5: Find the total for column (3) and (4).
Combing the entire steps one can construct the following frequency distribution.
Class Tally Frequency Percent
(1) (2) (3) (4)
M 5 20
////
S //// // 7 28
D //// // 7 28
W //// 6 24
2) Ungrouped frequency Distribution:
It is a table of all the potential raw score values that could possible occur in the data along
with the number of times each actually occurred. It is often constructed for small set or data
on discrete variable.

Constructing ungrouped frequency distribution:


 First find the smallest and largest raw score in the collected data.
 Arrange the data in order of magnitude and count the frequency.
 To facilitate counting one may include a column of tallies.
Example:
The following data represent the mark of 20 students.
80 76 90 85 80
70 60 62 70 85
65 60 63 74 75
76 70 70 80 85

Construct a frequency distribution, which is ungrouped.

Solution:
Step 1: Find the range, Range=Max-Min=90-60=30.
Step 2: Make a table as shown
Step 3: Tally the data.
Step 4: Compute the frequency.

Page 8 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Mark Tally Frequency
60 // 2
62 / 1
63 / 1
65 / 1
70 //// 4
74 / 1
75 // 2
76 / 1
80 /// 3
85 /// 3
90 / 1
Each individual value is presented separately, that is why it is named ungrouped
frequency distribution.
3) Grouped frequency Distribution:
When the range of the data is large, the data must be grouped in to classes that are more than
one unit in width.
Definitions:
 Grouped Frequency Distribution: a frequency distribution when several numbers are
grouped in one class.
 Class limits: Separates one class in a grouped frequency distribution from another. The limits
could actually appear in the data and have gaps between the upper limits of one class and
lower limit of the next.
 Units of measurement (U): the distance between two possible consecutive measures. It is
usually taken as 1, 0.1, 0.01, 0.001, -----.
 Class boundaries: Separates one class in a grouped frequency distribution from another. The
boundaries have one more decimal places than the row data and therefore do not appear in the
data. There is no gap between the upper boundary of one class and lower boundary of the
next class. The lower class boundary is found by subtracting U/2 from the corresponding
lower class limit and the upper class boundary is found by adding U/2 to the corresponding
upper class limit.
 Class width: the difference between the upper and lower class boundaries of any class. It is
also the difference between the lower limits of any two consecutive classes or the difference
between any two consecutive class marks.
 Class mark (Mid points): it is the average of the lower and upper class limits or the average
of upper and lower class boundary.
 Cumulative frequency: is the number of observations less than/more than or equal to a
specific value.
 Cumulative frequency above: it is the total frequency of all values greater than or equal to
the lower class boundary of a given class.
 Cumulative frequency blow: it is the total frequency of all values less than or equal to the
upper class boundary of a given class.
 Cumulative Frequency Distribution (CFD): it is the tabular arrangement of class interval
together with their corresponding cumulative frequencies. It can be more than or less than
type, depending on the type of cumulative frequency used.
 Relative frequency (rf): it is the frequency divided by the total frequency.
 Relative cumulative frequency (rcf): it is the cumulative frequency divided by the total
frequency.

Page 9 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Guidelines for classes

1. There should be between 5 and 20 classes.


2. The classes must be mutually exclusive. This means that no data value can fall
into two different classes
3. The classes must be all inclusive or exhaustive. This means that all data values
must be included.
4. The classes must be continuous. There are no gaps in a frequency distribution.
5. The classes must be equal in width. The exception here is the first or last class. It
is possible to have an "below ..." or "... and above" class. This is often used with
ages.

Steps for constructing Grouped frequency Distribution

1. Find the largest and smallest values


2. Compute the Range(R) = Maximum - Minimum
3. Select the number of classes desired, usually between 5 and 20 or use Sturges rule
k  1 3.32 log n where k is number of classes desired and n is total number of
observation.
4. Find the class width by dividing the range by the number of classes and rounding
R
up, not off. w  .
k
5. Pick a suitable starting point less than or equal to the minimum value. The starting
point is called the lower limit of the first class. Continue to add the class width to
this lower limit to get the rest of the lower limits.
6. To find the upper limit of the first class, subtract U from the lower limit of the
second class. Then continue to add the class width to this upper limit to find the
rest of the upper limits.
7. Find the boundaries by subtracting U/2 units from the lower limits and adding U/2
units from the upper limits. The boundaries are also half-way between the upper
limit of one class and the lower limit of the next class. !may not be necessary to
find the boundaries.
8. Tally the data.
9. Find the frequencies.
10. Find the cumulative frequencies. Depending on what you're trying to accomplish,
it may not be necessary to find the cumulative frequencies.
11. If necessary, find the relative frequencies and/or relative cumulative frequencies

Example*:

Construct a frequency distribution for the following data.

11 29 6 33 14 31 22 27 19 20
18 17 22 38 23 21 26 34 39 27

Page 10 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:

Step 1: Find the highest and the lowest value H=39, L=6

Step 2: Find the range; R=H-L=39-6=33

Step 3: Select the number of classes desired using Sturges formula;

k  1 3.32 log n =1+3.32log (20) =5.32=6(rounding up)

Step 4: Find the class width; w=R/k=33/6=5.5=6 (rounding up)

Step 5: Select the starting point, let it be the minimum observation.

 6, 12, 18, 24, 30, 36 are the lower class limits.

Step 6: Find the upper class limit; e.g. the first upper class=12-U=12-1=11

 11, 17, 23, 29, 35, 41 are the upper class limits.

So combining step 5 and step 6, one can construct the following classes.

Class limits
6 – 11
12 – 17
18 – 23
24 – 29
30 – 35
36 – 41

Step 7: Find the class boundaries;

E.g. for class 1 Lower class boundary=6-U/2=5.5

Upper class boundary =11+U/2=11.5

 Then continue adding w on both boundaries to obtain the rest boundaries. By


doing so one can obtain the following classes.

Class boundary
5.5 – 11.5
11.5 – 17.5
17.5 – 23.5
23.5 – 29.5
29.5 – 35.5
35.5 – 41.5

Page 11 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Step 8: tally the data.

Step 9: Write the numeric values for the tallies in the frequency column.

Step 10: Find cumulative frequency.

Step 11: Find relative frequency or/and relative cumulative frequency.

The complete frequency distribution follows:

Class Class boundary Class Tally Freq. Cf (less Cf (more rf. rcf (less
limit Mark than than type) than type
type)
6 – 11 5.5 – 11.5 8.5 // 2 2 20 0.10 0.10
12 – 17 11.5 – 17.5 14.5 // 2 4 18 0.10 0.20
18 – 23 17.5 – 23.5 20.5 7 11 16 0.35 0.55
//////
24 – 29 23.5 – 29.5 26.5 //// 4 15 9 0.20 0.75
30 – 35 29.5 – 35.5 32.5 /// 3 18 5 0.15 0.90
36 – 41 35.5 – 41.5 38.5 // 2 20 2 0.10 1.00

Diagrammatic and Graphic presentation of data.

These are techniques for presenting data in visual displays using geometric and pictures.

Importance:

 They have greater attraction.


 They facilitate comparison.
 They are easily understandable.

-Diagrams are appropriate for presenting discrete data.

-The three most commonly used diagrammatic presentation for discrete as well as qualitative
data are:

 Pie charts
 Pictogram
 Bar charts

Page 12 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Pie chart

- A pie chart is a circle that is divided in to sections or wedges according to the percentage of
frequencies in each category of the distribution. The angle of the sector is obtained using:

𝑉𝑎𝑙𝑢𝑒𝑠 𝑜𝑓 𝑡ℎ𝑒 𝑝𝑎𝑟𝑡


𝐴𝑛𝑔𝑙𝑒 𝑜𝑓 𝑠𝑒𝑐𝑡𝑜𝑟 = ∗ 360
𝑡ℎ𝑒 𝑤ℎ𝑜𝑙𝑒 𝑞𝑢𝑎𝑛𝑡𝑖𝑡𝑦

Example: Draw a suitable diagram to represent the following population in a town.

Men Women Girls Boys


2500 2000 4000 1500

Solutions:

Step 1: Find the percentage.

Step 2: Find the number of degrees for each class.

Step 3: Using a protractor and compass, graph each section and write its name corresponding
percentage.

Class Frequency Percent Degree


Men 2500 25 90
Women 2000 20 72
Girls 4000 40 144
Boys 1500 15 54

Page 13 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Pictogram

-In these diagram, we represent data by means of some picture symbols. We decide
abut a suitable picture to represent a definite number of units in which the variable is
measured.

Example1: The number of burgers sold in Cafeteria

Example: draw a pictogram to represent the following population of a town.

Year 1989 1990 1991 1992


Population 2000 3000 5000 7000

Solution

Page 14 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Bar Charts:
- A set of bars (thick lines or narrow rectangles) representing some magnitude over time
space.
- They are useful for comparing aggregate over time space.
- Bars can be drawn either vertically or horizontally.
- There are different types of bar charts. The most common being :
 Simple bar chart
 Deviation o0r two way bar chart
 Broken bar chart
 Component or sub divided bar chart.
 Multiple bar charts.

Simple Bar Chart

-Are used to display data on one variable.


-They are thick lines (narrow rectangles) having the same breadth. The magnitude of a quantity
is represented by the height /length of the bar.
Example: The following data represent sale by product, 1957- 1959 of a given company for three
products A, B, C.

Product Sales($) Sales($) Sales($)


In 1957 In 1958 In 1959
A 12 14 18
B 24 21 18
C 24 35 54

Solutions:

Sales by product in 1957

30
25
Sales in $

20
15
10
5
0
A B C
product

Page 15 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Component Bar chart
-When there is a desire to show how a total (or aggregate) is divided in to its component
parts, we use component bar chart.
-The bars represent total value of a variable with each total broken in to its component parts
and different colours or designs are used for identifications
Example:
Draw a component bar chart to represent the sales by product from 1957 to 1959.
Solutions:

SALES BY PRODUCT 1957-1959

100

80
Sales in $

Product C
60
Product B
40
Product A
20

0
1957 1958 1959
Year of production

Multiple Bar charts


- These are used to display data on more than one variable.
- They are used for comparing different variables at the same time.
Example:
Draw a component bar chart to represent the sales by product from 1957 to 1959.
Solutions:

Sales by product 1957-1959

60
50
Sales in $

40 Product A
30 Product B
20 Product C

10
0
1957 1958 1959
Year of production

Page 16 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Graphical Presentation of data
- The histogram, frequency polygon and cumulative frequency graph or ogive are most
commonly applied graphical representation for continuous data.
Procedures for constructing statistical graphs:
 Draw and label the X and Y axes.
 Choose a suitable scale for the frequencies or cumulative frequencies and label it on the Y
axes.
 Represent the class boundaries for the histogram or ogive or the mid points for the
frequency polygon on the X axes.
 Plot the points.
 Draw the bars or lines to connect the points.
Histogram

A graph which displays the data by using vertical bars of various height to represent
frequencies. Class boundaries are placed along the horizontal axes. Class marks and class limits
are some times used as quantity on the X axes.

Example: Construct a histogram to represent the previous data (example *).

Solution

Page 17 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Frequency Polygon:
- A line graph. The frequency is placed along the vertical axis and classes mid points are placed
along the horizontal axis. It is customer to the next higher and lower class interval with
corresponding frequency of zero, this is to make it a complete polygon.
Example:
Draw a frequency polygon for the above data (example *).
Solutions:
8

4
Value Frequency

0
2. 5 8. 5 14.5 20.5 26.5 32.5 38.5 44.5

Class Mid points

Ogive (cumulative frequency polygon)


- A graph showing the cumulative frequency (less than or more than type) plotted against upper
or lower class boundaries respectively. That is class boundaries are plotted along the horizontal
axis and the corresponding cumulative frequencies are plotted along the vertical axis. The
points are joined by a free hand curve.

Example:
Draw an Ogive curve(less than type) for the above data. (Example *)

4.2. Measures of central tendency, measures of Dispersion, and measures of location

Page 18 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

4.2.1. Measures of central tendency


Introduction
 When we want to make comparison between groups of numbers it is good to have a single
value that is considered to be a good representative of each group. This single value is
called the average of the group. Averages are also called measures of central tendency.
 An average which is representative is called typical average and an average which is not
representative and has only a theoretical value is called a descriptive average. A typical
average should posses the following:
 It should be rigidly defined.
 It should be based on all observation under investigation.
 It should be as little as affected by extreme observations.
 It should be capable of further algebraic treatment.
 It should be as little as affected by fluctuations of sampling.
 It should be ease to calculate and simple to understand.
Objectives:
 To comprehend the data easily.
 To facilitate comparison.
 To make further statistical analysis.

The Summation Notation:


 Let X1, X2 ,X3 …XN be a number of measurements where N is the total number of
observation and Xi is ith observation.
 Very often in statistics an algebraic expression of the form X1+X2+X3+...+XN is
used in a formula to compute a statistic. It is tedious to write an expression like this
very often, so mathematicians have developed a shorthand notation to represent a
sum of scores, called the summation notation.
N
 The symbol X
i 1
i is a mathematical shorthand for X1+X2+X3+...+XN

The expression is read, "the sum of X sub i from i equals 1 to N." It means "add up all the
numbers."
Example: Suppose the following were scores made on the first homework assignment for
five students in the class: 5, 7, 7, 6, and 8. In this example set of five numbers, where
N=5, the summation could be written:

The "i=1" in the bottom of the summation notation tells where to begin the sequence of
summation. If the expression were written with "i=3", the summation would start with the
third number in the set. For example:

In the example set of numbers, this would give the following result:

Page 19 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

The "N" in the upper part of the summation notation tells where to end the sequence of
summation. If there were only three scores then the summation and example would be:

Sometimes if the summation notation is used in an expression and the expression must be
written a number of times, as in a proof, then a shorthand notation for the shorthand
notation is employed. When the summation sign "" is used without additional notation,
then "i=1" and "N" are assumed.
For example:

PROPERTIES OF SUMMATION
n
1.  k  nk
i 1
where k is any constant
n n
2.  kX i  k X i where k is any constant
i 1 i 1
n n
3.  (a  bX
i 1
i )  na  b X i
i 1
where a and b are any constant
n n n
4. (X
i 1
i  Yi )   X i   Yi
i 1 i 1

The sum of the product of the two variables could be written:

Example: considering the following data determine


X Y
5 6
7 7
7 8
6 7
8 8
5 5 5

 Xi  ( X i  Yi ) X
2
a) d) g) i
i 1 i 1 i 1
5 5 5 5
b)  Yi
i 1
e)  ( X i  Yi )
i 1
h) ( X i )( Yi )
i 1 i 1
5 5
c) 10
i 1
f) X Y
i 1
i i

Solutions:

Page 20 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
5
a) X
i 1
i  5  7  7  6  8  33
5
b) Y
i 1
i  6  7  8  7  8  36
5
c) 10  5 *10  50
i 1
5
d) (X
i 1
i  Yi )  (5  6)  (7  7)  (7  8)  (6  7)  (8  8)  69  33  36
5
e) (X
i 1
i  Yi )  (5  6)  (7  7)  (7  8)  (6  7)  (8  8)  3  33  36
5
f) X Y
i 1
i i  5 * 6  7 * 7  7 * 8  6 * 7  8 * 8  241
5

X  5 2  7 2  7 2  6 2  8 2  223
2
g) i
i 1
5 5
h) ( X i )( Yi )  33 * 36  1188
i 1 i 1

Types of measures of central tendency


There are several different measures of central tendency; each has its advantage and
disadvantage.
 The Mean (Arithmetic, Geometric and Harmonic)
 The Mode
 The Median
 Quantiles (Quartiles, Deciles and Percentiles)
The choice of these averages depends up on which best fit the property under discussion.
The Arithmetic Mean
 Is defined as the sum of the magnitude of the items divided by the number of
items.
 The mean of X1, X2 ,X3 …Xn is denoted by A.M ,m or X and is given by:

𝑋1 + 𝑋2 + … + 𝑋𝑛
𝑋 =
𝑛

∑𝑛𝑖=1 𝑋𝑖
𝑋 =
𝑛
 If X1 occurs f1 times
 If X2occurs f2 times
 .
 .
 If Xn occurs fn times

Then the mean will be

Page 21 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
k

f X , i i
X  i 1
k

f
i 1
i

k
where k is the number of classes and f
i 1
i n

Example:
Obtain the mean of the following number
2, 7, 8, 2, 7, 3, 7
Solution:
Xi fi Xifi
2 2 4
3 1 3
7 3 21
8 1 8
Total 7 36
4

f i Xi
36
X  i 1
4
  5.15
f
7
i
i 1

Arithmetic Mean for Grouped Data

If data are given in the shape of a continuous frequency distribution, then the mean is
obtained as follows:
∑𝑘𝑖=1 𝑓𝑖 𝑋𝑖
𝑋̅ =
∑𝑘𝑖=1 𝑓𝑖

where Xi =the class mark of the ith class and fi = the frequency of the ith class

Example:
calculate the mean for the following age distribution.
Class Class boundaries frequency
6- 10 5.5 - 10.5 35
11- 15 10.5 - 15.5 23
16- 20 15.5 - 20.5 15
21- 25 20.5 - 25.5 12
26- 30 25.5 - 30.5 9
31- 35 30.5 - 35.5 6

Solutions:
 First find the class marks
 Find the product of frequency and class marks
 Find mean using the formula.

Page 22 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Class Class boundaries fi Xi Xifi
6- 10 5.5 - 10.5 35 8 280
11- 15 10.5 - 15.5 23 13 299
16- 20 15.5 - 20.5 15 18 270
21- 25 20.5 - 25.5 12 23 276
26- 30 25.5 - 30.5 9 28 252
31- 35 30.5 - 35.5 6 33 198
Total 10 1575
0

f i Xi
1575
X  i 1
6
  15.75
f
100
i
i 1

Exercises:
1. Marks of 75 students are summarized in the following frequency distribution:
Marks No. of students
40-44 7
45-49 10
50-54 22
55-59 f4
60-64 f5
65-69 6
70-74 3
If 20% of the students have marks between 55 and 59
i. Find the missing frequencies f4 and f5.
ii. Find the mean.
 If the values in a series or mid values of a class are large enough, coding of values is a
good device
to simplify the calculations.
 For raw data suppose we have used the following coding system.
di  X i  A
 X i  di  A
n n

 Xi  (d i  A)
X  i 1
 i 1

n n
n

d i
 X  A i 1

n
 X  Ad
Where A is an assumed mean and d is the mean of the coded data.

Page 23 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
 If the data are expressed in terms of ungrouped frequency distribution
di  X i  A
 X i  di  A
k k

fX i i  f (d i i  A)
X  i 1
 i 1

n n
k

fd i i
 X  A i 1

n
 X  Ad
 In both cases the true mean is the assumed mean plus the average of the deviations
from the assumed mean.
 Suppose the data is given in the shape of continuous frequency distribution with a
constant class size of w then the following coding is appropriate.
X A
d  i
i w
 X  wd  A
i i
k k
 f X  f ( wd  A)
i i i i
X  i 1  i 1
n n
k
 f wd
i i
 X  A i 1
n
 X  A  wd
Where: Xi is the original class mark for the ith class.
di is the transformed class mark for the ith class.
A is an assumed mean usually the mean of the class marks.
(i =1, 2… k)
Example:
1. Suppose the deviations of the observations from an assumed mean of 7 are:
1, -1, -2, -2, 0, -3, -2, 2, 0, -3.
a) Find the true mean b) Find the original observation
Solutions:
10
A  7,  d i  10
i 1

a)  d   10  1
10
 X  A  d  7 1  6
The true mean is 6.
b) Using Xi=A+di we obtain the following original observations:
8, 6, 5, 5, 7, 4, 5, 9, 7, 4.

Page 24 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Special properties of Arithmetic mean
1. The sum of the deviations of a set of items from their mean is always zero. i.e.
n

 ( X  X )  0.
i 1
i

2. The sum of the squared deviations of a set of items from their mean is the
n n
minimum. i.e.  ( Xi  X ) 2   ( X  A) 2 , A  X
i
i 1 i 1

3. If X 1 is the mean of n1 observations


If X 2 is the mean of n 2 observations
.
.
If X k is the mean of n k observations
Then the mean of all the observation in all groups often called the combined mean
is given by:
k

X1n1  X 2 n 2  .... X k n k 
Xini
Xc   i1k
n1  n 2  ...n k
n
i 1
i

Example:
In a class there are 30 females and 70 males. If females averaged 60 in an examination
and boys averaged 72, find the mean for the entire class.
Solutions:
Females
Males
X 1  60 X 2  72
n1  30 n2  70

X n
X 1 n1  X 2 n2 i 1 i i
Xc   2
n1  n2  nii 1

30(60)  70(72) 6840


 Xc    68.40
30  70 100

4. If a wrong figure has been used when calculating the mean the correct mean can be
obtained with out repeating the whole process using:
(CorrectValue  WrongValue)
CorrectMean  WrongMean
n
Where n is total number of observations.

Example:
An average weight of 10 students was calculated to be [Link] it was discovered
that one weight was misread as 40 instead of 80 k.g. Calculate the correct average
weight.

Page 25 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
(CorrectValue  WrongValue)
CorrectMean  WrongMean
n
(80  40)
CorrectMean  65   65  4  69k.g.
10

5. The effect of transforming original series on the mean.


a) If a constant k is added/ subtracted to/from every observation then the new
mean will be the old mean± k respectively.
b) If every observations are multiplied by a constant k then the new mean will
be k*old mean
Example:
1. The mean of n Tetracycline Capsules X1, X2, …,Xn are known to be 12 gm.
New set of capsules of another drug are obtained by the linear transformation
Yi = 2Xi – 0.5 ( i = 1, 2, …, n ) then what will be the mean of the new set of
capsules
Solutions:
NewMean  2 * OldMean 0.5  2 *12  0.5  23.5

2. The mean of a set of numbers is 500.


a) If 10 is added to each of the numbers in the set, then what will be the mean of
the new set?
b) If each of the numbers in the set are multiplied by -5, then what will be the
mean of the new set?
Solutions:
a).NewMean  OldMean 10  500  10  510
b).NewMean  5 * OldMean 5 * 500  2500

Weighted Mean
 When a proper importance is desired to be given to different data a weighted mean
is appropriate.
 Weights are assigned to each item in proportion to its relative importance.
 Let X1, X2, …Xn be the value of items of a series and W1, W2, …Wn their
corresponding weights , then the weighted mean denoted X w is defined as:
n

X W i i
Xw  i 1
n

W
i 1
i

Example:
A student obtained the following percentage in an examination:
English 60, Biology 75, Mathematics 63, Physics 59, and chemistry [Link] the
students weighted arithmetic mean if weights 1, 2, 1, 3, 3 respectively are allotted
to the subjects.

Page 26 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
5

X W i i
60 * 1  75 * 2  63 * 1  59 * 3  55 * 3 615
Xw  i 1
   61.5
1 2  1 3  3
5
10
W
i 1
i

Merits and Demerits of Arithmetic Mean


Merits:
 It is rigidly defined.
 It is based on all observation.
 It is suitable for further mathematical treatment.
 It is stable average, i.e. it is not affected by fluctuations of sampling to some extent.
 It is easy to calculate and simple to understand.
Demerits:
 It is affected by extreme observations.
 It can not be used in the case of open end classes.
 It can not be determined by the method of inspection.
 It can not be used when dealing with qualitative characteristics, such as intelligence,
honesty, beauty.
 It can be a number which does not exist in a serious.
 Some times it leads to wrong conclusion if the details of the data from which it is
obtained are not available.
 It gives high weight to high extreme values and less weight to low extreme values.

The Geometric Mean

 The geometric mean of a set of n observation is the nth root of their product.
 The geometric mean of X1, X2 ,X3 …Xn is denoted by G.M and given by:
𝐺. 𝑀 = 𝑛√𝑋1 ∗ 𝑋2 ∗ … ∗ 𝑋𝑛

 Taking the logarithms of both sides


1
log(G.M)  log(n X 1 * X 2 * ...* X n )  log(X 1 * X 2 * ...* X n ) n
1 1
 log(G.M)  log(X 1 * X 2 * ....* X n )  (logX 1  log X 2  ...  log X n )
n n
n
1
 log(G.M)   log X i
n i1
 The logarithm of the G.M of a set of observation is the arithmetic mean of
their logarithm.
1 n
 G.M  Anti log(  log X i )
n i1
Example:
Find the G.M of the numbers 2, 4, 8.
Solutions:
G.M  n X1 * X 2 * ...* X n  3 2 * 4 * 8  3 64  4
Remark: The Geometric Mean is useful and appropriate for finding averages of ratios.

Page 27 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Harmonic Mean
The harmonic mean of X1, X2 , X3 …Xn is denoted by H.M and given by:
n
H.M  n , This is called simple harmonic mean.
1

i 1 X i

In a case of frequency distribution:

n k
H.M  , n   fi
k
fi

i 1 X i
i 1

If observations X1, X2, …Xn have weights W1, W2, …Wn respectively, then their
harmonic mean is given by

W i
H.M  n
i 1
, This is called Weighted Harmonic Mean.
W
i 1
i Xi

Remark: The Harmonic Mean is useful and appropriate in finding average speeds and
average rates.

Example: A cyclist pedals from his house to his college at speed of 10 km/hr and back
from the college to his house at 15 km/hr. Find the average speed.
Solution: Here the distance is constant
The simple H.M is appropriate for this problem.
X1= 10km/hr X2=15km/hr
2
H.M   12km /hr
1 1

10 15

The Mode
- Mode is a value which occurs most frequently in a set of values
- The mode may not exist and even if it does exist, it may not be unique.
- In case of discrete distribution the value having the maximum frequency is the model
value.

Examples:
1. Find the mode of 5, 3, 5, 8, 9
Mode =5
2. Find the mode of 8, 9, 9, 7, 8, 2, and 5.
It is a bimodal Data: 8 and 9
3. Find the mode of 4, 12, 3, 6, and 7.
No mode for this data.
- The mode of a set of numbers X1, X2, …Xn is usually denoted by X̂ .

Page 28 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

Mode for Grouped data

If data are given in the shape of continuous frequency distribution, the mode is defined
as:

 1 
X̂  L mo  w 
 1   2 

Where:
Xˆ  the mod e of the distribution
w  the sizeof the mod al class
 1  f mo  f 1
 2  f mo  f 2
f mo  frequencyof the mod al class
f 1  frequencyof the classpreceedingthe mod al class
f 2  frequencyof the class followingthe mod al class

Note: The modal class is a class with the highest frequency.

Example: Following is the distribution of the size of certain farms selected at random
from a district. Calculate the mode of the distribution.

Size of farms No. of farms


5-15 8
15-25 12
25-35 17
35-45 29
45-55 31
55-65 5
65-75 3
Solutions:
45  55 is the mod al class, sin ce it is a class with the highest frequency .

Lmo  45 ˆ  45  10 2 
X
 2  26 
w  10  45.71
 1  f mo  f 1  2
 2  f mo  f 2  26
f mo  31
f 1  29
f2  5

Page 29 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Merits and Demerits of Mode
Merits:
 It is not affected by extreme observations.
 Easy to calculate and simple to understand.
 It can be calculated for distribution with open end class
Demerits:
 It is not rigidly defined.
 It is not based on all observations
 It is not suitable for further mathematical treatment.
 It is not stable average, i.e. it is affected by fluctuations of sampling to
some extent.
 Often its value is not unique.
Note: being the point of maximum density, mode is especially useful in finding the most
popular size in studies relating to marketing, trade, business, and industry. It is the
appropriate average to be used to find the ideal size.

The Median
- In a distribution, median is the value of the variable which divides it in to two
equal halves.
- In an ordered series of data median is an observation lying exactly in the middle
of the series. It is the middle most value in the sense that the number of values
less than the median is equal to the number of values greater than it.
- If X1, X2, …Xn be the observations, then the numbers arranged in ascending
order will be X[1], X[2], …X[n], where X[i] is ith smallest value.
 X[1]< X[2]< …<X[n]
-Median is denoted by 𝑋̃.
Median for ungrouped data
𝑿 𝒏+𝟏 , 𝒊𝒇 𝒏 𝒊𝒔 𝒐𝒅𝒅
[ ]
𝟐
̃=
𝑿
𝟏
(𝑿 𝒏 + 𝑿 𝒏+𝟐 ) , 𝒊𝒇 𝒏 𝒊𝒔 𝒆𝒗𝒆𝒏
{𝟐 [𝟐] [
𝟐
]
Example: Find the median of the following numbers.
a) 6, 5, 2, 8, 9, 4.
b) 2, 1, 8, 3, 5.
Solutions: ~ 1
X  (X n  X n )
a) First order the data: 2, 4, 5, 6, 8, 9 2 [2] [  1]
2
Here n=6 1
 (X [3]  X [ 4 ] )
2
1
 (5  6)  5.5
2
b) Order the data :1, 2, 3, 5, 8 ~ X
X n 1
[ ]
Here n=5 2

 X[3]
3

Page 30 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Median for grouped data

If data are given in the shape of continuous frequency distribution, the median is defined
~ w n
X  L med  (  c)
f med 2
Where :
L med  lower class boundaryof the medianclass.
as: w  the size of the medianclass
n  total numberof observations.
c  the cumulativefrequency(less than type) preceedingthe medianclass.
f med  thefrequency of the medianclass.

Remark:
The median class is the class with the smallest cumulative frequency (less than type) greater
n
than or equal to .
2
Example: Find the median of the following distribution.

Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3

Solutions:
 First find the less than cumulative frequency.
 Identify the median class.
 Find median using formula.

Class Frequency [Link](less


than type)
40-44 7 7
45-49 10 17
50-54 22 39
55-59 15 54
60-64 12 66
65-69 6 72
70-74 3 75

Page 31 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
n 75
  37.5
2 2
39 is the first cumulativefrequencyto be greaterthanor equalto 37.5
 50  54 is the median class.

L  49.5, w  5
med
n  75, c  17, f  22
med

~
 X L  w ( n  c)
med f 2
med
 49.5  5 (37.5  17)
22
 54.16

Merits and Demerits of Median


Merits:
 Median is a positional average and hence not influenced by extreme observations.
 Can be calculated in the case of open end intervals.
 Median can be located even if the data are incomplete.

Demerits:
 It is not a good representative of data if the number of items is small.
 It is not amenable to further algebraic treatment.
 It is susceptible to sampling fluctuations.

Quantiles

When a distribution is arranged in order of magnitude of items, the median is the value of the
middle term. Their measures that depend up on their positions in distribution quartiles, deciles,
and percentiles are collectively called quantiles.

Quartiles:
- Quartiles are measures that divide the frequency distribution in to four equal parts.
- The value of the variables corresponding to these divisions are denoted Q1, Q2, and
Q3 often called the first, the second and the third quartile respectively.
- Q1 is a value which has 25% items which are less than or equal to it. Similarly Q2 has
50%items with value less than or equal to it and Q3 has 75% items whose values are
less than or equal to it.
N 1
- To find Qi (i=1, 2, 3) we count i of the classes beginning from the lowest class.
4

Page 32 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
- For grouped data: we have the following formula
w iN
Q
i  LQ i  f ( 4  c) , i  1,2,3
Qi
Where :
L  lower class boundary of the quartile class.
Qi
w  the size of the quartile class
N  total number of observations.
c  the cumulative frequency (less than type) preceeding the quartile class.
f  thefrequency of the quartile class.
Qi

Remark:
The quartile class (class containing Qi ) is the class with the smallest cumulative frequency
iN
(less than type) greater than or equal to .
4
Deciles:
- Deciles are measures that divide the frequency distribution in to ten equal parts.
- The values of the variables corresponding to these divisions are denoted D1, D2,.. D9
often called the first, the second,…, the ninth decile respectively.
N 1
- To find Di (i=1, 2,..9) we count i of the classes beginning from the lowest class.
10
- For grouped data: we have the following formula
w iN
Di  L Di  (  c) , i  1,2,...,9
f Di 10
Where :
L Di  lower class boundaryof the decile class.
w  the size of the decileclass
N  total numberof observations.
c  the cumulativefrequency( less than type) preceedingthe decile class.
f Di  thefrequency of the decileclass.

Remark: The decile class (class containing Di )is the class with the smallest cumulative
iN
frequency (less than type) greater than or equal to .
10

Page 33 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Percentiles:
- Percentiles are measures that divide the frequency distribution in to hundred equal
parts.
- The values of the variables corresponding to these divisions are denoted P1, P2,.. P99
often called the first, the second,…, the ninety-ninth percentile respectively.
N 1
- To find Pi (i=1, 2,..99) we count i of the classes beginning from the lowest
100
class.

- For grouped data: we have the following formula


w iN
Pi  L Pi  (  c) , i  1,2,...,99
f Pi 100
Where :
L Pi  lower class boundaryof the percentile class.
w  the size of the percentile class
N  total numberof observations.
c  the cumulativefrequency(less than type) preceedingthe percentile class.
f Pi  thefrequency of the percentileclass.

Remark:
The percentile class (class containing Pi )is the class with the smallest cumulative
iN
frequency (less than type) greater than or equal to .
100
Example1: Considering the following data
a) 64,76, 77, 81, 62,64, 63, 70, 81, 72
b) 29, 40, 42, 25, 27, 26, 30, 41, 28
Calculate:
i. All quartiles.
ii. The 5th and 7th deciles
iii. The 50th and 90th percentiles

Solution:
a) Order: 62, 63, 64, 64, 70, 72, 76, 77, 81, 81
here: N = 10
i. All quartiles
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄1 = 1 ( ) 𝑣𝑎𝑙𝑢𝑒 = ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
(2.75) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒 + 0.75(3𝑟𝑑 𝑣𝑎𝑙𝑢𝑒 − 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒)
= 63 + 0.75(64 − 63) = 63 + 0.75(1)
= 63.75

Page 34 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄2 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
(5.5) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71

𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝑄3 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (8.25)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.25(9𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 77 + 0.25(81 − 77)
= 77 + 0.25(4)
= 77 + 1
= 78
ii. The 5th and 7th deciles
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝐷5 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (5.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71
𝑁 + 1 𝑡ℎ 11 𝑡ℎ
𝐷7 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (7.7)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.7(8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 7𝑡ℎ 𝑉𝑎𝑙𝑢𝑒 )
= 76 + 0.7(77 − 76)
= 76 + 0.7(1)
= 76.7

iii. The 50th and 90th percentiles

𝑁 + 1 𝑡ℎ 11 𝑡ℎ 11 𝑡ℎ
𝑃50 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
= (5.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(6𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 5𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 70 + 0.5(72 − 70)
= 70 + 0.5(2)
= 70 + 1
= 71

Page 35 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
𝑁 + 1 𝑡ℎ 11 𝑡ℎ 11 𝑡ℎ
𝑃90 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 9 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
(9.9) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 9𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.9(10𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 9𝑡ℎ 𝑉𝑎𝑙𝑢𝑒 )
= 81 + 0.9(81 − 81)
= 81 + 0.9(0)
= 81

b) Order: 25, 26, 27, 28, 29, 30, 40, 41, 42


here: N = 9
i. All quartiles

𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄1 = 1 ( ) 𝑣𝑎𝑙𝑢𝑒 = ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (2.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒 + 0.5(3𝑟𝑑 𝑣𝑎𝑙𝑢𝑒 − 2𝑛𝑑 𝑣𝑎𝑙𝑢𝑒)
= 26 + 0.5(27 − 26) = 26 + 0.5(1)
= 26.5

𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄2 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒 = 2 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29

𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝑄3 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒 = 3 ( ) 𝑣𝑎𝑙𝑢𝑒
4 4
= (7.5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 + 0.5(8𝑡ℎ 𝑣𝑎𝑙𝑢𝑒 − 7𝑡ℎ 𝑣𝑎𝑙𝑢𝑒)
= 40 + 0.5(41 − 40)
= 40 + 0.5(1)
= 40.5

ii. The 5th and 7th deciles

𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝐷5 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29

𝑁 + 1 𝑡ℎ 10 𝑡ℎ
𝐷7 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒 = 7 ( ) 𝑣𝑎𝑙𝑢𝑒
10 10
(7) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 40

Page 36 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
th th
iii. The 50 and 90 percentiles

𝑁 + 1 𝑡ℎ 10 𝑡ℎ 10 𝑡ℎ
𝑃50 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 50 ( ) 𝑣𝑎𝑙𝑢𝑒 = 5 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
= (5)𝑡ℎ 𝑣𝑎𝑙𝑢𝑒
= 29

𝑁 + 1 𝑡ℎ 10 𝑡ℎ 10 𝑡ℎ
𝑃90 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 90 ( ) 𝑣𝑎𝑙𝑢𝑒 = 9 ( ) 𝑣𝑎𝑙𝑢𝑒
100 100 10
(9) 𝑡ℎ
= 𝑣𝑎𝑙𝑢𝑒
= 42

Example2: Considering the following distribution


Calculate:
a) All quartiles.
b) The 7th decile.
c) The 90th percentile.

Values Frequency
140- 150 17
150- 160 29
160- 170 42
170- 180 72
180- 190 84
190- 200 107
200- 210 49
210- 220 34
220- 230 31
230- 240 16
240- 250 12

Solutions:
 First find the less than cumulative frequency.
 Use the formula to calculate the required quantile.
Values Frequency [Link](less than type)
140- 150 17 17
150- 160 29 46
160- 170 42 88
170- 180 72 160
180- 190 84 244
190- 200 107 351
200- 210 49 400
210- 220 34 434
220- 230 31 465
230- 240 16 481
240- 250 12 493

Page 37 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
a) Quartiles:
i. Q1
- determine the class containing the first quartile.
N
 123.25
4
 170  180 is the classcontainingthe first quartile.

LQ  170 , w 10 w N
1  Q1  LQ 1  (  c)
N  493 , c  88 , f Q  72 f Q1 4
1

10
 170  (123.25  88)
72
 174.90

ii. Q2
- determine the class containing the second quartile.
2* N
 246.5
4
 190  200 is the class containing the sec ond quartile.

LQ  190 , w 10 w 2* N
2
 Q2  LQ2  (  c)
N  493 , c  244 , f Q 107
2
f Q2 4
10
 190  (246.5  244)
107
 190.23

iii. Q3
- determine the class containing the third quartile.
3* N
 369.75
4
 200  210 is the class containing the third quartile.

LQ  200 ,
3
w 10  Q3  LQ 3 
w 3* N
(  c)
f Q3 4
N  493 , c  351 , f Q  49
3
10
 200  (369.75  351)
49
 203.83

Page 38 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
b) D7
- determine the class containing the 7th decile.
7* N
 345.1
10
190  200 is the class containing the seventh decile.

LD  190 , w 10 w 7* N
7
 D7  LD  (  c)
N  493 , c  244 , f D 107 f D 10
7

7
7

10
 190  (345.1  244)
107
 199.45

c) P90
- determine the class containing the 90th percentile.
90 * N
 443.7
100
 220  230 is the class containing the 90th percentile.

L P90  220 , w  10 w 90 * N
 P90  LP  (  c)
N  493 , c  434 , f P90  31 f P 100
90

90

10
 220  (443.7  434)
31
 223.13

Page 39 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.2.2. Measures of Dispersion (Variation)

Introduction and objectives of measuring Variation

-The scatter or spread of items of a distribution is known as dispersion or variation. In


other words the degree to which numerical data tend to spread about an average value is
called dispersion or variation of the data.
-Measures of dispersions are statistical measures which provide ways of measuring the
extent in which data are dispersed or spread out.
Objectives of measuring Variation:

 To judge the reliability of measures of central tendency


 To control variability itself.
 To compare two or more groups of numbers in terms of their variability.
 To make further statistical analysis.
Absolute and Relative Measures of Dispersion

The measures of dispersion which are expressed in terms of the original unit of a series
are termed as absolute measures. Such measures are not suitable for comparing the
variability of two distributions which are expressed in different units of measurement and
different average size. Relative measures of dispersions are a ratio or percentage of a
measure of absolute dispersion to an appropriate measure of central tendency and are thus
pure numbers independent of the units of measurement. For comparing the variability of
two distributions (even if they are measured in the same unit), we compute the relative
measure of dispersion instead of absolute measures of dispersion.

Types of Measures of Dispersion

Various measures of dispersions are in use. The most commonly used measures of
dispersions are:
1) Range and relative range
2) Quartile deviation and coefficient of Quartile deviation
3) Mean deviation and coefficient of Mean deviation
4) Standard deviation and coefficient of variation.

The Range (R)


The range is the largest score minus the smallest score. It is a quick and dirty measure of
variability, although when a test is given back to students they very often wish to know
the range of scores. Because the range is greatly affected by extreme scores, it may give a
distorted picture of the scores. The following two distributions have the same range, 13,
yet appear to differ greatly in the amount of variability.
Distribution 1: 32 35 36 36 37 38 40 42 42 43 43 45
Distribution 2: 32 32 33 33 33 34 34 34 34 34 35 45
For this reason, among others, the range is not the most important measure of variability.
R LS , L  l arg est observation
S  smallest observation

Page 40 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Range for grouped data:
If data are given in the shape of continuous frequency distribution, the range is computed
as:

R  UCLk  LCL1 , UCLk is upperclass lim it of the last class.


UCL1 is lower class lim it of the first class.
This is some times expressed as:
R  X k  X1 , X k is class mark of the last class.
X 1 is classmark of the first class.
Merits and Demerits of range

Merits:
 It is rigidly defined.
 It is easy to calculate and simple to understand.
Demerits:
 It is not based on all observation.
 It is highly affected by extreme observations.
 It is affected by fluctuation in sampling.
 It is not liable to further algebraic treatment.
 It can not be computed in the case of open end distribution.
 It is very sensitive to the size of the sample.
Relative Range (RR)
-it is also some times called coefficient of range and given by:
LS R
RR  
LS LS
Example:
1. Find the relative range of the above two distribution.(exercise!)
2. If the range and relative range of a series are 4 and 0.25 respectively. Then what is the
value of:
a) Smallest observation
b) Largest observation
Solutions :( 2)
R  4  L  S  4 _________________(1)
RR  0.25  L  S  16 _____________(2)
Solving (1) and (2) at the same time , one can obtain the following value
L  10 and S  6

Page 41 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Quartile Deviation (Semi-inter quartile range), Q.D

The inter quartile range is the difference between the third and the first quartiles of a set of
items and semi-inter quartile range is half of the inter quartile range.
Q3  Q1
Q.D 
2
Coefficient of Quartile Deviation (C.Q.D)

(Q3  Q1 2 2 * Q.D Q3  Q1
C. Q.D   
(Q3  Q1 ) 2 Q3  Q1 Q3  Q1
 It gives the average amount by which the two quartiles differ from the median.

Example: Compute Q.D and its coefficient for the following distribution.

Values Frequency
140- 150 17
150- 160 29
160- 170 42
170- 180 72
180- 190 84
190- 200 107
200- 210 49
210- 220 34
220- 230 31
230- 240 16
240- 250 12

Solutions:
In the previous chapter we have obtained the values of all quartiles as:
Q1= 174.90, Q2= 190.23, Q3=203.83

Q3  Q1 203.83  174.90
 Q.D    14.47
2 2
2 * Q.D 2 *14.47
C.Q.D    0.076
Q3  Q1 203.83  174.90
Remark: Q.D or C.Q.D includes only the middle 50% of the observation.

Page 42 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Mean Deviation (M.D):

The mean deviation of a set of items is defined as the arithmetic mean of the values of the
absolute deviations from a given average. Depending up on the type of averages used we
have different mean deviations.
a) Mean Deviation about the mean
 Denoted by M.D( X ) and given by

n
 Xi  X
M .D( X )  i 1
n
 For the case of frequency distribution it is given as:

k
 fi X i  X
M .D ( X )  i 1
n

Steps to calculate M.D ( X ):


1. Find the arithmetic mean, X
2. Find the deviations of each reading from X .
3. Find the arithmetic mean of the deviations, ignoring sign.

b) Mean Deviation about the median.


~
 Denoted by M.D( X ) and given by

n ~
~
 Xi  X
M .D ( X )  i 1
n
 For the case of frequency distribution it is given as:

k ~
~
 fi X i  X
M .D( X )  i 1
n
~
Steps to calculate M.D ( X ):
~
1. Find the median, X
~
2. Find the deviations of each reading from X .
3. Find the arithmetic mean of the deviations, ignoring sign.

Page 43 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

c) Mean Deviation about the mode.


 Denoted by M.D( X̂ ) and given by

X i
ˆ
X
ˆ)
M.D( X i 1
n

 For the case of frequency distribution it is given as:

k
 f i X i  Xˆ
M .D ( Xˆ )  i 1
n

Steps to calculate M.D ( X̂ ):


1. Find the mode, X̂
2. Find the deviations of each reading from X̂ .
3. Find the arithmetic mean of the deviations, ignoring sign.

Examples:
1. The following are the number of visit made by ten mothers to the local doctor’s surgery.
8, 6, 5, 5, 7, 4, 5, 9, 7, 4
Find mean deviation about mean, median and mode.

Solutions:
First calculate the three averages
~
X  6, X  5.5, Xˆ  5
Then take the deviations of each observation from these averages.
Xi 4 4 5 5 5 6 7 7 8 9 total
X 6
i
2 2 1 1 1 0 1 1 2 3 14

X i  5.5 1.5 1.5 0.5 0.5 0.5 0.5 1.5 1.5 2.5 3.5 14

Xi  5 1 1 0 0 0 1 2 2 3 4 14
10

X i  6)
14
 M .D( X )  i 1
  1.4
10 10
10

~
X i  5.5
14
M .D( X )  i 1
  1.4
10 10
10

X i  5)
14
M .D( Xˆ )  i 1
  1.4
10 10

Page 44 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
2. Find mean deviation about mean, median and mode for the following
distributions.(exercise)

Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3

Remark: Mean deviation is always minimum about the median.

Coefficient of Mean Deviation (C.M.D)

M .D
C.M .D 
Average about which deviations are taken

M .D( X )
 C.M .D( X ) 
X
~
~ M .D( X )
C.M .D( X )  ~
X

M .D( Xˆ )
C.M .D( Xˆ ) 

Example:
Calculate the C.M.D about the mean, median and mode for the data in example 1 above.

Solutions:
M .D
C.M .D 
Average about which deviations are taken

M .D( X ) 1.4
 C.M .D( X )    0.233
X 6
~
~ M .D( X ) 1.4
C.M .D( X )  ~   0.255
X 5.5
M .D( Xˆ ) 1.4
C.M .D( Xˆ )    0.28
Xˆ 5
Exercise:
Identify the merits and demerits of Mean Deviation

Page 45 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The Variance

Population Variance

If we divide the variation by the number of values in the population, we get


something called the population variance. This variance is the "average squared deviation
from the mean".
1
Population Varince   2   ( X i   ) 2 , i  1,2,.....N
N
 For the case of frequency distribution it is expressed as:
1
Population Varince   2   f i ( X i   ) 2 , i  1,2,.....k
N
Sample Variance
One would expect the sample variance to simply be the population variance with the
population mean replaced by the sample mean. However, one of the major uses of
statistics is to estimate the corresponding parameter. This formula has the problem that
the estimated value isn't the same as the parameter. To counteract this, the sum of the
squares of the deviations is divided by one less than the sample size.
1
Sample Varince  S 2   ( X i  X ) 2 , i  1,2,....., n
n 1
 For the case of frequency distribution it is expressed as:
1
Sample Varince  S 2   fi ( X i  X ) 2 , i  1,2,.....k
n 1
We usually use the following short cut formula.
n
 X i  nX 2
2

S 2  i 1 , for raw data.


n 1
k
 f i X i  nX 2
2

S 2  i 1 , for frequency distributi on.


n 1
Standard Deviation
There is a problem with variances. Recall that the deviations were squared. That means
that the units were also squared. To get the units back the same as the original data
values, the square root must be taken.

Population s tan dard deviation     2


Sample s tan dard deviation  s  S 2
The following steps are used to calculate the sample variance:

Page 46 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
1. Find the arithmetic mean.
2. Find the difference between each observation and the mean.
3. Square these differences.
4. Sum the squared differences.
5. Since the data is a sample, divide the number (from step 4 above) by the number of
observations minus one, i.e., n-1 (where n is equal to the number of observations in the
data set).

Examples: Find the variance and standard deviation of the following sample data
1. 5, 17, 12, 10.
2. The data is given in the form of frequency distribution.
No. Class Frequency
1. 40-44 7
2. 45-49 10
3. 50-54 22
4. 55-59 15
5. 60-64 12
6. 65-69 6
7. 70-74 3

Solutions:
1. X  11
Xi 5 10 12 17 Total
(Xi- X) 2 36 1 1 36 74

n
 ( X i  X )2 74
 S 2  i 1   24.67.
n 1 3
 S  S 2  24.67  4.97.

2. X  55
Xi(C.M) 42 47 52 57 62 67 72 Total
fi(Xi- X) 2 1183 640 198 60 588 864 867 4400

n
 fi ( X i  X )2 4400
 S 2  i 1   59.46.
n 1 74
 S  S 2  59.46  7.71.
Special properties of Standard deviations

Page 47 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction

1.
 ( X i  X )2   ( X i  A) 2 , A  X
n 1 n 1
2. For normal (symmetric distribution the following holds.
 Approximately 68.27% of the data values fall within one standard deviation of the
mean. i.e. with in ( X  S , X  S )
 Approximately 95.45% of the data values fall within two standard deviations of the
mean. i.e. with in ( X  2S , X  2S )
 Approximately 99.73% of the data values fall within three standard deviations of the
mean. i.e. with in ( X  3S , X  3S )

3. Chebyshev's Theorem
For any data set ,no matter what the pattern of variation, the proportion of the values that
fall with in k standard deviations of the mean or ( X  kS, X  kS) will be at least
1
1  2 , where k is an number greater than 1. i.e. the proportion of items falling beyond k
k
standard deviations of the mean is at most 1
k2
Example: Suppose a distribution has mean 50 and standard deviation
[Link] percent of the numbers are:
a) Between 38 and 62
b) Between 32 and 68
c) Less than 38 or more than 62.
d) Less than 32 or more than 68.
Solutions:
a) 38 and 62 are at equal distance from the mean,50 and this distance is 12
 ks  12
12 12
k   2
S 6

1
 Applying the above theorem at least (1  ) *100%  75% of the numbers lie
k2
between 38 and 62.

b) Similarly done.
1
c) It is just the complement of a) i.e. at most *100%  25% of the numbers lie
k2
less than 32 or more than 62.
d) Similarly done.

Example 2:

Page 48 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
The average score of a special test of knowledge of wood refinishing has a mean of 53
and standard deviation of 6. Find the range of values in which at least 75% the scores will
lie. (Exercise)

4. If the standard deviation of X 1 , X 2 , ..... X n is S , then the standard deviation of


a) X  k , X  k , ..... X  k will also be S
1 2 n

i.e., if Var ( X )  S , then


2

 Var ( X  k )  Var ( X )  Var (k )  S 2  0


 s tan dard deviation  S

b)
kX1 , kX 2 , .....kX n would be k S
i.e., if Var ( X )  S 2 , then
 Var (kX )  k 2 Var ( X )  k 2 S 2
 s tan dard deviation  k S
c)
a  kX1 , a  kX 2 , .....a  kX n would be k S
i.e., if Var ( X )  S 2 , then
 Var (a  kX )  Var (a)  Var (kX )  0  k 2 Var ( X )  k 2 S 2
 s tan dard deviation  k S
Examples:
1. The mean and standard deviation of n Tetracycline Capsules X 1 , X 2 , ..... X n are
known to be 12 gm and 3 gm respectively. New set of capsules of another drug are
obtained by the linear transformation Yi = 2Xi – 0.5 ( i = 1, 2, …, n ) then what will
be the standard deviation of the new set of capsules
2. The mean and the standard deviation of a set of numbers are respectively 500 and 10.
a. If 10 is added to each of the numbers in the set, then what will
be the variance and standard deviation of the new set?
b. If each of the numbers in the set are multiplied by -5, then what
will be the variance and standard deviation of the new set?
Solutions:
1. Using c) above the new standard deviation = |k|S = 2*3 = 6
2. a. They will remain the same. i.e., S = 10 and S 2  10 *10  100
b. New standard deviation = |k|S = |-5|*10 = 5*10 = 50

Page 49 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Coefficient of Variation (C.V)

 Is defined as the ratio of standard deviation to the mean usually expressed as percents.
S
C.V  *100
X
 The distribution having less C.V is said to be less variable or more consistent.
Examples:
1. An analysis of the monthly wages paid (in Birr) to workers in two firms A and B belonging to
the same industry gives the following results

Value Firm A Firm B


Mean wage 52.5 47.5
Median wage 50.5 45.5
Variance 100 121

In which firm A or B is there greater variability in individual wages?

Solutions:
Calculate coefficient of variation for both firms.
SA 10
[Link]  *100  *100  19.05%
XA 52.5
S 11
[Link]  B *100  *100  23.16%
XB 47.5
Since [Link] < [Link], in firm B there is greater variability in individual wages.
2. A meteorologist interested in the consistency of temperatures in three cities during a given
week collected the following data. The temperatures for the five days of the week in the three
cities were

City 1 25 24 23 26 17
City2 22 21 24 22 20
City3 32 27 35 24 28

Which city have the most consistent temperature, based on these data?
(Exercise)

Standard Scores (Z-scores)


 If X is a measurement from a distribution with mean X and standard
deviation S, then its value in standard units is
X 
Z , for population.

X X
Z , for sample
S

Page 50 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Z gives the deviations from the mean in units of standard deviation
Z gives the number of standard deviation a particular observation lie above or
below the mean.
 It is used to compare two observations coming from different groups.
Examples:
1. Two sections were given introduction to statistics examinations. The following
information was given.

Value Section 1 Section 2


Mean 78 90
[Link] 6 5

Student A from section 1 scored 90 and student B from section 2 scored [Link]
speaking who performed better?

Solutions:
Calculate the standard score of both students.
X A  X 1 90  78
ZA   2
S1 6
X B  X 2 95  90
ZB   1
S2 5
 Student A performed better relative to his section because the score of student A is
two standard deviation above the mean score of his section while, the score of student B
is only one standard deviation above the mean score of his section.
2. Two groups of people were trained to perform a certain task and tested to find out
which group is faster to learn the task. For the two groups the following information
was given:

Value Group one Group two

Mean 10.4 min 11.9 min

[Link]. 1.2 min 1.3 min

Relatively speaking:
a) Which group is more consistent in its performance
b) Suppose a person A from group one take 9.2 minutes while
person B from Group two take 9.3 minutes, who was faster in
performing the task? Why?

Page 51 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
a) Use coefficient of variation.
S1 1.2
C.V1  *100  *100  11.54%
X1 10.4
S 1.3
C.V2  2 *100  *100  10.92%
X2 11.9
Since C.V2 < C.V1, group 2 is more consistent.
b) Calculate the standard score of A and B

X A  X 1 9.2  10.4
ZA    1
S1 1.2
X B  X 2 9.3  11.9
ZB    2
S2 1.3
Child B is faster because the time taken by child B is two standard deviation shorter
than the average time taken by group 2 while, the time taken by child A is only one
standard deviation shorter than the average time taken by group 1.

Moments
- If X is a variable that assume the values X1, X2,…..,Xn then
1. The rth moment is defined as:
X  X 2  ...  X n
r r r
X  1
r
n
n
 Xi
r

 i 1
n
- For the case of frequency distribution this is expressed as:
k
 fi X i
r

X r  i 1
n
- If r  1,it is the simple arithmetic mean, this is called the first moment.
2. The rth moment about the mean ( the rth central moment)
n n
- Denoted by Mr and defined as:
 ( X i  X )r (n  1) (X i  X )r
Mr  i 1
 i 1

n n n 1

Page 52 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
- For the case of frequency distribution this is expressed as:
k
 fi ( X i  X )r
M r  i 1
n
- If r  2 , it is population variance, this is called the second central moment. If we
assume n 1  n ,it is also the sample variance.
3. The rth moment about any number A is defined as:

'
- Denoted by M r and
n n

(X i  A) r
(n  1) (X i  A) r
Mr  i 1
 i 1
'

n n n 1
- For the case of frequency distribution this is expressed as:
k
 f i ( X i  A) r
M r  i 1
'

n
Example:
1. Find the first two moments for the following set of numbers 2, 3, 7
2. Find the first three central moments of the numbers in problem 1
3. Find the third moment about the number 3 of the numbers in problem 1.
Solutions:
1. Use the rth moment formula.
n
 Xi
r

X r  i 1
n
237
 X1  4 X
3
2 2  32  7 2
X 
2
 20.67
3
2. Use the rth central moment formula.

Page 53 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
n
 ( X i  X )r
M r  i 1
n
(2  4)  (3  4)  (7  4)
 M1  0
3
(2  4) 2  (3  4) 2  (7  4) 2
M2   4.67
3
(2  4)3  (3  4)3  (7  4)3
M3  6
th
3
3. Use the r moment about A.
n
 ( X i  A) r
M r  i 1
n
(2  3)3  (3  3)3  (7  3)3
 M3   21
'

Page 54 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4.2.3. Measures of location
[Link]. Skewness

- Skewness is the degree of asymmetry or departure from symmetry of a distribution.


- A skewed frequency distribution is one that is not symmetrical.
- Skewness is concerned with the shape of the curve not size.
- If the frequency curve (smoothed frequency polygon) of a distribution has a longer tail
to the right of the central maximum than to the left, the distribution is said to be
skewed to the right or said to have positive skewness. If it has a longer tail to the left of
the central maximum than to the right, it is said to be skewed to the left or said to have
negative skewness.
- For moderately skewed distribution, the following relation holds among the three
commonly used measures of central tendency.
Mean  Mode  3 * (Mean  Median )
Measures of Skewness
- Denoted by  3
- There are various measures of skewness.
1. The Pearsonian coefficient of skewness
~
Mean  Mode X  Xˆ 3  X  Xˆ 
3     

S tan dard deviation S 2 S 
2. The Bowley’s coefficient of skewness ( coefficient of skewness based on quartiles)
(Q3  Q2 )  (Q2  Q1 ) Q3  Q1  2Q2
3  
Q3  Q1 Q3  Q1
3. The moment coefficient of skewness

M3 M3 M3
3    , Where  is the population s tan dard deviation.
M2
32
( ) 2 32
 3

The shape of the curve is determined by the value of  3


 If  3  0 then the distribution is positively skewed.
 If  3  0 then the distributi on is symmetric.
 If  3  0 then the distribution is negatively skewed.
Remark:
o In a positively skewed distribution, smaller observations are more
frequent than larger observations. i.e. the majority of the
observations have a value below an average.
o In a negatively skewed distribution, smaller observations are less
frequent than larger observations. i.e. the majority of the
observations have a value above an average.

Page 55 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Examples:
1. Suppose the mean, the mode, and the standard deviation of a certain distribution
are 32, 30.5 and 10 respectively. What is the shape of the curve representing the
distribution?
Solutions:
Use the Pearsonian coefficient of skewness
Mean  Mode 32  30.5
3    0.15
S tan dard deviation 10
 3  0  The distributi on is positively skewed.
2. In a frequency distribution, the coefficient of skewness based on the quartiles is
given to be 0.5. If the sum of the upper and lower quartile is 28 and the median is
11, find the values of the upper and lower quartiles.
Solutions:
~
Given:  3  0.5, X  Q2  11 Required: Q1 ,Q3

Q1  Q3  28...........................(*)

(Q3  Q2 )  (Q2  Q1 ) Q3  Q1  2Q2


3    0.5
Q3  Q1 Q3  Q1
Substituting the given values , one can obtain the following
Q3  Q1  12...................................(**)
Solving (*) and (**) at the same time we obtain the following values
Q1  8 and Q3  20

3. Some characteristics of annually family income distribution (in Birr) in two


regions is as follows:
Region Mean Median Standard Deviation
A 6250 5100 960
B 6980 5500 940
a) Calculate coefficient of skewness for each region
b) For which region is, the income distribution more skewed. Give your
interpretation for this Region
c) For which region is the income more consistent?

Solutions: (exercise)

Page 56 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
4. For a moderately skewed frequency distribution, the mean is 10 and the median is
8.5. If the coefficient of variation is 20%, find the Pearsonian coefficient of
skewness and the probable mode of the distribution. (exercise)

5. The sum of fifteen observations, whose mode is 8, was found to be 150 with
coefficient of variation of 20%
(a) Calculate the pearsonian coefficient of skewness and give appropriate
conclusion.
(b) Are smaller values more or less frequent than bigger values for this
distribution?
(c) If a constant k was added on each observation, what will be the new
pearsonian coefficient of skewness? Show your steps. What do you conclude
from this?
Solutions: (exercise)

[Link]. Kurtosis

Kurtosis is the degree of peakdness of a distribution, usally taken relative to a


normal distribution. A distribution having relatively high peak is called leptokurtic. If
a curve representing a distribution is flat topped, it is called platykurtic. The normal
distribution which is not very high peaked or flat topped is called mesokurtic.
Measures of kurtosis
The moment coefficient of kurtosis:
 Denoted by  4 and given by
M4 M4
4   4
M2
2

Where : M 4 is the fourth moment about the mean.
M 2 is the sec ond moment about the mean.
 is the population s tan dard deviation.
The peakdness depends on the value of  4 .
 If  4  3 then the curve is leptokurtic.
 If  4  3 then the curve is mesokurtic.
  If  4  3 then the curveis platykurtic.
Examples:
1. If the first four central moments of a distribution are:
M1  0, M 2  16, M 3  60, M 4  162
a) Compute a measure of skewness
b) Compute a measure of kurtosis and give your interpretation.

Page 57 of 58
Lecture notes on Biostatistics (EaBC 636) Chapter 1: Introduction
Solutions:
M3  60
3  32
  0.94  0
a) M2 163 2
 The distribution is negatively skewed .

M 4 162
4  2
 2  0.6  3
b) M2 16
 The curve is platykurtic.
2. The median and the mode of a mesokurtic distribution are 32 and 34 respectively. The
4th moment about the mean is 243. Compute the Pearsonian coefficient of skewness and
identify the type of skewness. Assume (n-1 = n).

3. If the standard deviation of a symmetric distribution is 10, what should be the value of
the fourth moment so that the distribution is mesokurtic?
Solutions (exercise).

Page 58 of 58

You might also like