0% found this document useful (0 votes)
27 views170 pages

Basic Biostatistics for MSc Students

Uploaded by

sayihmehari74
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
27 views170 pages

Basic Biostatistics for MSc Students

Uploaded by

sayihmehari74
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

University of Gondar

college of medicine and health science


department of Epidemiology and Biostatistics

Basic biostatistics for Msc. students

By Teresa Kisi (Bsc., MPH) 1


Course content

 Chapter one: Introduction to Biostatistics


Type of biostatistics
Types of variables
Methods of data organization and presentation
Measures of Central Tendency
Measures of Variation
 Chapter two: Probability and probability distribution
– Mutually exclusive events and the additive law
– Conditional Probability and the multiplicative law
– Random variablesByand probability distributions 2
Teresa K. (Bsc., MPH)
Course content cont’d…
 Chapter three: Sampling techniques and sample size determination
– Common terms used in sampling
– Sampling methods
– Errors in Sampling
– sample size determination
 Chapter four: Estimation and Hypothesis Testing
 Point estimation
 Sampling distribution of means
 Interval estimation (large samples)
 The Null and Alternative Hypotheses
 Level of significance
 Tests of significance on means and proportions (large samples)
 One tailed tests
 Comparing the means of small samples
 Confidence interval or P-value?
By Teresa K. (Bsc., MPH) 3
Course content cont’d…

 Chapter five: Test of association and independence


 Chapter six: Analysis of variance (ANOVA)
 Chapter seven: Correlation and linear regression
 Chapter eight: Logistic regression
 Chapter nine: survival analysis

By Teresa K. (Bsc., MPH) 4


Chapter one:
Introduction to Biostatistics

Objectives of the chapter


 After completing this chapter, we will be able to:
– Define Statistics and Biostatistics
– Enumerate the importance and limitations of
statistics
– Define and Identify the different types of variable
and list why we need to classify variables

By Teresa K. (Bsc., MPH) 5


Objectives cont’d…

– Identify the different methods of medical and


biological data organization and presentation

– Identify the criterion for the selection of a method


to organize and present data

By Teresa K. (Bsc., MPH) 6


Statistics?

By Teresa K. (Bsc., MPH) 7


Statistics

 The science of assembling and interpreting numerical


data (Bland, 2000)
 The discipline concerned with the treatment of
numerical data derived from groups of individuals
(Armitage et al., 2001).
 Generally the term statistics is used to mean either
statistical data or statistical methods.

By Teresa K. (Bsc., MPH) 8


Statistics cont’d…

Statistical data: refers to numerical


descriptions of things. These descriptions may
take the form of counts or measurements.

E.g. statistics of malaria cases include fever


cases, number of positives obtained, sex and
age distribution of positive cases, etc.

By Teresa K. (Bsc., MPH) 9


Statistics cont’d…

 NB: Even though statistical data always denote


figures (numerical descriptions), it must be
remembered that all 'numerical descriptions' are
not statistical data.

Why?

By Teresa K. (Bsc., MPH) 10


Statistics cont’d…

 Statistical methods: refers methods that are used


for collecting, organising, analyzing and
interpreting numerical data for understanding a
phenomenon or making wise decisions. In this sense
it is a branch of scientific method and helps us to
know in a better way the objective under study.

By Teresa K. (Bsc., MPH) 11


Biostatistics?

By Teresa K. (Bsc., MPH) 12


 Biostatistics: The tools of statistics are employed in
many fields - business, education, psychology,
agriculture, and economics, to mention only few.

 When the data being analyzed are derived from the


public health data, biological sciences and medicine,
we use the term biostatistics to distinguish this
particular application of statistical tools and
concepts.

By Teresa K. (Bsc., MPH) 13


–Types of biostatistics?

By Teresa K. (Bsc., MPH) 14


Types of biostatistics

Biostatistics

Descriptive Statistics Inferential Statistics

collection making inferences


organizing hypothesis testing
summarizing determining relationship
presenting of data making the prediction

By Teresa K. (Bsc., MPH) 15


Types of Biostatistics
1. Descriptive (exploratory) statistics: is the aspect of
organization, presentation and summarization of data.
These include techniques for tabular and graphical
presentation of data as well as the methods used to
summarize a body of data with one or two
meaningful figures
E.g. At our health centre, 50 patients were diagnosed
with angina last year.

By Teresa K. (Bsc., MPH) 16


Descriptive statistics cont’d …

 Some statistical summaries which are especially


common in descriptive analyses are:
Measures of central tendency
Measures of dispersion
Measures of association
Cross-tabulation /contingency table
Histogram
Quantile, Q-Q plot
Scatter plot
Box plot By Teresa K. (Bsc., MPH) 17
2. Inferential Statistics:

 This branch of statistics deals with techniques of making


conclusions about the population.

 Inferential statistics builds upon descriptive statistics.

 The inferences are drawn from particular properties of


sample to particular properties of population.

By Teresa K. (Bsc., MPH) 18


Inferential Statistics cont’d...

 Inferential statistics are used to make generalizations


from a sample to a population.

 They encompasses a variety of procedures to ensure


that the inferences are sound and rational, even
though they may not always be correct.

By Teresa K. (Bsc., MPH) 19


Statistical inference cont’d…

 In short, inferential statistics enables us to make


confident decisions in the face of uncertainty.

E.g. Antibiotics reduce the duration of viral throat


infections by 1-2 days.

Five percent of women aged 30-49 consult their GP


each year with heavy menstrual bleeding.

By Teresa K. (Bsc., MPH) 20


Summery

Descriptive statistical methods


– Provide summary indices for a given data, e.g.
arithmetic mean, median, standard deviation,
coefficient of variation, etc.
Inductive (inferential) statistical methods
– Produce statistical inferences about a population
based on information from a sample derived from
the population, need to take variation into account

sample Population

– Estimating population values from sample values 21


By Teresa K. (Bsc., MPH)
Summery

• E.g.
At our health centre, 50 patients were diagnosed
with angina last year. (descriptive )

Antibiotics reduce the duration of viral throat


infections by 1-2 days. (inferential)

Five percent of women aged 30-49 consult their GP


each year with heavy menstrual bleeding.
(inferential)
By Teresa K. (Bsc., MPH) 22
• Why we need biostatistics?

By Teresa K. (Bsc., MPH) 23


Why we need biostatistics?

 Main reason: handling variations:


o Biological variation
–Among individuals as well as within same
individual over time
»Example: height, weight, blood pressure,
eye color ...
o Sample variation:
Biomedical research projects are usually carried
out on small numbers of study subjects

By Teresa K. (Bsc., MPH) 24


Why need to learn biostatistics? Cont’d....

 Essential for scientific method of investigation


– Formulate hypothesis
– Design study to objectively test hypothesis
– Collect reliable and unbiased data
– Process and evaluate data rigorously
– Interpret and draw appropriate conclusions

By Teresa K. (Bsc., MPH) 25


Why need to learn biostatistics? Cont’d....

 Essential for understanding, appraisal and critique of


scientific literature
 Public health and medicine are becoming
increasingly quantitative.

By Teresa K. (Bsc., MPH) 26


limitations of statistics:

 It deals with only those subjects of inquiry that are


capable of being quantitatively measured and
numerically expressed.

 It deals on aggregates of facts and no importance is


attached to individual items – suited only if the group
characteristics are desired to be studied.

 Statistical data are only approximation and not


mathematically correct.
By Teresa K. (Bsc., MPH) 27
variables

 Variable: A variable is a characteristic under study


that assumes different values for different elements.
or it is a characteristic or attribute that can assume
different value.
Some examples of variables include:
 Diastolic blood pressure,
 heart rate, height,
 The weight and
 Stage of bladder cancer to list some

By Teresa K. (Bsc., MPH) 28


variables cont’d…

 Random variable: are varibles whose value are


determined by chance.
 Data: the measurements or observatuions (values)
for a variable

 Data set: it is a collection of observation on a


variable.

By Teresa K. (Bsc., MPH) 29


variables cont’d…

Values Many

variables Data Data set

Mrs. brown Mr. Patel Mr. Amanda


Age 32 24 20
Sex Female Male Male
Blood type O O A

By Teresa K. (Bsc., MPH) 30


Types of variables
 Depending on the characteristic of the measurement,
variable can be:
Qualitative(Categorical) variable
A variable or characteristic which cannot be
measured in quantitative form. But, can only be
identified by name or categories, or variable that
can be placed into distinct categories, according to
some characteristic or attribute.
 For instance place of birth, ethnic group, type of
drug, stages of breast cancer (I, II, III, or IV),
degree of pain (low, moderate, sever or
unbearable). By Teresa K. (Bsc., MPH) 31
Types of variables cont’d…

• The categories should be clear cut, not overlapping,


and cover all the possibilities. For example, sex (male
or female), vital status (alive or dead), disease stage
(depends on disease), ever smoked (yes or no).

By Teresa K. (Bsc., MPH) 32


Types of variables cont’d…

Quantitative(Numerical) variable:
 Is one that can be measured and expressed numerically.
 They can be of two types
Discrete Data
The values of a discrete variable are usually whole
numbers, such as the number of episodes of
diarrhoea in the first five years of life.
Observations can only take certain numerical values
Numerical discrete data occur when the observations
are integers that correspond with a count of some
sort. By Teresa K. (Bsc., MPH) 33
Discrete Data cont’d…

 Some common examples are:


 The number of bacteria colonies on a plate,
 The number of cells within a prescribed area
upon microscopic examination,
 The number of heart beats within a specified
time interval,
 A mother’s history of numbers of births ( parity)
and pregnancies (gravidity),
 The number of episodes of illness a patient
experiences during some time period, etc.
By Teresa K. (Bsc., MPH) 34
Continuous Data

A continuous variable is a measurement on a


continuous scale
Each observation theoretically falls somewhere
along a continuum.
One is not restricted, in principle, to particular
values such as the integers of the discrete scale.
most clinical measurements, such as:
 blood pressure,
 serum cholesterol level,
 height, weight, age etc. are on a numerical
continuous scale.
By Teresa K. (Bsc., MPH) 35
Continuous Data cont’d…
Continuous data are used to report a measurement
of the individual that can take on any value within
an acceptable range.
Data

Qualitative

Quantitative

Discrete Continuous
By Teresa K. (Bsc., MPH) 36
Scales of measurement

Data comes in various sizes and shapes and it is


important to know about these so that the proper
analysis can be used on the data.
There are four at which we measure:
Nominal scales of measurement
It may be thought of as "naming" level. This level of
measurement do not put subjects in any particular
order. There is no logical basis for saying one
category is higher or less than the other category. In
research activities a YES/NO scale is nominal.

By Teresa K. (Bsc., MPH) 37


Nominal scales of cont’d…

The simplest data consist of unordered,


dichotomous, or "either - or" types of
observations, i.e., either the patient lives or the
patient dies, either he has some particular
attribute or he does not.

 Examples are: Blood group, Gender, religious


affiliation

By Teresa K. (Bsc., MPH) 38


Nominal scales cont’d…

 The nominal level of measurement classifies data


into mutually exclusive (non over lapping),
exhaustive categories in which no order or ranking
can be imposed on the data

By Teresa K. (Bsc., MPH) 39


Ordinal Scales of Measurement

An ordinal scale is next up the list in terms of power of


measurement. The simplest ordinal scale is a ranking.
At this level we put subjects in order from lowest to
height.
It is important to know that ranks do not tell us by
how much subjects differ.
There is no objective distance between any two points
on your subjective scale.
Hence, an ordinal scale only lets you interpret gross
order and not the relative positional distances.

By Teresa K. (Bsc., MPH) 40


Ordinal Scales cont’d…

 E.g. If we told that third students have better


knowledge than first year student, then we do not
know by how much they are better.
To measure the amount of the difference between
subjects we need the next level of measurement.

By Teresa K. (Bsc., MPH) 41


Ordinal Scales cont’d…

Some of the examples under this scales of


measurement includes:
• Academic status, job satisfaction index,
employment status, response to treatment
(none, slow, moderate, fast)
e.g. 1. strongly agree
2. agree
3. no opinion
4. disagree
5. strongly disagree
By Teresa K. (Bsc., MPH) 42
Ordinal Scales cont’d…

 The ordinal level of measurement classifies data


into categories that can be ranked; however, precise
differences between the ranks do not exist.

By Teresa K. (Bsc., MPH) 43


Interval Scales of Measurement

 It is more powerful than nominal and ordinal as it not


only orders or ranks or rates but also shows exact
distances in between.

 On interval measurement scales, one unit on the scale


represents the same magnitude on the trait or
characteristic being measured across the whole range
of the scale.
 They do not have a "true" zero point, however, and
therefore it is not possible to make statements about
how many times higher one score is than another.
By Teresa K. (Bsc., MPH) 44
Interval Scales cont’d …

 A good example of an interval scale is the Fahrenheit


scale for temperature.
 Equal differences on this scale represent equal
differences in temperature, but the scale is not a RATIO
Scale. Thus, a temperature of 30 degrees is not twice
as warm as one of 15 degrees
The interval level of measurement ranks data, and
precise differences between units of measure do
exist; however, there is no meaningful zero.

By Teresa K. (Bsc., MPH) 45


Ratio Scales of Measurement

 The highest level of measurement


 This has the properties of an interval scale together
with a fixed origin or zero point.
 Examples of variables which are ratio scaled include
weights, lengths and times.

By Teresa K. (Bsc., MPH) 46


Ratio Scales cont’d…

 Ratio scales permit the researcher to compare both


differences in scores and the relative magnitude of
scores.
– For instance the difference between 5 and 10
minutes is the same as that between 10 and 15
minutes, and 10 minutes is twice as long as 5
minutes.

By Teresa K. (Bsc., MPH) 47


Ratio Scales cont’d…

 The ratio level of measurement possesses all the


characteristics of interval measurement, and there
is exists a true zero. In addition, true ratio exist
between different units of measure.

By Teresa K. (Bsc., MPH) 48


Summary table for the four scales of measurement

Lowest scale Scale characterstics


Nominal Naming

Ordinal Ordering

Interval Equal interval without absolute zero

Ratio Equal interval with absolute zero

Highest

By Teresa K. (Bsc., MPH) 49


Categorize the following variables into nominal,
ordinal, interval or ratio

 Gender
Height
 Grade(A, B, C, D and F ) Weight
 Rating scale(poor, good, excelent) Time
 Eye colour Age
 Political affilation IQ
 Religious affilation Temprature
 Ranking of tennis players Salary
 Majour field
 Nationality

By Teresa K. (Bsc., MPH) 50


ASSIGNMENT 1
Exercise 1: Table 1.6 contains the characteristics of cases and controls
from a case-control study into stressful life events and breast cancer
in women (Protheroeet al.1999). Identify the type of each variable
in the table.

Exercise 2: Table 1.7 is from a cross-section study to determine the


incidence of pregnancy-related venous thromboembolic events and
their relationship to selected risk factors, such as maternal age,
parity, smoking, and so on (Lindqvistet al.1999). Identify the type of
each variable in the table.

Exercise 3: Table 1.8 is from a study to compare two lotions, Malathion


and d-phenothrin, in the treatment of head lice (Chosidowet
al.1994). In 193 schoolchil-dren, 95 children were given Malathion
and 98 d-phenothrin. Identify the type of each variable in the table.

By Teresa K. (Bsc., MPH) 51


By Teresa K. (Bsc., MPH) 52
By Teresa K. (Bsc., MPH) 53
By Teresa K. (Bsc., MPH) 54
 Response and Explanatory variables
 A variable can be also either response (dependant,
outcome) variables or explanatory (independent,
predictor) variables.
 Response (dependent, outcome) variables: are
variables which can be affected by explanatory
variable and it is the outcome of a study.
A variable you would be interested in predicting or
forecasting.
 While explanatory variables are any variables that
explain the response variable.

By Teresa K. (Bsc., MPH) 55


exercise 1:

In a study to determine whether surgery or


chemotherapy results in higher survival rates for a
certain type of cancer,

Which variable is the explanatory variable and which


one is the response variable?

By Teresa K. (Bsc., MPH) 56


• What is the importance of
variable classification?

By Teresa K. (Bsc., MPH) 57


• Source of Data?

By Teresa K. (Bsc., MPH) 58


Primary source of data

It needs the involvement of the researcher


himself. Census and sample survey are sources of
primary types of data

Experiments is also another means of getting the


data needed to answer a question

By Teresa K. (Bsc., MPH) 59


Source of Data…
secondary data.
The data needed to answer a question may already
exit in the form of published reports, commercially
available data banks, or the research literature.

In this case data were obtained from already collected


sources like newspaper, magazines, DHS, hospital
records and existing data like;
Mortality reports
Morbidity reports
Epidemic reports
Reports of laboratory utilization (including
laboratory test results)
By Teresa K. (Bsc., MPH) 60
Methods of data organization and presentation

 The data collected in a survey is called raw data. In most


cases, useful information is not immediately evident
from the mass of unsorted data.

 Collected data need to be organized in such a way as to


condense the information they contain in a way that will
show patterns of variation clearly.

By Teresa K. (Bsc., MPH) 61


1. Frequency Distributions

 Quite often, the presentation of data in a meaningful


way is done by preparing a frequency distribution. If
this is not done the raw data will not present any
meaning and any pattern in them, may not be
detected.

 Given a set of scores, constructing a frequency


distribution includes proportion(P)/ percentages.

By Teresa K. (Bsc., MPH) 62


Frequency Distributions cont’d …

 Frequency distribution determines the number of


units (e.g., people) which fall into a series of specified
categories.
 The Frequency is the count of the number of times
that a particular combination occurred in a data set.
 The relative frequency is the frequency of the
event/value/category divided by the total number of
data points.
Frequency distribution can be grouped or ungrouped

By Teresa K. (Bsc., MPH) 63


Ungrouped Frequency Distribution

 It uses to present categorical variable in simplified


and easily understandable way
 This frequency table can be constructed by listing all
possible categories of the variable and then counting
the number laying on each category of the variable
as a frequency.

By Teresa K. (Bsc., MPH) 64


Example
Example : the following data is about current age of
women and it was collected from 240 women ( data
1).

By Teresa K. (Bsc., MPH) 65


Example: Consider the data collected on age at first
marriage of 240 women. One of the variable in this
dataset is religion followed by the women. Hence, for
such types of variable, we can use ungrouped frequency
distribution to summarize the data as follows:
religion frequency Relative frequency(%)
Orthodox 103 42.9
Muslim 33 13.8
Protestant 97 40.4
Others* 7 2.9
Total 240 100
*catholic, none religious
By Teresa K. (Bsc., MPH) 66
Grouped Frequency Distribution

In order to present data using grouped frequency


distribution, it is not as simple as that of ungrouped. In
this case we need to compute some values. These
values are given below:
Number of class(K): The number of categories
the table will have
Number of class can be computed/ estimated using
Sturge’s rule as:
K = 1+3.322log(n)
Where:
K= number of class
n=sample size.
By Teresa K. (Bsc., MPH) 67
Grouped Frequency cont’d…

• Then the width of each class, W, can be computed


as:

By Teresa K. (Bsc., MPH) 68


Grouped Frequency cont’d…

Class limit: The range for each class/ The smallest


and largest values that can go into any class; they
can be either lower or upper class limits.
Lower class limit: Smallest observation of the
category
Upper class limit: Smallest observation plus
width of the class minus one.

By Teresa K. (Bsc., MPH) 69


Grouped Frequency cont’d…

 When forming classes, always make sure that each item


(measurement or observation) goes into one and only
one class, i.e. classes should be mutually exclusive
(namely, that successive classes have no values in
common).
 To this end we must make sure that the smallest and
largest values fall within the classification, that none of
the values can fall into possible gaps between successive
classes.

By Teresa K. (Bsc., MPH) 70


Grouped Frequency cont’d…

 Note that: the Sturges rule should not be regarded


as final, but should:
Be considered as a guide only. The number of
classes specified by the rule should be increased
or decreased for convenient or clear presentation.

By Teresa K. (Bsc., MPH) 71


Grouped Frequency cont’d…

 Class Boundaries/True Limits: are those limits, which


are determined mathematically to make an interval of a
continuous variable continuous in both directions, and
no gap exists between classes. It is obtained by
subtracting and adding 0.5 from lower and upper class
limit respectively
 Lower class boundary
Upper class boundary

By Teresa K. (Bsc., MPH) 72


Grouped Frequency cont’d…

 Class mark/ Mid-point (Xc) of an interval: is the value


of the interval which lies mid-way between the lower
true limit (LTL) and the upper true limit (UTL) of a
class.
It is calculated as: The average of lower and upper
class limit.

By Teresa K. (Bsc., MPH) 73


NB: The constructed grouped frequency distribution
expected to be:
– Class intervals should be continuous (for
continuous data), non overlapping(mutually
exclusive) and exhaustive.
– Class intervals should generally be of the same
width
– Open indeed class intervals should be avoided.
These are classes like less then 10, greater than
65, and so on.

By Teresa K. (Bsc., MPH) 74


Grouped Frequency cont’d…
Example for data 1
 The number of classes(k) can be computed using
Sturg's rule as:

 Therefore, the width W of each class can be


computed as:

 Thus the width of each class can be 4 and the lower


class limit for the first class will be the minimum
observation from the dataset.
By Teresa K. (Bsc., MPH) 75
Example for data 1
 Thus, the grouped frequency distribution of current age of women can be
constructed as:
Class Class boundary Class mark Frequency RF(%) CF
limit
15-18 14.5-18.5 16.5 15 6.25 15
19-22 18.5-22.5 20.5 49 20.41 64
23-26 22.5-26.5 24.5 51 21.25 115
27-30 26.5-30.5 28.5 40 16.67 155
31-34 30.5-34.5 32.5 21 8.75 176
35-38 34.5-38.5 36.5 22 9.17 198
39-42 38.5-42.5 40.5 18 7.50 216
43-46 42.5-46.5 44.5 15 6.25 231
47-50 46.5-50.5 48.5
By Teresa K. (Bsc., MPH)
9 3.75 240
76
Example for data 1 cont’d …

Where RF and CF are relative frequency and cumulative


frequency respectively.
 Note that: the value to be added or subtracted on the
class limits to get class boundaries depends on the
decimal number of the dataset that we want to
summarize.
The width of a class is found from the true class limit by
subtracting the true lower limit from the upper true limit
of any particular class.
For example, the width of the above distribution is (let's
take the fourth class) ( w = 30.5 - 26.5 = 4).
By Teresa K. (Bsc., MPH) 77
Statistical Tables

A statistical table is an orderly and systematic


presentation of data in rows and columns.
Rows : are horizontal arrangements.
Columns: are vertical arrangements.

By Teresa K. (Bsc., MPH) 78


Statistical Tables cont’d…

 Based on the purpose for which the table is designed


and the complexity of the relationship, a table could
be either of simple frequency table or cross
tabulation.
Simple frequency table is used when the
individual observations involve only to a single
variable.
Cross tabulation is used to obtain the frequency
distribution of one variable by the subset of
another variables.

By Teresa K. (Bsc., MPH) 79


Statistical tables cont’d…

Construction of tables
There are no hard and fast rules to follow, the
following general principles should be addressed in
constructing tables.
Tables should be as simple as possible.
Tables should be self-explanatory:
 Title should be clear and to the point (a good
title answers: what? when? where? how
classified ?) and it should be placed above the
table.
 Each row and column should be labeled.
By Teresa K. (Bsc., MPH) 80
Statistical tables cont’d …

 Numerical entities of zero should be explicitly


written rather than indicated by a dash. Dashed
are reserved for missing or unobserved data.

 If data are not original, their source should be


given in a footnote.

By Teresa K. (Bsc., MPH) 81


Tables cont’d…
One-variable/ Simple frequency table
– Most basic table is a simple frequency distribution with one
variable
Title Column
Example,
Fig 3. Blood group of voluntary blood donors examined in Red Cross Blood bank,
within a day, May 2006 (n=548)
Rows
Eample 2: simple table cont’d...

T a b le 5 . C lin ic a l s y m p to m s a m o n g 5 4 p a tie n ts w ith S


T y p h im u riu m -in f e c tio n , O s lo , N o rw a y , M a y 1 9 9 8

S y m p to m s C ases
n %
D ia rrh o e a 54 100
Fever 35 65
H eadache 12 22
J o in t p a in 4 7
M u s c le p a in 4 7
Two and three variable table

If two variables are cross tabulated, it is a two


variable table

If the tabulation is among three variables, it is


three variable table

In cross tabulated frequency distributions where


there are row and column totals, the decision for
the denominator is based on the variable of interest
to be compared over the subset of the other
variable.
Two and three variable table cont’d…
Table 1. Distribution of variable 1 by variable 2,
population X (n=58), place Y, period Z
Variable 2

Variable 1 Value 1 Value 2 Value 3 Total

Value 1 2 4 7 13
Value 2 3 5 3 11
Value 3 4 5 4 13
Value 4 5 6 2 13
Unkown 3 2 3 8

Total 17 22 19 58

Explanation of acronyms, units used, …


Two and three variable table Cont….

Age Male Female Total


15-24 34 76 110
25-34 48 56 104
35-44 65 54 119
45-54 56 58 114
55-64 78 53 131
65-74 46 47 93
74+ 42 51 93

369 395 764


Two and three variable table cont’d...

T a b le 1 . C a s e s o f S a lm o n e lla
T y p h im u r iu m - in f e c t io n b y a g e - g r o u p a n d s e x ,
H e rø y , N o rw a y , 1 9 9 9

A g e g ro u p S e x T o ta l
(y e a rs ) M a le F e m a le
0 - 9 7 5 1 2
1 0 - 1 9 5 5 1 0
2 0 - 2 9 5 5 1 0
3 0 - 3 9 1 4 5
4 0 - 4 9 2 3 5
5 0 - 5 9 0 3 3
6 0 - 6 9 2 1 3
7 0 - 2 4 6
T o ta l 2 4 3 0 5 4
Two and three variable table cont’d...
Distribution participants by age, sex and residency

Residence Age Male Female Total


Urban 15-24 34 76 110
25-34 48 56 104
35-44 65 54 119

Rural 15-24 56 58 114


25-34 78 53 131
35-44 46 47 93
Total 369 395 764
Common form of a two by two variable

It is a special form of table favorite among


epidemiologist
It is used to compare whether there is relationship
between the two variables

Number of Total
Cases Controls
Exposure

Exposed 23 23 46

Non exposed 4 139 143

Total 27 162 189


Composite/ Higher Order Table

It is a large table combining several separate


variable/tables

Age, sex and other demographic variables may be


combined to form a single table
– Example of composite table
Characteristics Number Percent
Marital status
Single 50 67.6
Married 20 27.0
Divorced/ widowed 4 5.4
Current Residence (n=73)
Within the PA 40 54.8
Within the PA (H. Post) 25 34.2
Within the nearest town 8 11.0
Residence of origin
Within the PA 4 5.4
Outside the PA 24 32.4
Outside the Woreda 46 62.2
Training TVETI
Axum 19 25.7
Makele 55 74.3
Totals 74 100
Graphical Presentation

Graphs are often easier to interpret than tables,


perhaps at the expense of detail.
A variety of graphs are used depending on the type
of data.
If we want to present categorical/qualitative or
quantitative discrete data/variable using graph, then
pie chart and bar chart are the appropriate ones,
however if the variable is numerical/quantitative
continuous data in nature, then we can use histogram,
frequency polygon and box plot.

By Teresa K. (Bsc., MPH) 92


Graphical Presentation cont’d…

There are, however, general rules that are commonly


accepted about construction of graphs.
Every graph should be self-explanatory and as
simple as possible.
Titles are usually placed below the graph and it
should answer again question like: what ? Where?
When? How classified?
Legends or keys should be used to differentiate
variables if more than one is shown.

By Teresa K. (Bsc., MPH) 93


Graphical Presentation cont’d…

The units in to which the scale is divided should


be clearly indicated.
The numerical scale representing frequency must
start at zero or a break in the line should be
shown.

By Teresa K. (Bsc., MPH) 94


Examples of graphs:

Bar Chart
Bar diagrams are used to represent and compare the
frequency distribution of discrete variables and
attributes or categorical series. When we represent
data using bar diagram, all the bars must have equal
width and the distance between bars must be equal.
Each category of variable is represented by a bar

Variables are qualitative, or treated as qualitative

It can be displayed as horizontal or vertical

By Teresa K. (Bsc., MPH) 95


Types of bar charts

There are different types of bar diagrams:


A. Simple bar chart: It is a one-dimensional
diagram in which the bar represents the
whole of the magnitude.
The height or length of each bar indicates
the size (frequency) of the figure
represented.

– one variable

– It can be displayed as horizontal or vertical


Types of bar charts cont’d…

Figure 1: immunization status of children in adami Tulu Wereda,


1995
Type of bar chart cont’d ...

B. Grouped bar chart


– Data from 2-variable or 3 variable tables
– Distinct colours or shading is used to
differentiate
– Legend is necessary
E.g. Grouped/ joined bar chart

Cell separated
One cell By a space

The meaning of
each bar is shown
in a legend

Figure 2: TT immunization status by marital status of women 15-49 years,


Asendabo town, 1996.
Type of bar chart cont’d ...

C. Stacked bar chart


– It is used to show the same data as a grouped bar
chart using a single bar

– Different groups are differentiated by different


segments within a single bar

– You are able to see the overall change easier, but


changes between groups may be difficult than
grouped bars
FEg Stacked
ig ure 1 . C as bar
e s o chart
f S T yp him urium -inf e c tio n
b y ag e -g ro up and s e x, H e rø y, N o rw ay, 1 9 9 9
A g e -g ro up
70 -
M a le
60 - 69
F e m a le
50 - 59

40 - 49

30 - 39

20 - 29

10 - 19

0 - 9

0 2 4 6 8 10 12 14
N um b e r o f c a s e s

Figure 1: cases of S. Typhimurium-infection by age group and


sex, Heroy, Norway, 1999
Type of bar chart cont’d ...

D. 100% component bar chart


– It is a variant of stacked bar chart , where bars are
pulled to 100% rather than their real values;

– It is helpful for comparing the contribution of


different subgroups within the categories of the
main variable
Eg 100% Component bar chart

Proportional distribution by sex Male Female


100 %

80 %

60 %

40 %

20 %

0%
0-9 10 - 19 20 - 29 30 - 39 40 - 49 50 - 59 60 - 69 70 -
Age-group
Figure 1. Cases of S Typhimurium-infection by age-group and sex, Herøy, Norway,
1999
Pie -Charts;

 It is a circle divided into sectors so that the areas of the


sectors are proportional to the frequencies.
 It is split into segments to show percentages or the
relative contributions of categories of data.
 It is a good method of representation if you wish to
compare a part of group with the whole group.
 The number of categories should not be too much.

By Teresa K. (Bsc., MPH) 104


e.g. Pie chart

Fig.1. Distribution of religion of participants from Kunama ethnic group


among Eritrean Refugees in Shimelba Camp, July 2006
Quantitative continuous data

Histograms: is the graph of the frequency distribution


of continuous measurement variables.
It is constructed on the basis of the following
principles:
The horizontal axis is a continuous scale running
from one extreme end of the distribution to the
other. It should be labeled with the name of the
variable and the units of measurement.

By Teresa K. (Bsc., MPH) 106


Histograms cont’d …

For each class in the distribution a vertical rectangle


is drawn with
Its base on the horizontal axis extending from
one class boundary of the class to the other class
boundary, there will never be any gap between
the histogram rectangles.
The bases of all rectangles will be determined by
the width of the class intervals.

By Teresa K. (Bsc., MPH) 107


Histograms cont’d …

Area of each column is proportional to the number of


observations in that interval
In constructing
– Use equal class intervals
– Do not use scale breaks

It could show second variable by shading


Figure 1: Age distribution of women in a reproductive age group
included in a study of violence against women in Butajira, 1984.
Frequency polygon

 If we join the midpoints of the tops of the adjacent


rectangles of the histogram with line segments a
frequency polygon is obtained.

Note: it is not essential to draw histogram in order


to obtain frequency polygon. It can be drawn with
out erecting rectangles of histogram as follows:

By Teresa K. (Bsc., MPH) 110


Frequency polygon cont’d…

 The scale should be marked in the numerical values of


the midpoints of intervals

 Erect ordinates on the midpoints of the interval - the


length or altitude of an ordinate representing the
frequency of the class on whose mid-point it is erected.

 Join the tops of the ordinates and extend the connecting


lines to the scale of sizes.

By Teresa K. (Bsc., MPH) 111


Construction of a frequency polygon from a histogram
15 cases
14

13 1 case patient
12 1 case staff member
11

10

0
00- 06- 12- 18- 00- 06- 12- 18- 00- 06- 12- 18- 00- 06- 12- 18- 00-

27 August 28 August 29 August 30 August

Date and time of onset


Mid point/ class mark
By Teresa K. (Bsc., MPH) 113
Ogive or cumulative frequency curve:

 To construct an Ogive curve:


Compute the cumulative frequency of the
distribution.
 Prepare a graph with the cumulative frequency on
the vertical axis and the true upper class limits (class
boundaries) of the interval scaled along the X-axis
(horizontal axis).
The true lower limit of the lowest class interval with
lowest scores is included in the X-axis scale; this is
also the true upper limit of the next lower interval
having a cumulative frequency of 0.
By Teresa K. (Bsc., MPH) 114
By Teresa K. (Bsc., MPH) 115
Summarizing Data

 The first step in looking at data is to describe the


data at hand in some concise way.
 One type of measure useful for summarizing data
defines the center, or middle, of the sample.

By Teresa K. (Bsc., MPH) 116


Measures of Central Tendency/ Measures of Location

 Measures of central Tendency: the various methods of


determining the actual value at which the data tend to
concentrate. Hence, measures of central Tendency is a
value which tends to sum up or describe the mass of the
data.
 These central tendency includes:
Mean ,
Median and
Mode .

By Teresa K. (Bsc., MPH) 117


Arithmetic Mean/simple Mean ( X )

Definition: the arithmetic mean is the sum of all


observations divided by the number of observations. it
is usually denoted by

 Let us consider X1, X2, ..., XN are the list of N


measurements obtained from N subjects. Then the
mean for ungrouped number of measurements for N
subjects is defined as:

By Teresa K. (Bsc., MPH) 118


The mean for Grouped data can be computed as
follows:

 Where: k=the number of classes


Xci=class mark for the ith class and
fi=frequency of the ith class

By Teresa K. (Bsc., MPH) 119


properties of Mean

Individual extreme values (also known as 'outliers')


can distort its ability to represent the typical value of
a variable (which is The main weakness of the
mean.)
It is unique for the given set of data
The value of the arithmetic mean is determined by
every item in the series.
The sum of the deviations about it is zero.

By Teresa K. (Bsc., MPH) 120


Example 1

 Consider the data on birth weight of 10 new born


children in kilo gram at university of Gondar hospital:
2.51, 3.01, 3.25, 2.02,1.98, 2.33, 2.33, 2.98, 2.88,
2.43.
Then the average birth weight can be computed as:

By Teresa K. (Bsc., MPH) 121


Example 2

 Compute mean for the grouped frequency


distribution given bellow:
The grouped frequency distribution for current
age of women

By Teresa K. (Bsc., MPH) 122


Class Class boundary Class mark Frequency RF(%) CF
limit

15-18 14.5-18.5 16.5 15 6.25 15


19-22 18.5-22.5 20.5 49 20.41 64
23-26 22.5-26.5 24.5 51 21.25 115
27-30 26.5-30.5 28.5 40 16.67 155
31-34 30.5-34.5 32.5 21 8.75 176
35-38 34.5-38.5 36.5 22 9.17 198
39-42 38.5-42.5 40.5 18 7.50 216
43-46 42.5-46.5 44.5 15 6.25 231
47-50 46.5-50.5 48.5 9 3.75 240
By Teresa K. (Bsc., MPH) 123
Example 2 cont’d…

 Where as: fi = frequency distribution of ith class


Xc = is the mid-point
n = total sample size

By Teresa K. (Bsc., MPH) 124


Median

 An alternative measure of central location, perhaps


second in popularity to the arithmetic mean.
 Suppose there are n observations in a sample. If these
observations are ordered from smallest to largest,
then the median is defined as follows:
The median, is a value such that at least half of the
observations are less than or equal to median and
at least half of the observations are greater than or
equal to median .
The median is the midpoint of the data array.

By Teresa K. (Bsc., MPH) 125


Median cont’d …

 To find the median of a data set:


Arrange the data in ascending order.
Find the middle observation of this ordered
data.

By Teresa K. (Bsc., MPH) 126


Median cont’d…

 If the number of data is ODD, then the median is the


middle data point:

Median =

 If the number of data is EVEN, then the median is the


average of the two values around the middle.

Median =

By Teresa K. (Bsc., MPH) 127


Median cont’d…

• Extreme values do NOT affect the median, making


the median a good alternative to the mean to
measure central tendencies when such values occur.

By Teresa K. (Bsc., MPH) 128


Example:

 Consider the data on the weight of 10 new born


children at university of Gondar hospital within a
month:
2.51, 3.01, 3.25, 2.02,1.98, 2.33, 2.33, 2.98, 2.88, 2.43.

– Find median for the data.

By Teresa K. (Bsc., MPH) 129


Example cont’d…

 First arrange the data in to ascending order as:


1.98, 2.02, 2.33, 2.33, 2.43,2.51, 2.88, 2.98, 3.01,
3.25.
 As 10 is even we need to take the middle two
observations and the median will be the average of
this two middle observations.

By Teresa K. (Bsc., MPH) 130


Median cont’d…

Median for grouped data:


 The median for grouped data is defined by:

Where as:
LCB= lower class boundary of the median class
Fc= cumulative frequency just before the median
class
fc=frequency of the median class
W =class width and n=number of observations. 131
By Teresa K. (Bsc., MPH)
Example median for grouped data 1

 Consider the example on age of women we presented


using frequency distribution bellow. Compute median
for grouped data?
 To compute median for grouped data, we need first
find the median class. In this example half of the
observation is 120.
 Let us see the distribution with the cumulative
frequency:

By Teresa K. (Bsc., MPH) 132


Example median for grouped data 1 cont’d…

By Teresa K. (Bsc., MPH) 133


 As we can see from the distribution, the class which
contains 120 observation for the first time is the class
with cumulative frequency 155 as 120 is under 155. So,
the median class is the 4th class

By Teresa K. (Bsc., MPH) 134


Mode

 Mode is the value appearing most frequently


 It can be obtained by counting the number of appearance for
each observation from the list.
 Important for summarising nominal/categorical types of data
 disadvantage,
 In small number of observations, there may be no mode.
 In addition, sometimes, there may be more than one mode
such as when dealing with a bimodal (two-peaks)
distribution.
 Example
a. 22, 66, 69, 70, 73. (no modal value)
b. 1.8, 3.0, 3.3, 2.8, 2.9, 3.6, 3.0, 1.9, 3.2, 3.5 (modal
value = 3.0 kg) By Teresa K. (Bsc., MPH) 135
By Teresa K. (Bsc., MPH) 136
Mode cont’d…

NB: The mode for grouped data is modal class. Modal


class is the class with the largest frequency.

By Teresa K. (Bsc., MPH) 137


Skewness:
 If extremely low or extremely high observations are present
in a distribution, then the mean tends to shift towards those
scores.
 Based on the type of skewness, distributions can be:
 Symmetrical distribution: when data values are
evenly distributed on both sides of the three
measures of central tendency (Mean, Median and
Mode).
 It is neither positively nor negatively skewed. A curve
is symmetrical if one half of the curve is the mirror
image of the other half.
 If the distribution is symmetric and has only one
mode, all three measures are the same, an example
being the normal Bydistribution.
Teresa K. (Bsc., MPH) 138
By Teresa K. (Bsc., MPH) 139
Positively skewed distribution: Occurs when the
majority of scores are at the left end of the curve
and a few extreme large scores are scattered at
the right end.
 For positively skewed distributions (where the
upper, or left tail of the distribution is longer
(“fatter”) than the lower, or right tail) the
measures are ordered as follows:
Mode < median < mean.

By Teresa K. (Bsc., MPH) 140


By Teresa K. (Bsc., MPH) 141
Negatively skewed distribution: occurs when
majority of scores are at the right end of the curve
and a few small scores are scattered at the left
end.
For negatively skewed distributions (where the
right tail of the distribution is longer than the
left tail), the reverse ordering occurs:
Mean < median < mode.

By Teresa K. (Bsc., MPH) 142


By Teresa K. (Bsc., MPH) 143
Measures of Dispersion/ Variation

 Measures of dispersion or variability will give us


information about the spread of the scores in our
distribution.
 Without knowing something about how the data is
dispersed, measures of central tendency may be
misleading.
 Most common measures of dispersion includes
Range,
Inter-quartile range,
Variance,
Standard deviation and
Coefficient of variation.
By Teresa K. (Bsc., MPH) 144
Measures of Dispersion/ Variation

 Consider the following three datasets


Dataset 1:7, 7, 7, 7, 7, 7 Mean=7, s.d=0
Dataset 2: 6, 7, 7, 7, 7, 8, mean=7, s.d=0.63
Dataset 3: 3, 2, 7, 8, 9, 13, mean=7, s.d=4.04

By Teresa K. (Bsc., MPH) 145


Measures of Dispersion cont’d…

 RANGE: It is the difference between the largest and


smallest observation from the data

EXAMPLE: Consider the data on the weight (in Kg) of


10 new born children at university of Gondar hospital
within a month:
2.51, 3.01, 3.25, 2.02,1.98, 2.33, 2.33, 2.98, 2.88,
2.43.

By Teresa K. (Bsc., MPH) 146


Measures of Dispersion cont’d…

Then the range for the dataset can be computed by


first arranging all observation in to ascending order
as:
1.98, 2.02, 2.33, 2.33, 2.43, 2.51, 2.88, 2.98, 3.01, 3.25.
Range = Maximum-Minimum=3.25-1.98=1.27
 It is based upon two extreme cases in the entire
distribution, the range may be considerably changed if
either of the extreme cases happens to drop out, while
the removal of any other case would not affect it at all.
 It wastes information , it takes no account of the entire
data. By Teresa K. (Bsc., MPH) 147
Measures of Dispersion cont’d…

 The extremes values may be unreliable; that is, they


are the most likely to be faulty

 Not suitable with regard to the mathematical


treatment required in driving the techniques of
statistical inference.

By Teresa K. (Bsc., MPH) 148


Quantiles

The Pth percentile is the value Vp such that P percent


of the sample points are less than or equal to Vp.
The median, being the 50th percentile, is a special case
of a quantile.
As was the case for the median, a different definition is
needed for the pth percentile, depending on whether
np/100 is an integer or not.

By Teresa K. (Bsc., MPH) 149


Quantiles cont’d …

The pth percentile is defined by:


1. The (k+1)th largest sample point if np/100 is not
an integer (where k is the largest integer less
than np/100)
2. The average of the (np/100)th and (np/100 + 1)th
larges observation if np/100 is an integer.

By Teresa K. (Bsc., MPH) 150


Quintiles cont’d …

Example 1: Compute the 10th and 90th percentile for


the birth weight data below.
Suppose the sample consists of birth weights (in
grams) of all live born infants born at a private
hospital in a city, during a 1-week period. This sample
is shown in the following table:

3265, 3323, 2581, 2759, 3260, 3649,2841


3248, 3245, 3200, 3609, 3314, 3484, 3031
2838, 3101, 4146, 2069, 3541, 2834
By Teresa K. (Bsc., MPH) 151
Quintiles cont’d …

By sorting the data from the smallest to highest

2069 2581 2759 2834 2838 2841 3031 3101 3200


3245 3248 3260 3265 3314 3323 3484 3541 3609
3649 4146

Solution: Since 20×0.1=2 and 20×0.9=18 are integers,


the 10th and 90th percentiles are defined by:

By Teresa K. (Bsc., MPH) 152


Quintiles cont’d …

10th percentile = the average of the 2nd and 3rd


largest values = (2581+2759)/2 = 2670 g
90th percentile=the average of the18th and 19th
largest values = (3609+3649)/2 = 3629 grams.

We would estimate that 80 percent of birth weights


would fall between 2670 g and 3629 g, which gives us
an overall feel for the spread of the distribution.

By Teresa K. (Bsc., MPH) 153


Quintiles cont’d …

 Quartiles: is Other quantiles which divide the


distribution into four equal parts. The second
quartile is the median.
 The interquartile range (IQR): is the difference
between the first and the third quartiles.
 To compute it, we first sort the data, in ascending
order, then find the data values corresponding to the
first quarter of the numbers (first quartile), and then
the third quartile.

By Teresa K. (Bsc., MPH) 154


Quintiles cont’d …

example 2:
Given the following data set (age of patients) find the
interquartile range!
18,59,24,42,21,23,24,32
1. sort the data from lowest to highest
18 21 23 24 24 32 42 59

2. Find the bottom and the top quarters of the data


3. Find the difference (interquartile range) between
the two quartiles.
By Teresa K. (Bsc., MPH) 155
Quintiles cont’d …

 1st quartile = The {(n+1)/4}th observation = (2.25) th


observation = 21 + (23-21)x 0.25 = 21.5
 3rd quartile = {3/4 (n+1)}th observation = (6.75)th
observation = 32 + (42-32)x 0.75 = 39.5
Hence, IQR = 39.5 - 21.5 = 18

 The interquartile range is a preferable measure to the


range. Because it is less prone to distortion by a single
large or small value. That is, outliers in the data do not
affect the inerquartile range. Also, it can be computed
when the distribution has open-end classes.
By Teresa K. (Bsc., MPH) 156
outlier

An outlier is an observation that lies an abnormal


distance from other values in a random sample from a
population.
Before abnormal observations can be singled out, it is
necessary to characterize normal observations.
Two activities are essential for characterizing a set of
data:
Examination of the overall shape of the graphed data
for important features, including symmetry and
departures from assumptions.
Examination of the data for unusual observations that
are far from the mass ofK. (Bsc.,
By Teresa data. MPH) 157
Standard Deviation and Variance

 Variance:
While the inter-quartile range eliminates the
problem of outliers it creates another problem in
that you are eliminating half of your data.
The solution to both problems is to measure
variability from the center of the distribution.
Variance measure how far on average scores
deviate or differ from the mean.
Variance is the average of the square of the
distance each value is from
By Teresa K. (Bsc.,the
MPH) mean 158
Mathematically the formula for population
variance is defined as:
N

 ( xi   ) 2
 2
 i 1
N
The mathemetical formula for sample variance is
defined as:
n

2
 (x i  x) 2

s  i 1
n 1
By Teresa K. (Bsc., MPH) 159
Variance cont’d…

Short cut formula for the sample variance

2
( x ) 2
x 
 2
= n
n 1
By Teresa K. (Bsc., MPH) 160
Standard Deviation

 The sample and population standard deviations are


denoted by S and σ (by convention) respectively.
 The standard deviation, S.D., is just the positive square
root of the variance.
 It expresses exactly the same information as the variance,
but re-scaled to be in the same units as the mean.
 Mathematically: Population standard deviation

 ( xi   )2
  i 1
N
By Teresa K. (Bsc., MPH) 161
Standard Deviation cont’d…

 Sample standard deviation can be defined as:


n 

 ( xi  x )2
s  i1
n  1

 Example1 The Areas of spray able surfaces with DDT


from a sample of 15 houses are measured as follows (in
m2) :
101,105,110,114,115,124,125,125,130,133,135,136,13
7,140,145

By Teresa K. (Bsc., MPH) 162


Example 1 cont’d …

 Find the variance and standard deviation of the


above distribution.
 Solutions
The mean of the sample is 125 m2.
Variance (sample) = s2 = Σ(xi –x)2/n-1 = {(101-125) 2
+(105-125) 2 + ….(145-125) 2 } / (15-1)
= 2502/14
= 178.71 m4
Hence, the standard deviation
= 178.71
= 13.37 m2
By Teresa K. (Bsc., MPH) 163
Variance for grouped frequency distribution
 In a grouped frequency distribution, the variance is
computed as:

k k
n( fi x2ci )  ( fi xci )2
2
s  i 1 i 1
n(n 1)
 Where as
fi =frequency of ith class
Xci =class mark of ith class
n = total number of the sample
By Teresa K. (Bsc., MPH) 164
Example of Variance for grouped frequency
distribution
 Consider the following data of time spend by college
students for leisure activities. Compute standard
deviation.

By Teresa K. (Bsc., MPH) 165


k k
n( fi x 2
ci )  (  f i x ci ) 2

s= i 1 i 1
n ( n  1)

By Teresa K. (Bsc., MPH) 166


Coefficient of variance

 The standard deviation is an absolute measure of deviation of


observations around their mean and is expressed with the same
unit of the data.
 Due to this nature of the standard deviation it is not directly
used for comparison purposes with respect to variability.
 Coefficient of variation, is often used for this purpose
 The coefficient of variation (CV) is defined by:

CV =

 The coefficient of variation is most useful in comparing the


variability of several different samples, each with different
means. By Teresa K. (Bsc., MPH) 167
Coefficient of variance cont’d…

 CV is a relative measure free from unit of measurement.


example
Weights of newborn Weights of newborn
elephants (kg) mice (kg)
929 853 0.72 0.42
878 939 0.63 0.31
895 972 0.59 0.38
937 841 0.79 0.96
801 826 1.06 0.89
Mice show
n=10, X = 0.68 greater birth-
n=10, X = 887.1
s = 0.255 weight variation
s = 56.50
CV = 0.375
CV = 0.0637
By Teresa K. (Bsc., MPH) 168
When to use coefficient of variance
 When comparison groups have very different means
(CV is suitable as it expresses the standard deviation
relative to its corresponding mean)

 When different units of measurements are involved,


e.g. group 1 unit is mm, and group 2 unit is gm (CV is
suitable for comparison as it is unit-free)

 In such cases, standard deviation should not be used


for comparison

By Teresa K. (Bsc., MPH) 169


By Teresa K. (Bsc., MPH) 170

You might also like