SANGA
CHE 233 HEALTH STATISTICS
INTRODUCTION
Population changes reflect the natural facts of life: births and deaths. Births, in
turn, have long been largely governed by the mechanisms of family formation.
Vital statistics are compilations of data on marriage, divorce, birth, and death.
Births and deaths directly determine changes in the size of a population; marriages
and divorces create and dissolve, respectively, the conditions under which most
births occur. The surplus of births over deaths is called natural increase; under
unfavorable demographic conditions, deaths may exceed births, in which case a
natural decrease occurs.
Definition of Terms
Statistics is a branch of science that deals with collecting, organising, analyzing
and interpreting numerical data for understanding a phenomenon or making wise
decisions.
Vital statistics are quantitative information about a population's "vital events"
such as the number of births (natality), deaths (mortality), marriages (nuptiality)
and divorces.
1
SANGA
Health statistics is a branch of statistics which deals with collecting, organizing,
analyzing and interpreting numerical data for health and disease in human
population.
Population: is the entire collection, or set, of individuals or objects whose
properties are to be analyzed.
Biostatistics - When the different statistical methods are applied in biological,
medical and public health data they constitute the discipline of biostatistics.
Sources of Health Statistics
i. Surveys
ii. Administrative and medical records
iii. Health care claims data
iv. Vital records
v. Disease registries
vi. literatures
The Importance of Statistics in Healthcare
Describe the level of community health
Diagnose community illnesses
2
SANGA
Discover solutions to health problems and find clues for administrative
action
Determine priorities for health programmes
Promote health legislation
Determine the met and unmet health needs
Disseminate information on the health situation and health programmes
Determine success or failure of specific health programmes
Demand public support for health work
DATA AND DATA COLLECTION METHODS
These are raw facts or unprocessed information. There are basically two types of
data:
i. Numeric or quantitative information and
ii. Non-numeric or qualitative information.
There are two classifications of numeric data:
i. Discrete numeric data: These are countable whole numbers or integers,
e.g. the number of Patients, nurses or doctors in a hospital
ii. Continuous numeric data: These are numbers within an interval or
uncountable range of numbers, e.g. a measure of a quantity will usually
be continuous, i.e. weight, mass, litres etc.
3
SANGA
Non-numeric data are values that cannot be quantified. For example, matriculation
number, age group, blood group, DNA, RNA, sex, tribe, country, etc. Data in this
form are either categorical or ordinal. Examples of ordinal non-numeric data are
students‟ height, age group, while DNA, RNA, sex, country, tribe are examples of
categorical non-numeric data.
Methods of Data Collection
Data collected for investigation can either be primary data or secondary data. Data
collected directly from the source or respondents are known as primary data. These
are data, which are collected by the investigator for the purpose of a specific
inquiry or study. Such data are original in character and are mostly generated by
surveys conducted by individuals or research institutions. When an investigator
uses data, which have already been collected by others or those from established
data bank are known as secondary data. Such data are primary data for the agency
that collected them, and become secondary for someone else who uses these data
for his own purposes. Secondary data are less expensive to collect both in money
and time.
How to Collect Primary Data
i. Observation: is a technique that involves systematically selecting,
watching and recoding behaviors of people or other phenomena and
4
SANGA
aspects of the setting in which they occur, for the purpose of getting
(gaining) specified information.
ii. Interviews: A good interviewer can stimulate and maintain the
respondent’s interest, and can create a rapport (understanding, concord)
and atmosphere conducive to the answering of questions.
iii. Experiments
iv. Self-administered Questionnaire
Sources of Secondary Data include:
i. Official publications of Central Statistical Authority
ii. Publication of Ministry of Health and Other Ministries
iii. Internet, Websites, News Papers and Journals
iv. International Publications like Publications by WHO, World Bank,
UNICEF
Data analysis and Presentation
The presentation of data in a meaningful way is done by preparing a frequency
distribution. A table listing all classes and their frequencies is called frequency
distribution.
Let us demonstrate the concept of a frequency distribution by using the following
set.
5
SANGA
1 5 3 4 1 3 2 5
2 4 1 3 2 0 1 2
1 2 0 2 1 4 5 3
Let x represent these data values, we can use a frequency distribution to represent
this set of data by listing the x values with their frequencies in Table.
Frequency distribution
X 0 1 2 3 4 5
F 2 6 6 4 3 3
GRAPHICAL METHODS
Graphical (pictorial) representation of the data reveals patterns of behaviour of the
variable being studied. There are several graphic (pictorial) ways to describe data.
Data can be presented graphically in many ways as, line graph, dot plot display,
bar chart, pie chart, histogram, cumulative frequency curve (Ogive) and stem-and-
leaf display.
Dot Plot Display
Dot plots display the data of a sample by representing each piece of data with a dot
positioned along a scale. This scale can be either horizontal or vertical. The
6
SANGA
frequency of the values is represented along the other scale. They are usually used
to represent the frequency distribution of a discrete variable. The dot plot display is
a convenient technique to use as you first begin to analyze the data. It results in a
picture of the data as well as sorts the data into numerical order.
Example
A random sample of 20 children took their weights in kilogram in a hospital are
presented below:
23 22 26 28 22 29 30 25 26 27
21 23 27 26 25 29 30 26 25 28
Construct a dot plot of these data.
Bar Chart
Bar chart/diagrams are used to represent and compare the frequency distribution of
discrete variables and attributes or categorical series. When we represent data
using bar diagram, all the bars must have equal width and the distance between
bars must be equal.
There are different types of bar diagrams, the most important ones are:
7
SANGA
Simple bar chart
It is a one-dimensional diagram in which the bar represents the whole of the
magnitude. The height or length of each bar indicates the size (frequency) of the
figure represented.
Series 1
4.5
3.5
2.5
1.5
0.5
0
21 22 23 25 26 27 28 29 30
Multiple bar chart
In this type of chart, the component figures are shown as separate bars adjoining
each other. The height of each bar represents the actual value of the component
figure. It depicts distributional pattern of more than one variable.
8
SANGA
Example
The following table shows the intake through JAMB by the Faculty of Science of a
certain University in three consecutive years.
Department 2002 2003 2004
Botany 43 40 35
Chemistry 28 35 42
Zoology 45 40 35
Computer Science 33 25 28
Physics 40 35 38
Mathematics 35 42 45
Biology 37 40 42
Total 261 257 265
Draw (i) multiple bar chart department by department for the three years.
50
45
40
35
30
2002
25
2003
20 2004
15
10
0
Botany chemistry zoology computer science physics mathematics biology
9
SANGA
Pie-chart
The pie chart (circle graph) is used to display relative frequencies rather than actual
frequencies for the data (qualitative or quantitative discrete data). We draw a circle
and then divide it into a series of wedges or slices to represent each class in the
relative frequency distribution. The size of each slice is proportional to the
percentage of the data that fall into the corresponding class.
Example
The population of five towns in a country X in 1986 is as follows:
Town Population Sectorial Angle
A 50,000 90
B 100,000 180
C 25,000 45
D 12500 22.5
E 12500 22.5
SOLUTION
POPULATION
A
B
C
D
E
10
SANGA
VITAL RATIO
A ratio is a general term given to any numerator – denominator relationship
between two numbers. Example 5/6 59/10 23/60 etc
General fertility rate
This is the number of live births per 1000 women aged 15-49 in a given year.
i.e. GFR = live births x 1000
Pf15-49
Age – specific Fertility Rate
This is the number of births per annum per 1000 women of specific age group.
Example
Age popn of females live births ASFR
(births/Pfax100)
15 – 19 1611090 463631 288
20 – 24 1558276 427298 274
25 – 29 1425242 412878 290
30 – 34 1381174 380778 276
35 – 39 1632695 308671 189
40 – 44 1581373 281161 178
45 – 49 1400555 239701 171
10590405 2514118 1666
GFR = 2514118 x 1000 = 237 per 1000 women
10590405
11
SANGA
Total fertility rate
This is the average number of children that would be born to a woman by the time
she ended child bearing if she were to pass through all her child bearing years
conforming to the age-specific fertility rates of a given year.
i.e. TFR = 5x∑ASFR / 1000
= 5(1666/1000) = 5x1.66 = 8 per woman
Mortality
Mortality is a relationship of death cases to the whole population. Two basic types
of mortality include;
i. General (crude) mortality rate or death rate
ii. Specific mortality rate or death rate
Crude death rate (or mortality rate) is the number of death cases in a year per
1000 of the population.
Specific Mortality Rate or Death Rate
Age and sex related mortality rates can be computed for both genders and age
groups.
Sex related mortality rate is the number of death cases of the cohort by 1000
mid-year population of the cohort.
12
SANGA
FREQUENCY OF DEATHS BY AGE
Perinatal mortality rate is the number of infant deaths in the first 24hours by
1000 live births in the given year.
Postnatal mortality rate is the number of infant deaths in the first 0-6 days by
1000 live births in the given year.
Neonatal mortality rate is the number of neonatal death (0-28 days) by 1000 live
births in the given year.
Infant mortality rate is the number of infant death 0 – 365 days by 1000 live
births.
CAUSE SPECIFIC DEATH RATES (CSDR)
The CSDR is the number of deaths attributed to a particular cause (c) divided by
population at risk (p), usually expressed in deaths per 100,000
i.e. CSDE = Dc x 100,000
P
Example:
Community popn at risk TB Malaria Diabetes
A 304848 278 252 188
B 211075 170 170 154
C 416418 193 180 228
13
SANGA
Maternal Mortality Rate
Maternal mortality is the death of a woman while pregnant or within 42 days of
termination of pregnancy, irrespective of the duration and site of the pregnancy
from any cause related or aggravated by the pregnancy or its management but not
from accidental or incidental causes.
Example:
Region live births Abortion Lack of ANC Lack of PNC labour
A 56403 67 32 10 193
B 97166 81 14 6 104
C 93415 72 18 10 107
Life expectancy rate is the estimate of the average number of additional years that
a person of a given age can expect to live. The most common measure of life
expectancy is life expectancy at birth.
SAMPLE
Sample can be defined as a finite part of a statistical population whose properties
are studied to gain information about the whole population.
14
SANGA
SAMPLING TECHNIQUE: Sampling technique is the act, process, or technique
of selecting a representative part of a population for the purpose of determining
parameters or characteristics of the whole population.
Types of Sampling Technique
The types of sampling are referred to as sampling procedures or methods of
sampling. They are also grouped as:
Scientific or probability sampling
Non-scientific or non-probability sampling
Scientific or non-probability sampling method: This is a sampling method where
every member of the population has a greater chance of being selected as a sample.
The different methods in scientific or probability sampling include:
i. Simple random sampling – every member or element of the population
has equal chance of being selected. The steps involved identifying each
member of the population, for example listing with name or number and
then selecting the sample through any of table of random numbers, ballot
method, or throw of dice.
ii. Systematic random sampling –the researcher starts randomly, that is at
any point, having listed the population, and selects the ‘nth’ member until
15
SANGA
the sample size is reached. A way of getting the interval is often by
dividing the total population by the sample size required.
iii. Stratified random sampling – the researcher starts by subdividing the
population into homogenous subgroups or strata (stratum for singular),
by some known unifying qualities for those groups, then draws a sample
using simple random sampling or systematic sampling methods.
iv. Cluster or area sampling -with the cluster sampling, instead of drawing
individual members of the population, groups or clusters are selected
from the population. The population is divided into clusters along
geographical boundaries, and clusters randomly selected. All the
members of selected clusters are studied.
Non-scientific/Non-probability Sampling
Non probability or non-scientific sampling method is the method of sampling
where not every member of the population is given a chance to be part of the
sample. Choice usually depends on the judgment of the researcher. The major
types of non-scientific or non-probability sampling are:
1. Convenience (accidental) sampling –the researcher primarily works with a
group that is accessible and convenient; hence this method is prone to bias.
16
SANGA
There is no evidence that the sample is representative of the population of
interest.
2. Purposive (judgmental) sampling – where the researcher uses his/her
discretion to determine the members of the population that may suit the
purposes and requirements of the study.
3. Quota (proportional) sampling is the nonprobability equivalent of stratified
sampling. Like stratified sampling, the researcher first identifies the stratums
and their proportions as they are represented in the population. Then
convenience or judgment sampling is used to select the required number of
subjects from each stratum. This differs from stratified sampling, where the
stratums are filled by random sampling.
VARIABLE
A variable is any entity that can take on different values. Variables are of interest
in research because they are the main reasons for the research.
Examples of variables include age, sex, weight, height, educational
attainment/qualification, experience, weather, socio-economic status, temperature,
state of health etc.
17
SANGA
Types of Variables
Independent variable – the independent variable is that variable in the research,
which the researcher manipulates, and could be equated to a cause or stimulus or
treatment in research that has to establish cause-effect relationship.
Dependent variable - is that variable not manipulated by the researcher, but which
the researcher expects will change once the independent variable is introduced. It
can be equated to the effect, response or the result.
MEASURES OF CENTRAL TENDENCY
A measure of central tendency is a single value that attempts to describe a set of
data by identifying the central position within that set of data. As such, measures of
central tendency are sometimes called measures of central location. They are also
classed as summary statistics. The mean (often called the average) is most likely
the measure of central tendency that you are most familiar with, but there are
others, such as the median and the mode.
The mean, median and mode are all valid measures of central tendency, but under
different conditions, some measures of central tendency become more appropriate
to use than others. In the following sections, we will look at the mean, mode and
median, and learn how to calculate them and under what conditions they are most
appropriate to be used.
18
SANGA
Mean (Arithmetic)
The mean (or average) is the most popular and well known measure of central
tendency. It can be used with both discrete and continuous data, although its use is
most often with continuous data. The mean is equal to the sum of all the values in
the data set divided by the number of values in the data set. So, if we have n values
in a data set and they have values x 1, x 2 …, x n, the sample mean, usually denoted by
x (pronounced "x bar"), is:
x 1+ x 2+ …+ xn
x=
n
This formula is usually written in a slightly different manner using the Greek
capitol letter, ∑, pronounced "sigma", which means "sum of...":
∑x
x=
n
Example, consider the wages of staff at a factory below:
Staff 1 2 3 4 5 6 7 8 9 10
Salar 15k 18k 16k 14k 15k 15k 12k 17k 90k 95k
The mean salary for these ten staff is $30.7k. However, inspecting the raw data
suggests that this mean value might not be the best way to accurately reflect the
typical salary of a worker, as most workers have salaries in the $12k to 18k range.
19
SANGA
The mean is being skewed by the two large salaries. Therefore, in this situation, we
would like to have a better measure of central tendency. As we will find out later,
taking the median would be a better measure of central tendency in this situation.
Mean for grouped data
In Tim’s office, there are 25 employees. Each travels to work every morning in
his/her own car. The distribution of the driving times (in minutes) from home to
work for he employees is shown in the table below:
Class 0 – 10 10 – 20 20 – 30 30 – 40 40 – 50
Frequency 3 10 6 4 2
Solution
Find the class mark for each class by adding the lower and upper class limits of the
class and divide by 2.
Class Interval (x) Frequency (f) Fx
0 – 10 5 3 15
10 – 20 15 10 150
20 – 30 25 6 150
30 – 40 35 4 140
40 – 50 45 2 90
∑ 25 545
∑ fx 545
x=
∑f
= 25 = 21.8
20
SANGA
Median
The median is the middle score for a set of data that has been arranged in order of
magnitude. In order to calculate the median, suppose we have the data below:
65 55 89 56 35 14 56 55 87 45 92
We first need to rearrange that data into order of magnitude (smallest first):
14 35 45 55 55 56 56 65 87 89 92
Our median mark is the middle mark - in this case, 56 (highlighted in bold). It is
the middle mark because there are 5 scores before it and 5 scores after it. This
works fine when you have an odd number of scores, but what happens when you
have an even number of scores? What if you had only 10 scores? Well, you simply
have to take the middle two scores and average the result. So, if we look at the
example below:
65 55 89 56 35 14 56 55 87 45
We again rearrange that data into order of magnitude (smallest first):
14 35 45 55 55 56 56 65 87 89
Only now we have to take the 5th and 6th score in our data set and average them to
get a median of 55.5.
21
SANGA
Mode
The mode is the most frequent score in our data set. On a histogram it represents
the highest bar in a bar chart or histogram. You can, therefore, sometimes consider
the mode as being the most popular option. An example of a mode is presented
below:
14 35 45 55 55 55 56 65 87 89
Assignment
1. Stephen has been working on programming and updating a web site for his
company for the past 15 months. The following numbers represent the past 7
months:
24, 25, 31, 50, 53, 66, 78
What is the mean (average) number of hours that Stephen worked on thei
website each month?
2. a. find the median of the following data
12, 2, 16, 8, 14, 10, 6
b. find the median of the following data:
1, 1, 3, 4, 4, 6, 7, 8, 9, 11
3. a. The ages of 12 randomly selected students are listed below:
23, 21, 29, 24, 31, 21, 27, 23, 24, 32, 33, 19
22
SANGA
b. find the mode of the following data:
76, 81, 79, 80, 78, 83, 77, 79, 82, 75
Bibliography
Last, J. M. (1997). Public Health and Human Ecology, 2nd edition. Stamford, CT:
Appleton and Lange.
Slee, V. N.; Slee, D. A.; and Schmidt, H. J. (2000). The Endangered Medical
Record. St. Paul, MN: Tringa Press.
Webb, E. J.; Campbell, D. T.; and Schwartz, R. D. et al. (1988). Unobtrusive
Measures: Non-interactive Research in the Social Sciences. Chicago: Rand
McNally.
23