Understanding Health Statistics Basics
Understanding Health Statistics Basics
Health
It’s a state of complete physical, mental and social well-being and not merely the absence of
diseases or infirmity.
Statistics
It’s the study of how to collect, organize, analyze and interpret numerical information from data.
It’s the body of theory, concepts, methods and methodologies of collection, analysis,
interpretation and representation of numerical data for decision making.
Health statistics
It’s the science of collecting, organizing, summarizing/analyzing, presenting and interpreting
health data and using them for decision making.
Uses of statistics
- Used in the medical investigation as it provides a way of organizing information instead
of relying on the exchange of anecdotes and personal experience.
- Interpretation of data helps to draw general conclusions about phenomena.
- Statistical data helps health administrators to predict health events and outbreak of
diseases.
- Analysis of statistical data helps to provide explanation to some of the causal factors of
underlying complex situations.
- Description of large masses of data is made possible by use of statistical methods which
makes them more meaningful and easy to understand.
- Helps service managers and administrators to use statistical data to plan, organize and
discharge services in the health facility.
Type of statistics
Statistics has 3 distinct parts namely:
1. Descriptive
2. Inferential
3. Experimental
Descriptive statistics
Summarizes population data numerically or graphically.
1
It can be defined as those methods involving the collection, presentation and characterization of
a set of data to properly describe the various features of that set of data.
It’s a type of statistics that describes large masses of data.
e.g. statistics pertaining to central tendency such as mean, median and mode.
- Statistics pertaining to dispersion around the central tendency such us range, standard
deviation, variance, and quartile deviation.
- Statistics of graphs depicting the shape of a distribution.
Inferential Statistics
They include;
i. Estimation
ii. Modelling relationships
iii. Hypothesis testing.
It’s a type of statistics that infers or induces from a small group (sample) and generalizes to the
whole group (population).
Also known as inductive statistics. It deals with the method of drawing conclusions from
numbers observed.
Involves drawing the right conclusions from the statistical analysis that has been performed
using descriptive statistics.
Most predictions of the future and generalizations about a population by studying a smaller
sample.
i. Estimation
It’s the group of statistics which allows for the estimation about population values
based upon sample data. E.g population parameter estimates and confidence inter
ii. Modelling relationships
Allows us to develop mathematical equations which describe the interrelations
between 2 or more variables.
iii. Hypothesis testing
Allows us to test for whether a particular hypothesis we’ve developed is supported by
a systematic analysis of the data.
As built on the descriptive statistics by going a step further to make interpretation
with a view to population upon which a decision would be based e.g cli-square, t-test,
f-test etc.
2
Experimental statistics
Relates to the design of experiments to establishing causes and effects of such designs as
experimental, anasin experiments etc.
Types of data.
There are 2 types of data, namely;
1. Primary data.
2. Secondary data.
Primary data
Its data especially collected for the purpose for which its required.
Data that one collects by himself.
Data can be collected the whole population or use of sampling by investigating only part of the
population.
3
Where the information can be measured or counted. We can arrange for ourselves or
someone else to take the necessary measurements directly and using the 3 rd party
influence. This method has the distinct advantage that accurate measurements can be
obtained.
It can however turn out to be very time consuming and costly if misused.
ii. Experimentation
It’s a method where a particular Rx is given to the population or the people to
determine the effect it has on the population. E.g applying fertilizer on different crops
to determine their yields etc.
iii. Observation.
It involves watching or counting events as they happen.
It involves getting data when investigation concerns attitude and behaviour.
Its applicable when all other methods cannot be effectively used due to the situation.
Data can be obtained by either directly or indirectly. Its sometimes referred to as
participatory or non-participatory.
a) Direct/participatory observation.
The investigator goes to the study situation and behaves like the respondent.
For instance, if you are carrying out an investigation on alcoholic behaviour, you
will go and mix with alcoholics in their drinking places and behave like them
recording your observations.
Respondents may or may not know that you are there as an investigator.
Another example can be taken from (CID) who does a passive but participatory
role in data collection.
The observer may ask a series of questions informally or administer a test.
b) Indirect/ non participatory observation
In this method, the respondent is observed through a device e.g through one way
mirror, tape recorders, cameras, closed circuits or even a key hole.
In most cases the respondent is not aware that he is being observed.
iv. Interviewing
This is a method of data collection in which you or your assistant conducts an
interview conversation with respondents while completing a pre-designed interview
schedule.
There are advantages and disadvantages of conducting the interview by yourself or
employing an interviewer.
4
- Reduces chances of wrong recording of information by the investigator is more
knowledgeable about the investigation than the interviewer.
Disadvantages.
- Information collected may not be as accurate as the one the investigator collects, this is
because the interviewer may misinterpret the questions wrongly when conducting the
interview.
- The assistant may record the information wrongly taking the editing session very difficult
and expensive.
v. Questionnaire
This is where an investigator/interviewer has a set of questions (the questionnaires),
designed questions which he/she administers to the respondent (the person chosen to
be in the sample).
It has the advantages that the results can be obtained quickly and reliably. Due to high
response rate obtained when using this method, its usually considered the most
effective way of obtaining data.
This method is subdivided into 3 parts;
1. By post
2. Face to face interview
3. By phone.
1. Post.
The questions are sent to the respondent through the post office, the respondent fills the
questions and posts back.
It requires that:
i. Write an introduction letter
ii. Give an explanation on how questions should be answered.
5
iii. A statement that ensures the respondent that the information given by them will
be treated highly confidential.
iv. Enclose a stamp and addressed envelope for return postage.
Advantages of postal questionnaires
- Can be sent to a very large no. of people in the population at a low cost.
- There is no risk of bias or mistake that can originate from the interviewer.
N/B; the method of data collection that you need to use depends on:
i. Available resources that include money, man power and time. However
the method that gives you the best results utilizing the available resources
to its maximum. The method therefore should enable you to conclude your
investigation effectively.
ii. A thorough consideration of advantages and disadvantages should be
considered so that you can apply the method that has less disadvantages.
Secondary Data
This is the data which had already been collected by ourselves or some other agencies e.g the
central bureau of statistics publishers. Such data include economic statistical tables for a different
purpose but which is relevant to our present investigation.
It’s the data taken from some other sources.
This is data that already exists in some form.
6
investigating the figure to the accuracy we require and finally the trustworthiness of the
figures.
- Is the information actual, seasonal, adjusted, estimated or projected?
- The reason for collecting the data maybe unknown.
- The data maybe incomplete.
Variables
This is any quality that can have a number of values which may be either discrete or continuous.
It’s a property that can take on different values. It’s any attribute, property or gender that
changes.
It therefore refers to all those characteristics of people, events, objects or situations that change
e.g students in a class may differ in sex, age, intelligence, height etc. These properties are
variables.
Therefore variables could vary in quality and quantity.
Type of variables
a. Organismic variables
They are human oriented independent variables e.g gender (sex) or age.
b. Qualitative variables
This type of variable differs in terms of quality (kind).
They are classified according to that attribute (quality or characteristics) e.g
gender, nationality, social economic status, academic qualification, marital
status etc.
Taste can be classified as salt, sour.
Smell can be classified as odour, pungent.
c. Quantitative variables
This variables assumes values that vary in terms of magnitude. They differ by
how much or how many thus means they can be given numerical basis.
They are very easy to measure/count and compare with others e.g weight,
height, age, distance, marks obtained in a test etc.
Quantitative values can be sub-divided into:
i. Continuous variables
ii. Discrete variables.
i. Continuous variables
These are variables that can take any value (within a certain range) and are not therefore
restricted.
7
They are characterized by being related to some numerical scale of measurement, any interval of
which may of desired, be an infinite number (any number that contains an endless line of digits
after a decimal point). E.g length, height, weight, temperature, volume, time etc. because they
take fractions or decimals e.g 51/2 ft, 2.6kg etc.
ii. Discrete variables
They are variables that can take finite or defined (whole numbers)
They can be counted or for which there is a fixed set of values e.g 1,2,3,4 and 1.2, 21/2, 31/2 etc
Independent variables
These are variables that can be manipulated or treated.
The effect is reflected on the dependent variable. The value of the dependent variable thus
depends on that of the independent variable. It’s a variable that which is subject to the researcher
e.g a pupil’s performance (y) depends on age (x).
Y-dependent variable
X-independent variable
Bumper harvest depends on the rainfall.
Note that in graphing, the dependent variable is placed on the vertical (y-axis) while the
independent variable is placed on the horizontal (x-axis).
Dependent variable
These are variables where the effects of independent variable are felt/seen (the one being
influenced) e.g rational use of medicine depends on educational level, occupational or socio-
economic status.
Distribution
This is the arrangement of a set of numbers classified according to some properties or attributes
such as age, height, weight etc.
Population
This is any defined whole set of people, situations or objects. The definition of a populatin is
done in terms of certain characteristics e,g if there are 100 students in health records class, we
say the population size is 100. This measurement is of interest representing the aggregate of units
to be covered which could be finite of infinite. When the population can be easily counted then
its said to be finite e.g the number of contestants for a political post but if the population under
consideration is large e.g the grain of sand, then we say it’s infinite.
8
Sample
This is a sub-set of a population. It’s a sub-group or sub-aggregate draw from the population. i.e
the proportion appropriately selected out of the population by the same statistical method of
observation. It’s a small set of a population. It’s a group of items or part of a population selected
to represent such a population in a statistical investigation.
A sample must be representative of the target population. It must have all or most of the
characteristics of the population.
Sampling frame
Refers to a list of all items that constitute a given population e.g voters register may be used as a
sampling frame for the voting population.
Large sample
This is the sample with more than thirty items.
Small sample
This is the sample with less than thirty items.
N/B: the size of your sample depends on the size of your population. The larger the population,
the larger the sample size should be. It should be proportional to the population to reduce
sampling error.
Parameter
Any numerical value describing a characteristic of a population. It’s a situation when mean
(average), standard deviation or variance of a population are computed for statistical analysis. Its
any number or value calculated from the data of a population.
Statistic
This refers to a descriptive measure of a sample i.e a numerical value or function computed to
describe a sample of a population.
Constant
It’s the opposite of variable, therefore it does not vary (change)
9
- Focus group discussion guides
- Key informants guide
i. Questionnaire
This is a research instrument that gathers data over a large sample. Each item in the
questionnaire is developed to address each specific objective.
A questionnaire that is well thought of saves time. There is no opportunity for
interview bias and can cover a wide area.
Open-ended questionnaires can generate large amount of data that can take long time
to process and analyze.
Advantages
a. Large amount of information can be collected from a large no of people in a short
period of time and in a relatively cost effective way.
b. Can be carried out by the researcher or by any no. of people with a limited effect
to its validity and reliability.
c. The results of the questionnaire can usually be quickly and easily quantified by
either a researcher or through the use of software package.
d. Can be analyzed more scientifically and objectively than other forms of research.
Disadvantages
a. They may confuse the respondent as to the nature of the information required. The
respondent will not answer right as he does not know the purpose of the study.
b. May discourage respondents to the extent of discarding the questionnaires.
c. May leave important information required in the study; some of the questions may
be omitted. The respondent may not understand the question well or was bored.
d. Response rates can be very low, the respondents will not post back the
questionnaires.
e. The participants may forget important issues because they occur after the event.
f. Questionnaire is standardized so its not possible to explain any points in the
question that participants might mis-interpret. This could be partially solved by
piloting.
g. There is no direct contact so the researcher cannot deal with any other mis-
understanding.
h. There is no opportunity for the researcher to ask for further information or probe
deeper to the answers given to the respondents.
There are two broad categories of questions that are used in the questionnaires
a. Structured or close ended.
b. Unstructured or open-ended
Structured
10
These are questions which are accompanied by a list of all possible answers from which
the respondents select the answer that describes their situation.
Leave allowance for others e.g others (specify)
Advantages.
1. Saves time for the researcher. Economical to use in terms time and money.
2. The researcher gets the information he requires.
3. It’s easier to analyze since they are in an immediate usable form.
4. It’s simple for the respondent to understand.
5. They are easier to administer because each item is followed by an alternative answers.
Disadvantages.
1. It’s difficult to construct because the answers must be well thought of.
2. Responses are limited and the respondents are compelled to answer questions
according to the researcher’s choices.
Unstructured
This refers to questions that give the respondent complete freedom of expressions. These
free expressions permit an individual/respondent to use his or her own words.
Advantages.
1. Are simple to formulate because the researcher does not have to labour to come up
with appropriate response categories.
2. The respondent’s responses may give an insight into his feelings, background, hidden
motivation, interests and decisions.
3. It gives the researcher the real feelings or picture of the respondents.
4. It gives respondent’s freedom to excess him/herself in his/her own words.
5. It saves time for the researcher.
6. Can stimulate the respondent to think and express what he/she considers most
important.
7. It permits a greater depth of the respondent when the respondent is allowed to give a
personal response usually reasons for the response given may be directly or indirectly
included.
Disadvantages.
1. It’s time consuming for the respondent.
2. It’s tedious and cumbersome for the researcher to compile the responses.
3. There is tendency to provide information that does not answer the stipulated research
questions or objectives.
4. Responding to open ended questions is time consuming. This may put off some
respondents.
11
Contingency questions
These are follow up questions needed to get further information from the relevant sample. These
questions probe for more information. They also simplify the respondent’s task in that they will
not be required to answer questions that are not relevant to them.
Matrix questions
These are questions that share the same set of response categories e.g
Were the learning objectives achieved?
1. Agree
2. Do not agree
3. No
4. Uncertain
5. Strongly disagree
6. Strongly agree
Time for lecture is well distributed?
1. Strongly agree
2. Agree
3. Uncertain
4. Disagree
Circle a number to represent your response.
Advantages
- Simple to answer.
- Space is efficiently used.
Disadvantages
- Some respondents especially the ones that may not be too keen to give the right responses
might form a pattern of agree or disagree with the statement.
12
classes about history and culture? What motivates your work? Pleasant work and nice co-
workers.
5. Leading or biased questions should be avoided e.g. asking for the gender of a child-like is
he a boy when you see by yourself. Were you at the KCs bar on the night of 15th?
6. Very personal and sensitive questions should be avoided as the respondent may be
dishonest in answering them.
7. Simple words that are easily understandable should be used. Difficult words that are not
familiar will discourage the respondent. Avoid jargon words e.g. short hand etc.
8. Questions that assume facts with no evidence should be avoided. Such questions offend
and discourage the respondent e.g. asking a mother if she affords formula milk for
breastfeeding by looking at her social class. This may discourage her to respond. You got
HIV because you were unfaithful.
9. Avoid psychologically threatening questions e.g. asking a mother, are you afraid that
your child who is HIV positive is about to die?
Interview Schedule
This is a data collection tool which is very much like questionnaire. The little difference is that
the tool is filled by the enumerator/investigator/researcher/interviewer who are specially
appointed for the purpose.
N/B: In certain situations the tool may be handed over to the respondent and the enumerator
helps in recording the answers.
Enumerator explains the aims and also removes the difficulties which the respondent may fail in
understanding the definition or concept of difficult times.
13
Advantages
- More information and in greater depth can be obtained. The enumerator has a chance to
probe for more information.
- There is greater flexibility as there is an opportunity to restructure questions.
- Personal information can be obtained more easily.
- There is no missing returns and non-response is low.
- There is control on who will answer the questions.
- The language of the interview can be adopted to the ability or educational level of the
respondent and as such misinterpretations concerning questions can be avoided.
Disadvantages.
- There may arise a communication barrier between the interviewer and the respondent.
- The program of the interviewer may not coincide with that of the respondent.
- Its time consuming as the interviewer can take a lot of time in one respondent.
- There is a tendency of biasness (interview bias).
- There is a probability of shunning down of the interviewer by the respondent e.g a young
person interviewing an old person about his sex life.
- Very expensive especially when large and widely spread geographical sample is taken.
- Certain types of respondents such VIPs may not be approachable under this method.
2. Rating schedule
Set of questions that helps guide a psychologist or sociologist to measure the attitude and
behaviour of an individual.
3. Survey schedule
Formulated for a surveyor to guide him on his information collection.
4. Interview schedule
Set of questions with structured answers to guide an interview.
14
2. Address whether or not the responses will be recorded. Recording responses has the
advantage that all materials are captured but the disadvantage that many respondents may
speak less freely if they are being recorded obtain a consent before recording.
3. It must contain questions that will answer the research questions.
4. It must not contain questions that are not related to the research questions.
5. Each question must address a single use.
6. Topics that might be covered include:
a. Demographic questions-age, education, position (occupation),gender can be
observed.
b. Knowledge- questions about what a person knows about a specific topic.
c. Behaviour- questions about what a person does in general or has done on a specific
occasion and/or what a person plans to do in future.
d. Opinions or values- questions as to what a person thinks about a topic.
e. Feelings- questions about how a person feels about an issue.
f. Sensory- questions as to what a person has seen, touched, heard, tasted or smelt.
Wording of questions
Wording should be open ended. Questions that require a one word answer, such as yes or no,
should only be used as an introduction to more searching questions or for specific facts.
Questions must be neutral. Avoid using many questions that might lead the respondents to
provide a specific answer or that might generate a strong emotional reaction.
Questions should be worded clearly, using simple words, and not using any jargon, abbreviations
or technical terms.
Order of questions
To get the respondent involved in the interview, start with some non-controversial questions such
as facts.
Ask questions about the present before asking about the past or the future.
Use the last question of the interview to allow the respondent to talk about any issue that he/she
thinks is important and has not been addressed or has not been addressed in sufficient depth.
15
- If the respondent does not understand a question, the interviewer can interpret it.
- The respondent is not provided with potential resources so that respondent’s genuine
views are obtained.
- The interviewer can notice that a respondent is distressed and either terminate the
interview or take steps to reassure the respondent.
- The data are complete as the interviewer can check that all issues have been covered and
all relevant demographic data is recorded.
Checklist
This involves a schedule containing a set of questions which are filled by the researcher or
enumerators as he/she directly observes things around him/her which are of interest by asking the
respondent.
The checklist is systematically planned and includes all the items or points that must be
considered during observation in a field or when extracting data from existing records. E.g
ventilated pit latrine, you want to observe the functionality.
Set of questions will be:
a. Evidence of use YES NO SOMETIMES
Advantages
- Subjective bias is eliminated if observation is done accurately.
- It’s less demanding of active cooperation on the part of the respondent.
- The information obtained under this method relates to what is currently happening or
what recently happened.
Disadvantages
- It’s expensive in terms of labour.
- Information provided by this method is very limited.
- Sometimes unforeseen factors may interfere with the observation task.
16
- Participants are randomly selected. To obtain meaningful information, a highly skilled and
trained facilitator must guide the group but careful not to leave it in pre-determined direction.
- It’s of 6-8 participants.
17
impossible or are inappropriate to use.
- Access to people in real life situations. - time consuming
- Good for explaining meaning and - depends on the role of researcher.
content.
- Can be strong in validity and in depth - Overt may affect the situation and thus
understanding. validity of findings.
18
Key informants guide
This is another type of data collection tool that the study involves interview.
Open ended or unstructured questions are used to collect the information from the opinion
leaders or knowledgeable community leaders.
These people may either confirm or reject data that has been featured using other research
methods.
Sampling methods
Population
In statistics, population means total number of items in a specific field of inquiry. It’s the entire
group of individuals or objects under consideration.
The examples of population are number of wild animals in a national park, total number of cars
in a country, total number of students in a college, total number of human beings in a country
etc. population is also known as universe.
In order to collect statistical data, census method and sampling method can be used.
In census method, all the units of a population or universe are contracted e.g in the census of
population of Kenya, all the individuals residing in Kenya is included. Although this technique
provides complete and accurate information but it’s very expensive and inconvenient.
A sample.
When the sample are selected from a population, the units selected must be taken random.
Sample is a portion, piece or segment that represents the whole (population).
According to this method, a few units from the whole population must have equal chance to be
selected. The units selected are just by chance or coincidence. If these units selected are not
taken at random then bias will take place and undue importance will be given to some units. In
this case the sampling will not be fair and representative e.g if from a class students selected to
be part of the study will be only intelligent students. This will not be representative regarding the
performance.
Sampling
The process of selecting a number of individuals for a study in such a way that the selected
individuals represent the large group from which they were selected.
The purpose of sampling is to secure a representative group which will enable the researcher to
gather information about a population.
Advantages of sampling method
- As only a small part of the whole population is studied, its cheaper to collect the data.
19
- The data are collected and analyzed more quickly, thus sampling saves a lot of time.
- Since only a part of the whole population is to be studied a good quality of labour with
better supervision can be provided.
- An investigation of a small part of the population gives us more detailed information.
Sampling frame
Refers to a list of all items that constitute given population.
A list of entire population from which items can be selected to form a sample. E.g voters register
may be used as a sampling frame for the voting population.
Types of sampling
1. Simple random sampling
The name comes from the fact that no complexities are involved.
Random expresses the idea of chance being the only criterion for selection.
It’s therefore a sampling procedure that provides equal opportunity of selection from each
element in a population. All is needed is clearly defined population in a sample frame
(boundaries should be defined).
Random samples are satisfactory when the population is homogeneous (uniformity of certain
characteristics)
The main objective of the simple random is to eliminate any form of bias in the selection and to
obtain a representative sample.
There are various techniques of selecting randomly, the most common is lottery technique,
where a symbol of each unit population is placed in a container mixed well and then the likely
numbers are drawn, that constitute the sample.
A more sophisticated method particularly used for large populations is the use of random
number tables. These tables are mathematically prepared so that numbers are written in a
random way and therefore each item has an equal probability of being selected.
Assignment
Read and make notes on how to get a sample using random number tables.
Example
Suppose you want to investigate the socio-economic status of patients with diabetes in Kisumu
county, here are the steps to follow;
1. Get all the information on the total number of people with this condition (sampling
frame). Such information can be obtained from existing health records or previous
investigation.
2. Decide on your sample size
20
3. Give a number to each patient in the location within Kisumu County.
4. Write those numbers on pieces of paper: fold them properly making sure that the numbers
cannot be seen.
5. Put all the papers in the basket and shake them properly.
6. Pick any of the papers at random, repeat several times until you reach your sample size.
7. You will then go to the location and interview those patients whose numbers you have
picked.
K=N/n
K=10000
1000
K= 10 interval.
- Determine the starting point from which you will start picking every 10 th item e.g using
simple random, if the first randomly selected sample is number 3 and K is 10 th item after
number 3,then the items or samples we shall come with are 3,13,23,33 etc all selected
systematically till you reach your sample size of 1000.
21
Assignment
Advantages and disadvantages of simple random sampling and systematic sampling.
Advantages of simple random sampling
- Its straight forward and probably the simplest method of sampling.
- It prevents the sampling bias and error.
- The method is fair.
- It’s simple and makes data interpretation easier.
Disadvantages
- It’s more time consuming than non-representative methods.
- It’s expensive.
- Each person has to be located and questioned.
Advantages of systematic random sampling
- Easy to organize.
- More precise than simple random sampling and more evenly spread over population.
- Simple to apply the analysis of data and has a sound mathematical basis.
- Biasness is eliminated.
Disadvantages
- There is no guarantee that the behaviour of these people represents the behaviour of the
other groups.
Stratified sampling.
This is a method of obtaining a sample from a population when distinct groups (strata) of a
population can be identified. The principle of this sampling is to divide a population into
different groups called strata such that each element of the population belongs to one and only
stratum.
The population divided into groups should be in such a way that units within each group are as
similar as possible. The groups should be homogeneous.
The composition of groups can be for instance different tribes, religions, gender, socio-economic
groups, occupation, age, income groups etc.
The chosen variables should be one that result in internally homogeneous stratum, then within
each stratum, random sampling is applied using either simple or systematic random method to
choose the sample.
Stratified sampling can be proportional or non-proportional. In proportional sampling the
participants are chosen in proportion to the number in each group. Non-proportional occurs when
the response weight of the sub-group is not a factor.
22
Examples
The ministry of health wants to determine if geographical location has a significant effect on
health care workers support for a merit pay plan
Draw later
Advantages
It allows representation of district groups in a population proportionally, this is because the
samples are taken in accordance to this proportion of the constituent parts of the whole
population.
Example 2
1. Assume that Kisumu town is divided into two district locations. It can therefore stratify
its population into two groups according to their locations.
2. Suppose Kisumu town has a population of 100 000 and this population is distributed
between the two locations namely: Kisumu East and Kisumu West. The proportion
representation of this locations in town is 3:7 .
23
Therefore if we wish to study for instance malaria prevalence as a health condition which
has affected 10,000 people in Kisumu town, we shall take our sample in accordance to
3:7 representation of the location in the town.
If our sample is 1000 cases only
How do you obtain 1000 cases from 10,000
Steps
1. Obtain information about the total population and population of a location in Kisumu
town which is 100,000 and locations (Kisumu East 30,000 and Kisumu West 70,000)
2. Obtain proportion representation of a location, this is worked in the form of a ratio as
follows
=7/10
3. Ten is common in both locations and therefore we shall say that the proportional
representation is 3:7. This means that for every 3 people in Kisumu East location there
are 7 people in Kisumu West location.
4. If the sample size is 1000, apply the same ratio3:7 in which the population occurs in
order to get smaller sub- samples which are 3/10 of 1000 for Kisumu East and 7/10 of
1000 for Kisumu West, this smaller sub –samples are called stratified samples.
5. After establishing the sub-samples, then apply simple random sampling to get the
actual cases to be investigated.
Sampling Method
-The population is divided into first sampling units which are sampled by simple random
sampling
- The selected units are therefore subdivided into second stage units and the sample is
selected out of these units.
-The selected second stage units are then further sub divided into 3 rd stage units and again
a sample is selected out of 3rd stage units
Example
Suppose we had a division known as Rera Division, there is a justification that a control
programme on the spread of malaria must be started in the whole of Rera Division,
however there is need to start a pilot programme in two villages in order to provide us
with information on how effective the major control programme will be.
We are required to select two villages in which the pilot programme will be introduced.
Rera division has 24 villages spread equally over the 3 locations and 12 sub locations.
24
Steps
1. Show the layout of Rera division in order to depict the constituent locations, sub-
locations and villages.
RERA DIVISION
LOCATION 1
SUBLOCATION 1 2 3 4
VILLAGES 1 2 1 2 1 2 1 2
RERA DIVISION
LOCATION 2
SUBLOCATION 1 2 3 4
VILLAGES 1 2 1 2 1 2 1 2
RERA DIVISION
LOCATION 3
SUBLOCATION 1 2 3 4
VILLAGES 1 2 1 2 1 2 1 2
-.
-Out of the three locations in Rera division, select two of them using simple random sampling e.g
1&2,2&3, 1and 3.
-In each selected locations there are 4 sub-locations, in each sub-location there are two villages.
- Apply the same random sampling method to select one village in each of the two sub-locations
e.g
1 -Division
3 –Location -2 locations picked
12-Sub-locations -2 sub-locations picked
24-Villages -2 villages picked
-Having selected the two villages they are your representative sample where the control
programme would be started on a pilot basis
25
N/B
The programme will be introduced to affect everybody in the two villages .Each person in the
two villages is considered to be a true representative sample of the population in Rera division,
this gives everybody in this two villages a right to be interviewed.
Cluster sampling method
This is a method of sampling which follows the same basic steps as in multi stage sampling
method. The difference in cluster sampling and multi-stage sampling is that, cluster sampling
deals with only specific selected places or characters referred as clusters.
It involves selecting some known members of a population for the study e.g footballers,
prostitutes, adolescents, women etc.
Eg In Rera division we would take the following steps as in;
1 Follow the same steps as explained in multi stage sampling
2 Issue instructions specifying who will be interviewed in the already identified areas e.g people
living within 100m from the ponds, dams, lake etc
The number of specification to be included in your investigation forms a group of clusters within
the population to be interviewed. This is with biased method of sampling, because there is a
deliberate focus on some groups to the exclusion of others. When a particular strategy is selected
for the study, then, either cluster or focus group of sampling is said to have been used in
selecting the sample points.
Quota sampling
This is a method of sampling in which interviewers are sent to different survey locations with
instructions to interview all people up to a given number, this given numbers are what we call
quotas
The enumerators are given a quota of say 400 people and are told to interview all the people they
can until their quota has been met.
Such quota is nearly always divided up into different types of people with sub-quotas for each
type e.g
Out of 400, the enumerator may be told to interview 250 working wives, 100 non-working wives
and 50 unmarried women.
Sample size determination
There is no specific size for a statistical study. Sample sizes depend on the type of study being
conducted and the population being studied.
Sample size for a given population
26
NUMBER (N) S N S
10 10 150 108
20 19 200 132
30 28 300 169
40 36 500 217
50 44 1000 278
60 52 2000 322
70 59 5000 357
80 66 10000 370
90 73 50000 381
100 80 100000 384
General Rules
-A smaller percentage is required for a larger population.
Studies using a population less than 100 should use the entire population
Population over approximately 300 require a sample size of 50/
Populations over 100,000 would require only 384 in the sample population.
Data cording
A systematic way in which to condense extensive data sets in to smaller analyzable units through
the creation of categories and concepts derived from data.
When data have been collected in research they have to edited and coded in a numerical form
ready to be summarized to table, charts, diagram or grouped into frequencies before calculations
are made.
Levels of Coding
Open
Breakdown compare and categorize data.
AXIAL
Make connections with categories after open coding
SELECTIVE
Select the core category, relate it to other categories and confirm the explanation to those
relationships.
27
Why coding
It lets you make sense of and analyze your data.
For qualitative studies it can have you generate a general theory.
The type of statistical analysis you can use depends on the type data you collect, how you collect
it and how its coded.
Coding facilitates the organization, retrieval and interpretation of data and leads to conclusions
on the basis of that interpretation.
Editing the collected data
It’s the examination or scrutiny of the collected data in order to find out mistakes, errors and
omissions. The aspects of accuracy, approximation and errors are analyzed.
Accuracy
It involves describing a phenomenon exactly as it is.
Absolute or perfect accuracy cannot be obtained therefore in statistics relative accuracy is
required and not absolute.
The degree of accuracy depends on the nature and purpose of inquiry and also on the materials of
measurement. E.g when measuring the height of men and women it should be accurate up to
inches or centimeters and while measuring the length between two cities it should be accurate up
to kilometer.
Approximation
This is the basis of rounding off the figures with a view to simplify them thus affecting the
standard of reasonable accuracy.
There are 3 methods of rounding:
a. Round up
b. Round down
c. Round to the nearest whole number.
If the last figure is 75, then 1 is added to the last significant figure [Link] 25’268 rounded to two
decimal places it will be 25.27
If the last figure 5 then last significant figure is taken as it is e,g if 25.263 rounded to two
decimal places then it will be 25.26.
If the last figure is 5 then 1 can be added to the last significant figure or it can be taken as it [Link]
of 25.265 is to be rounded off two decimal places, then it can be taken as 25.27 or 25.26. its
correct in both ways.
28
If a figure is rounded to the nearest whole figure in thousands then the portion being left is
ignored if it’s less than 500, next higher figure in thousands is taken, if it’s more than 500 and so
on. E.g
1. 15,856,432- 15,856,000
2. 15,856,854-15,857,000
3. 15,856,500-15,857,000 or 15,856,000
Errors
The word error has a special meaning in statistics. We can distinguish between mistake and
error.
A mistake, means incorrect presentation or man factors. Can occur in the collection of the data.
E.g the respondent may have mistakenly ticked the ‘yes’ box instead of ‘no’ box.
An error means the difference between the actual figures. The deviation is just by chance and it’s
not due to carelessness of human beings.
Normally the errors arise due to rounding off or approximation.
Main sources of statistical errors
a. Errors of origin
They crop up due to faulty definition of the units, bias or errastic trends in the data.
b. Errors of inadequacy
Errors due to inadequacy of samples or incomplete information.
c. Errors of manipulation
Errors which result while measuring, weighing and counting unconsciously.
Types of errors
1. Sampling errors
The difference between the estimates of a value as obtained from the sample and the actual
value.
Even when the sample is chosen in a correct manner, it cannot exactly be representative of the
population from which it’s chosen because all samples will not be similar.
The amount of the sample error will depend on the size of the sample.
The greater the size of the sample the smaller the size of sampling errors and vice versa.
2. Non- sampling errors
29
Those errors which take place when samples are not selected at random. When a sample is
chosen at random, it cannot be exactly representative of the population from which it’s chosen.
The difference between the characteristics of a sample is known as sample error.
When the errors arise due to other reasons, those are known as non-sampling errors.
3. Biased errors
These errors arise due to the bias on the part of investigator, enumerator or instrument etc. They
are of cumulative in nature. E.g the instrument is more or less than actual or the investigator
ignores some digits with giving any weight/reason etc.
4. Unbiased errors
These errors arise by chance in the usual course. These are of compensatory nature because
positive and negative errors cancel each other and mostly the estimated value is just equal to true
and actual value.
5. Positive and negative errors.
Errors may be positive or negative if the true value is greater than estimated value, the error is
said to be positive , on the other hand, if it’s less then it’s said to be negative.
Measurement of errors
Errors can be measured either absolutely or relatively
a. Absolute error
This is the difference between the actual value and estimated value.
Ae= A-E
Where;
Ae = Absolute error.
A = Actual Value
E = Estimated Value
e.g Assume the population of a town was estimated as 1,424,880 where as the actual population
was 1,578,620.
30
Ae = 1,578,620- 1,424,880
= 153,740.
b. Relative error
This the ratio between absolute error and actual value.
Re = Ac
A
Where: Re = relative error
31
- You can also extract data from published statistics. These are obtained in special surveys
and publications from individual research projects and central bureau of statistics such as
demographic health survey and contraceptives prevalence surveys.
- Data obtained from each of the above named sources have their advantages and
disadvantages.
Study using the whole population or part of the population
Advantages Disadvantages
- Gives very reliable data because every - Very expensive to conduct using
member of the population is whole population.
interviewed and when the sample is
used detailed information can be
obtained because of dealing with
fewer people.
- Gives information that includes those - Difficult to follow up of individuals
who attend and those who do not who might not be in during the study
attend a health facility. period.
- Sampling may be the only feasible - Impossible to show the trend of the
method of collecting the information. condition
- Has a short term value, this is because
its value ends immediately after the
end of investigation.
- There is always a sampling error that
will tend to reduce the strength of
investigation when samples are used.
32
- The collected data are mostly large in quantity and its necessary to organize the data in
such a way that further analysis and interpretation of data are made easily and correctly.
- Before the tabulation of data takes place, it’s necessary to classify the data into
homogeneous groups.
Classification of data
- It means the act of arranging the data in groups or classes according to some resemblance
of the data in each group or class.
- In classification of data, the elements which posses the same characteristics are grouped
in one class. In this way, the whole data is divided into a number of classes.
- L.R Connor has defined classification as ‘the process of arranging things in groups or
classes according to their resemblances and affinities (natural liking for or attraction to a
person, thing, idea etc). Thus classification is the sorting of the data into homogeneous
groups according to their observed characteristics.
- It’s the process of keeping data with common characteristics together i.e data concerning
males, females, children age groupings need to be put separately from others.
- This enables one to analyze data concerning a specific character or items more easily.
This means its important to know what items go together when you have classified
information.
- There are two main ways of classifying items from the questionnaire or data collection
instrument. They are classified in terms of variable (quantitative) or an attribute
(qualitative) data.
1. Quantitative classification
It refers to the classification of data according to some characteristics that can be
measured such as height, weight, age etc. in this, data are classified by assigning arbitrary
limits called class-limits.
2. Qualitative classification
Classification made according to some attribute or quality such as sex, literacy, religion
etc.
Other classification
a. Temporary classification
Classification with respect to time is called temporary classification.
b. Spatial classification
When classification is made with respect to places e.g production of sugarcane in Kenya
region wise or production of tea in different areas.
Classification may be further divided as under;
a. Simplified classification
b. Manifold classification
Simplified classification is that type of classification in which data are divided into two classes
e.g persons may be divided into males and females or literate and illiterates.
33
Data is divided into or according to one criterion.
Manifold classification is that type of classification when data is divided into a number of
classes and sub-classes. E.g the population of Kenya may be classified as males and females.
Again, males and females may be further sub-divided into male literate and female literates.
These may be further sub-divided according to their professions. In this case, data are classified
according to various criteria.
Main objectives of data classification
- To eliminate unnecessary details.
- To bring out clearly points of similarities and dissimilarities.
- To enable one to intake comparisons and draw inferences.
- To reflect the important aspects of the data.
- To utilize the data for further statistical analysis.
Organisation of data
The process of organization the raw data into a compact and readily comprehensible form known
as data summary.
Data summary involves the preparation of two kinds of tables i.e
a. Frequency distribution table
b. Contingency tables (age and sex) - relationship between two or more
variables.
Before these tables can be prepared, the raw data needs to be sorted using any of the simplest
methods of sorting data such as: a) hand tally b) hand sorting c) sorting coded information (use
of computer programming).
Tabulation of data
- The systematic arrangement of the statistical data in columns and raws.
- The main advantages of the tabulation is that a mass of data which is confusing is
presented in a logical sequence giving the shape of statistical tables which answers all the
questions of the problem under investigation.
- It’s the process between the collection of data and its final analysis. It prepares the
ground for the analysis and interpretation of data.
- It’s the process in which information is extracted from either the questionnaire, interview
h bschedule or classified data and entering it on separate summary sheet.
Purpose
- It simplifies the details from different classifications so that salient features on the data
are obtained/seen.
- It facilitates the interpretation of the assembled data.
34
Advantages of tabulation
- Tabulated data can be understood easily as compared to data given in narrative form.
- The comparisons between different classes of data can be made easily.
- The required data can easily be located.
- The unnecessary details are avoided.
- Tabulated data takes less space.
Principles of table construction
The format of a table depends upon the nature of data and purpose of constructing a table. A
table should be constructed in such a way that it achieves its purpose in a best manner.
35
Recruited during the year 10 3 13 16 1 17
Left during the year (13) (4) (17) (10) - (10)
Total 76 10 86 82 11 93
During 2007, wastage declined by 3 among men compared with 2006 and no women left, 6 more
men but fewer women were recruited than in the previous year. The total number employed in 1 st
year 2008 amounted to 93.
Arrange the above information in concise tabular form showing all relevant totals and sub-totals.
Example 2
The following report was prepared by an examination officer on the performance of health
records students in Nairobi KMTC. Out of 3,500 male candidates below 20 years of age, 500
passed and 3000 failed.
Of the 1000 male candidates 20 years old and over 200 passed and 800 failed. As regards the
female candidates, out of 500 below 20 years of age 100 passed and 400 failed. Of the 340
females 20 years old and over, 80 passed and 260 failed.
Present the above information in a tabular form.
36
In the year 2003, there were 50 000, 30 000 and 25 000 patients treated in medical, surgical and
pediatrics clinic respectively.
In the year 2004, there were 80 000, 60 000 and 50 000 patients treated in the medical, surgical
and pediatrics clinic respectively.
Clinic attendances
YEAR Medical Surgical Peadiatrics Total
2001 61 000 40 000 5 000 106 000
2002 75 000 50 000 30 000 155 000
2003 50 000 30 000 25 000 105 000
2004 80 000 60 000 50 000 190 000
Total 266 000 180 000 110 000 556 000
Source: health records department
Array presentation
This is the arrangement of a group of data in a pleasing way so that they are in order.
When you are conducting an investigation you will reach a stage when you will have
several figures referring to different classifications.
When these figures are written down the way they were obtained from the original
classification, they are referred to as an array.
Frequency distribution
37
When the statistical data is grouped according to size or magnitude, the series formed is
known as frequency distribution which generally consists of class interval and their
corresponding frequencies.
Number of classes
The number of classes in a frequency distribution depends upon the number of items of a
series.
A frequency distribution should not have less than 6-8 classes and should not have more
than 20-25 classes
Class interval
To find out the class interval deduct the minimum value from the maximum value (range)
and divide by the
Number of classes or groups.
The class intervals of different classes should be equal. If they are not equal then these
will give misleading results.
Example
Suppose the maximum value is 100 and minimum value is 20 and desired number of
classes is 8. Calculate the class interval?
40 22 50 61 30 58 51 75
58 70 93 59 49 55 63 38
87 57 23 41 60 57 52 77
37 62 53 83 48 73 28 31
21 76 32 57 53 25 42 63
95 54 64 39 82 54 33 45
48 22 53 65 26 65 87 43
51 66 34 78 55 44 27 74
89 46 67 45 30 57 97 81
43 28 99 47 79 56 68 35
Steps
38
1. Obtain the range, range is a measure of dispersion or spread and is obtained by
subtracting the lowest (minimum) figure from the highest (maximum) figure in the
distribution.
1- 21-30
2- 31-40
3- 41-50
4- 51-60
5- 61-70
39
6- 71-80
7- 81-90
8- 91-100
5 Classes are not supposed to overlap each other i.e
One number should not appear in two or more classes
Example
For discrete data
20-30
30-40
40-50
50-60
This is wrong because it does not indicate clear boundaries of the classes. Should you wish to
have such groupings you should indicate very clearly the boundaries of upper class limits for
instance
20 but under 30
30 but under 40
40 but under 50
Types of frequency distribution
There are two types of frequency distribution:
1. Discrete series
In discrete series the various units are capable of exact measurement and each unit of the
data is separate and complete and definite breaks are visible between different units.
e.g we can count the number of students whose scores in statistics is exactly 80%.
40
2. Continuous series
In this series the statistical units are arranged in groups or classes because they are not
exactly measurable and are only approximations.
Example
Marks No of students
0-20 25
20-40 35
40-60 45
60-80 26
80-100 16
Continuous variables are measured
Example
2, 4 3 ,1 5 ,7 ,9 ,21 ,13 ,15 , 18, 17, 14, 10, 12, 16, 7 ,6, 19, 22, 11, 23, 22, 24, 2, 5, 3, 4, 3, 2.
Group the data taking the class interval as 5 in ;
1 Inclusive form of grouping
2 Exclusive form of grouping
TALLY SHEETS
INCLUSIVE FORM
Both the units are included in the class while taking the items in a group.
In first group (1-5), both 1 and 5 values will be included and in the second group (6-10)
both values will be included and so on.
41
EXCLUSIVE FORM
The lower limit is included but the upper limit is excluded while taking the items in a
group.
In the first group ( 0-5), zero or more than zero but below 5 will be included and in the
second group (5-10), 5 or more than 5 but below 10 will be included and so on.
Example
From the following observations, prepare a frequency distribution using both inclusive
and exclusive method, starting with 5 -10 and 0-50
12 36 40 30 28 20 19 10 10 16
19 27 15 26 20 19 7 45 33 21
26 37 6 20 11 17 37 30 20 5
Tally Frequency
0 but under 5 9
5 but under 10 6
10 but under 15 5
15 but under 20 5
20 but under 25 5
PRESENTATION OF DATA
Diagrams
The main objective of statistics is to simplify the complexity of the quantitative data and make
them easily understandable.
Diagrammatic representation is best suited to spatial series (in relation to geographical locations)
and data split into different categories whenever a comparison of the same data at different
places is to be made, diagrams will be the best way to do that.
42
Advantages of diagrams
- They provide an easy and attractive means of representing data.
- They make the information contained in data readily understandable.
- They facilitate comparisons.
- They save time and labour.
- They give an effective impression.
- They have great memorizing value as compared to mere figures.
Limitations of diagrams
- Diagrams do not give accurate result but rough ideas.
- A technical man can construct a diagram so a common man cannot do it correctly.
- Comparisons of diagrams cannot be made of the unit not common or the phenomena is
not the same.
- They can be misused very easily.
- This method of data presentation is very expensive.
Principles of diagram construction
The diagrams should be neat and clean so that they may have an attractive impression on the
mind of the reader.
A brief heading on the top of the diagram should be given so that the reader may create an idea
about the diagram before he studies it. The relative data should be given near diagram so it can
give a correct view.
The scale to be used should be suitable and mentioned on the right hand top or left hand bottom.
All types of symbols used should be explained.
Types of diagrams
1. One dimensional diagrams e.g. bar diagrams
2. Two dimensional diagrams e.g. rectangles, squares and circles.
3. Pictograms and maps.
Bar charts
This is a means of presenting information visually by drawing bars that represent specific data
frequency. All charts have common principles of constructing.
General principles of bar charts construction:
a. It should have a clear title.
b. Scale used should be indicated clearly.
c. Show the attribute (quality) or variable (quantity) on the horizontal axis (x-axis).
d. Frequencies should be appear in the vertical axis (Y-axis).
e. The height of the bars should be proportional to the frequencies.
43
f. The source of information should be indicated normally as footnotes.
g. The axis should be clearly labelled either X or Y axis.
h. The Y axis should be ¾ of the length of X-axis.
Simple bar charts
This type of charts are used to present data which has been presented in a simple table. The
height of each bar indicates the size of the figure represented.
The width of the bars is not taken into account and it should be uniform for all bars.
The number of bars depend on the number of figures.
Medical clinic attendances.
YEAR Pts
2001 61000
2002 75000
2003 50000
2004 80000
Total 266000
Source: out-pt attendance register.
N/B
- The height of the bars shows the frequency (patients of each year)
- Bars are distinctly drawn are clearly separate from one another.
- It’s easy to notice the year which had the highest number of attendance (2004) and the
one with the lowest (2003).
- The pattern of the bars can also be seen.
- Downward slope of frequencies was realized in 2003.
- Gradual downward slope.
- Upward slope of frequencies was realized in2004.
- Sharp rise.
Multiple component bar chart.
- The component figures are shown as separate bar charts adjoining each other.
- The height of each bar represents the actual value of this component figure.
- They are useful if we have two or more sets of comparable data and wish to compare and
contrast them.
CLINIC ATTENDANCE
YEAR MEDICAL SURGICAL PEDIATRICS TOTAL
44
2001 61000 40000 5000 106000
2002 75000 50000 30000 155000
2003 50000 30000 25000 105000
2004 80000 60000 50000 190000
TOTAL 266000 180000 110000 556000
600000
500000
400000
clinic attendance
300000 MEDICAL
SURGICAL
200000 PEDIATRICS
TOTAL
100000
0
2001 2002 2003 2004 TOTAL
years
Multiple chart
This is a bar chart similar to a simple bar chart but has separate bars which are independent for
each component in the same period (year).
Component/Sectional bar chart
This is also called subdivided bar diagrams. Instead of placing the bars for each component side
by side we may place these one on top of the other.
Each bar is subdivided into two or more components across the years.
45
1200000
1000000
800000
clinic attendance
600000 TOTAL
PEDIATRICS
400000 SURGICAL
MEDICAL
200000
0
2001 2002 2003 2004 TOTAL
years
46
Advantages of bar charts.
1. They are easy to construct.
2. They can depict data more accurately.
3. They can be used to indicate the sizes of component figures.
Disadvantages.
1. They are not more informative.
2. They are restricted to three or four component figures only.
Pie charts
A pie-chart is a circle divided by radial lines into sections so that the area of each section is
proportional to the size of the figure represented.
A pie chart is particularly useful where it’s desired to show the relative proportions of the figures
that are obtained to make up a single overall total.
In order to construct a sub-divided circle, first find out the various angles which represent the
various sectors by the following formula;
Angle of each component is equal to component divided by total components times 360 degrees.
Example
In 2004, medical clinic had 80000 patients, surgical had 60000 patients and pediatrics had 50000
patients.
Present the information of a pie chart.
47
Patients attendance
pediatrics
26%
medical clinic
42%
surgical
32%
48
2006 1000
2007 600
Source: production department.
Features
- Information can reach a wider audience.
- Difficult to produce the symbols exactly.
- Part symbols are difficult to interpret.
- Often difficult shaped symbols give different impressions of size and are thus misleading.
Examples
- The statistics of housing stock are represented by a house.
- Statistics about bread could be represented by a sliced loaf.
- Budget statistics could be represented by a pile of coins etc.
Time series graphs
- In time series, values of a variable are given at different periods of time.
- When a graph of such a series is drawn it would give changes in the value of a variable
with the passage of time.
- The graphical presentation of such a series is called histogram.
The main aim of drawing such graphs is to have comparison to study the following;
1. Changes in one variable over a period of time.
2. Changes of two or more variables over a period of time.
Factors influencing time series.
- Whether patterns/seasonal variations.
- Cyctical variations (ups and downs) e.g inflation.
- Trends.
While constructing a histogram, time is taken along X-axis and the values along Y-axis, then the
data is plotted and points are joined by means of straight lines to get a histogram.
49
Main examples of time series.
a. Population of a country over a specific period of time.
b. Sales of a business enterprise over a period of one year.
c. Prices of some specific commodities over a period of time.
d. Temperature over a period of time.
Types of time series graphs
1. Time series graphs or histograms
2. Z- charts.
3. Lorenze curves
4. Graphs of frequency distribution
Example 1
Monthly sales of EPZ stores for the year 2007 were as follows;
Month Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec
Sales 50 40 60 70 50 80 100 90 110 80 70 120
(shs)
120
100
sales (shillings)
80
60
Series2
Series1
40
20
0
Mon Jan Feb Mar Apr May Jun Jul Aug Sep Oct Nov Dec
th
year 2007
Example 2
The following table gives the sales of a certain firm in 6 years. Draw a graph of time series.
50
YEAR 2001 2002 2003 2004 2005 2006
SALES 820 950 1000 950 900 1050
(SHS)
Z- charts
A Z-chart is simply a time series chart incorporating three curves for:
1. Individual monthly figures.
2. Monthly cumulative figures for the year.
3. A moving annual total.
Z- chart takes its name from the fact that the three curves together tend to look like the letter Z.
A Z-chart is of great importance for presenting business data over a period of one year.
It includes the following information:
a. Monthly totals
These simply show the monthly results at a glance together with any rising or falling trends and
seasonal variations.
b. Cumulative totals
These show the performance to date and can be easily compared with planned or budgeted
performance.
c. Annual moving totals
51
These show a comparison of the current levels of performance with those of the previous year. If
the line is rising then this year’s monthly results are better than the results of the corresponding
month last year and vice versa.
Sometimes separate vertical scales are used to plot the monthly data and the data for cumulative
and moving totals.
In some cases, the same vertical scale is used to plot the monthly data and the data for the
cumulative and the moving annual totals.
The decision to take same vertical scale or separate vertical scales should be made in view of the
nature of the given data.
Moving totals
Direct comparison of the figures for one month with the same month of the previous year is
possible. But this sort of comparison does not allow an overall trend to be easily observed.
A better solution is to use a moving value.
Examination of the hospitals cost sharing collection shows that the patient attendance is seasonal.
(for example figures in the following table gives a false impression). This problem of seasonal
influence frequently arises in time series and an excellent method of eliminating such influences
to add together 12 consecutive months.
If we add the 12 months immediately preceding the end of each and every month in the table, we
shall obtain a series of totals, one for each month and each total will be the total for the year
immediately preceding the end of that month. Such series is called a moving total or more
specifically since the totals are yearly totals it’s referred to as a moving annual total.
52
Constructing of a Z-chart
1. A double scale is used on the vertical axis.
N/B: for the monthly figures
For the MAT and cumulative figures.
2. The points where each curve starts must be clearly marked. These are:
a. Monthly figures for December of the previous year.
b. Cumulative figures at zero.
c. MAT at the MAT figures for the December of the previous year.
3. Since a MAT is the total of the 12 immediately preceding months the MAT for the final
month must be the same as the cumulative total. The two curves will therefore meet at the
last month of the graph.
The following are the sales of EPZ Ltd for the year 2005 and 2006.
Months 2005 2006 Monthly cumulative for MAT
2006
January 400 420 420 7200
February 480 450 870 7170
March 420 600 1470 7350
April 580 640 2110 7410
May 600 580 2690 7390
June 800 700 3390 7290
July 750 800 4190 7340
August 600 750 4940 7490
September 550 600 5540 7540
October 500 480 6020 7520
November 600 550 6570 7470
December 900 950 7520 7520
7180 7520
Monthly cumulative totals are obtained as follows:
Feb = 420+450= 870
Mar = 870+ 600= 1470
Apr = 1470 + 640 =2110 and so on.
Moving annual totals are obtained as follows.
January = total sales for the year + sales for January 2006 – sales for January 2005
= 7180 + 420 - 400 = 7200
February = previous month MAT + sales for February 2006- sales for 2005
= 7200 + 450 – 480 = 7170
March = 7170 + 600 + 420 = 7350 and so on.
53
Draw graph.
Lorenz curve
This is a graph to measure dispersion. It measures inequalities of wealth distribution.
Important use of the Lorenz curve is in the measurement of the extent to which income is
unevenly distributed between the various income groups.
Examples of the uses of such curve.
- To show the degree of concentration of a particular industry in the hands of a few firms.
- The extent to which the income of various groups of people vary.
- The extent to which the wealth of profit of various firms vary.
Steps
1. Write down the values of the two variables being plotted.
2. Express the variables as % of the total.
3. Compute the cumulative % of each variable.
4. Draw a horizontal and vertical axis and plot 0% to 100% on each axis.
5. Mark the cumulative % on the graph and join the points together by a free hand curve.
This is Lorenz curve.
6. Draw the line of equal distribution by joining 0% to 100% point by a straight line or by
dividing the graph on 45 degrees.
If the Lorenz curve is a way from the line of equal distribution, there is greater disparity or
inequality and vice versa.
Example
The following figures are taken from the survey on ‘business prospects’ for 2007.
Maize flour sales.
Number of establishments Net output ()
23 104
26 450
24 860
19 1350
14 2190
6 3125
Draw a Lorenz curve using the above data.
Solution
54
Number of % cumulative Net output % Cumulative
establishments
23 20.5 20.5 104 1.3 1.3
26 23.2 43.7 450 5.6 6.9
24 21.4 65.1 860 10.6 17.5
19 17.0 82.1 1350 16.7 34.2
14 12.5 94.6 2190 27.1 61.3
6 5.4 100 3125 38.7 100.0
112 100 8079 100.0
Draw graph
Number of establishments
(Cumulative %)
This curve shows the greater disparity between the number of establishments and the net output.
20.5% establishments have only 1.3% net output and 5.4% establishments have 38.7% share of
net output.
Graphs of frequency distributions
The graphs of frequency distribution of continuous type are as follows:
a. Dgive curve
b. Histogram
c. Frequency polygram
d. Frequency curve.
To understand how to draw the above graphs its important to note the following terms:
a. Class interval
A set of classes that are used to define the raw data or size of the group chosen.
It can be determined by finding the range of the raw data obtained and dividing it by the value of
number of classes we desire to have.
b. Class limit
These are the end number of class interval. The lower value for class interval is called lower
class limit while the upper value for class interval is called upper class limit.
c. Class boundaries
Class boundaries are easily gotten by subtracting 0.5 from the lower class limit or lower value of
class interval and adding 0.5 to upper class limit or upper value of class interval e.g 20-29 will be
19.5-29.5.
55
d. Class width
This is the size of the class interval and it’s obtained by subtracting lower class boundaries from
upper class boundaries e.g 19.5-29.5 the class width will be 29.5-19.5=10
e. Class mark
This is the mid-point or value of the class interval. It can be derived by adding lower and upper
class limit and dividing by 2 or adding lower and upper class boundaries and dividing the sum by
2.
Histogram
Its graph in which class boundaries or class interval is marked on the horizontal axis and the
corresponding class frequency on the vertical axis.
It assumes continuous data.
Steps
- The ratios between the areas of the rectangles must be equal to the ratios between
frequencies.
- The bases of the rectangles must be equal to the width of the corresponding class
intervals.
- The first bar must start from the lower true (real limit) and end in the upper time (upper
real limit)
- If the class width are equal, the heights should be preferably equal to the class
frequencies.
e.g
Class boundaries F
0.5 – 5.5 6
5.5 – 10.5 8
10.5 – 15.5 5
15.5 – 20.5 4
20.5 – 25.5 2
Draw graph
56
50 -60 15
60 -70 13
70 -80 10
80 -90 5
90 -100 2
Draw graph
Frequency polygon
- Are very useful if we want to compare two distributions.
- This is a graph that joins the mid-point of the tops of column of the histogram through
straight lines.
- To draw a frequency polygon, there is need to obtain class mark (mid-point) for each
class boundary against the corresponding frequency.
- The points to draw frequency polygon will be joined with the help of a ruler. Sometimes
its used to derive mode.
Use the example above to draw a frequency polygon
Draw a histogram and superimpose frequency polygon from the above data.
Frequency polygon curve
- It can be drawn on the same lines as frequency polygon but the mid-point at the height of
class interval rectangles will be joined smoothly. The difference between frequency
polygon and frequency curve is joining the points.
57
d. Plot the cumulative frequencies on the graph at the upper class limits of the classes to
which they refer.
e. Then join all these points by the help of a curve.
N/B An ogive curve is used to find out the values of median, quartiles, deciles and percentiles
graphically.
There are two types of ogives
a. Less than
b. More than.
Less than
Its usually of ‘S’ shape
Interval Cumulative
frequency
0-10 2 2
10-20 8 10
20-30 12 22
30-40 18 40
40-50 28 68
50-60 22 90
60-70 6 96
70-80 4 100
- Plot the points with coordinates having x-coordinates as actual limit and y-coordinate as
the cumulative frequency.
Draw graph
More than
Marks frequency C.F of more than
More than 1 3 60
More than 11 8 57
More than 21 12 49
More than 31 14 37
More than 41 10 23
More than 51 6 13
More than 61 5 7
More than 71 2 2
58
- Plot the points with coordinates having x-coordinates as actual lower limits and y-
coordinates as cumulative frequencies.
Marks Frequency Cumulative Class boundary
frequency
1-10 3 60 0.5
11-20 8 57 10.5
21-30 15 49 20.5
31-40 14 37 30.5
41-50 10 23 40.5
51-60 6 13 50.5
61-70 5 7 60.5
71-80 2 2 70.5
Draw graph
59
The formal process whereby a person is accepted by a hospital for the process of Rx as an in-
patient.
If an in-patient is formally discharged from the hospital and then returns for further Rx, the
admission process is repeated and a second admission is recorded in the statistics.
Live births, in the hospital are cons.
Collection of health data
Identification of health problems in a community
- Questionnaire
- Interviews
- Observation
- Health records
- News e.g radios, newspapers etc
- Attending chief barazas.
Collection of health data
Hospital administration statistics
- Health administrative data is generated from health institutions where health services are
provided to the patients or clients.
- The information obtained indicates a detailed account of the activities that are carried out
in a health institution.
- It forms a data base which is used by health managers/administrators in the Mx and
planning of health services.
Examples
1. Mx of health services
Managers of health services require data concerning patients in order to manage them more
effectively and efficiently. This include the mx of scarce resources such as patients food, linen,
transport, staff, finances available, physical facilities such as beds and theatre etc
Data concerning family planning and maternal child health services (MCH/FP) are also required
for evaluation and better mx for such services. The requirements of doctors and other health
workers needs statistical information concerning diagnosis, operations, Rx and investigations.
Adequate information would improve the efficiency of the process involved in arriving at the
diagnosis and subsequent Rx, thereby making it possible to eliminate wasteful process.
2. Planners
60
Health information aids planners in making accurate plans for day to day and future requirements
of health activities with accurate and timely data, planners are able to estimate for service
requirements such as drugs, finances and physical facility requirements.
The health information data bank enables health workers and information officers to create an
analytical frame work commonly referred to as Health Activity Analysis.
3. Health Activity Analysis
This is a systematic way of collection, analysis and presentation of health data. It concerns data
obtained from patients/clients attendances, staff, bed use, physical facilities, service delivery
data, finances etc.
4. Patient/client attendance
Patients/clients make use of health facilities either as in-patient or out-patient clients.
61
- All patients admitted from home, transferred from another ward, coming from parole or
returned from having absconded are indicated on admission section.
- All patients indicated on this section are added to patients who remained in the ward the
previous day and the total is obtained which is the total number of patients in that ward.
Discharges
- These are discharges from the wards which are discharged home, those who are dead,
transferred to another ward, absconders (run away from hospitals) and those who are
granted hospital leaves (parole). The total discharges are subtracted from the totals
obtained in admission section.
- The results obtained must be equal to the physical count of patients who are in the ward
at the time of physical count.
- This is the figure entered in section E of the DBR in the column of patients who are either
sleeping on the bed or cots (accommodated).
- Section F of the DBR is normally for records use only. The officer in charge of health
information completes this section.
- This section is used to balance the DBR to check the accuracy of entries in all other
sections. The last total that’s obtained after adding all the admissions to the previous day,
number of patients who remained in the ward and subtracting all discharges must be
equal to patients indicated in section E that is today’s daily return numbers of patients.
- When DBR is returned to the records department is properly checked, summarized on
daily summary in patients statistics form.
- The daily summary form, summarizes the day statistics. It indicates the total admissions
and discharges in the ward.
- There is a column for OBD (occupied Bed Days), this is the total number of patients
remaining in the ward each day.
63
Analytical Formulae
To be able to obtain the measures explained above we need to apply the following:
a. Available bed days = available beds times days in period
ABD = AB times DIP (month, year)
b. Available beds = total authorized beds in a hospital (ward) during a specified period or
available bed days divided by days in period
ABD = ABD divided by DIP
c. Occupied bed days = total patients remaining each day or in-patient days added together
for the days in the period or % cc times available bed days divided by one hundred.
OBD = % CC times ABD divided by 100
d. % CC = OBD divided by ABD times 100
e. Average number of occupied beds (average number of patients)
OB = OBD divided by DIP
f. Vacant bed days (excess patient days)
VBD = ABD – OBD
g. ABD = VBD PLUS OBD
h. OBD = ABD –VBD
i. Average length of stay
ALOS = OBD divided by D&D (Discharge & deaths)
j. OBD = ALOS times D&D
k. D & D = OBD divided by ALOS
l. Turnover per bed/discharges & death per available bed/through put per bed
TOB=D&D divided by AB
m. D&D = TOB times AB
n. AB = D&D divided by TOB
o. Turn over interval (TOI) = VBD divided by D&D
p. VBD = TOI times D&D
q. D&D = VBD divided by TOI
Example
a. In 2005 a hospital recorded the following data.
In-patient hospital days = 10,676
Number of allocated beds = 25
(throughout the year)
Discharges & deaths = 124
Calculate:
1. Excess in-patient days
2. The extra number of beds required to cater for excess patient days through the year.
64
3. % occupation
4. ALOS
5. Average daily population/average number of patients.
6. TOI
7. TOB
b. A word had 25 beds available throughout the year in 2012 and 650 patients were
discharged during this period and a total of 8125 OBD.
Calculate:
- Average daily bed occupancy
OBD divided by DIP
8125/364
- % CC
- D&D per available bed
= D&D divided by 25
= 650/25
The following data was extracted from the DBR of a hospital which has 50 beds and 10 cots in
the month of October, November, and December 2015. 427 patients discharged and 50 died.
October
41 41 44 46 48 47
47 47 49 48 49 43
41 46 41 43 46 41 (1359)
40 43 43 40 47 423
42 44 41 423 40 40
November
46 47 48 40 41 43
48 46 41 43 42 42
43 44 40 44 48 46 (1323)
44 42 43 46 40 49
49 41 40 48 41 48
65
December
41 43 44 45 46 45
43 45 43 41 41 47
44 47 49 42 43 48 (1375)
48 46 46 45 44 40
42 49 45 40 49 41
OBD – 4047
Calculate
a. OBD
b. % occupancy
c. ALOS
d. Average daily occupancy
e. TOI
f. TOB
g. D&D per bed
h. VBD
i. Interpret the above results from manager’s view
Level of measurements
Lowest : - nominal – naming system, categorical data.
- ordinal – categorical data arranged in order.
- interval – data which lies in an arbitrary scale e.g temperature
Highest: - ratio – a multiple or comparison between the groups of data.
Evaluation techniques
a. Program techniques
It’s a way of measuring whether programs are doing what they are supposed to do.
Evaluation means collecting information about program activity effectiveness and then
comparing it to the actual results.
Evaluation helps communities to see if their programs are achieving the expected results.
Reasons for evaluation of a health program
- Evaluation helps communities understand how the health programs are working.
- Evaluation provides a platform to show the success or challenges of a given program.
66
- Evaluation shows what works in another community that can be adopted by a different
community.
- Evaluation will show us how staffs in a given program are working.
Community needs assessment
- Needs assessment are part of the planning stage of programs and help in identifying
community health priorities.
- Needs assessment helps the community to understand the following:
a. What kind of health problems communities are experiencing.
b. What causes this health problems.
c. What resources are available to address this health problems.
d. What goals and objectives community need to write or formulate to help out this
problems.
e. Which community members have the most urgent needs and how best to meet the
needs of the community members.
Terminologies used in evaluation
a. Goal
These are broad statements that describe what programs or activities should be achieved.
b. Objective
These are identifiable and measurable actions to be completed in a specific time.
c. Activity
It’s something that communities do to meet program objectives. They are building blocks that
make up a program.
d. Indicator
Evaluation indicators are signs, events or statistics that measure the success of programs or
activities in meeting the objective.
1. Hard or soft indicators
Hard indicators based on numbers and it’s quantitative.
Soft indicators not based on numbers and it’s qualitative.
2. Short term or long term indicators
Short term- they appear within a few weeks or months after activity starts and show progress
towards meeting the objective.
Long term- occur many months or years to show progress towards meeting the objectives.
Main evaluation tasks.
67
a. Collection of data.
b. Analyzing the evaluation data.
Examining findings enables evaluators to answer evaluation questions explaining the conclusion.
This analysis explains why the activity was successful or not.
Types of evaluation
1. Traditional evaluation
Traditional evaluation in health were focused on assessing the impact of specific program
activities of defined programs.
2. Economic evaluation
Types:
- Cost minimization
- Cost effectiveness analysis
- Cost benefit analysis
- Cost utility analysis
They combine program effectiveness information with economic resources that is costs and
benefits in quantitative terms.
They allow decision makers to prioritize health activities in the face of finite or scarce resources.
3. Process evaluation
It refers to evaluation that are focused on outputs.
A relasp is assumed between outputs and outcomes and evidence of a change in output is taken
as indirect evidence of an impact on the desired outcome.
4. Formative evaluation
Refer to efforts to identify the best use of the available resources prior to the traditional program
evaluation.
5. Empowerment evaluation
Involves an approach whereby programs takes stock of their existing strength and weaknesses,
focus on key goals and program improvement, develop self-initiated strategies to achieve this
goals and determine the type of evidence that can document credible progress.
6. Performance measures
Use of statistical methods and other evaluation tools on an going basis to assure accountability
for public health programs and to improve performance.
Evaluation circle
68
Evaluation techniques or methods in MCH
- Evaluation of MCH program should be an undergoing process conducted throughout the
various stages of an intervention, starting with the project design and ending with an
assessment of final outcome.
- A well designed evaluative strategy involves the following steps:
a. A formative evaluation
During the project developmental phase to clarify objectives while taking into account the
cultural environment and other local factors that influence project execution.
Collection of baseline data and existing historical trends. Found out where we are.
b. Process evaluation
Done throughout project implementation phase to provide timely feedback on how the
intervention has been put into operation and what could be done to improve its operationality to
achieve the desired outcomes.
c. Impact evaluation
Assess the net effects/total impact of the intervention and whether the stated intervention goals
are reached.
Assess both short term outcomes and long term system impacts.
Its usually very useful as it informs donor funding organization and policy decision makers
whether to invest in or scale up the particular program.
Maternal child health indicators
1. Impact indicators
- Maternal mortality rate or ratio
- Under 5 child mortality rates.
- Starting prevalence
2. Coverage indicators
- Ante-natal care
- ARV for HIV positive pregnant women
- Skilled attendance at birth
- EBF
- Post-natal care for mothers and babies within 2 days of birth.
69
- Immunization.
- Antibiotic Rx of pneumonia.
Methods of evaluating impact of family planning methods.
1. Demographic and statistics techniques
- Standardization and decomposition
- Trend analysis
2. Analysis of acceptance data
- Couple year of protection.
- Analysis of the reproduction process.
- Component projection.
- Simulation.
3. Experimental designs
- Random experimental designs
- Quasi experimental designs
- Matching studies
4. Multi-valid analysis
- Sub-divide within a country e.g Kisii and Kisumu county
- Country
70
- It should be rigidly defined.
- It should be based on all values.
- It should be easily understood and calculated.
- It should be least affected by the functuations of sampling.
- It should be capable of further algebraic or statistical Rx.
- It should be least affected by extreme values.
Types of average
- Arithmetic mean or simple average
- Median
- Mode
- Geometric mean
- Harmonic mean
They are also called measures of central tendency, measures of location and measures of
description.
Arithmetic mean
- Its also called mean or simple average.
- As obtained by summing up the values of all the items of a series and dividing this sum
by the number of items.
X = X1 + X2 + X3 + X4…………….Xn =
n
X = mean or A.M n = number of items
X1, X2, X3 ………………….Xn = value of items.
(greek letter called sigma) = sum of all items.
Computation of A.M
- Can be calculated by:
a. Direct method
b. Short cut method
Ungrouped data
a. Direct method
Example 1.
Kisii level 5 has got the following cases over a period of 5 months: 75,55,48,72,60
Find the average of malaria cases
71
Total cases=75 + 55 + 48 + 72 + 60 = 310
Number of months = 5
X = sigma X = 310 = 62
n 5
b. short-cut method
- A specific value from the given values is assumed as mean and its known as provisional
mean (P.M)
- The differences between values of various items of the series and its provisional mean are
known as the deviations.
X = P.M + sigmaDx
n
where P.M = provisional mean
Dx = deviations from P.M
Sigma Dx = the sum of deviations from P.M.
This method is applied when the data is too large and complicated.
Example 2
The number of patients attended Kenyatta National Hospital over a period of 10 months are:
1000, 1200, 1300, 1100, 1090, 1010, 1500, 1900, 1700, 2000.
Calculate the average number of attendances or A.M. by:
a. Direct method
b. Shortcut method
- Direct = sigma X = 13,800 = 1380
n 10
- Short cut method
72
1500 1500-1500 0
1900 1900-1500 400
1700 1700-1500 200
2000 200-1500 500
Sigma X =13800 Sigma Dx=-1200
Grouped data.
Direct method
X = sigmafx
n
where
X = arithmetic mean
f = frequency
n = total of frequencies
sigma = sum of all items
short cut method
X= P.M/A.M + sigma fDx
N
Direct method
73
Values (x) Frequency (f) Product (fx)
5 20 100
10 43 430
15 75 1125
20 67 1340
25 72 1800
30 45 1350
35 39 1365
40 9 360
45 8 360
50 6 300
Total Sigma f =384 Sigma fx = 8530
X= P.M/A.M + sigmafdx
n
= 25 + -1070
384
= 25 - 1070
384
= 25-2.786
=22.2
74
Direct method
Marks M.P. (x) F fx
0-20 10 5 50
20-40 30 7 210
40-60 50 13 650
60-80 70 8 560
80-100 90 7 630
Sigmaf = 40 Sigmafx=2100
P.M/A.M =50
X= P.M/A.M + sigma fdx
n
= 50+ 100
40
50 + 2.5
52.5
Ungrouped data changing to grouped data.
Advantages of A.M
- It can be easily understood.
- It takes into account all the items of the series.
75
- Its not necessary to arrange the data first and then calculate A.M.
- Its capable of algebraic Rx.
- Its good method of comparison.
- Its not indefinite. Its determined.
- Its used frequently.
Disadvantages
- It’s affected by extreme values to a great extent.
- It may be a figure which does not exist in a series.
- It cannot be calculated if all the items of a series are not known.
- It cannot be used in case of qualitative data.
Median
This is the value of middle item of a series when these items are arranged in ascending or
descending order.
The formula for finding median item is n+1 divided by 2, where n=number of items. When the
half value is required the median is more suitable.
Computation of the median
When the number of items is odd, the formula will be n+1 divided of the item.
Example
In a hospital, there are 5 workers whose ages are 20, 15,19,21,17 years. Find the median age.
Solution
First arrange the data in ascending order
S/N VALUES
1 15
2 17
3 19
4 20
5 21
76
3rd item = 19 years
When the data is in even number, we apply:
½ (n/2
Example
The marks of 6 students in a class are 80,70,75,85,60,80
Solution
First arrange the data in ascending order.
S/N VALUES
1 60
2 70
3 75
4 80
5 80
6 85
Median
Median = size of
78
Computing of median in a continuous series
In order to calculate the median of the continuous frequency distribution, there is one difficulty
i.e the value of the median lies in a class interval.
Median = L + i/f (m-c)
Where L= lower class boundary of the median group
i= class interval of the median group
f = frequency of the median group
m = the middle item i.e (n/2)th item or (n+1/2)th item.
c = cf of the group preceding the median group.
Example
Find the median
Marks Students Cumulative frequency
(x) (f) (cm.f)
0-10 2 2
10-20 18 20
20-30 30 50
30-40 45 95
40-50 35 130
50-60 20 150
60-70 6 156
70-80 3 159
M = (n+1/2)th item
= (159+1/2) th item
= 80th item.
79
The 80th item lies in 95, so the median group is 30-40 marks group.
Median = L + i/f (m-c)
= 30 + 10/45 (80-50)
= 36.67.
Obtaining median graphically
- A cumulative frequency curve is drawn along vertical axis.
- Draw a horizontal line from the median item along y-axis.
- Draw another vertical line from the point where this horizontal line intersects CF curve.
- The point where this vertical line touches x-axis mark that point.
- The distance of that point along x-axis from the origin is median.
Example
Find out the value of median graphically from the following data.
Class interval F cf
0-10 5 5
10-20 10 15
20-30 15 30
30-40 8 38
40-50 7 45
Median item = n+1/2
= 45+1/2
= 46/2
= 23rd item.
Draw line graph
Advantages of median
- It’s easy to calculate.
- It’s simple and is understood easily.
- It’s less affected by the values of extreme items.
- It can be calculated by inspection in some cases.
80
- It’s especially useful in the study of those phenomena which are of qualitative in nature.
Disadvantages
- It’s not a suitable representative of a series in most of the cases.
- It’s not suitable for further algebraic Rx.
- It’s not used frequently like for mean.
- It cannot be determined exactly in the case of a continuous series.
Quartiles, Deciles and Percentiles
- The method by which the median is determined can be further extended to divide a series
into more than two parts.
- It is usually the tendency to divide the series into 4, 10 or 100 parts.
- The value of the item which divide the series into 4 equal parts is called ‘quartiles’.
- The value that divides the series into 10 equal parts is called ‘deciles’.
- The value that divides the series into 100 equal parts is called ‘percentile’.
- The 2nd quartile, 5th decile and 50th percentile is the median.
- The value of the item dividing the 1 st half of a series into 2 equal parts is called 1 st
quartile or lower quartile and the value of the item dividing the latter half of a series into
2 equal parts is called 3rd quartile or upper quartile.
- The process of calculating the quartiles, deciles and percentiles is the same as of median.
Q1 (1st quartile) = (n/4)th item, Q3 (3rd quartile) = (3n/4)th item
D4 (4th decile) = (4n/10)th item, P65 (65th percentile) = (65n/100)th item.
Example
Calculate Q1, Q3, D7, P85 from the following data.
Earnings Employees
Shs F c.m.f
11 3 3
81
12 6 9
13 10 19
14 15 34
15 24 58
16 42 100
17 75 175
18 90 265
19 79 344
20 55 399
21 36 435
22 26 461
23 19 480
24 13 493
25 7 500
Q1 = (n/4)th item
= (500/4)th item = 125th item
The 125th item lies in cf of 175 so:
Q1 = 17 approximately.
Q3 = (3n/4)th item = (500*3/4) = 375th item = 20 approximately.
Q7 = (7n/10)th item = (500*7/10) = 350th item = 20 approximately.
P85 = (85n/100) = (500*85/100) = 425th item = 21 approximately.
Example 2
Marks Number of students c.f
0-10 2 2
10-20 7 9
20-30 21 30
30-40 25 55
40-50 30 85
50-60 35 120
60-70 28 148
70-80 12 160
Q1 = L + j/f (n/4 – c)
= n/4 = 160/4 = 40, 40 lies in cf 55, class = 30-40
L = lower class interval
j = class interval
82
f = frequency of quartile class
c = cf of previous class.
= 30 + 10/25 (40-30)
=30+10/25(10)
= 30 + 100/25 = 30+4 = 34
Q3 = L+ i/f (3n/4 – c)
= 3n/4 = 3*160/4 = 120, 50-60
50 + 10/35 (120-85)
= 60
D3 = 3n/10 = 3 * 160/10 = 48, 30-40
= L + i/f (3n/10 – c)
= 30+10/25(48-30) =37.2
P4o = (40n/100) = 40*160/100 =64, 40-50
= L+ i/f(40n/100 –c)
40+10/30 (64-55)
=43.
Graphic calculation of quartiles, deciles, and percentiles
Calculate the same way the median is calculated.
Example
Marks F c.f
0-10 2 2
10-20 7 9
20-30 21 30
30-40 25 55
40-50 30 85
50-60 35 120
60-70 28 148
70-80 12 160
84
Class interval F
0-10 2
10-20 7
20-30 11
30-40 6
40-50 4
85
30-40 6
40-50 4
Draw graph
Advantages
- It’s easy to understand.
- Extreme items do not affect its value.
- It possesses the merit of simplicity.
Disadvantages
- Its often not clearly defined.
- Exact location is often uncertain.
- Its unsuitable for further algebraic Rx.
- It does not take into account extreme values.
Geometric mean
Also known as geometric average.
It’s the nth root of the product of items of a series.
Its calculated with the help of logarithmic tables.
Ungrouped data.
G.M = Anti – log of sigma logx divided by n.
Grouped data
G.M = Anti- log of sigma flogx divided by n.
Where:
G.M = geometric mean.
Logx = log values of variables.
n = number of items.
f = frequencies
calculate the G.M of the following data:
130,135,140,145,146,148,149,150,157
Solution
X Log x
130 2.1139
86
135 2.1303
140 2.1461
145 2.1614
146 2.1614
148 2.1703
149 2.1732
150 2.1761
157 2.1959
Total 19.4316
G.M = Anti-log of sigma log x divided by n
=anti-log of 19.4316/9
= Anti-log of 92.159
= 144.2
Advantages of G.M
- It’s rigidly defined.
- It takes account all the items of a series and cannot be calculated even if a single item is
missing.
- It can be placed for algebraic Rx.
- Its not affected by fluctuations of sampling.
- It gives more weight to smaller items and less weight to big items, hence, it’s a better
measure than A.M.
Disadvantages
- Its not easy calculate.
- Its not easy to understand.
87
- If one value of the item in a series is zero, the G.M will be zero.
- The value of the G.M may not necessarily exist in the series.
- It does not give equal weight to every item.
Harmonic mean
It’s the reciprocal of the A.M of the reciprocals of the values of items in a given series.
Ungrouped data = H.M = n/sigma 1/x
Grouped data
H.M = n/sigma f 1/x
Where:
H.M= harmonic mean
n = number of items
f = frequencies
1/x = reciprocal of values
Calculate the H.M
1,2,4,5,8,10,10
Advantages of H.M
- It’s based on all the items of a series.
- It’s capable for further algebraic Rx.
- It’s not affected by the fluctuations of sampling.
Disadvantages
- It’s not easily understood.
- It’s a value which does not exist in a series.
- It’s not a good representative of a series.
Relationship between mean, mode and median
- In a symmetrical distribution the mean, mode and median coincide are the same.
- A distribution which is not symmetrical is either asymmetrical or moderate asymmetrical.
88
- In a highly asymmetrical distribution, its impossible to forecast the relationship between
the averages.
- In a moderately asymmetrical distribution, the median would be somewhere between the
mean and mode.
- Usually in such a distribution, the difference between the mean and the median is one
third of the difference between the mean and mode.
Therefore:
Median = mean - 1/3 (mean – mode)
Mode = mean -3 (mean – median)
Or
Mean = 3 median – 2 mode
Draw line graphs
Measures of dispersion/variation/spread
Dispersion
- It’s the extent of the scatteredness of items around a measure of central tendency.
- The degree to which numerical data tend to spread about an average value is called the
variation or dispersion of the data.
- A measure of dispersion indicates the extent to which the individual observes differ on
average from the mean or from any other measure of central tendency.
Significance of measuring dispersion
- To determine the reliability of an average.
- To serve as a basis for the control of the variability.
- To compare two or more series with regard to their variability.
- To facilitate the use of other statistical measures.
Properties
- It should be simple to understand.
- It should be easy to compute.
- It should be rigidly defined.
- It should be based on each and every item of the distribution.
- It should be amendable to further algebraic calculation.
- It should have sampling stability.
- It should not be unduly affected by extreme items.
Methods of measuring dispersion.
89
- Range
- Quartile deviation or inter- quartile range.
- Mean deviation or average deviation.
- Standard deviation.
- Lorens curve.
The first two are positional measures because they depend on the values at a particular position
in the distribution.
The other 2 i.e average deviation and the standard deviation are called calculation measures of
deviation because all of the values are employed in their calculation and the last one is a
graphical method.
Range
The difference between the smallest value and the largest value of a series.
Where the data are grouped into a frequency distribution, the range is equal to the difference
between the upper boundary of the highest class and the lower boundary of the lowest class.
Range (absolute range) = L – S
Co- efficient of range = L – S
L+S
Where L = largest value
S = smallest value
If the averages of two distributions are about the same, a comparison of the range indicates that
the distribution with the smallest range has less distribution and the average of that distribution is
more typical of the group.
Example
200, 210, 208, 160, 200, 250
Range = L – S
= 250 - 160 = 90
Co – efficient of range = L – S
L+S
= 250 – 160
250 + 160
= 90
90
140
= 0.22
Example 2
Class interval f m.d
10-20 8 15
20-30 10 25
30-40 12 35
40-50 8 45
50-60 4 55
Range = L – S
55 – 15
= 40
Co-efficient of range = L – S
L+S
= 40/70
= 0.57
Uses of range
- Quality control – this is to keep check on the quality of the product with 100%
inspection. If the range increases beyond a certain point, the production machine should
be examined to find out why the items produced have not followed their usual more
constant pattern.
- Fluctuations in the prices – range is useful in studying the variations in the prices of
stocks and shares and other commodities that are sensitive to price changes from one
period to another.
- Weather forecast – the meteorological department does make use of the range in
ascertaining, say the difference between the minimum and maximum temperature. This
information is vital to the general public because they as to within what limits the
temperature is likely to vary on a particular day.
Advantages of range
a. It’s the simplest understand and compute.
b. It takes minimum time to calculate the value of the range.
Disadvantages
91
a. It’s not based on each and every value of distribution.
b. It’s subject to fluctuations of considerable magnitude from sample to sample.
c. Cannot tell us anything about the character of the distribution within two extreme
observations.
d. It cannot be computed in case of open – end distributions.
Quartile deviation/ semi- interquartile range
- It’s based on the upper quartile (q3) and the lower quartile (q1).
- The difference between upper quartile and lower quartile is called interquartile range and
one – half of this range is the quartile deviation
Formula
Inter-quartile range = Q3 –Q1
Quartile deviation (QD) = Q3 – Q1 / 2
Co –efficient of QD = Q3 – Q1 / Q3 + Q1
Where Q3= upper quartile. Q1 = lower quartile
QD is an absolute measure of dispersion where as co-efficient of QD is a relative measure.
Where Q.D is very small it describes high uniformity or small variation of the central 50% of the
items and high Q.D means that the variation among the central item is large.
Example 3
Find out the value of quartile deviation and its co-efficient from the following data of marks
obtained in a class. 20, 28, 40, 12, 30, 15, 50.
Solution
Arrange the marks in ascending order.
12, 15, 20, 28, 30, 40, 50
Q1 = size of (n + 1 /4)th item = (7 + 1 / 4 ) = 2nd item
Q1 = 15
Advantages of Q.D
- It’s superior to range as a measures of dispersion.
- It has special utility in measuring variation in case of open and distributions or one in
which the data may be ranked but measured quantitatively.
- Its useful in badly skewed distribution.
- Its not affected at all by extreme observations.
92
Disadvantages
- It ignores 50% items.
- Its not capable for mathematical Rx.
- It’s value is very much affected by sampling fluctuations.
- It’s infact not a measure of dispersion as it really does not show the scatter and an
average but rather a distance on a scale i.e quartile deviations is not itself measured from
an average, but it’s a positional average.
Mean deviation/ Average deviation
This is the average of deviations of the items from Arithmetic mean, median or mode.
It’s usually calculated either from mean or median. All the deviations are taken positively despite
the negative sign.
Formula
Ungouped data
M.D = sigmadx / n
Where M.D= mean deviation
dx = deviations from mean
n = number of items.
If it’s calculated from median then:
M.D = sigmadm / n
Where dm = deviations from median
Grouped data
M.D = sigmafdx/n or M.D = sigmafdm/n
N/B:
If the distribution is symmetrical the average (mean or median) +- mean deviation is the range
that will include 57.5% of the observations in the series.
Co-efficient of M.D
Example
Calculate the M.D and its co-efficient.
Income (ksh. 3000, 4000, 4200, 4400, 4600, 4800, 5800)
Solution
93
Median = size of (n + 1 /2) th item = ( 7 + 1 /2)th item = 4th item
Median = 4400
STANDARD DEVIATION
- It’s the square root of the arithmetic mean of the squares of the deviations from the mean.
- It’s denoted by small sigma.
- Also known as root mean square deviation for the reason that is the square root of the
means of square deviations from the A.M.
- It’s a measure of dispersion, it’s a measure of how much ‘spread’ or ‘variability’ is
present in the sample.
- If all the numbers in the sample are very close to each other, the standard deviation is
close to zero.
- If the numbers are well dispersed, the standard deviation will tend to be large. It shows
that a small standard deviation means a high degree of uniformity of the observations as
well as homogeneity of a distribution and vice versa.
Standard deviation of ungrouped data (discrete)
Example 1, 2, 3,3,4,5
Steps
a. Calculate the mean. = mean = 1+2+3+3+4+5 /6 = 18/6 = 3
b. Subtract the mean from every score
c. Square the results (deviations)
d. Add the squared deviations.
e. Divide the sum of squared deviations by N.
Scores X – mean X X2
1 1-3 -2 4
2 2-3 -1 1
3 3-3 0 0
3 3-3 0 0
4 4-3 +1 1
5 5-3 +2 4
SigmaX2=10
= sigma X2 /N
= 10/6 = 1.67 = 1.7 (variance)
Variance is a measure of variability which takes into account the difference between each
observation and sample mean. The larger this deviations are from the mean the larger the
variability.
94
S = square root of variance = 1.3
Variance
- It combines all the values in a data set to produce a measure of spreads.
- The variance (S2) and standard deviation (the square root of the variance) are the most
commonly used measures of spreads.
- It’s calculated as the average squared deviation of each number from the mean of a data
set.
- Calculating variance involves squaring deviations, so it does not have the same unit of
measurement as the original observations.
- Taking the square root of the variance gives us the units used in the original scale and this
is standard deviation.
- Variance (S2) = average squared deviation of values from mean.
- Standard deviation is the measure of spread most commonly used in statistical practice
when the mean is used to calculate central tendency. Thus, it measures spread around the
mean.
- Because of its close links with the mean, standard deviation can be greatly affected if the
mean gives a poor measure of central tendency.
Properties of standard deviation
- Only used to measure spread or dispersion and the mean of a data set.
- Standard deviation is never negative.
- Standard deviation is sensitive to outliers. A single outlier can raise the standard
deviation and in turn, distort the picture of spread.
- For data with approximately the same mean, the greater the spread, the greater the
standard deviation.
- If all data set are the same, the standard deviation is zero (because each value is equal to
the mean).
Advantages of standard deviation
- Its rigidly defined and is based on all the observations of the distribution.
- Its amendable to algebraic treatment.
- It finds wide application in other statistical techniques like skewness, correlation and
regression analysis, sampling theory and tests of significance etc.
- Its possible to calculate the combined standard deviation of two or more groups.
Disadvantages
- It cannot be used for comparing the dispersion of two or more series observations given
in different units.
95
- Its difficult to understand and compute.
- It gives more weight to extreme figures/values
Co-efficient of variation
- Its also known as relative standard variation.
- It’s a standardized measure of dispersion of a probability distribution or frequency
distribution.
- Its often expressed as a percentage and is defined as a ratio of the standard deviation to
the mean.
- Its commonly used when comparing the relative variability of two or more sets of figures.
- The dispersion of two or more series cannot be compared unless we have a relative
measure of variation.
- Karl Pearson introduced a relative measure of variation.
- It expresses the standard deviation as percentage of arithmetic mean.
- It is calculated as follows:
C.V = STD DEVIATION divided by MEAN.
N/B:
A distribution with smaller co-efficient of variation is said to be more homogeneous or uniform
than the other series and vice versa.
In order to calculate C.V, first calculate the mean and standard deviation of the distribution.
Calculate the co-efficient of variation from the following data:
Basic statistics symbols
1. ∑ – sigma (summation) sum of
2. π - mean/average/arithmetic mean
3. ×1– assumed mean
4. (x or n) – individual items/number of items in the distribution.
5. ԁ1 – deviation
6. f – frequency
7. Z – mode
8. Q.d – quartile deviation
9. S- class interval
10. S.k – skewness
11. C.V- coefficient of variation.
12. ~_ – approximately.
13. + - addition
14. - - subtraction
15. ÷ – division
16. × – multiplication
17. = - equals to
96
18. n – sample size
19. N – population size
20. – mean – population mean
21. – sample mean
22. S- sample standard deviation
23. (sigma) – population standard deviation.
97