Basic Statistics
1
Objectives
At the end of this lecture the students will be able to:
Define statistics
Explain what is meant by descriptive and inferential statistics.
Distinguish between a qualitative and a quantitative variables.
Explain the roles of statistics
Describe the types of data and scales of measurement
Identify different methods of data collection
Organize data into a frequency distribution
2
What is statistics?
The scientific study of numerical data based on variation in nature.
A set of procedures and rules for reducing large masses of data into
manageable proportions allowing us to draw conclusions from those data.
Statistics is the art and science of making decisions in the face of
uncertainty
Statistics is the science of collecting, summarizing, presenting, interpreting
data, and of using them to test hypotheses.
We can define it in two senses :
In the plural sense: Statistics are the raw data themselves
In the singular sense: as scientific methods for collecting, organizing,
summarizing, presenting and analyzing data as well as deriving valid
conclusions and making reasonable decisions.
Statistics can be divided in to two main areas or branches.
Descriptive statistics is concerned with summary calculations, graphs,
charts and tables. Generally describes a set of data elements by
graphically displaying the information or describing its central
tendencies and how it is distributed. This also include Collection,
organization, summarization, and presentation of data.
Inferential statistics consists of Generalizing from samples to
populations using probabilities.
performing estimations and hypothesis tests,
Determining relationships between variables,
Making predictions.
4
1.2 Stages in Statistical Investigation
The subject statistics follows five stages:
Objective and planning: clearly describe a well defined objective based on the
problem encountered. Well defined objectives includes data needed, type of data
needed, source of data, etc. are the foundations for proper planning and application
of different statistical procedures so as to make the data more informative and
conclusive.
Collection of Data: Having planed the objectives, the next task is to collect or
combine the relevant data. Data can be combined from the existing sources, or these
can be generated from experiments conducted for the purpose adopting:
i. Complete enumeration (or census) and Sampling technique
Organization of data:
i. Edition of data: Once the data are collected, these need to be checked for
correctness at the first instance. Thus, data sets collected (raw data) should be put
under rigorous checking before these are subjected to further analysis.
Classifications: data arrangement according to their homogeneity
5
Data tabulation : arranging data in table form
Diagramatic and graphical presentation of data: Upon collection of the
data following a definite procedure of collection from the population,
having specific objectives in mind, and on being edited, the next step can
be presented in tabular and graphics.
Analysis of data: Various statistical measures/tools, central tendency,
measures of dispersion, association, probability distribution, testing of
hypothesis, modeling, and other analyses are now applied on the data
collected, edited, and tabulated to answer the objective of the study.
Statistical Inference: Based on the results as revealed from the analysis
of data, statistical implications practical conclusion are drawn in relation to
the objectives of the study framed earlier.
6
Statistical data unless it possesses the following criteria.
The data must be aggregate of facts
They must be affected to a marked extent by a multiplicity of causes
They must be estimated according to reasonable standards of
accuracy
The data are collected in a systematic manner for predefined purpose
The data should be placed in relation to each other
7
Definitions of some Statistical terms
Population: is a data set representing the entire entity of interest.
A Population Size (N) refers to the number of observations in the population.
A Finite Population is a population having fixed number of elements. population of
students in a university, germplasms of mango in a mango garden and so on.
Infinite Population is a population having infinite number of element. For
example, fishes in a particular river, population of hairs on a person’s head, etc.
A Target Population is The entire group of people or objects to which the
researcher wishes to generalize the study findings. Example: All people with AIDS
in specific study area.
Sampling: The process or method of sample selection from the population.
Census is Complete enumeration of the population. Data are collected from each
and every individual unit of the targeted population.
8
A sample is a data set consisting of a portion (or subset) of a population. It is a
selection of individuals, events, or objects taken from a well-de fined population.
Sample is a representative of the population.
Data: are the values (measurements or observations) that the variables can assume
Survey: A collection of quantitative information about members of a population
Sample survey: A survey that include only a portion of the population
Census: A collection of information about every member of a population
Sample Size ( n) is defined as the number of elements/units with which
the sample is constituted of. There is no hard and fast rule, but generally a
sample is recognized as large sample if the sample size n 30, otherwise
small sample.
A set of data is a collection of observed values representing one or more
characteristics of some objects or units. Each value in the data set is called
9
a data value (or a datum).
Parameter: Characteristic or measure obtained from a population.
Statistic: Characteristic or measure obtained from a sample.
Variable: is a characteristic or attribute that can assume different values.
Variables can be classified in two:
Qualitative variable or Quantitative variable
Qualitative variables: are nonnumeric variables and can't be measured,
variables have no numerical meaning. e.g. gender, religious affiliation
Quantitative variables: are numerical variables and can be measured, e.g.
number of children in family, height, weight, income, yield, age, etc.
is naturally measured as a number for which meaningful arithmetic
operations make sense
Quantitative variables are classified as continuous or discrete, according to
the
10 number of values they can take as count ( discrete ) or as measured
(continuous) value.
Continuous variable can take any value within a specified interval. No
gaps between possible values. They are obtained by measuring.
Discrete variables can assume certain numerical values. That is, there are
gaps between the possible values. It can assumes values that can be
counted. Discrete variables can be assigned as countable values such as 0,
1, 2, 3, . . . .
Example: number of children in a family, the number of students in a
classroom, and the number of calls received by a switchboard operator
each day for a month.
Continuous variables can assume an infinite number of values between
any two specific intervals. They are obtained by measuring. They often
include fractions and decimals.
Example: Length, Temperature, Time, Mass, BP, Hg level, etc, 11
12
13
Example of Population and Sample
1. An insurance company would like to determine the proportion of all medical doctors
who have been involved in one or more times to treatment of COVID-19. The
company selects 500 doctors at random from a professional directory and determines
the number in the sample who have been involved in a treatment of COVID-19.
Solution
The population is all medical doctors listed in the professional directory.
The sample is the 500 doctors selected at random from the professional directory.
2. We want to know the average amount of money first year college students spend at
ABC College on school supplies that do not include books. We randomly survey 100
first year students at the college. Three of those students spent $150, $200, and $225,
respectively.
Solution
The population is all first year students attending ABC College this term.
The sample is the 100 first year students surveyed at the college.
The variable could be the amount of money spent by one first year student.
The data would be the actual dollar amounts spent by the first year students.
Examples of the data would be $150, $200, and $225.
14
Uses of Statistics
Statistics presents fact in the form of numerical data
It condenses and summarizes a mass of data.
It facilitates comparison of data
It helps in formulating and testing hypothesis
It helps in predicting future trend
It helps in formulating polices.
Limitations of Statistics
Statistics is not suitable to the study of qualitative phenomenon
Statistics does not study individuals
Statistical laws are not exact
Statistics table may be misused
15
15 Statistics is only, one of the methods of studying a problem
2. Scales of measurement refer to ways in which variables or numbers are
defined and categorized.
Use of level of measurements
Helps you decide how to interpret the data from the variable.
Helps you decide what statistical analysis is appropriate on the values
that were assigned.
There are four levels of measurement scales; Nominal, ordinal, interval
and Ratio scales
Nominal scales:
Mutually exclusive unordered categories
No arithmetic and relational operation can be applied.
Examples
– Sex (male, female)
– Race/ethnicity (white, black, latino, asian, native american.)
– Marital status Blood type
Can be summarize by: Tables – using counts and percentages and Bar
16 chart
Ordinal Scales
Ordered Categories
Differences between the ranks do not exist.
Arithmetic operations are not applicable but relational operations
Ordering is the sole property of ordinal scale.
Examples: Disease state
Mild
Moderate
Severe
Agreement questions
Strongly agree
Agree
Indifferent
Disagree
Strongly disagree
Income
Low, medium, high
17
Interval Scales
data that can be ranked and differences are meaningful
There is no meaningful zero, so ratios are meaningless.
All arithmetic operations except division are applicable.
Relational operations are also possible.
E.g. IQ , Temperature in oF.
Ratio Scales
• Data that can be ranked, differences are meaningful, and there is a
true zero.
• - E.g. Age, weight, height, pulse rate
-classifies data that can be ranked, differences are meaningful,
and there is a true zero. True ratios exist between the different
units of measure.
All arithmetic and relational operations are applicable.
18
Exercise
Classify the ff measurement systems into one of the four types
of scales
I. Times for swimmers to complete a 50-meter race
II. Blood type of patients.
III. Hemoglobin level
IV. Blood pressure
V. The net wages of a group of workers;
VI. Regions numbers of Ethiopia
VII. Socioeconomic status of a family when classified as low,
middle and upper classes
19
Exercise
1. A social researcher in a particular city wishes to obtain information on the number
of children in households that receive welfare support. A random sample of 400
households is selected from the city welfare rolls. A check on welfare recipient data
provides the number of children in each household.
a) Identify the population of measurements that is of interest to the researcher.
b) Identify the sample.
c) What characteristics of the population are of interest to the researcher?
2. The faculty senate at a major university with 35,000 students is considering changing
the current grading policy from A, B, C, D, F to a plus and minus system—that is, B-
, B, B+ rather than just B. The faculty is interested in the students’ opinions
concerning this change and will sample 500 students.
a) What is the population of interest?
b) Identify the population size
c) What is the sample and sample size?
d) How could the sample be selected?
20
Methods of data collection and Organization
Types of Data: There are two types (sources) of data.
Primary data
are the first hand information collected, compiled and published by
organization for some purpose.
Are original data in character and have not undergone any sort of
statistical treatment.
Refer to those that are collected by conducting survey to meet the
specific problem needs at hand.
Secondary data
are the second hand information collected by someone (organization)
for some purpose. The secondary data are not pure in character and
have undergone some treatment at least once.
21 data taken from already available published or unpublished source.
Methods of collection
o Collecting Primary data
Observation Interview
Use of self administered questionnaire
o Collecting secondary data Use of documentary sources
Primary sources Secondary sources
- Sample units (respondents) - Previous research
- Individuals or groups - Official statistics
Key informants - Mass media products
- Merchants - Diaries
- Religious people - Letters
- Elder people - Government reports
- Experts in the study area - Web information
22
- Historical data and information
Observation
Systematically selecting , watching and recording behaviours of people or
other phenomena and aspects of the settings in which they occur.
For the purpose of obtaining specified observation
Includes
Visual observation
Radiographic, Biomedical, x-ray, microscope, clinical examinations, etc
It can also be used in observing behaviour of people, culture etc.
It could be
Participant observation or
Non-participant observation
23
o Advantage
More accurate data on behaviour or activity
o Disadvantages
Observer bias
Prejudice
Desirability bias
Needs skilled human power in high level machines
Interviews
Face to face interview
Telephone interview
Group interview or Focused Group Discussion (FGD)
Face to face interview
o Advantage
Permits detailed & in-depth questions & responses
Minimizes non-response
o Disadvantage
24
Costly, Interviewer bias, Investigator bias, Interviewer cheating
Telephone interview
Advantage
Convenient
Saves time
Relatively inexpensive
Less interviewer & investigator bias than personal interview
Disadvantage
Non-coverage
Limited length & depth of questions and responses
25
Self-administered Questionnaire
Questionnaire is the main data collection instrument in formal sample
survey
Open-ended questions: - allows the respondent to respond freely in his or
her own words
Closed – ended questions:- Predetermined list of alternate responses is
presented to the respondent for checking the appropriate one(s)
Advantage
Cost effective for large areas
Minimizes interviewer bias
Promotes accurate answers
Sensitive issues can be gathered
Disadvantage
Low response rates
Unanswered questions
Incorrect answers
26
Use of documentary sources
These include
Clinical & other personal records
Vital statistics
Census data
Sources
Official publications of different organizations
News papers & journals
International publications (WHO, UNICEF, etc)
Health facilities’ records
27
Data Organization
Editing of Data means the examination of collected data to discover any
error and mistake before presenting it
Classification of Data
The process of arranging data into homogenous group or classes according to
some common characteristics present in the data is called classification.
The bases of classification:
(1) Qualitative Base (3) Geographical Base
(2) Quantitative Base (4) Chronological or Temporal Base
Tabulation of Data
The process of placing classified data into tabular form is known as
tabulation.
28
A table is a symmetric arrangement of statistical data in rows and columns.
Frequency distributions
Is a table that shows data classified in to a number of classes with a
corresponding number of times falling in each categories (frequency)
Frequency is the number of times a certain value of the variable is
separated in a given class.
There are three types frequency distribution
Categorical frequency distribution
Ungrouped frequency distribution
Grouped frequency distributions.
29
Categorical frequency distribution
• Used for data that can be placed in specific categories
• Used for nominal & Ordinal
– E.g. Blood type, marital status etc.
• Example: A health worker collected data on blood type of 30
individuals and recorded as follows (Hypothetical)
• O, A, AB, B, O, O, O, A, B, O, AB, B, B, A, AB, O, O, O, B, AB, O,
A, AB, B, O, O, O, A, B, O
30
Ungrouped frequency distribution
Used to organize discrete quantitative data in tabular form
Is a table of all the potential raw score values along with the number of
times each actually occurred
Often used for small set of data on discrete variables
The major components of this type of FD are class, tally, frequency and
cumulative frequency.
Example:- The following data represent the number of days of sick leave
taken by each of 50 workers of a company over the last 6 weeks.
2 0 0 5 8 3 4 1 0 0
7 1 7 1 5 4 0 4 0 1
8 9 7 0 1 7 2 5 5 4
3 3 0 0 2 5 1 3 0 2
31 4 5 0 5 7 5 1 1 0 2
Construct ungrouped frequency distribution
How many workers had at least 1 day of sick leave?
How many workers had between 3 and 5 days of sick leave?
A frequency table which presents each distinct value along with its
frequency of occurrences is given below
Since 12 of the 50 workers had no days of sick leave, the answer is 50-12=38
The answer is the sum of the frequencies for values 3, 4 and 5 that is
32
4+5+8=17
Grouped FD:- Some of basic terms used in FD
Lower class limits (LCL)- the smallest number that can belong to the
different class
Upper class limits (UCL)- are the largest number that can belong to the
different classes.
Class Boundaries are the number used to separate classes, but without
the gaps created by class limits.
Class marks are the midpoints of the classes. It can be found by adding
the lower class limit to the upper class limit and dividing the sum by 2.
Class width is the difference between two consecutive lower class limits
or two consecutive lower class boundaries.
33
Steps to construct grouped FD:
Select the number of classes, use Struge’s rule.
Κ=1+ 3.32 log(n) (rounding up)
Find the highest and the lowest values.
Range (R)=maximum-minimum
Find the class width, width(w)= R/K, (rounding up)
Select the starting point as the lowest class limit. Add the width to that
score to get the lower class limit of the next class. Keep adding until you
achieve the number of desired classes.
Find the upper class limit, UCL1=LCL1+Cw-U/2; U is unit of
measurement. Then add the width to each upper class limit to get all
upper class limits.
Unit of measurement is the difference between the next expected
upcoming value.
34
Guidelines for creating classes
1. There should be b/n 5-20 classes
2. The classes must be mutually exclusive.
i.e. no data value fall into two d/t classes
3. The classes must be all inclusive or exhaustive. i.e. all data values must
be included
4. The classes must be continues
i .e. no gaps in a frequency distribution
5. The classes must be equal in width.
35
Example: Consider the following marks of 30 Accounting and Finance
students out of 50%
24 30 36 35 42 40 26 23
36 36 12 45 29 21 34 40
16 47 28 32 33 44 19 34
30 36 35 47 20 14
Construct grouped frequency distribution for the above data
Steps
1. K=1+3.32log(30)=5.9 ≈ 6
2. R= 47-12 =35
3. Width (w)= 35/6 = 5.8, rounded to 6
Select the starting point (Small value) = 12, so,
1st lower limit = 12 4th lower limit = 24+6 = 30
2nd lower limit = 12 +6 =18 5th lower limit= 30+6 = 36
36
3rd lower limit = 18+6 = 24 6th lower limit = 36+6 = 42
Find the upper class limit (UCL) by UCL1=LCL1+Cw-U/2
Class Tally Freq Class Class
1st UCL = 12 + 6-1 = 17 mark boundary
12-17 /// 3 14.5 11.5-17.5
2nd UCL = 17 + 6 = 23
18-23 //// 4 20.5 17.5-23.5
3rd UCL = 23 + 6 = 29
24-29 //// 4 26.5 23.5-29.5
4th UCL = 29 + 6 = 35
30-35 //// /// 8 32.5 29.5-35.5
5th UCL = 35 + 6 = 41 36-41 //// / 6 38.5 35.5-41.5
6th UCL = 41 + 6 = 47 42-47 //// 5 44.5 41.5-47.5
Total 30
37
Cumulative frequency distributions
There are two types of cumulative frequency distributions namely the
―less than CF‖ and the ―more than CF ‖ distribution.
Less than CF: it is the sum of all frequencies lying below the upper class
boundaries of each class.
Shows the number of data sets less than upper class boundary of each
classes
More than CF: : it is the sum of all frequencies lying above the lower
class boundaries of each class.
Shows the number of data sets greater than lower class boundary of each
classes
38
Classes Tally Freq LCf MCf (greater
(less than) than)
12-17 /// 3 3 30
18-23 //// 4 7 27
24-29 //// 4 11 23
30-35 //// /// 8 19 19
36-41 //// / 6 25 11
42-47 //// 5 30 5
Total 30
39
Diagrammatic/Graphical presentation of data
Techniques for presenting data in visual displays using geometric and
pictures.
Importance
Greater attraction Easily understandable
Facilitate comparison Greater memorizing value
May reveal unsuspected patterns in complex set of data
Diagrammatic display of data:-
Pie chart
Barcharts
Pictograms
Pie chart: is a circular diagram and the area of the sector of a circle is
used in pie chart.
The angles of each component are calculated by the formula.
Angle of sector= component part/Total *360◦
40
Example: The following table gives the details of monthly budget
of a family.
41
Bar Charts
Simple bar chart
Component bar chart
Multiple bar chart
It uses vertical bars to represent the frequencies of a distribution.
Make the bars the same width
simple bar chart is used to represents data involving only one variable
42
Stratified (Stacked or component) Bar Chart is used to represent data in which
the total magnitude is divided into different or components.
Example : The table below shows the quantity in hundred kgs of Wheat, Barley
and Oats produced on a certain form during the years 1991 to 1994. Draw stratified
bar chart/Component.
Years Wheat Barley Oats
1991 34 18 27
1992 43 14 24
1993 43 16 27
1994 45 13 34
43
44
Multiple bar charts are used for two or more sets of inter-related data are
represented
multiple bar diagram facilities comparison between more than one phenomenon.
Example 2.6: Draw a multiple bar chart to represent the import and export of
Canada (values in $) for the years 1991 to 1995.
Years Imports Exports
1991 7930 4260
1992 8850 5225
1993 9780 6150
1994 11720 7340
1995 12150 8145
45
46
Graphical presentation of data:
The three most commonly used graphs in research are
1. The histogram.
2. The frequency polygon.
3. The cumulative frequency graph, or ogive
Histogram is a special type of bar graph in which the horizontal scale
represents classes of data values and the vertical scale represents
frequencies.
The height of the bars correspond to the frequency values,
The bars drawn adjacent to each other (without gaps).
47
Example : Consider the data on marks of 30 Accounting and finance
students
Marks Frequency Class boundary
12-17 3 11.5-17.5
18-23 4 17.5-23.5
24-29 4 23.5-29.5
30-35 8 29.5-35.5
36-41 6 35.5-41.5
42-47 5 41.5-47.5
Total 30
48
Frequency Polygon
Join the mid points of the tops of the adjacent rectangles of the
histogram with segments
When it is joined with x-axis the area under the polygon is equal to
the area under the histogram.
The scales should be marked in the numerical values of the
midpoints (Xc)
The length of the ordinates represent the class frequency.
49
Cumulative frequency polygon (Ogive
Line graph obtained by plotting the cumulative frequency distribution (Y-
axis) against class boundaries (x-axis)
Two types
Cumulative frequency Less than the UCB (Lcf)or
Cumulative frequency More than the LCB (Mcf)
Classes Class boundaries Freq Less than cF More than cf
12-17 11.5-17.5 3 3 30
18-23 17.5-23.5 4 7 27
24-29 23.5-29.5 4 11 23
30-35 29.5-35.5 8 19 19
36-41 35.5-41.5 6 25 11
42-47 41.5-47.5 5 30 5
Total 30
50
Less than & More than Ogive
51
Thank you!
52