0% found this document useful (0 votes)
20 views52 pages

Introduction to Basic Statistics Concepts

This document outlines the fundamentals of statistics, including definitions, branches (descriptive and inferential), and types of data. It details the stages of statistical investigation, methods of data collection, and the importance of organizing data into frequency distributions. Additionally, it discusses scales of measurement and provides examples of populations and samples, along with the uses and limitations of statistics.

Uploaded by

TEFERI BEZU
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
20 views52 pages

Introduction to Basic Statistics Concepts

This document outlines the fundamentals of statistics, including definitions, branches (descriptive and inferential), and types of data. It details the stages of statistical investigation, methods of data collection, and the importance of organizing data into frequency distributions. Additionally, it discusses scales of measurement and provides examples of populations and samples, along with the uses and limitations of statistics.

Uploaded by

TEFERI BEZU
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Basic Statistics

1
Objectives
 At the end of this lecture the students will be able to:
 Define statistics

 Explain what is meant by descriptive and inferential statistics.

 Distinguish between a qualitative and a quantitative variables.

 Explain the roles of statistics

 Describe the types of data and scales of measurement

 Identify different methods of data collection

 Organize data into a frequency distribution

2
What is statistics?
 The scientific study of numerical data based on variation in nature.

 A set of procedures and rules for reducing large masses of data into
manageable proportions allowing us to draw conclusions from those data.
 Statistics is the art and science of making decisions in the face of
uncertainty
 Statistics is the science of collecting, summarizing, presenting, interpreting
data, and of using them to test hypotheses.
We can define it in two senses :
 In the plural sense: Statistics are the raw data themselves

 In the singular sense: as scientific methods for collecting, organizing,


summarizing, presenting and analyzing data as well as deriving valid
conclusions and making reasonable decisions.
Statistics can be divided in to two main areas or branches.
 Descriptive statistics is concerned with summary calculations, graphs,
charts and tables. Generally describes a set of data elements by
graphically displaying the information or describing its central
tendencies and how it is distributed. This also include Collection,
organization, summarization, and presentation of data.
 Inferential statistics consists of Generalizing from samples to
populations using probabilities.
 performing estimations and hypothesis tests,
 Determining relationships between variables,
 Making predictions.

4
 1.2 Stages in Statistical Investigation
The subject statistics follows five stages:
 Objective and planning: clearly describe a well defined objective based on the

problem encountered. Well defined objectives includes data needed, type of data
needed, source of data, etc. are the foundations for proper planning and application
of different statistical procedures so as to make the data more informative and
conclusive.
 Collection of Data: Having planed the objectives, the next task is to collect or
combine the relevant data. Data can be combined from the existing sources, or these
can be generated from experiments conducted for the purpose adopting:
i. Complete enumeration (or census) and Sampling technique
 Organization of data:
 i. Edition of data: Once the data are collected, these need to be checked for
correctness at the first instance. Thus, data sets collected (raw data) should be put
under rigorous checking before these are subjected to further analysis.
 Classifications: data arrangement according to their homogeneity
5
 Data tabulation : arranging data in table form
 Diagramatic and graphical presentation of data: Upon collection of the

data following a definite procedure of collection from the population,


having specific objectives in mind, and on being edited, the next step can
be presented in tabular and graphics.

 Analysis of data: Various statistical measures/tools, central tendency,

measures of dispersion, association, probability distribution, testing of


hypothesis, modeling, and other analyses are now applied on the data
collected, edited, and tabulated to answer the objective of the study.

 Statistical Inference: Based on the results as revealed from the analysis

of data, statistical implications practical conclusion are drawn in relation to


the objectives of the study framed earlier.

6
Statistical data unless it possesses the following criteria.
The data must be aggregate of facts
They must be affected to a marked extent by a multiplicity of causes
They must be estimated according to reasonable standards of
accuracy
The data are collected in a systematic manner for predefined purpose
The data should be placed in relation to each other

7
Definitions of some Statistical terms
Population: is a data set representing the entire entity of interest.
A Population Size (N) refers to the number of observations in the population.
A Finite Population is a population having fixed number of elements. population of
students in a university, germplasms of mango in a mango garden and so on.
Infinite Population is a population having infinite number of element. For
example, fishes in a particular river, population of hairs on a person’s head, etc.
A Target Population is The entire group of people or objects to which the
researcher wishes to generalize the study findings. Example: All people with AIDS
in specific study area.
Sampling: The process or method of sample selection from the population.
Census is Complete enumeration of the population. Data are collected from each
and every individual unit of the targeted population.

8
A sample is a data set consisting of a portion (or subset) of a population. It is a
selection of individuals, events, or objects taken from a well-de fined population.
Sample is a representative of the population.
Data: are the values (measurements or observations) that the variables can assume
Survey: A collection of quantitative information about members of a population
Sample survey: A survey that include only a portion of the population
Census: A collection of information about every member of a population
Sample Size ( n) is defined as the number of elements/units with which
the sample is constituted of. There is no hard and fast rule, but generally a
sample is recognized as large sample if the sample size n 30, otherwise
small sample.
A set of data is a collection of observed values representing one or more
characteristics of some objects or units. Each value in the data set is called
9
a data value (or a datum).
 Parameter: Characteristic or measure obtained from a population.

 Statistic: Characteristic or measure obtained from a sample.

 Variable: is a characteristic or attribute that can assume different values.


 Variables can be classified in two:
Qualitative variable or Quantitative variable
 Qualitative variables: are nonnumeric variables and can't be measured,

variables have no numerical meaning. e.g. gender, religious affiliation

 Quantitative variables: are numerical variables and can be measured, e.g.

number of children in family, height, weight, income, yield, age, etc.

 is naturally measured as a number for which meaningful arithmetic

operations make sense


 Quantitative variables are classified as continuous or discrete, according to
the
10 number of values they can take as count ( discrete ) or as measured
(continuous) value.
 Continuous variable can take any value within a specified interval. No

gaps between possible values. They are obtained by measuring.

 Discrete variables can assume certain numerical values. That is, there are

gaps between the possible values. It can assumes values that can be
counted. Discrete variables can be assigned as countable values such as 0,
1, 2, 3, . . . .

Example: number of children in a family, the number of students in a


classroom, and the number of calls received by a switchboard operator
each day for a month.

 Continuous variables can assume an infinite number of values between

any two specific intervals. They are obtained by measuring. They often
include fractions and decimals.
 Example: Length, Temperature, Time, Mass, BP, Hg level, etc, 11
12
13
Example of Population and Sample
1. An insurance company would like to determine the proportion of all medical doctors
who have been involved in one or more times to treatment of COVID-19. The
company selects 500 doctors at random from a professional directory and determines
the number in the sample who have been involved in a treatment of COVID-19.
Solution
 The population is all medical doctors listed in the professional directory.
 The sample is the 500 doctors selected at random from the professional directory.
2. We want to know the average amount of money first year college students spend at
ABC College on school supplies that do not include books. We randomly survey 100
first year students at the college. Three of those students spent $150, $200, and $225,
respectively.
Solution
 The population is all first year students attending ABC College this term.
 The sample is the 100 first year students surveyed at the college.
 The variable could be the amount of money spent by one first year student.
 The data would be the actual dollar amounts spent by the first year students.
Examples of the data would be $150, $200, and $225.
14
Uses of Statistics
 Statistics presents fact in the form of numerical data

 It condenses and summarizes a mass of data.

 It facilitates comparison of data

 It helps in formulating and testing hypothesis

 It helps in predicting future trend

 It helps in formulating polices.


Limitations of Statistics
 Statistics is not suitable to the study of qualitative phenomenon

 Statistics does not study individuals

 Statistical laws are not exact


 Statistics table may be misused
15

15 Statistics is only, one of the methods of studying a problem
2. Scales of measurement refer to ways in which variables or numbers are
defined and categorized.
 Use of level of measurements
 Helps you decide how to interpret the data from the variable.
 Helps you decide what statistical analysis is appropriate on the values
that were assigned.
There are four levels of measurement scales; Nominal, ordinal, interval
and Ratio scales
 Nominal scales:
 Mutually exclusive unordered categories
 No arithmetic and relational operation can be applied.
Examples
– Sex (male, female)
– Race/ethnicity (white, black, latino, asian, native american.)
– Marital status Blood type
 Can be summarize by: Tables – using counts and percentages and Bar
16 chart
Ordinal Scales
 Ordered Categories
 Differences between the ranks do not exist.
 Arithmetic operations are not applicable but relational operations
 Ordering is the sole property of ordinal scale.
Examples: Disease state
Mild
Moderate
Severe
Agreement questions
Strongly agree
Agree
Indifferent
Disagree
Strongly disagree
Income
Low, medium, high
17
Interval Scales
 data that can be ranked and differences are meaningful
 There is no meaningful zero, so ratios are meaningless.
 All arithmetic operations except division are applicable.
 Relational operations are also possible.
 E.g. IQ , Temperature in oF.
Ratio Scales
• Data that can be ranked, differences are meaningful, and there is a
true zero.
• - E.g. Age, weight, height, pulse rate
 -classifies data that can be ranked, differences are meaningful,
and there is a true zero. True ratios exist between the different
units of measure.
 All arithmetic and relational operations are applicable.
18
Exercise
 Classify the ff measurement systems into one of the four types
of scales
I. Times for swimmers to complete a 50-meter race
II. Blood type of patients.
III. Hemoglobin level
IV. Blood pressure
V. The net wages of a group of workers;
VI. Regions numbers of Ethiopia
VII. Socioeconomic status of a family when classified as low,
middle and upper classes

19
Exercise
1. A social researcher in a particular city wishes to obtain information on the number
of children in households that receive welfare support. A random sample of 400
households is selected from the city welfare rolls. A check on welfare recipient data
provides the number of children in each household.
a) Identify the population of measurements that is of interest to the researcher.
b) Identify the sample.
c) What characteristics of the population are of interest to the researcher?
2. The faculty senate at a major university with 35,000 students is considering changing
the current grading policy from A, B, C, D, F to a plus and minus system—that is, B-
, B, B+ rather than just B. The faculty is interested in the students’ opinions
concerning this change and will sample 500 students.
a) What is the population of interest?
b) Identify the population size
c) What is the sample and sample size?
d) How could the sample be selected?

20
Methods of data collection and Organization
Types of Data: There are two types (sources) of data.
 Primary data
 are the first hand information collected, compiled and published by
organization for some purpose.
 Are original data in character and have not undergone any sort of

statistical treatment.

 Refer to those that are collected by conducting survey to meet the

specific problem needs at hand.


Secondary data
 are the second hand information collected by someone (organization)
for some purpose. The secondary data are not pure in character and
have undergone some treatment at least once.

21 data taken from already available published or unpublished source.
Methods of collection
o Collecting Primary data
Observation Interview
Use of self administered questionnaire
o Collecting secondary data Use of documentary sources

Primary sources Secondary sources


- Sample units (respondents) - Previous research
- Individuals or groups - Official statistics
Key informants - Mass media products
- Merchants - Diaries
- Religious people - Letters
- Elder people - Government reports
- Experts in the study area - Web information
22
- Historical data and information
Observation
 Systematically selecting , watching and recording behaviours of people or

other phenomena and aspects of the settings in which they occur.

 For the purpose of obtaining specified observation

 Includes

 Visual observation

 Radiographic, Biomedical, x-ray, microscope, clinical examinations, etc

 It can also be used in observing behaviour of people, culture etc.

 It could be

 Participant observation or

 Non-participant observation
23
o Advantage
 More accurate data on behaviour or activity
o Disadvantages
 Observer bias
 Prejudice
 Desirability bias
 Needs skilled human power in high level machines
Interviews
 Face to face interview
 Telephone interview
 Group interview or Focused Group Discussion (FGD)
Face to face interview
o Advantage
 Permits detailed & in-depth questions & responses
 Minimizes non-response
o Disadvantage
24
 Costly, Interviewer bias, Investigator bias, Interviewer cheating
Telephone interview
 Advantage
 Convenient
 Saves time
 Relatively inexpensive
 Less interviewer & investigator bias than personal interview
 Disadvantage
 Non-coverage
 Limited length & depth of questions and responses

25
Self-administered Questionnaire
 Questionnaire is the main data collection instrument in formal sample
survey
 Open-ended questions: - allows the respondent to respond freely in his or
her own words
 Closed – ended questions:- Predetermined list of alternate responses is
presented to the respondent for checking the appropriate one(s)
 Advantage
 Cost effective for large areas
 Minimizes interviewer bias
 Promotes accurate answers
 Sensitive issues can be gathered
 Disadvantage
 Low response rates
 Unanswered questions
 Incorrect answers
26
Use of documentary sources
 These include
 Clinical & other personal records
 Vital statistics
 Census data
 Sources
 Official publications of different organizations
 News papers & journals
 International publications (WHO, UNICEF, etc)
 Health facilities’ records

27
Data Organization
 Editing of Data means the examination of collected data to discover any

error and mistake before presenting it

 Classification of Data

 The process of arranging data into homogenous group or classes according to

some common characteristics present in the data is called classification.


The bases of classification:
(1) Qualitative Base (3) Geographical Base

(2) Quantitative Base (4) Chronological or Temporal Base

 Tabulation of Data

 The process of placing classified data into tabular form is known as

tabulation.
28
 A table is a symmetric arrangement of statistical data in rows and columns.
Frequency distributions

 Is a table that shows data classified in to a number of classes with a

corresponding number of times falling in each categories (frequency)

 Frequency is the number of times a certain value of the variable is

separated in a given class.

There are three types frequency distribution


 Categorical frequency distribution

 Ungrouped frequency distribution

 Grouped frequency distributions.

29
Categorical frequency distribution
• Used for data that can be placed in specific categories
• Used for nominal & Ordinal
– E.g. Blood type, marital status etc.
• Example: A health worker collected data on blood type of 30
individuals and recorded as follows (Hypothetical)
• O, A, AB, B, O, O, O, A, B, O, AB, B, B, A, AB, O, O, O, B, AB, O,
A, AB, B, O, O, O, A, B, O

30
Ungrouped frequency distribution
 Used to organize discrete quantitative data in tabular form

 Is a table of all the potential raw score values along with the number of

times each actually occurred

 Often used for small set of data on discrete variables

 The major components of this type of FD are class, tally, frequency and
cumulative frequency.
Example:- The following data represent the number of days of sick leave
taken by each of 50 workers of a company over the last 6 weeks.
2 0 0 5 8 3 4 1 0 0
7 1 7 1 5 4 0 4 0 1
8 9 7 0 1 7 2 5 5 4
3 3 0 0 2 5 1 3 0 2
31 4 5 0 5 7 5 1 1 0 2
Construct ungrouped frequency distribution
How many workers had at least 1 day of sick leave?
How many workers had between 3 and 5 days of sick leave?
A frequency table which presents each distinct value along with its
frequency of occurrences is given below

Since 12 of the 50 workers had no days of sick leave, the answer is 50-12=38
The answer is the sum of the frequencies for values 3, 4 and 5 that is
32
4+5+8=17
Grouped FD:- Some of basic terms used in FD

 Lower class limits (LCL)- the smallest number that can belong to the

different class

 Upper class limits (UCL)- are the largest number that can belong to the

different classes.

 Class Boundaries are the number used to separate classes, but without

the gaps created by class limits.

 Class marks are the midpoints of the classes. It can be found by adding

the lower class limit to the upper class limit and dividing the sum by 2.

 Class width is the difference between two consecutive lower class limits

or two consecutive lower class boundaries.


33
Steps to construct grouped FD:
 Select the number of classes, use Struge’s rule.
Κ=1+ 3.32 log(n) (rounding up)
 Find the highest and the lowest values.
Range (R)=maximum-minimum
 Find the class width, width(w)= R/K, (rounding up)

 Select the starting point as the lowest class limit. Add the width to that
score to get the lower class limit of the next class. Keep adding until you
achieve the number of desired classes.
 Find the upper class limit, UCL1=LCL1+Cw-U/2; U is unit of
measurement. Then add the width to each upper class limit to get all
upper class limits.
 Unit of measurement is the difference between the next expected
upcoming value.
34
Guidelines for creating classes

1. There should be b/n 5-20 classes

2. The classes must be mutually exclusive.


i.e. no data value fall into two d/t classes

3. The classes must be all inclusive or exhaustive. i.e. all data values must
be included

4. The classes must be continues

i .e. no gaps in a frequency distribution

5. The classes must be equal in width.

35
Example: Consider the following marks of 30 Accounting and Finance
students out of 50%

24 30 36 35 42 40 26 23
36 36 12 45 29 21 34 40
16 47 28 32 33 44 19 34
30 36 35 47 20 14
Construct grouped frequency distribution for the above data
Steps
1. K=1+3.32log(30)=5.9 ≈ 6
2. R= 47-12 =35
3. Width (w)= 35/6 = 5.8, rounded to 6
Select the starting point (Small value) = 12, so,
1st lower limit = 12 4th lower limit = 24+6 = 30
2nd lower limit = 12 +6 =18 5th lower limit= 30+6 = 36
36
3rd lower limit = 18+6 = 24 6th lower limit = 36+6 = 42
Find the upper class limit (UCL) by UCL1=LCL1+Cw-U/2
Class Tally Freq Class Class
1st UCL = 12 + 6-1 = 17 mark boundary
12-17 /// 3 14.5 11.5-17.5
2nd UCL = 17 + 6 = 23
18-23 //// 4 20.5 17.5-23.5
3rd UCL = 23 + 6 = 29
24-29 //// 4 26.5 23.5-29.5
4th UCL = 29 + 6 = 35
30-35 //// /// 8 32.5 29.5-35.5
5th UCL = 35 + 6 = 41 36-41 //// / 6 38.5 35.5-41.5
6th UCL = 41 + 6 = 47 42-47 //// 5 44.5 41.5-47.5
Total 30

37
Cumulative frequency distributions
 There are two types of cumulative frequency distributions namely the

―less than CF‖ and the ―more than CF ‖ distribution.

 Less than CF: it is the sum of all frequencies lying below the upper class

boundaries of each class.

 Shows the number of data sets less than upper class boundary of each

classes

 More than CF: : it is the sum of all frequencies lying above the lower

class boundaries of each class.

 Shows the number of data sets greater than lower class boundary of each

classes

38
Classes Tally Freq LCf MCf (greater
(less than) than)
12-17 /// 3 3 30
18-23 //// 4 7 27
24-29 //// 4 11 23
30-35 //// /// 8 19 19
36-41 //// / 6 25 11
42-47 //// 5 30 5
Total 30

39
Diagrammatic/Graphical presentation of data
 Techniques for presenting data in visual displays using geometric and
pictures.
 Importance
 Greater attraction Easily understandable
 Facilitate comparison Greater memorizing value
 May reveal unsuspected patterns in complex set of data
 Diagrammatic display of data:-
Pie chart
Barcharts
Pictograms
 Pie chart: is a circular diagram and the area of the sector of a circle is
used in pie chart.
 The angles of each component are calculated by the formula.
 Angle of sector= component part/Total *360◦
40
Example: The following table gives the details of monthly budget
of a family.

41
Bar Charts
 Simple bar chart
Component bar chart
Multiple bar chart
 It uses vertical bars to represent the frequencies of a distribution.
 Make the bars the same width
 simple bar chart is used to represents data involving only one variable

42
Stratified (Stacked or component) Bar Chart is used to represent data in which
the total magnitude is divided into different or components.
Example : The table below shows the quantity in hundred kgs of Wheat, Barley
and Oats produced on a certain form during the years 1991 to 1994. Draw stratified
bar chart/Component.

Years Wheat Barley Oats


1991 34 18 27
1992 43 14 24
1993 43 16 27
1994 45 13 34

43
44
Multiple bar charts are used for two or more sets of inter-related data are
represented
multiple bar diagram facilities comparison between more than one phenomenon.
Example 2.6: Draw a multiple bar chart to represent the import and export of
Canada (values in $) for the years 1991 to 1995.

Years Imports Exports


1991 7930 4260
1992 8850 5225
1993 9780 6150
1994 11720 7340
1995 12150 8145

45
46
Graphical presentation of data:
The three most commonly used graphs in research are
1. The histogram.
2. The frequency polygon.
3. The cumulative frequency graph, or ogive
 Histogram is a special type of bar graph in which the horizontal scale
represents classes of data values and the vertical scale represents
frequencies.
 The height of the bars correspond to the frequency values,
 The bars drawn adjacent to each other (without gaps).

47
Example : Consider the data on marks of 30 Accounting and finance
students
Marks Frequency Class boundary
12-17 3 11.5-17.5
18-23 4 17.5-23.5
24-29 4 23.5-29.5
30-35 8 29.5-35.5
36-41 6 35.5-41.5
42-47 5 41.5-47.5
Total 30

48
Frequency Polygon
 Join the mid points of the tops of the adjacent rectangles of the
histogram with segments
 When it is joined with x-axis the area under the polygon is equal to
the area under the histogram.
 The scales should be marked in the numerical values of the
midpoints (Xc)
 The length of the ordinates represent the class frequency.

49
Cumulative frequency polygon (Ogive
 Line graph obtained by plotting the cumulative frequency distribution (Y-
axis) against class boundaries (x-axis)
 Two types
 Cumulative frequency Less than the UCB (Lcf)or
 Cumulative frequency More than the LCB (Mcf)

Classes Class boundaries Freq Less than cF More than cf


12-17 11.5-17.5 3 3 30
18-23 17.5-23.5 4 7 27
24-29 23.5-29.5 4 11 23
30-35 29.5-35.5 8 19 19
36-41 35.5-41.5 6 25 11
42-47 41.5-47.5 5 30 5
Total 30

50
Less than & More than Ogive

51
Thank you!

52

You might also like