BRIEF REVIEW
OF
BASIC STATISTICS
1
What is statistics?
Statistics is a very broad subject, with applications in a vast number
of different fields. In generally one can say that statistics is the
methodology for collecting, analyzing, interpreting and drawing
conclusions from collected data.
Statistics consists of a body of methods for collecting and analyzing
data.
• What kind and how much data need to be collected?
• How should we organize and summarize the data?
• How can we analyse the data and draw conclusions from it?
• How can we assess the strength of the conclusions and evaluate
their uncertainty?
2
Statistics provides methods for
1. Design: Planning and carrying out research studies.
2. Description: Summarizing and exploring data.
3. Inference: Making predictions and generalizing about
phenomena represented by the data.
3
Population and Sample
Population and sample are two basic concepts of statistics.
Population can be characterized as the set of individual persons or
objects in which an investigator is primarily interested during his or
her research problem. Sometimes wanted measurements for all
individuals in the population are obtained, but often only a set of
individuals of that population are observed; such a set of
individuals constitutes a sample. This gives us the following
definitions of population and sample.
Population is the collection of all individuals or items under
consideration in a statistical study. (Weiss, 1999)
Sample is that part of the population from which information is
collected. (Weiss, 1999)
4
5
A (statistical) population is the set of measurements (or record of
some qualitative trait) corresponding to the entire collection of
units for which inferences are to be made.
A sample from statistical population is the set of measurements
that are actually collected in the course of an investigation.
6
Descriptive and Inferential Statistics
Descriptive statistics consist of methods for organizing and
summarizing information.
Inferential statistics consist of methods for drawing and measuring
the reliability of conclusions about population based on
information obtained from a sample of the population.
7
Parameters and Statistics
Usually the features of the population under investigation can be
summarized by numerical parameters. Hence the research problem
usually becomes as on investigation of the values of parameters.
These population parameters are unknown and sample statistics
are used to make inference about them. That is, a statistic
describes a characteristic of the sample which can then be used
to make inference about unknown parameters.
A parameter is an unknown numerical summary of the population.
A statistic is a known numerical summary of the sample which can
be used to make inference about parameters.
8
Statistical data analysis
9
Variables and organization of the data
Variables
A characteristic that varies from one person or thing to another is
called a variable, i.e, a variable is any characteristic that varies from
one individual member of the population to another.
Examples of variables for humans are height, weight, number of
siblings, sex, marital status, and eye colour. The first three of these
variables yield numerical information (yield numerical
measurements) and are examples of quantitative (or numerical)
variables, last three yield non-numerical information (yield non-
numerical measurements) and are examples of qualitative (or
categorical) variables.
Quantitative variables can be classified as either discrete or
continuous.
10
Discrete variables
Some variables, such as the numbers of children in family, the
numbers of car accident on the certain road on different days, or
the numbers of students taking basics of statistics course are the
results of counting and thus these are discrete variables. Typically,
a discrete variable is a variable whose possible values are some or
all of the ordinary counting numbers like 0, 1, 2, 3, . . . . As a
definition, we can say that a variable is discrete if it has only a
countable number of distinct possible values. That is, a variable is
discrete if it can assume only a finite numbers of values or as many
values as there are integers.
11
Continuous variables
Quantities such as length, weight, or temperature can in principle
be measured arbitrarily accurately. There is no invisible unit.
Weight may be measured to the nearest gram, but it could be
measured more accurately, say to the tenth of a gram. Such a
variable, called continuous, is intrinsically different from a discrete
variable.
12
Organization of the data
Observing the values of the variables for one or more people or
things yield data. Each individual piece of data is called an
observation and the collection of all observations for particular
variables is called a data set or data matrix. Data set are the values
of variables recorded for a set of sampling units.
Data is presented in a matrix form (data matrix). All the values of
particular variable is organized to the same column; the values of
variable forms the column in a data matrix. Observation, i.e.
measurements collected from sampling unit, forms a row in a data
matrix. Consider the situation where there are k numbers of
variables and n numbers of observations (sample size is n). Then
the data set should look like
13
where xij is a value of the jth variable collected from ith observation,
i=1, 2, . . . , n and j =1, 2, . . . , k.
14
Describing data by tables and graphs
Qualitative variable
The number of observations that fall into particular class (or
category) of the qualitative variable is called the frequency (or
count) of that class. A table listing all classes and their frequencies
is called a frequency distribution.
In addition of the frequencies, we are often interested in the
percentage of a class. We find the percentage by dividing the
frequency of the class by the total number of observations and
multiplying the result by 100. The percentage of the class,
expressed as a decimal, is usually referred to as the relative
frequency of the class.
15
A table listing all classes and their relative frequencies is called a
relative frequency distribution. The relative frequencies provide
the most relevant information as to the pattern of the data. One
should also state the sample size, which serves as an indicator of
the creditability of the relative frequencies. Relative frequencies
sum to 1 (100%).
A cumulative frequency (cumulative relative frequency) is obtained
by summing the frequencies (relative frequencies) of all classes up
to the specific class. In a case of qualitative variables, cumulative
frequencies makes sense only for ordinal variables, not for nominal
variables.
16
Frequency distribution
Shops Employ Employee Shops
ee (value) (frequency)
1 6
3 2
2 6
4 1
3 10
6 3
4 10
7 1
5 7 10 2
6 3 How many shop have 3 employee? 2
7 3 How many shop have 4 employee? 1
How many shop have 6 employee? 3
8 6 How many shop have 7 employee? 1
9 4 How many shop have 10 employee? 2
17
Relative frequency and percentage
Employee Absolute Relative % relative
frequenc frequency frequency
y
3 2 2/9=0,22 22,2
The sum of
4 1 1/9=0,11 11,1
percentage relative
6 3 3/9=0,34 33,3 frequency is equal to
7 1 1/9=0,11 11,1 100
10 2 2/9=0,22 22,2
Tot n=9 1,00 100,0
3 shops have 2 employee (abs. freq.)
They are the 22% of the shops.
18
Comparison between frequency distribution
Campania Catalogna
Employee Abs. Employee Abs.
Freq. Freq.
3 2 3 4
4 1 4 4
6 3 5 11
7 1 6 15
10 2 8 8
Tot n=9 10 6
Tot n=48
Looking at the absolute frequencies we conclude that the shops
with 3 employee are less in Campania that in Catalogna (2 vs 4)
and the some for 4 employee (1 vs 4).
This kind of comparison is WRONG! 19
Comparison between frequency distribution
Campania Catalogna
Employ Abs. Rel. Employee Abs. Rel. Freq.
ee Freq . Freq. Freq. perc.
perc.
3 4 8,3
3 2 22,2
4 4 8,3
4 1 11,1
5 11 22,9
6 3 33,3
6 15 31,3
7 1 11,1
8 8 16,7
10 2 22,2
10 6 12,5
Tot n=9 100
Tot n=48 100,0
In terms of relative percentage frequencies the shops with 3
employee are 22,2% in Campania and only 8,3% in Catalogna
12
20
Cumulative frequencies
Employee No. shops Cumulative Cumulative
(values) (freq.) frequency
frequencies
are defined
3 2 2
only if the
4 1 3
values of the
6 3 6
variable
7 1 7
(ordinal) are
10 2 9
sorted.
How many shops have max 3 employee? 2
How many shops have max 4 employee? 2+1
How many shops have max 6 employee? 2+1+3
How many shops have max 7 employee? 2+1+3+1
How many shops have max 10 employee? 2+1+3+1+2
13
21
Cumulative frequencies
Employee Abs. CUM. Rel. Perc. cum.
Freq. nj Abs. cum. Freq. Pj
Freq. Nj Freq. Fj
3 2 2 0,22 22,2
4 1 3 0,33 33,3
6 3 6 0,67 66,6
7 1 7 0,78 77,7
10 2 9 1,00 100,0
Looking at perc. cum. freq. Pj,
We can see that 66,6% of the shops (i.e. 2/3) has
a number of employee equal or less than 6.
14
22
If the discrete variable can have a lot of different values or the
quantitative variable is a continuous variable, then the data must
be grouped into classes (categories) before the table of frequencies
can be formed. The main steps in a process of grouping
quantitative variable into classes are:
(a) Find the minimum and the maximum values variable have in the
data set
(b) Choose intervals of equal length that cover the range between
the minimum and the maximum without overlapping. These are
called class intervals, and their end points are called class limits.
(c) Count the number of observations in the data that belongs to
each class interval. The count in each class is the class frequency.
(c) Calculate the relative frequencies of each class by dividing the
class frequency by the total number of observations in the data.
23
Frequency distribution in classes
We choose 3 classes:
Revenues
(sorted
- <= 250
values) - From 250 to 350
180
- > 350
200 Revenues Abs.
205 classes Freq.
270 (0 – 250] 3
280 (250 – 350] 4
340 Oltre 350 2
350 How many shops have a revenue
500 less then 250?
600
17
24