Introduction (Basic) Statistics
Introduction (Basic) Statistics
Contents
1. Introduction
1.1 Definitions and Classification of Statistics
1.2 Stages in Statistical Investigation
1.3 Definition of Some Statistical terms
1.4 Use, Scope, Limitation & Misuse of Statistics
1.4.1 Uses of Statistics
1.4.2 Scope of Statistics
1.4.3 Limitations of Statistics
1.4.4 Misuses of Statistics
1.5 Scales of Measurement
1.6 Introduction to Methods of Data Collection
1.6.1 Methods of Primary Data Collection
1.6.2 Methods of Secondary Data Collection
INTRODUCTION
“Statistical thinking will one day be as necessary for efficient citizenship as the ability to read
and write.”
H. [Link]
In the modern world of computers and information technology, the importance of statistics is
very well recognized by all the disciplines. Statistics has originated as a science of statehood and
found applications slowly and steadily in Agriculture, Economics, Commerce, Biology,
Medicine, Industry, planning, education and so on. In the meantime, there is no other human
walk of life, where statistics cannot be applied. Hence, we are constantly being bombarded with
statistics and statistical information.
The word “Statistics” and “Statistical” are all derived from Latin word status which means a
political state. Statistics is defined differently by different authors over a period of time. In the
olden days statistics was confined to only state affairs but in modern days it embraces almost
every sphere of human activity. Therefore, a number of old definitions, which was confined to
1|Page minilikderse@[Link]
Introduction to statistics
narrow field of enquiry, were replaced by more definitions, which are much more comprehensive
and exhaustive. Let us examine different way of defining statistics by different authors and
Dictionaries.
Despite these, the word statistics can have two different senses while we use it as plural and
singular verb. Statistics in singular verb is defined as the branch of mathematics that deals with
the collection, organization, analysis, and interpretation of numerical data. Statistics is especially
useful in drawing general conclusions about a set of data from a sample of the data. But statistics
in plural verb is defined as numerical data which has been collected, classified, and interpreted.
Based on the usage of statistical data statistics is defined broadly in to two mutually exclusive
groups so called Descriptive statistics and inferential statistics.
Descriptive statistics are used to describe the basic features of the data in a study. They provide
simple summaries about the sample and the measures. Together with simple graphics analysis,
they form the basis of virtually every quantitative analysis of data. Various techniques that are
commonly used are classified as:
2|Page minilikderse@[Link]
Introduction to statistics
Example-1: Of 350 randomly selected people in the town of Addis Ababa 280 people had the
last name Abebe. An example of descriptive statistics is the following statement: "80% of these
people have the last name Abebe."
Example-2: On the last 3 Sundays, Hiwot Car salesman sold 2, 1, and 0 new cars respectively.
An example of descriptive statistics is the following statement: "Hiwot averaged 1 new car sold
for the last 3 Sundays."
These are both descriptive statements because they can actually be verified from the information
provided.
Inferential statistics (statistical induction) comprise the use of statistics to make inferences or
conclusions and determine the relationships concerning about some unknown aspect of a
population parameters based on the data which are obtained from the sample. That is., inferential
statistics aim to make inferences from the data in order to make conclusions that go beyond the
data.
Example-3: Of 350 randomly selected people in the town of Addis Ababa 280, Ethiopia,
people had the last name Abebe. An example of inferential statistics is the following statement:
"80% of all people living in Ethiopia have the last name Abebe."
We have no information about all people living in Ethiopia, just about the 350 living in Addis
Ababa. We have taken that information and generalized it to talk about all people living in
Ethiopia.
Example-4: On the last 3 Sundays, Hiwot. Car salesman sold 2, 1, and 0 new cars respectively.
An example of inferential statistics is the following statements: "Hiwot never sells more than 2
cars on a Sunday."
3|Page minilikderse@[Link]
Introduction to statistics
Although this statement is true for the last 3 Sundays, we do not know that this is true for all
Sundays.
Before we deal with statistical investigation, let us see what statistical data mean. Each and every
numerical data can’t be considered as statistical data unless it possesses the following criteria.
These are:
A statistician should be involved at all the different stages of statistical investigation. This
includes formulating the problem, and then collecting, organizing and classifying, presenting,
analyzing and interpreting of statistical data. Let’s see each stage in detail
I. Formulating the problem: First research must emanate if there is a problem. At this
stage the investigator must be sure to understand the problem and then formulate it in
statistical term. Clarify the objectives very carefully. Ask as many questions as
necessary because “An approximate answer to the right question is worth a great deal
more than a precise answer to the wrong question.” -The first golden rule of
applied mathematics-
Therefore, the first stage in any statistical investigation should be to:
4|Page minilikderse@[Link]
Introduction to statistics
In this section, we will define those terms which will be used most frequently. These are:
5|Page minilikderse@[Link]
Introduction to statistics
Data set: Facts or figures collected for a particular study. Each value in the data set is called data
value or datum.
Raw Data: Data sheets are where the data are originally recorded. Original data are called raw
data. Data sheets are often hand drawn, but they can also be printouts from database programs
like Microsoft Excel.
Population: The totality of all subjects with certain common characteristics that are
being studied in a specified time and place.
Sample: Is a portion of a population which is selected using some technique of sampling. Sample
must be representative of the population so that it must be selected by any of the developed
technique.
Sampling: Is the process of selecting units (e.g., people, organizations) from a population of
interest so that by studying the sample we may fairly generalize our results back to the
population from which they were chosen. There are two types of sampling techniques namely
random sampling technique and non-random sampling technique.
Random sampling technique or probability sampling technique gives a non- zero chance for all
elements to be included in the sample. In other words, there is no personal bias regarding the
selection. The five common random sampling techniques are:
Quota sampling
6|Page minilikderse@[Link]
Introduction to statistics
Convenience sampling
Volunteer sampling
Purposive sampling
Haphazard sampling
Snow ball sampling etc…
Sample size: The number of elements or observation to be included in the sample.
Variable: Is an attribute of a physical or an abstract system which may change its value while it
is under observation. Variables are often specified according to their type and intended use and
hence variable can be classified in to two namely qualitative and quantitative variables.
7|Page minilikderse@[Link]
Introduction to statistics
8|Page minilikderse@[Link]
Introduction to statistics
Of social sciences, economics leans most heavily on statistical methods for analyses of data
relating to micro as well as to macro economics, from demand analyses up to national income
analyses.
9|Page minilikderse@[Link]
Introduction to statistics
Statistics is only, one of the methods of studying a problem: Statistical method does
not provide complete solution of the problems because problems are to be studied
taking the background of the countries culture, philosophy or religion into
consideration. Thus the statistical study should be supplemented by other evidences.
At times, association or relationship between two or more variables is studied in
statistics, but such a relationship does not indicate ‘cause and effect’ relationship. It
simply shows the similarity or dissimilarity in the movement of the two variables. In
such cases, it is the user who has to interpret the results carefully, pointing out the
type of relationship obtained.
Source of data not given: At times, the source of data not given. In the absence of
the source, the reader does not know how far the data are reliable. Further, if he
wants refer to the original source, he is unable to do so.
Defective data: Another misuse is that sometimes one gives inaccurate data. This
may be done knowingly in order to defend one’s position or to prove a particular
point. This apart, the definition used to denote a certain phenomenon may be
defective.
Unrepresentative sample: In statistics, several times one has to conduct a survey,
which necessitates to choose a sample from a given population or universe. The
sample may turn out to be unrepresentative of the universe. One may choose a
sample just on the basis of convenience. He may collect the desired information
10 | P a g e minilikderse@[Link]
Introduction to statistics
from either his friends or nearby respondents in his neighborhood even though
such respondents do not constitute a representative sample.
Inadequate sample: Earlier, we have seen a sample that is unrepresentative of the
universe is a major misuse of statistics. This apart, at times one may conduct a
survey based on an extremely inadequate sample. For example, in a city we may
find that are 100,000 households. When we have to conduct a household survey,
we may take a sample of merely 100 households comprising only 0.1 percent of
the universe. A survey based on such a small sample may not yield right
information.
Unfair comparison: An important misuse of statistics is making unfair
comparisons from the data collected. For instance, one may construct an index of
production choosing the base year where the production was much less. Then he
may compare the subsequent year’s production from this low base. Such a
comparison will undoubtedly give a rosy picture of the production though the
reality it is not so. Another source of unfair comparison could be when one makes
absolute comparisons instead of relative ones. An absolute comparison of two
figures say, production or export, may show a good increase, but in relative terms
it may turn out to be very negligible. Another example of unfair comparison is
when the population of the two cities is different; a comparison of over all death
and death rate by a particular disease is attempted.
Unwarranted conclusion: Another misuse of statistics may be on account of
unwanted conclusions. This may be as a result of making false assumptions. For
example, while making projection of population in the next five years, one may
assume a lower rate of growth though the past two years indicate otherwise.
Another source of unwarranted conclusion may be the use of wrong average.
Suppose in a series there are extreme values, one is very high while the other is
too low, such as 800 and 38. The use of an arithmetic average in such a case may
give a wrong idea.
Confusion of correlation and causation: In statistics, several times one has to
examine the relationship between two variables. A close relationship between the
two variables may not establish a cause-and-effect relationship in the sense that
11 | P a g e minilikderse@[Link]
Introduction to statistics
one variable is the cause and the other variable is the effect. It should be taken as
something that is measures degrees of association rather than try to find out casual
relationship.
Suppression of unfavorable results: Another wrong use of statistics may be on
account of suppressing results that are unfavorable to the organization or an
individual. Revealing such results may expose the concerned organization or
individual in bad light. In order to avoid such a situation, one may be attempted to
hide unfavorable, though true, facts emerging from statistical study.
Mistake in arithmetic: Finally, one may come across certain mistakes in
calculation or in the application of wrong formula. This human error may result in
grossly wrong figures, leading to wrong conclusion.
1.5 Scales of Measurement
Normally, when one hears the term measurement, they may think in terms of measuring the
length of something (i.e. the length of a piece of wood) or measuring a quantity of something
(i.e. a cup of flour). This represents a limited use of the term measurement. In statistics, the term
measurement is used more broadly and is more appropriately termed scales of measurement.
Scales of measurement refer to ways in which variables or numbers are defined and categorized.
Each scale of measurement has certain properties which in turn determine the appropriateness for
use of certain statistical analyses. The four scales of measurement are nominal, ordinal, interval,
and ratio.
Nominal scale allows for only qualitative classification (categorical data). That is, it
can be measured only in terms of whether the individual items belong to some
distinctively different categories, but we cannot quantify or even rank order those
categories. For example, all we can say is that two individuals are different in terms
of variable A (e.g., they are of different race), but we cannot say which one "has
more" of the quality represented by the variable. Typical examples of nominal
variables are gender, race, color, etc.
Ordinal scale allows us to rank or order the items we measure in terms of which has
less and which has more of the quality represented by the variable, but still they do
not allow us to say "how much more." A typical example of an ordinal variable is the
12 | P a g e minilikderse@[Link]
Introduction to statistics
Interval scale allows us not only to rank or order the items that are measured, but also
to quantify and compare the sizes of differences between them. For example,
temperature, as measured in degrees Fahrenheit or Celsius, constitutes an interval
scale. We can say that a temperature of 40 degrees is higher than a temperature of 30
degrees, and that an increase from 20 to 40 degrees is twice as much as an increase
from 30 to 40 degrees.
Ratio scale is very similar to interval variables; in addition to all the properties of
interval variables, they feature an identifiable absolute zero point, thus they allow for
statements such as is two times more than y. typical examples of ratio scales are
measures of time or space. For example, as the Kelvin temperature scale is a ratio
scale, not only can we say that a temperature of 200 degrees is higher than one of 100
degrees; we can correctly state that it is twice as high. Interval scales do not have the
ratio property. Most statistical data analysis procedures do not distinguish between
the interval and ratio properties of the measurement scales.
13 | P a g e minilikderse@[Link]
Introduction to statistics
1.6
Contents
We have already explained what it means by statistical data. Numerical facts or measurements
obtained in the course of enquiry in to a phenomenon, marked by uncertainty, constitute
statistical data. The statistical data may be already available or may have to be collected by an
investigator or an agency. Data termed primary when the reference is to data collected for the
first time by the investigator and is termed secondary when the data are taken from records or
data already available.
In primary data collection, you collect the data yourself using methods such as interviews and
questionnaires. The key point here is that the data you collect is unique to you and your research
and, until you publish, no one else has access to it. There are many methods of collecting
primary data and the main methods include:
Advantages:
Relatively cheap.
14 | P a g e minilikderse@[Link]
Introduction to statistics
No interviewer bias.
Disadvantages:
Design problems
Advantages:
15 | P a g e minilikderse@[Link]
Introduction to statistics
Disadvantages:
Time consuming.
Geographic limitations.
Can be expensive.
16 | P a g e minilikderse@[Link]
Introduction to statistics
Secondary data analysis can be literally defined as second-hand analysis and is the analysis of
data or information that was either gathered by someone else (e.g., researchers, institutions, other
NGOs, etc.) or for some other purpose than the one currently being considered, or often a
combination of the two.
Some of the sources of secondary data are government document, official statistics, technical
report, scholarly journals, trade journals, review articles, reference books, research institutes,
universities, libraries, library search engines, computerized data base and world wide web (
).
Data availability
Level of observation
Quality of documentation
Data quality control
Outdated data
17 | P a g e minilikderse@[Link]
Introduction to statistics
“I've come loaded with statistics, for I've noticed that a man can't prove
anything without statistics”
M. TWAIN
This chapter introduces tabular and graphical methods commonly used to summarize both
qualitative and quantitative data. Tabular and graphical summaries of data can be obtained in
annual reports, newspaper articles and research studies. Everyone is exposed to these types of
presentations, so it is important to understand how they are prepared and how they will be
interpreted.
Modern statistical software packages provide extensive capabilities for summarizing data and
preparing graphical presentations. MINITAB, SPSS and STATA are three packages that are
widely available.
Some of basic terms that are most frequently used while we deal with frequency distribution are
the following:
Lower Class Limits are the smallest number that can belong to the different class.
Upper Class Limits are the largest number that can belong to the different classes.
18 | P a g e minilikderse@[Link]
Introduction to statistics
Class Boundaries are the number used to separate classes, but without the gaps created
by class limits.
Class midpoints are the midpoints of the classes. Each class midpoint can be found by
adding the lower class limit to the upper class limit and dividing the sum by 2.
Class width is the difference between two consecutive lower class limits or two
consecutive lower class boundaries.
The categorical frequency distribution is used for data which can be placed in specific categories
such as nominal or ordinal level data. For example, data such as political affiliation, religious
affiliation, or major field of study would use categorical frequency distribution.
The major components of categorical frequency distribution are class, tally and frequency.
Moreover, even if percentage is not normally a part of a frequency distribution, it will be added
since it is used in certain types of graphical presentations, such as pie graph.
1. You have to identify that the data is in nominal or ordinal scale of measurement
2. Make a table as show below
19 | P a g e minilikderse@[Link]
Introduction to statistics
Example 2.1: Twenty-five army inductees were given a blood test to determine their blood type.
The data set is given as follows:
A B B AB O
O O B AB B
B B O A O
A O O O AB
AB A O B A
Solution:
When the data are numerical interested of categorical, the range of data is small and each class is
only one unit, this distribution is called an ungrouped frequency distribution.
The major components of this type of frequency distributions are class, tally, frequency and
cumulative frequency. The steps are almost similar with that of categorical frequency
distribution.
Cumulative frequencies are used to show how many values are accumulated up to and including
a specific class.
20 | P a g e minilikderse@[Link]
Introduction to statistics
Example 2.2: The following data represent the number of days of sick leave taken by each of 50
workers of a company over the last 6 weeks.
2 0 0 5 8 3 4 1 0 0
7 1 7 1 5 4 0 4 0 1
8 9 7 0 1 7 2 5 5
4 3 3 0 0 2 5 1 3
0 2 4 5 0 5 7 5 1
1 0 2
Solution:
A. Since this data set contains only a relatively small number of distinct or different
values, it is convenient to represent it in a frequency table which presents each distinct
value along with its frequency of occurrence.
B. Since 12 of the 50workers had no days of sick leave, the answer is 50-12=38
C. The answer is the sum of the frequencies for values 3, 4 and 5 that is 4+5+8=17
21 | P a g e minilikderse@[Link]
Introduction to statistics
When the range of the data is large, the data must be grouped in which each class has more than
one unit in width. While we construct this frequency distribution, we have to follow the
following steps.
I. Use Struge’s rule. That is, where is the number of class and
way. If you fail to calculate by Struge’s rule, this method is more appropriate.
When we choose the number of classes, we have to think about the following criteria
The classes must be mutually exclusive. Mutually exclusive classes have non
overlapping class limits so that values can’t be placed in to two classes.
The classes must be continuous. Even if there are no values in a class, the class
must be included in the frequency distribution. There should be no gaps in a
frequency distribution. The only exception occurs when the class with a zero
frequency is the first or last. A class width with a zero frequency at either end
can be omitted with out affecting the distribution.
The classes must be equal in width. The reason for having classes with equal
width is so that there is not a distorted view of the data. One exception occurs
when a distribution is open-ended. i.e., it has no specific beginning or end values.
4. Find the class width by dividing the range by the number of classes
22 | P a g e minilikderse@[Link]
Introduction to statistics
Note that: Round the answer up to the nearest whole number if there is a reminder. For
instance, and
5. Select the starting point as the lowest class limit. This is usually the lowest score
(observation). Add the width to that score to get the lower class limit of the next class.
Keep adding until you achieve the number of desired classes calculated in step 3.
6. Find the upper class limit; subtract unit of measurement from the lower class limit of
the second class in order to get the upper class limit of the first class. Then add the width
to each upper class limit to get all upper class limits.
Unit of measurement: Is the next expected value. For instance, 28, 23, 52, and then the
unit of measurement of this data set is one. Because take one datum arbitrarily, say 23,
then the next value will be 24. Therefore, . If the data set is 24.12, 30,
21.2, then give priority to the datum with more decimal place. Take 24.12 and guess the
next possible value. It is 24.13. Therefore,
Note that: U=1 is the maximum value of unit of measurement and is the value when we
don’t have a clue about the data.
8. Tally the data and write the numerical values for tallies in the frequency column
9. Find cumulative frequency. We have two type of cumulative frequency namely less than
cumulative frequency and more than cumulative frequency. Less than cumulative
frequency is obtained by adding successively the frequencies of all the previous classes
including the class against which it is written. The cumulate is started from the lowest to
the highest size. More than cumulative frequency is obtained by finding the cumulate
total of frequencies starting from the highest to the lowest class.
23 | P a g e minilikderse@[Link]
Introduction to statistics
For example, the following frequency distribution table gives the marks obtained by 40
students:
The above table shows how to find less than cumulative frequency and the table shown
below shows how to find more than cumulative frequency.
Example 2.3: Consider the following set of data and construct the frequency distribution.
11 29 6 33 14 21 18 17 22 38
31 22 27 19 22 23 26 39 34 27
Steps
3.
4.
24 | P a g e minilikderse@[Link]
Introduction to statistics
5. Select starting point. Take the minimum which is 6 then add width 6 on it to get the next
class LCL.
6. Upper class limit. Since unit of measurement is one. . So 11 is the UCL of the
first class. Therefore, is the first class
8. 9 and 10
An important variation of the basic frequency distribution uses relative frequencies, which are
easily found by dividing each class frequency by the total of all frequencies. A relative frequency
25 | P a g e minilikderse@[Link]
Introduction to statistics
distribution includes the same class limits as a frequency distribution, but relative frequencies are
used instead of actual frequencies. The relative frequencies are sometimes expressed as percent.
Relative frequency distribution enables us to understand the distribution of the data and to
compare different sets of data.
We have discussed the techniques of classification and tabulation that help us in organizing the
collected data in a meaningful fashion. However, this way of presentation of statistical data does
not always prove to be interesting to a layman. Too many figures are often confusing and fail to
convey the massage effectively.
One of the most effective and interesting alternative way in which a statistical data may be
presented is through diagrams and graphs. There are several ways in which statistical data may
be displayed pictorially such as different types of graphs and diagrams.
Pie chart can used to compare the relation between the whole and its components. Pie chart is a
circular diagram and the area of the sector of a circle is used in pie chart. Circles are drawn with
radii proportional to the square root of the quantities because the area of a circle is .
26 | P a g e minilikderse@[Link]
Introduction to statistics
To construct a pie chart (sector diagram), we draw a circle with radius (square root of the total).
The total angle of the circle is .
These angles are made in the circle by mean of a protractor to show different components. The
arrangement of the sectors is usually anti-clock wise.
Example2.4: The following table gives the details of monthly budget of a family. Represent
these figures by a suitable diagram.
27 | P a g e minilikderse@[Link]
Introduction to statistics
2 Bar Charts
The bar graph (simple bar chart, multiple bar chart and stratified or stacked bar chart) uses
vertical or horizontal bins to represent the frequencies of a distribution. While we draw bar chart,
we have to consider the following two points. These are
Make the bars the same width
Make the units on the axis that are used for the frequency equal in size
A simple bar chart is used to represents data involving only one variable classified on spatial,
quantitative or temporal basis. In simple bar chart, we make bars of equal width but variable
length, i.e. the magnitude of a quantity is represented by the height or length of the bars.
Following steps are undertaken in drawing a simple bar diagram:
28 | P a g e minilikderse@[Link]
Introduction to statistics
Draw two perpendicular lines one horizontally and the other vertically at an appropriate
place of the paper.
Take the basis of classification along horizontal line (X-axis) and the observed variable
along vertical line (Y-axis) or vice versa.
Marks signs of equal breath for each class and leave equal or not less than half breath in
between two classes.
Finally, marks the values of the given variable to prepare required bars.
Example 2.5: Draw simple bar diagram to represent the profits of a bank for 5 years.
Multiple bar charts are used two or more sets of inter-related data are represented (multiple
bar diagram facilities comparison between more than one phenomenon). The technique of
29 | P a g e minilikderse@[Link]
Introduction to statistics
simple bar chart is used to draw this diagram but the difference is that we use different
shades, colors, or dots to distinguish between different phenomena.
Example 2.6: Draw a multiple bar chart to represent the import and export of Canada (values
in $) for the years 1991 to 1995.
Stratified (Stacked) Bar Chart is used to represent data in which the total magnitude is divided
into different or components. In this diagram, first we make simple bars for each class taking
30 | P a g e minilikderse@[Link]
Introduction to statistics
total magnitude in that class and then divide these simple bars into parts in the ratio of various
components. This type of diagram shows the variation in different components within each class
as well as between different classes. Sub-divided bar diagram is also known as component bar
chart.
Example 2.7: The table below shows the quantity in hundred kgs of Wheat, Barley and Oats
produced on a certain form during the years 1991 to 1994. Draw stratified
bar chart.
Solution: To make the component bar chart, first of all we have to take year wise total
production.
31 | P a g e minilikderse@[Link]
Introduction to statistics
3 Histogram
Histogram is a special type of bar graph in which the horizontal scale represents classes of data
values and the vertical scale represents frequencies. The height of the bars correspond to the
frequency values, and the drawn adjacent to each other (without gaps).
We can construct a histogram after we have first completed a frequency distribution table for a
data set. The axis is reserved for the class boundaries.
32 | P a g e minilikderse@[Link]
Introduction to statistics
7.0
6.0
Frequency
5.0
4. 0
3.0
2.0
1.0
0.0 35.5
5.5 11.5 17.5 23.5 29.5 41.5
Class boundaries
Relative frequency histogram has the same shape and horizontal ( ) scale as a histogram,
but the vertical ( ) scale is marked with relative frequencies instead of actual frequencies.
4 Frequency Polygon
A frequency polygon uses line segment connected to points located directly above class midpoint
values. The heights of the points correspond to the class frequencies, and the line segments are
extended to the left and right so that the graph begins and ends on the horizontal axis with the
same distance that the previous and next midpoint would be located.
33 | P a g e minilikderse@[Link]
Introduction to statistics
7.0
6.0
5.0
4.0
3.0
2.0
5 Ogive Graph
An Ogive (pronounced as “oh-jive”) is a line that depicts cumulative frequencies, just as the
cumulative frequency distribution lists cumulative frequencies. Note that the Ogive uses class
boundaries along the horizontal scale, and graph begins with the lower boundary of the first class
and ends with the upper boundary of the last class. Ogive is useful for determining the number of
values below some particular value. There are two type of Ogive namely less than Ogive and
more than Ogive. The difference is that less than Ogive uses less than cumulative frequency and
Example 2.10: Take the data in example 2.3 and draw less than and more than Ogive
34 | P a g e minilikderse@[Link]
Introduction to statistics
20
Less than Ogive
15
10
“The way to make sense out of raw data is to compare and contrast, to understand difference.”
G. BATESON
Researchers are often interested in defining a value that best describes some attribute of the
population. The best way to reduce a set of data and still retain part of the information is to
summarize the set with a single value. Therefore, measures of central tendency are one of
descriptive statistics.
35 | P a g e minilikderse@[Link]
Introduction to statistics
Our objective in this chapter is to develop measures that can be used to summarize a data set.
When describing, exploring, and comparing data sets, the following characteristics are usually
extremely important.
1. Center: A representative or average value that indicated where the middle of the
data set is located.
2. Variation: A measure of the amount that the data values vary among themselves
3. Distribution: The nature or shape of the distribution of the data (such as bell-
shaped, uniform, or skewed)
4. Outliers: Sample values that lie very far away from the vast majority of the other
sample values.
5. Time: Changing characteristics of the data over time.
The above five characteristics are so important that they might be better remembered by using a
mnemonic for the first letters CVDOT, such as “Computer Viruses Destroy Or Terminate”. A
measure of center is a value at the center or middle of a data set. The three major objective of
measure of central tendency are
Let the symbol (read “ sub ”) denotes any of the value assumed by a
variable . The letter in , which can stand for any of the numbers is called a
subscript, or index. Clearly any letter other than , such as , could have been used as
well.
The symbol is used to denote the sum of all the from to ; by definition,
36 | P a g e minilikderse@[Link]
Introduction to statistics
1.
2. Where is
4. , where b is constant
5.
6.
Example 3.1: Express each of the following by using the summation notations
A.
B.
C.
D.
E.
Solutions:
A. B. C.
37 | P a g e minilikderse@[Link]
Introduction to statistics
D. E.
An average is a value that is typical, or representative, of a set of data since such typical values
tend to lie centrally within a set of data arranged according to magnitude, averages are also
called measures of central tendency.
Several types of averages can be defined, the most common being the arithmetic mean, the
median, the mode, the geometric mean, and the harmonic mean. Each has advantages and
disadvantages, depending on the data and the intended purpose.
The (arithmetic) mean is generally the most important of all numerical measurements used to
The arithmetic mean of asset of values is the measure of center found by adding the values and
dividing the total by the number of values. The mean is denoted by (pronounced “ ”) if
38 | P a g e minilikderse@[Link]
Introduction to statistics
the data set is a sample from a larger population; if all values of the population are used, then we
population; if all values of the population are used, then we denote the mean by (lower case
Greek mu).
Note that sample statistics are usually represented by English letters, such as , and population
parameters are usually represented by Geek letters, such as .
Solution:
Frequently one uses the term average synonymously with arithmetic mean. Strictly speaking,
however, this is incorrect since there are averages other than arithmetic mean.
, where
mean is
The only new concept in calculating mean for grouped data is that find mid-points or each class
and label it as .
39 | P a g e minilikderse@[Link]
Introduction to statistics
Midpoint =
Example 3.4: Take the data in example 2.3 in chapter two. The midpoint of each class is given as
follow:
Midpoint (
8.5 2 17
14.5 2 29
20.5 7 143.5
26.5 4 106
32.5 3 97.5
38.5 2 77
Total
The algebraic sum of the deviation of a set of numbers from their arithmetic mean is zero.
That is,
The sum of the square of deviations of a set of numbers from the mean is always the
40 | P a g e minilikderse@[Link]
Introduction to statistics
If numbers have mean , numbers have mean … numbers have mean , then
the mean of all numbers is , combined mean. That is, a weighted arithmetic
Or
Weighted Mean is a special type arithmetic mean and it will be functional when values have its
41 | P a g e minilikderse@[Link]
Introduction to statistics
Example 3.5: A teacher attaches weights 2 to homework 3 to mid term exam and 5 for final
exam. If a student score 90, 50 and 60 for HM, MT and FE, respectively,
what is his/ her average academic performance?
2 90
3 50
5 60
Geometric mean is often used in business and economics for finding average rates of change
average rates of growth, or average ratio.
Example 3.6: A price of a commodity increased by 5%, 8% and 77% for the three consecutive
years. What was the average yearly price increase?
Solution:
42 | P a g e minilikderse@[Link]
Introduction to statistics
It is useful in averaging ratios and percentages and determining rates of increase and
decrease
as follows
It gives less weight to large items and more to small ones than does the arithmetic
average. It is because of this reason that geometric mean is never larger than the
arithmetic mean, on occasions it may turn out to be same as the arithmetic mean, but it is
usually smaller
It is based on each and every item of the series
It is rigidly defined
It is difficult to understand
It is difficult to compute
It can’t be computed when there are both negative and positive values in a series
It is biased for small values as it gives more weight to small values
Its computation becomes difficult especially when the values of items are large
43 | P a g e minilikderse@[Link]
Introduction to statistics
No value can be negative. The harmonic mean is often used as a measure of central tendency for
data sets consisting of rates of change, such as speeds.
Example 3.7: Four students drive from Jimma to Addis Ababa at a speed of 40 km/hr. Because
they need to reach statistics class on time, they return at a speed of 60
km/hr. What is their average speed for the round trip?
Solution:
It is difficult to understand
It is difficult to calculate
It can’t be computed when one or more items are zero
It gives largest weight to smallest items. Hence, it is not useful for analyzing the
economic data
44 | P a g e minilikderse@[Link]
Introduction to statistics
A well known inequality concerning arithmetic, geometric, and harmonic means for any set of
equal.
The mode of a set observation is that value which occurs with the greatest frequency; that is, it is
the most common value. The value that occurs most frequently in a data set is called mode. The
mode may not exist; even if it does exist it may not be unique.
Example 3.8: The set 2, 2, 5, 7, 8, 9, 9, 9, 10, 10, and 11 has mode 9. The set 3, 5, 8, 10, 12, 15,
and 16 has no mode. The set 2, 3, 4, 4, 4, 5, 5, 7, 7, 7, and 9 has two
modes, 4 and 7 and is called bimodal.
A distribution having only one mode is called unimodal. In the case of grouped data where a
frequency curve has been constructed to fit the data, the mode will be the value (or values) of
corresponding to the maximum point (or points) on the curve. This value of is sometimes
denoted by (Pronounced as “ hat”). From the frequency distribution the mode can be obtained
Where
is lower class boundary of the modal class. Modal class is a class which contains the mode and
has the highest frequency. is excess of modal frequency over frequency of next lower class.
is excess of modal frequency over frequency of the next higher class. is size of the modal
class interval
45 | P a g e minilikderse@[Link]
Introduction to statistics
Example 3.8: Take the data in example 2.3 and find the mode of the frequency distribution.
Solution:
The median of a data set is the measure of center that is the middle value when the original data
values are arranged in order of increasing (or decreasing) magnitude. The median is often
To find the median, first sort the values, and then follow one of these two procedures:
1. If the number of values is odd, the median is the number located in the exact middle of
the list.
46 | P a g e minilikderse@[Link]
Introduction to statistics
2. If the number of values is even, the median if found by computing the mean of the two
middle numbers.
Where
is the lower class boundary of the median class. Median class is a class which accommodates
observation. is the total of all frequencies. is the sum of frequencies of all classes
lower than the median class. is the frequency of the median class and is the width of
median class.
Example 3.9: Take the data from example 2.3 and calculate the median.
The median is used when one must find the center or middle value of a data set
The median is used when one must determine whether the data values fall in tot the upper
half or lower half of the distribution
47 | P a g e minilikderse@[Link]
Introduction to statistics
Note that:
MEASURES OF DISPERSION
This chapter deals with the second most important characteristics of a distribution
so called variation which belongs to CVDOT. Without knowing something about
how data is dispersed, measures of central tendency may be misleading. Measures of
dispersion provide a more complete picture.
4.1 Introduction and Objectives of Measuring Dispersion
48 | P a g e minilikderse@[Link]
Introduction to statistics
The term dispersion is generally used in two senses. Firstly, dispersion refers to the variations of
the items among themselves. If the value of all the items of a series is the same, there will be no
variation among different items of a series; the more will be the dispersion. Secondly, dispersion
refers to the variation of the items around an average. If the difference between the value of
items and the average is large, the dispersion will be high and on the other hand if the difference
between the value of the items and averaging is small, the dispersion will be low. Thus,
dispersion is defined as scatteredness or spreadness of the individual items in a given series.
The measures of dispersion are helpful in statistical investigation. Some of the main objectives
of dispersion are as under:
49 | P a g e minilikderse@[Link]
Introduction to statistics
the same units and of the same averaging size. These measures are not suitable for
comparing the variability in two distributions having variables expressed in different
units.
50 | P a g e minilikderse@[Link]
Introduction to statistics
Where R=Range, L= Largest value in the series, S= smallest value in the series
Example 4.1: five students obtained the following marks in statistics: . Find the
Range and coefficient of range
Solution: Here,
51 | P a g e minilikderse@[Link]
Introduction to statistics
Coefficient of Range =
Example 4.2: Find out range and coefficient of range of the following series
Merits of Range
It is simple to understand
It is easy to compute
It is well-defined
It helps in giving an idea about the variation, just by giving the lowest value and the
greatest value of variable
Demerits:
52 | P a g e minilikderse@[Link]
Introduction to statistics
Inter-quartile range and quartile deviation are other measures of dispersion. The difference
between the upper quartile and lower quartile is called inter-quartile range.
Symbolically,
The inter-quartile ranges covers dispersion of middle 50% of the items of the series. Quartile
deviation, also called semi-inter-quartile range is half of the difference between the upper and
lower quartile. That is, half of the inter-quartile range. Its formula as:
The relative measure of quartile deviation also called the coefficient of quartile deviation is
defined as:
Example 4.3: Find inter-quartile deviation, quartile deviation and coefficient of quartile
deviation from the following data.
Solution: First arrange the data in ascending order. 25, 18, 20, 24, 27, 28, 30
53 | P a g e minilikderse@[Link]
Introduction to statistics
Example 4.4: Find inter-quartile range, quartile deviation and coefficient of quartile deviation
from the following data
Marks 2 3 4 5 6 7 8 9
No. Of students 10 11 12 13 5 12 7 5
Solution:
2 10 10
3 11 21
4 12 33
5 13 46
6 5 51
7 12 63
8 7 70
9 5 75=N
Total N=75
54 | P a g e minilikderse@[Link]
Introduction to statistics
Any applied statistician who has analyzed a number of sets of real data is likely to have come
across outliers. The intuitive definition of an outlier would be ‘an observation deviates so much
from other observations as to arouse suspicious that it was generated by different mechanism.’
That is, outliers are observations that are distinct from the main body of the data and are
incompatible with the rest of data. These values may be genuine observations from individuals
with very extreme levels of the variable.
A simple approach to detect outlier is that print the data and visually checks them by eye. This is
suitable of the number of observations is not too large and if the potential outlier is much lower
than or higher than the rest of the data.
When the number of observations gets larger and larger, we can check the presence of outlier by
the 1.5 IQR rule. The steps to identify outliers are presented as follows:
considered as outlier
1 2 5 5 7 8 10 11 11 12 15 25
55 | P a g e minilikderse@[Link]
Introduction to statistics
Solution: The first step is arranging the data in ascending order then let us calculate the first and
third quartile
Therefore, the observation less than -4 and greater than 20 are considered as outlier. That is, 25 is
outlier.
By using the concept of 1.5IQR rule, we can draw box plot which is used to give five-number
summaries. Five-number summaries contains minimum, quartile one, median, quartile three and
maximum.
1. Notice that you must have ordered data before you can find the Five – Number
Summaries.
2. Find the median first. It’s the middle point
3. Then find the quartiles, Q1 and Q3 and the 1.5 IQR outlier limits
4. Draw a “box" from Q1 to Q3 with bars at Q1, Q3 and the median. (In the below
example the box is horizontal, but it could also be vertical.)
56 | P a g e minilikderse@[Link]
Introduction to statistics
Solution:
Merits of
It is simple to understand
It is easy to compute
It is well-defined
It helps in studying the middle 50% item in the series
It is not affected by the extreme items
It is useful in the case of open-ended
Demerits of
57 | P a g e minilikderse@[Link]
Introduction to statistics
by
Where denotes the absolute value of the deviation. Generally, arithmetic mean and
median are used in calculating mean deviation. So, stands for the average used for calculating
. That is,
Where is the frequency of . is the midpoint in the case of grouped frequency distribution
The relative measure of mean deviation, also called the coefficient of mean deviation is obtained
by dividing mean deviation by the particular average used in computing mean deviation. Thus,
58 | P a g e minilikderse@[Link]
Introduction to statistics
Exercise: Find all coefficients of mean deviations for the following frequency distribution:
Merits of
It is simple to understand
It is easy to compute
It is well-defined
It is based on all observations
It is not unduly affected by the extreme items
It can be calculated by using any average
Demerits of
Note that: of all the mean deviations taken about different averages or any arbitrary value, the
mean deviation about the median has the smallest value.
4.3.4 The Variance, the Standard Deviation and the Coefficient of Variation
Standard deviation is the most important and widely used measure of dispersion. It was first
used by Karl Pearson in 1893. The standard deviation of a statistical data is defined as the
59 | P a g e minilikderse@[Link]
Introduction to statistics
positive square root of the mean of the squared deviations of items from the mean of the series
under consideration.
Where is the frequency of . is the midpoint in the case of grouped frequency distribution
For comparing two or more series for variability, the corresponding relative measure, called
coefficient of variation is calculated. This measure is defined as:
Just as it is possible to calculate combined mean of two or more groups, similarly the combined
or pooled standard deviation of two or more groups can be calculated. The combined standard
60 | P a g e minilikderse@[Link]
Introduction to statistics
Where is the group sample size and is the variance of the group.
Example 4.7: Two samples of size 100 and 150, respectively, have means 50 and 60 and
standard deviations 5 and 6. Find the mean and standard of the combined
sample of size 250.
Solution: Given
Exercise: Find the shortcut formula to find standard deviation, mean deviation and combined
deviation.
In certain cases, mean and standard deviation are calculated by using one or two incorrect values
of the variable. Just as we can correct an incorrect mean, similarly, there is a procedure of
correcting an incorrect standard deviation.
1. Find out incorrect sum of square values of the variable. That is,
61 | P a g e minilikderse@[Link]
Introduction to statistics
2. Find corrected . To do so, we subtract the square of the incorrect item from incorrect
is approximately equal to
is approximately equal to
The standard deviation of the first natural numbers can be found from the following
formula:
For example, the standard deviation of the first 5 natural numbers is given as:
62 | P a g e minilikderse@[Link]
Introduction to statistics
remains unaffected
Example 4.8: If the value of standard deviation in moderately symmetrical distribution is 24,
find the value of mean deviation and quartile deviation.
Example 4.9: If the mean and the standard deviation of 25 boys’ weight are 50 and 5,
63 | P a g e minilikderse@[Link]
Introduction to statistics
Solution:
the interval
When the observations are grouped into classes, all observations in a class are equal to the
midpoint of the class. This introduces some error known as grouping error. Sheppard suggests a
Merits of
It is simple to understand
It is well-defined
It is based on all items
It is suitable for further algebraic treatment
It has sampling stability
It is very useful in the study of “Tests of Significant”
Demerits of
It is easy to calculate
It is unduly affected by extreme values
64 | P a g e minilikderse@[Link]
Introduction to statistics
A standard score for sample vale in a data set is obtained by the mean of the data set from the
value and dividing the result by the standard deviation of the data set. Basically, the standard
score (z-score) tells us how many standard deviations a specific value is above or below the
mean value of the data set. That is, the z-score is the number of standard deviations the data
value falls above (positive z-score) or below (negative z-score) the mean for the data set.
Note that: The Z-score is affected by an outlying value in the data set because the outlier (very
small or very large value) directly affects the value of the mean and the standard values by
eliminating the unit of measurement.
Example 4.10: what is the Z-score for the value of 14 in the following sample data set?
3 8 6 14 4 12 7 10
Solution:
= 8, SD = 3.8173 thus, Z =
The data value of 14 is located 1.57 standard deviations above the mean 8 because the
z-score is positive.
65 | P a g e minilikderse@[Link]
Introduction to statistics
Moments are statistical tools used in statistical investigation. The moments of a distribution are
the arithmetic mean of the various powers of the deviations of items from some number. In our
course, we shall use it in the study of Skewness and Kurtosis of statistical distribution.
Where
Moments about the origin for grouped frequency distribution and for ungrouped frequency
distribution
Where is the frequency of . is the midpoint in the case of grouped frequency distribution
Note that: ,
Moments about the mean for grouped frequency distribution and for ungrouped frequency
distribution
Where is the frequency of . is the midpoint in the case of grouped frequency distribution
66 | P a g e minilikderse@[Link]
Introduction to statistics
Moments about any arbitrary constant for grouped frequency distribution and for ungrouped
frequency distribution
Example 4.11: Find the first four moments about the mean for the following individual series
3 6 8 10 18
Solution:
[Link]
1 3 -6 36 -216 1296
2 6 -3 9 -27 81
3 8 -1 1 -1 1
4 10 1 1 1 1
5 18 9 81 729 6561
Tota
l
Now
67 | P a g e minilikderse@[Link]
Introduction to statistics
Exercise: Find the relationship between moments about any arbitrary constant and moments
Skewness
Skewness is a measure of symmetry, or more precisely, the lack of symmetry or departure from
absolute terms by taking the difference between arithmetic mean and mode. Therefore, the
formula for absolute skewness is given as:
If the value of arithmetic mean is greater than mode, Skewness is positive and if the value of
mode is greater than mean, the skewness is negative. This absolute measure of skewness is not
free from unit of measurement. Further, the difference between arithmetic mean and mode in
absolute terms may be higher in one situation, although the frequency curves of two distributions
are similarly skewed. It is because of this reason; it is desirable to have a measure that can be
68 | P a g e minilikderse@[Link]
Introduction to statistics
directly used for comparisons. If the absolute differences are expressed in relation to some
measure of spread, the resultant measure will be a relative measure of skewness. The two most
commonly used measures are Karl Pearson’s coefficient skewness and Bowley’s coefficient of
skewness.
It has been indicated above that the distance between the arithmetic mean and the mode can be
used as a measure of skewness. However, since the measure of skewness should be a pure
number, free from units of measurement, we define
When the distribution is symmetrical, mean, median and mode coincide and therefore, the
coefficient of skewness will be zero. When the coefficient of skewness is positive, it indicates
that the distribution is positively skewed and when the coefficient is negative, the distribution is
negatively skewed. The numerical value say 0.9 or 0.3 etc, indicates the degree of skewness.
Therefore, Karl Pearson’s formula for skewness indicates both direction as well as the extent of
skewness.
Note that: In moderately skewed distributions the averages have the following relationship.
Hence, for moderately skewed distribution the following formula is used for skewness.
69 | P a g e minilikderse@[Link]
Introduction to statistics
This measure is also referred to as the quartile measure skewness and the value of the coefficient
lies between .
Kurtosis
Kurtosis in Greek language mean ‘bulginess’, it measures the flatness of the curve. Three terms
are used for indicating flatness, mesokurtic stands for a normal curve, leptokurtic for a peaked
curve and platykurtic for a curve less peaked than normal.
70 | P a g e minilikderse@[Link]
Introduction to statistics
Note that:
ELEMENTARY PROBABILITY
W. BAGEHOT
The notion that chance, or probability, can be treated numerically is relatively recent. Indeed, for
most of recorded history it was felt that what occurred in life was determined by forces that were
beyond one’s ability to understand. It was only during the first half of the 17th century, near the
end of Renaissance, that people become curious about the world and the laws governing its
operation. Among the curious were the gamblers.
5.1 Introduction
A cynical person once said, “The only two sure things are death and taxes.” This philosophy no
doubt arose because so much in people’s lives is affected by chance.
Probability as a general concept can be defined as the chance of an event occurring. Most people
are familiar with probability from observing or plying games of chance, such as card games or
lotteries. Probability is the basis of inferential statistics.
71 | P a g e minilikderse@[Link]
Introduction to statistics
Sample Space: It is the set of all possible outcomes of a probability experiment and
denoted by .
Event: It is a subset of sample space (contains one or more outcomes which are in the
sample space) and is defined for a particular purpose. An event can be one outcome or
more than one outcome. Simple event is an event having only single outcome. Compound
event consisting of one or more outcomes or simple events. Event is denoted by capital
letters such as A, B, F etc.
Mutually exclusive events: Suppose you have two events, say A and B. if these events
have no common sample point(s) or do not occur simultaneously, then the two events are
called mutually exclusive events.
Equally-likely events: It is a situation where the probability of the occurrence of one
event as likely as the other event. That is, they must have equal probability of occurrence.
Exhaustive events: It is a satiation where the events contain all elements based on the
definition of the events.
Union of events: The union of two events A and B, denoted by , consists of all
72 | P a g e minilikderse@[Link]
Introduction to statistics
If the choices can’t be performed together then the number of ways in which you can make a
choice in different ways. For example, if there are two way of bus to
voyage Awassa from Addis Ababa and three railways, then collectively we have
Multiplication Rule
Suppose that two experiments are to be performed. Then if experiment one can result in any one
of possible outcomes and if for each outcome of experiment one there are possible outcomes
of experiment two, then together there are possible outcomes of the two experiments.
Example 5.1: A small community consists of 10 women, each of whom has three children. If
one women and one of her children are to be chosen as mother and child
of the year, how many different choices are possible?
Solution:
By regarding the choice of the woman as the outcome of the first experiment and the subsequent
choice of one of her children as the outcome of the second experiment, we see from the basic
The generalized basic principle of multiplication is that if experiments that are to be performed
are such that the first one may result in any of possible outcomes there are possible
outcomes the second experiment, and if for each of the possible outcomes of the first two
experiments there are possible outcomes of the third experiment, and if …, then there is a
total of .
73 | P a g e minilikderse@[Link]
Introduction to statistics
Example 5.2: How many different 7-place license plates are possible if the first 3 places are to
be occupied by letters and the final 4 by numbers?
Example 5.3: In the above example, how many license plates would be possible if repetition
among letters or numbers were prohibited?
Permutation
How many different ordered arrangements of letters are possible? By direct enumeration
permutation. Thus, there are six possible permutations of a set of 3 objects. This result could also
have been obtained from the basic principle, since the first object in the permutation can be any
of the 3, the second object in the permutation can then be chosen any of the remaining 2, and the
third object in the permutation is then chosen the remaining one. Thus there are
possible permutations.
Suppose now that we have objects. Reasoning, similar to that we have just used for the 3 letter
Example 5.4: A class of stat 173 consists of 6 men and 4 women. An examination is given, and
the students are ranked according to their performance. Assume that no
two students obtain the same score.
74 | P a g e minilikderse@[Link]
Introduction to statistics
Solution:
B. As there are possible rankings of the men among themselves and possible rankings
of the women among themselves, it follows from the basic principle that the two groups
arrange themselves; it follows the basic principle that the two groups arrange themselves
We shall now determine the number of permutations of a set of objects when certain of the
objects are indistinguishable from each other. Then the formula is:
Different permutations of objects, of which are alike are alike, …, are alike.
Example 5.5: How many different letter arrangements can be formed using the letter PEPPER?
Generally, if we are asked to arrange objects among objects, then we will have the following
total arrangements
75 | P a g e minilikderse@[Link]
Introduction to statistics
Example 5.6: Suppose a business man has a choice of five locations in which to establish his
business. He wishes to arrange only the top three locations. How many
different ways can he arrange them?
Solution:
Combination
We are often interested in determining the number of different groups of objects that could be
formed from a total of objects. A selection of objects without regard to order is called a
combination. That is, combinations are used when the order or arrangement is not important. The
the formula
Example 5.7: From a group of 5 women and 7 men, how many different committees consisting
of 2 women and 3 men can be performed? What if 2 of the men are
feuding and refuse to serve on the committee together?
Solution:
As there are possible groups of 2 women, and possible groups of 3 me, it follows from
76 | P a g e minilikderse@[Link]
Introduction to statistics
Possible committees consisting of 2 women and 3 men. On the other hand, if 2 of the men refuse
to serve on the committee together, then, as there are possible group of 3 men not
containing either of the 2 feuding men and groups of 3 men containing exactly 1 of the
feuding men, it follows that there are groups of 3 men not containing
both of the feuding men. Since there are ways to choose the 2 women, it follows that in this
If a procedure has different simple events, each with an equal chance of occurring, and event A
77 | P a g e minilikderse@[Link]
Introduction to statistics
Example 5.8: Toss a fair coin once and find the probability of the occurrence of head
Solution: Since the sample space is finite i.e., either head or tail and the outcomes are
equally-likely
Example 5.9: For a card drawn from an ordinary deck, find the probability of getting a queen.
If one of the assumptions stated above is violated, the classical approach no longer valid
Frequentist Approach
very large, an event is observed to occur in of these, then the probability of an event is or
conduct an experiment a large number of times, and count the number of times event A actually
78 | P a g e minilikderse@[Link]
Introduction to statistics
Example 5.10: Suppose a coin was tossed 1000 times and the result was 587 tails. The relative
frequency of tails is . Another 1000 tosses lead to 511 tails. Then the
we obtain a sequence of numbers, which gets closer and closer to the number
defined as the probability of a trial in a single toss.
Therefore,
Axiomatic Approach
Both the classical and frequentist approaches have serious drawbacks, the first because the words
“equally likely” are vague and the second because the “large number” involved is vague.
Because of these difficulties, statisticians have been led to an axiomatic approach of probability.
79 | P a g e minilikderse@[Link]
Introduction to statistics
Subjective Approach
A probability derived from an individual's personal judgment about whether a specific outcome
is likely to occur. Subjective probabilities contain no formal calculations and only reflect the
subject's opinions and past experience.
Subjective probabilities differ from person to person. Because the probability is subjective, it
contains a high degree of personal bias. An example of subjective probability could be asking
Arsenal fan, before the football season starts, the chances of Arsenal winning the world
champions. While there is no absolute mathematical proof behind the answer to the example,
fans might still reply in actual percentage terms, such as the Arsenal having a 95% chance of
winning the world champions.
Rule 3: For the empty set, i.e. the impossible event has probability zero.
Rule 6:
Example 5.11: Suppose we toss two coins and suppose that each of the four points in the
sample space is equally likely and hence has
probability . Let E is the event that the first coin falls head, and F is the
event that the second coin falls heads.
80 | P a g e minilikderse@[Link]
Introduction to statistics
In words, this is saying that the probability that both A and B occur is equal to the probability
that A occurs times the probability that B occurs given that has occurred. We call the
conditional probability of B given A, i.e. the probability that B will occur given that A has
occurred.
Example 5.12: A jar contains black and white marbles. Two marbles are chosen without
replacement. The probability of selecting a black marble and then a white
marble is 0.34, and the probability of selecting a black marble on the first
draw is 0.47. What is the probability of selecting white marble on the second
draw, given that the first marble drawn was black?
Solution:
Example 5.13: The probability that it is Friday and that a student is absent is 0.03. Since there
are 5 schooldays in a week, the probability that it is Friday is 0.2. What is
the probability that a student is absent given that today is Friday?
Solution:
81 | P a g e minilikderse@[Link]
Introduction to statistics
It often happens that the knowledge that a certain event E has occurred has no effect on the
probability that some other event F has occurred, that is, that . One would
expect that in this case, the equation would also be true. If these equations are
true, we might say the F is independent of E.
Definition: Two events E and F are independent if both E and F have positive probability and if
Example 5.14: Suppose that we roll a pair of fail dice, so each of the 36 possible out come is
equally likely. Let A denotes the event that the first die lands on 3, let C be the
event that the sum of the dice is 7
A. Since is the event that the first die lands on 3 and the second on 5, we see that
and
82 | P a g e minilikderse@[Link]
Introduction to statistics
Exercise:
1. A box contains 3 white balls and 5 black balls. We extract 2 balls from the box,
one after the other. Give a sample space for this experiment and the probabilities
of the elementary events of this sample space.
PROBABILITY DISTRIBUTION
“Free yourself from the rigid conduct of tradition and open yourself to the new
forms of probability.”
HANS BENDER
A random variable is a variable whose values are determined by chance. A random variable can
be either continuous or discrete. By define the random variable; we can assign the number to a
random variable.
Example 6.1: Suppose we are about to learn the sexes of the three children of a certain family.
The sample space of this experiment consists of the following 8 outcomes.
83 | P a g e minilikderse@[Link]
Introduction to statistics
The outcomes means, for instance that the youngest child is a girl, the next youngest is a
boy, and the oldest is a boy. Suppose that each of these 8 possible outcomes is equally likely,
and so each has probability 1/8. If we let x denote the number of female children in this family,
then the value of x is determined by the outcomes of the experiment. That is, x is a random
variable whose value will be . We now determine the probabilities that x will equal
each of these four values. Since X will equal 0 if the outcome is we see that
Similarly,
A probability distribution consists of the values a random variable can assume and the
corresponding probabilities of the values. The probabilities are determined theoretically or by
observation. The probability distribution can be denoted by is used to represent the
probability that is equal to . The sum of the probability distribution is one. That is,
Example 6.2: Suppose we toss a coin three times, the sample space is represented as
and if the random
variable for the number of heads.
84 | P a g e minilikderse@[Link]
Introduction to statistics
Probability
Example 6.3: Suppose that is a random variable that takes on one of the value . If
and . What is ?
Example 6.4: A sales women has scheduled two appointments to sell encyclopedias. She feels
her first appointments will lead to a sale with probability 0.3. She also
feels that the second will lead to a sale with probability 0.6 and that the
results from the two appointments are independent. What is the probability distribution
of , the number of sales made?
Solution: The random variable can take on any of the value . It will equal 0 if neither
appointment leads to a sale, and so
85 | P a g e minilikderse@[Link]
Introduction to statistics
The random variable will equal 1 either if there is a sale on the first and not on the second
appointment or if there is no sale on the first and one sale on the second appointment. Since these
two events are disjoint, we have
Finally, the random variable will equal 2 if both appointments result in sales; thus
1. The sum of the probabilities of all the events in the sample space must equal 1; that
is,
2. The probability of each event in the sample space must be between or equal to 0
and 1. that is,
86 | P a g e minilikderse@[Link]
Introduction to statistics
Where is probability density function in the case of discrete random variable its name will
change to probability mass function (pmf).
Example 6.5: Find the expected value of the following random variable
0 1 2 3 4
0.18 0.34 0.23 0.21 0.04
Solution:
Note that: The expected value of a random variable is the same as with the mean of a
random variable
Suppose that we are given a discrete random variable random variable along with its probability
mass function, and that we want to compute the expected value of some function of , say .
How can we accomplish this? One way is follows: Since is determined from the probability
mass function of . Once we have determined the probability mass function of we can
Example 6.6: Let X denote a random variable that takes on any of the values -1, 0, 1 with
respective probability
87 | P a g e minilikderse@[Link]
Introduction to statistics
Compute
Hence
The expected value of a random variable is also referred to as the mean or the first
By definition
Example 6.7: The following are the annual income of 7 men and 7 women residents of a certain
community.
88 | P a g e minilikderse@[Link]
Introduction to statistics
Men Women
33.5 24.2
25.0 19.5
28.6 27.4
41.0 28.6
30.5 32.2
29.6 22.4
32.8 21.6
Suppose that a woman and a man randomly chosen. Find the expected value of the sum
of their incomes.
Solution: Let be the man’s income and Y is the woman’s income. Since is equally likely to
Similarly,
89 | P a g e minilikderse@[Link]
Introduction to statistics
That is, the expected value of the sum of their incomes is approximately $ 56,700.
That is,
Example 6.8: The return from a certain investment is a random variable X with probability
distribution.
90 | P a g e minilikderse@[Link]
Introduction to statistics
we have
Therefore,
Properties of Variance
91 | P a g e minilikderse@[Link]
Introduction to statistics
3. The square root of the is called the standard deviation of , and we denote it by
. That is,
Many types of probability problems have only two outcomes, or they can be reduced to two
outcomes. For example, when a coin is tossed, it can land heads or tails.
1. Each trial can have only two outcomes or outcomes that can be reduced to two
outcomes.
4. The probability of a success must remain the same for each trial
The out comes of a binomial experiment and the corresponding probabilities of these outcomes
are called a binomial distribution
The probability mass function of a binomial random variable having parameter (n, p) is given by
Example 6.9: Five fair coins are flipped. If the outcomes are assumed independent, find
the probability of the number of heads obtained
92 | P a g e minilikderse@[Link]
Introduction to statistics
Hence,
Example 6.10:
and
Solution:
A.
B.
93 | P a g e minilikderse@[Link]
Introduction to statistics
Solution: The number of effective screws in the shipment of size 1000 is a Binomial random
variable with parameters , . Hence, the expected number of
defective screws is and the variance of
the number of detective screws is
A discrete probability distribution that is useful when is large and is small and when the independent
variable occurs over a period of time is called the Poisson distribution, name for Simeon D. Poisson
(1781-1840). In addition to being used for the stated conditions (i.e. is large, is small, and the
variable occur over a period of time), the Poisson distribution can be used when a density of items is
distributed over a given area or volume, such as the number of plants growing per acre of woods or the
number of defects in a given length of videotape.
Solution:
Both the expected value and the variance of a Poisson random variable are equal to . That is, we have the
following. If is a Poisson random variable with parameter , ; then
94 | P a g e minilikderse@[Link]
Introduction to statistics
Example6.13: Suppose the average number of accidents occurring weekly on a particular high way is
equal to 1.2. Approximate the probability that there is at least one accident this
week.
Solution: Let denote the number of accidents because it is reasonable to suppose that there are a large
number of cars passing along the high way, each having a small probability of being involved in
an accident, the number of such accidents should be approximately a Poisson random variable.
That is, if denotes the number of accidents that will occur this week, then is approximately
Poisson random variable with mean value . The desired probability is now obtained as
follows.
Therefore, there is approximately a 70% chance that there will be at least one accident
this week.
We can approximate Binomial distribution to Poisson distribution if is large and is too small.
Thus, the approximately Poisson distribution has a parameter.
Example 6.14: Suppose that items produced by a certain machine are independently
defective with probability [Link] is the Poisson approximation for this
probability?
Solution: If we let denote the number of defective items, then is a Binomial random variable
with parameters and . Thus the desired probability is
95 | P a g e minilikderse@[Link]
Introduction to statistics
Thus, even in this case, where is equal to 10 (which is not that large) and is equal
to 0.1 (which is not that small), the Poisson approximation to the Binomial
probability is quite accurate.
Since must assume some value, it follows that the total area under the density curve must
equal 1. Also, since the area under the graph of the probability density function between points
and is the same regardless of whether the end points and are themselves included. That is,
96 | P a g e minilikderse@[Link]
Introduction to statistics
The most important type of random variable is the normal random variable. The probability
density function of a normal random variable is determined by two parameters: the expected
value and the standard deviation of . We designate these values as and , respectively.
And
The normal probability density function is a bell-shaped density curve that is symmetric about
the value ; its variability is measured by . The larger is, the more variability there is in the
curve.
Since the probability density function of a normal random variable is symmetric about its
expected value ; it follows that is equally likely to be on either side of . That is,
Not all bell-shaped symmetric density curves are normal. The normal density curves are
specified by a particular formula:
A normal random variable having mean value 0 and standard deviation 1 is called a standard
normal variable, and its density curve is called the standard normal curve. The letter represents
97 | P a g e minilikderse@[Link]
Introduction to statistics
Once the values are transformed by using the above formula, they are called value is
actually the number of standard deviations that a particular value is a way from the mean.
1. Between 0 and any value: Look up the value in the table to get the area
2. In any tail
98 | P a g e minilikderse@[Link]
Introduction to statistics
General procedure is
Example 6.15: Find the area under the normal distribution curve between and
Since table gives the area between 0 and any value to the right of 0, one need look up the
value in the table. Find 2.3 in the left column and 0.04 in the top row. The value where the column and
99 | P a g e minilikderse@[Link]
Introduction to statistics
2.2
2.3 0.4904
A.
B.
Solution:
0
0 1.5
100 | P a g e minilikderse@[Link]
Introduction to statistics
0 0.8
0 0.8 0 0 0.8
A.
B.
Solution:
0 1 2
101 | P a g e minilikderse@[Link]
Introduction to statistics
0 1 2 0 1
0 2 2
-1.5 0 2.5
0 2.5
-1.50
-1.50 2.5
102 | P a g e minilikderse@[Link]
Introduction to statistics
Let be a normal random variable with mean and standard deviation . We can determine
Example 6.18: IQ examination scores for sixth-graders are normally distributed with mean value
100 and standard deviation 14.2.
Solution: Let denote the score of a randomly chosen student. We compute probabilities
103 | P a g e minilikderse@[Link]
Introduction to statistics
A.
Or equivalently,
Therefore,
Normal distribution is a limiting case of the Binomial probability distribution under the
following condition:
104 | P a g e minilikderse@[Link]
Introduction to statistics
De-Moivre provide that under the above two conditions, the distribution of standard Binomial
variate
tends to the distribution of standard normal distribution. If and are nearly equal (i.e., is
nearly 0.5), then the normal approximation is surprisingly good even for small values of .
It has been proved that this variate tends to be a standard normal variate if
105 | P a g e minilikderse@[Link]
Introduction to statistics
Chi-square Distribution:
The square of a standard normal variable is called a chi-square variate with one degree of
freedom. Thus if is a random variable following normal distribution with mean and standard
of freedom.
this is the sum of the square of independent standard normal variates, follows chi-square
Chi-square distribution has a number of applications. Some of which are listed below
Student’s distribution
106 | P a g e minilikderse@[Link]
Introduction to statistics
It is often the case that one wants to calculate the size of sample needed to obtain a certain level
of confidence in survey results. Unfortunately, this calculation requires prior knowledge of the
be conducted so that a reasonable estimate of this critical population parameter can be made. If
such a preliminary sample is not made, but confidence intervals for the population mean are to
be constructing using an unknown , then the distribution known as the Student t distribution can
be used.
First, a little history about this curious name. William Gosset (1876-1937) was a Guinness
Brewery employee who needed a distribution that could be used with small samples. Since the
Irish brewery did not allow publication of research results, he published under the pseudonym of
Student. We know that large samples approach a normal distribution. What Gosset showed was
that small samples taken from an essentially normal population have a distribution characterized
by the sample size. The population does not have to be exactly normal, only unimodal and
basically symmetric. This is often characterized as heap-shaped or mound shaped.
5. The variance is greater than one, but approaches one from above as the sample size
107 | P a g e minilikderse@[Link]
Introduction to statistics
108 | P a g e minilikderse@[Link]