0% found this document useful (0 votes)
12 views108 pages

Introduction (Basic) Statistics

The document provides an introduction to statistics, covering definitions, classifications, and the importance of statistical methods across various fields. It outlines the stages of statistical investigation, including problem formulation, data collection, organization, presentation, analysis, and interpretation. Additionally, it discusses the uses, scope, limitations, and potential misuses of statistics, emphasizing the need for careful application and understanding of statistical principles.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
12 views108 pages

Introduction (Basic) Statistics

The document provides an introduction to statistics, covering definitions, classifications, and the importance of statistical methods across various fields. It outlines the stages of statistical investigation, including problem formulation, data collection, organization, presentation, analysis, and interpretation. Additionally, it discusses the uses, scope, limitations, and potential misuses of statistics, emphasizing the need for careful application and understanding of statistical principles.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOC, PDF, TXT or read online on Scribd

Introduction to statistics

Contents

1. Introduction
1.1 Definitions and Classification of Statistics
1.2 Stages in Statistical Investigation
1.3 Definition of Some Statistical terms
1.4 Use, Scope, Limitation & Misuse of Statistics
1.4.1 Uses of Statistics
1.4.2 Scope of Statistics
1.4.3 Limitations of Statistics
1.4.4 Misuses of Statistics
1.5 Scales of Measurement
1.6 Introduction to Methods of Data Collection
1.6.1 Methods of Primary Data Collection
1.6.2 Methods of Secondary Data Collection
INTRODUCTION

“Statistical thinking will one day be as necessary for efficient citizenship as the ability to read
and write.”

H. [Link]

In the modern world of computers and information technology, the importance of statistics is
very well recognized by all the disciplines. Statistics has originated as a science of statehood and
found applications slowly and steadily in Agriculture, Economics, Commerce, Biology,
Medicine, Industry, planning, education and so on. In the meantime, there is no other human
walk of life, where statistics cannot be applied. Hence, we are constantly being bombarded with
statistics and statistical information.

1.1 Definitions and Classification of Statistics

The word “Statistics” and “Statistical” are all derived from Latin word status which means a
political state. Statistics is defined differently by different authors over a period of time. In the
olden days statistics was confined to only state affairs but in modern days it embraces almost
every sphere of human activity. Therefore, a number of old definitions, which was confined to

1|Page minilikderse@[Link]
Introduction to statistics

narrow field of enquiry, were replaced by more definitions, which are much more comprehensive
and exhaustive. Let us examine different way of defining statistics by different authors and
Dictionaries.

The American Heritage Dictionary defines statistics as “The mathematics of collection,


organization and interpretation of numerical data, especially the analyses of population
characteristics by inference from sampling.”
The Merriam-Webster’s collegiate Dictionary defines statistics as “A branch of
mathematics dealing with the collection, analyses, interpretation, and presentation of
masses of numerical data.”
The former American Statistical Association president Jon Kettering define statistics as
“… the science of learning from data …It presents exciting opportunities for those who
work as professional statisticians. Statistics is essential for the proper running of
government, central to decision making in industry and a core component of modern
educational curricula at all level.”

Despite these, the word statistics can have two different senses while we use it as plural and
singular verb. Statistics in singular verb is defined as the branch of mathematics that deals with
the collection, organization, analysis, and interpretation of numerical data. Statistics is especially
useful in drawing general conclusions about a set of data from a sample of the data. But statistics
in plural verb is defined as numerical data which has been collected, classified, and interpreted.

Based on the usage of statistical data statistics is defined broadly in to two mutually exclusive
groups so called Descriptive statistics and inferential statistics.

Descriptive statistics are used to describe the basic features of the data in a study. They provide
simple summaries about the sample and the measures. Together with simple graphics analysis,
they form the basis of virtually every quantitative analysis of data. Various techniques that are
commonly used are classified as:

Graphical description in which we use graphs to summarize data.

2|Page minilikderse@[Link]
Introduction to statistics

Tabular description in which we use tables to summarize data.

Summary statistics in which we calculate certain values to summarize data.

Example-1: Of 350 randomly selected people in the town of Addis Ababa 280 people had the
last name Abebe. An example of descriptive statistics is the following statement: "80% of these
people have the last name Abebe."

Example-2: On the last 3 Sundays, Hiwot Car salesman sold 2, 1, and 0 new cars respectively.
An example of descriptive statistics is the following statement: "Hiwot averaged 1 new car sold
for the last 3 Sundays."

These are both descriptive statements because they can actually be verified from the information
provided.

Inferential statistics (statistical induction) comprise the use of statistics to make inferences or
conclusions and determine the relationships concerning about some unknown aspect of a
population parameters based on the data which are obtained from the sample. That is., inferential
statistics aim to make inferences from the data in order to make conclusions that go beyond the
data.

Example-3: Of 350 randomly selected people in the town of Addis Ababa 280, Ethiopia,
people had the last name Abebe. An example of inferential statistics is the following statement:
"80% of all people living in Ethiopia have the last name Abebe."

We have no information about all people living in Ethiopia, just about the 350 living in Addis
Ababa. We have taken that information and generalized it to talk about all people living in
Ethiopia.

Example-4: On the last 3 Sundays, Hiwot. Car salesman sold 2, 1, and 0 new cars respectively.
An example of inferential statistics is the following statements: "Hiwot never sells more than 2
cars on a Sunday."

3|Page minilikderse@[Link]
Introduction to statistics

Although this statement is true for the last 3 Sundays, we do not know that this is true for all
Sundays.

1.2 Stages in Statistical Investigation

Before we deal with statistical investigation, let us see what statistical data mean. Each and every
numerical data can’t be considered as statistical data unless it possesses the following criteria.
These are:

The data must be aggregate of facts


They must be affected to a marked extent by a multiplicity of causes
They must be estimated according to reasonable standards of accuracy
The data must be collected in a systematic manner for predefined purpose
The data should be placed in relation to each other

A statistician should be involved at all the different stages of statistical investigation. This
includes formulating the problem, and then collecting, organizing and classifying, presenting,
analyzing and interpreting of statistical data. Let’s see each stage in detail

I. Formulating the problem: First research must emanate if there is a problem. At this
stage the investigator must be sure to understand the problem and then formulate it in
statistical term. Clarify the objectives very carefully. Ask as many questions as
necessary because “An approximate answer to the right question is worth a great deal
more than a precise answer to the wrong question.” -The first golden rule of
applied mathematics-
Therefore, the first stage in any statistical investigation should be to:

Get a clear understanding of the physical background to the situation


under study;
Clarify the objectives;

4|Page minilikderse@[Link]
Introduction to statistics

Formulate the objective in statistical terms


II. Proper collection of data: In order to draw valid conclusions, it is important to have
‘good’ data. Data are gathered with aim to meet predetermine objectives. In other
words, the data must provide answers to problems. The data itself form the foundation
of statistical analyses and hence the data must be carefully and accurately collected. In
section 1.6, we will see the methods of data collection.
III. Organization and classification of data: In this stage the collected data organized in a
systematic manner. That means the data must be placed in relation to each other. The
classification or sorting out of data is, by itself, a kind of organization of data.
IV. Presentation of data: The purpose of putting the organized data in graphs, charts and
tables is two-fold. First, it is a visual way to look at the data and see what happened and
make interpretations. Second, it is usually the best way to show the data to others.
Reading lots of numbers in the text puts people to sleep and does little to convey
information.
V. Analyses of data: Is the process of looking at and summarizing data with the intent to
extract useful information and develop conclusions. Data analysis is closely related to
data mining, but data mining tends to focus on larger data sets, with less emphasis on
making inference, and often uses data that was originally collected for a different
purpose. In this stage different types of inferential statistical methods will apply. For
instance, hypothesis testing such as test of association.
VI. Interpretation of data: Interpretation means drawing valid conclusions from data
which form the basis of decision making. Correct interpretation requires a high degree
of skill and experience.
Note that: Analyses and interpretation of data are the two sides of the same
coin.

1.3 Definition of Some Statistical Terms

In this section, we will define those terms which will be used most frequently. These are:

Data: Facts or figures from which the conclusion can be drawn.

5|Page minilikderse@[Link]
Introduction to statistics

Data set: Facts or figures collected for a particular study. Each value in the data set is called data
value or datum.

Raw Data: Data sheets are where the data are originally recorded. Original data are called raw
data. Data sheets are often hand drawn, but they can also be printouts from database programs
like Microsoft Excel.

Population: The totality of all subjects with certain common characteristics that are
being studied in a specified time and place.

Sample: Is a portion of a population which is selected using some technique of sampling. Sample
must be representative of the population so that it must be selected by any of the developed
technique.

Sampling: Is the process of selecting units (e.g., people, organizations) from a population of
interest so that by studying the sample we may fairly generalize our results back to the
population from which they were chosen. There are two types of sampling techniques namely
random sampling technique and non-random sampling technique.

Random sampling technique or probability sampling technique gives a non- zero chance for all
elements to be included in the sample. In other words, there is no personal bias regarding the
selection. The five common random sampling techniques are:

Simple Random sampling


Systematic Random sampling
Stratified Random sampling
Cluster Random sampling
Multi-stage sampling
Non-random sampling technique is mostly known as non-probability sampling techniques and
in this case not all elements of a population have a known chance of inclusion or if some
outcomes have a zero chance of being selected as a sample. The most familiar examples of non-
random sampling techniques are

Quota sampling

6|Page minilikderse@[Link]
Introduction to statistics

Convenience sampling
Volunteer sampling
Purposive sampling
Haphazard sampling
Snow ball sampling etc…
Sample size: The number of elements or observation to be included in the sample.

Parameter: Any measure computed from the data of a population.

Example-5: Populations mean and population standard deviation

Statistic: Any measure computed from the sample.

Example-6: sample mean , sample standard deviation

Survey: A collection of quantitative information about members of a population when no special


control is exercised over any of the factors influencing the variable of interest.

Sample survey: A survey that include only a portion of the population.

Census: A collection of information about every member of a population

Sample survey has the following advantages over census

Sample survey saves time and cost


Has great accuracy
Avoid wastage of material

Variable: Is an attribute of a physical or an abstract system which may change its value while it
is under observation. Variables are often specified according to their type and intended use and
hence variable can be classified in to two namely qualitative and quantitative variables.

A quantitative variable is naturally measured as a number for which meaningful


arithmetic operations make sense. Examples: Height, age, crop yield, GPA, salary,
temperature, area, air pollution index (measured in parts per million), etc.

7|Page minilikderse@[Link]
Introduction to statistics

Qualitative variable: Any variable that is not quantitative is qualitative. Qualitative


variables take a value that is one of several possible categories. As naturally measured,
qualitative variables have no numerical meaning. Examples: Hair color, gender, field of
study, college attended, political affiliation, status of disease infection.
Quantitative variables can be classified as discrete and continuous variable. Discrete variables
can assume certain numerical values. That is, there are gaps between the possible values. Such as
0, 1, 2...It may be countable finite or countable infinite. For example, the number of students in a
class room, number of children in a family. Continuous variable can take any value within a
specified interval with a finite enough measuring device. No gaps between possible values. They
are obtained by measuring. For example, consider the heights of two people no matter how close
it is we can find another person whose height falls some where between the two heights is a
continuous variable.

1.4 Use, Scope, Limitation and Misuse of Statistics

1.4.1 Uses of Statistics


Statistics presents fact in the form of numerical data
It condenses and summarizes a mass of data in to a few presentable and precise
figures.
It facilitates comparison of data
It helps in formulating and testing hypothesis
It helps in predicting future trend
It helps in formulating polices.

1.4.2 Scope of Statistics


The scope of statistics is indeed very vast. Apart from helping elicit an intelligent assessment
from a body of figures and facts, statistics is indispensable tool for any scientific enquiry-right
from the stage of planning enquiry to the stage of conclusion. It applies almost all sciences: pure
and applied, physical natural, biological, medical, agricultural and engineering. It also finds
applications in social and management sciences, in commerce, business and industry.

8|Page minilikderse@[Link]
Introduction to statistics

Of social sciences, economics leans most heavily on statistical methods for analyses of data
relating to micro as well as to macro economics, from demand analyses up to national income
analyses.

1.4.3 Limitations of Statistics


Statistics, with all its wide application in every sphere of human activity, has its own limitation.
Some of them are given below

Statistics is not suitable to the study of qualitative phenomenon: Since statistics is


basically a science that deals with a set of numerical data, it is applicable to the study
of only these subjects of enquiry, which can be expressed in terms of quantitative
measurements. As a matter of fact, qualitative phenomenon like honesty, poverty,
beauty, intelligence etc, cannot be expressed numerically and any statistical analysis
cannot be directly applied on these qualitative phenomenons. Nevertheless,
statistical techniques may be applied indirectly by first reducing the qualitative
expressions to accurate quantitative terms. For example, the intelligence of a group
of students can be studied on the basis of their marks in a particular examination.
Statistics does not study individuals: Statistics does not give any specific
importance to the individual items; in fact it deals with an aggregate of objects.
Individual items, when they are taken individually do not constitute any statistical
data and do not serve any purpose for any statistical enquiry.
Statistical laws are not exact: It is well known that mathematical and physical
sciences are exact. But statistical laws are not exact and statistical laws are only
approximations. Statistical conclusions are not universally true. They are true only
on an average.
Statistics table may be misused: Statistics must be used only by experts; otherwise,
statistical methods are the most dangerous tools on the hands of the inexpert. The
use of statistical tools by the inexperienced and untraced persons might lead to
wrong conclusions. Statistics can be easily misused by quoting wrong figures of
data. As King says aptly ‘statistics are like clay of which one can make a God or
Devil as one pleases.’

9|Page minilikderse@[Link]
Introduction to statistics

Statistics is only, one of the methods of studying a problem: Statistical method does
not provide complete solution of the problems because problems are to be studied
taking the background of the countries culture, philosophy or religion into
consideration. Thus the statistical study should be supplemented by other evidences.
At times, association or relationship between two or more variables is studied in
statistics, but such a relationship does not indicate ‘cause and effect’ relationship. It
simply shows the similarity or dissimilarity in the movement of the two variables. In
such cases, it is the user who has to interpret the results carefully, pointing out the
type of relationship obtained.

1.4.4 Misuse of Statistics


Apart from the limitation of statistics mentioned above, there are misuses of statistics. Many
people, knowingly or unknowingly, use data in wrong manner. Let us see what the main misuses
of statistics are so that the same could be avoided when one has to use statistical data. The
misuse of statistics may take several forms some of which are explained below.

Source of data not given: At times, the source of data not given. In the absence of
the source, the reader does not know how far the data are reliable. Further, if he
wants refer to the original source, he is unable to do so.
Defective data: Another misuse is that sometimes one gives inaccurate data. This
may be done knowingly in order to defend one’s position or to prove a particular
point. This apart, the definition used to denote a certain phenomenon may be
defective.
Unrepresentative sample: In statistics, several times one has to conduct a survey,
which necessitates to choose a sample from a given population or universe. The
sample may turn out to be unrepresentative of the universe. One may choose a
sample just on the basis of convenience. He may collect the desired information

10 | P a g e minilikderse@[Link]
Introduction to statistics

from either his friends or nearby respondents in his neighborhood even though
such respondents do not constitute a representative sample.
Inadequate sample: Earlier, we have seen a sample that is unrepresentative of the
universe is a major misuse of statistics. This apart, at times one may conduct a
survey based on an extremely inadequate sample. For example, in a city we may
find that are 100,000 households. When we have to conduct a household survey,
we may take a sample of merely 100 households comprising only 0.1 percent of
the universe. A survey based on such a small sample may not yield right
information.
Unfair comparison: An important misuse of statistics is making unfair
comparisons from the data collected. For instance, one may construct an index of
production choosing the base year where the production was much less. Then he
may compare the subsequent year’s production from this low base. Such a
comparison will undoubtedly give a rosy picture of the production though the
reality it is not so. Another source of unfair comparison could be when one makes
absolute comparisons instead of relative ones. An absolute comparison of two
figures say, production or export, may show a good increase, but in relative terms
it may turn out to be very negligible. Another example of unfair comparison is
when the population of the two cities is different; a comparison of over all death
and death rate by a particular disease is attempted.
Unwarranted conclusion: Another misuse of statistics may be on account of
unwanted conclusions. This may be as a result of making false assumptions. For
example, while making projection of population in the next five years, one may
assume a lower rate of growth though the past two years indicate otherwise.
Another source of unwarranted conclusion may be the use of wrong average.
Suppose in a series there are extreme values, one is very high while the other is
too low, such as 800 and 38. The use of an arithmetic average in such a case may
give a wrong idea.
Confusion of correlation and causation: In statistics, several times one has to
examine the relationship between two variables. A close relationship between the
two variables may not establish a cause-and-effect relationship in the sense that

11 | P a g e minilikderse@[Link]
Introduction to statistics

one variable is the cause and the other variable is the effect. It should be taken as
something that is measures degrees of association rather than try to find out casual
relationship.
Suppression of unfavorable results: Another wrong use of statistics may be on
account of suppressing results that are unfavorable to the organization or an
individual. Revealing such results may expose the concerned organization or
individual in bad light. In order to avoid such a situation, one may be attempted to
hide unfavorable, though true, facts emerging from statistical study.
Mistake in arithmetic: Finally, one may come across certain mistakes in
calculation or in the application of wrong formula. This human error may result in
grossly wrong figures, leading to wrong conclusion.
1.5 Scales of Measurement

Normally, when one hears the term measurement, they may think in terms of measuring the
length of something (i.e. the length of a piece of wood) or measuring a quantity of something
(i.e. a cup of flour). This represents a limited use of the term measurement. In statistics, the term
measurement is used more broadly and is more appropriately termed scales of measurement.
Scales of measurement refer to ways in which variables or numbers are defined and categorized.
Each scale of measurement has certain properties which in turn determine the appropriateness for
use of certain statistical analyses. The four scales of measurement are nominal, ordinal, interval,
and ratio.

Nominal scale allows for only qualitative classification (categorical data). That is, it
can be measured only in terms of whether the individual items belong to some
distinctively different categories, but we cannot quantify or even rank order those
categories. For example, all we can say is that two individuals are different in terms
of variable A (e.g., they are of different race), but we cannot say which one "has
more" of the quality represented by the variable. Typical examples of nominal
variables are gender, race, color, etc.
Ordinal scale allows us to rank or order the items we measure in terms of which has
less and which has more of the quality represented by the variable, but still they do
not allow us to say "how much more." A typical example of an ordinal variable is the

12 | P a g e minilikderse@[Link]
Introduction to statistics

socioeconomic status of families. For example, we know that upper-middle is higher


than middle but we cannot say that it is, for example, 18% higher. Also this very
distinction between nominal, ordinal, and interval scales itself represents a good
example of an ordinal variable. For example, we can say that nominal measurement
provides less information than ordinal measurement, but we cannot say "how much
less" or how this difference compares to the difference between ordinal and interval
scales.

Interval scale allows us not only to rank or order the items that are measured, but also
to quantify and compare the sizes of differences between them. For example,
temperature, as measured in degrees Fahrenheit or Celsius, constitutes an interval
scale. We can say that a temperature of 40 degrees is higher than a temperature of 30
degrees, and that an increase from 20 to 40 degrees is twice as much as an increase
from 30 to 40 degrees.

Ratio scale is very similar to interval variables; in addition to all the properties of
interval variables, they feature an identifiable absolute zero point, thus they allow for
statements such as is two times more than y. typical examples of ratio scales are

measures of time or space. For example, as the Kelvin temperature scale is a ratio
scale, not only can we say that a temperature of 200 degrees is higher than one of 100
degrees; we can correctly state that it is twice as high. Interval scales do not have the
ratio property. Most statistical data analysis procedures do not distinguish between
the interval and ratio properties of the measurement scales.

Note that: Permissible Arithmetic operations of measurement of scales are given


below

Nominal Ordinal Interval Ratio


Counting Greater than or Addition and Multiplication and
less than subtraction of division of scale
operations. scale values. values.

13 | P a g e minilikderse@[Link]
Introduction to statistics

1.6

Contents

2. Methods of Data Collection and Data Presentation

2.1 Methods of Data Collection

We have already explained what it means by statistical data. Numerical facts or measurements
obtained in the course of enquiry in to a phenomenon, marked by uncertainty, constitute
statistical data. The statistical data may be already available or may have to be collected by an
investigator or an agency. Data termed primary when the reference is to data collected for the
first time by the investigator and is termed secondary when the data are taken from records or
data already available.

2.1.1 Method of primary data collection

In primary data collection, you collect the data yourself using methods such as interviews and
questionnaires. The key point here is that the data you collect is unique to you and your research
and, until you publish, no one else has access to it. There are many methods of collecting
primary data and the main methods include:

Questionnaire: It is a popular means of collecting data, but is difficult to design


and often require many rewrites before an acceptable questionnaire is produced.

Advantages:

Can be used as a method in its own right or as a basis for


interviewing or a telephone survey.

Can be posted, e-mailed or faxed.

Can cover a large number of people or organizations

Wide geographic coverage.

Relatively cheap.

No prior arrangements are needed

14 | P a g e minilikderse@[Link]
Introduction to statistics

Avoids embarrassment on the part of the respondent.

Respondent can consider responses

Possible anonymity of respondent.

No interviewer bias.

Disadvantages:

Design problems

Historically low response rate (although inducements may help).

Time delay whilst waiting for responses to be returned

Require a return deadline.

Several reminders may be required.

Assumes no literacy problems.

No control over who completes it.

Not possible to give assistance if required.

Replies not spontaneous and independent of each other.

Respondent can read all questions beforehand and then decide


whether to complete or not. For example, perhaps because it is too
long, too complex, uninteresting, or too personal.

Interviewing is a technique that is primarily used to gain an understanding of the


underlying reasons and motivations for people’s attitudes, preferences or
behavior. Interviews can be undertaken on a personal one-to-one basis or in a
group. They can be conducted at work, at home, in the street or in a shopping
center, or some other agreed location.

Advantages:

Serious approach by respondent resulting in accurate information.

Good response rate.

15 | P a g e minilikderse@[Link]
Introduction to statistics

Completed and immediate.

Possible in-depth questions.

Interviewer in control and can give help if there is a problem.

Can investigate motives and feelings.

Can use recording equipment.

Characteristics of respondent assessed – tone of voice, facial


expression, hesitation, etc.

Can use props.

If one interviewer used, uniformity of approach.

Used to pilot other methods.

Disadvantages:

Need to set up interviews.

Time consuming.

Geographic limitations.

Can be expensive.

Normally need a set of questions.

Respondent bias – tendency to please or impress, create false


personal image, or end interview quickly.

Embarrassment possible if personal questions.

Transcription and analysis can present problems– subjectivity.

If many interviewers, training required.

Observation: It involves recording the behavioral patterns of people, objects and


events in a systematic manner.
Diaries: A diary is a way of gathering information about the way individuals
spend their time on professional activities. They are not about records of
engagements or personal journals of thought! Diaries can record either

16 | P a g e minilikderse@[Link]
Introduction to statistics

quantitative or qualitative data, and in management research can provide


information about work patterns and activities.

2.1.2 Methods of secondary data collection

Secondary data analysis can be literally defined as second-hand analysis and is the analysis of
data or information that was either gathered by someone else (e.g., researchers, institutions, other
NGOs, etc.) or for some other purpose than the one currently being considered, or often a
combination of the two.

Some of the sources of secondary data are government document, official statistics, technical
report, scholarly journals, trade journals, review articles, reference books, research institutes,
universities, libraries, library search engines, computerized data base and world wide web (

).

Advantage of secondary data

Saves time and money


Unobtrusive
Avoid data collection problem
Provide bases for comparison
Disadvantage of secondary data

Data availability
Level of observation
Quality of documentation
Data quality control
Outdated data

17 | P a g e minilikderse@[Link]
Introduction to statistics

2.2 METHODS OF DATA PRESENTATION

“I've come loaded with statistics, for I've noticed that a man can't prove
anything without statistics”
M. TWAIN
This chapter introduces tabular and graphical methods commonly used to summarize both
qualitative and quantitative data. Tabular and graphical summaries of data can be obtained in
annual reports, newspaper articles and research studies. Everyone is exposed to these types of
presentations, so it is important to understand how they are prepared and how they will be
interpreted.

Modern statistical software packages provide extensive capabilities for summarizing data and
preparing graphical presentations. MINITAB, SPSS and STATA are three packages that are
widely available.

2.2.1 Frequency Distribution


A frequency distribution is the organization of row data in table form, using classes and
frequencies. There are three basic types of frequency distributions, and there are specific
procedures for constructing each type. The three types are categorical, ungrouped and grouped
frequency distributions.

The reasons for constructing a frequency distribution are as follows

To organize the data in a meaningful, intelligible way.


To enable the reader to determine the nature or shape of the distribution
To facilitate computational procedures for measures of average and spread
To enable the researcher to draw charts and graphs for the presentation of data
To enable the reader to make comparisons between different data set

Some of basic terms that are most frequently used while we deal with frequency distribution are
the following:

Lower Class Limits are the smallest number that can belong to the different class.
Upper Class Limits are the largest number that can belong to the different classes.

18 | P a g e minilikderse@[Link]
Introduction to statistics

Class Boundaries are the number used to separate classes, but without the gaps created
by class limits.
Class midpoints are the midpoints of the classes. Each class midpoint can be found by
adding the lower class limit to the upper class limit and dividing the sum by 2.
Class width is the difference between two consecutive lower class limits or two
consecutive lower class boundaries.

Categorical Frequency Distribution

The categorical frequency distribution is used for data which can be placed in specific categories
such as nominal or ordinal level data. For example, data such as political affiliation, religious
affiliation, or major field of study would use categorical frequency distribution.

The major components of categorical frequency distribution are class, tally and frequency.
Moreover, even if percentage is not normally a part of a frequency distribution, it will be added
since it is used in certain types of graphical presentations, such as pie graph.

Steps of constructing categorical frequency distribution

1. You have to identify that the data is in nominal or ordinal scale of measurement
2. Make a table as show below

3. Put distinct values of a data set in column A


4. Tally the data and place the result in column B
5. Count the tallies and place the results in column C
6. Find the percentage of values in each class by using the formula

19 | P a g e minilikderse@[Link]
Introduction to statistics

Where f frequency and n is total number of values

Example 2.1: Twenty-five army inductees were given a blood test to determine their blood type.
The data set is given as follows:

A B B AB O
O O B AB B
B B O A O
A O O O AB
AB A O B A

Construct a frequency distribution for the above data.

Solution:

Ungrouped Frequency Distribution

When the data are numerical interested of categorical, the range of data is small and each class is
only one unit, this distribution is called an ungrouped frequency distribution.

The major components of this type of frequency distributions are class, tally, frequency and
cumulative frequency. The steps are almost similar with that of categorical frequency
distribution.

Cumulative frequencies are used to show how many values are accumulated up to and including
a specific class.

20 | P a g e minilikderse@[Link]
Introduction to statistics

Example 2.2: The following data represent the number of days of sick leave taken by each of 50
workers of a company over the last 6 weeks.

2 0 0 5 8 3 4 1 0 0
7 1 7 1 5 4 0 4 0 1
8 9 7 0 1 7 2 5 5
4 3 3 0 0 2 5 1 3
0 2 4 5 0 5 7 5 1
1 0 2

A. Construct ungrouped frequency distribution


B. How many workers had at least 1 day of sick leave?
C. How many workers had between 3 and 5 days of sick leave?

Solution:

A. Since this data set contains only a relatively small number of distinct or different
values, it is convenient to represent it in a frequency table which presents each distinct
value along with its frequency of occurrence.

B. Since 12 of the 50workers had no days of sick leave, the answer is 50-12=38
C. The answer is the sum of the frequencies for values 3, 4 and 5 that is 4+5+8=17

21 | P a g e minilikderse@[Link]
Introduction to statistics

Grouped Frequency Distribution

When the range of the data is large, the data must be grouped in which each class has more than
one unit in width. While we construct this frequency distribution, we have to follow the
following steps.

1. Find the highest and the lowest values


2. Find the range; or
3. Select the number of classes desired. Here, we have two choices to get the desired
number of classes:

I. Use Struge’s rule. That is, where is the number of class and

is the number of observations.


II. Select the number of classes arbitrarily between 5 and 20. This is a conventional

way. If you fail to calculate by Struge’s rule, this method is more appropriate.

When we choose the number of classes, we have to think about the following criteria

The classes must be mutually exclusive. Mutually exclusive classes have non
overlapping class limits so that values can’t be placed in to two classes.
The classes must be continuous. Even if there are no values in a class, the class
must be included in the frequency distribution. There should be no gaps in a
frequency distribution. The only exception occurs when the class with a zero
frequency is the first or last. A class width with a zero frequency at either end
can be omitted with out affecting the distribution.
The classes must be equal in width. The reason for having classes with equal
width is so that there is not a distorted view of the data. One exception occurs
when a distribution is open-ended. i.e., it has no specific beginning or end values.
4. Find the class width by dividing the range by the number of classes

22 | P a g e minilikderse@[Link]
Introduction to statistics

Note that: Round the answer up to the nearest whole number if there is a reminder. For

instance, and

5. Select the starting point as the lowest class limit. This is usually the lowest score
(observation). Add the width to that score to get the lower class limit of the next class.

Keep adding until you achieve the number of desired classes calculated in step 3.

6. Find the upper class limit; subtract unit of measurement from the lower class limit of

the second class in order to get the upper class limit of the first class. Then add the width
to each upper class limit to get all upper class limits.
Unit of measurement: Is the next expected value. For instance, 28, 23, 52, and then the
unit of measurement of this data set is one. Because take one datum arbitrarily, say 23,

then the next value will be 24. Therefore, . If the data set is 24.12, 30,

21.2, then give priority to the datum with more decimal place. Take 24.12 and guess the
next possible value. It is 24.13. Therefore,
Note that: U=1 is the maximum value of unit of measurement and is the value when we
don’t have a clue about the data.

7. Find the class boundaries. and

.In short, and

8. Tally the data and write the numerical values for tallies in the frequency column
9. Find cumulative frequency. We have two type of cumulative frequency namely less than
cumulative frequency and more than cumulative frequency. Less than cumulative
frequency is obtained by adding successively the frequencies of all the previous classes
including the class against which it is written. The cumulate is started from the lowest to
the highest size. More than cumulative frequency is obtained by finding the cumulate
total of frequencies starting from the highest to the lowest class.

23 | P a g e minilikderse@[Link]
Introduction to statistics

For example, the following frequency distribution table gives the marks obtained by 40
students:

The above table shows how to find less than cumulative frequency and the table shown
below shows how to find more than cumulative frequency.

Example 2.3: Consider the following set of data and construct the frequency distribution.

11 29 6 33 14 21 18 17 22 38
31 22 27 19 22 23 26 39 34 27

Steps

1. Highest value=39, Lowest value=6


2.

3.

4.

24 | P a g e minilikderse@[Link]
Introduction to statistics

5. Select starting point. Take the minimum which is 6 then add width 6 on it to get the next
class LCL.

6. Upper class limit. Since unit of measurement is one. . So 11 is the UCL of the
first class. Therefore, is the first class

7. Find the class boundaries. Take the formula in step 7. and

8. 9 and 10

Relative Frequency Distribution

An important variation of the basic frequency distribution uses relative frequencies, which are
easily found by dividing each class frequency by the total of all frequencies. A relative frequency

25 | P a g e minilikderse@[Link]
Introduction to statistics

distribution includes the same class limits as a frequency distribution, but relative frequencies are
used instead of actual frequencies. The relative frequencies are sometimes expressed as percent.

Relative frequency distribution enables us to understand the distribution of the data and to
compare different sets of data.

2.2.2 Graphical Presentation of Data

We have discussed the techniques of classification and tabulation that help us in organizing the
collected data in a meaningful fashion. However, this way of presentation of statistical data does
not always prove to be interesting to a layman. Too many figures are often confusing and fail to
convey the massage effectively.

One of the most effective and interesting alternative way in which a statistical data may be
presented is through diagrams and graphs. There are several ways in which statistical data may
be displayed pictorially such as different types of graphs and diagrams.

General steps in constructing graphs

1. Draw and label the and axes


2. Choose a suitable scale for the frequencies or cumulative frequencies and label it on the
axis.
3. Represent the class boundaries for the histogram or Ogive or the mid point for the
frequency polygon on the axis.
4. Plot the points
5. Draw the bars or lines
1 Pie Chart

Pie chart can used to compare the relation between the whole and its components. Pie chart is a
circular diagram and the area of the sector of a circle is used in pie chart. Circles are drawn with
radii proportional to the square root of the quantities because the area of a circle is .

26 | P a g e minilikderse@[Link]
Introduction to statistics

To construct a pie chart (sector diagram), we draw a circle with radius (square root of the total).
The total angle of the circle is .

The angles of each component are calculated by the formula

These angles are made in the circle by mean of a protractor to show different components. The
arrangement of the sectors is usually anti-clock wise.

Example2.4: The following table gives the details of monthly budget of a family. Represent
these figures by a suitable diagram.

Solution: The necessary computations are given below:

27 | P a g e minilikderse@[Link]
Introduction to statistics

2 Bar Charts
The bar graph (simple bar chart, multiple bar chart and stratified or stacked bar chart) uses
vertical or horizontal bins to represent the frequencies of a distribution. While we draw bar chart,
we have to consider the following two points. These are
Make the bars the same width
Make the units on the axis that are used for the frequency equal in size
A simple bar chart is used to represents data involving only one variable classified on spatial,
quantitative or temporal basis. In simple bar chart, we make bars of equal width but variable
length, i.e. the magnitude of a quantity is represented by the height or length of the bars.
Following steps are undertaken in drawing a simple bar diagram:

28 | P a g e minilikderse@[Link]
Introduction to statistics

Draw two perpendicular lines one horizontally and the other vertically at an appropriate
place of the paper.
Take the basis of classification along horizontal line (X-axis) and the observed variable
along vertical line (Y-axis) or vice versa.

Marks signs of equal breath for each class and leave equal or not less than half breath in
between two classes.

Finally, marks the values of the given variable to prepare required bars.

Example 2.5: Draw simple bar diagram to represent the profits of a bank for 5 years.

Multiple bar charts are used two or more sets of inter-related data are represented (multiple
bar diagram facilities comparison between more than one phenomenon). The technique of

29 | P a g e minilikderse@[Link]
Introduction to statistics

simple bar chart is used to draw this diagram but the difference is that we use different
shades, colors, or dots to distinguish between different phenomena.

Example 2.6: Draw a multiple bar chart to represent the import and export of Canada (values
in $) for the years 1991 to 1995.

Stratified (Stacked) Bar Chart is used to represent data in which the total magnitude is divided
into different or components. In this diagram, first we make simple bars for each class taking

30 | P a g e minilikderse@[Link]
Introduction to statistics

total magnitude in that class and then divide these simple bars into parts in the ratio of various
components. This type of diagram shows the variation in different components within each class
as well as between different classes. Sub-divided bar diagram is also known as component bar
chart.
Example 2.7: The table below shows the quantity in hundred kgs of Wheat, Barley and Oats
produced on a certain form during the years 1991 to 1994. Draw stratified
bar chart.

Solution: To make the component bar chart, first of all we have to take year wise total
production.

The required diagram is given below:

31 | P a g e minilikderse@[Link]
Introduction to statistics

3 Histogram

Histogram is a special type of bar graph in which the horizontal scale represents classes of data
values and the vertical scale represents frequencies. The height of the bars correspond to the
frequency values, and the drawn adjacent to each other (without gaps).

We can construct a histogram after we have first completed a frequency distribution table for a
data set. The axis is reserved for the class boundaries.

Example2.8: Take the data in example 2.3.

32 | P a g e minilikderse@[Link]
Introduction to statistics

7.0

6.0
Frequency

5.0

4. 0

3.0

2.0

1.0

0.0 35.5
5.5 11.5 17.5 23.5 29.5 41.5

Class boundaries

Relative frequency histogram has the same shape and horizontal ( ) scale as a histogram,

but the vertical ( ) scale is marked with relative frequencies instead of actual frequencies.

4 Frequency Polygon

A frequency polygon uses line segment connected to points located directly above class midpoint
values. The heights of the points correspond to the class frequencies, and the line segments are
extended to the left and right so that the graph begins and ends on the horizontal axis with the
same distance that the previous and next midpoint would be located.

Example 2.9: Take the data in example 2.3.

33 | P a g e minilikderse@[Link]
Introduction to statistics

7.0

6.0

5.0

4.0

3.0

2.0

2.5 8.5 14.5 20.5 26.5 32.5 38.5 44.5


Midpoints

5 Ogive Graph

An Ogive (pronounced as “oh-jive”) is a line that depicts cumulative frequencies, just as the
cumulative frequency distribution lists cumulative frequencies. Note that the Ogive uses class
boundaries along the horizontal scale, and graph begins with the lower boundary of the first class
and ends with the upper boundary of the last class. Ogive is useful for determining the number of
values below some particular value. There are two type of Ogive namely less than Ogive and
more than Ogive. The difference is that less than Ogive uses less than cumulative frequency and

more than Ogive uses more than cumulative frequency on axis.

Example 2.10: Take the data in example 2.3 and draw less than and more than Ogive

34 | P a g e minilikderse@[Link]
Introduction to statistics

20
Less than Ogive

15

10

More than Ogive

5.5 11.5 17.5 23.5 29.5 35.5 41.5


Class Boundaries

MEASURES OF CENTRAL TENDENCY

“The way to make sense out of raw data is to compare and contrast, to understand difference.”

G. BATESON

Researchers are often interested in defining a value that best describes some attribute of the
population. The best way to reduce a set of data and still retain part of the information is to
summarize the set with a single value. Therefore, measures of central tendency are one of
descriptive statistics.

3.1 Introduction and Objectives of Measure of Central Tendency

35 | P a g e minilikderse@[Link]
Introduction to statistics

Our objective in this chapter is to develop measures that can be used to summarize a data set.
When describing, exploring, and comparing data sets, the following characteristics are usually
extremely important.

1. Center: A representative or average value that indicated where the middle of the
data set is located.
2. Variation: A measure of the amount that the data values vary among themselves
3. Distribution: The nature or shape of the distribution of the data (such as bell-
shaped, uniform, or skewed)
4. Outliers: Sample values that lie very far away from the vast majority of the other
sample values.
5. Time: Changing characteristics of the data over time.
The above five characteristics are so important that they might be better remembered by using a
mnemonic for the first letters CVDOT, such as “Computer Viruses Destroy Or Terminate”. A
measure of center is a value at the center or middle of a data set. The three major objective of
measure of central tendency are

To summarize a set of data by single value


To facilitate comparison among different data sets
To use for further statistical analysis or manipulation
3.2 The Summation Notation

Let the symbol (read “ sub ”) denotes any of the value assumed by a

variable . The letter in , which can stand for any of the numbers is called a

subscript, or index. Clearly any letter other than , such as , could have been used as

well.

The symbol is used to denote the sum of all the from to ; by definition,

36 | P a g e minilikderse@[Link]
Introduction to statistics

When no confusion can result, we often denote this sum simply by


.
The symbol is the Greek capital letter sigma, denoting sum.

Basic properties of summation

1.

2. Where is

constant. More simply,

3. If are any constants, then

4. , where b is constant

5.

6.
Example 3.1: Express each of the following by using the summation notations

A.

B.

C.

D.
E.
Solutions:

A. B. C.

37 | P a g e minilikderse@[Link]
Introduction to statistics

D. E.

3.3. Properties of measures of central tendency

Each measure of central tendency should have the following properties

It should be easy to calculate and understand


It should be rigidly defined. In other words, it should have one and only one
interpretation so that personal bias of the investigator does not affect the value of its use
fullness.
It should be representative of the data under consideration
It should have sampling stability. In other words, it should not be affected by sampling
fluctuations.
It should not be affected by extreme values
It should be amenable for further algebraic manipulation

3.4. Types of Measures of Central Tendency

An average is a value that is typical, or representative, of a set of data since such typical values
tend to lie centrally within a set of data arranged according to magnitude, averages are also
called measures of central tendency.

Several types of averages can be defined, the most common being the arithmetic mean, the
median, the mode, the geometric mean, and the harmonic mean. Each has advantages and
disadvantages, depending on the data and the intended purpose.

3.4.1. Arithmetic Mean (simple and weighted)

The (arithmetic) mean is generally the most important of all numerical measurements used to

describe data, and it is what most people call and average.

The arithmetic mean of asset of values is the measure of center found by adding the values and

dividing the total by the number of values. The mean is denoted by (pronounced “ ”) if

38 | P a g e minilikderse@[Link]
Introduction to statistics

the data set is a sample from a larger population; if all values of the population are used, then we
population; if all values of the population are used, then we denote the mean by  (lower case
Greek mu).

Note that sample statistics are usually represented by English letters, such as , and population
parameters are usually represented by Geek letters, such as .

Example 3.2: The grades of a student on six examinations were Find

the arithmetic mean of the grades.

Solution:

Frequently one uses the term average synonymously with arithmetic mean. Strictly speaking,
however, this is incorrect since there are averages other than arithmetic mean.

If the numbers occur times, respectively, the arithmetic mean is

, where

Example 3.3: If occurs with frequencies respectively, the arithmetic

mean is

Arithmetic mean for grouped data

The only new concept in calculating mean for grouped data is that find mid-points or each class

and label it as .

39 | P a g e minilikderse@[Link]
Introduction to statistics

Then where is the mid-point of the class limit/boundary

Midpoint =

Example 3.4: Take the data in example 2.3 in chapter two. The midpoint of each class is given as
follow:

Midpoint (

8.5 2 17
14.5 2 29
20.5 7 143.5
26.5 4 106
32.5 3 97.5
38.5 2 77
Total

Properties of the Mean

The algebraic sum of the deviation of a set of numbers from their arithmetic mean is zero.

That is,

The sum of the square of deviations of a set of numbers from the mean is always the

least. That is, where

40 | P a g e minilikderse@[Link]
Introduction to statistics

If numbers have mean , numbers have mean … numbers have mean , then

the mean of all numbers is , combined mean. That is, a weighted arithmetic

mean of all the means.


If A is any guessed or assumed arithmetic mean ( which may be any number) and if

are the deviations of from A, then

Or

Advantages and uses of arithmetic mean

The mean is computed by using all the values of the data


The mean varies less when samples are taken from the same population and all three
measures are computed for these samples.
The mean is used in computing other statistics, such as variances
The mean for the data is unique
The mean can’t be computed for an open-ended frequency distribution
The mean is affected by extremely high or low values and may not be the appropriate
average to use in this situation.

Weighted Mean is a special type arithmetic mean and it will be functional when values have its

own weight. Suppose we give a weight to , to to , then the weighted mean

can be calculated by using the formula

41 | P a g e minilikderse@[Link]
Introduction to statistics

Example 3.5: A teacher attaches weights 2 to homework 3 to mid term exam and 5 for final
exam. If a student score 90, 50 and 60 for HM, MT and FE, respectively,
what is his/ her average academic performance?

Solution: Takeout the givens as follow

2 90
3 50
5 60

3.4.2 The Geometric Mean

The geometric mean of a set of positive numbers (observations) is the root

of the product of the numbers.

Geometric mean is often used in business and economics for finding average rates of change
average rates of growth, or average ratio.

Example 3.6: A price of a commodity increased by 5%, 8% and 77% for the three consecutive
years. What was the average yearly price increase?

Solution:

42 | P a g e minilikderse@[Link]
Introduction to statistics

Let be the price of the original year .

Average yearly price increase was 12.6%

Merits of Geometric Mean

It is useful in averaging ratios and percentages and determining rates of increase and
decrease

It is capable of algebraic treatment. Its formula can be extended to calculate combined

as follows

It gives less weight to large items and more to small ones than does the arithmetic
average. It is because of this reason that geometric mean is never larger than the
arithmetic mean, on occasions it may turn out to be same as the arithmetic mean, but it is
usually smaller
It is based on each and every item of the series
It is rigidly defined

Limitations of Geometric Mean

It is difficult to understand
It is difficult to compute
It can’t be computed when there are both negative and positive values in a series
It is biased for small values as it gives more weight to small values
Its computation becomes difficult especially when the values of items are large

43 | P a g e minilikderse@[Link]
Introduction to statistics

3.4.3 The Harmonic Mean

The harmonic mean of a set of numbers is the reciprocal of the arithmetic

mean of the reciprocals of the numbers or observations.

No value can be negative. The harmonic mean is often used as a measure of central tendency for
data sets consisting of rates of change, such as speeds.

Example 3.7: Four students drive from Jimma to Addis Ababa at a speed of 40 km/hr. Because
they need to reach statistics class on time, they return at a speed of 60
km/hr. What is their average speed for the round trip?

Solution:

Merits of Harmonic Mean

It is based on all the items of the series


It is capable of further algebraic treatment
It is rigidly defined
It is the most suitable average for measuring the time, speed etc

Limitations of Harmonic Mean

It is difficult to understand
It is difficult to calculate
It can’t be computed when one or more items are zero
It gives largest weight to smallest items. Hence, it is not useful for analyzing the
economic data

44 | P a g e minilikderse@[Link]
Introduction to statistics

A well known inequality concerning arithmetic, geometric, and harmonic means for any set of

positive numbers is the equality holds when the observations are

equal.

3.4.4 The Mode

The mode of a set observation is that value which occurs with the greatest frequency; that is, it is
the most common value. The value that occurs most frequently in a data set is called mode. The
mode may not exist; even if it does exist it may not be unique.

Example 3.8: The set 2, 2, 5, 7, 8, 9, 9, 9, 10, 10, and 11 has mode 9. The set 3, 5, 8, 10, 12, 15,
and 16 has no mode. The set 2, 3, 4, 4, 4, 5, 5, 7, 7, 7, and 9 has two
modes, 4 and 7 and is called bimodal.

A distribution having only one mode is called unimodal. In the case of grouped data where a

frequency curve has been constructed to fit the data, the mode will be the value (or values) of

corresponding to the maximum point (or points) on the curve. This value of is sometimes

denoted by (Pronounced as “ hat”). From the frequency distribution the mode can be obtained

from the formula

Where

is lower class boundary of the modal class. Modal class is a class which contains the mode and

has the highest frequency. is excess of modal frequency over frequency of next lower class.

is excess of modal frequency over frequency of the next higher class. is size of the modal

class interval

45 | P a g e minilikderse@[Link]
Introduction to statistics

Example 3.8: Take the data in example 2.3 and find the mode of the frequency distribution.

Solution:

The modal class is a class with highest frequency which is and

Advantages and uses of Mode

The mode is used when the most typical case is desired


The mode is easiest to compute
The mode can be used when the data is normal, such as religious preference, and political
affiliation
The mode is not always unique. A data can have more than one mode
The mode does not always exist for a data set

3.4.5 The Median and Other Quantiles

The median of a data set is the measure of center that is the middle value when the original data
values are arranged in order of increasing (or decreasing) magnitude. The median is often

denoted by (pronounced as “ tilde”).

To find the median, first sort the values, and then follow one of these two procedures:

1. If the number of values is odd, the median is the number located in the exact middle of
the list.

46 | P a g e minilikderse@[Link]
Introduction to statistics

2. If the number of values is even, the median if found by computing the mean of the two
middle numbers.

Median for grouped data can be computed as

Where

is the lower class boundary of the median class. Median class is a class which accommodates

observation. is the total of all frequencies. is the sum of frequencies of all classes

lower than the median class. is the frequency of the median class and is the width of

median class.

Example 3.9: Take the data from example 2.3 and calculate the median.

Solution: , the median class is then ,

Advantages and uses of Median

The median is used when one must find the center or middle value of a data set
The median is used when one must determine whether the data values fall in tot the upper
half or lower half of the distribution

47 | P a g e minilikderse@[Link]
Introduction to statistics

The median is used to find the average of an open-ended distribution


The median is affected less than the mean by extremely high or extremely low values

Quartile divides a given set of data in to four equal parts

Note that:

Decile divides a given set of data in to ten equal parts

Percentile divides a give set of data in to hundred equal parts

MEASURES OF DISPERSION

“Statistics are designed to keep you safe.”


[Link]

This chapter deals with the second most important characteristics of a distribution
so called variation which belongs to CVDOT. Without knowing something about
how data is dispersed, measures of central tendency may be misleading. Measures of
dispersion provide a more complete picture.
4.1 Introduction and Objectives of Measuring Dispersion

48 | P a g e minilikderse@[Link]
Introduction to statistics

The term dispersion is generally used in two senses. Firstly, dispersion refers to the variations of
the items among themselves. If the value of all the items of a series is the same, there will be no
variation among different items of a series; the more will be the dispersion. Secondly, dispersion
refers to the variation of the items around an average. If the difference between the value of
items and the average is large, the dispersion will be high and on the other hand if the difference
between the value of the items and averaging is small, the dispersion will be low. Thus,
dispersion is defined as scatteredness or spreadness of the individual items in a given series.

The measures of dispersion are helpful in statistical investigation. Some of the main objectives
of dispersion are as under:

1. To determine the reliability of an average: The measures of dispersion help in


determining the reliability of an average. It points out as to how far an average is
representative of a statistical series. If the dispersion or variation is small, the average
will closely represent the individual values and it is highly representative on the other
hand, if the dispersion or variation is large, the average will be quite unreliable.
2. To compare the variability of two or more series: The measures of dispersion help in
comparing the variability of two or more series. It is also useful to determine the
uniformity or consistency of two or more series. A high degree of variation would mean
less consistency or less uniformity as compared to the data having less variation.
3. For facilitating the use of other statistical measures: Measures of dispersion serve the
basis of many other statistical measures such as correlation, regression, testing of
hypothesis etc.
4. Basis of statistical quality control: The measure of dispersion is the basis of statistical
quality control. The extent of the dispersion gives indication to the management as to
whether the variation in the quality of the product is due to random factors or there is
some defect in the manufacturing process.
4.2 Absolute and Relative Measures
Measures dispersion may be either absolute or relative

1. Absolute measures of dispersion: Absolute measure is expressed in the same statistical


unit in which the original data are given such as kilograms, tones etc. These measures are
suitable for comparing the variability in two distributions having variables expressed in

49 | P a g e minilikderse@[Link]
Introduction to statistics

the same units and of the same averaging size. These measures are not suitable for
comparing the variability in two distributions having variables expressed in different
units.

2. Relative measures of dispersion: A relative measure of dispersion is the ratio of a


measure of absolute dispersion to an appropriate average or the selected items of the data.

50 | P a g e minilikderse@[Link]
Introduction to statistics

4.3 Types of Measures of Variation


4.3.1 The Range and Relative Range
It is the simplest measures of dispersion. It is defined as the difference between the largest and
smallest value in the series. Its formula is:

Where R=Range, L= Largest value in the series, S= smallest value in the series

The relative measures of range, also called coefficient of range, is defined as

Example 4.1: five students obtained the following marks in statistics: . Find the
Range and coefficient of range

Solution: Here,

51 | P a g e minilikderse@[Link]
Introduction to statistics

Coefficient of Range =

Example 4.2: Find out range and coefficient of range of the following series

Size 5-10 11-15 16-20 21-25 26-30


Frequency 4 9 15 30 40
Solution: Here,

Merits of Range

It is simple to understand
It is easy to compute
It is well-defined
It helps in giving an idea about the variation, just by giving the lowest value and the
greatest value of variable

Demerits:

It can’t be calculated in case of open-ended distribution


It is not based on all observations of the series
It is affected by sampling fluctuation
It is affected by extreme values in the series

4.3.2 The Quartile Deviation and Coefficient of Quartile Deviation

52 | P a g e minilikderse@[Link]
Introduction to statistics

Inter-quartile range and quartile deviation are other measures of dispersion. The difference

between the upper quartile and lower quartile is called inter-quartile range.

Symbolically,

The inter-quartile ranges covers dispersion of middle 50% of the items of the series. Quartile
deviation, also called semi-inter-quartile range is half of the difference between the upper and
lower quartile. That is, half of the inter-quartile range. Its formula as:

The relative measure of quartile deviation also called the coefficient of quartile deviation is
defined as:

Example 4.3: Find inter-quartile deviation, quartile deviation and coefficient of quartile
deviation from the following data.

28, 18, 20, 24, 27, 27, 30, 15

Solution: First arrange the data in ascending order. 25, 18, 20, 24, 27, 28, 30

53 | P a g e minilikderse@[Link]
Introduction to statistics

Example 4.4: Find inter-quartile range, quartile deviation and coefficient of quartile deviation
from the following data

Marks 2 3 4 5 6 7 8 9
No. Of students 10 11 12 13 5 12 7 5
Solution:

Marks No. Of students CF

2 10 10
3 11 21
4 12 33
5 13 46
6 5 51
7 12 63
8 7 70
9 5 75=N
Total N=75

54 | P a g e minilikderse@[Link]
Introduction to statistics

Percentile range is defined as

IQR and Outlier

Any applied statistician who has analyzed a number of sets of real data is likely to have come
across outliers. The intuitive definition of an outlier would be ‘an observation deviates so much
from other observations as to arouse suspicious that it was generated by different mechanism.’
That is, outliers are observations that are distinct from the main body of the data and are
incompatible with the rest of data. These values may be genuine observations from individuals
with very extreme levels of the variable.

1.5 IQR Rule

A simple approach to detect outlier is that print the data and visually checks them by eye. This is
suitable of the number of observations is not too large and if the potential outlier is much lower
than or higher than the rest of the data.

When the number of observations gets larger and larger, we can check the presence of outlier by
the 1.5 IQR rule. The steps to identify outliers are presented as follows:

1. Arrange the data in ascending order


2. Calculate the first, the third quartiles and inter-quartile range

3. Compute and any observation outside this range is

considered as outlier

Example 4.5: Consider the following data

1 2 5 5 7 8 10 11 11 12 15 25

Check the presence of outliers.

55 | P a g e minilikderse@[Link]
Introduction to statistics

Solution: The first step is arranging the data in ascending order then let us calculate the first and
third quartile

The next step is computing

Therefore, the observation less than -4 and greater than 20 are considered as outlier. That is, 25 is
outlier.

By using the concept of 1.5IQR rule, we can draw box plot which is used to give five-number
summaries. Five-number summaries contains minimum, quartile one, median, quartile three and
maximum.

Steps to draw box plot are”

1. Notice that you must have ordered data before you can find the Five – Number
Summaries.
2. Find the median first. It’s the middle point
3. Then find the quartiles, Q1 and Q3 and the 1.5 IQR outlier limits
4. Draw a “box" from Q1 to Q3 with bars at Q1, Q3 and the median. (In the below
example the box is horizontal, but it could also be vertical.)

5. Draw a straight line from Q3 to either the largest observation or the

upper outlier bound, whichever is smaller.

56 | P a g e minilikderse@[Link]
Introduction to statistics

6. Draw a straight line from Q1 to either the smallest observation or the

lower outlier bound, whichever is larger.


7. Any remaining observations (the outliers) are shown as individual points on the plot.
Example 4.6: Take the data 1, 2, 5, 5, 7, 8, 10, 11, 12, 12, 18, 25 and draw box plot

Solution:

Merits of

It is simple to understand
It is easy to compute
It is well-defined
It helps in studying the middle 50% item in the series
It is not affected by the extreme items
It is useful in the case of open-ended

Demerits of

It is not based on all the items


It is not capable of further algebraic treatment
It doesn’t have sampling stability

4.3.3 Mean Deviation & Coefficient of Mean Deviation

57 | P a g e minilikderse@[Link]
Introduction to statistics

Consider a set of observations . The mean or average deviation is defined

by

Where denotes the absolute value of the deviation. Generally, arithmetic mean and

median are used in calculating mean deviation. So, stands for the average used for calculating

. That is,

For a frequency distribution

Where is the frequency of . is the midpoint in the case of grouped frequency distribution

or class value in the case of ungrouped frequency distribution.

The relative measure of mean deviation, also called the coefficient of mean deviation is obtained
by dividing mean deviation by the particular average used in computing mean deviation. Thus,

58 | P a g e minilikderse@[Link]
Introduction to statistics

Exercise: Find all coefficients of mean deviations for the following frequency distribution:

Marks 10-20 21-30 31-40 41-50 51-60


No. Of students 4 8 20 12 6

Merits of

It is simple to understand
It is easy to compute
It is well-defined
It is based on all observations
It is not unduly affected by the extreme items
It can be calculated by using any average

Demerits of

It is not capable of further algebraic treatment


It does not take in to account the signs of the deviations of items from the average

Note that: of all the mean deviations taken about different averages or any arbitrary value, the
mean deviation about the median has the smallest value.

4.3.4 The Variance, the Standard Deviation and the Coefficient of Variation

Standard deviation is the most important and widely used measure of dispersion. It was first

used by Karl Pearson in 1893. The standard deviation of a statistical data is defined as the

59 | P a g e minilikderse@[Link]
Introduction to statistics

positive square root of the mean of the squared deviations of items from the mean of the series
under consideration.

For an individual series, the is given by

For a frequency distribution the is given by

Where is the frequency of . is the midpoint in the case of grouped frequency distribution

or class value in the case of ungrouped frequency distribution.

For comparing two or more series for variability, the corresponding relative measure, called
coefficient of variation is calculated. This measure is defined as:

Note that: The data set with minimum is less dispersed

Just as it is possible to calculate combined mean of two or more groups, similarly the combined
or pooled standard deviation of two or more groups can be calculated. The combined standard

deviation of two groups is denoted by and is computed as follow:

60 | P a g e minilikderse@[Link]
Introduction to statistics

For more than two groups (say groups)

Where is the group sample size and is the variance of the group.

Example 4.7: Two samples of size 100 and 150, respectively, have means 50 and 60 and
standard deviations 5 and 6. Find the mean and standard of the combined
sample of size 250.

Solution: Given

Exercise: Find the shortcut formula to find standard deviation, mean deviation and combined
deviation.

Correcting incorrect values of standard deviation

In certain cases, mean and standard deviation are calculated by using one or two incorrect values
of the variable. Just as we can correct an incorrect mean, similarly, there is a procedure of
correcting an incorrect standard deviation.

Steps in calculating the corrected standard deviation

1. Find out incorrect sum of square values of the variable. That is,

61 | P a g e minilikderse@[Link]
Introduction to statistics

2. Find corrected . To do so, we subtract the square of the incorrect item from incorrect

and add the square of correct item to incorrect . thus,

3. Apply the following formula:

Relationship between measures of dispersion

is approximately equal to

is approximately equal to

Note that: These relationship hold if the distribution is moderately symmetrical

Properties of standard deviation

The important mathematical properties of standard deviation are as follow:

The standard deviation of the first natural numbers can be found from the following

formula:

For example, the standard deviation of the first 5 natural numbers is given as:

62 | P a g e minilikderse@[Link]
Introduction to statistics

We can calculate combined standard deviation for two or more groups

If a constant amount ‘ ’ is added or subtracted from each item of a series, then

remains unaffected

If each item of a series is multiplied or divided by a constant ‘ ’ , then is affected by

the same amount


The standard deviation has the following relation to the arithmetic mean in a symmetrical
distribution:

 includes 68.27% of the observations

 includes 95.45% of the observations

 includes 99.73% of the observations

In general, for any distribution with the interval contains at

least fraction of the total number of observations.

Example 4.8: If the value of standard deviation in moderately symmetrical distribution is 24,
find the value of mean deviation and quartile deviation.

Solution: For moderate symmetrical series

Example 4.9: If the mean and the standard deviation of 25 boys’ weight are 50 and 5,

respectively, at least how many boys will in the interval ?

63 | P a g e minilikderse@[Link]
Introduction to statistics

Solution:

Therefore, at least or 50% of the boys have body weight with in

the interval

Sheppard’s Correction for Variance

When the observations are grouped into classes, all observations in a class are equal to the
midpoint of the class. This introduces some error known as grouping error. Sheppard suggests a

correction known as Sheppard’s correction. It is given by where is the class width.

Merits of

It is simple to understand
It is well-defined
It is based on all items
It is suitable for further algebraic treatment
It has sampling stability
It is very useful in the study of “Tests of Significant”

Demerits of

It is easy to calculate
It is unduly affected by extreme values

64 | P a g e minilikderse@[Link]
Introduction to statistics

4.3.5 Standard Score

A standard score for sample vale in a data set is obtained by the mean of the data set from the
value and dividing the result by the standard deviation of the data set. Basically, the standard
score (z-score) tells us how many standard deviations a specific value is above or below the
mean value of the data set. That is, the z-score is the number of standard deviations the data
value falls above (positive z-score) or below (negative z-score) the mean for the data set.

Z-score computed from the sample

Z-score computed from the population

Note that: The Z-score is affected by an outlying value in the data set because the outlier (very
small or very large value) directly affects the value of the mean and the standard values by
eliminating the unit of measurement.

Example 4.10: what is the Z-score for the value of 14 in the following sample data set?

3 8 6 14 4 12 7 10

Solution:

= 8, SD = 3.8173 thus, Z =

 The data value of 14 is located 1.57 standard deviations above the mean 8 because the
z-score is positive.

4.4 Moments, Skewness and Kurtosis

65 | P a g e minilikderse@[Link]
Introduction to statistics

Moments are statistical tools used in statistical investigation. The moments of a distribution are
the arithmetic mean of the various powers of the deviations of items from some number. In our
course, we shall use it in the study of Skewness and Kurtosis of statistical distribution.

Moments about the origin

Where

Moments about the origin for grouped frequency distribution and for ungrouped frequency
distribution

Where is the frequency of . is the midpoint in the case of grouped frequency distribution

or class value in the case of ungrouped frequency distribution.

Note that: ,

Moments about the Mean (Central Moments)

Moments about the mean for grouped frequency distribution and for ungrouped frequency
distribution

Where is the frequency of . is the midpoint in the case of grouped frequency distribution

or class value in the case of ungrouped frequency distribution.

66 | P a g e minilikderse@[Link]
Introduction to statistics

Note that: if it is assumed

Moments about any arbitrary constant

Moments about any arbitrary constant for grouped frequency distribution and for ungrouped

frequency distribution

Example 4.11: Find the first four moments about the mean for the following individual series

3 6 8 10 18

Solution:

[Link]

1 3 -6 36 -216 1296
2 6 -3 9 -27 81
3 8 -1 1 -1 1
4 10 1 1 1 1
5 18 9 81 729 6561

Tota
l
Now

67 | P a g e minilikderse@[Link]
Introduction to statistics

Exercise: Find the relationship between moments about any arbitrary constant and moments

about the mean.

Skewness

Skewness is a measure of symmetry, or more precisely, the lack of symmetry or departure from

symmetry. For a positively skewed distribution, and for a

negatively skewed distribution, . Skewness can be measured in

absolute terms by taking the difference between arithmetic mean and mode. Therefore, the
formula for absolute skewness is given as:

If the value of arithmetic mean is greater than mode, Skewness is positive and if the value of
mode is greater than mean, the skewness is negative. This absolute measure of skewness is not
free from unit of measurement. Further, the difference between arithmetic mean and mode in
absolute terms may be higher in one situation, although the frequency curves of two distributions
are similarly skewed. It is because of this reason; it is desirable to have a measure that can be

68 | P a g e minilikderse@[Link]
Introduction to statistics

directly used for comparisons. If the absolute differences are expressed in relation to some
measure of spread, the resultant measure will be a relative measure of skewness. The two most
commonly used measures are Karl Pearson’s coefficient skewness and Bowley’s coefficient of
skewness.

Karl-Pearson’s Coefficient of Skewness

It has been indicated above that the distance between the arithmetic mean and the mode can be
used as a measure of skewness. However, since the measure of skewness should be a pure
number, free from units of measurement, we define

When the distribution is symmetrical, mean, median and mode coincide and therefore, the
coefficient of skewness will be zero. When the coefficient of skewness is positive, it indicates
that the distribution is positively skewed and when the coefficient is negative, the distribution is
negatively skewed. The numerical value say 0.9 or 0.3 etc, indicates the degree of skewness.
Therefore, Karl Pearson’s formula for skewness indicates both direction as well as the extent of
skewness.

Note that: In moderately skewed distributions the averages have the following relationship.

Hence, for moderately skewed distribution the following formula is used for skewness.

Bowley’s Coefficient of Skewness

Bowley’s coefficient of skewness is based on quartiles. Thus a measure of skewness based on


the distance from the median is defined as follows:

69 | P a g e minilikderse@[Link]
Introduction to statistics

This measure is also referred to as the quartile measure skewness and the value of the coefficient

lies between .

Kurtosis

Kurtosis in Greek language mean ‘bulginess’, it measures the flatness of the curve. Three terms
are used for indicating flatness, mesokurtic stands for a normal curve, leptokurtic for a peaked
curve and platykurtic for a curve less peaked than normal.

The relative measure of kurtosis is defined as:

70 | P a g e minilikderse@[Link]
Introduction to statistics

Note that:

The relative measure of skewness based on central moment is defined as:

ELEMENTARY PROBABILITY

“Life is a school of probability. “

W. BAGEHOT

The notion that chance, or probability, can be treated numerically is relatively recent. Indeed, for
most of recorded history it was felt that what occurred in life was determined by forces that were
beyond one’s ability to understand. It was only during the first half of the 17th century, near the
end of Renaissance, that people become curious about the world and the laws governing its
operation. Among the curious were the gamblers.

5.1 Introduction
A cynical person once said, “The only two sure things are death and taxes.” This philosophy no
doubt arose because so much in people’s lives is affected by chance.

Probability as a general concept can be defined as the chance of an event occurring. Most people
are familiar with probability from observing or plying games of chance, such as card games or
lotteries. Probability is the basis of inferential statistics.

5.2 Definitions of Some Probability Terms


Terms that are most frequently used and cornerstone of probability are defined as follows:

Probability experiment: It is a process that leads to well-defined results called outcomes.


Outcomes: It is the result of a single trial of probability experiment. It is sometimes
called sample point.

71 | P a g e minilikderse@[Link]
Introduction to statistics

Sample Space: It is the set of all possible outcomes of a probability experiment and

denoted by .

Event: It is a subset of sample space (contains one or more outcomes which are in the
sample space) and is defined for a particular purpose. An event can be one outcome or
more than one outcome. Simple event is an event having only single outcome. Compound
event consisting of one or more outcomes or simple events. Event is denoted by capital
letters such as A, B, F etc.
Mutually exclusive events: Suppose you have two events, say A and B. if these events
have no common sample point(s) or do not occur simultaneously, then the two events are
called mutually exclusive events.
Equally-likely events: It is a situation where the probability of the occurrence of one
event as likely as the other event. That is, they must have equal probability of occurrence.
Exhaustive events: It is a satiation where the events contain all elements based on the
definition of the events.

Union of events: The union of two events A and B, denoted by , consists of all

outcomes that are in A or in B or both A and B.

Intersection of events: The intersection of event A and B, denoted by , consists of

all outcomes that are in both A and B.

Compliment of an event: The compliment of event A, denoted by , consists of all

outcomes that are not in A.


Null event: The event containing no outcomes. It is the compliment o0f the sample space.

Probability of an event: The probability of event A, denoted by , is the probability

the outcome of the experiment is contained in A.


Independent: Two events said to be independent if knowing whether a specific one has
occurred does not change the probability that the other occurs.

5.3 Basic Principles of Counting


The following principle of counting will be basic to all our work

72 | P a g e minilikderse@[Link]
Introduction to statistics

The Addition Principle

If the choices can’t be performed together then the number of ways in which you can make a

choice in different ways. For example, if there are two way of bus to

voyage Awassa from Addis Ababa and three railways, then collectively we have

different way to arrive Awassa.

Multiplication Rule

Suppose that two experiments are to be performed. Then if experiment one can result in any one

of possible outcomes and if for each outcome of experiment one there are possible outcomes

of experiment two, then together there are possible outcomes of the two experiments.

Example 5.1: A small community consists of 10 women, each of whom has three children. If
one women and one of her children are to be chosen as mother and child
of the year, how many different choices are possible?

Solution:

By regarding the choice of the woman as the outcome of the first experiment and the subsequent
choice of one of her children as the outcome of the second experiment, we see from the basic

principle that there are possible choices.

The generalized basic principle of multiplication is that if experiments that are to be performed

are such that the first one may result in any of possible outcomes there are possible

outcomes the second experiment, and if for each of the possible outcomes of the first two

experiments there are possible outcomes of the third experiment, and if …, then there is a

total of .

73 | P a g e minilikderse@[Link]
Introduction to statistics

Example 5.2: How many different 7-place license plates are possible if the first 3 places are to
be occupied by letters and the final 4 by numbers?

Solution: By the generalized version of the basic principle the answer is

Example 5.3: In the above example, how many license plates would be possible if repetition
among letters or numbers were prohibited?

Solution: In this case there would be possible plates.

Permutation

How many different ordered arrangements of letters are possible? By direct enumeration

we see that there are 6: namely, and . Each arrangement is known as a

permutation. Thus, there are six possible permutations of a set of 3 objects. This result could also
have been obtained from the basic principle, since the first object in the permutation can be any
of the 3, the second object in the permutation can then be chosen any of the remaining 2, and the

third object in the permutation is then chosen the remaining one. Thus there are

possible permutations.

Suppose now that we have objects. Reasoning, similar to that we have just used for the 3 letter

shows that there are

Different permutations of the objects

Example 5.4: A class of stat 173 consists of 6 men and 4 women. An examination is given, and
the students are ranked according to their performance. Assume that no
two students obtain the same score.

74 | P a g e minilikderse@[Link]
Introduction to statistics

A. How many different rankings are possible?


B. If the men are ranked just among themselves and women among themselves,
how many different rankings are possible?

Solution:

A. As each ranking corresponds to a particular ordered arrangement of the 10 people, we see

that the answer to this part is

B. As there are possible rankings of the men among themselves and possible rankings

of the women among themselves, it follows from the basic principle that the two groups
arrange themselves; it follows the basic principle that the two groups arrange themselves

in way so that we have a total of possible rankings.

We shall now determine the number of permutations of a set of objects when certain of the

objects are indistinguishable from each other. Then the formula is:

Different permutations of objects, of which are alike are alike, …, are alike.

Example 5.5: How many different letter arrangements can be formed using the letter PEPPER?

Solution: possible teller arrangements.

Generally, if we are asked to arrange objects among objects, then we will have the following

total arrangements

75 | P a g e minilikderse@[Link]
Introduction to statistics

Example 5.6: Suppose a business man has a choice of five locations in which to establish his
business. He wishes to arrange only the top three locations. How many
different ways can he arrange them?

Solution:

Combination

We are often interested in determining the number of different groups of objects that could be

formed from a total of objects. A selection of objects without regard to order is called a

combination. That is, combinations are used when the order or arrangement is not important. The

number of combinations of objects selected from objects is denoted by and is given by

the formula

Example 5.7: From a group of 5 women and 7 men, how many different committees consisting
of 2 women and 3 men can be performed? What if 2 of the men are
feuding and refuse to serve on the committee together?

Solution:

As there are possible groups of 2 women, and possible groups of 3 me, it follows from

the basic principle that there are

76 | P a g e minilikderse@[Link]
Introduction to statistics

Possible committees consisting of 2 women and 3 men. On the other hand, if 2 of the men refuse

to serve on the committee together, then, as there are possible group of 3 men not

containing either of the 2 feuding men and groups of 3 men containing exactly 1 of the

feuding men, it follows that there are groups of 3 men not containing

both of the feuding men. Since there are ways to choose the 2 women, it follows that in this

case there are possible committees.

5.4 Probability of an Event


The probability of an event is denoted by where stands for probability and the dot stands

for any event, say A, B, G etc.

There are four approaches to calculate a probability of an event. These are

The classical approach


The frequentist approach
The axiomatic approach and
The subjective approach
The Classical Approach

If a procedure has different simple events, each with an equal chance of occurring, and event A

can occur in of these ways, then

77 | P a g e minilikderse@[Link]
Introduction to statistics

Assumptions in classical approach

The out comes must be equally-likely


The experiment should never be repeated more than once
The sample space should be finite

Example 5.8: Toss a fair coin once and find the probability of the occurrence of head

Solution: Since the sample space is finite i.e., either head or tail and the outcomes are
equally-likely

Example 5.9: For a card drawn from an ordinary deck, find the probability of getting a queen.

Solution: Since there are 4 queens and 52 cards

If one of the assumptions stated above is violated, the classical approach no longer valid

Frequentist Approach

This approach is called empirical approach. If after repetition of an experiment, where is

very large, an event is observed to occur in of these, then the probability of an event is or

conduct an experiment a large number of times, and count the number of times event A actually

occurs, then an estimate of is

78 | P a g e minilikderse@[Link]
Introduction to statistics

Example 5.10: Suppose a coin was tossed 1000 times and the result was 587 tails. The relative

frequency of tails is . Another 1000 tosses lead to 511 tails. Then the

relative frequency of tails is . Proceeding, in this manner

we obtain a sequence of numbers, which gets closer and closer to the number
defined as the probability of a trial in a single toss.

Therefore,

Axiomatic Approach

Both the classical and frequentist approaches have serious drawbacks, the first because the words
“equally likely” are vague and the second because the “large number” involved is vague.
Because of these difficulties, statisticians have been led to an axiomatic approach of probability.

Axiom 1: For every event A,

Axiom 2: For the sure or certain event,

Axiom 3: For any number of mutually exclusive events

In particular, for two mutually exclusive events

79 | P a g e minilikderse@[Link]
Introduction to statistics

Subjective Approach

A probability derived from an individual's personal judgment about whether a specific outcome
is likely to occur. Subjective probabilities contain no formal calculations and only reflect the
subject's opinions and past experience.

Subjective probabilities differ from person to person. Because the probability is subjective, it
contains a high degree of personal bias. An example of subjective probability could be asking
Arsenal fan, before the football season starts, the chances of Arsenal winning the world
champions. While there is no absolute mathematical proof behind the answer to the example,
fans might still reply in actual percentage terms, such as the Arsenal having a 95% chance of
winning the world champions.

5.5 Some Probability Rules


Rule 1: If then

Rule 2: For every event i.e. a probability between 0 and 1.

Rule 3: For the empty set, i.e. the impossible event has probability zero.

Rule 4: If is the complement of A, then

Rule 5: If and are any two events, then

More generally, if are three events, then

Rule 6:

, for independent event

Example 5.11: Suppose we toss two coins and suppose that each of the four points in the
sample space is equally likely and hence has
probability . Let E is the event that the first coin falls head, and F is the
event that the second coin falls heads.

Solution: Then the probability of either the first or the second


coin falls head.

80 | P a g e minilikderse@[Link]
Introduction to statistics

5.5 Conditional Probability and Independence


Let A and B two events such that . Denote the probability of B given that A has
occurred since A is know to have occurred; it becomes the new sample replacing the original S.
From this we are led to the definition

In words, this is saying that the probability that both A and B occur is equal to the probability
that A occurs times the probability that B occurs given that has occurred. We call the
conditional probability of B given A, i.e. the probability that B will occur given that A has
occurred.

Example 5.12: A jar contains black and white marbles. Two marbles are chosen without
replacement. The probability of selecting a black marble and then a white
marble is 0.34, and the probability of selecting a black marble on the first
draw is 0.47. What is the probability of selecting white marble on the second
draw, given that the first marble drawn was black?

Solution:

Example 5.13: The probability that it is Friday and that a student is absent is 0.03. Since there
are 5 schooldays in a week, the probability that it is Friday is 0.2. What is
the probability that a student is absent given that today is Friday?

Solution:

81 | P a g e minilikderse@[Link]
Introduction to statistics

It often happens that the knowledge that a certain event E has occurred has no effect on the
probability that some other event F has occurred, that is, that . One would
expect that in this case, the equation would also be true. If these equations are
true, we might say the F is independent of E.

Definition: Two events E and F are independent if both E and F have positive probability and if

Note that: If then E and F are independent if and only if

Example 5.14: Suppose that we roll a pair of fail dice, so each of the 36 possible out come is
equally likely. Let A denotes the event that the first die lands on 3, let C be the
event that the sum of the dice is 7

A. Are A and B independent?


B. Are A and C independent
Solution:

A. Since is the event that the first die lands on 3 and the second on 5, we see that

On the other hand

and

Therefore, since we see that and so events A

and B are not independent

82 | P a g e minilikderse@[Link]
Introduction to statistics

B. Events A and C are independent. This is seen by noting that

While and .Therefore, and so

events A and C are independent.

Exercise:

1. A box contains 3 white balls and 5 black balls. We extract 2 balls from the box,
one after the other. Give a sample space for this experiment and the probabilities
of the elementary events of this sample space.

PROBABILITY DISTRIBUTION

“Free yourself from the rigid conduct of tradition and open yourself to the new
forms of probability.”
HANS BENDER

6.1 The Definition of Random Variable and Probability Distribution


Before probability distribution is defined formally, the definition of reviewed. In the first
chapter, a variable was defined as a characteristic or attribute that can assume different values
various letter of the alphabet are used to represent the variables. Since the variables in this
chapter are associated with probability, they are called random variables.

A random variable is a variable whose values are determined by chance. A random variable can
be either continuous or discrete. By define the random variable; we can assign the number to a
random variable.

Example 6.1: Suppose we are about to learn the sexes of the three children of a certain family.
The sample space of this experiment consists of the following 8 outcomes.

83 | P a g e minilikderse@[Link]
Introduction to statistics

The outcomes means, for instance that the youngest child is a girl, the next youngest is a

boy, and the oldest is a boy. Suppose that each of these 8 possible outcomes is equally likely,
and so each has probability 1/8. If we let x denote the number of female children in this family,
then the value of x is determined by the outcomes of the experiment. That is, x is a random

variable whose value will be . We now determine the probabilities that x will equal

each of these four values. Since X will equal 0 if the outcome is we see that

Since X will equal 1 if the outcome is or , we have

Similarly,

A probability distribution consists of the values a random variable can assume and the
corresponding probabilities of the values. The probabilities are determined theoretically or by
observation. The probability distribution can be denoted by is used to represent the
probability that is equal to . The sum of the probability distribution is one. That is,

Example 6.2: Suppose we toss a coin three times, the sample space is represented as
and if the random
variable for the number of heads.

A. Assign a value for a random variable

84 | P a g e minilikderse@[Link]
Introduction to statistics

B. Find the probability distribution for A


Solution:

A. Once a random variable, say , is defined as the number of heads,


B.
Number of heads 0 1 2 3

Probability

We can check that

Example 6.3: Suppose that is a random variable that takes on one of the value . If
and . What is ?

Solution: Since the probability must sum to 1, we have

Example 6.4: A sales women has scheduled two appointments to sell encyclopedias. She feels
her first appointments will lead to a sale with probability 0.3. She also
feels that the second will lead to a sale with probability 0.6 and that the
results from the two appointments are independent. What is the probability distribution
of , the number of sales made?

Solution: The random variable can take on any of the value . It will equal 0 if neither
appointment leads to a sale, and so

85 | P a g e minilikderse@[Link]
Introduction to statistics

The random variable will equal 1 either if there is a sale on the first and not on the second
appointment or if there is no sale on the first and one sale on the second appointment. Since these
two events are disjoint, we have

Finally, the random variable will equal 2 if both appointments result in sales; thus

As check on this result, we note that

Two requirements for a probability distribution

1. The sum of the probabilities of all the events in the sample space must equal 1; that
is,
2. The probability of each event in the sample space must be between or equal to 0
and 1. that is,

6.2 Introduction to Expectation- Mean and Variance of a


Random Variable
A key concept in probability is the expected value of a random variable. If is a discrete random
variable that takes on one of the possible values then the expected value of ,
denoted by , is defined by

If is continuous random variable

86 | P a g e minilikderse@[Link]
Introduction to statistics

Where is probability density function in the case of discrete random variable its name will
change to probability mass function (pmf).

Example 6.5: Find the expected value of the following random variable

0 1 2 3 4
0.18 0.34 0.23 0.21 0.04

Solution:

Note that: The expected value of a random variable is the same as with the mean of a
random variable

Suppose that we are given a discrete random variable random variable along with its probability

mass function, and that we want to compute the expected value of some function of , say .

How can we accomplish this? One way is follows: Since is determined from the probability

mass function of . Once we have determined the probability mass function of we can

compute by using the definition of expected value.

Example 6.6: Let X denote a random variable that takes on any of the values -1, 0, 1 with
respective probability

87 | P a g e minilikderse@[Link]
Introduction to statistics

Compute

Solution: Letting , it follows that the probability mass function of is given by

Hence

The reader should note that

If and are constants then

The expected value of a random variable is also referred to as the mean or the first

moment of . The quantity is called the moment of .

By definition

Example 6.7: The following are the annual income of 7 men and 7 women residents of a certain
community.

Annual income (in $ 1000)

88 | P a g e minilikderse@[Link]
Introduction to statistics

Men Women

33.5 24.2

25.0 19.5

28.6 27.4

41.0 28.6

30.5 32.2

29.6 22.4

32.8 21.6

Suppose that a woman and a man randomly chosen. Find the expected value of the sum
of their incomes.

Solution: Let be the man’s income and Y is the woman’s income. Since is equally likely to

be any of the values in the men’s column, we see that

Similarly,

Therefore, the expected value of the sum of their incomes is

89 | P a g e minilikderse@[Link]
Introduction to statistics

That is, the expected value of the sum of their incomes is approximately $ 56,700.

If is a random variable with mean μ, then the variance of x, denoted by is defined by

An alternative formula for is derived as follows

That is,

Example 6.8: The return from a certain investment is a random variable X with probability
distribution.

Find the variance of the return.

Solution: Let us first compute that expected return as follows:

90 | P a g e minilikderse@[Link]
Introduction to statistics

To compute we use the formula

Now, since will equal with respective probabilities of

we have

Therefore,

Properties of Variance

1. For any random variance X and constant C, it can be shown that

2. If and are independent random variable

91 | P a g e minilikderse@[Link]
Introduction to statistics

3. The square root of the is called the standard deviation of , and we denote it by

. That is,

6.3 Common Discrete Probability Distribution


Binomial Distribution

Many types of probability problems have only two outcomes, or they can be reduced to two
outcomes. For example, when a coin is tossed, it can land heads or tails.

A probability experiment is a probability experiment that satisfies the following four


requirements:

1. Each trial can have only two outcomes or outcomes that can be reduced to two
outcomes.

2. There must be a fixed number of trials

3. The outcomes of each trial must be independent

4. The probability of a success must remain the same for each trial

The out comes of a binomial experiment and the corresponding probabilities of these outcomes
are called a binomial distribution

The probability mass function of a binomial random variable having parameter (n, p) is given by

Example 6.9: Five fair coins are flipped. If the outcomes are assumed independent, find
the probability of the number of heads obtained

92 | P a g e minilikderse@[Link]
Introduction to statistics

Solution: If we let equal the number of heads (successes) parameters .

Hence,

Example 6.10:

A. Determine when is a Binomial random variable with parameters

and

B. Determine when is a Binomial random variable with parameters and

Solution:

A.

B.

If is Binomial random variable with parameter and , then

93 | P a g e minilikderse@[Link]
Introduction to statistics

Example6.11: Suppose that each screw produced is independently defective with


probability 0.01. Find the expected value and variance of the number of
defective screws in a shipment of size 1000.

Solution: The number of effective screws in the shipment of size 1000 is a Binomial random
variable with parameters , . Hence, the expected number of
defective screws is and the variance of
the number of detective screws is

The Poisson Distribution

A discrete probability distribution that is useful when is large and is small and when the independent
variable occurs over a period of time is called the Poisson distribution, name for Simeon D. Poisson

(1781-1840). In addition to being used for the stated conditions (i.e. is large, is small, and the

variable occur over a period of time), the Poisson distribution can be used when a density of items is
distributed over a given area or volume, such as the number of plants growing per acre of woods or the
number of defects in a given length of videotape.

If is Poisson random variable with parameter , then

Example6.12: If is a Poisson random variable with parameter , find

Solution:

Using the fact that , we obtain

Both the expected value and the variance of a Poisson random variable are equal to . That is, we have the
following. If is a Poisson random variable with parameter , ; then

94 | P a g e minilikderse@[Link]
Introduction to statistics

Example6.13: Suppose the average number of accidents occurring weekly on a particular high way is
equal to 1.2. Approximate the probability that there is at least one accident this
week.

Solution: Let denote the number of accidents because it is reasonable to suppose that there are a large
number of cars passing along the high way, each having a small probability of being involved in
an accident, the number of such accidents should be approximately a Poisson random variable.
That is, if denotes the number of accidents that will occur this week, then is approximately
Poisson random variable with mean value . The desired probability is now obtained as
follows.

Therefore, there is approximately a 70% chance that there will be at least one accident
this week.

We can approximate Binomial distribution to Poisson distribution if is large and is too small.
Thus, the approximately Poisson distribution has a parameter.

Example 6.14: Suppose that items produced by a certain machine are independently
defective with probability [Link] is the Poisson approximation for this
probability?

Solution: If we let denote the number of defective items, then is a Binomial random variable
with parameters and . Thus the desired probability is

95 | P a g e minilikderse@[Link]
Introduction to statistics

Since , the Poisson approximation yields the value.

Thus, even in this case, where is equal to 10 (which is not that large) and is equal
to 0.1 (which is not that small), the Poisson approximation to the Binomial
probability is quite accurate.

6.4 Common Continuous Probability Distribution


Every continuous random variable has a curve associated with it. This curve, formally known
as a probability density function, can be used to obtain probabilities associated with the random
variable. This is accomplished as follows, consider any two points and , where is less than
. The probability that assumes a value that lies between and is equal to the area under the
curve between and . That is,

Since must assume some value, it follows that the total area under the density curve must
equal 1. Also, since the area under the graph of the probability density function between points
and is the same regardless of whether the end points and are themselves included. That is,

Normal Random Variables

96 | P a g e minilikderse@[Link]
Introduction to statistics

The most important type of random variable is the normal random variable. The probability
density function of a normal random variable is determined by two parameters: the expected
value and the standard deviation of . We designate these values as and , respectively.

And

The normal probability density function is a bell-shaped density curve that is symmetric about
the value ; its variability is measured by . The larger is, the more variability there is in the
curve.

Since the probability density function of a normal random variable is symmetric about its

expected value ; it follows that is equally likely to be on either side of . That is,

Not all bell-shaped symmetric density curves are normal. The normal density curves are
specified by a particular formula:

A normal random variable having mean value 0 and standard deviation 1 is called a standard

normal variable, and its density curve is called the standard normal curve. The letter represents

a standard normal random variable.

Probabilities Associated with a Standard Normal Random Variable

97 | P a g e minilikderse@[Link]
Introduction to statistics

Once the values are transformed by using the above formula, they are called value is

actually the number of standard deviations that a particular value is a way from the mean.

Steps to find areas under the normal distribution curve

1. Between 0 and any value: Look up the value in the table to get the area

2. In any tail

a. Look up the value to get the area

b. Subtract the area from 0.5

3. Between values on the same side of the mean

a. Look up both values to get the area

b. Subtract the smaller area from the larger area

4. Between two values on opposite sides of the mean

a. Look up both values to get the area

b. Add the areas

5. Less than any value to get the right of the mean

a. Look up the value to get the area

b. Add 0.5 to the area

6. Greater than any value to the left of the mean

a. Look up the value in the table to get the area

b. Add 0.5 to the area


7. In any two tailed

98 | P a g e minilikderse@[Link]
Introduction to statistics

a. Look up values in the table to get the areas

b. Subtract both areas from 0.5


c. Add the answer

General procedure is

Draw the picture


Shade the area desired
Find the correct figure
Follow the direction

Example 6.15: Find the area under the normal distribution curve between and

Solution: Draw the area as follows:

Since table gives the area between 0 and any value to the right of 0, one need look up the

value in the table. Find 2.3 in the left column and 0.04 in the top row. The value where the column and

row meet in the table is the answer, 0.4904.

0.00 0.01 0.02 0.03 0.04 …


0.0
0.1
0.2

99 | P a g e minilikderse@[Link]
Introduction to statistics

2.2
2.3 0.4904

Example 6.16: Find

A.

B.

Solution:

A. Draw the area as follows:

0
0 1.5

100 | P a g e minilikderse@[Link]
Introduction to statistics

B. Draw the required area as follows:

0 0.8

0 0.8 0 0 0.8

Example 6.17: Find

A.

B.

Solution:

A. Draw the graph as follows:

0 1 2

101 | P a g e minilikderse@[Link]
Introduction to statistics

0 1 2 0 1
0 2 2

B. Draw the graph as follows:

-1.5 0 2.5

0 2.5
-1.50
-1.50 2.5

Since , due to symmetric property of normal distribution

Finding Normal Probabilities: Conversion to the Standard Normal

102 | P a g e minilikderse@[Link]
Introduction to statistics

Let be a normal random variable with mean and standard deviation . We can determine

probabilities concerning by using the fact that the variable defined by

has a standard normal distribution.

We can compute any probability statements in terms of . For example,

where is a standard normal random variable

Example 6.18: IQ examination scores for sixth-graders are normally distributed with mean value
100 and standard deviation 14.2.

A. What is the probability a randomly chosen sixth-grader has a score greater


than 130?
B. What is the probability a randomly chosen sixth-grader has score between
90 and 115?

Solution: Let denote the score of a randomly chosen student. We compute probabilities

concerning by making use of the fact that the standardized variable

has a standard normal distribution

103 | P a g e minilikderse@[Link]
Introduction to statistics

A.

B. The inequality is equivalent to

Or equivalently,

Therefore,

Properties of the Normal distribution

1. The normal distribution curve is bell-shaped


2. The mean, median and mode are equal and located at the center of the distribution
3. The normal distribution curve is unimodal
4. The curve is symmetrical about the mean, which is equivalent to saying that is shape the
same on both sides of vertical line passing through the center
5. The curve is continuous. That is, no gaps or holes

6. The curve never touches the axis

7. The total area under the normal distribution curve is equal to 1

Relation between Binomial and Normal Distribution

Normal distribution is a limiting case of the Binomial probability distribution under the
following condition:

104 | P a g e minilikderse@[Link]
Introduction to statistics

I. , the number of trial is indefinitely large

II. Neither and is very small

We know that for a Binomial variate with parameters and

De-Moivre provide that under the above two conditions, the distribution of standard Binomial
variate

tends to the distribution of standard normal distribution. If and are nearly equal (i.e., is

nearly 0.5), then the normal approximation is surprisingly good even for small values of .

Relation between Poisson and Normal Distribution

If is a random variable following Poisson distribution with parameter , then

Thus standard Poisson variate becomes

It has been proved that this variate tends to be a standard normal variate if

105 | P a g e minilikderse@[Link]
Introduction to statistics

Chi-square Distribution:

The square of a standard normal variable is called a chi-square variate with one degree of

freedom. Thus if is a random variable following normal distribution with mean and standard

deviation , then is a standard normal variate. is a chi-square variate with 1 degree

of freedom.

If are independent random variables following normal distribution with means

and standard deviations respectively then the variate

this is the sum of the square of independent standard normal variates, follows chi-square

distribution with degree of freedom.

Applications of chi-square distribution

Chi-square distribution has a number of applications. Some of which are listed below

Chi-square test of goodness of fit


Chi-square test for independence of attributes
To test whether the population has a specified value of the variance

Student’s distribution

106 | P a g e minilikderse@[Link]
Introduction to statistics

It is often the case that one wants to calculate the size of sample needed to obtain a certain level
of confidence in survey results. Unfortunately, this calculation requires prior knowledge of the

population standard deviation ( ). Realistically, is unknown. Often a preliminary sample will

be conducted so that a reasonable estimate of this critical population parameter can be made. If
such a preliminary sample is not made, but confidence intervals for the population mean are to

be constructing using an unknown , then the distribution known as the Student t distribution can

be used.

First, a little history about this curious name. William Gosset (1876-1937) was a Guinness
Brewery employee who needed a distribution that could be used with small samples. Since the
Irish brewery did not allow publication of research results, he published under the pseudonym of
Student. We know that large samples approach a normal distribution. What Gosset showed was
that small samples taken from an essentially normal population have a distribution characterized
by the sample size. The population does not have to be exactly normal, only unimodal and
basically symmetric. This is often characterized as heap-shaped or mound shaped.

Following are the important properties of the Student t distribution.

1. The Student t distribution is different for different sample sizes.


2. The Student t distribution is generally bell-shaped, but with smaller sample sizes shows
increased variability (flatter). In other words, the distribution is less peaked than a normal
distribution and with thicker tails. As the sample size increases, the distribution
approaches a normal distribution. For n > 30, the differences are negligible.

3. The mean is zero (much like the standard normal distribution).

4. The distribution is symmetrical about the mean.

5. The variance is greater than one, but approaches one from above as the sample size

increases ( =1 for the standard normal distribution).

6. The population standard deviation is unknown.

107 | P a g e minilikderse@[Link]
Introduction to statistics

7. The population is essentially normal (unimodal and basically symmetric)

108 | P a g e minilikderse@[Link]

You might also like