Module of Basic Statistics
Module of Basic Statistics
SCIENCE
DEPARTMENT OF STATISTICS
BU, BONGA
I
Basic Statistics, BU, CNSc, 2015
Table of Contents Page
CHAPTER 1 ................................................................................................................... 1
1. THE NATURE OF PROBABILITY AND STATISTICS .................................... 1
1.1 INTRODUCTION ................................................................................................. 1
1.2 Objectives .............................................................................................................. 1
1.3 Definition and classifications of statistics ............................................................... 1
1.3.1Definition......................................................................................................... 1
1.3.2 Classifications ................................................................................................. 2
1.4 Stages in Statistical Investigation ........................................................................... 2
1.5 Definitions of some terms ...................................................................................... 3
1.6 Types of Variables or Data .................................................................................... 3
1.7 Applications, Uses and Limitations of statistics ..................................................... 4
1.7.1 Applications of statistics ................................................................................. 4
1.7.2 Uses of statistics ............................................................................................. 4
1.7.3 Limitations of statistics ................................................................................... 4
1.8 Scales of measurement ........................................................................................... 5
1.8.1 Order .............................................................................................................. 5
1.8.2 Distance .......................................................................................................... 5
1.8.3 Fixed Zero ...................................................................................................... 6
1.8.4 SCALE TYPES .............................................................................................. 6
1.9. Introduction to Methods of Data Collection .......................................................... 8
Summary ......................................................................................................................... 9
Exercise 1...................................................................................................................... 10
CHAPTER 2 ............................................................................................................. 13
METHODS OF DATA COLLECTION AND PRESNTATION ................................ 13
2.1 Introduction ......................................................................................................... 13
2.2 Objectives ............................................................................................................ 13
2.3 Methods of Data Presentation .............................................................................. 13
2.3.1Categorical frequency Distribution ................................................................. 14
2.3.2Ungrouped frequency Distribution: ................................................................ 15
2.2.3 Grouped frequency Distribution: ................................................................... 17
2.4 Diagrammatic and Graphic presentation of data ................................................... 21
2.3.1Diagrammatic presentation of data ................................................................. 21
2.3.2 Graphical Presentation of data ....................................................................... 26
Summary ....................................................................................................................... 28
II
Basic Statistics, BU, CNSc, 2015
Exercise 2...................................................................................................................... 29
CHAPTER 3 ............................................................................................................. 31
3. MEASURES OF CENTERAL TENDENCY ......................................................... 31
3.1 Introduction ......................................................................................................... 31
3.2 Objectives ............................................................................................................ 31
3.3. The Summation Notation .................................................................................... 31
3.3.1 Properties of Summation ............................................................................... 33
3.4 Types of measures of central tendency ................................................................. 34
3.4.1 Mean............................................................................................................. 35
3.4.2The Mode ..................................................................................................... 46
3.4.3The Median.................................................................................................... 48
3.4.4 Quantiles ....................................................................................................... 52
Summary ....................................................................................................................... 60
Exercise- 3 .................................................................................................................... 61
CHAPTER 4 ................................................................................................................. 63
4. Measures of Dispersion (Variation) ........................................................................... 63
4.1 Introduction ......................................................................................................... 63
4.2 Objectives ............................................................................................................ 63
4.3 Absolute and Relative Measures of Dispersion .................................................... 63
4.4 Types of Measures of Dispersion ......................................................................... 63
4.4.1The Range (R) ............................................................................................... 64
4.4.2 Relative Range (RR) ..................................................................................... 65
4.4.3 The Quartile Deviation (Semi-inter quartile range) ........................................ 65
4.4.4 Coefficient of Quartile Deviation (C.Q.D)..................................................... 66
4.4.5 The Mean Deviation (M.D): .......................................................................... 67
4.4.6 Coefficient of Mean Deviation (C.M.D) ........................................................ 70
4.4.7 The Variance................................................................................................. 71
Summary ....................................................................................................................... 85
Exercise 4...................................................................................................................... 85
CHAPTER 5 ................................................................................................................. 87
5. ELEMENTARY PROBABILITY ............................................................................. 87
5.1 Introduction ......................................................................................................... 87
5.2 Objectives ............................................................................................................ 87
5.3 Definitions of some probability terms .................................................................. 87
5.4 Counting Rules .................................................................................................... 89
III
Basic Statistics, BU, CNSc, 2015
5.4.1 The Multiplication Rule: ............................................................................... 90
5.4.2 Permutation................................................................................................... 91
5.4.3 Combination ................................................................................................. 93
5.5 Approaches to measuring Probability ................................................................... 95
5.5.1 The classical approach .................................................................................. 95
5.5.2 The Frequents Approach ............................................................................... 99
5.5.3 Axiomatic Approach: .................................................................................... 99
5.6 Conditional probability and Independency ......................................................... 100
5.6.1 Conditional probability of an event ............................................................. 101
5.6.2 Probability of Independent Events ............................................................... 102
Summary ..................................................................................................................... 103
CHAPTER 6 ............................................................................................................... 105
6. RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS....................... 105
6.1 Introduction ....................................................................................................... 105
6.2 Objectives .......................................................................................................... 105
6.3 Discrete random variable: .................................................................................. 106
6.4 Continuous random variable .............................................................................. 106
6.5 Properties of Probability Distribution: ................................................................ 107
6.6Introduction to expectation ................................................................................. 108
6.7 Mean and Variance of a random variable ........................................................... 110
6.8 Common Discrete Probability Distributions ....................................................... 111
6.8.1 Binomial Distribution ................................................................................. 111
6.8.2Poisson Distribution ..................................................................................... 115
6.9 Common Continuous Probability Distributions .................................................. 117
6.9.1 Normal Distribution .................................................................................... 117
6.9.2 Properties of the Standard Normal Distribution: .......................................... 118
Summary ..................................................................................................................... 124
Exercise 6.................................................................................................................... 125
CHAPTER 7 ............................................................................................................... 128
7. SAMPLING AND SAMPLING DISTRIBUTION .................................................. 128
7.1 Introduction ....................................................................................................... 128
7.2 Objectives .......................................................................................................... 128
7.2Errors in sample survey ...................................................................................... 129
7.3 Random Sampling or probability sampling. ....................................................... 129
7.3.1Simple Random Sampling: .......................................................................... 129
IV
Basic Statistics, BU, CNSc, 2015
7.3.2Stratified Random Sampling: ....................................................................... 130
7.3.3Cluster Sampling ......................................................................................... 130
7.3.4 Systematic Sampling: ................................................................................. 130
7.4 Non Random Sampling or non-probability sampling. ......................................... 131
7.4.1 Judgment Sampling ..................................................................................... 131
7.4.2 Convenience Sampling................................................................................ 131
7.4.3 Quota Sampling .......................................................................................... 131
7.5 Sampling Distribution of the sample mean ......................................................... 132
7.6 Steps for the construction of Sampling Distribution of the mean ........................ 132
7.7 Central Limit Theorem ...................................................................................... 135
Summary ..................................................................................................................... 135
Exercises 7 .................................................................................................................. 136
CHAPTER 8 ............................................................................................................... 138
8. ESTIMATION AND HYPOTHESIS TESTING...................................................... 138
8.1 Introduction ....................................................................................................... 138
8.2 Objectives .......................................................................................................... 139
8.3 Statistical Estimation ......................................................................................... 139
8.3.1Point Estimation........................................................................................... 140
8.3.2 Interval estimation ...................................................................................... 140
8.4Point and Interval estimation of the population mean: µ ...................................... 141
8.4.1 Point Estimation .......................................................................................... 141
8.5 Hypothesis Testing ............................................................................................ 145
8.5.1 Null hypothesis: .......................................................................................... 145
8.5.2Alternative hypothesis:................................................................................. 145
8.6 Types and size of errors: .................................................................................... 145
8.6.1 General steps in hypothesis testing: ............................................................. 146
8.7 Hypothesis testing about the population means: ................................................. 146
8.8 Test of Association ............................................................................................ 151
8.9 Decision Rule .................................................................................................... 153
Summary ..................................................................................................................... 157
Exercise 8.................................................................................................................... 159
CHAPTER 9 ............................................................................................................... 162
9. SIMPLE LINEAR REGRESSION AND CORRELATION ..................................... 162
9.1 Introduction ....................................................................................................... 162
9.2 Objectives .......................................................................................................... 162
V
Basic Statistics, BU, CNSc, 2015
9.3 Simple Correlation ............................................................................................. 162
Steps........................................................................................................................ 166
9.4 Simple Linear Regression .................................................................................. 168
Summary ..................................................................................................................... 175
Exercise 9.................................................................................................................... 176
ANSWER FOR SELECTED EXERCISE ................................................................... 177
References ................................................................................................................... 183
Appendix: Tables ........................................................................................................ 184
A. The Standard Normal Distribution Table............................................................. 184
B. The Student‟s t-distribution Table ....................................................................... 185
C. The Chi-Square distribution Table ...................................................................... 186
VI
Basic Statistics, BU, CNSc, 2015
CHAPTER 1
1. THE NATURE OF PROBABILITY AND STATISTICS
1.1 INTRODUCTION
The first unit of this module introduces you the concept of statistics and methods of data
collection. You may be familiar with probability and statistics through radio, television,
newspapers, and magazines. The word statistics is derived from the Latin word status
which means a political state or government. It was originally applied in connection
with kings and monarchs collecting data on their citizenry which pertained to state
wealth, taxes collected population and so on.
1.2 Objectives
After completing this chapter, Student should be able to
1
Basic Statistics, BU, CNSc, 2015
Statistics is defined as the science of collecting, organizing, presenting, analyzing
and interpreting numerical data for the purpose of assisting in making a more
effective decision.
1.3.2 Classifications
Depending on how data can be used statistics is sometimes divided in to two main areas
or branches.
1. Collection of data: the process of measuring, gathering, assembling the raw data up
on which the statistical investigation is to be based.
Data can be collected in a variety of ways; one of the most common methods
is through the use of survey. Survey can also be done in different methods,
three of the most common methods are:
Telephone survey
Mailed questionnaire
Personal interview.
Activity: discuss the advantage and disadvantage of the above three methods with respect
to each other.
2
Basic Statistics, BU, CNSc, 2015
4. Analysis of data: The process of extracting relevant information from the
summarized data, mainly through the use of elementary mathematical operation.
5. Inference of data: The interpretation and further observation of the various statistical
measures through the analysis of the data by implementing those methods by which
conclusions are formed and inferences made.
Statistical techniques based on probability theory are required.
1.5 Definitions of some terms
a. Statistical Population: It is the collection of all possible observations of a specified
characteristic of interest (possessing certain common property) and being under
study. An example is all of the students in BHU those who take stat 2011 course.
b. Sample: It is a subset of the population, selected using some sampling technique in
such a way that they represent the population.
c. Sampling: The process or method of sample selection from the population.
d. Sample size: The number of elements or observation to be included in the sample.
e. Census: Complete enumeration or observation of the elements of the population. Or
it is the collection of data from every element in a population
f. Parameter: Characteristic or measure obtained from a population.
g. Statistic: Characteristic or measure obtained from a sample.
h. Variable: It is an item of interest that can take on many different numerical values.
1.6 Types of Variables or Data
1. Qualitative Variables are nonnumeric variables and can't be measured. Examples
include gender, religious affiliation, and state of birth.
2. Quantitative Variables are numerical variables and can be measured.
3
Basic Statistics, BU, CNSc, 2015
1.7 Applications, Uses and Limitations of statistics
1.7.1 Applications of statistics
Statistics can be applied in any field of study which seeks quantitative evidence. The
applications areas of statistics are:
The main function of statistics is to enlarge our knowledge of complex phenomena. The
following are some uses of statistics:
4
Basic Statistics, BU, CNSc, 2015
1.8 Scales of measurement
Proper knowledge about the nature and type of data to be dealt with is essential in order to
specify and apply the proper statistical method for their analysis and inferences.
Measurement scale refers to the property of value assigned to the data based on the properties
of order, distance and fixed zero. In mathematical terms measurement is a functional
mapping from the set of objects {Oi} to the set of real numbers {M(Oi)}.
The goal of measurement systems is to structure the rule for assigning numbers to objects
in such a way that the relationship between the objects is preserved in the numbers
assigned to the objects. The different kinds of relationships preserved are called
properties of the measurement system.
1.8.1 Order
The property of order exists when an object that has more of the attribute than another
object, is given a bigger number by the rule system. This relationship must hold for all
objects in the "real world".
The property of ORDER exists when for all i, j if Oi > Oj, then M(Oi) > M(Oj).
1.8.2 Distance
The property of distance is concerned with the relationship of differences between
objects. If a measurement system possesses the property of distance it means that the unit
of measurement means the same thing throughout the scale of numbers. That is, an inch
is an inch, no matters where it falls -immediately ahead or a mile downs the road.
5
Basic Statistics, BU, CNSc, 2015
More precisely, an equal difference between two numbers reflects an equal difference in
the "real world" between the objects that were assigned the numbers. In order to define
the property of distance in the mathematical notation, four objects are required: O i, Oj,
Ok, and Ol . The difference between objects is represented by the "-" sign; Oi - Oj refers to
the actual "real world" difference between object i and object j, while M(O i) - M(Oj)
refers to differences between numbers. The property of DISTANCE exists, for all i, j, k,
l, if Oi-Oj ≥ Ok- Ol then M(Oi)-M(Oj) ≥ M(Ok)-M( Ol ).
1.8.3 Fixed Zero
A measurement system possesses a rational zero (fixed zero) if an object that has none of
the attribute in question is assigned the number zero by the system of rules. The object
does not need to really exist in the "real world", as it is somewhat difficult to visualize a
"man with no height". The requirement for a rational zero is this: if objects with none of
the attribute did exist would they be given the value zero. Defining O 0 as the object with
none of the attribute in question, the definition of a rational zero becomes:
The property of FIXED ZERO exists if M(O0) = 0. The property of fixed zero is
necessary for ratios between numbers to be meaningful.
1.8.4 SCALE TYPES
Measurement is the assignment of numbers to objects or events in a systematic fashion.
Four levels of measurement scales are commonly distinguished: nominal, ordinal,
interval, and ratio. Each possessed different properties of measurement systems.
[Link] Nominal Scales
Nominal scales are measurement systems that possess none of the three properties stated
above. In nominal scales:
The level of measurement which classifies data into mutually exclusive, all
inclusive categories in which no order or ranking can be imposed on the data.
There is no arithmetic and relational operation can be applied.
Examples:
o Political party preference (Republican, Democratic, or Other,)
o Sex (Male or Female.)
o Marital status(married, single, widow, divorce)
o Country code
o Regional differentiation of Ethiopia.
6
Basic Statistics, BU, CNSc, 2015
[Link] Ordinal Scales
Ordinal Scales are measurement systems that possess the property of order, but not the
property of distance. The property of fixed zero is not important if the property of
distance is not satisfied. In ordinal scales:
The level of measurement which classifies data into categories that can be ranked
and differences between the ranks do not exist.
Arithmetic operations are not applicable but relational operations are applicable.
Ordering is the sole property of ordinal scale.
Examples:
o Letter grades (A, B, C, D, F).
o Rating scales (Excellent, Very good, Good, Fair, poor).
o Military status.
[Link] Interval Scales
Interval scales are measurement systems that possess the properties of Order and
distance, but not the property of fixed zero. In Interval scales:
The level of measurement which classifies data that can be ranked and
differences are meaningful. However, there is no meaningful zero, so ratios
are meaningless.
All arithmetic operations except division are applicable.
Relational operations are also possible.
Examples:
o IQ
o Temperature in Fo.
7
Basic Statistics, BU, CNSc, 2015
All arithmetic and relational operations are applicable.
Examples:
o Weight
o Height
o Number of students
o Age
1.9. Introduction to Methods of Data Collection
What are the sources of data you know?
1. Primary Data
Data measured or collected by the investigator or the user directly from
the source.
Two activities involved: planning and measuring.
a) Planning:
Identify source and elements of the data.
Decide whether to consider sample or census.
If sampling is preferred, decide on sample size, selection method,… etc
Decide measurement procedure.
Set up the necessary organizational structure.
b) Measuring: There are different options.
Focus Group
Telephone Interview
Mail Questionnaires
Door-to-Door Survey
Mall Intercept
New Product Registration
Personal Interview and
Experiments are some of the sources for collecting the primary data.
2. Secondary Data: are data gathered or compiled from published and unpublished
sources or files.
When our source is secondary data check:
8
Basic Statistics, BU, CNSc, 2015
The type and objective of the situations.
The purpose for which the data are collected and compatible with the
present problem.
The nature and classification of data is appropriate to our problem.
There are no biases and misreporting in the published data.
Note: Data which are primary for one may be secondary for the other.
Summary
• The two major areas of statistics are descriptive and inferential. Descriptive statistics
includes the collection, organization, summarization, and presentation of data. Inferential
statistics includes making inferences from samples to populations, estimations and
hypothesis testing, determining relationships, and making predictions. Inferential
statistics is based on probability theory.
• Since in most cases the populations under study are large, statisticians use subgroups
called samples to get the necessary data for their studies. There are four basic methods
used to obtain samples: random, systematic, stratified, and cluster.
• There are two basic types of statistical studies: observational studies and experimental
studies. When conducting observational studies, researchers observe what is happening or
what has happened and then draw conclusions based on these observations. They do not
attempt to manipulate the variables in any way.
• Finally, the applications of statistics are many and varied. People encounter them in
everyday life, such as in reading newspapers or magazines, listening to the radio, or
watching television. Since statistics is used in almost every field of endeavor, the
9
Basic Statistics, BU, CNSc, 2015
educated individual should be knowledgeable about the vocabulary, concepts, and
procedures of statistics. Also, everyone should be aware that statistics can be misused.
Exercise 1
1. Define the following terms.
a. Statistics.
b. Descriptive statistics
c. Qualitative variable
d. Nominal scale
e. Sample survey
2. To assess the opinion of students at the Bule Hora University about Cafeteria
Safety, the Ethiopian television reporter interviews 20 students he meets walking on the
campus late at night that are willing to give their opinion.
The report of head of the minister about the American Soldiers in Iraq terrorist attack
mission dismissed terrorists 30% at first campaign, 60% at second campaign and 59% at
third campaign.
b. Assume that in your class there are 50 students. Take their CGPA for all 50 students
and analysis mean CGPA; that is assumed 3.00.
d. Of 1100 voters in Senator‟s district: 400 strongly favor his bill; 300 favor; 200 neutral;
150 do not favor, and 50 strongly do not favor his bill.
e. Patients may be characterized as unimproved, improved & much improved.
f. The height of the men in Bule Hora University.
g. Your score on an individual intelligence test as a measure of your intelligence.
7. The following present a list of different attributes and rules for assigning numbers to
objects. Try to classify the different measurement systems into one of the four types of
scales.
a. Your checking account number as a name for your account.
b. Your checking account balance as a measure of the amount of money you have in that
account.
11
Basic Statistics, BU, CNSc, 2015
c. The order in which you were eliminated in a spelling bee as a measure of your
spelling ability.
d. Your score on the first statistics test as a measure of your knowledge of statistics.
e. Your score on an individual intelligence test as a measure of your intelligence.
f. The distance around your forehead measured with a tape measure as a measure of
your intelligence.
g. A response to the statement "Abortion is a woman's right" where "Strongly Disagree"
= 1, "Disagree" = 2, "No Opinion" = 3, "Agree" = 4, and "Strongly Agree" = 5, as a
measure of attitude toward abortion.
h. Times for swimmers to complete a 50-meter race
i. Months of the year September, October…
j. Socioeconomic status of a family when classified as low, middle and upper classes.
k. Blood type of individuals, A, B, AB and O.
l. Pollen counts provided as numbers between 1 and 10 where 1 implies there is almost
m. No pollen and 10 that it is rampant, but for which the values do not represent an
actual counts of grains of pollen.
n. Regions numbers of Ethiopia (1, 2, 3 etc.)
o. The number of students in a college;
p. The net wages of a group of workers;
q. The height of the men in the same town;
12
Basic Statistics, BU, CNSc, 2015
CHAPTER 2
METHODS OF DATA COLLECTION AND PRESNTATION
2.1 Introduction
The second Chapter of this module introduces the methods of data collection and
presentation. This unit will deal how to collect and present the data you have
collected so that they can be of use. Thus the collected data also known as raw data are
always in an unorganized form and need to be organized and presented in a
meaningful and readily comprehensible form in order to facilitate further statistical
analysis.
2.2 Objectives
At the end of this chapter students will be able to:
• Arrange raw data in an array and then classified data to construct a frequency table and
a cumulative frequency table.
13
Basic Statistics, BU, CNSc, 2015
Frequency distribution: is the organization of raw data in table form using classes and
frequencies. There are three basic types of frequency distributions. These are:
Categorical frequency distribution
Ungrouped frequency distribution
Grouped frequency distribution
There are specific procedures for constructing each type.
2.3.1Categorical frequency Distribution
Categorical frequency Distribution is used for data that can be place in specific categories
such as nominal, or ordinal. E.g. Marital status
Example: A social worker collected the following data on marital status for 25
persons. (M=married, S=single, W=widowed, D=divorced)
M S D W D
S S M M M
W D S M M
W D D S S
S W W D D
Solution:
Since the data are categorical, discrete classes can be used. There are four types of marital
status M, S, D, and W. These types will be used as class for the distribution. We follow
procedure to construct the frequency distribution.
Step 1: Make a table as shown.
Step 2: Tally the data and place the result in column (2).
Step 3: Count the tally and place the result in column (3).
Step 4: Find the percentages of values in each class by using;
14
Basic Statistics, BU, CNSc, 2015
f
% * 100 Where f= frequency of the class, n=total number of value.
n
Percentages are not normally a part of frequency distribution but they can be added since
they are used in certain types diagrammatic such as pie charts.
Step 5: Find the total for column (3) and (4). Combing the entire steps one can construct the
following frequency distribution.
70 60 62 70 85
65 60 63 74 75
76 70 70 80 85
15
Basic Statistics, BU, CNSc, 2015
Construct a frequency distribution, which is ungrouped.
Solution:
60 // 2
62 / 1
63 / 1
65 / 1
70 //// 4
74 / 1
75 // 2
76 / 1
80 /// 3
85 /// 3
90 / 1
16
Basic Statistics, BU, CNSc, 2015
2.2.3 Grouped frequency Distribution:
When the range of the data is large, the data must be grouped in to classes that are more than
one unit in width.
Definitions:
Grouped Frequency Distribution: a frequency distribution when several numbers
are grouped in one class.
Class limits: Separates one class in a grouped frequency distribution from another.
The limits could actually appear in the data and have gaps between the upper limits of
one class and lower limit of the next.
Units of measurement (U): the distance between two possible consecutive measures.
It is usually taken as 1, 0.1, 0.01, 0.001, -----.
Class boundaries: Separates one class in a grouped frequency distribution from
another. The boundaries have one more decimal places than the row data and
therefore do not appear in the data. There is no gap between the upper boundary of
one class and lower boundary of the next class.
The lower class boundary is found by subtracting U/2 from the corresponding lower class
limit and the upper class boundary is found by adding U/2 to the corresponding upper
class limit.
Class width: the difference between the upper and lower class boundaries of any
class. It is also the difference between the lower limits of any two consecutive classes
or the difference between any two consecutive class marks.
Class mark (Mid points): it is the average of the lower and upper class limits or the
average of upper and lower class boundary.
Cumulative frequency: is the number of observations less than/more than or equal to
a specific value.
Cumulative frequency above: it is the total frequency of all values greater than or
equal to the lower class boundary of a given class.
Cumulative frequency blow: it is the total frequency of all values less than or equal
to the upper class boundary of a given class.
Cumulative Frequency Distribution (CFD): it is the tabular arrangement of class
interval together with their corresponding cumulative frequencies. It can be more than
or less than type, depending on the type of cumulative frequency used.
17
Basic Statistics, BU, CNSc, 2015
Relative frequency (rf): it is the frequency divided by the total frequency.
Relative cumulative frequency (rcf): it is the cumulative frequency divided by the
total frequency.
Class limits
6 – 11
12 – 17
18 – 23
24 – 29
30 – 35
36 – 41
Class Class boundary Class Tally Freq. Cf (less Cf (more rf. rcf (less
limit Mark than than type) than type
type)
6 – 11 5.5 – 11.5 8.5 // 2 2 20 0.10 0.10
12 – 17 11.5 – 17.5 14.5 // 2 4 18 0.10 0.20
18 – 23 17.5 – 23.5 20.5 7 11 16 0.35 0.55
//////
24 – 29 23.5 – 29.5 26.5 //// 4 15 9 0.20 0.75
30 – 35 29.5 – 35.5 32.5 /// 3 18 5 0.15 0.90
36 – 41 35.5 – 41.5 38.5 // 2 20 2 0.10 1.00
20
Basic Statistics, BU, CNSc, 2015
2.4 Diagrammatic and Graphic presentation of data
These are techniques for presenting data in visual displays using geometric and pictures.
Importance:
They have greater attraction.
They facilitate comparison.
They are easily understandable.
21
Basic Statistics, BU, CNSc, 2015
Class Frequency Percent Degree
Men 2500 25 90
Women 2000 20 72
Girls 4000 40 144
Boys 1500 15 54
Pie-Chart
Class
C ategory
Boy s
Men
15.0%
Women
Girls
40.0%
20.0%
25.0%
[Link] Pictogram
In this diagram, we represent data by means of some picture symbols. We decide about a
suitable picture to represent a definite number of units in which the variable is measured.
Example 2.5
The following table shows the orange production in a plantation from production year
1990-1993. Represent the data by using a pictogram.
22
Basic Statistics, BU, CNSc, 2015
Solution
Activity
Draw a pictogram to represent the following population of a town.
Year 1989 1990 1991 1992
Population 2000 3000 5000 7000
2.3.1.3Bar Charts
A set of bars (thick lines or narrow rectangles) representing some magnitude over time space.
They are useful for comparing aggregate over time space. Bars can be drawn either vertically
or horizontally.
There are different types of bar charts. The most common being:
23
Basic Statistics, BU, CNSc, 2015
Simple Bar Charts are used to display data on one variable. They are thick lines (narrow
rectangles) having the same breadth. The magnitude of a quantity is represented by the
height /length of the bar.
Example: The following data represent sale by product, 1957- 1959 of a given company for
three products A, B, C.
A 12 14 18
B 24 21 18
C 24 35 54
Solutions:
30
25
Sales in $
20
15
10
5
0
A B C
product
24
Basic Statistics, BU, CNSc, 2015
Example:
Draw a component bar chart to represent the sales by product from 1957 to 1959.
Solution:
100
80
Sales in $
Product C
60
Product B
40
Product A
20
0
1957 1958 1959
Year of production
Example:
Draw a component bar chart to represent the sales by product from 1957 to 1959.
Solution:
25
Basic Statistics, BU, CNSc, 2015
Sales by product 1957-1959
60
50
Sales in $
40 Product A
30 Product B
20 Product C
10
0
1957 1958 1959
Year of production
26
Basic Statistics, BU, CNSc, 2015
[Link] Frequency Polygon:
It is a line graph. The frequency is placed along the vertical axis and classes mid points are
placed along the horizontal axis. It is customer to the next higher and lower class interval with
corresponding frequency of zero, this is to make it a complete polygon.
Example: Draw a frequency polygon for the above data (example *).
Solution:
4
Value Frequency
0
2. 5 8. 5 14.5 20.5 26.5 32.5 38.5 44.5
Activity
27
Basic Statistics, BU, CNSc, 2015
Summary
• When data are collected, the values are called raw data. Since very little knowledge can
be obtained from raw data, they must be organized in some meaningful way. A frequency
distribution using classes is the common method that is used.
• Finally, Other graphs such as the bar chart, pictogram and pie chart can also be used.
Some of these graphs are frequently seen in newspapers, magazines, and various
statistical reports.
28
Basic Statistics, BU, CNSc, 2015
Exercise 2
1. Define cumulative frequency distribution? Point out its special advantages and uses.
3. State the different types of bar diagrams. What are their merits and demerits?
4. A social worker collected the following data on marital status for 25 persons.
M S D W D
S S M M M
W D S M M
W D D S S
S W W D D
5. A demographer is interested in the number of children a family may have, took sample
of 30 families and obtained the following observations.
4 2 4 3 2 8
3 4 4 2 2 8
5 3 4 5 4 5
4 3 5 2 7 3
3 6 7 3 8 4
6. The following data are on age of 20 women who attended health education in a certain
hospital. Construct frequency distribution.
30, 25, 23, 41, 39, 27, 41, 24, 32, 29, 35, 31, 36, 33, 36, 42, 35, 37, 41, and 29
29
Basic Statistics, BU, CNSc, 2015
7. The following data represent the mark of 20 students.
80 76 90 85 80
70 60 62 70 85
65 60 63 74 75
76 70 70 80 85
8. Draw (a) histogram (b) frequency polygon (c) Ogive for the following
30
Basic Statistics, BU, CNSc, 2015
CHAPTER 3
3. MEASURES OF CENTERAL TENDENCY
3.1 Introduction
When we want to make comparison between groups of numbers it is good to have a single
value that is considered to be a good representative of each group. This single value is
called the average of the group. Averages are also called measures of central tendency. An
average which is representative is called typical average and an average which is not
representative and has only a theoretical value is called a descriptive average. A typical
average should possess the following:
Summarize data, using measures of central tendency, such as the mean, median,
mode, and midrange.
Describe data, using measures of variation, such as the range, variance, and standard
deviation.
Identify the position of a data value in a data set, using various measures of position,
such as percentiles, deciles, and quartiles.
Use the techniques of exploratory data analysis
3.3. The Summation Notation
Let X1, X2 ,X3 …XN be a number of measurements where N is the total number of
observation and Xi is ith observation.
Very often in statistics an algebraic expression of the form X 1+X2+X3+...+XN is used
in a formula to compute a statistic. It is tedious to write an expression like this very
31
Basic Statistics, BU, CNSc, 2015
often, so mathematicians have developed a shorthand notation to represent a sum of
scores, called the summation notation.
N
The symbol X
i 1
i is a mathematical shorthand for X1+X2+X3+...+XN
The expression is read, "the sum of X sub i from i equals 1 to N." It means "add up all the
numbers."
Example: Suppose the following were scores made on the first homework assignment for
five students in the class: 5, 7, 7, 6, and 8. In this example set of five numbers, where
N=5, the summation could be written:
The "i=1" in the bottom of the summation notation tells where to begin the sequence of
summation. If the expression were written with "i=3", the summation would start with the
third number in the set. For example:
In the example set of numbers, this would give the following result:
The "N" in the upper part of the summation notation tells where to end the sequence of
summation. If there were only three scores then the summation and example would be:
Sometimes if the summation notation is used in an expression and the expression must be
written a number of times, as in a proof, then a shorthand notation for the shorthand
notation is employed. When the summation sign " " is used without additional
32
Basic Statistics, BU, CNSc, 2015
For example:
n n
2. kX i k X i where k is any constant
i 1 i 1
n n
3. (a bX
i 1
i ) na b X i
i 1
where a and b are any constant
n n n
4. (X
i 1
i Yi ) X i Yi
i 1 i 1
X Y
5 6
7 7
7 8
6 7
8 8
5
a) X
i 1
i
5
b) Y
i 1
i
5
c) 10
i 1
5
d) (X
i 1
i Yi )
33
Basic Statistics, BU, CNSc, 2015
5
e) (X
i 1
i Yi )
5
f) X Y
i 1
i i
X
2
g) i
i 1
5 5
h) ( X i )( Yi )
i 1 i 1
Solutions:
5
a) X
i 1
i 5 7 7 6 8 33
5
b) Y
i 1
i 6 7 8 7 8 36
5
c) 10 5 *10 50
i 1
5
d) (X
i 1
i Yi ) (5 6) (7 7) (7 8) (6 7) (8 8) 69 33 36
5
e) (X
i 1
i Yi ) (5 6) (7 7) (7 8) (6 7) (8 8) 3 33 36
5
f) X Y
i 1
i i 5 * 6 7 * 7 7 * 8 6 * 7 8 * 8 241
X 5 2 7 2 7 2 6 2 8 2 223
2
g) i
i 1
5 5
h) ( X i )( Yi ) 33 * 36 1188
i 1 i 1
34
Basic Statistics, BU, CNSc, 2015
The Median
Quantiles (Quartiles, Deciles and Percentiles)
The choice of these averages depends up on which best fit the property under discussion.
3.4.1 Mean
i. Arithmetic Mean
It is defined as the sum of the magnitude of the items divided by the number of items.
X 1 X 2 ... X n
X
n
n
X i
X i 1
If X1 occurs f1 times
If X2occurs f2 times
If Xn occurs fn times
fX i i
Then the mean will be X i 1
k , where k is the number of classes and
f
i 1
i
f
i 1
i n
35
Basic Statistics, BU, CNSc, 2015
Example: Obtain the mean of the numbers 2, 7, 8, 2, 7, 3, 7
Solution:
Xi fi Xifi
2 2 4
3 1 3
7 3 21
8 1 8
Total 7 36
f i Xi
36
X i 1
4
5.15
f
7
i
i 1
f i Xi
X i 1
k
f i 1
i
Where Xi = the class mark of the ith class and fi = the frequency of the ith class
Class Frequency
6- 10 35
36
Basic Statistics, BU, CNSc, 2015
11- 15 23
16- 20 15
21- 25 12
26- 30 9
31- 35 6
Solutions:
Class fi Xi Xifi
6- 10 35 8 280
11- 15 23 13 299
16- 20 15 18 270
21- 25 12 23 276
26- 30 9 28 252
31- 35 6 33 198
f X i i
1575
X i 1
6
15.75
f
100
i
i 1
37
Basic Statistics, BU, CNSc, 2015
Activity
40-44 7
45-49 10
50-54 22
55-59 f4
60-64 f5
65-69 6
70-74 3
di X i A
X i di A
n n
Xi (d i A)
X i 1
i 1
n n
n
d i
X A i 1
n
X Ad
38
Basic Statistics, BU, CNSc, 2015
Where A is an assumed mean and d is the mean of the coded data. If the data are
expressed in terms of ungrouped frequency distribution,
di X i A
X i di A
k k
fi X i f (d i i A)
X i 1
i 1
n n
k
fd i i
X A i 1
n
X Ad
In both cases the true mean is the assumed mean plus the average of the deviations from
the assumed mean. Suppose the data is given in the shape of continuous frequency
distribution with a constant class size of “w” then the following coding is appropriate.
X A
d i
i w
X wd A
i i
k k
f X f ( wd A)
i i i i
X i 1 i 1
n n
k
f wd
i i
X A i 1
n
X A wd
A is an assumed mean usually the mean of the class marks (i =1, 2… k).
39
Basic Statistics, BU, CNSc, 2015
Example:
1. Suppose the deviations of the observations from an assumed mean of 7 are: 1, -1,
-2, -2, 0, -3, -2, 2, 0, -3.
a) Find the true mean
b) Find the original observation.
Solutions:
10
A 7, d i 10
i 1
10
a) d 1
10
X A d 7 1 6
The true mean is 6.
( X X ) 0.
i 1
i
2. The sum of the squared deviations of a set of items from their mean is the
n n
minimum. i.e. ( Xi X ) 2 ( X i A) 2 , A X
i 1 i 1
40
Basic Statistics, BU, CNSc, 2015
Then the mean of all the observation in all groups often called the combined mean is
given by:
X n X 2 n 2 .... X k n k X n i i
Xc 1 1 i 1
n1 n 2 ...n k
k
n
i 1
i
Solution:
Females Males
X 1 60 X 2 72
n1 30 n2 70
X n X 2 n2 X n i i
Xc 1 1 i 1
n1 n2 2
n
i 1
i
(CorrectValue WrongValue)
CorrectMean WrongMean
n
41
Basic Statistics, BU, CNSc, 2015
Solution:
(CorrectValue WrongValue)
CorrectMean WrongMean
n
(80 40)
CorrectMean 65 65 4 69k.g.
10
1. The mean of n Tetracycline Capsules X1, X2, …,Xn are known to be 12 gm. New set
of capsules of another drug are obtained by the linear transformation
Yi = 2Xi – 0.5 ( i = 1, 2, …, n ) then what will be the mean of the new set of capsules
Solution:
It is rigidly defined.
It is based on all observation.
42
Basic Statistics, BU, CNSc, 2015
It is suitable for further mathematical treatment.
It is stable average, i.e. it is not affected by fluctuations of sampling to some extent.
It is easy to calculate and simple to understand.
Demerits:
X W i i
Xw i 1
n
W
i 1
i
Example:
A student obtained the following percentage in an examination: English 60, Biology 75,
Mathematics 63, Physics 59, and chemistry [Link] the students weighted arithmetic
mean if weights 1, 2, 1, 3, 3 respectively are allotted to the subjects.
Solution:
43
Basic Statistics, BU, CNSc, 2015
5
X W i i
60 * 1 75 * 2 63 * 1 59 * 3 55 * 3 615
Xw i 1
61.5
1 2 1 3 3
5
10
W
i 1
i
The geometric mean of a set of n observation is the nth root of their product.
The geometric mean of X1, X2 ,X3 …Xn is denoted by G.M and given by:
G.M n X1 * X2 * ... * Xn
The logarithm of the G.M of a set of observation is the arithmetic mean of their
logarithm.
1 n
G.M Anti log( log X i )
n i1
Example:
Solution:
G.M n X1 * X2 * ... * Xn 3 2 * 4 * 8 3 64 4
Remark: The Geometric Mean is useful and appropriate for finding averages of
ratios.
44
Basic Statistics, BU, CNSc, 2015
iv The Harmonic Mean
The harmonic mean of X1, X2 , X3 …Xn is denoted by H.M and given by:
n
H.M n , This is called simple harmonic mean.
1
i 1 X i
k
n
H.M k , n fi
fi
i 1 X i
i 1
If observations X1, X2, …Xn have weights W1 , W2, …Wn respectively, then their
harmonic mean is given by
W i
H.M n
i 1
, This is called Weighted Harmonic Mean.
W i 1
i Xi
Remark: The Harmonic Mean is useful and appropriate in finding average speeds and
average rates.
Example: A cyclist pedals from his house to his college at speed of 10 km/hr and back
from the college to his house at 15 km/hr. Find the average speed.
3.4.2The Mode
Examples:
1
X̂ L mo w
1 2
Where:
46
Basic Statistics, BU, CNSc, 2015
Xˆ the mod e of the distribution
w the size of the mod al class
1 f mo f 1
2 f mo f 2
f mo frequencyof the mod al class
f 1 frequencyof the class preceedingthe mod al class
f 2 frequencyof the class following the mod al class
Example: The Following is the distribution of the size of certain farms selected at
random from a district. Calculate the mode of the distribution.
5-15 8
15-25 12
25-35 17
35-45 29
45-55 31
55-65 5
65-75 3
Solution:
47
Basic Statistics, BU, CNSc, 2015
45 55 is the mod al class,sin ce it is a class with thehighestfrequency.
L mo 45
w 10
1 f mo f 1 2
2 f mo f 2 26
f mo 31
f 1 29
f2 5
Xˆ 45 10
2
2 26
45.71
Demerits:
3.4.3The Median
In a distribution, median is the value of the variable which divides it in to two equal
halves. In an ordered series of data median is an observation lying exactly in the middle
48
Basic Statistics, BU, CNSc, 2015
of the series. It is the middle most value in the sense that the number of values less than
the median is equal to the number of values greater than it. If X1, X2, …Xn be the
observations, then the numbers arranged in ascending order will be X[1], X[2], …X[n],
where X[i] is ith smallest value.
a) 6, 5, 2, 8, 9, 4.
b) 2, 1, 8, 3, 5, 8.
Solutions:
~ 1
X X n X n
2 [2] [ 1]
2
X [ 3] X [ 4 ]
1
2
5 6 5.5
1
2
~ X
X n 1
[ ]
2
X [3]
3
49
Basic Statistics, BU, CNSc, 2015
[Link] Median for grouped data
If data are given in the shape of continuous frequency distribution, the median is defined
as:
~ w n
X L med ( c)
f med 2
Where :
L med lower class boundary of the median class.
w the size of the median class
n total number of observations.
c the cumulativefrequency(less than type) preceeding the median class.
f med thefrequency of the median class.
Remark:
The median class is the class with the smallest cumulative frequency (less than type) greater
n
than or equal to .
2
Example: Find the median of the following distribution.
Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
65-69 6
70-74 3
50
Basic Statistics, BU, CNSc, 2015
Solution:
40-44 7 7
45-49 10 17
50-54 22 39
55-59 15 54
60-64 12 66
65-69 6 72
70-74 3 75
n 75
37.5
2 2
39 is the first cumulative frequencyto be greater thanor equalto 37.5
50 54 is the median class.
51
Basic Statistics, BU, CNSc, 2015
L 49.5, w 5
med
n 75, c 17, f 22
med
~
X L w ( n c)
med f 2
med
49.5 5 (37.5 17)
22
54.16
[Link] Quartiles:
Quartiles are measures that divide the frequency distribution in to four equal parts. The value
of the variables corresponding to these divisions are denoted Q 1, Q2, and Q3 often called the
first, the second and the third quartile respectively. Q1 is a value which has 25% items which
are less than or equal to it. Similarly Q2 has 50%items with value less than or equal to it and Q 3
has 75% items whose values are less than or equal to it.
iN
To find Qi (i=1, 2, 3) we count of the classes beginning from the lowest class.
4
52
Basic Statistics, BU, CNSc, 2015
For grouped data we have the following formula
Q L Q w ( iN c) ,i 1,2,3
i i fQ 4
i
Where :
L Q lower class boundary of thequartile class.
i
w thesize of thequartile class
N total numberof observations.
c thecumulativefrequency(lessthantype) preceedingthequartile class.
f Q thefrequency of thequartile class.
i
Remark:
The quartile class (class containing Qi ) is the class with the smallest cumulative frequency
iN
(less than type) greater than or equal to .
4
[Link] Deciles:
- Deciles are measures that divide the frequency distribution in to ten equal parts.
- The values of the variables corresponding to these divisions are denoted D 1, D2,.. D9
often called the first, the second,…, the ninth decile respectively.
iN
- To find Di (i=1, 2,..9) we count of the classes beginning from the lowest class.
10
- For grouped data: we have the following formula
w iN
Di L Di ( c ) , i 1,2,...,9
f Di 10
Where :
L Di lower class boundaryof the decile class.
w the size of the decileclass
N total number of observations.
c the cumulativefrequency(less than type) preceedingthe decile class.
f Di thefrequency of the decile class.
53
Basic Statistics, BU, CNSc, 2015
Remark:
The decile class (class containing Di )is the class with the smallest cumulative frequency
iN
(less than type) greater than or equal to .
10
[Link] Percentiles:
Percentiles are measures that divide the frequency distribution in to hundred equal parts.
The values of the variables corresponding to these divisions are denoted P1, P2,.. P99 often
called the first, the second,…, the ninety-ninth percentile respectively. To find Pi (i=1,
iN
2,..99) we count of the classes beginning from the lowest class. For grouped data:
100
we have the following formula:
w iN
Pi L Pi ( c) , i 1,2,...,99
f Pi 100
Where :
L Pi lower class boundary of the percentile class.
w the size of the percentile class
N total number of observations.
c the cumulativefrequency( less than type) preceedingthe percentile class.
f Pi thefrequency of the percentileclass.
Remark:
The percentile class (class containing Pi )is the class with the smallest cumulative
iN
frequency (less than type) greater than or equal to .
100
54
Basic Statistics, BU, CNSc, 2015
Example: Considering the following distribution:
Values Frequency
140- 150 17
150- 160 29
160- 170 42
170- 180 72
180- 190 84
200- 210 49
210- 220 34
220- 230 31
230- 240 16
240- 250 12
Calculate:
a) All quartiles.
b) The 7th decile.
c) The 90th percentile.
55
Basic Statistics, BU, CNSc, 2015
Solutions:
140- 150 17 17
150- 160 29 46
160- 170 42 88
a) Quartiles:
i. Q1
56
Basic Statistics, BU, CNSc, 2015
N
123.25
4
170 180 is the class containingthe first quartile.
LQ 170 ,
1
w 10
N 493 , c 88 , f Q 72
1
w N
Q1 LQ1 ( c)
fQ 41
10
170 (123.25 88)
72
174.90
ii. Q2
Determine the class containing the second quartile.
2* N
246.5
4
190 200 is the class containingthe 2 nd quartile.
LQ 190 ,
2
w 10
N 493 , c 244 , f Q 107
2
w 2* N
Q2 LQ ( c)
2
fQ2
4
10
170 (246.5 244)
72
190.23
iii. Q3
Determine the class containing the third quartile.
57
Basic Statistics, BU, CNSc, 2015
3* N
369.75
4
200 210 is the class containingthe third quartile.
LQ 200 ,
3
w 10
N 493 , c 351 , f Q 49
3
w 3* N
Q3 LQ 3 ( c)
fQ 4
3
10
200 (369.75 351)
49
203.83
b) D7
Determine the class containing the 7th decile.
7* N
345.1
10
190 200 is the class containingthe seventh decile.
LD 190 ,
7
w 10
N 493 , c 244 , f D 107
7
w 7* N
D7 LD ( c)
7
f D 10
7
10
190 (345.1 244)
107
199.45
58
Basic Statistics, BU, CNSc, 2015
c) P90
90 * N
443.7
100
220 230 is the class containingthe90th percentile.
LP 220 ,
90
w 10
N 493 , c 434 , f P 3107
90
w 90 * N
P90 LP ( c)
90
f P 100
90
10
220 (443.7 434)
31
223.13
59
Basic Statistics, BU, CNSc, 2015
Summary
• This chapter explains the basic ways to summarize data. These include measures of
central tendency. They are the mean, median, mode, and midrange. The weighted mean
can also be used.
• There are several measures of the position of data values in the set. There are standard
scores, percentiles, quartiles, and deciles. Sometimes a data set contains an extremely
high or extremely low data value, called an outlier.
60
Basic Statistics, BU, CNSc, 2015
Exercise- 3
1. What is meant by central tendency? Briefly describe the methods of measurement
central tendency. Point out the merits and limitations of each method.
2. Why the arithmetic mean is most commonly used measure of a central value?
3. Compare the arithmetic mean, median, mode, geometric mean as to the manner in
which they are affected by extreme values.
5. Discuss briefly the use of weighted average in statistics, describing the cases in which
the weighted average is better than un weighted average.
Calculate the median, quartiles, 6th deciles, and 75th percentile from the following data.
Show that the value of 75thpercentile is the same as that of Q3.
Marks 80 70 60 50 40 30 20 10
No of Students 100 90 80 60 32 20 13 5
6. A quality control inspector at the lap top computer assembly plant found the
following number of defective computers on 20 consecutive working days.
10 14 10 12 18 19 25 22 17
28 13 17 18 14 18 21 15 16
7. Calculate Q1, Q3, D4, D6, P40 and P80 for the following tables
Frequency 4 6 10 15 12 7 6
8. The following are the scores for the midterm exam given to 13 students in
statistics.
42, 42, 68, 80, 75, 54, 62, 89, 72, 80, 80, 75, 65
61
Basic Statistics, BU, CNSc, 2015
Calculate the mean, median and mode.
Family A B C D E F G H I J
10. from the table given below find out the Median and Quartiles:
Frequency 7 10 13 26 35 22 11 5
11. If arithmetic mean of two items is 5 and G.M is 4, find their H.M. and also the
quantities.
12. Consider the following distribution, and then determine modal value of the
distribution.
X 1 2 3 4 5 6 7 8 9
F 3 1 18 25 40 30 22 10 6
13. The following frequency distribution is the distribution of profit earned by 15
companies during 2003 – 2004.
62
Basic Statistics, BU, CNSc, 2015
CHAPTER 4
4. Measures of Dispersion (Variation)
4.1 Introduction
The scatter or spread of items of a distribution is known as dispersion or variation. In
other words the degree to which numerical data tend to spread about an average value is
called dispersion or variation of the data. Measures of dispersions are statistical measures
which provide ways of measuring the extent in which data are dispersed or spread out.
4.2 Objectives
At the end of this Chapter student should be able to:
63
Basic Statistics, BU, CNSc, 2015
Mean deviation and coefficient of Mean deviation
Standard deviation and coefficient of variation.
For this reason, among others, the range is not the most important measure of variability.
R LS , L l arg est observation
S smallestobservation
It is rigidly defined.
It is easy to calculate and simple to understand.
Demerits:
64
Basic Statistics, BU, CNSc, 2015
It is not based on all observation.
It is highly affected by extreme observations.
It is affected by fluctuation in sampling.
It is not liable to further algebraic treatment.
It cannot be computed in the case of open end distribution.
LS R
RR
LS LS
Example:
1. If the range and relative range of a series are 4 and 0.25 respectively. Then what is the
value of:
a) Smallest observation
b) Largest observation
Solutions :( 2)
R 4 L S 4 __________ _______(1)
RR 0.25 L S 16 __________ ___( 2)
Solving (1) and ( 2) at the same time , one can obtain the following value
L 10 and S 6
Activity
Q3 Q1
Q.D
2
65
Basic Statistics, BU, CNSc, 2015
4.4.4 Coefficient of Quartile Deviation (C.Q.D)
(Q3 Q1 2 2 * Q.D Q3 Q1
C. Q.D
(Q3 Q1 ) 2 Q3 Q1 Q3 Q1
It gives the average amount by which the two quartiles differ from the median.
Example: Compute Q.D and its coefficient for the following distribution.
Values Frequency
140- 149 17
150- 159 29
160- 169 42
170- 179 72
180- 189 84
200- 209 49
210- 219 34
220- 229 31
230- 239 16
240- 249 12
Solution:
In the previous chapter we have obtained the values of all quartiles as:
Q3 Q1 203.83 174.90
Q.D 14.47
2 2
2 * Q.D 2 *14.47
C.Q.D 0.076
Q3 Q1 203.83 174.90
66
Basic Statistics, BU, CNSc, 2015
Remark: Q.D or C.Q.D includes only the middle 50% of the observation.
n
Xi X
M .D ( X ) i 1
n
k
fi X i X
M .D ( X ) i 1
n
n ~
~
Xi X
M .D( X ) i 1
n
k ~
~
fi X i X
M .D ( X ) i 1
n
67
Basic Statistics, BU, CNSc, 2015
~
Steps to calculate M.D ( X ):
~
Find the median, X
~
Find the deviations of each reading from X .
Find the arithmetic mean of the deviations, ignoring sign
c) Mean Deviation about the mode.
X i
ˆ
X
ˆ)
M.D( X i 1
n
k
f i X i Xˆ
M .D ( Xˆ ) i 1
n
1. The following are the number of visit made by ten mothers to the local doctor‟s surgery.
8, 6, 5, 5, 7, 4, 5, 9, 7, 4
Find mean deviation about mean, median and mode.
Solutions:
~
X 6, X 5.5, Xˆ 5
68
Basic Statistics, BU, CNSc, 2015
Then take the deviations of each observation from these averages.
Xi 4 4 5 5 5 6 7 7 8 9 total
Xi 6 2 2 1 1 1 0 1 1 2 3 14
X i 5.5 1.5 1.5 0.5 0.5 0.5 0.5 1.5 1.5 2.5 3.5 14
Xi 5 1 1 0 0 0 1 2 2 3 4 14
10
X i 6) 14
M .D( X ) i 1
1.4
10 10
10
~
X i 5.5 14
M .D ( X ) i 1
1.4
10 10
10
X i 5) 14
M .D( Xˆ ) i 1
1.4
10 10
2. Find mean deviation about mean, median and mode for the following
distributions.(exercise)
Class Frequency
40-44 7
45-49 10
50-54 22
55-59 15
60-64 12
69
Basic Statistics, BU, CNSc, 2015
65-69 6
70-74 3
Remark: Mean deviation about the mean is always minimum than mean deviation about
the median.
M .D( X )
C.M .D( X )
X
~
~ M .D( X )
C.M .D( X ) ~
X
M .D( Xˆ )
C.M .D( Xˆ )
Xˆ
Example: calculate the C.M.D about the mean, median and mode for the data in example
1 above.
Solutions:
M .D
C.M .D
Average about which deviationsare taken
M .D( X ) 1.4
C.M .D( X ) 0.233
X 6
~
~ M .D( X ) 1.4
C.M .D( X ) ~ 0.255
X 5.5
ˆ ) 1.4
M .D( X
C.M .D( Xˆ ) 0.28
ˆ
X 5
70
Basic Statistics, BU, CNSc, 2015
Activity
1
Population Varince 2 ( X i ) 2 , i 1,2,.....N
N
1
Population Varince 2 f i ( X i ) 2 , i 1,2,.....k
N
[Link] Sample Variance
One would expect the sample variance to simply be the population variance with the
population mean replaced by the sample mean. However, one of the major uses of
statistics is to estimate the corresponding parameter. This formula has the problem that
the estimated value isn't the same as the parameter. To counteract this, the sum of the
squares of the deviations is divided by one less than the sample size.
1
Sample Varince S 2 ( X i X ) 2 , i 1,2,....., n
n 1
For the case of frequency distribution it is expressed as:
1
Sample Varince S 2 f i ( X i X ) 2 , i 1,2,.....k
n 1
We usually use the following short cut formula.
n
X nX 2
2
i
S2 i 1 , for raw data.
n 1
k
f X i nX
2 2
i
S 2
i 1
, frequency distribution.
n 1
71
Basic Statistics, BU, CNSc, 2015
[Link] Standard Deviation
There is a problem with variances. Recall that the deviations were squared. That means
that the units were also squared. To get the units back the same as the original data
values, the square root must be taken.
Step 1: Find the difference between each observation and the mean.
Step 1: Since the data is a sample, divide the number (from step 4 above) by the number
of observations minus one, i.e., n-1 (where n is equal to the number of observations in the
data set).
Example: Find the variance and standard deviation of the following sample data.
1. X 11
Xi 5 10 12 17 Total
(Xi- X ) 2 36 1 1 36 74
n
( X i X )2 74
S2 i 1
24.67.
n 1 3
S S2 24.67 4.97.
72
Basic Statistics, BU, CNSc, 2015
2. X 55
Xi(C.M) 42 47 52 57 62 67 72 Total
n
fi ( X i X )2 4400
S2 i 1
59.46.
n 1 74
S S2 59.46 7.71.
1. ( X i X )2 ( X i A) 2 ,A X
n 1 n 1
2. For normal (symmetric distribution) the following holds:
Approximately 68.27% of the data values fall within one standard deviation of the
mean. i.e. with in ( X S , X S )
Approximately 95.45% of the data values fall within two standard deviations of the
mean. i.e. with in ( X 2 S , X 2 S )
Approximately 99.73% of the data values fall within three standard deviations of the
mean. i.e. with in ( X 3S , X 3S )
Chebyshev's Theorem
For any data set ,no matter what the pattern of variation, the proportion of the values that
fall within k standard deviations of the mean or ( X kS , X kS ) will be at least
1 , where k is an number greater than 1. i.e. the proportion of items falling beyond
1
k2
k standard deviations of the mean is at most 1
k2
Example 1: Suppose a distribution has mean 50 and standard deviation 6. What percent
of the numbers are?
a) Between 38 and 62
73
Basic Statistics, BU, CNSc, 2015
b) Between 32 and 68
c) Less than 38 or more than 62.
d) Less than 32 or more than 68.
Solutions:
a) 38 and 62 are at equal distance from the mean,50 and this distance is 12
ks 12
12 12
k 2
S 6
Applying the above theorem at least (1 1 ) *100% 75% of the numbers lie
k2
between 38 and 62.
c) It is just the complement of a) i.e. at most 1 *100% 25% of the numbers lie
k2
less than 32 or more than 62.
d) Done in similarly way as „c‟.
Activity
The average score of a special test of knowledge of wood refinishing has a mean of 53
and standard deviation of 6. Find the range of values in which at least 75% the scores will
lie.
74
Basic Statistics, BU, CNSc, 2015
Examples:
b. If each of the numbers in the set are multiplied by -5, then what will be the variance
and standard deviation of the new set?
Solutions:
The distribution having less C.V is said to be less variable or more consistent.
Examples:
1. An analysis of the monthly wages paid (in Birr) to workers in two firms A and B belonging to
the same industry gives the following results.
Value Firm A Firm B
75
Basic Statistics, BU, CNSc, 2015
In which firm A or B is there greater variability in individual wages?
Solutions:
SA 10
[Link] *100 *100 19.05%
XA 52.5
SB 11
[Link] *100 *100 23.16%
XB 47.5
Since [Link] < [Link], in firm B there is greater variability in individual wages.
Activity
City 1 25 24 23 26 17
City2 22 21 24 22 20
City3 32 27 35 24 28
Which city have the most consistent temperature, based on these data? (Exercise)
76
Basic Statistics, BU, CNSc, 2015
X X
Z , for sample
S
Mean 78 90
[Link] 6 5
Student A from section 1 scored 90 and student B from section 2 scored 95. Relatively
speaking who performed better?
Solutions:
X A X 1 90 78
ZA 2
S1 6
X B X 2 95 90
ZB 1
S2 5
Therefore, Student A performed better relative to his section because the score of student
A is a two standard deviation above the mean score of his section, while the score of
student B is only one standard deviation above the mean score of his section.
2. Two groups of people were trained to perform a certain task and tested to find out
which group is faster to learn the task. For the two groups the following information
was given:
77
Basic Statistics, BU, CNSc, 2015
Value Group one Group two
Relatively speaking:
S2 1.3
C.V2 *100 *100 10.92%
X2 11.9
Hence, Child B is faster because the time taken by child B is two standard deviation
shorter than the average time taken by group 2 while, the time taken by child A is only
one standard deviation shorter than the average time taken by group 1.
[Link] Moments
If X is a variable that assume the values X1, X2,…..,Xn then
78
Basic Statistics, BU, CNSc, 2015
X X 2 ... X n
r r r
X 1r
n
n
Xi
r
i 1
n
k
fi X i
r
Xr i 1
n
2. The rth moment about the mean ( the rth central moment)
It is denoted by Mr and defined as:
n n
( X i X )r (n 1) i
( X i X )r
Mr i 1
1
n n n 1
k
fi ( X i X )r
Mr i 1
n
n n
( X i A) r
(n 1) i 1
( X i A) r
Mr
' i 1
n n n 1
79
Basic Statistics, BU, CNSc, 2015
k
f i ( X i A) r
Mr i 1
'
Example:
1. Find the first two moments for the following set of numbers 2, 3, 7
2. Find the first three central moments of the numbers in problem 1
3. Find the third moment about the number 3 of the numbers in problem 1.
Solutions:
i 1
Xr
n
23 7
X1 4 X
3
2 2 32 7 2
X2 20.67
3
80
Basic Statistics, BU, CNSc, 2015
[Link] Skewness
Skewness is the degree of asymmetry or departure from symmetry of a distribution. A
skewed frequency distribution is one that is not symmetrical. Skewness is concerned with
the shape of the curve not size. If the frequency curve (smoothed frequency polygon) of a
distribution has a longer tail to the right of the central maximum than to the left, the
distribution is said to be skewed to the right or said to have positive skewness. If it has a
longer tail to the left of the central maximum than to the right, it is said to be skewed to
the left or said to have negative skewness.
For moderately skewed distribution, the following relation holds among the three
commonly used measures of central tendency.
Measures of Skewness
Measures of Skewness are denoted by 3 and there are various measures of skewness.
1. Suppose the mean, the mode, and the standard deviation of a certain distribution are
32, 30.5 and 10 respectively. What is the shape of the curve representing the
distribution?
Solution:
Given: ~ Required: Q1 , Q3
3 0.5, X Q2 11
Q1 Q3 28...........................(*)
82
Basic Statistics, BU, CNSc, 2015
Solving (*) and (**) at the sametime we obtain the following values
Q1 8 and Q3 20
Some characteristics of annually family income distribution (in Birr) in two regions is as
follows:
Activity
83
Basic Statistics, BU, CNSc, 2015
representing a distribution is flat topped, it is called platykurtic. The normal
distribution which is not very high peaked or flat topped is called mesokurtic.
Measures of kurtosis
M4 M
4 44
M2
2
Where : M 4 is the fourth moment aboutthe mean.
M 2 is the sec ond moment aboutthe mean.
is the populations tan dard deviation.
Examples:
M3 60
a) 3 32
0.94 0
M2 16 3 2
The distribution is negatively skewed .
b) 4
M4
162
0.6 3
2
M2 162
84
Basic Statistics, BU, CNSc, 2015
Activity
1. The median and the mode of a mesokurtic distribution are 32 and 34 respectively.
The 4th moment about the mean is 243. Compute the Pearsonian coefficient of
skewness and identify the type of skewness. Assume (n-1 = n).
2. If the standard deviation of a symmetric distribution is 10, what should be the value
of the fourth moment so that the distribution is mesokurtic?
Summary
• To summarize the variation of data, statisticians use measures of variation or dispersion.
The three most common measures of variation are the range, variance, and standard
deviation. The coefficient of variation can be used to compare the variation of two data
sets. The data values are distributed according to Chebyshev‟s theorem on the empirical
rule.
Exercise 4
1. The following data represents the price-earning (P/E) ratio of 50 stocks selected at
random from the stocks listed with New York stock exchange (NYSE), during a
given period of time. (The ratio has been rounded to the whole number.)
10 11 13 13 10 11 26 16 12 11
10 12 11 9 15 9 12 18 12 19
11 11 13 21 11 13 14 10 13 12
9 19 8 13 15 10 13 18 10 13
11 8 17 11 10 9 13 11 18 10
Compute: The mean, median, mode, range, inter-quartile range, variance, standard
deviation, coefficient of variation.
85
Basic Statistics, BU, CNSc, 2015
2. The sum of fifteen observations, whose mode is 8, was found to be 150 with
coefficient of variation of 20%
(a) Calculate the pearsonian coefficient of skewness and give appropriate conclusion.
(b) Are smaller values more or less frequent than bigger values for this
distribution?
(c) If a constant k was added on each observation, what will be the new pearsonian
coefficient of skewness? Show your steps. What do you conclude from this?
3. Some characteristics of annually family income distribution (in Birr) in two regions are
as follows:
86
Basic Statistics, BU, CNSc, 2015
CHAPTER 5
5. ELEMENTARY PROBABILITY
5.1 Introduction
Probability theory is the foundation upon which the logic of inference is built. It helps us
to cope up with uncertainty. In general, probability is the chance of an outcome of an
experiment. It is the measure of how likely an outcome is to occur.
5.2 Objectives
After completing this chapter, the student should be able to:
Determine sample spaces and find the probability of an event, using classical
probability or empirical probability.
Find the probability of compound events, using the addition rules.
Find the probability of compound events, using the multiplication rules.
Find the conditional probability of an event.
Find the total number of outcomes in a sequence of events, using the fundamental
counting rule.
Find the number of ways that r objects can be selected from n objects, using the
permutation rule.
Find the number of ways that r objects can be selected from n objects without
regard to order, using the combination rule.
Find the probability of an event, using the counting rules.
5.3 Definitions of some probability terms
1. Experiment: Any process of observation or measurement or any process which
generates well defined outcome.
2. Probability Experiment: It is an experiment that can be repeated any number of times under
similar conditions and it is possible to enumerate the total number of outcomes without predicting
an individual out come. It is also called random experiment.
Example: If a fair die is rolled once it is possible to list all the possible outcomes
i.e.1, 2, 3, 4, 5, 6 but it is not possible to predict which outcome will occur.
87
Basic Statistics, BU, CNSc, 2015
5. Event: It is a subset of sample space. It is a statement about one or more outcomes of a
random experiment .They are denoted by capital letters.
Example: Considering the above experiment let A be the event of odd numbers, B be the event
of even numbers, and C be the event of number 8.
A 1,3,5
B 2,4,6
C or empty spaceor impossibleevent
Remark:
If S (sample space) has n members then there are exactly 2n subsets or events.
6. Equally Likely Events: Events which have the same chance of occurring.
7. Complement of an Event: the complement of an event A is non-occurrence of A and is
'
denoted by A , or Ac , or A contains those points of the sample space which do not belong to an
event A.
8. Elementary Event: an event having only a single element or sample point.
9. Mutually Exclusive Events: Two events which cannot happen at the same time.
10. Independent Events: Two events are independent if the occurrence of one does not affect
the probability of the other occurring.
11. Dependent Events: Two events are dependent if the first event affects the outcome or
occurrence of the second event in a way the probability is changed.
Example: .What is the sample space for the following experiment
a) S={1,2,3,4,5,6}
88
Basic Statistics, BU, CNSc, 2015
b) S={(HH),(HT),(TH),(TT)}
c) S={t /t≥0}
Sample space can be
Countable ( finite or infinite)
Uncountable.
5.4 Counting Rules
In order to calculate probabilities, we have to know
Example: A student goes to the nearest snack to have a breakfast. He can take tea, coffee, or
milk with bread, cake and sandwich. How many possibilities does he have?
Solution:
Tea
Bread
Cake
Sandwich
Coffee
Bread
Cake
89
Basic Statistics, BU, CNSc, 2015
Sandwitch
Milk
Bread
Cake
Sandwich
There are nine possibilities.
Example: The digits 0, 1, 2, 3, and 4 are to be used in 4 digit identification card. How many
different cards are possible if
a)
1st digit 2nd digit 3rd digit 4th digit
5 5 5 5
90
Basic Statistics, BU, CNSc, 2015
5 * 5 * 5 * 5 625 differentcards are possible.
b)
1st digit 2nd digit 3rd digit 4th digit
5 4 3 2
5.4.2 Permutation
An arrangement of n objects in a specified order is called permutation of the objects.
the formula is
n!
n Pr
(n r )!
3. The number of permutations of n objects in which k1 are alike k2 are alike ---- etc
is:
91
Basic Statistics, BU, CNSc, 2015
n!
n Pr
k1!*k2 * ...* kn
Example:
1.
a)
Here n 4, there are four disnict object
There are 4! 24 permutations.
Here n 4, r 2
b) 4! 24
There are 4 P2 12 permutations.
(4 2)! 2
Here n 10
Of which 2 are C , 2 are O, 2 are R ,1E ,1T ,1I ,1N
2. K1 2, k 2 2, k3 2, k 4 k5 k6 k7 1
U sin g the 3rd rule of permutatio n , there are
10!
453600 permutatio ns.
2!*2!*2!*1!*1!*1!*1!
Activity
1. Six different statistics books, seven different physics books, and 3 different
Economics books are arranged on a shelf. How many different arrangements are
possible if;
i. The books in each particular subject must all stand together
92
Basic Statistics, BU, CNSc, 2015
ii. Only the statistics books must stand together
2. If the permutation of the word WHITE is selected at random, how many of the
permutations
i. Begins with a consonant?
ii. Ends with a vowel?
iii. Has a consonant and vowels alternating?
5.4.3 Combination
A selection of objects without regard to order is called combination.
Example: Given the letters A, B, C, and D list the permutation and combination for
selecting two letters.
Solution:
Permutation Combination
AB BA CA DA AB BC
AC BC CB DB AC BD
AD BD CD DC AD DC
Note that in permutation AB is different from BA. But in combination AB is the same as BA.
n
C
n r or and is given by the formula:
r
n n!
r (n r )!*r!
Examples:
93
Basic Statistics, BU, CNSc, 2015
Solution:
n9 , r 5
n n! 9!
126 ways
r ( n r )!* r! 4!* 5!
2. Among 15 clocks there are two defectives .In how many ways can an inspector chose
three of the clocks for inspection so that:
a) There is no restriction.
b) None of the defective clock is included.
c) Only one of the defective clocks is included.
d) Two of the defective clock is included.
Solution:
a) If there is no restriction select three clocks from 15 clocks and this can be done in :
n 15 , r 3
n n! 15!
455 ways
r ( n r )!* r! 12!* 3!
2 13
* 286 ways.
0 3
c) Only one of the defective clocks is included.
This is equivalent to one defective and two non defective, which can be done in:
94
Basic Statistics, BU, CNSc, 2015
2 13
* 156 ways.
1 2
d) Two of the defective clock is included.
This is equivalent to two defective and one non defective, which can be done in:
2 13
* 13 ways.
2 1
Activity
2. If 3 books are picked at random from a shelf containing 5 novels, 3 books of poems,
and a dictionary, in how many ways this can be done if
a) There is no restriction.
b) The dictionary is selected?
c) 2 novels and 1 book of poems are selected?
5.5 Approaches to measuring Probability
There are four different conceptual approaches to the study of probability theory. These
are:
95
Basic Statistics, BU, CNSc, 2015
- All outcomes are equally likely.
- Total number of outcome is finite, say N.
Definition: If a random experiment with N equally likely outcomes is conducted and out
of these NA outcomes are favourable to the event A, then the probability that event A
Example:
S 1, 2, 3, 4, 5, 6
N n( S ) 6
A 4
N A n( A) 1
n( A)
P ( A) 1 6
n( S )
96
Basic Statistics, BU, CNSc, 2015
A 1,3,5
N A n( A) 3
n( A)
P( A) 3 6 0.5
n( S )
A 2,4,6
N A n( A) 3
n( A)
P( A) 3 6 0.5
n( S )
N A n( A) 0
n( A)
P ( A) 0 60
n( S )
80
Total selection N n( S )
10
a) Let A be the event that all will be defective.
97
Basic Statistics, BU, CNSc, 2015
30 50
Total way in which A occur * N A n( A)
10 0
30 50
*
n( A) 10 0
P( A) 0.00001825
n( S ) 80
10
b) Let A be the event that 6 will be non defective.
30 50
Total way in which A occur * N A n( A)
4 6
30 50
*
n( A) 4 6
P( A) 0.265
n( S ) 80
10
c) Let A be the event that all will be non-defective.
30 50
Total way in which A occur * N A n( A)
0 10
30 50
*
n( A) 0 10
P( A) 0.00624
n( S ) 80
10
Activity
1. What is the probability that a waitress will refuse to serve alcoholic beverages to
only three minors if she randomly checks the I.D‟s of five students from among
ten students of which four are not of legal age?
2. If 3 books are picked at random from a shelf containing 5 novels, 3 books of
poems, and a dictionary, what is the probability that
a) The dictionary is selected?
98
Basic Statistics, BU, CNSc, 2015
b) 2 novels and 1 book of poems are selected?
Short coming of the classical approach:
NA
P( A) lim
N N
Example: If records show that 60 out of 100,000 bulbs produced are defective. What is the
probability of a newly produced bulb to be defective?
Solution:
NA 60
P( A) lim 0.0006
N N 100,000
1. P( A) 0
2. P( S ) 1, S is the sure event.
99
Basic Statistics, BU, CNSc, 2015
3. If A and B are mutually exclusive events, the probability that one or the other occur
equals the sum of the two probabilities. i. e.
P( A B) P( A) P( B)
4. P( A' ) 1 P( A)
5. 0 P( A) 1
6. P(ø) =0, ø is the impossible event.
5.6 Conditional probability and Independency
Conditional Events: If the occurrence of one event has an effect on the next occurrence of
the other event then the two events are conditional or dependent events.
Example: Suppose we have two red and three white balls in a bag
2
Let A= the event that the first draw is red p ( A)
5
2
B= the event that the second draw is red p( B)
5
A and B are independent.
2
Let A= the event that the first draw is red p ( A)
5
This is conditional.
Let B= the event that the second draw is red given that the first draw is red
p( B) 1 4
100
Basic Statistics, BU, CNSc, 2015
5.6.1 Conditional probability of an event
The conditional probability of an event A given that B has already occurred, denoted
p( A B) is:
p( A B)
p( A B) = , p( B) 0
p( B)
(2) p( B ' A) 1 p( B A)
Examples
1. For a student enrolling at freshman at certain university the probability is 0.25 that
he/she will get scholarship and 0.75 that he/she will graduate. If the probability is 0.2
that he/she will get scholarship and will also graduate. What is the probability that a
student who get a scholarship graduate?
Solution: Let A= the event that a student will get a scholarship
2. If the probability that a research project will be well planned is 0.60 and the
probability that it will be well planned and well executed is 0.54, what is the
probability that it will be well executed given that it is well planned?
Solution; Let A= the event that a research project will be well Planned
101
Basic Statistics, BU, CNSc, 2015
given p( A) 0.60, p A B 0.54
Re quired pB A
p A B 0.54
p B A 0.90
p A 0.60
Activity
A lot consists of 20 defective and 80 non-defective items from which two items are
chosen without replacement. Events A & B are defined as A = the first item chosen is
defective, B = the second item chosen is defective
pB pB A. p A p B A' . p A'
5.6.2 Probability of Independent Events
Here p A B p A, P B A p B
Example; A box contains four black and six white balls. What is the probability of
getting two black balls in drawing one after the other under the following conditions?
Required p A B
• There are four basic types of probability. They are classical, frequentist (Empirical),
axiomatic and subjective probability. Classical probability uses samples spaces. The
probability of any event is a number from 0 to 1. If an event cannot occur, the probability
is 0. If an event is certain, the probability is 1. The sum of the probability of all the events
in the sample space is 1. To find the probability of the complement of an event, subtract
the probability of the event from 1.
• Two events are mutually exclusive if they cannot occur at the same time; otherwise, the
events are not mutually exclusive. To find the probability of two mutually exclusive
events occurring, add the probability of each event. To find the probability of two events
when they are not mutually exclusive, add the possibilities of the individual events and
then subtract the probability that both events occur at the same time. These types of
probability problems can be solved by using the addition rules.
• Two events are independent if the occurrence of the first event does not change the
probability of the second event occurring. Otherwise, the events are dependent. To find
the probability of two independent events occurring, multiply the probabilities of each
event. To find the probability that two dependent events occur, multiply the probability
that the first event occurs by the probability that the second event occurs given that the
first event has already occurred. The complement of an event is found by selecting the
outcomes in the sample space that are not involved in the outcomes of the event.
These types of problems can be solved by using the multiplication rules and the
complementary event rules.
• Finally, when a large number of events can occur, the fundamental counting rule, the
permutation rule, and the combination rule can be used to determine the number of ways
that these events can occur.
• The counting rules and the probability rules can be used to solve more-complex
probability problems.
103
Basic Statistics, BU, CNSc, 2015
Exercise 5
1. Two urns contain 3 white 7 black balls and 10 white, 7 black balls respectively. A
ball is transferred from the first urn to the second and then a ball is drawn from the
second urn. Find the probability that it is black
2. Three persons write their names on individual slip of papers and deposit the slip in a
box. Each of three persons draws at random a slip from the box. Find the
probability that each person has drawn the slip bearing his own name.
3. A box contains 6 red and 4 black balls. Two draws of two balls each are made
without replacement. Find the probability that the first two balls are both red and the
next two are one red and one black.
4. Two urns contain: 3 white, 2 black and 4 white in the first urn and 3 black balls in
second urn. One ball is transferred from the first urn to the second. The ball is drawn
from the second urn. Find the probability that it is white.
5. If an event A3 is independent of two mutually exclusive events A1and A2 Show that
A3 is independent of A1U A2.
6. The probabilities that A and B solve a given problem independently are 2/3 and 3/5
respectively. If both of them attempt the problem, find the probability that the
problem will be solved.
7. An urn contains 8 red and 12 black balls. 3 balls are drawn at random and returned to
the urn. Again 3 balls are drawn. Find the probability that the 6 balls in the two draws
include 2 black balls.
8. A company has two machines M1and M2. M1 produces 60% of its product and M2
produces 40% of its product. M1 produces 5% defective units and M2 produces 4%
defective units. A unit is selected at random from the whole product. Find the
probability that it is defective.
104
Basic Statistics, BU, CNSc, 2015
CHAPTER 6
6. RANDOM VARIABLES AND PROBABILITY DISTRIBUTIONS
6.1 Introduction
What is probability?
Probability theory is a part of our everyday life. We may hear a doctor say that a patient
has a 50-50 chance of survival or a meteorologist predict heavy rain with 80%
chance. Probability theory is concerned with the study of random (or chance)
phenomena. Such phenomena are characterized by the fact that their future
behavior is not predictable in a deterministic fashion.
measure of the degree of uncertainty associated with random events. The study of
probability dates back to the 17th century and the work of two mathematicians Blaise
Pascal (1623-1662) and pierre de Fermat (1601-1665). Today the mathematical
theory of probability is the basis for statistical applications in social, economic and
decision making areas.
The element of chance plays a vital role in our life. Many important business
decisions are made on this basis. The entire business of insurance and share market is
based on probability theory. Quality control, reliability theory, queuing theory, system
failure, radar detection, noise, birth and death rates and games of chance are some other
fields where probability finds application.
6.2 Objectives
After studying this chapter, you should be able to:
105
Basic Statistics, BU, CNSc, 2015
Definition: A random variable is a numerical description of the outcomes of the experiment or
a numerical valued function defined on sample space, usually denoted by capital letters.
Example: If X is a random variable, then it is a function from the elements of the sample space
to the set of real numbers. i.e. R
Example: Flip a coin three times, let X be the number of heads in three tosses.
106
Basic Statistics, BU, CNSc, 2015
Examples:
Height of students at certain college.
Mark of a student.
Life time of light bulbs.
Length of time required to complete a given training.
Definition: a probability distribution consists of a value a random variable can assume and
the corresponding probabilities of the values.
Example: Consider the experiment of tossing a coin three times. Let X be the number of
heads. Construct the probability distribution of X.
Solution:
First identify the possible value that X can assume.
Calculate the probability of each possible distinct value of X and express X in the
form of frequency distribution.
X x 0 1 2 3
P X x 1 8 38 38 18
Probability distribution is denoted by P for discrete and by f for continuous random
variable.
6.5 Properties of Probability Distribution:
1.
P( x) 0, if X is discrete.
f ( x) 0, if X is continuous.
2.
P X x 1 , if X is discrete.
x
f ( x)dx 1 , if is continuous.
x
Note:
1. If X is a continuous random variable then
107
Basic Statistics, BU, CNSc, 2015
b
P(a X b) f ( x)dx
a
E ( X ) X 1P( X 1 ) X 2 P( X 2 ) .... X n P( X n )
n
X i P( X i )
i 1
Again let X be a continuous random variable assuming the values in the interval (a, b)
b
such that f ( x)dx 1,then
a
b
E ( X ) x f ( x)dx
a
108
Basic Statistics, BU, CNSc, 2015
Examples:
1. What is the expected value of a random variable X obtained by tossing a coin three
times where is the number of heads
Solution:
First construct the probability distribution of X
X x 0 1 2 3
P X x 1 8 38 38 18
E ( X ) X 1 P( X 1 ) X 2 P( X 2 ) .... X n P( X n )
0 *1 8 1* 3 8 ..... 2 *1 8
1.5
2. Suppose a charity organization is mailing printed return-address stickers to over one million
homes in the Ethiopia. Each recipient is asked to donate either $1, $2, $5, $10, $15, or $20.
Based on past experience, the amount a person donates is believed to follow the following
probability distribution:
6
E ( X ) xi P( X xi ) $7.25
i 1
109
Basic Statistics, BU, CNSc, 2015
6.7 Mean and Variance of a random variable
Let X be a given random variable, then
1. The expected value of X which also known as mean is given as:
Mean of X E (X )
2. The variance of X is also given by:
n
E ( X ) xi P ( X xi ) , if X is discrete
2 2
i 1
Where:
x 2 f ( x)dx , if X is continuous .
x
Example:
1. Find the mean and the variance of a random variable X in example 2 above.
Solutions:
X x $1 $2 $5 $10 $15 $20 Total
E ( X ) 7.25
Var( X ) E ( X 2 ) [ E ( X )]2 82.15 7.252 29.59
Activity
Two dice are rolled. Let X be a random variable denoting the sum of the numbers on the
two dice.
110
Basic Statistics, BU, CNSc, 2015
Let X and Y are random variables and k be a constant.
RULE 1
E (k ) k
RULE 2
Var ( k ) 0
RULE 3
E (kX ) kE ( X )
RULE 4
111
Basic Statistics, BU, CNSc, 2015
Let P the probability of success
q 1 p the probability of failureon any given trial
n
P( X x) p x q n x , x 0,1,2,....,n
x
And this is sometimes written as:
X ~ Bin (n, p)
When using the binomial formula to solve problems, we have to identify three things:
1. What is the probability of getting three heads by tossing a fair con four times?
Solution:
X ~ Bin (n 4, p 0.50)
n
P( X x) p x q n x , x 0,1,2,3,4
x
4
0.5 x 0.54 x
x
4
0.54
x
4
P( X 3) 0.54 0.25
3
112
Basic Statistics, BU, CNSc, 2015
2. Suppose that an examination consists of six true and false questions, and assume that
a student has no knowledge of the subject matter. The probability that the student will
guess the correct answer to the first question is 30%. Likewise, the probability of
guessing each of the remaining questions correctly is also 30%.
a) What is the probability of getting more than three correct answers?
b) What is the probability of getting at least two correct answers?
X ~ Bin (n 6, p 0.30)
a) P( X 3) ?
n
P( X x) p x q n x , x 0,1,2,..6
x
6
0.3 x 0.7 6 x
x
P( X 3) P( X 4) P( X 5) P( X 6)
0.060 0.010 0.001
0.071
Thus, we may conclude that if 30% of the exam questions are answered by guessing,
the probability is 0.071 (or 7.1%) that more than four of the questions are answered
correctly by the student.
b) P( X 2) ?
P( X 2) P( X 2) P( X 3) P( X 4) P( X 5) P( X 6)
0.324 0.185 0.060 0.010 0.001
0.58
P( X 3) ?
113
Basic Statistics, BU, CNSc, 2015
P( X 3) P( X 0) P( X 1) P( X 2) P( X 3)
0.118 0.303 0.324 0.185
0.93
c) P( X 5) ?
P( X 5) 1 P( X 5)
1 {P( X 5) P( X 6)}
1 (0.010 0.001)
0.989
Activity
1. Suppose that 4% of all TVs made by A&B Company in 2000 are defective. If eight of
these TVs are randomly selected from across the country and tested, what is the
probability that exactly three of them are defective? Assume that each TV is made
independently of the others.
2. An allergist claims that 45% of the patients she tests are allergic to some type of
weed. What is the probability that
a. Exactly 3 of her next 4 patients are allergic to weeds?
E ( X ) np , Var ( X ) npq
114
Basic Statistics, BU, CNSc, 2015
6.8.2Poisson Distribution
A random variable X is said to have a Poisson distribution if its probability distribution is
given by:
x e
P( X x) , x 0,1,2,......
x!
Where the averagenumber.
The Poisson distribution depends only on the average number of occurrences per unit
time of space. The Poisson distribution is used as a distribution of rare events, such as:
Number of misprints.
Natural disasters like earth quake.
Accidents.
Hereditary.
Arrivals
The process that gives rise to such events are called Poisson process.
Examples:
1. If 1.6 accidents can be expected an intersection on any given day, what is the
probability that there will be 3 accidents on any given day?
1.6 x e 1.6
X poisson1.6 p X x
x!
1.63 e 1.6
p X 3 0.1380
3!
115
Basic Statistics, BU, CNSc, 2015
Activity
On the average, five smokers pass a certain street corners every ten minutes, what is the
probability that during a given 10minutes the number of smokers passing will be
a. 6 or fewer
b. 7 or more
c. Exactly 8
If X is a Poisson random variable with parameters then
E (X ) , Var (X )
Note:
The Poisson probability distribution provides a close approximation to the binomial probability
distribution when n is large and p is quite small or quite large with np .
(np) x e ( np )
P( X x) , x 0,1,2,......
x!
Where np the averagenumber.
Example:
1. Find the binomial probability P(X=3) by using the Poisson distribution if p 0.01
and n 200
Solution:
116
Basic Statistics, BU, CNSc, 2015
U sin g Poisson , np 0.01 * 200 2
23 e 2
P ( X 3) 0.1804
3!
U sin g Binomial , n 200, p 0.01
200
P ( X 3) (0.01)3 (0.99)99 0.1814
3
1 x 2
1
f ( x) e 2
, x , , 0
2
Where E ( X ), 2 Variance( X )
and 2 are the Parametersof the Normal Distribution.
2. It is asymptotic to the axis, i.e., it extends indefinitely in either direction from the mean.
3. It is a continuous distribution.
4. It is a family of curves, i.e., every unique pair of mean and standard deviation defines a
different normal distribution. Thus, the normal distribution is completely described by two
parameters: mean and standard deviation.
5. Total area under the curve sums to 1, i.e., the area of the distribution on each side of the
mean is 0.5. f ( x)dx 1
117
Basic Statistics, BU, CNSc, 2015
7. Mean Median mod e
8. The probability that a random variable will have a value between any two points is equal to
the area under the curve between those points.
Note: To facilitate the use of normal distribution, the following distribution known as the
standard normal distribution was derived by using the transformation,
X
Z
1
1 2z 2
f ( z) e
2
6.9.2 Properties of the Standard Normal Distribution:
Same as a normal distribution, but also...
Mean is zero
Variance is one
Standard Deviation is one
Areas under the standard normal distribution curve have been tabulated in various ways.
The most common ones are the areas between
Z 0 and a positive value of Z .
a X b
P ( a X b) P ( )
a b
P ( a X b) P ( Z )
118
Basic Statistics, BU, CNSc, 2015
Note:
P ( a X b) P ( a X b)
P ( a X b)
P ( a X b)
Examples:
1. Find the area under the standard normal distribution which lies
Solution:
Area P (1.45 Z 0)
P (0 Z 1.45)
0.4265
Area P( Z 0.35)
P(0.35 Z 0) P( Z 0)
P(0 Z 0.35) P( Z 0)
0.1368 0.50 0.6368
119
Basic Statistics, BU, CNSc, 2015
Area P( Z 0.35)
1 P ( Z 0.35)
1 0.6368 0.3632
120
Basic Statistics, BU, CNSc, 2015
P( Z z ) 0.9868
P( Z 0) P(0 Z z )
0.50 P(0 Z z )
P(0 Z z ) 0.9868 0.50 0.4868
and from table
P(0 Z 2.2) 0.4868
z 2.2
3. A random variable X has a normal distribution with mean 80 and standard deviation
4.8. What is the probability that it will take a value
a) Less than 87.2
b) Greater than 76.4
c) Between 81.2 and 86
Solution:
a)
X 87.2
P( X 87.2) P( )
87.2 80
P( Z )
4.8
P( Z 1.5)
P( Z 0) P(0 Z 1.5)
0.50 0.4332 0.9332
121
Basic Statistics, BU, CNSc, 2015
b)
X 76.4
P( X 76.4) P( )
76.4 80
P( Z )
4.8
P( Z 0.75)
P( Z 0) P(0 Z 0.75)
0.50 0.2734 0.7734
c)
81.2 X 86.0
P(81.2 X 86.0) P( )
81.2 80 86.0 80
P( Z )
4.8 4.8
P(0.25 Z 1.25)
P(0 Z 1.25) P(0 Z 1.25)
0.3934 0.0987 0.2957
4. A normal distribution has mean [Link] its standard deviation if 20.0% of the area
under the normal curve lies to the right of 72.9
Solution
122
Basic Statistics, BU, CNSc, 2015
X 72.9
P( X 72.9) 0.2005 P( ) 0.2005
72.9 62.4
P( Z ) 0.2005
10.5
P( Z ) 0.2005
10.5
P (0 Z ) 0.50 0.2005 0.2995
And from table P(0 Z 0.84) 0.2995
10.5
0.84
12.5
5. A random variable has a normal distribution with 5 .Find its mean if the
probability that the random variable will assume a value less than 52.5 is 0.6915.
Solution
52.5
P( Z z ) P( Z ) 0.6915
5
P(0 Z z ) 0.6915 0.50 0.1915.
But from the table
P(0 Z 0.5) 0.1915
52.5
z 0.5
5
50
Activity
Of a large group of men, 5% are less than 60 inches in height and 40% are between 60 &
65 inches. Assuming a normal distribution, find the mean and standard deviation of
heights.
123
Basic Statistics, BU, CNSc, 2015
Summary
• A discrete probability distribution consists of the values a random variable can assume
and the corresponding probabilities of these values. There are two requirements of a
probability distribution: the sum of the probabilities of the events must equal 1, and the
probability of any single event must be a number from 0 to [Link] distributions can
be graphed.
• The mean, variance, and standard deviation of a probability distribution can be found.
The expected value of a discrete random variable of a probability distribution can also be
found. This is basically a measure of the average.
• A binomial experiment has four requirements. There must be a fixed number of trials.
Each trial can have only two outcomes. The outcomes are independent of each other, and
the probability of a success must remain the same for each trial.
The probabilities of the outcomes can be found by using the binomial formula.
• In addition to the binomial distribution, there are some other commonly used
probability distributions such as Poisson distribution.
124
Basic Statistics, BU, CNSc, 2015
Exercise 6
1. A bag contains 3 red and 4 white balls. Find the probability distribution of the number
of red balls in 3 draws with replacement from the bag.
2. An experiment consists of three independent tosses of a fair coin. Let X denote the
number of heads, Y denote the number of head runs, Z denote the length of
head runs, a head run being defined as consecutive occurrence of at least two heads,
its length being the number of heads occurring together in three tosses of the coin.
Find the probability function of i) X, ii) Y, iii) Z, iv) X+Y, v) XY and
construct the probability table.
3. A continuous random variable X follows the probability law f(x) = Ax2, 0 < x < 1.
Determine A and find the probability that X lies between 0.2 and 0.5.
4. The amount of bread (in hundreds of pounds) X that a certain bakery is able to sell in
a day is found to be a numerical valued random phenomenon, with a probability
function specified by the pdf f(x) given by:
kx, 0 x 5
f x k 10 x , 5 x 10
0, other wise
i. Find the value of K such that f(x) is a pdf.
ii. What is the probability that the number of pounds of bread that will be sold
tomorrow is:
a. More than 500 Pounds
b. Less than 500 Pounds
c. Between 250 and 750 Pounds
iii. Denoting by A, B, C the events that the pounds of bread sold are as in (a), (b) and
(c), respectively. Find P(A/B), P(A/C). Also check that whether:
a. A and B are independent
b. A and C are independent events.
5. Find the expectation of the number of failures preceding the first success in an infinite
series of independent trials with constant probability of success in each trial.
125
Basic Statistics, BU, CNSc, 2015
6. The density function of a random variable X is given by :
kx2 x , 5 x 2
2
f x
0, other wise
Find
i. K
ii. Mean and Variance of the distribution.
7. The elementary probability law of continuous random variable X is
126
Basic Statistics, BU, CNSc, 2015
11. There is rainfall in a certain place is 10 days in every thirty days. Find the probability
that:
i. There is rainfall on at least 3 days of a given week.
ii. The first four days of a given week will be wet and the remaining days dry.
12. A department in a workshop has10 machines which may need adjustment from time
to time during the day. Three of these machines are odd; each having a probability of
1/11 of needing adjustment during the day and 7 are new, having corresponding
probabilities of 1/21. Assuming that the machine needs adjustments on the same
day, determine the probability that on a particular day.
i. Just 2 old and no new machines need adjustment.
ii. Just 2 machines need adjustment which is of the same type.
13. A wireless set is manufactured with 25 soldered joints each. On an average one
joint in 500 are defective. How many sets can be expected to be free from
defective joints in a consignment of 10,000 sets?
14. Red blood deficiency may be determined by examining a specimen of the blood
under the microscope. Suppose a certain small fixed volume contains on an
average 20 red calls for a normal person. Using poison distribution, obtain the
probability that a specimen from a normal person will contain less than 15 red
cells.
15. An insurance company has discovered that only 0.1% of the population is
involved in a certain type of accident each year. If its 10,000 policy holders
more randomly selected from the population, what is the probability that not more
than 5 of its clients will be involved is such an accident next year?
16. A company finds that the time taken by one of its engineers to complete a repair
job has a normal distribution with mean 40 minutes and S.D 5 minutes. State what
proportion of jobs take:
i. Less than 35 minutes.
ii. More than 48 minutes.
127
Basic Statistics, BU, CNSc, 2015
CHAPTER 7
7. SAMPLING AND SAMPLING DISTRIBUTION
7.1 Introduction
Given a variable X, if we arrange its values in ascending order and assign probability to
each of the values or if we present Xi in a form of relative frequency distribution the
result is called Sampling Distribution of X.
7.2 Objectives
After completing this chapter, you should be able to
- List of households.
- List of students in the registrar office.
128
Basic Statistics, BU, CNSc, 2015
7.2Errors in sample survey
There are two types of errors
a) Sampling error:
It is the discrepancy between the population value and sample value. May arise due to in
appropriate sampling techniques applied
b) Non sampling errors: are errors due to procedure bias such as:
- Due to incorrect responses
- Measurement
- Errors at different stages in processing the data.
The Need (reason) for Sampling
- Reduced cost
- Greater speed
- Greater accuracy
- Greater scope
- More detailed information can be obtained.
There are two types of sampling.
Examples:
7.3.3Cluster Sampling
The population is divided in to non-overlapping groups called clusters. A simple random
sample of groups or cluster of elements is chosen and all the sampling units in the
selected clusters will be surveyed. Clusters are formed in a way that elements within a
cluster are heterogeneous, i.e. observations in each cluster should be more or less
dissimilar. Cluster sampling is useful when it is difficult or costly to generate a simple
random sample. For example, to estimate the average annual household income in a large
city we use cluster sampling, because to use simple random sampling we need a complete
list of households in the city from which to sample. To use stratified random sampling,
we would again need the list of households. A less expensive way is to let each block
within the city represent a cluster. A sample of clusters could then be randomly selected,
and every household within these clusters could be interviewed to find the average annual
household income.
A complete list of all elements within the population (sampling frame) is required. The
procedure starts in determining the first element to be included in the sample. Then the
technique is to take the kth item from the sampling frame.
130
Basic Statistics, BU, CNSc, 2015
Let
N
N population size, n sample size, k sampling int erval.
n
Chose any number between 1 and k . Suppose it is j (1 j k ) .
The j th unit is selected at first and then ( j k )th , ( j 2k )th ,....etc until the
required sample size is reached.
Example:
Judgment sampling.
Convenience sampling
Quota Sampling.
7.4.1 Judgment Sampling
In this case, the person taking the sample has direct or indirect control over which items
are selected for the sample.
Note:
131
Basic Statistics, BU, CNSc, 2015
N
We have possible samples if sampling is without replacement.
n
2. After this on wards we consider that samples are drawn from a given
population using simple random sampling.
7.5 Sampling Distribution of the sample mean
Sampling distribution of the sample mean is a theoretical probability distribution that
shows the functional relationship between the possible values of a given sample mean
based on samples of size n and the probability associated with each value, for all
possible samples of size n drawn from that particular population.
There are commonly three properties of interest of a given sampling distribution.
Its Mean
Its Variance
Its Functional form.
7.6 Steps for the construction of Sampling Distribution of the mean
1. From a finite population of size N , randomly draw all possible samples of size n .
Populationmean 10
population Variance 2 8
Take samples of size 2 with replacement and construct sampling distribution of the
sample mean.
Solution:
N 5, n 2
132
Basic Statistics, BU, CNSc, 2015
We have N n 52 25 possible samples since sampling is with replacement.
Step 1: Draw all possible samples:
6 8 10 12 14
6 8 10 12 14
6 6 7 8 9 10
8 7 8 9 10 11
10 8 9 10 11 12
12 9 10 11 12 13
14 10 11 12 13 14
X Frequency
133
Basic Statistics, BU, CNSc, 2015
6 1
7 2
8 3
9 4
10 5
11 4
12 3
13 2
14 1
X
X i f i 250
10
i
f 25
( X i X ) 2 f i 100
X 2
4 2
fi 25
Remark:
2
X 2
n
2. If sampling is without replacement
2 N n
X 2
n N 1
134
Basic Statistics, BU, CNSc, 2015
3. In any case the sample mean is unbiased estimator of the population mean.i.e
X E (X ) (Show!)
2
i.e. X
2
n
2
X ~ N ( , )
n
X
Z ~ N (0,1)
n
2
approximately normally distributed with mean and variance , when the sample size is
n
large.
Summary
• To obtain information and make inferences about a large population, researchers select
a sample. A sample is a subgroup of the population. Using a sample rather than a
population, researchers can save time and money, get more detailed information, and get
information that otherwise would be impossible to obtain.
135
Basic Statistics, BU, CNSc, 2015
• The four most common methods researchers use to obtain samples are random,
systematic, stratified, and cluster sampling methods. In random sampling, some type of
random method (usually random numbers) is used to obtain the sample. In systematic
sampling, the researcher selects every kth person or item after selecting the first one at
random. In stratified sampling, the population is divided into subgroups according to
various characteristics, and elements are then selected at random from the subgroups. In
cluster sampling, the researcher selects an intact group to use as a sample. When the
population is large, multistage sampling (a combination of methods) is used to obtain a
subgroup of the population.
• Researchers must use caution when conducting surveys and designing questionnaires;
otherwise, conclusions obtained from these will be inaccurate.
• Most sampling methods use random numbers, which can also be used to simulate many
real-life problems or situations.
The purpose of simulation is to duplicate situations that are too dangerous, too costly, or
too time-consuming to study in real life. Most simulation techniques can be done on the
computer or calculator, since they can rapidly generate random numbers, count the
outcomes, and perform the necessary computations. Sampling and simulation are two
techniques that enable researchers to gain information that might otherwise be
unobtainable.
Exercises 7
1) Suppose that the population distribution of the gripping strengths of industrial
workers is known to have a mean of 110 and standard deviation of 10. For a random
sample of 75 workers, what is the probability that the sample mean gripping strength
will be
a) Between 109 and 112
b) Greater than112?
2) The amount of sulphur in a daily emission from a factory has a normal distribution
with mean of 134 pounds and a standard deviation of 22pounds. For a day selected
randomly, find the probability that the mean amount of sulphur emission will be less
than 130 pounds.
136
Basic Statistics, BU, CNSc, 2015
3) A population consists of the four numbers, 3,7,11, 13 and 15. Consider all possible
samples of size 2 drawn from this population without replacement.
Find
137
Basic Statistics, BU, CNSc, 2015
CHAPTER 8
8. ESTIMATION AND HYPOTHESIS TESTING
8.1 Introduction
Researchers are interested in answering many types of questions. For example, a scientist
might want to know whether the earth is warming up. A physician might want to know
whether a new medication will lower a person‟s blood pressure. An educator might wish
to see whether a new teaching technique is better than a traditional one. A retail merchant
might want to know whether the public prefers a certain color in a new line of fashion.
Automobile manufacturers are interested in determining whether seat belts will reduce
the severity of injuries caused by accidents. These types of questions can be addressed
through statistical hypothesis testing, which is a decision-making process for evaluating
claims about a population. In hypothesis testing, the researcher must define the
population under study, state the particular hypotheses that will be investigated, give the
significance level, select a sample from the population, collect the data, perform the
calculations required for the statistical test, and reach a conclusion. Hypotheses
concerning parameters such as means and proportions can be investigated.
Inference is the process of making interpretations or conclusions from sample data for the
totality of the population. It is only the sample data that is ready for inference. In statistics
there are two ways though which inference can be made.
Statistical estimation
Statistical hypothesis testing.
138
Basic Statistics, BU, CNSc, 2015
Inference
Analyzed
Populatio
n
Numerica
Sample
l data
Data analysis is the process of extracting relevant information from the summarized data.
8.2 Objectives
After completing this chapter, the student should be able to:
139
Basic Statistics, BU, CNSc, 2015
8.3.1Point Estimation
It is a procedure that results in a single value as an estimate for a parameter.
Definitions:
Confidence Level: The percent of the time the true value will lie in the interval estimate
given.
Consistent Estimator: An estimator which gets closer to the value of the parameter as
the sample size increases.
Degrees of Freedom: The number of data values which are allowed to vary once a
statistic has been determined.
Relatively Efficient Estimator: The estimator for a parameter with the smallest
variance.
Unbiased Estimator: An estimator whose expected value is the value of the parameter
being estimated.
140
Basic Statistics, BU, CNSc, 2015
8.4Point and Interval estimation of the population mean: µ
8.4.1 Point Estimation
Another term for statistic is point estimate, since we are estimating the parameter value.
A point estimator is the mathematical way we compute the point estimate. For instance,
sum of xi over n is the point estimator used to compute the estimate of the population
xi
means, .That is X is a point estimator of the population mean.
n
i. Confidence interval estimation of the population mean
Although X possesses nearly all the qualities of a good estimator, because of sampling
error, we know that it's not likely that our sample statistic will be equal to the population
parameter, but instead will fall into an interval of values. We will have to be satisfied
knowing that the statistic is "close to" the parameter. That leads to the obvious question,
what is "close"?
We can phrase the latter question differently: How confident can we be that the value of
the statistic falls within a certain "distance" of the parameter? Or, what is the probability
that the parameter's value is within a certain range of the statistic's value? This range is
the confidence interval. The confidence level is the probability that the value of the
parameter falls within the range specified by the confidence interval surrounding the
statistic. There are different cases to be considered to construct confidence intervals.
Case 1: If sample size is large or if the population is normal with known variance
Recall the Central Limit Theorem, which applies to the sampling distribution of the mean
of a sample. Consider samples of size n drawn from a population, whose mean is and
standard deviation is with replacement and order important. The population can have
any frequency distribution. The sampling distribution of X will have a mean x
and a standard deviation x , and approaches a normal distribution as n gets
n
large.
141
Basic Statistics, BU, CNSc, 2015
This allows us to use the normal distribution curve for computing confidence intervals.
X
Z has a normal distribution with mean 0 and var iance 1
n
X Z n
X , where is a measureof error.
Z n
For the interval estimator to be good the error should be small. How it be small?
By making n large
Small variability
Taking Z small
To obtain the value of Z, we have to attach this to a theory of chance. That is, there is an area of
size 1 such
P( Z 2 Z Z 2 ) 1
Where is the probability that the parameterlies outsidethe int erval
Z 2 s tan ds for the s tan dard normal var iableto the right of which
2 probability lies, i.e P( Z Z 2 ) 2
X
P( Z 2 Z 2 ) 1
n
P( X Z 2 n X Z 2 n) 1
142
Basic Statistics, BU, CNSc, 2015
Here are the z values corresponding to the most commonly used confidence levels.
100(1 ) % 2 Z 2
X
t has t distributi on with n 1 deg rees of freedom.
S n
The unit of measurement of the confidence interval is the standard error. This is just the
standard deviation of the sampling distribution of the statistic.
Example:
1. From a normal sample of size 25 a mean of 32 was found .Given that the population
standard deviation is 4.2. Find
a) A 95% confidence interval for the population mean.
b) A 99% confidence interval for the population mean.
Solution:
143
Basic Statistics, BU, CNSc, 2015
b)
2. A drug company is testing a new drug which is supposed to reduce blood pressure.
From the six people who are used as subjects, it is found that the average drop in
blood pressure is 2.28 points, with a standard deviation of .95 points. What is the 95%
confidence interval for the mean change in pressure?
Solution:
That is, we can be 95% confident that the mean decrease in blood pressure is between 1.28 and
3.28 points.
144
Basic Statistics, BU, CNSc, 2015
8.5 Hypothesis Testing
This is also one way of making inference about population parameter, where the
investigator has prior notion about the value of the parameter.
Definitions:
Statistical hypothesis: is an assertion or statement about the population whose
plausibility is to be evaluated on the basis of the sample data.
Test statistic: is a statistics whose value serves to determine whether to reject or accept
the hypothesis to be tested. It is a random variable.
Statistic test: is a test or procedure used to evaluate a statistical hypothesis and its value
depends on sample data.
There are two types of hypothesis:
8.5.1 Null hypothesis:
It is the hypothesis to be tested. It is the hypothesis of equality or the hypothesis of no
difference. Usually denoted by H0.
8.5.2Alternative hypothesis:
It is the hypothesis available when the null hypothesis has to be rejected. It is the
hypothesis of difference. Usually denoted by H1 or Ha.
Decision
145
Basic Statistics, BU, CNSc, 2015
Type I error: Rejecting the null hypothesis when it is true.
NOTE:
1. There are errors that are prevalent in any two choice decision making problems.
2. There is always a possibility of committing one or the other errors.
3. Type I error ( ) and type II error ( ) have inverse relationship and therefore,
cannot be minimized at the same time. In practice we set at some value and
design a test that minimize . This is because a type I error is often considered to be
more serious, and therefore more important to avoid, than a type II error.
8.6.1 General steps in hypothesis testing:
1. The first step in hypothesis testing is to specify the null hypothesis (H0) and the
alternative hypothesis (H1).
[Link] next step is to select a significance level,
[Link] the sampling distribution of the estimator.
[Link] fourth step is to calculate a statistic analogous to the parameter specified by the
null hypothesis.
[Link] the critical region.
[Link] decision.
[Link] of the result.
8.7 Hypothesis testing about the population means:
Suppose the assumed or hypothesized value of is denoted by 0 , then one can formulate
1. H 0 : 0 vs H1 : 0
2. H 0 : 0 vs H1 : 0
3. H 0 : 0 vs H1 : 0
Case 1: When sampling is from a normal distribution with 2 known
The relevant test statistic is Z X
n
146
Basic Statistics, BU, CNSc, 2015
After specifying we have the following regions (critical and acceptance) on the
standard normal distribution corresponding to the above three hypothesis.
Summary Table for the decision rule
H0 Reject H0 if Accept H0 if Inconclusive if
X 0
Where: Z cal
n
Case 2: When sampling is from a normal distribution with 2 unknown and small
sample size
X
t ~ t with n 1 deg rees of freedom.
S n
After specifying we have the following regions on the student t-distribution
147
Basic Statistics, BU, CNSc, 2015
0 tcal t tcal t tcal t
X 0
Where: t cal
S n
If a sample size is large one can perform a test hypothesis about the mean by using:
X 0
Z cal , if 2 is known.
n
X 0
, if 2 is unknown.
S n
Examples:
1. Test the hypotheses that the average height content of containers of certain lubricant is 10
liters if the contents of a random sample of 10 containers are 10.2, 9.7, 10.1, 10.3, 10.1, 9.8,
9.9, 10.4, 10.3, and 9.8 liters. Use the 0.01 level of significance and assume that the
distribution of contents is normal.
Solution:
H 0 : 10 vs H1 : 10
t- Statistic is appropriate because population variance is not known and the sample size is also
small.
148
Basic Statistics, BU, CNSc, 2015
Step 4: identify the critical region.
Here we have two critical regions since we have two tailed hypothesis.
Step 5: Computations:
X 10.06, S 0.25
X 0 10.06 10
t cal 0.76
S n 0.25 10
Step 6: Decision
Step 7: Conclusion
At 1% level of significance, we have no evidence to say that the average height content of
containers of the given lubricant is different from 10 litters, based on the given sample data.
2. The mean life time of a sample of 16 fluorescent light bulbs produced by a company is
computed to be 1570 hours. The population standard deviation is 120 hours. Suppose the
hypothesized value for the population mean is 1600 hours. Can we conclude that the life time
of light bulbs is decreasing?
(Use 0.05 and assume the normality of the population)
Solution:
H 0 : 1600 vs H1 : 1600
149
Basic Statistics, BU, CNSc, 2015
Step 2: select the level of significance, 0.05 ( given)
Step 3: Select an appropriate test statistics
Step 5: Computations:
X 0 1570 1600
Z cal 1.0
n 120 16
Step 6: Decision
Step 7: Conclusion
At 5% level of significance, we have no evidence to say that that the life time of light bulbs is
decreasing, based on the given sample data.
Activity
It is known in a pharmacological experiment that rats fed with a particular diet over a certain
period gain an average of 40 gms in weight. A new diet was tried on a sample of 20 rats yielding
a weight gain of 43 gms with variance 7 gms2 . Test the hypothesis that the new diet is an
improvement assuming normality.
150
Basic Statistics, BU, CNSc, 2015
8.8 Test of Association
Suppose we have a population consisting of observations having two attributes or
qualitative characteristics say A and B. If the attributes are independent then the
probability of possessing both A and B is PA*PB
A B1 B2 . . Bj . Bc Total
. .
. .
. .
. .
. .
. .
Total C1 C2 Cj N
151
Basic Statistics, BU, CNSc, 2015
The chi-square procedure test is used to test the hypothesis of independency of two
attributes .For instance we may be interested
(Oij eij ) 2
~ ( r 1)(c 1)
r c
cal
2 2
i 1 j 1 eij
given by:
Ri * C j
eij
n
Remark:
r c r c
n Oij eij
i 1 j 1 i 1 j 1
152
Basic Statistics, BU, CNSc, 2015
H 0 : Thereis no association between A and B.
H1 : not H 0 ( Thereis association between A and B).
(Oij eij ) 2
2 ( r 1)(c 1) at
r c
Reject H 0 if cal
2
i 1 j 1 eij
Examples:
1. A geneticist took a random sample of 300 men to study whether there is association
between father and son regarding boldness. He obtained the following results.
Son
Bold 85 59
Not 65 91
Using 5% test whether there is association between father and son regarding
boldness.
Solution:
153
Basic Statistics, BU, CNSc, 2015
Then calculate the expected frequencies( eij‟s)
Ri * C j
eij
n
R1 * C1 144 *150
e11 72
n 300
R1 * C2 144 *150
e12 72
n 300
R2 * C1 156 *150
e21 78
n 300
R2 * C2 156 *150
e22 78
n 300
Obtain the calculated value of the chi-square.
2 2 (Oij eij ) 2
2
cal
i 1 j 1
eij
(85 72) 2 (59 72) 2 (65 78) 2 (91 78) 2
9.028
72 72 78 78
Obtain the tabulated value of chi-square
0.05
Degrees of freedom (r 1)(c 1) 1*1 1
02.05 (1) 3.841 from table.
154
Basic Statistics, BU, CNSc, 2015
Conclusion: At 5% level of significance we have evidence to say there is association
between father and son regarding boldness, based on this sample data.
2. Random samples of 200 men, all retired were classified according to education and number
of children is as shown below
Education Number of children
level
0-1 2-3 Over 3
Elementary 14 37 32
Secondary 31 59 27
and above
Test the hypothesis that the size of the family is independent of the level of education
attained by fathers. (Use 5% level of significance)
Solution:
H 0 : There is no associatio n between the size of the family and the level of
education attained by fathers.
H1 : not H 0 .
First calculate the row and column totals
Ri * C j
eij
n
155
Basic Statistics, BU, CNSc, 2015
2 3 (Oij eij ) 2
2
cal
i 1 j 1 e
ij
0.05
Degrees of freedom (r 1)(c 1) 1* 2 2
02.05 (2) 5.99 from table.
156
Basic Statistics, BU, CNSc, 2015
Summary
This chapter introduces the basic concepts of hypothesis testing. A statistical
hypothesis is a conjecture about a population. There are two types of statistical
hypotheses: the null and the alternative hypotheses. The null hypothesis states that
there is no difference, and the alternative hypothesis specifies a difference. To test
the null hypothesis, researchers use a statistical test. Researchers compute a test
value from the sample data to decide whether the null hypothesis should be
rejected. Statistical tests can be one-tailed or two-tailed, depending on the
hypotheses.
The null hypothesis is rejected when the difference between the population
parameter and the sample statistic is said to be significant. The difference is
significant when the test value falls in the critical region of the distribution. The
critical region is determined by a, the level of significance of the test. The level is
the probability of committing a type I error. This error occurs when the null
hypothesis is rejected when it is true.
A second kind of error, the type II error, can occur when the null hypothesis is not
rejected when it is false.
• There are two common methods used to test hypotheses; they are the traditional method
and the P-value method.
• All hypothesis-testing situations using the traditional method should include the
following steps:
1. State the null and alternative hypotheses and identify the claim.
2. State an alpha level and find the critical value(s).
3. Compute the test value.
4. State critical value
5. State rejection region
6. Make the decision to reject or not reject the null hypothesis.
7. Summarize the results.
157
Basic Statistics, BU, CNSc, 2015
• The z-test is used to test a mean when the population standard deviation is known.
When the sample size is less than 30, the population values need to be normally
distributed.
Important Formulas
• When the population standard deviation is not known, researchers use a t-test to test a
claim about a mean. If the sample size is less than 30, the population values need to be
normally or approximately normally distributed.
• There is a relationship between confidence intervals and hypothesis testing. When the
null hypothesis is rejected, the confidence interval for the mean using the same level of
significance will not contain the hypothesized mean. When the null hypothesis is not
rejected, the confidence interval, using the same level of significance, will contain the
hypothesized mean.
158
Basic Statistics, BU, CNSc, 2015
Exercise 8
1. An electrical firm manufactures light bulbs that have a length of life that is
approximately normally distributed with a standard deviation of 40 hours. If a random
sample of 30 bulbs has an average life of 780 hours, find a 99% confidence interval
for the population mean of all bulbs produced by this firm.
2. A random sample of 400 households was drawn from a town and a survey generated
data on weekly earning. The mean in the sample was Birr 250 with a standard
deviation Birr 80. Construct a 95% confidence interval for the population mean
earning.
3. A major truck has kept extensive records on various transactions with its
customers. If a random sample of 16 of these records shows average sales of
290 liters of diesel fuel with a standard deviation of 12 liters, construct a 95%
confidence interval for the mean of the population sampled.
4. The manufacturer of a certain type of battery is trying to estimate the lifetime of the
battery. He believes each battery will last for a random amount of time that has a
N(µ,100) distribution. (The lifetimes are measured in hours.) He carries out an
experiment to estimate µ. A sample of 400 batteries is tested and their lifetimes are
measures. The (sample) mean lifetime is found to be 74.2 hours. Calculate a 95%
confidence interval for µ. How do you interpret this interval?
5. A biostatistician intends to estimate µ, the mean blood pressure of women
between the ages of 45 and 50. She takes a random sample of 20 women and
measures their blood pressure. Based on past experience she believes the
measurements will follow a N(µ, 100) distribution. (Measurements are in mm
mercury.) Suppose she discovers the sample mean is equal to 136.9 mm mercury.
Find a 95% confidence interval for µ.
6. A biologist measured a random sample of 12 fossil skeletons of an extinct species of
bird. He found that their skulls had a mean length of 6.34cm and a standard
deviation of 0.45cm. He believes that the lengths of the skulls follow a
normal distribution. Use the data to obtain a 95% confidence interval for the mean of
this distribution.
159
Basic Statistics, BU, CNSc, 2015
7. A machine is set up such that the average content of juice per bottle equals µ. A
sample of 100 bottles yields an average content of [Link] a 90% and a
95% confidence interval for the average content. Assume that the population
standard deviation σ= 5cl.
8. What sample size is required to estimate the average contents to within 0.5cl at the
95% confidence level? (= + or - 0.5 cl) Assume that the population standard
deviation σ= 5cl.
9. A machine is set up such that the average content of juice per bottle equals µ. A
sample of 36 bottles yields an average content of48.5cl. Test the hypothesis that the
average content per bottle is 50cl at the 5% significance level. Assume that the
population standard deviation σ= 5cl.
10. A machine is set up such that the average content of juice per bottle equals µ. A
sample of 100 bottles yields an average content of 48.8cl. Test the hypothesis that the
average content per bottle is 50cl at the 5% significance level. Compare the
conclusion to that based on the 36 bottles sample. Assume that the population
standard deviation σ= 5cl.
11. A machine is set up such that the average content of juice per bottle equals µ. A
sample of 36 bottles yields an average content of [Link] you reject the
hypothesis that the average content per bottle is less than or equal to 45cl in favor of
the alternative that it exceeds 45cl (5% significance level)? Assume that the
population standard deviation σ= 5cl.
12. The manager claims that the average content of juice per bottle is less than 50cl. The
machine operator disagrees. A sample of 100 bottles yields an average content of 49cl
per bottle. Does this sample allow the manager to claim he is right (5%
significance level)? Assume that the population standard deviation σ= 5 cl.
13. Out of a sample of 80 customers 60 of them reply they are satisfied with the
service they received .Calculate a 95% confidence interval for the proportion
of satisfied customers .
160
Basic Statistics, BU, CNSc, 2015
14. Random samples of 200 men, all retired were classified according to education and
number of children is as shown below.
Education level
No. of Children
0-1 2-3 Over 3
Elementary 14 37 32
Secondary and above 31 59 27
Test the hypothesis that the size of the family is independent of the level of
education attained by fathers. (Use 5% level of significance)
15. From a normal population with the standard deviation is 4.2. A sample of size
25 are taken with mean of 32. Find a 99% confidence interval for the population
mean.
161
Basic Statistics, BU, CNSc, 2015
CHAPTER 9
9. SIMPLE LINEAR REGRESSION AND CORRELATION
9.1 Introduction
Linear regression and correlation is studying and measuring the linear relationship among
two or more variables. When only two variables are involved, the analysis is referred to
as simple correlation and simple linear regression analysis, and when there are more than
two variables the term multiple regression and partial correlation is used.
Regression Analysis: is a statistical technique that can be used to develop a
mathematical equation showing how variables are related.
Correlation Analysis: deals with the measurement of the closeness of the relationship
which are described in the regression equation.
We say there is correlation when the two series of items vary together directly or
inversely.
9.2 Objectives
After completing this chapter, students should be able to:
162
Basic Statistics, BU, CNSc, 2015
When higher values of X are associated with higher values of Y and lower values
of X are associated with lower values of Y, then the correlation is said to be
positive or direct.
Examples:
1. One variable being the cause of the other. The cause is called “subject” or
“independent” variable, while the effect is called “dependent” variable.
2. Both variables being the result of a common cause. That is, the correlation that
exists between two variables is due to their being related to some third force.
Example:
163
Basic Statistics, BU, CNSc, 2015
Y2=be the rate of getting a scholar ship.
Both X1&Y1 and X1&Y2 have high positive correlation, likewise Y1 & Y2 have positive
correlation but they are not directly related, but they are related to each other via X1.
[Link]:
The correlation that arises by chance is called spurious correlation.
Examples:
r
( X i X )(Yi Y ) and the short cut formula is
i
( X X ) 2
i
(Y Y ) 2
n XY ( X )( Y )
r
[n X 2 ( X ) 2 ] [n Y 2 ( Y ) 2
r
XY nXY
[ X 2 nX 2 ] [ Y 2 nY 2 ]
Remark:
164
Basic Statistics, BU, CNSc, 2015
[Link] negative linear relationship ( if r 1)
Examples:
1. Calculate the simple correlation between mid semester and final exam scores of 10
students (both out of 50)
(X) (Y)
1 31 31
2 23 29
3 41 34
4 32 35
5 29 25
6 33 35
7 28 33
8 31 42
9 31 31
10 33 34
Solution:
165
Basic Statistics, BU, CNSc, 2015
r
XY nXY
[ X 2 nX 2 ] [ Y 2 nY 2 ]
10331 10(31.2)(32.9)
(9920 10(973.4)) (11003 10(1082.4))
66.2
0.363
182.5
This means mid semester exam and final exam scores have a slightly positive correlation.
Activity
The following data were collected from a certain household on the monthly income (X)
and consumption (Y) for the past 10 months. Compute the simple correlation coefficient.(
X: 650 654 720 456 536 853 735 650 536 666
Y: 450 523 235 398 500 632 500 635 450 360
The above formula and procedure is only applicable on quantitative data, but when we
have qualitative data like efficiency, honesty, intelligence, etc
Steps
i. Rank the different items in X and Y.
ii. Find the difference of the ranks in a pair , denote them by D i
iii. Use the following formula
6 Di
2
rs 1
n(n 2 1)
Where rs coefficien t of rank correlatio n
D the difference between paired ranks
n the number of pairs
Example:
166
Basic Statistics, BU, CNSc, 2015
Aster and Almaz were asked to rank 7 different types of lipsticks, see if there is
correlation between the tests of the ladies.
Lipsticks A B C D E F G
Aster 2 1 4 3 5 7 6
Almaz 1 3 2 4 5 6 7
Solution:
X Y R1-R2 D2
2 1 1 1
1 3 -2 4
4 2 2 4
3 4 -1 1
5 5 0 0
7 6 1 1
6 7 -1 1
Total 12
6 Di
2
6(12)
rs 1 1 0.786
n(n 1)
2
7(48)
167
Basic Statistics, BU, CNSc, 2015
9.4 Simple Linear Regression
Simple linear regression refers to the linear relationship between two variables. We
usually denote the dependent variable by Y and the independent variable by X. A simple
regression line is the line fitted to the points plotted in the scatter diagram, which would
describe the average relationship between the two variables. Therefore, to see the type of
relationship, it is advisable to prepare scatter plot before fitting the model.
Yˆ a bX
Where a is a constant which gives the value of Y when X=0 .It is called the Y-
intercept. b is a constant indicating the slope of the regression line, and it gives a
measure of the change in Y for a unit change in X. It is also regression coefficient of Y
on X.
168
Basic Statistics, BU, CNSc, 2015
Where : Yi observed value
Yˆi estimated value a bX i
b
( X i X )(Yi Y ) XY nXY
( X i X )2 X 2 nX 2
a Y bX
Example 1: The following data shows the score of 12 students for Accounting and Statistics
Examinations.
Accounting Statistics
X Y
1 74.00 81.00
2 93.00 86.00
3 55.00 67.00
4 41.00 35.00
169
Basic Statistics, BU, CNSc, 2015
5 23.00 30.00
6 92.00 100.00
7 64.00 55.00
8 40.00 52.00
9 71.00 76.00
10 33.00 24.00
11 30.00 48.00
12 71.00 87.00
a.
170
Basic Statistics, BU, CNSc, 2015
Accounting Statistics
X2 Y2 XY
X Y
b)
171
Basic Statistics, BU, CNSc, 2015
The Coefficient of Correlation (r) has a value of 0.92. This indicates that the two
variables are positively correlated (Y increases as X increases).
c)
Using OLS:
172
Basic Statistics, BU, CNSc, 2015
Figure 9.2 Scatter Diagram and Regression Line
Yˆ 7.0194 0.9560 X
7.0194 0.9560(85) 88.28
Activity
A car rental agency is interested in studying the relationship between the distance
driven in kilometer (Y) and the maintenance cost for their cars (X in birr). The
following summarized information is given based on samples of size 5.
2
i 1 X i 147,000,000 i 1Yi 314
5 5 2
To know how far the regression equation has been able to explain the variation in Y we
2
use a measure called coefficient of determination ( r )
(Yˆ Y ) 2
i.e r 2
(Y Y ) 2
Where r the simple correlatio n coefficient.
r2 gives the proportion of the variation in Y explained by the regression of Y on X.
SX Y
( X i X )(Yi Y ) XY nXY
n 1 n 1
174
Basic Statistics, BU, CNSc, 2015
Summary
• Many relationships among variables exist in the real world. One way to determine
whether a linear relationship exists is to use the statistical techniques known as
correlation and regression. The strength and direction of a linear relationship are
measured by the value of the correlation coefficient. It can assume values between and
including -1 and +1. The closer the value of the correlation coefficient is to -1 or +1, the
stronger the linear relationship is between the variables. A value of -1 or +1 indicates a
perfect linear relationship. A positive relationship between two variables means that for
small values of the independent variable, the values of the dependent variable will be
small, and that for large values of the independent variable, the values of the dependent
variable will be large. A negative relationship between two variables means that for small
values of the independent variable, the values of the dependent variable will be large, and
that for large values of the independent variable, the values of the dependent variable will
be small.
• Remember that a significant relationship between two variables does not necessarily
mean that one variable is a direct cause of the other variable. In some cases this is true,
but other possibilities that should be considered include a complex relationship involving
other (perhaps unknown) variables, a third variable interacting with both variables, and a
relationship due solely to chance.
• Relationships can be linear or nonlinear. To determine the shape, you draw a scatter plot
of the variables. If the relationship is linear, the data can be approximated by a straight
line, called the regression line, or the line of best fit. The closer the value of r is to -1 or
+1, the more closely the points will fit the line.
175
Basic Statistics, BU, CNSc, 2015
Exercise 9
1. The following are advertised sale prices of color televisions in Addis with different
size.
Size (inches) 9 20 27 31 35 40 60
Sale Price ($) 147 197 297 447 1177 2177 2497
a) Decide which variable should be the independent variable and which should be
the dependent variable.
b) Make a scatter plot of the data.
c) Does it appear from inspection that there is a relationship between the
variables?
d) Calculate the least squares line. Put the equation in the form of: y= a+ bx
e) Find and interpret the correlation coefficient.
f) Find the estimated sale price for a 32 inch television
g) What is the slope of the least squares (best-fit) line? Interpret the slope.
1. The monthly income (X) and monthly food expenditure (Y) of 11 households (in
hundreds of (birr) are taken randomly to fit linear relationship between the two
variables.
X 3.8 4.5 2.5 4.8 7.7 5.0 12.6 8.5 5.5 7.1 3.5
Y 3.1 3.6 2.3 3.7 4.6 4.1 6.5 5.1 4.0 4.1 3.2
176
Basic Statistics, BU, CNSc, 2015
ANSWER FOR SELECTED EXERCISE
Exercise 1
3. C. Yes. The implied population is the responses (watch/not watch) of all TV owners in
Addis Ababa
4. a. The weight of each pineapple in the experimental field and/or the maximum
girth of each pineapple in the experimental field.
a. We doubt about the mechanisms how the mission is measured and quantified.
This leads miss use of statistical figures.
3.00. There is a student who has scored above 3.00 and below
3.00.
i. Ordinal scale
h. Nominal scale
i. Nominal scale
j. Ordinal scale
k. Ordinal scale
177
Basic Statistics, BU, CNSc, 2015
l. Ratio Scales
m. Interval scale.
Exercise 2
4.
5.
Classes Frequency
23 - 26 3
27 - 30 4
31 - 34 3
35 - 38 5
39 - 42 5
Total 20
7.
178
Basic Statistics, BU, CNSc, 2015
70 //// 4
74 / 1
75 // 2
76 / 1
80 /// 3
Exercise 3
12. Mode = 5
13. 22.5
Exercise 5
1. 97/200
2. 1/6
3. 4/21
4. 23/40
6.13/15
7. 60/36!
179
Basic Statistics, BU, CNSc, 2015
8. 0.046
9. 0.5
10. 9/19
15. 4/65
Exercise 6
8 i) 2 ii) 0.6778
13. 9512
e 20 20
14 x
14.
x 0 x!
15. 0.06651
Exercise 7
1. a) 0.7633 b)0.0418
2. 0.57
Exercise 8
180
Basic Statistics, BU, CNSc, 2015
1. (761.19, 798.81)
2. (242.16, 257.84)
3. (283.61, 296.39
8. 384
9. H0 is not rejected
11. H0 is rejected
12. H0 is rejected
2
cal 63, tabulate
2
02.05, 2 5.99
15. A 99% confidence interval for the population mean is (29.8328, 34.1672)
Exercise 9
1. c) Yes
f) 1008.6
g) 54.8 as the size of the television increases by one inch, the average price increases
by 54.8 units.
181
Basic Statistics, BU, CNSc, 2015
2. b) yes c) y =65.0876+7.0948
f) 72.2 cm
j) Slope = 7.0948. As the age of boy increases by one year, the average height
182
Basic Statistics, BU, CNSc, 2015
References
Eshetu Wencheko (2000). Introduction to Statistics. Addis Ababa University Press.
Bluman, A.G. (1995). Elementary Statistics: A Step by Step Approach (2nd Ed.).Wm.
C. Brown Communications, Inc.
Freund, J.E and Simon, G.A. (1998). Modern Elementary Statistics (9th Ed.). .
Gupta, C.B. and Gupta, V. (2004). An Introduction to Statistical Methods. Vikas
Publishing House, Pvt. Ltd, India.
Spiegel, M.R. and Stephens, L.J. (2007). Schaum's Outline Series (4th Ed.). McGraw-
Hill, New York.
183
Basic Statistics, BU, CNSc, 2015
Appendix: Tables
A. The Standard Normal Distribution Table
184
Basic Statistics, BU, CNSc, 2015
B. The Student’s t-distribution Table
185
Basic Statistics, BU, CNSc, 2015
C. The Chi-Square distribution Table
186
Basic Statistics, BU, CNSc, 2015
187
Basic Statistics, BU, CNSc, 2015