Module 6: Introduction to Statistics
Learning Objectives
After careful study of this chapter, you should be able to do the
following:
1. Compute and interpret the sample mean, sample variance, sample
standard deviation, sample median, and sample range.
2. Explain the concepts of sample mean, sample variance, population
mean, and population variance.
3. Construct and interpret visual data displays, including the stem-
and-leaf display, the histogram, and the box plot.
6. Construct and interpret normal probability plots
7. Explain how to use box plots and other data displays to visually
compare two or more samples of data
8. Know how to use simple time series plots to visually display the
important features of time-oriented data
9. Know how to construct and interpret scatter diagrams of two or
more variables.
What is Statistics?
According to the Encyclopedia Britannica:
“Statistics is the art and science of gathering, analyzing, and making
inferences from data”
In his book “Statistical Models”, A. C. Davison answers the question
more thorough way:
“Statistics concerns what can be learned from data. Applied statistics
comprises a body of methods for data collection and analysis across the
whole range of science, and in areas such as engineering, medicine, business,
and law - wherever variable data must be summarized, or used to test or
confirm theories, or to inform decisions. Theoretical statistics underpins
this by providing a framework for understanding the properties and scope of
methods used in applications.”
Therefore, Statistics is the art of learning from data. It is concerned with
the collection of data, its subsequent description, and its analysis, which
often leads to the drawing of conclusions.
2
The Role of Probability in Statistics
• Because data used in statistical analyses often involves some
amount of "chance" or random variation, understanding probability
helps us to understand statistics and how to apply it.
• Elements of probability allow us to quantify the strength or
“confidence” in our conclusions.
• Major component that supplements statistical methods and help
gauge the strength of the statistical inference.
• The discipline of probability provides the transition between
descriptive statistics and inferential methods.
Probability vs Inferential Statistics
For a statistical problem, the
sample along with inferential
statistics allows us to draw
conclusions about the population,
with inferential statistics
making clear use of elements of
probability.
Problems in probability allow us
to draw conclusions about
characteristics of hypothetical
data taken from the population
based on known features of the
population.
4
Sample vs Population
• In statistics, we are interested in obtaining information about a
total collection of elements, which we will refer to as the
population.
• The population is often too large for us to examine each of its
members.
• For instance, we might have all the residents of a given state, or all
the television sets produced in the last year by a particular
manufacturer etc.
• In such cases, we try to learn about the population by choosing and
then examining a subgroup of its elements.
• This subgroup of a population is called a sample.
In Summary;
• Information is gathered in the form of samples, or collections of
observations.
• Samples are collected from populations that are collections of all
individuals or individual items of a particular type. 5
Descriptive Statistics Vs Inferential Statistics
• Once the data have been collected, we can organize and
summarize it in such a manner as to arrive at their orderly
presentation and conclusion.
• This procedure can be called Descriptive Statistics, which
are either numerical or graphical summaries and
representations of data.
• Inferential statistics are techniques that allow us to use
samples to make generalizations about the populations from
which the samples were drawn. It is, therefore, important
that the sample accurately represents the population.
• In a nutshell, descriptive statistics focus on describing the
visible characteristics of a dataset (a population or sample).
Whereas inferential statistics focus on making predictions
or generalizations about a larger dataset, based on a 6sample
of those data.
Descriptive Statistics Vs Inferential
Statistics
Descriptive:
• Organizing and Inferential:
summarizing data using • Use the sample data to
numbers and graphs. make an inference or
• Data summary: draw a conclusion of the
Bargraph, Histograms, population.
pie charts etc.
• Measure of central • Use probability to
Tendency: determine how confident
Mean, median and we can be that the
mode. conclusions we make are
• Measures of Variability: correct. (Confidence
Range, Variance and intervals and margins of
Standard deviation. Errors).
A Typical Dataset
Official golf balls must conform to a number of rules. In particular they
are expected to all have similar aerodynamical properties.
To test these properties 100 balls from the brand “HIT-ME” are “hit” by
a machine, and the distance reached by each ball is measured (in meters)
261.3 ; 258.1 ; 254.2 ; 257.7 ; 237.9 ; 255.8 ; 241.4 ; 256.8 ; 255.3 ; 255 ; 259.4 ; 270.5 ; 270.7 ; 272.6 ;
272.2 ; 251.1 ; 232.1 ; 273 ; 266.6 ; 273.2 ; 265.7 ; 264.5 ; 245.5 ; 280.3 ; 248.3 ;267 ; 271.5 ; 240.8 ;
268.9 ; 263.5 ; 262.2 ; 244.8 ; 279.6 ; 272.7 ; 278.7 ; 255.8 ; 276.1 ; 274.2 ; 267.4 ; 244.5 ; 252 ; 264 ;
247.7 ; 273.6 ; 264.5 ; 285.3 ; 277.8 ; 261.4 ; 253.6 ; 278.5 ; 260 ; 271.2 ; 254.8 ; 256.1 ; 264.5 ; 255.4 ;
259.5 ; 274.9 ; 272.1 ; 273.3 ; 279.3 ; 279.8 ; 272.8 ; 268.5 ; 283.7 ; 263.2 ; 257.5 ; 233.7 ; 260.2 ; 263.7 ;
244.3 ; 241.2 ; 254.4 ; 274 ; 260.7 ; 260.6 ; 255.1 ; 233.7 ; 253.7 ; 250.2 ; 251.4 ; 270.6 ; 273.4 ; 242.9 ;
276.6 ; 237.8 ; 261 ; 236 ; 251.8 ; 280.3 ; 268.3 ; 266.8 ; 254.5 ; 234.3 ; 251.6 ; 226.8 ; 240.5 ; 252.1 ;
245.6 ; 270.5
Our goal is to make “meaningful” statements about all the
“HIT-ME” balls, but using only this sample…
7
A Typical Dataset
Sample
(a small number of golf balls
randomly taken from the
Population population)
(all “HIT-ME” golf
balls) Our hope is that the sample is
somewhat representative of the
entire population of golf balls…
Before trying to do this, let’s see if we can “understand” the
data a bit better, and summarize it in nice ways… 8
Numerical Summaries – Sample Mean
Definition: Sample Mean/Sample Average
For the golf ball dataset we can easily compute this (with a
computer) and see that
Clearly this is good information to have, but it would be good to
know also if the distances are similar for all balls, or vary
wildly… 9
Data Summary and Display
Population Mean
For a finite population with N measurements, the mean is
The sample mean is a reasonable estimate of the
population mean.
Sample Variance/Standard Deviation
Definition: Sample Variance/Standard Deviation
Notice the
In our example
units are
squared !!!
The sample standard deviation is given by
An intuitive interpretation of what the sample standard
deviation represents is not so easy, but we can still understand
why it does measure variability: 10
Sample Variance/Standard Deviation
always non-negative
Properties: Sample Variance/Standard Deviation
The last expression makes computations on paper easier (but it
11
is often numerically not good)
Data Summary and Display
Population Variance
When the population is finite and consists of N values,
we may define the population variance as
The sample variance is a reasonable estimate of the
population variance.
Toyish Example
Four measurements were made of the inside diameter of
forged pistons rings used in a scooter engine. The data, in
millimeters is
12
The Sample Range
Another way to assess variability:
Definition: Sample Range
In our example
This is easy to calculate, but ignores a lot of information in the
sample !!!
13
Other Numerical Summaries
There are many other numerical summaries that are important
(we’ll encounter them later, in the context of graphical
representations of data)
Definition: Order Statistics
If one is in a setting where the order of the data is not
important then all the information one might need is contained
in the order statistics ! 15
Sample Median
Definition: Sample Median
Order the values of a data set of size n from smallest to largest. If n
is odd, the sample median is the value in position (n+1)/2; if n is even, it
is the average of the values in positions n/2 and (n/2)+1.
This is essentially the value the “splits” the dataset in two:
approximately half of the data is below the median and half is above
the median.
Graphical Representations
Over the years it has been found that tables and graphs are particularly useful
ways of presenting data, often revealing important features such as the range,
the degree of concentration, and the symmetry of the data.
1. Frequency Tables and Graphs
A data set having a relatively small number of distinct values can be
conveniently presented in a frequency table.
For Example:
LINE GRAPH, BAR GRAPH & FREQUENCY
POLYGON
• Data from a frequency table can
be graphically represented by a Line graph for starting salaries
line graph. data:
This plots the distinct data
values on the horizontal axis and
indicates their frequencies by the
heights of vertical lines.
• When the lines in a line graph are
given added thickness, the graph is
called a bar graph.
• Another type of graph used to
represent a frequency table is the
frequency polygon.
This plots the frequencies of
the different data values on the
vertical axis, and then connects the
plotted points with straight lines.
LINE GRAPH, BAR GRAPH & FREQUENCY
POLYGON
Bar graph for starting salaries Frequency polygon for starting
data salaries data
Relative Frequency Tables and Graphs
• Consider a data set consisting of n values. If f is the
frequency of a particular value, then the ratio f/n is
called its relative frequency.
• That is, the relative frequency of a data value is the
proportion of the data that have that value.
• The relative frequencies can be represented graphically
by a relative frequency line or bar graph or by a relative
frequency polygon.
• But if the data are not numerical in nature, then a pie
chart is often used to indicate relative frequencies.
Grouped Data and Histograms
• Using a line or a bar graph to plot the frequencies of data values
is often an effective way of portraying a data set.
• However, for some data sets the number of distinct values is
too large to utilize this approach.
• Instead, in such cases, it is useful to divide the values into
groupings, or class intervals, and then plot the number of data
values falling in each class interval.
• The number of class intervals chosen should be a trade-off
between:
(1). Choosing too few classes at a cost of losing too much
information about the actual data values in a class and
(2) choosing too many classes, which will result in the
frequencies of each class being too small for a pattern to be
discernible.
• The appropriate number of class intervals is a subjective choice.
• The endpoints of a class interval are called the class boundaries.
Grouped Data and Histograms
Example:
Create a Frequency Histogram using 6 classes for the data given below;
76, 84, 76, 103, 92, 47, 98, 54, 80, 91, 69, 86, 83, 75, 93, 89, 96, 65, 94 and
85.
Histogram
A bar graph plot of class data, with the
bars placed adjacent to each other, is
called a histogram.
Frequency Distributions
You can also describe the histogram numerically
15
Frequency
5
0
230 240 250 260 270 280 290
Important: you always lose information when binning, but it will
often make it much easier to qualitatively understand the
data… Keep this in mind… 21
NORMAL DATA SETS
• Many of the large data sets observed in practice have
histograms that are similar in shape.
• These histograms often reach their peaks at the sample
median and then decrease on both sides of this point in a bell-
shaped symmetric fashion.
• Such data sets are said to be normal and their histograms are
called normal histograms.
• Any data set that is not approximately symmetric about its
sample median is said to be skewed.
• It is “skewed to the right” if it has a long tail to the right and
“skewed to the left” if it has a long tail to the left.
• It follows from the symmetry of the normal histogram that a
data set that is approximately normal will have its sample mean
and sample median approximately equal.
SKEWED DATA SETS
Box and Whisker Plots
• A box plot is often used to plot some of the summarizing
statistics of a data set.
• A straight line segment stretching from the smallest to
the largest data value is drawn on a horizontal axis;
imposed on the line is a “box,” which starts at the first
and continues to the third quartile, with the value of the
second quartile indicated by a vertical line.
• The length of the line segment on the box plot, equal to
the largest minus the smallest data value, is called the
range of the data. Also, the length of the box itself,
equal to the third quartile minus the first quartile, is
called the interquartile range.
• Boxplots are especially helpful in visually comparing two
or more groups
Box and Whisker Plots
Suppose that in the golf ball dataset there were two elements
that were quite extreme, three extra balls with distance
220m, 215m, 310m.
First Quartile Sample Median Third Quartile
Outliers Outliers
220 240 260 280 300
Whisker extends to the Whisker extends to the
smallest point within 1.5 InterQuartile Range largest point within 1.5
IQR (IQR) IQR
These plots are easy to understand, and are therefore quite
useful, we can even compare different datasets easily… 23
Boxplots
Min Q1 Median Q3 Max
Lower Upper
Quartile Quartile
Measures of Spread
How much do values typically
vary from the center?
Range - is the difference of the maximum
and minimum value - spread of the entire
data set
Interquartile Range (IQR) - is the
difference of the upper quartile (Q3)
& the lower quartile (Q1) –
spread of the middle 50% of the data
Box-and-Whisker Plot Practice
Make a box-and-whisker plot using the following set of data:
1, 6, 10, 4, 2, 8, 15, 6, 3
Step #1: Put the numbers in order from least to greatest:
1, 2, 3, 4, 6, 6, 8, 10, 15
Step #2: Create a number line…that extends beyond the upper/lower extremes.
Step #3: Find the Median…draw a line above the median on the number line.
Step #4: Find the Upper Quartile…draw a line above the upper quartile.
Step #5: Find the Lower Quartile…draw a line above the lower quartile.
Step #6: Create a box (a rectangle) above the number line with the upper and lower
quartiles as two of the sides.
Step #7: Create whiskers (straight lines) that extend from the box to the upper and lower
extremes.
Box-and-Whisker Plot Practice
There are 12 people are in a Habanera Pepper eating contest who ate the
following amount of peppers:
4, 4, 4, 9, 15, 2, 5, 0, 10, 12, 1, 18
Put the numbers in order from least to greatest:
0, 1, 2, 4, 4, 4, 5, 9, 10, 12, 15, 18
BELOW THE MEDIAN ABOVE THE MEDIAN
(4.5) (4.5)
Exercise (Box plot)
Find the minimum, maximum, median (Q2), lower quartile (Q1),
and upper quartile (Q3) for the following sets of data and draw
the Box and Whisker plot (or Boxplot)
1) 32, 40, 35, 29, 14, 32
2) 6, 1, 7, 6, 5, 5, 0, 1, 0, 8, 4
3) 121, 143, 98, 144, 165, 118
Multiple Box Plots
Example inspired by the Figure 6.15 of Montgomery:
Data of the manufacturing quality index of three companies:
Comparative boxplots of Quality Index
C3
C2
C1
70 80 90 100 110
Which company should you prefer?
24
Stem and Leaf plot
• An efficient way of organizing a small- to moderate-sized data
set is to utilize a stem and leaf plot.
• It provides a partial sorting of the data and allows you to
detect the distributional pattern of the data.
• Such a plot is obtained by first dividing each data value into
two parts — its stem and its leaf.
• For instance, if the data are all two-digit numbers, then we
could let the stem part of a data value be its tens digit and let
the leaf be its ones digit.
• Thus, for instance, the value 62 is expressed as
Stem Leaf
6 2
• If numbers have more then 2 digits then the leaf always
represents the last digit in the number(ones digit) and stem
represents the other digits.
Stem and Leaf plot
Example 1: A group of students measured their pulse rates in beats per
minute. The results were.
66 69 62 58 74 56 67 72 61 62 59 60 72 58 63
Draw a Stem and Leaf diagram to show this.
We need a title
Pulse rates (beats per minute)
We need the stems 5 8 66 89 88 9
6 6 09 12 27 21 32 60 73 9
We need the Leaves 7 4 22 22 4
We need to reorder the leaves
We need to write in n n = 15
We need the key 6 / 2 means 62 beats per minute
Stem and Leaf plot
Example 2: A group of students were asked how much pocket money they
get. The results were.
£6.60 £6.90 £6.20 £5.80 £7.40 £5.60 £6.70 £7.20
£6.10 £6.20 £5.90 £6.00 £7.20 £5.80 £6.30
Draw a Stem and Leaf diagram to show this
Pocket money
5 6 8 8 9
Note this gives us the same diagram
as the previous example. This shows 6 0 1 2 2 3 6 7 9
the importance of the title and key.
7 2 2 4
n = 15 Key: 6 / 2 means £6.20
Time Series Plots
• A time series or time sequence is a data set in which the
observations are recorded in the order in which they occur.
• A time series plot is a graph in which the vertical axis
denotes the observed value of the variable (say x) and the
horizontal axis denotes the time (which could be minutes,
days, years, etc.).
• When measurements are plotted as a time series, we
often see
•trends,
•cycles, or
•other broad features of the data
Time Series Plots
PAIRED DATA SETS AND THE SAMPLE
CORRELATION COEFFICIENT
• We are often concerned with data sets that consist of pairs
of values that have some relationship to each other.
• If each element in such a data set has an x value and a y value,
then we represent the ith data point by the pair (xi, yi).
• A useful way of portraying a data set of paired values is to
plot the data on a two dimensional graph, with the x-axis
representing the x value of the data and the y-axis
representing the y value. Such a plot is called a scatter
diagram/scatterplot.
• Therefore, when one wishes to show a relationship between
two continuous variables, a scatterplot can be employed.
PAIRED DATA SETS AND THE SAMPLE
CORRELATION COEFFICIENT
Scatterplot:
Figure below shows a scatterplot of birthweight by maternal age.
Sample correlation coefficient.
The sample correlation coefficient (r) is a measure of the closeness of
association of the points in a scatter plot to a linear regression line
based on those points.
Interpreting Scatter Diagrams
• The relationship between two variables is called a
Correlation.
• A line of best-fit is a line which helps us to identify the
type of correlation (positive, negative, no correlation) &
make predictions.
• The line of best fit is drawn so that the points are
evenly distributed on either side of the line.
• The closer the dots to the line, the stronger the
correlation.
Remember:
1. The line of best fit is a STRAIGHT LINE.
2. It DOES NOT have to pass through the origin.
3. It DOES NOT have to go through each point.
Types of Correlation
Correlation describes a relationship between two data sets. A graph
may show the correlation between data. The correlation can help
you analyze trends and make predictions.
Correlation can be strong or weak
Strong Positive Correlation Weak Positive Correlation
All the points lie close to the The points are well spread out
line of best fit. from the line of best fit but still
follow the trend.
End of Module 6: Introduction to statistics
Next!! Statistical Inference and regression