Basic Statistics for
Statistical Process Control
Purpose of Basic Statistics
The purpose of Basic Statistics is to:
• Provide a numerical summary of the data being analyzed.
• Provide the foundation for assessing process capability.
• Provide a common objective language to be used throughout an organization
to describe processes.
Data:
• Factual information organized for analysis.
• Numerical or other information represented in a form
suitable for processing by computer
Relax….it won’t
be that bad!
Population vs. Sample
Population: All the items that have the “property of interest” under study.
Sample: A significantly smaller subset of the population used to study the population and
make inferences.
Population
Sample
Sample
Sample
Population Parameters: Sample Statistics:
• Arithmetic descriptions of a population • Arithmetic descriptions of a sample
• N, µ, , 2, P • n, X-bar , s, s2, p
Statistical Notation – Cheat Sheet
An individual value, an observation The Standard Deviation of sample data
Sample size
The Standard Deviation of population data
Population size The variance of sample data
Summation The variance of population data
The Mean, average of sample data The range of data
The Mean of population data
Types of Data
Numeric Variable
Attribute Data Count Data
Data
Measuring Classifying Counting
something something something
If you are
measuring a If you are
Is often binary,
characteristic counting
i.e. there are
and the data has something. Data
only two possible
decimal can only be
values (e.g.
subdivisions (e.g. discrete numbers
pass/fail,
dimensions, (e.g. number of
good/bad).
temperature, defects per unit).
voltage, …).
Numerical variable Data
Numeric variable data could represent for example:
• The diameter of a metal cylinder (millimeters)
• Oven temperature (degrees)
• Wire resistance (Ohm)
Statistical Summary:
• Numeric variable data is usually summarized by statistics such as
mean, standard deviation and range.
Statistical Model:
• Numeric variable data often follow (under certain conditions) the
Normal Distribution.
Attribute Data
Attribute data results from classifying things. It could represent:
• Conformity of production parts
• Customer satisfaction
• On time delivery
Statistics:
• Attribute data is usually summarized as a proportion or a
percentage.
Statistical Model:
• The number of defective parts (for samples of constant size)
follow (under certain conditions) the Binomial Distribution.
3 defective
parts in a
sample of 12
parts
Count Data
Attribute data results from counting things. It could represent:
• Leaks in a fuel tank
• Errors on an invoice
• Scratches on a sheet of metal
Statistics:
• Count data is usually summarized by calculating the average
‘Defects per Unit’ (DPU).
Statistical Model:
• The number of defects per unit (for samples of constant size)
follow (under certain conditions) the Poisson Distribution.
Descriptive Statistics
Descriptive Statistics
Measures of Location (central tendency)
• Mean
• Median
• Mode
Measures of Variation (dispersion)
• Range
• Standard Deviation
• Variance
Descriptive Statistics
• Open the MINITAB worksheet: “Basic_Statistics.MTW”
• Calculate descriptive statistics:
Stat > Statistiques élémentaires > Afficher les statistiques descriptives…
Measures of Location
Mean is:
• Commonly referred to as the average.
• The arithmetic balance point of a distribution of data.
Stat > Statistiques élémentaires > Afficher les statistiques descriptives… > Statistiques…
Sample Population
Measures of Location
Median is:
• The mid-point, or 50th percentile, of a distribution of data.
• Arrange the data from low to high, or high to low:
o It is the single middle value in the ordered list if there is an odd number of observations
o It is the average of the two middle values in the ordered list if there are an even number
of observations
Stat > Statistiques élémentaires > Afficher les statistiques descriptives… > Statistiques…
Measures of Location
Mode is:
• The most frequently occurring value in a distribution of data.
• Relevant only for discrete data (count data)
Stat > Statistiques élémentaires > Afficher les statistiques descriptives… > Statistiques…
0
Measures of Variation
Range is:
• The Difference between the largest observation and the smallest observation
in the data set.
• A small range would indicate a small amount of variability and a large range a
large amount of variability.
Stat > Statistiques élémentaires > Afficher les statistiques descriptives… > Statistiques…
Measures of Variation
Standard Deviation is:
• Equivalent of the average deviation of values from the Mean for a distribution
of data.
Sample Population
Minitab provides
sample standard
deviation
Measures of Variation
Variance is the:
• Average squared deviation of each individual data point from the Mean.
Sample Population
The Normal Distribution
Normal Distribution
• The Normal Distribution is the most recognized distribution in
statistics.
• What are the characteristics of a Normal Distribution?
o Only random error is present
o Process free of assignable cause
o Process free of drifts and shifts
The Normal Curve
• The normal curve is a smooth, symmetrical, bell-shaped curve,
generated by the density function.
• It is the most useful continuous probability model as many naturally
and industrially occurring measurements are approximately
Normally Distributed.
Normal Distribution
• The area under the curve between any 2 points represents the
proportion of the distribution between those points.
The area between the
Mean and any other point
depends upon the Standard
Deviation.
m x
Normal Distribution
• Each combination of Mean and Standard Deviation generates a unique
Normal curve:
• The “Standard” Normal Distribution
o Has a μ = 0, and σ = 1
o Data from any Normal Distribution can be made to fit the standard Normal
by converting raw scores to standard scores.
o Z-scores measure how many Standard Deviations from the mean a
particular data-value lies.
Normal Distribution
-6 -5 -4 -3 -2 -1 +1 +2 +3 +4 +5 +6
68.27 % of the data will fall within +/- 1 standard deviation
95.45 % of the data will fall within +/- 2 standard deviations
99.73 % of the data will fall within +/- 3 standard deviations
99.9937 % of the data will fall within +/- 4 standard deviations
99.999943 % of the data will fall within +/- 5 standard deviations
99.9999998 % of the data will fall within +/- 6 standard deviations
Process Sigma Performance
LSL Target USL
Process Performance Process Performance
1 sigma 4 sigma
USL – LSL = 2. USL – LSL = 8.
Cp = 0,33 Cp = 1,33
2 sigma 5 sigma
USL – LSL = 4. USL – LSL = 10.
Cp = 0,66 Cp = 1.66
3 sigma 6 sigma
USL – LSL = 6. USL – LSL = 12.
Cp = 1 Cp = 2
Normal Distribution
Centered Process 1.5 σ shifted
Non-
Conforming Conforming Non-Conforming
conforming
1 sigma (Cp=0,33) 68,27 % 31.73 % 30.9 % 69.1 %
2 sigma (Cp=0,66) 95.45 % 4.55 % 69.2 % 30.8 %
0.27 % 6.68 %
3 sigma (Cp=1) 99.73 % 93.32 %
(2 700 ppm) (66 800 ppm)
0.006 4 % 0.62 %
4 sigma (Cp=1,33) 99.993 6 % 99.38 %
(64 ppm) (6 200 ppm)
0.000 058 % 0.023 %
5 sigma (Cp=1,67) 99.999 942 % 99.977 %
(0.58 ppm) (230 ppm)
0.000 000 2 % 0.000 34 %
6 sigma (Cp=2) 99.999 999 8 % 99.999 66 %
(0.002 ppm) (3.4 ppm)
Normality Assessment
Why Assess Normality?
• While many processes in nature behave according to the normal distribution,
many processes in business, particularly in the areas of service and transactions,
do not.
• There are many types of distributions:
• There are many statistical tools that assume Normal Distribution properties in
their calculations.
• If we know that the basic structure of the data should follow a Normal
Distribution, but plots from our data shows otherwise, we know the data contain
Special Causes.
Special Causes = Opportunity
Normality Assessment: Anderson-Darling Test
The Anderson-Darling test uses an empirical density function
100
Expected for Normal Distribution
Actual Data
Departure of the actual 20%
80
data from the expected C
u
Normal Distribution. m
u
l
The Anderson-Darling a 60
t
Goodness-of-Fit test i
v
e
assesses the magnitude P
e 40
of these departures r
c
using an Observed e
n
t
minus Expected formula. 20
20%
0
3.0 3.5 4.0 4.5 5.0 5.5
Raw Data Scale
Anderson-Darling Test
Open the MINITAB worksheet: “[Link]”
Stat > Statistiques élémentaires > Test de normalité…
If the P-value is more than 0.05
your data are Normal enough
for most purposes
Graphical Summary
Stat > Statistiques élémentaires > Récapitulatif graphique…
The Anderson-Darling test also appears in this
output. Again, if the P-value is greater than 0.05,
assume the data are Normal.
Rapport récapitulatif pour Dist A
Test de normalité d'Anderson-Darling
A au carré 0.18
Valeur de P 0.921
Moyenne 50.031
EcTyp 4.951
Variance 24.511
Asymétrie -0.061788
Aplatissement -0.180064
N 500
Minimum 35.727
1er quartile 46.800
Médiane 50.006
3e quartile 53.218
Maximum 62.823
Intervalle de confiance = 95 % pour la moyenne
36 40 44 48 52 56 60 49.596 50.466
Intervalle de confiance = 95 % pour la médiane
49.663 50.500
Intervalle de confiance = 95 % pour l'écart type
4.662 5.278
Intervalles de confiance = 95 %
Moyenne
Médiane
49.50 49.75 50.00 50.25 50.50
Exercise
Exercise: Normality Assessment
Exercise:
1. Generate Normal Probability Plots and the graphical summary
using the “[Link]” file.
2. Use only the columns Dist B and Dist C.
3. Answer the following quiz questions based on your analysis of
this data set:
• Is Distribution B Normal?
• Is Distribution C Normal?
English-French Terminology
Basic Statistics: English-French Terminology
English French
Statistical Process Control Maitrise statistique des procédés
Population Population
Sample Échantillon
Sample size Taille de l'échantillon / Effectif de l'échantillon
Frame Sous-population
Qualitative data / Attribute data Données qualitatives / Données d'attribut
Continuous data / Variable data Données continues / Données variables
Count data / Discrete data Données de comptage / données discrètes
Measures of central tendency Mesures de tendance centrale
Measures of dispersion Mesures de dispersion
Mean Moyenne
Median Médiane
Mode Mode
Basic Statistics: English-French Terminology
English French
Range Étendue
Standard deviation Écart type
Variance Variance
Normal Distribution Loi normale
Standard normal distribution Loi normale centrée réduite
Common causes Causes communes
Special causes Causes spéciales
Normality test Test de normalité