Basic Statistical Methods
Basic Statistical Methods
Statistical Methods
Dr Arunangshu Mukhopadhyay
Professor, NIT Jalandhar
1
Two areas of statistics:
Descriptive Statistics: Collection, presentation, and description of sample
data.
Inferential Statistics: Making decisions and drawing conclusions about
populations. 2
3
Descriptive Statistics
4
5
6
For continuous functional data, mostly we deal with mean and standard
deviation. Median is a point above and below of which 50% of the cases
fall.
49%
49%
7
‘Mode’ by graphical method 8
Inferential Statistics
Can your experiment make a statement about the general population?
Two types
1. Parametric
◼ Interval or ratio measurements
◼ Continuous variables
◼ Usually assumes that data is normally distributed
2. Non- Parametric
◼ Ordinal or nominal measurements
◼ Discreet variables
◼ Makes no assumption about how data is distributed
9
10
Different dimensions of Quality
in case of textiles
➢ Functional Performance
➢ Durability
➢ Psychological aspect
➢ Conformance to standards
11
Confidence Interval Tolerance Interval
12
Confidence Interval to Estimate when is Known
13
Mediocre incoming quality due to multiple vendors
14
Quality
Quality of
Conformance
Quality of Quality of
Design Performance
20
21
22
Quality Management
Quality:
◼ In manufacturing, a measure of excellence or a state of being free from
defects, deficiencies and significant variations.
◼ It is brought about by strict and consistent commitment to certain
standards that achieve uniformity of a product in order to satisfy specific
customer or user requirements.
◼ ISO 8402-1986 standard defines quality as, "the totality of features and
characteristics of a product or service that bears its ability to satisfy
stated or implied needs."
◼ If an automobile company finds a defect in one of their cars and makes a
product recall, customer reliability and therefore production will decrease
because trust will be lost in the car's quality. 23
Universal Process for Managing Quality
Spider chart for gap analysis
Organizational
Performance
culture of
B measurement
empowerment
E
N
C
Strategic H Training for
commitment M technical skills
A
R
K
Motivation through
I Resource
reward and N commitment
recognition G
Strategic and
Benchmarking operations
planning
30
31
32
Quality Concepts
33
34
35
Customer
Stakeholders Product Design
Financial
Information Process Design
Enterprise Database
Customer support
And service Supplier
Sales and
Delivery Manufacturing
Employees
39
40
41
42
43
44
Kaoru Ishikawa’s Basic Seven QC Tools
1. Flow chart
2. Check Sheet
3. Cause-Effect Diagram
4. Pareto Chart
5. Control Chart
6. Histogram
7. Scatter Diagram
45
Seven Quality Control Tools
1. Flowchart
Flow Chart of Spinning Process
46
Seven Quality Control Tools
2. Check sheet
47
Seven Quality Control Tools
Check sheet
◼ A check sheet is nothing but a form used to collect data in such a way
that it makes not only the collection of data easy, but also the analysis of
that data automatic.
◼ Each mark in the check sheet indicates a defect.
◼ The type of defects, number of defects, and their distribution can be seen
at a glance, which makes analysis of data very quick and easy.
◼ Check sheets provide a logical display of data that are manually derived
and yield results from which conclusions can be easily drawn.
48
Seven Quality Control Tools
Check sheet
49
Seven Quality Control Tools
3. Cause and Effect (Fish Bone) Diagram
50
Seven Quality Control Tools
‘Cause’, which ultimately lead to create an adverse ‘Effect”. Effect is the
quality problem. CauseEffect analysis is a tool for analyzing and illustrating
a process by showing the main causes and subcauses leading to an effect.
Sub-branches are known as twiglets.
51
Seven Quality Control Tools
Another example in textile for Cause Effect Diagram
52
Seven Quality Control Tools
4. Pareto Chart
53
Seven Quality Control Tools
5. Control Charts
54
Control Charts
Sl. No. Type of data X R P nP C
5. 1Control Charts
Fibre properties such as strength, length and fineness ◘
2 The strength of yarn, cord, and fabric ◘
3 The amount of variation in strength within the breaks or number from a ◘
bobbin, cone, or spool of yarn
4 Yarn, rove and sliver linear density, lap weight ◘
5 Variation of linear density of yarn within a bobbin ◘
6 Percent bad quality ◘
7 Number of bad pieces or units ◘
8 Imperfections in yarn, number of particular defects in fabric ◘
56
Control Charts
Mean
58
Control Charts
Upper Specification
limit
Control Specification
limit limit
Lower
Specification limit
59
Control Charts
Type of Control Chart
Variables Control Chart Attributes Control Charts
“Variable” means measuring “Attributes” counts the number of defective
characteristics such as time, weight, items in a sample or the number of defects
volume, length, pressure drop, associated with a particular type of item
concentration, etc
n is number of sample
61
R chart
To draw R-chart, k different samples, each of size n are selected and the
points Ri (range of the ith sample) are plotted and according to these points
the decision is made.
FOR KNOWN mean µ and standard deviation σ FOR UNKNOWN mean µ and standard deviation
σ
In this case σ is replaced by 𝑅ത /d2
65
Control Charts
66
Seven Quality Control Tools
6. Histogram
67
Bar chart
68
Seven Quality Control Tools
7. Scatter Diagrams
Regression coefficient for CVm
Predicted CVm (%)
r = 0.9
r = 0.01
r = -0.9
70
An interpretation of the size of the coefficient has been described by
Cohen (1992) as:
Correlation coefficient value Relationship
73
74
75
Design of Experiments (Not part of 7 Quality Control Tools)
76
Road Map to Design of Experiment
Customer Requirement
Control Chart
Process Capability
Design of Experiment
77
Impact of innovation and continuous improvement
Temp. 1550F
◼ Kaizen
◼ Benchmarking
◼ QS 9000
◼ ISO 14000
Impact at the bottom line
◼ 6σ Process Control
83
Total Quality Management (TQM)
84
85
Customer needs & Differences between Customer
expectations Expectations and satisfaction
Customer
Company vision
and mission
Management
commitment
Process People
Organizational
culture
The current ISO 9001, ISO 9002 and ISO 9003 will be consolidated into the
single revised ISO 9001 standard. The revised ISO 9004 is based on eight
quality management principles: customer focus, leadership, involvement of
people, process approach, system approach to management, continual
improvement, factual approach to decision making, and mutually beneficial
supplier relationships.
90
91
92
What about Lean, TOC, TQM
◼ Six Sigma
• Remove defects, minimize variance
◼ Lean
• Remove waste, shorten the flow
◼ TOC
• Remove and manage constraints
◼ TQM
• Continuous Improvement
93
6σ Process Control
◼ Six Sigma aims, at producing not more than 3.4 defects per million of
parts produced in a manufacturing process. It uses a variety of statistics
to determine the best practices for any given process.
◼ Statisticians and Six Sigma consultants study the existing processes and
determine the methods that produce the best overall results Six Sigma
statistically ensures that 99.9997% of all products produced in a process
are of acceptable quality, it allows only 3.4 defects per million
opportunities.
◼ If a given process fails to meet this criterion it is re-analyzed, altered and
tested to find out if there are any improvements. If no improvement is
found, the process is reanalyzed, altered and tested again. This cycle is
repeated until an Improvement becomes visible. 94
Excellent
Poor Process Process
Capability Capability
Very High Very High Very Low Very Low
Probability Probability Probability Probability
of Defects of Defects of Defects of Defects
95
96
97
The Motorola Six Sigma concept; (a) Centered at target, (b) Mean shifted by 1.5σ
98
99
Data Driven Decisions
Y= f (X)
To get results, should we focus our behavior on the Y or X ?
• Y • X1 . . . XN
• Dependent • Independent
• Output • Input-Process
• Effect • Cause
• Symptom • Problem
• Monitor • Control
• Response • Factor
101
102
Basic Implementation
Roadmap
Identify Customer Requirements
Control
-Sustain Improvement
-Drive Towards Perfection
103
Next Project Define
Customers, Value, Problem Statement Validate
Scope, Timeline, Team Project $
Celebrate Primary/Secondary & OpEx Metrics
Project $ Current Value Stream Map
Measure
Voice Of Customer (QFD)
Control Assess specification / Demand
Document process (WIs, Std Work) Measurement Capability (Gage R&R)
Mistake proof, TT sheet, CI List Correct the measurement system
Analyze change in metrics Process map, Spaghetti, Time obs.
Value Stream Review Measure OVs & IVs / Queues
Prepare final report
Validate
Project $
Validate
Project $
104
Six Sigma Team • Own vision, direction,
integration, results
Executive Leadership • Lead change
• Project owner
• Implement solutions
• Part-time
• Black Belt managers
• Help Black Belts
Project Champions
Green Belts
Black Belts
Master Black
• Full time Belts
• Train and coach • Devote 50% - 100% of time to Black Belt activities
Black and Green Belts • Facilitate and practice problem solving
• Statistical problem solving experts • Train and coach Green Belts and project teams
105
6σ Process Control
◼ Once an improvement is found, it is documented and the knowledge is
spread across other units in the company so they can implement this new
process and reduce their defects per million opportunities.
◼ Six Sigma methodology improves any existing business process by
constantly reviewing and re-tuning the process. To achieve this, Six
Sigma uses a methodology known as DMAIC (Define opportunities,
Measure performance, Analyze opportunity, Improve performance,
Control performance).
◼ Even though Six Sigma was initially implemented at Motorola to
improve the manufacturing process, all types of businesses can profit
from implementing Six Sigma.
106
Six Sigma Methodology (DMAIC)
Define
Measure
Control
Analyse
Improve
107
DMAIC Steps
1. Define
112
113
114
115
116
117
6σ Process Control
118
6σ Process Control
Quality Function Deployment (QFD)
◼ With QFD, Six Sigma teams can more effectively focus on the activities
that mean the most to the customer, beat the competition, and align with
the mission of the organization.
119
6σ Process Control
Cause & Effect Matrix
◼ The C&E Matrix helps Six Sigma project leaders facilitate team
decision-making.
◼ The C&E Matrix is a fool that helps Six Sigma teams select, prioritize,
and analyze the data they collect over the course of a project to identify
problems in that process. Six Sigma teams typically use the C&E Matrix
in the Measure phase of the DMAIC methodology.
120
6σ Process Control
Failure Mode and Effect Analysis (FMEA)
◼ FMEA helps Six Sigma teams to identify and address weaknesses in a
product or process before they occur. Before implementing new
products, processes or services.
◼ Six Sigma teams use FMEA to identify ways. An effective FMEA
identifies corrective actions required to prevent failures from reaching
the customer and will improve performance, quality, and reliability.
121
6σ Process Control
t-Test
◼ The t-test helps Six Sigma teams validate test results using small sample
sizes. The t-test is used to determine the statistical difference between
two groups, not just a difference due to random chance.
◼ Six Sigma teams might use it to determine if a plan for a comparative
analysis of patient blood pressures, before and after they receive a drug,
is likely to provide reliable results.
122
6σ Process Control
Control Charts
◼ Six Sigma teams use Control Charts to assess process stability. Control
Charts are a simple but highly effective tool for monitoring and
improving process performance over time because they help Six Sigma
teams to observe and analyze venation.
◼ The three basic components of any control chart are a center-line, upper
and lower statistically determined control limits, and performance data
plotted over time.
123
6σ Process Control
Design of Experiment (DOE)
◼ DOE helps Six Sigma Black Belts make the most of valuable resources.
◼ DOE is a statistical technique that encompasses the planning, design,
data collection, analysis and interpretation strategy used by Six Sigma
professionals! Six Sigma teams use DOE to determine the relationship
between factors (X) affecting a process and the output of that process
(Y).
124
Understanding Data
125
Statistical
Data Level Meaningful Operations
Methods
127
Data Type
Ordinal Data
Ordinal values represent discrete and ordered units. It is therefore nearly the
same as nominal data, except that it’s ordering matters. The main limitation
of ordinal data, the differences between the values is not really known.
Because of that, ordinal scales are usually used to measure non-numeric
features like happiness, customer satisfaction and so on. You can see an
example below:
128
Data Type
Interval Data
Interval values represent ordered units that have the same difference.
Therefore we speak of interval data when we have a variable that contains
numeric values that are ordered and where we know the exact differences
between the values. An example would be a feature that contains
temperature of a given place like you can see below:
129
Data Type
Ratio Data
Ratio values are also ordered units that have the same difference. Ratio
values are the same as interval values, with the difference that they do
have an absolute zero. Good examples are height, weight, length etc.
130
Type of statistical tools which can be used
Nominal Ordinal Interval Ratio
Frequency distribution. Yes Yes Yes Yes
Nominal
Qualitative
Ordinal
Variables
Discrete
Quantitative
Continuous
133
Simple Statistics
134
Statistics
Operations
research
Population Acceptance Quality/process Design of
prediction/ sampling control experiment
acceptance
F - Test
Chi-square
Z - Test
ANOVA
Statistical
Quality
Control
Inductive
Statistics
Staistical
Design of
Process
Experiment
Control
Acceptance
Sampling
139
Inductive
statistics
Parametric Nonparametric
statistics statistics
Estimation Hypothesis
testing
Point Interval
estimation estimation
Frequency
Characteristic
Quantitative Qualitative
(Through objective (Through subjective
assessment) assessment) 141
Road map for Simple Statistics
OBJECTIVE ASSESSMENT SUBJECTIVE ASSESSMENT
(Parametric) (Nonparametric)
Testing of Testing of
means variance
142
Chart on organizing a raw data set
Raw data set
146
Frequency Polygon
◼ When there are huge number of observations and
◼ The histogram is constructed using reduced class intervals
◼ Then, the frequency polygon tends to be less angular and more smooth
◼ This is known as ‘frequency curve’.
◼ The most known type of frequency curve is the ‘normal curve’.
147
Cumulative Frequency Curve
◼ It is a graph of a cumulative distribution, with data values on the
horizontal plane axis and either the cumulative relative frequencies, the
cumulative frequencies on the vertical axis.
◼ Cumulative frequency is defined as the sum of all the previous
frequencies up to the current point. To find the popularity of the given
data or the likelihood of the data that fall within the certain frequency
range, this curve helps in finding those details accurately.
◼ Create this by plotting the point corresponding to the cumulative
frequency of each class interval. Most of the Statisticians use this curve,
to illustrate the data in the pictorial representation. It helps in estimating
the number of observations which are less than or equal to the particular
value. 148
Cumulative Frequency Curve
149
Box Plot
◼ A box plot (also known as box and whisker plot) is a type of chart often
used in explanatory data analysis to visually show the distribution of
numerical data and skewness through displaying the data quartiles (or
percentiles) and averages.
◼ Box plots show the five-number summary of a set of data: including the
minimum score, first (lower) quartile, median, third (upper) quartile, and
maximum score.
150
Box Plot
Minimum Score; The lowest score, excluding outliers (shown at the end of
the left whisker).
Lower Quartile; Twenty-five percent of scores fall below the lower
quartile value (also known as the first quartile).
Median; The median marks the mid-point of the data and is shown by the
line that divides the box into two parts (sometimes known as the second
quartile). Half the scores are greater than or equal to this value and half are
less.
Upper Quartile; Seventy-five percent of the scores fall below the upper
quartile value (also known as the third quartile). Thus, 25% of data are
above this value.
151
Box Plot
Maximum Score; The highest score, excluding outliers (shown at the end
of the right whisker).
Whiskers; The upper and lower whiskers represent scores outside the
middle 50% (i.e. the lower 25% of scores and the upper 25% of scores).
The Interquartile Range (or IQR); This is the box plot showing the
middle 50% of scores (i.e., the range between the 25th and 75th percentile).
Note: Box plots divide the data into sections that each contain
approximately 25% of the data in that set.
152
Box Plot
This box plot, comparing four
machines for energy output,
shows that machine has a
significant effect on energy with
respect to both location and
variation. Machine 3 has the
highest energy response (about
72.5); machine 4 has the least
variable energy response with
about 50% of its readings being
within 1 energy unit.
Box plots are an excellent tool for conveying location and variation information in
data sets, particularly for detecting and illustrating location and variation changes
between different groups of data. 153
Box Plot
The box plot shape will show if a statistical data set is normally distributed
or skewed.
◼ When the median is in the middle of the box, and the whiskers are about
the same on both sides of the box, then the distribution is symmetric.
◼ When the median is closer to the bottom of the box, and if the whisker is
shorter on the lower end of the box, then the distribution is positively
skewed (skewed right).
◼ When the median is closer to the top of the box, and if the whisker is
shorter on the upper end of the box, then the distribution is negatively
skewed (skewed left).
154
Pictorial representation of results
155
Graph with error bar
156
Methods of Studying Variation
1. Range
2. Quartile deviation
3. Mean deviation
4. Standard deviation
5. Lorenz curve
157
158
Methods of Studying Variation
Range
◼ Range is the simplest measure of studying dispersion. It is the difference
between the largest and smallest value in the distribution.
◼ It is given by the formula-
◼ Range = L – S Where, L = Largest value S = Smallest value
159
Methods of Studying Variation
Quartile Deviation
◼ A quartile is a measure that divides the data into four quarters. The first
quartile, denoted by Q1 lies in the middle of the first half of the data set. It
covers the first 25 percent of the data set. The second quartile, denoted by
Q2, divides the data such that 50 percent of the data lies below it and 50
percent of the data lies above it.
◼ This is called as the median. The third quartile, denoted by Q3, lies in the
middle of the second half of the data set. 75 percent of the data would lie
below the third quartile and 25 percent of the data would be greater than
the third quartile.
160
Methods of Studying Variation
◼ The interquartile range is a measure of absolute dispersion. It is calculated based on
the lower quartile and the upper quartile, that is, the first quartile and the third quartile
respectively. The interquartile range is the difference between the third quartile and
the first quartile.
◼ Interquartile range = Q3 – Q1
161
Methods of Studying Variation
Mean or Average Deviation:
◼ The mean deviation, also known as the average deviation, is the average
difference between the values in the distribution and the mean or the
median.
◼ This method shows the average scatteredness of the values in the
distribution around the mean or the median.
◼ This means that the mean deviation can be calculated in the following
two ways:
1. Mean Deviation (M.D.) about the mean value, and
2. Mean Deviation (M.D.) about the median value.
162
Methods of Studying Variation
Standard Deviation:
◼ Standard deviation is the square root of the average of the squared
deviations from the mean. It measures the absolute deviation of the
values from the mean. Greater the value of standard deviation, greater is
the deviation of the values from the mean.
𝑋−µ 2
◼ Standard deviatio𝑛 = σ = σ
𝑛
where, µ = Population mean,
x = any particular observation in data
n = no. of observation in data
164
165
Methods of Studying Variation
Coefficient of Variation:
◼ Coefficient of Variation (C.V.) is measured by the ratio of the standard
deviation to the mean. While the standard deviation is an absolute
measure, the coefficient of variation is a relative measure. It is useful
in comparing the variability between two sets of data.
σ
◼ Coefficient of Variation (C.V.) = × 100
x
where, x = sample mean,
& σ = standard deviation
166
Methods of Studying Variation
Lorenz Curve:
◼ Lorenz curve is a graphical method of studying the dispersion of data,
named after Dr. Max O. Lorenz, who developed it in 1905.
◼ In order to construct a Lorenz curve, the items as well as the frequencies
are cumulated and the total is considered as 100 percentages.
◼ Then percentages are calculated for the cumulated values. These
percentages are plotted on a graph paper.
◼ If there is equal distribution of frequencies, the points would lie on a
straight line. This line is known as the line of equal distribution or the
line of equality.
167
Methods of Studying Variation
◼ However, if the distribution is unequal, the curve would be away from the
line of equality.
◼ The farther the curve from the line of equal distribution, the higher is the
inequality or dispersion. Given below is the Lorenz curve depicting
income distribution among households.
168
Thank you
169