100% found this document useful (1 vote)
18 views169 pages

Basic Statistical Methods

The document discusses statistical methods in quality management, distinguishing between descriptive and inferential statistics. It outlines various quality control tools, such as control charts and check sheets, and emphasizes the importance of Total Quality Management (TQM) and ISO 9000 standards in ensuring product quality. Additionally, it highlights the role of benchmarking and continuous improvement in enhancing organizational performance.

Uploaded by

Prem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
100% found this document useful (1 vote)
18 views169 pages

Basic Statistical Methods

The document discusses statistical methods in quality management, distinguishing between descriptive and inferential statistics. It outlines various quality control tools, such as control charts and check sheets, and emphasizes the importance of Total Quality Management (TQM) and ISO 9000 standards in ensuring product quality. Additionally, it highlights the role of benchmarking and continuous improvement in enhancing organizational performance.

Uploaded by

Prem
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

B.

Statistical Methods

Dr Arunangshu Mukhopadhyay
Professor, NIT Jalandhar

1
Two areas of statistics:
Descriptive Statistics: Collection, presentation, and description of sample
data.
Inferential Statistics: Making decisions and drawing conclusions about
populations. 2
3
Descriptive Statistics

4
5
6
For continuous functional data, mostly we deal with mean and standard
deviation. Median is a point above and below of which 50% of the cases
fall.

49%

49%

7
‘Mode’ by graphical method 8
Inferential Statistics
Can your experiment make a statement about the general population?

Two types
1. Parametric
◼ Interval or ratio measurements
◼ Continuous variables
◼ Usually assumes that data is normally distributed
2. Non- Parametric
◼ Ordinal or nominal measurements
◼ Discreet variables
◼ Makes no assumption about how data is distributed

9
10
Different dimensions of Quality
in case of textiles
➢ Functional Performance

➢ Durability

➢ Aesthetics (visual appearance of the product)

➢ Psychological aspect

➢ Conformance to standards
11
Confidence Interval Tolerance Interval

Concept of confidence interval vs. tolerance interval

12
Confidence Interval to Estimate  when  is Known

13
Mediocre incoming quality due to multiple vendors
14
Quality

Quality of
Conformance

Quality of Quality of
Design Performance

The three aspects of quality


15
16
17
18
19
Responsibility of Quality

20
21
22
Quality Management
Quality:
◼ In manufacturing, a measure of excellence or a state of being free from
defects, deficiencies and significant variations.
◼ It is brought about by strict and consistent commitment to certain
standards that achieve uniformity of a product in order to satisfy specific
customer or user requirements.
◼ ISO 8402-1986 standard defines quality as, "the totality of features and
characteristics of a product or service that bears its ability to satisfy
stated or implied needs."
◼ If an automobile company finds a defect in one of their cars and makes a
product recall, customer reliability and therefore production will decrease
because trust will be lost in the car's quality. 23
Universal Process for Managing Quality
Spider chart for gap analysis
Organizational
Performance
culture of
B measurement
empowerment
E
N
C
Strategic H Training for
commitment M technical skills
A
R
K
Motivation through
I Resource
reward and N commitment
recognition G

Role of benchmarking in implementing best practices


Change
management
Current profile

Time- based Gap


competition Benchmarking analysis
Competitive
profile
Technological
development

Strategic and
Benchmarking operations
planning

Influences on benchmarking and its outcomes


Quality Function Deployment

◼ Quality function deployment (QFD) is a method developed in Japan


beginning in 1966 to help transform the voice of the customer into
engineering characteristics for a product.
◼ Yoji Akao, the original developer, described QFD as a "method to
transform qualitative user demands into quantitative parameters, to
deploy the functions forming quality, and to deploy methods for achieving
the design quality into subsystems and component parts, and ultimately to
specific elements of the manufacturing process“.
◼ The author combined his work in quality assurance and quality control
points with function deployment used in value engineering.
28
29
Quality Costs

◼ Prevention Costs; Prevention costs are costs incurred to ensure that


defects are minimized and prevented at the earliest...
◼ Appraisal Costs; Appraisal costs are costs incurred to identify defective
products before they are shipped off. These...
◼ Internal Failure Costs; Internal failure costs refer to costs incurred on
the defective units before they are identified...
◼ External Failure Costs; External failure costs are cost associated with
defective units which are shipped to customers.

30
31
32
Quality Concepts

Taguchi Quality loss function

33
34
35
Customer
Stakeholders Product Design

Financial
Information Process Design

Enterprise Database
Customer support
And service Supplier

Sales and
Delivery Manufacturing
Employees

Enterprise wide information needs


36
37
38
Quality Management

39
40
41
42
43
44
Kaoru Ishikawa’s Basic Seven QC Tools

1. Flow chart
2. Check Sheet
3. Cause-Effect Diagram
4. Pareto Chart
5. Control Chart
6. Histogram
7. Scatter Diagram

45
Seven Quality Control Tools
1. Flowchart
Flow Chart of Spinning Process

46
Seven Quality Control Tools
2. Check sheet

47
Seven Quality Control Tools
Check sheet
◼ A check sheet is nothing but a form used to collect data in such a way
that it makes not only the collection of data easy, but also the analysis of
that data automatic.
◼ Each mark in the check sheet indicates a defect.
◼ The type of defects, number of defects, and their distribution can be seen
at a glance, which makes analysis of data very quick and easy.
◼ Check sheets provide a logical display of data that are manually derived
and yield results from which conclusions can be easily drawn.

48
Seven Quality Control Tools
Check sheet

49
Seven Quality Control Tools
3. Cause and Effect (Fish Bone) Diagram

50
Seven Quality Control Tools
‘Cause’, which ultimately lead to create an adverse ‘Effect”. Effect is the
quality problem. CauseEffect analysis is a tool for analyzing and illustrating
a process by showing the main causes and subcauses leading to an effect.
Sub-branches are known as twiglets.

51
Seven Quality Control Tools
Another example in textile for Cause Effect Diagram

52
Seven Quality Control Tools
4. Pareto Chart

53
Seven Quality Control Tools
5. Control Charts

54
Control Charts
Sl. No. Type of data X R P nP C
5. 1Control Charts
Fibre properties such as strength, length and fineness ◘
2 The strength of yarn, cord, and fabric ◘
3 The amount of variation in strength within the breaks or number from a ◘
bobbin, cone, or spool of yarn
4 Yarn, rove and sliver linear density, lap weight ◘
5 Variation of linear density of yarn within a bobbin ◘
6 Percent bad quality ◘
7 Number of bad pieces or units ◘
8 Imperfections in yarn, number of particular defects in fabric ◘

9 Number of bad units with a reasonably constant sample size ◘


10 Number of bad units with a sample size that varies considerably from lot to ◘
lot, day to day, or week to week
55
Control Charts
◼ Is it possible to make a fabric same as before 1 year?
◼ If we practically see this question we find that it is not possible because
there can be some variation related to men or machine or variation due to
variation in raw material( cotton fibre).
◼ If we closely observe we can understand that variation is a day to day
process. But these variation affect quality and hence we make a chart
where we decide to keep a control of variation in process called Control
chart.
◼ The central idea in Control chart is that, if the production process is under
control, then the values of the statistic will lie in between “mean ± 3σ.”

56
Control Charts

UCL (mean + 3σ)

Mean

LCL (mean - 3σ)

A typical control chart


57
Control Charts
Specification Limit
Sometimes, the customer provides the control limits for the production
process, such control limits are called the upper specification limit (USL)
and lower specification limit (LSL) and the decision regarding the control
of the production process is taken by taking care of these specification
limits.

58
Control Charts

Upper Specification
limit

Control Specification
limit limit

Lower
Specification limit

59
Control Charts
Type of Control Chart
Variables Control Chart Attributes Control Charts
“Variable” means measuring “Attributes” counts the number of defective
characteristics such as time, weight, items in a sample or the number of defects
volume, length, pressure drop, associated with a particular type of item
concentration, etc

1. nP chart (control chart for number of


1. X bar chart (control chart for
mean). defectives).
2. R chart (control chart for 2. P chart (control chart for proportion
range). of defectives).
3. C chart (control chart for number of
defects). 60
ഥ chart
𝑿
For process having mean µ and standard deviation σ
FOR KNOWN mean µ and standard FOR UNKNOWN mean µ and standard
deviation σ deviation σ

n is number of sample

61
R chart
To draw R-chart, k different samples, each of size n are selected and the
points Ri (range of the ith sample) are plotted and according to these points
the decision is made.
FOR KNOWN mean µ and standard deviation σ FOR UNKNOWN mean µ and standard deviation
σ
In this case σ is replaced by 𝑅ത /d2

d2, D1 and D2 are tabulated in the statistical table


for the different values of n.

Ri is the range of the ith sample and k denotes total


samples of size n selected for study D3 and D4 are also tabulated in the statistical 62table
np chart
Variable X represents number of defective articles/products observed in the
sample of size n. Here X will be binomially distributed random variable with
parameters “n” and “p,” where p is the probability of getting defective
article produced by the process. NOTE THAT k is different samples, each
of size n. “p” is unknown.
In this case p is replaced by 𝑃ത
“p” is known.

Where , q = 1 − p di be the number of defective articles/products


observed in the ith sample.
63
p chart
Variable X represents number of defective articles/products observed in the
sample of size n. Here X will be binomially distributed random variable with
parameters “n” and “p,” where p is the probability of getting defective
article produced by the process. k different samples, each of size n
‘p’ is known. ‘p’ is unknown.

di = (number of defective articles)/(products observed in the ith sample)


pi = di /n = proportion of (defective articles)/(products produced by the process). These points pi are
plotted on the p chart and according to these plotted points the decision is made 64
C Chart
To draw c-chart, k different sample units are selected from the production
process and are inspected. Let ci be the number of defects observed in the ith
sample unit produced by the process. These points cj are plotted on the c-
chart and according to these plotted points the decision is made.
Variable X = C represents number of defects observed per article/product.
Here X will follow Poisson probability distribution with parameter “λ” where
λ is the average number of defects per article produced by the process.
“λ” is known. “λ” is unknown. λ is replaced by 𝐶ҧ

65
Control Charts

66
Seven Quality Control Tools
6. Histogram

67
Bar chart

68
Seven Quality Control Tools
7. Scatter Diagrams
Regression coefficient for CVm
Predicted CVm (%)

Experimental CVm (%) 69


Correlation Coefficient r
 Measures strength of a relationship between two continuous
variables -1 ≤ r ≤ 1

r = 0.9

r = 0.01

r = -0.9

70
An interpretation of the size of the coefficient has been described by
Cohen (1992) as:
Correlation coefficient value Relationship

-0.3 to +0.3 Weak


-0.5 to -0.3 or 0.3 to 0.5 Moderate
-0.9 to -0.5 or 0.5 to 0.9 Strong
-1.0 to -0.9 or 0.9 to 1.0 Very strong
Revi
ewer
:
Jean
Russ
ell
Univ
ersit
y of
Shef
field
Relationship Correlation Interpretation
Average IQ and chocolate 0.27 Weak positive relationship. More
consumption chocolate per capita = higher average IQ
Road fatalities and Nobel 0.55 Strong positive. More accidents = more
winners prizes!
Gross Domestic Product and 0.7 Strong positive. Wealthy countries =
Nobel winners more prizes
Mean temperature and Nobel -0.6 Strong negative. Colder countries = Revi
ewer
winners more prizes. :
Jean
Russ
ell
Univ
ersit
y of
Shef
field
Confounding

73
74
75
Design of Experiments (Not part of 7 Quality Control Tools)

76
Road Map to Design of Experiment
Customer Requirement

Process Flow Diagram

Check Sheet, Histogram, Pareto Analysis, Fishbone Diagram and


Scatter Plot

Control Chart

Process Capability

Design of Experiment
77
Impact of innovation and continuous improvement

Reviewer: Jean Russell


University of Sheffield
Experiment based on one factor at a time

Temp. 1550F

Yield versus reaction time with temperature constant at 155º F.


79
Time 1.7 hrs

Yield versus temperature with reaction time constant at 1.7 hours 80


81
82
Other Approaches
◼ TQM

◼ Kaizen

◼ Benchmarking

◼ QS 9000

◼ ISO 14000
Impact at the bottom line
◼ 6σ Process Control
83
Total Quality Management (TQM)

◼ Total Quality management is defined as a continuous effort by the


management as well as employees of a particular organization to
ensure long term customer loyalty and customer satisfaction.

◼ TQM has three principle components-customer focus, continuous


improvement, and top management commitment. Thus TQM has
more to do with the culture and mindset of an organization than
anything else.

84
85
Customer needs & Differences between Customer
expectations Expectations and satisfaction

Customer

Company vision
and mission

Management
commitment
Process People

Organizational
culture

Self- directed Process analysis Open channels


Integration of
Cross- functional and continuous Empowerment of
vendors
teams improvement communication

Features of a TQM model


87
ISO 9000

◼ The ISO 9000 family of quality management systems (QMS) is a set of


standards that helps organizations ensure they meet customers and other
stakeholder needs within statutory and regulatory requirements related to
a product or service.
◼ ISO 9000 deals with the fundamentals of quality management systems,
including the seven quality management principles that underlie the
family of standards.
◼ ISO 9001 deals with the requirements that organizations wishing to meet
the standard must fulfil.
Some description of ISO 9000 family is given in next slide.
88
◼ ISO 9000 Quality Management and Quantity Assurance
Standards - Guidelines for Selection and Use
◼ ISO 9001 Quality systems - Model for Quality Assurance
in Design, Development, Production, Installation and
Servicing
◼ ISO 9002 Quality systems - Model for Quality Assurance
in Production, Installation, and Servicing
◼ ISO 9003 Quality systems - Model for Quality Assurance
in Final Inspection and Test
◼ ISO 9004 Quality Management and Quality System
Elements-Guidelines
89
Revision to ISO 9000 Series Standards
ISO 9000 series standards are revised every five years. The revised ISO 9000
standards expected to be out in the year 2000 will be:
◼ ISO 9000: Quality management systems - Concepts and vocabulary
◼ ISO 900l: Quality management systems - Requirements
◼ ISO 9004: Quality management systems - Guidelines
◼ ISO 10011: Guidelines for auditing quality systems

The current ISO 9001, ISO 9002 and ISO 9003 will be consolidated into the
single revised ISO 9001 standard. The revised ISO 9004 is based on eight
quality management principles: customer focus, leadership, involvement of
people, process approach, system approach to management, continual
improvement, factual approach to decision making, and mutually beneficial
supplier relationships.
90
91
92
What about Lean, TOC, TQM
◼ Six Sigma
• Remove defects, minimize variance
◼ Lean
• Remove waste, shorten the flow
◼ TOC
• Remove and manage constraints
◼ TQM
• Continuous Improvement
93
6σ Process Control
◼ Six Sigma aims, at producing not more than 3.4 defects per million of
parts produced in a manufacturing process. It uses a variety of statistics
to determine the best practices for any given process.
◼ Statisticians and Six Sigma consultants study the existing processes and
determine the methods that produce the best overall results Six Sigma
statistically ensures that 99.9997% of all products produced in a process
are of acceptable quality, it allows only 3.4 defects per million
opportunities.
◼ If a given process fails to meet this criterion it is re-analyzed, altered and
tested to find out if there are any improvements. If no improvement is
found, the process is reanalyzed, altered and tested again. This cycle is
repeated until an Improvement becomes visible. 94
Excellent
Poor Process Process
Capability Capability
Very High Very High Very Low Very Low
Probability Probability Probability Probability
of Defects of Defects of Defects of Defects

LSL USL LSL USL

95
96
97
The Motorola Six Sigma concept; (a) Centered at target, (b) Mean shifted by 1.5σ
98
99
Data Driven Decisions
Y= f (X)
To get results, should we focus our behavior on the Y or X ?

• Y • X1 . . . XN
• Dependent • Independent
• Output • Input-Process
• Effect • Cause
• Symptom • Problem
• Monitor • Control
• Response • Factor

Why should we test or inspect Y, if we know this relationship?


100
Sigma and % accuracy

Defects per Million % Accuracy


Opportunities (DPMO)
One Sigma 691,500 30.85%
Two Sigma 308,500 69.15%
Three Sigma 66,810 93.32%
Four Sigma 6,210 99.38%
Five Sigma 233 99.977%
Six Sigma 3.4 99.9997%
Seven Sigma 0.020 99.999998%

101
102
Basic Implementation
Roadmap
Identify Customer Requirements

Understand and Define


Entire Value Streams
Vision (Strategic Business Plan)
Deploy Key Business Objectives
- Measure and target (metrics)
- Align and involve all employees
- Develop and motivate
Continuous Improvement (DMAIC)
Define, Measure, Analyze, Improve
Identify root causes, prioritize, eliminate waste,
make things flow and pulled by customers

Control
-Sustain Improvement
-Drive Towards Perfection

103
Next Project Define
Customers, Value, Problem Statement Validate
Scope, Timeline, Team Project $
Celebrate Primary/Secondary & OpEx Metrics
Project $ Current Value Stream Map
Measure
Voice Of Customer (QFD)
Control Assess specification / Demand
Document process (WIs, Std Work) Measurement Capability (Gage R&R)
Mistake proof, TT sheet, CI List Correct the measurement system
Analyze change in metrics Process map, Spaghetti, Time obs.
Value Stream Review Measure OVs & IVs / Queues
Prepare final report

Validate
Project $
Validate
Project $

Improve Analyze (and fix the obvious)


Optimize KPOVs & test the KPIVs Root Cause (Pareto, C&E, brainstorm)
Redesign process, set pacemaker Find all KPOVs & KPIVs
Validate
5S, Cell design, MRS FMEA, DOE, critical Xs, VA/NVA
Project $
Visual controls Graphical Analysis, ANOVA
Value Stream Plan Future Value Stream Map

104
Six Sigma Team • Own vision, direction,
integration, results
Executive Leadership • Lead change

• Project owner
• Implement solutions
• Part-time
• Black Belt managers
• Help Black Belts

Project Champions
Green Belts

Black Belts
Master Black
• Full time Belts
• Train and coach • Devote 50% - 100% of time to Black Belt activities
Black and Green Belts • Facilitate and practice problem solving
• Statistical problem solving experts • Train and coach Green Belts and project teams

105
6σ Process Control
◼ Once an improvement is found, it is documented and the knowledge is
spread across other units in the company so they can implement this new
process and reduce their defects per million opportunities.
◼ Six Sigma methodology improves any existing business process by
constantly reviewing and re-tuning the process. To achieve this, Six
Sigma uses a methodology known as DMAIC (Define opportunities,
Measure performance, Analyze opportunity, Improve performance,
Control performance).
◼ Even though Six Sigma was initially implemented at Motorola to
improve the manufacturing process, all types of businesses can profit
from implementing Six Sigma.
106
Six Sigma Methodology (DMAIC)
Define

Measure

Control

Analyse
Improve

107
DMAIC Steps
1. Define

1. Define 2. Measure 3. Analyze 4. Improve 5. Control

◼ Identify projects that are measurable


◼ Define projects including the demands
of the customer and the content of the
internal process.
◼ Develop team charter
◼ Define process map 108
DMAIC Steps
2. Measure
5.0
Control

1. Define 2. Measure 3. Analyze 4. Improve 5. Control

◼ Define performance standards


◼ Measure current level of quality into
Sigma. It precisely pinpoints the area
causing problems.
◼ Identify all potential causes for such
problems. 109
DMAIC Steps
3. Analyse

1. Define 2. Measure 3. Analyse 4. Improve 5. Control

◼ Establish process capability


◼ Define performance objectives
◼ Identify variation sources
Tools for analysis
3.0⚫ Process Mapping
Analyze

⚫ Failure Mode & Effect Analysis


⚫ Statistical Tests
⚫ Design of Experiments
⚫ Control charts
⚫ Quality Function
110
Deployment (QFD)
DMAIC Steps
4. Improve

1. Define 2. Measure 3. Analyse 4. Improve 5. Control

◼ Screen potential causes


◼ Discover variable relationships among causes and effects
◼ Establish operating tolerances
◼ Pursue a method to resolve and ultimately eliminate problems. It is also a
phase to explore the solution how to change, fix and modify the process.
◼ Carryout a trial run for a planned period of time to ensure the revisions and
improvements implemented in the process result in achieving the targeted
values. 111
DMAIC Steps
5. Control

1. Define 2. Measure 3. Analyse 4. Improve 5. Control

◼ Monitor the improved process continuously to ensure long term sustainability


of the new developments.
◼ Share the lessons learnt
◼ Document the results and accomplishments of all the improvement activities
for future reference.

112
113
114
115
116
117
6σ Process Control

Quality Control Tools Used in Six Sigma:


◼ Quality Function Deployment (QFD)
◼ Cause & Effect Matrix
◼ Failure Mode and Effect Analysis (FMEA)
◼ t-Test
◼ Control Charts
◼ Design of Experiment (DOE)

118
6σ Process Control
Quality Function Deployment (QFD)

◼ QFD helps Six Sigma Black Belts drive customer-focused development


across the design process. QFD is a system. It consists of a set of
procedures to identify, communicate, and prioritize customer
requirements.

◼ With QFD, Six Sigma teams can more effectively focus on the activities
that mean the most to the customer, beat the competition, and align with
the mission of the organization.
119
6σ Process Control
Cause & Effect Matrix
◼ The C&E Matrix helps Six Sigma project leaders facilitate team
decision-making.
◼ The C&E Matrix is a fool that helps Six Sigma teams select, prioritize,
and analyze the data they collect over the course of a project to identify
problems in that process. Six Sigma teams typically use the C&E Matrix
in the Measure phase of the DMAIC methodology.

120
6σ Process Control
Failure Mode and Effect Analysis (FMEA)
◼ FMEA helps Six Sigma teams to identify and address weaknesses in a
product or process before they occur. Before implementing new
products, processes or services.
◼ Six Sigma teams use FMEA to identify ways. An effective FMEA
identifies corrective actions required to prevent failures from reaching
the customer and will improve performance, quality, and reliability.

121
6σ Process Control
t-Test
◼ The t-test helps Six Sigma teams validate test results using small sample
sizes. The t-test is used to determine the statistical difference between
two groups, not just a difference due to random chance.
◼ Six Sigma teams might use it to determine if a plan for a comparative
analysis of patient blood pressures, before and after they receive a drug,
is likely to provide reliable results.

122
6σ Process Control
Control Charts
◼ Six Sigma teams use Control Charts to assess process stability. Control
Charts are a simple but highly effective tool for monitoring and
improving process performance over time because they help Six Sigma
teams to observe and analyze venation.
◼ The three basic components of any control chart are a center-line, upper
and lower statistically determined control limits, and performance data
plotted over time.

123
6σ Process Control
Design of Experiment (DOE)
◼ DOE helps Six Sigma Black Belts make the most of valuable resources.
◼ DOE is a statistical technique that encompasses the planning, design,
data collection, analysis and interpretation strategy used by Six Sigma
professionals! Six Sigma teams use DOE to determine the relationship
between factors (X) affecting a process and the output of that process
(Y).

124
Understanding Data

125
Statistical
Data Level Meaningful Operations
Methods

Nominal Classifying and Counting Nonparametric

Ordinal All of the above plus Ranking Nonparametric

Interval All of the above plus Addition, Parametric


Subtraction, Multiplication, and Division
Ratio Parametric
All of the above
126
Data Type
Nominal Data
Nominal values represent discrete units and are used to label variables, that
have no quantitative value. Just think of them as “labels”. Note that nominal
data that has no order. Therefore if you would change the order of its
values, the meaning would not change. You can see two examples of
nominal features below:

127
Data Type
Ordinal Data
Ordinal values represent discrete and ordered units. It is therefore nearly the
same as nominal data, except that it’s ordering matters. The main limitation
of ordinal data, the differences between the values is not really known.
Because of that, ordinal scales are usually used to measure non-numeric
features like happiness, customer satisfaction and so on. You can see an
example below:

128
Data Type
Interval Data
Interval values represent ordered units that have the same difference.
Therefore we speak of interval data when we have a variable that contains
numeric values that are ordered and where we know the exact differences
between the values. An example would be a feature that contains
temperature of a given place like you can see below:

129
Data Type
Ratio Data
Ratio values are also ordered units that have the same difference. Ratio
values are the same as interval values, with the difference that they do
have an absolute zero. Good examples are height, weight, length etc.

130
Type of statistical tools which can be used
Nominal Ordinal Interval Ratio
Frequency distribution. Yes Yes Yes Yes

Median and percentiles. No Yes Yes Yes


Add or subtract. No No Yes Yes
Mean, standard deviation,
standard error of the No No Yes Yes
mean.
Ratio, or coefficient of
No No No Yes
variation.
131
Descriptive statistics
Measures of shape of distribution Graphs
◼ Frequency distribution Interval, Ratio level
◼ Relative frequency of occurrence • The histogram: frequency histogram
→ proportion of values & relative frequency histogram
Nominal, Ordinal level • Frequency polygon: midpoint of class
interval
◼ Bar chart
• Pareto chart: bar chart with
◼ Pie chart descending sorted frequency
• Cumulative frequency
• Cumulative relative frequency →
OGIVE graph (Ojiv or Oh’-jive 132
graph)
Qualitative and quantitative variables may be further subdivided:

Nominal
Qualitative
Ordinal
Variables
Discrete
Quantitative
Continuous

133
Simple Statistics

134
Statistics

Descriptive Inductive Probability


statistics statistics theory

Operations
research
Population Acceptance Quality/process Design of
prediction/ sampling control experiment
acceptance

Road map of different statistics 135


Inferential Statistics
Non-parametric Statistics Parametric Statistics
Subjective Ranks t - Test

F - Test

Chi-square

Z - Test

ANOVA

◼ Parametric statistics are the category of values which follow a


particular rule
◼ Non-parametric statistics values cant be generalized in set of rules.
136
◼ Our discussion is limited to parametric statistics
Summarizing a raw
data set on a
quantitative variable

Study of central tendency of Study the variability (or


data set dispersion) of the data set

Quantifying the disparity


Locating the central
among the data entries
value of the data set
Range
Mode
Inter-quartile range
Median
Variance
Mean
Standard deviation
137
Chart on summarizing a raw data set on a quantitative variable
138
Inductive statistics is the branch of statistics that deals with
generalizations, predictions, estimations, and decisions about a population
from data sampled from that population.

Statistical
Quality
Control

Population Large amount of data is used in


Probability
prediction &
Theory Operational Research
acceptance

Inductive
Statistics
Staistical
Design of
Process
Experiment
Control

Acceptance
Sampling
139
Inductive
statistics

Parametric Nonparametric
statistics statistics

Estimation Hypothesis
testing

Point Interval
estimation estimation

Road map of inductive statistics 140


Statistical Quality Control
Statistical data represented by this data and frequency represents the
characteristic.
Statistical Unit

Frequency

Characteristic

Quantitative Qualitative
(Through objective (Through subjective
assessment) assessment) 141
Road map for Simple Statistics
OBJECTIVE ASSESSMENT SUBJECTIVE ASSESSMENT
(Parametric) (Nonparametric)

Continuous Discrete Rank Multiple


functional data functional data correlation ranking

Testing of Testing of
means variance

142
Chart on organizing a raw data set
Raw data set

Ungrouped data Grouped data

Categorical or qualitative Quantitative data set


data set

One may remove the


Make a frequency table where Obtain the modified range possible outliers and
frequencies of counts for different and then divide it into several then proceed with the
categories are found classes or subintervals modified range

Make a frequency table


where frequencies or counts
for different classes are
recorded 143
Presentation of Data
Frequency Polygon
◼ A frequency polygon is almost identical to a histogram, which is used to
compare sets of data or to display a cumulative frequency distribution. It
uses a line graph to represent quantitative data.
◼ Statistics deals with the collection of data and information for a particular
purpose. The tabulation of each run for each ball in cricket gives the
statistics of the game. Tables, graphs, pie-charts, bar graphs, histograms,
polygons etc. are used to represent statistical data pictorially.
◼ Frequency polygons are a visually substantial method of representing
quantitative data and its frequencies. Let us discuss how to represent a
frequency polygon.
144
Frequency Polygon
To draw frequency polygons, first we need to draw histogram and then
follow the below steps:
◼ Step 1- Choose the class interval and mark the values on the horizontal
axes
◼ Step 2- Mark the mid value of each interval on the horizontal axes.
◼ Step 3- Mark the frequency of the class on the vertical axes.
◼ Step 4- Corresponding to the frequency of each class interval, mark a
point at the height in the middle of the class interval
◼ Step 5- Connect these points using the line segment.
◼ Step 6- The obtained representation is a frequency polygon.
145
Frequency Polygon
Example Frequency Polygon

146
Frequency Polygon
◼ When there are huge number of observations and
◼ The histogram is constructed using reduced class intervals
◼ Then, the frequency polygon tends to be less angular and more smooth
◼ This is known as ‘frequency curve’.
◼ The most known type of frequency curve is the ‘normal curve’.

147
Cumulative Frequency Curve
◼ It is a graph of a cumulative distribution, with data values on the
horizontal plane axis and either the cumulative relative frequencies, the
cumulative frequencies on the vertical axis.
◼ Cumulative frequency is defined as the sum of all the previous
frequencies up to the current point. To find the popularity of the given
data or the likelihood of the data that fall within the certain frequency
range, this curve helps in finding those details accurately.
◼ Create this by plotting the point corresponding to the cumulative
frequency of each class interval. Most of the Statisticians use this curve,
to illustrate the data in the pictorial representation. It helps in estimating
the number of observations which are less than or equal to the particular
value. 148
Cumulative Frequency Curve

149
Box Plot
◼ A box plot (also known as box and whisker plot) is a type of chart often
used in explanatory data analysis to visually show the distribution of
numerical data and skewness through displaying the data quartiles (or
percentiles) and averages.
◼ Box plots show the five-number summary of a set of data: including the
minimum score, first (lower) quartile, median, third (upper) quartile, and
maximum score.

150
Box Plot
Minimum Score; The lowest score, excluding outliers (shown at the end of
the left whisker).
Lower Quartile; Twenty-five percent of scores fall below the lower
quartile value (also known as the first quartile).
Median; The median marks the mid-point of the data and is shown by the
line that divides the box into two parts (sometimes known as the second
quartile). Half the scores are greater than or equal to this value and half are
less.
Upper Quartile; Seventy-five percent of the scores fall below the upper
quartile value (also known as the third quartile). Thus, 25% of data are
above this value.
151
Box Plot
Maximum Score; The highest score, excluding outliers (shown at the end
of the right whisker).
Whiskers; The upper and lower whiskers represent scores outside the
middle 50% (i.e. the lower 25% of scores and the upper 25% of scores).
The Interquartile Range (or IQR); This is the box plot showing the
middle 50% of scores (i.e., the range between the 25th and 75th percentile).
Note: Box plots divide the data into sections that each contain
approximately 25% of the data in that set.

152
Box Plot
This box plot, comparing four
machines for energy output,
shows that machine has a
significant effect on energy with
respect to both location and
variation. Machine 3 has the
highest energy response (about
72.5); machine 4 has the least
variable energy response with
about 50% of its readings being
within 1 energy unit.
Box plots are an excellent tool for conveying location and variation information in
data sets, particularly for detecting and illustrating location and variation changes
between different groups of data. 153
Box Plot
The box plot shape will show if a statistical data set is normally distributed
or skewed.
◼ When the median is in the middle of the box, and the whiskers are about
the same on both sides of the box, then the distribution is symmetric.
◼ When the median is closer to the bottom of the box, and if the whisker is
shorter on the lower end of the box, then the distribution is positively
skewed (skewed right).
◼ When the median is closer to the top of the box, and if the whisker is
shorter on the upper end of the box, then the distribution is negatively
skewed (skewed left).

154
Pictorial representation of results

155
Graph with error bar

156
Methods of Studying Variation
1. Range

2. Quartile deviation

3. Mean deviation

4. Standard deviation

5. Lorenz curve

157
158
Methods of Studying Variation
Range
◼ Range is the simplest measure of studying dispersion. It is the difference
between the largest and smallest value in the distribution.
◼ It is given by the formula-
◼ Range = L – S Where, L = Largest value S = Smallest value

159
Methods of Studying Variation
Quartile Deviation
◼ A quartile is a measure that divides the data into four quarters. The first
quartile, denoted by Q1 lies in the middle of the first half of the data set. It
covers the first 25 percent of the data set. The second quartile, denoted by
Q2, divides the data such that 50 percent of the data lies below it and 50
percent of the data lies above it.
◼ This is called as the median. The third quartile, denoted by Q3, lies in the
middle of the second half of the data set. 75 percent of the data would lie
below the third quartile and 25 percent of the data would be greater than
the third quartile.

160
Methods of Studying Variation
◼ The interquartile range is a measure of absolute dispersion. It is calculated based on
the lower quartile and the upper quartile, that is, the first quartile and the third quartile
respectively. The interquartile range is the difference between the third quartile and
the first quartile.
◼ Interquartile range = Q3 – Q1

161
Methods of Studying Variation
Mean or Average Deviation:
◼ The mean deviation, also known as the average deviation, is the average
difference between the values in the distribution and the mean or the
median.
◼ This method shows the average scatteredness of the values in the
distribution around the mean or the median.
◼ This means that the mean deviation can be calculated in the following
two ways:
1. Mean Deviation (M.D.) about the mean value, and
2. Mean Deviation (M.D.) about the median value.

162
Methods of Studying Variation
Standard Deviation:
◼ Standard deviation is the square root of the average of the squared
deviations from the mean. It measures the absolute deviation of the
values from the mean. Greater the value of standard deviation, greater is
the deviation of the values from the mean.
𝑋−µ 2
◼ Standard deviatio𝑛 = σ = σ
𝑛
where, µ = Population mean,
x = any particular observation in data
n = no. of observation in data
164
165
Methods of Studying Variation
Coefficient of Variation:
◼ Coefficient of Variation (C.V.) is measured by the ratio of the standard
deviation to the mean. While the standard deviation is an absolute
measure, the coefficient of variation is a relative measure. It is useful
in comparing the variability between two sets of data.
σ
◼ Coefficient of Variation (C.V.) = × 100
x
where, x = sample mean,
& σ = standard deviation

166
Methods of Studying Variation
Lorenz Curve:
◼ Lorenz curve is a graphical method of studying the dispersion of data,
named after Dr. Max O. Lorenz, who developed it in 1905.
◼ In order to construct a Lorenz curve, the items as well as the frequencies
are cumulated and the total is considered as 100 percentages.
◼ Then percentages are calculated for the cumulated values. These
percentages are plotted on a graph paper.
◼ If there is equal distribution of frequencies, the points would lie on a
straight line. This line is known as the line of equal distribution or the
line of equality.
167
Methods of Studying Variation
◼ However, if the distribution is unequal, the curve would be away from the
line of equality.
◼ The farther the curve from the line of equal distribution, the higher is the
inequality or dispersion. Given below is the Lorenz curve depicting
income distribution among households.

168
Thank you

169

You might also like