100% found this document useful (2 votes)
42 views25 pages

Understanding Basic Statistics Concepts

Statistics

Uploaded by

phúc nguyễn
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
100% found this document useful (2 votes)
42 views25 pages

Understanding Basic Statistics Concepts

Statistics

Uploaded by

phúc nguyễn
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

BASIC STATISTICS

Click here to see this complete module at [Link]

[Link] ©. All rights reserved. 1


Basic Statistics
There are primarily two branches in which basic statistics are studied:

Descriptive Statistics
 Applied to describe the data using numbers, charts, and graphs. Terms such as mean,
median, mode, variance, standard deviation are values that summarize data. Descriptive
statistics describe the entire group for which the numbers were obtained. These are the
actual values for the entire group.

Inferential Statistics
 Uses sample statistics to infer relationships of the population parameters. This is most
often done in the Analyze and Improve phases using hypothesis testing, correlation
analysis, regression analysis, and design of experiments (DOE).

[Link] ©. All rights reserved. 2


Basic Statistics
Inferential Statistics

It is rarely possible to analyze the entire


population (such as the average weight
of all sharks in the ocean). For this reason
a sampling strategy is applied. The
analysis of the sample is used to apply
inferences to the entire population.

The values may or may not be the same


values for the entire population so they
are applied with a confidence interval.
:

[Link] ©. All rights reserved. 3


Basic Statistics
Notation
A comparison of means, variance, or proportions is initiated with a hypothesis statement about a
population or populations.

The sample statistics are studied to determine, with a certain level of confidence and power,
that the hypothesis (null hypothesis) is to be proven false or not false but not necessarily true.
See Hypothesis Testing for more information.

[Link] ©. All rights reserved. 4


Basic Statistics
Samples & Populations

[Link] ©. All rights reserved. 5


Basic Statistics
Types of Data
Attribute (Discrete) Data (Qualitative)
 Categories
• Pass/Fail (binomial)
• Machine 1, Machine 2, Machine 3
• Political party affiliation
• Temperature (hot, cold)
• Counts
• Ratings

Variable Data (Quantitative)


 Continuous Data (decimal subdivisions are meaningful)
• Temperature (F)
• Pressure (psi)
• Time (seconds, minutes, hours)
• Speed (m/s))

Continuous data is more informative and preferred but may be more difficult to obtain
[Link] ©. All rights reserved. 6
Basic Statistics
Data Measurement Levels
Nominal - Lowest level. A numerical label that represents a qualitative
description. These numbers are labels or assignments of numbers that
represent a category or classification.

Ordinal - The next level higher of data classification than nominal data.
Numerical data where number is assigned to represent a qualitative
description similar to nominal data. These are measures by only the rank

Interval - Numerical data where the data can be arranged in an order and
the differences between the values ​are meaningful but not necessarily a
zero point. These are measures using equal intervals.

Ratio – similar to interval data EXCEPT has a defined absolute zero point
and is the highest level of data measurement. Ratio data can be both
continuous and discrete. Highest level of usage and can be analyzed in
more ways than the other three types of data.

[Link] ©. All rights reserved. 7


Basic Statistics
Variation
Every process displays variation. There are two types of variation:

Controlled Variation
 Characterized by a stable and consistent pattern of variation over time
 Common causes create this type of variation

Uncontrolled Variation
 Characterized by variation that changes over time
 Special Causes create this type of variation

One goal of Six Sigma is to control the process variation


and make sure it is within the customer specification limits.
It’s a good thing to accomplish control and reduce the
variation, but it must meet the Voice of the Customer (VOC)
(i.e. the LCL and UCL should be within the LSL and USL)

[Link] ©. All rights reserved. 8


Basic Statistics
Basic Equations

Measures of Central Tendency


 Mean
 Median
 Mode

Measures of Dispersion
 Range
 Standard Deviation
 Variance
 Coefficient of Variation

[Link] ©. All rights reserved. 9


Basic Statistics
Measures of Central Tendency
Measurements of central tendency describe the centering of the distribution. Let’s look at some examples using the data:

{1, 3, 8, 3, 7, 11, 8, 3, 9, 10}


MEAN: Since most populations exhibit normality (bell-shaped curve) or can be assumed to be normal
 The most common measure for central tendency.
 Used to describe normal data.
 Strongly influenced by extreme values and those values would make the data set non-normal where the mode and
median are most suitable measures.
 Easier said as the sum of all sample values divided by the number of samples: 63 / 10 = 6.3

[Link] ©. All rights reserved. 10


Basic Statistics
Measures of Central Tendency (continued)
Measurements of central tendency describe the centering of the distribution. Let’s look at some examples using the data:

{1, 3, 8, 3, 7, 11, 8, 3, 9, 10}


MEDIAN: The median is the midpoint, the middle value or observation of the data set. If the set of data has an even count,
the median is the average of the middle two values. This is a measure for skewed or non-normal data.

Arrange the numbers in ascending or descending order: {1, 3, 3, 3, 7, 8, 8, 9, 10, 11}

Since the sample is an even set of data (10 samples) and the middle two values are 7 and 8, then the average of the two
middle values is 7.5.

Median = 7.5

[Link] ©. All rights reserved. 11


Basic Statistics
Measures of Central Tendency (continued)
Measurements of central tendency describe the centering of the distribution. Let’s look at some examples using the data:

{1, 3, 8, 3, 7, 11, 8, 3, 9, 10}


MODE: The mode is the most commonly occurring value in the data set. Not commonly used as a measure of central
location but can be found in the tallest bar of a vertical histogram chart.

Mode = 3

since it occurs more than any other value.

In summary as quick observation, the mode, median, and mean are not close to each other. This data set is probably not
representative of a normal distribution (where the mean, median, and mode would be near the same value)

[Link] ©. All rights reserved. 12


Basic Statistics
Measures of Dispersion
The measurements of dispersion describe the width of the distribution. The results for the measures of dispersion are
calculated below for the data set shown below.

Let’s look at some examples using the same data: {1, 3, 8, 3, 7, 11, 8, 3, 9, 10}
RANGE: The range, R, of the data is the difference of the highest and smallest values.
R = 11 - 1 = 10

DEVIATION: The deviation is the difference of each value from the mean. This is used in the calculation of the standard
deviation and variance. "x" is the point of interest.

The sum of the deviations from the mean is always zero.

[Link] ©. All rights reserved. 13


Basic Statistics
Measures of Dispersion (continued)
Data: {1, 3, 8, 3, 7, 11, 8, 3, 9, 10}

STANDARD DEVIATION: The standard deviation is shown by the following formulas for a sample and population. It also
equals the square root of the variance. "x" is the point of interest. This is the most common measure for dispersion.

"N" and "n" represent the sample size, although one is capitalized it is done for notation, they both represent the same
value (the sample size). As the sample size approaches infinity, the denominators in both formulas become equal. The "n-
1" is to remove bias from very small sample sizes such as less than 30 samples.

VARIANCE: The standard deviation squared and expressed as s2 or σ2


COEFFICIENT OF VARIATION: Ratio of the standard deviation to the mean. This value is expressed as a percentage and
helps to determine comparing variability when mean is also known.
This is not the same as Covariance.

[Link] ©. All rights reserved. 14


Basic Statistics
Measures of Dispersion

From the previous data set

Standard Deviation = 3.5


Variance = 12.23
Coefficient of Variation = 3.5/6.3 = 55%

[Link] ©. All rights reserved. 15


Basic Statistics
Normal Distribution
The Normal Distribution is a distribution of data that has a “bell-shape” and has certain consistent and predictable properties.

This x-axis represents the number of standard deviations (σ) )from the mean. These properties are very useful in our
understanding of the characteristics of the underlying process from which the data were obtained.

Most natural phenomena and man-made processes are


distributed normally, or can be “assumed” normal.

A normal distribution can be described completely


y knowing only the mean and standard deviation.

A standard normal distribution has a mean = 0 and standard


deviation = 1. Mean = Median = Mode

: The area under sections of the curve can be used to estimate the cumulative probability of a certain “event” occurring

[Link] ©. All rights reserved. 16


Basic Statistics
Empirical Rule
The previous rules of cumulative probability apply even when a set of data is not perfectly
normally distributed Let’s compare the values for a theoretical (perfect) normal distributions to
empirical (real-world) distributions

Number of standard Theoretical Normal Empirical Normal


deviations
±1σ 68% 60-75%
±2σ 95% 90-98%
± 3σ 99.7% 99-100%

[Link] ©. All rights reserved. 17


Basic Statistics
Visual Tools
Visualize the data whenever possible. There are many tools available. Look at the next several depictions that all represent
the same data. Through a combination of Descriptive Stats, Histograms, Box Plot, CI Interval, Runs Chart, Probability Plot,
Dot Plot you can get a good understanding of the process. This Box Plot is a small help, the next page adds more insight.

[Link] ©. All rights reserved. 18


Basic Statistics

[Link] ©. All rights reserved. 19


Basic Statistics
SPC
Used to analyze process performance by plotting data points, control limits, and a center line. A process should be in control
before assessing process capability. Data does not necessarily have to be normal to assess process capability.

For CONTINUOUS data:


• I-MR – no subgroups
• Xbar-R – subgroups <=8. Uses the range (hence the R) to estimate the process variation
• Xbar-S – subgroups >8. Uses the standard deviation (hence the S) to estimate the process variation

For DISCRETE data


• C-Chart
• U-Chart
• P-Chart
• NP-Chart

[Link] ©. All rights reserved. 20


Basic Statistics
Common Types of Data Distributions (there are many more)

DISCRETE DISTRIBUTIONS
Binomial Distribution
Poisson Distribution
Hypergeometric Distribution

CONTINUOUS DISTRIBUTIONS:
Uniform Distribution
Normal Distribution
Exponential Distribution
t Distribution
Chi-square Distribution
F Distribution

[Link] ©. All rights reserved. 21


Basic Statistics
Hypothesis Testing
Statistical test to analyze sample data, properly handle uncertainty, minimize subjectivity, understand the risk of decision
errors about the population.

Terms:
Ho = Null Hypothesis – assumed true until evidence can prove to reject it
Ha = Alternative Hypothesis – inferred if the null hypothesis is rejected
p-value = probability value of results could occur when Ho is true

Examples
Ho : Mean Group A = Mean Group B
Ha : Mean Group A = Mean Group B

Ho : Slope of the line is +1


Ha : Slope of the line is not +1

Ho : Variance Group A = Variance Group B


Ha : Variance Group A > Variance Group B

[Link] ©. All rights reserved. 22


Basic Statistics
Hypothesis Testing Decision Matrix

[Link] ©. All rights reserved. 23


Basic Statistics
Confidence Intervals (CI)
CI is range, such as {0.42 – 2.39}, in which population parameters are likely to fall given a certain level of confidence.

 The most common choices are 90%, 95%, and 99% confidence levels for confidence intervals.
 A 95% confidence interval (alpha-risk of 0.05) suggests that approximately 95 out of 100 confidence
intervals will contain the population parameter. There is a 5.0% risk that the population parameter is not
contained within the interval.

The goal is to test whether sample statistics such as the mean, standard deviation, and proportion (x, s, p) are only
estimates of the population parameters (,  and P). Confidence intervals quantify the uncertainty since there is
variability in the estimates.
What the difference between a Confidence Level and Confidence Interval?
Confidence level refers to the percentage of probability, or certainty, that the confidence interval would contain the true
population parameter when you draw a random sample repeatedly.
Often results are shown for a “95% Confidence Interval” …. this means there is a 95% confidence level that the confidence
interval contains the population parameter and 5% risk (alpha) of the CI not containing the population parameter.

[Link] ©. All rights reserved. 24


Basic Statistics

[Link] ©. All rights reserved. 25

You might also like