0% found this document useful (0 votes)
4 views36 pages

Chapter Three 1

Chapter 3 discusses statistical evaluation of analytical data, focusing on measures of central tendency such as mean, median, and range, as well as standard deviation and variance. It also covers errors affecting accuracy, categorizing them into determinate and indeterminate errors, and explains how to evaluate these errors through confidence intervals and significance tests. The chapter concludes with methods for comparing sample means and variances using statistical tests.

Uploaded by

adiyammezgebu16
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views36 pages

Chapter Three 1

Chapter 3 discusses statistical evaluation of analytical data, focusing on measures of central tendency such as mean, median, and range, as well as standard deviation and variance. It also covers errors affecting accuracy, categorizing them into determinate and indeterminate errors, and explains how to evaluate these errors through confidence intervals and significance tests. The chapter concludes with methods for comparing sample means and variances using statistical tests.

Uploaded by

adiyammezgebu16
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 3

Statistical evaluation of analytical data


Mean
The mean, 𝑋,ത is measure of central tendency and it is the numerical average
for a data set. The mean calculated by dividing the sum of the individual
values by the size of the data set.

where Xi is the ith measurement, and n is the number of independent


measurements
Median
The median is the middle value when we order our data from the
smallest to the largest value. When the data has an odd number of
values, the median is the middle value. For an even number of
values, the median is the average of the n/2 and the (n/2) + 1 values,
where n is the size of the data set.
When n = 5, the median is the third value in the ordered data set; for
n = 6, the median is the average of the third and fourth members of
the ordered data set.
Range
The range, w, is the difference between a data set’s largest and
smallest values.
w =Xlargest−Xsmallest
Standard Deviation
The standard deviation, s, describes the spread of individual values
about their mean, and is given as

where Xi is one of the n individual values in the data set, and 𝑋ത is the
data set's mean value.

Frequently, we report the relative standard deviation, sr, instead of the


absolute standard deviation.

The percent relative standard deviation, % sr, is srx100


Variance
Variance is the square of the standard deviation.
Errors That Affect Accuracy
Accuracy is how close a measure of central tendency is to its expected
value, μ. We express accuracy either as an absolute error, e

e = 𝑋−μ
or as a percent relative error, %e
𝑋ഥ −μ
%e = μ x100

Errors affecting the accuracy of an analysis are called determinate and


are characterized by a systematic deviation from the true value; that is,
all the individual measurements are either too large or too small. A
positive determinate error results in a central value that is larger than
the true value, and a negative determinate error leads to a central value
that is smaller than the true value. Both positive and negative
determinate errors may affect the result of an analysis, with their
cumulative effect leading to a net positive or negative determinate
error.
Determinate errors may be divided into four categories: sampling
errors, method errors, measurement errors, and personal errors.
Sampling error: An error introduced during the process of collecting
a sample for analysis.
Method error: An error due to limitations in the analytical method
used to analyze a sample.
Measurement error: An error due to limitations in the equipment
and instruments used to make measurements.
Personal error: An error due to biases introduced by the analyst.
Identifying Determinate Errors
Constant determinate error: A determinate error whose value is the
same for all samples. The presence of a constant determinate error can
be detected by running several analyses using different amounts of
sample, and looking for a systematic change in the property being
measured.
Proportional determinate error: A determinate error whose value
depends on the amount of sample analyzed. A proportional determinate
error is more difficult to detect since the result of an analysis is
independent of the amount of sample.
Potential determinate errors also can be identified by analyzing a
standard sample containing a known amount of analyte in a matrix
similar to that of the samples being analyzed.
Precision

Precision is a measure of the spread of data about a central value and


may be expressed as the range, the standard deviation, or the
variance. Precision is commonly divided into two categories:
repeatability and reproducibility.
Repeatability is the precision obtained when all measurements are
made by the same analyst during a single period of laboratory work,
using the same solutions and equipment.
Reproducibility, on the other hand, is the precision obtained under
any other set of conditions, including that between analysts, or
between laboratory sessions for a single analyst.
Errors affecting the distribution of measurements around a central
value are called indeterminate and are characterized by a random
variation in both magnitude and direction.
Evaluating Indeterminate Error
Although it is impossible to eliminate indeterminate error, its effect
can be minimized if the sources and relative magnitudes of the
indeterminate error are known. Indeterminate errors may be estimated
by an appropriate measure of spread. Typically, a standard deviation is
used, although in some cases estimated values are used.
Sources of Indeterminate Error
Indeterminate errors can be traced to several sources, including the
collection of samples, the manipulation of samples during the
analysis, and the making of measurements.
Error and Uncertainty
Analytical chemists make a distinction between error and uncertainty.
Error is the difference between a single measurement or result and its
true value. In other words, error is a measure of bias.
Uncertainty expresses the range of possible values that a measurement
or result might reasonably be expected to have. Uncertainty accounts
for all errors, both determinate and indeterminate, that might affect our
result.
Propagation of Uncertainty
Uncertainty When Adding or Subtracting
When measurements are added or subtracted, the absolute
uncertainty in the result is the square root of the sum of the squares
of the absolute uncertainties for the individual measurements. Thus,
for the equations R = A + B + C or R = A + B – C, or any other
combination of adding and subtracting A, B, and C, the absolute
uncertainty in R is
Example: The class A 10-mL pipet characterized with a result mean
9.992 mL and standard deviation of 0.006 is used to deliver two
successive volumes. Calculate the absolute and relative uncertainties
for the total delivered volume.
SOLUTION
The total delivered volume is obtained by adding the volumes of each
delivery; thus
Vtot = 9.992 mL + 9.992 mL = 19.984 mL
Using the standard deviation as an estimate of uncertainty, the
uncertainty in the total delivered volume is

its absolute uncertainty as 19.984 ± 0.008 mL. The relative


uncertainty in the total delivered volume is
Uncertainty When Multiplying or Dividing
When measurements are multiplied or divided, the relative
uncertainty in the result is the square root of the sum of the
squares of the relative uncertainties for the individual
measurements.
Thus, for the equations R = AxBxC or R = A xB/C, or any
other combination of multiplying and dividing A, B, and C,
the relative uncertainty is
Example: The quantity of charge, Q, in coulombs passing through an
electrical circuit is
Q = Ixt
where I is the current in amperes and t is the time in seconds. When a
current of 0.15 ± 0.01 A passes through the circuit for 120 ± 1 s, the
total charge is
Q = (0.15 A)x(120 s) = 18 C
Calculate the absolute and relative uncertainties for the total charge.
SOLUTION
Since charge is the product of current and time, its relative uncertainty
is

or ±6.7%. The absolute uncertainty in the charge is


sR = Rx0.0672 = (18) x (±0.0672) = ±1.2
Thus, we report the total charge as 18 C ± 1 C.
Populations and Samples
A population is the set of all objects in the system being investigated. If
we analyze every member of a population, we can determine the
population’s true central value, µ, and spread, σ.
The probability of occurrence for a particular value, P(V), is given as

where V is the value of interest, M is the value’s frequency of


occurrence in the population, and N is the size of the population.
In most circumstances, populations are so large that it is not feasible to
analyze every member of the population. Instead, we select and analyze
a limited subset, or sample, of the population. Sample: those members
of a population that we actually collect and analyze.
Confidence Intervals for Populations
Confidence interval: Range of results around a mean value that
could be explained by random error and reported as
Xi = µ ± zσ
where the factor z accounts for the desired level of confidence.

Example: What is the 95% confidence interval for the amount of


aspirin in a single analgesic tablet drawn from a population where µ is
250 mg and σ2 is 25?
SOLUTION
According to Table, the 95% confidence interval for a single member
of a normally distributed population is
Xi = µ ± 1.96σ = 250 mg ± (1.96)(5) = 250 mg ± 10 mg
Thus, we expect that 95% of the tablets in the population contain
between 240 and 260 mg of aspirin.
Alternatively, a confidence interval can be expressed in terms of the
population’s standard deviation and the value of a single member
drawn from the population. Thus, the above equation can be rewritten
as a confidence interval for the population mean
µ = Xi ± zσ

Example: The population standard deviation for the amount of aspirin in a


batch of analgesic tablets is known to be 7 mg of aspirin. A single tablet is
randomly selected, analyzed, and found to contain 245 mg of aspirin. What is
the 95% confidence interval for the population mean?
SOLUTION
The 95% confidence interval for the population mean is given as
µ = Xi ± zσ = 245 ± (1.96)(7) = 245 mg ± 14 mg
There is, therefore, a 95% probability that the population’s mean, µ, lies
within the range of 231–259 mg of aspirin.
Confidence intervals also can be reported using the mean for a sample
of size n, drawn from a population of known σ.

Example: What is the 95% confidence interval for the analgesic


tablets, if an analysis of five tablets yields a mean of 245 mg of
aspirin?
SOLUTION
In this case the confidence interval is given as

Thus, there is a 95% probability that the population’s mean is between


239 and 251 mg of aspirin. As expected, the confidence interval based
on the mean of five members of the population is smaller than that
based on a single member.
Degrees of Freedom
Unlike the population’s variance, the variance of a sample includes the
term n – 1 in the denominator, where n is the size of the sample

The denominators of the variance in the equations is commonly called


the degrees of freedom for the sample.
Confidence Intervals for Samples

To account for the uncertainty in estimating σ2, the term z in equation


is replaced with the variable t, where t is defined such that t > z at all
confidence levels.

Values for t at the 95% confidence level are shown in Table


Significance test
A statistical test to determine if the difference between two values is
significant. The first step in constructing a significance test is to state
the experimental problem as a yes -or- no question. A null hypothesis
and an alternative hypothesis provide answers to the question. The
null hypothesis, HO, is that indeterminate error is sufficient to explain
any difference in the values being compared. The alternative
hypothesis, HA, is that the difference between the values is too great
to be explained by random error and, therefore, must be real. A
significance test is conducted on the null hypothesis, which is either
retained or rejected. If the null hypothesis is rejected, then the
alternative hypothesis must be accepted.
After stating the null and alternative hypotheses, a significance level
for the analysis is chosen. The significance level is the confidence level
for retaining the null hypothesis or, in other words, the probability that
the null hypothesis will be incorrectly rejected. In the former case the
significance level is given as a percentage (e.g., 95%), whereas in the
latter case, it is given as α, where α is defined as

Thus, for a 95% confidence level, α is 0.05.


ഥ to µ
Comparing 𝑿
One approach for validating a new analytical method is to analyze a
standard sample containing a known amount of analyte, µ . The
method’s accuracy is judged by determining the average amount of
analyte in several samples, ഥX, and using a significance test to compare
it with µ. The null hypothesis is that ഥ
X and µ are the same and that any
difference between the two values can be explained by indeterminate
errors affecting the determination of ഥ X. The alternative hypothesis is
that the difference between ഥ X and µ is too large to be explained by
indeterminate error.
The equation for the test (experimental) statistic, texp, is derived from
the confidence interval for µ
The value of texp is compared with a critical value, t(α,v), which is
determined by the chosen significance level,α, the degrees of freedom
for the sample, v, and whether the significance test is one-tailed or two-
tailed. If texp is greater than t(α,v), then the confidence interval for the
data is wider than that expected from indeterminate errors. In this case,
the null hypothesis is rejected and the alternative hypothesis is
accepted. If texp is less than or equal to t(α,v), then the confidence
interval for the data could be attributed to indeterminate error, and the
null hypothesis is retained at the stated significance level.
Example: Before determining the amount of Na2CO3 in an unknown
sample, a student decides to check her procedure by analyzing a
sample known to contain 98.76% w/w Na2CO3. Five replicate
determinations of the %w/w Na2CO3 in the standard were made with
the following results
98.71% 98.59% 98.62% 98.44% 98.58%
Is the mean for these five trials significantly different from the
accepted value at the 95% confidence level (α = 0.05)?

SOLUTION
The mean and standard deviation for the five trials are

X= 98.59, s = 0.0973
The null and alternative hypotheses are
H0: ഥX=µ, HA: ഥ
X≠µ
The test statistic is

The critical value for t(0.05,4), as found in table, is 2.78. Since texp is
greater than t(0.05, 4), we must reject the null hypothesis and accept
the alternative hypothesis. At the 95% confidence level the difference
between ഥ X and µ is significant and cannot be explained by
indeterminate sources of error. There is evidence, therefore, that the
results are affected by a determinate source of error.

If evidence for a determinate error is found, its source should be


identified and corrected before analyzing additional samples.
F-test: Comparing σ2 to s2
Statistical test for comparing two variances to see if their difference
is too large to be explained by indeterminate error. The test statistic
for evaluating the null hypothesis is called an F-test, and is given as
either

depending on whether s2 is larger or smaller than σ2.


Note that Fexp is defined such that its value is always greater than or
equal to 1. If the null hypothesis is true, then Fexp should equal 1.
Example: A manufacturer’s process for analyzing aspirin tablets has a
known variance of 25. A sample of ten aspirin tablets is selected and
analyzed for the amount of aspirin, yielding the following results
254 249 252 252 249 249 250 247 251 252
Determine whether there is any evidence that the measurement process
is not under statistical control at α = 0.05.
SOLUTION
The variance for the sample of ten tablets is 4.3. The null hypothesis
and alternative hypotheses are
H0: s2 = σ2 HA: s2 ≠ σ2
The test statistic is

The critical value for F(0.05, ∞, 9) from table is 3.33. Since Fexp is
greater than F(0.05,∞, 9), we reject the null hypothesis and accept the
alternative hypothesis that the analysis is not under statistical control.
Comparing Two Sample Variances
The F-test can be extended to the comparison of variances for two
samples, A and B, by rewriting equation as

where A and B are defined such that s2A is greater than or equal to
s2B.
Comparing Two Sample Means
Consider two samples, A and B, for which mean values, 𝑋ത A and
𝑋ത B, and standard deviations, sA and sB, have been measured.
Confidence intervals for µA and µB can be written for both samples

where nA and nB are the number of replicate trials conducted on


samples A and B. A comparison of the mean values is based on the
null hypothesis that 𝑋ത A and 𝑋ത B are identical, and an alternative
hypothesis that the means are significantly different.
A test statistic is derived by letting µA equal µB, and combining
equations

The value of texp is compared with a critical value, t(α, v), as


determined by the chosen significance level, α, the degrees of freedom
for the sample, v, and whether the significance test is one-tailed or two-
tailed.
If the variances s2A and s2B estimate the same σ2, then the two standard
deviations can be factored out of the equation and replaced by a
pooled standard deviation, spool, which provides a better estimate for
the precision of the analysis.

with the pooled standard deviation given as


Q-test - Outliers
Outlier: Data point whose value is much larger or smaller than the
remaining data.
Dixon’s Q-test: Statistical test for deciding if an outlier can be
removed from a set of data.
The Q-test compares the difference between the suspected outlier and
its nearest numerical neighbor to the range of the entire data set. Data
are ranked from smallest to largest so that the suspected outlier is
either the first or the last data point.
The test statistic, Qexp, is calculated using equation if the suspected
outlier is the smallest value (X1)

or using equation if the suspected outlier is the largest value (Xn)


If Qexp is greater than Q(α, n), then the null hypothesis is rejected and
the outlier may be rejected. When Qexp is less than or equal to Q(α, n)
the suspected outlier must be retained.
Example: The following masses, in grams, were recorded in an
experiment to determine the average mass of a U.S. penny.
3.067 3.049 3.039 2.514 3.048 3.079 3.094 3.109 3.102
Determine if the value of 2.514 g is an outlier at α= 0.05.

SOLUTION
To begin with, place the masses in order from smallest to largest
2.514 3.039 3.048 3.049 3.067 3.079 3.094 3.102 3.109
and calculate Qexp

The critical value for Q(0.05, 9) is 0.493. Since Qexp > Q(0.05, 9) the
value is assumed to be an outlier, and can be rejected
Detection Limits
A method’s detection limit is the smallest amount or concentration of
analyte that can be detected with statistical confidence. The
International Union of Pure and Applied Chemistry (IUPAC) defines
the detection limit as the smallest concentration or absolute amount of
analyte that has a signal significantly larger than the signal arising
from a reagent blank. Mathematically, the analyte’s signal at the
detection limit, (SA)DL, is
(SA)DL = Sreag + zsreag
where Sreag is the signal for a reagent blank, sreag is the known standard
deviation for the reagent blank’s signal, and z is a factor accounting
for the desired confidence level.
Limit of quantitation (LOQ): The smallest concentration or
absolute amount of analyte that can be reliably determined.
The American Chemical Society’s Committee on Environmental
Analytical Chemistry recommends the limit of quantitation, (SA)LOQ,
which is defined as
(SA)LOQ = Sreag + 10sreag

You might also like