Standardizing Analytical Methods Guide
Standardizing Analytical Methods Guide
Standardizing Analytical
Methods
Chapter Overview
5A Analytical Standards
5B Calibrating the Signal (Stotal)
5C Determining the Sensitivity (kA)
5D Linear Regression and Calibration Curves
5E Compensating for the Reagent Blank (Sreag)
5F Using Excel and R for a Regression Analysis
5G Key Terms
5H Chapter Summary
5I Problems
5J Solutions to Practice Exercises
1 ACS Committee on Environmental Improvement “Guidelines for Data Acquisition and Data Quality Evaluation in
Environmental Chemistry,” Anal. Chem. 1980, 52, 2242–2249.
147
148 Analytical Chemistry 2.1
5A Analytical Standards
To standardize an analytical method we use standards that contain known
amounts of analyte. The accuracy of a standardization, therefore, depends
on the quality of the reagents and the glassware we use to prepare these
standards. For example, in an acid–base titration the stoichiometry of the
acid–base reaction defines the relationship between the moles of analyte
and the moles of titrant. In turn, the moles of titrant is the product of the
See Chapter 9 for a thorough discussion of titrant’s concentration and the volume of titrant used to reach the equiva-
titrimetric methods of analysis.
lence point. The accuracy of a titrimetric analysis, therefore, is never better
than the accuracy with which we know the titrant’s concentration.
2 (a) Smith, B. W.; Parsons, M. L. J. Chem. Educ. 1973, 50, 679–681; (b) Moody, J. R.; Green-
burg, P. R.; Pratt, K. W.; Rains, T. C. Anal. Chem. 1988, 60, 1203A–1218A.
3 Committee on Analytical Reagents, Reagent Chemicals, 8th ed., American Chemical Society:
Washington, D. C., 1993.
Chapter 5 Standardizing Analytical Methods 149
(a) (b)
Figure 5.1 Two examples of packaging labels for reagent grade chemicals. The label in (a) pro-
vides the manufacturer’s assay for the reagent, NaBr. Note that potassium is flagged with an
asterisk (*) because its assay exceeds the limit established by the American Chemical Society
(ACS). The label in (b) does not provide an assay for impurities; however it indicates that the
reagent meets ACS specifications by providing the maximum limits for impurities. An assay for
the reagent, NaHCO3, is provided.
150 Analytical Chemistry 2.1
assumed relationship
actual relationship
Ssamp
Example 5.1
A spectrophotometric method for the quantitative analysis of Pb2+ in
blood yields an Sstd of 0.474 for a single standard for which the concentra-
tion of lead is 1.75 ppb. What is the concentration of Pb2+ in a sample of
blood for which Ssamp is 0.361?
Solution
Equation 5.5 allows us to calculate the value of kA using the data for the
single external standard.
kA = CS std = 0.474 = 0.2709 ppm -1
std 1.75 ppb
Having determined the value of kA, we calculate the concentration of Pb2+
in the sample of blood is calculated using equation 5.6.
S samp 0.361
CA = = = 1.33 ppb
kA 0.2709 ppm -1
Chapter 5 Standardizing Analytical Methods 153
Example 5.2
A second spectrophotometric method for the quantitative analysis of Pb2+
in blood has a normal calibration curve for which
S std = (0.296 ppb -1) # C std + 0.003
What is the concentration of Pb2+ in a sample of blood if Ssamp is 0.397?
0.25
0.20
0.15
Sstd Figure 5.3 The photo at the top of the figure shows
0.10 a reagent blank (far left) and a set of five external
standards for Cu2+ with concentrations that in-
0.05 crease from left-to-right. Shown below the external
0 standards is the resulting normal calibration curve.
0 0.0020 0.0040 0.0060 0.0080 The absorbance of each standard, Sstd, is shown by
Cstd (M) the filled circles.
154 Analytical Chemistry 2.1
Solution
To determine the concentration of Pb2+ in the sample of blood, we replace
Sstd in the calibration equation with Ssamp and solve for CA.
S samp - 0.003
CA = = 0.397 - 0.003 = 1.33 ppb
0.296 ppb -1 0.296 ppb -1
It is worth noting that the calibration equation in this problem includes
an extra term that does not appear in equation 5.6. Ideally we expect
our calibration curve to have a signal of zero when CA is zero. This is the
purpose of using a reagent blank to correct the measured signal. The extra
term of +0.003 in our calibration equation results from the uncertainty
in measuring the signal for the reagent blank and the standards.
standard’s
matrix
sample’s
Ssamp
matrix
S spike = k A a C A V
Vo + C Vstd k dard, respectively.
f
std
Vf 5.8
As long as Vstd is small relative to Vo, the effect of the standard’s matrix on
the sample’s matrix is insignificant. Under these conditions the value of kA
is the same in equation 5.7 and equation 5.8. Solving both equations for
kA and equating gives
S samp S spike
Vo = Vo + C Vstd
CA V CA V 5.9
f f
std
Vf
which we can solve for the concentration of analyte, CA, in the original
sample.
156 Analytical Chemistry 2.1
dilute to Vf
Figure 5.5 Illustration showing the method of stan-
dard additions. The volumetric flask on the left con-
tains a portion of the sample, Vo, and the volumetric
flask on the right contains an identical portion of the
sample and a spike, Vstd, of a standard solution of the
analyte. Both flasks are diluted to the same final vol-
ume, Vf. The concentration of analyte in each flask is
shown at the bottom of the figure where CA is the ana- Vo Vo V
Concentration
lyte’s concentration in the original sample and Cstd is CA × CA × + C std × std
the concentration of analyte in the external standard. of Analyte Vf Vf Vf
Example 5.3
A third spectrophotometric method for the quantitative analysis of Pb2+ in
blood yields an Ssamp of 0.193 when a 1.00 mL sample of blood is diluted
to 5.00 mL. A second 1.00 mL sample of blood is spiked with 1.00 mL of
a 1560-ppb Pb2+ external standard and diluted to 5.00 mL, yielding an
Sspike of 0.419. What is the concentration of Pb2+ in the original sample
of blood?
Solution
We begin by making appropriate substitutions into equation 5.9 and solv-
ing for CA. Note that all volumes must be in the same units; thus, we first
covert Vstd from 1.00 mL to 1.00 × 10–3 mL.
0.193 = 0.419
1 . 00 mL
C A 5.00mL C A 5.00mL + 1560 ppb 1.005.#
1 . 00 mL 10 -3 mL
00 mL
0.193 0.419
0.200C A = 0.200C A + 0.3120 ppb
C A = 1.33 ppb
The concentration of Pb2+ in the original sample of blood is 1.33 ppb.
Chapter 5 Standardizing Analytical Methods 157
S spike = k A a C A V +Vo
V + C std V +Vstd k
Vstd 5.10
o std o
S samp S spike
CA = V Vstd 5.11
C A V +o V + C std V + Vstd
o std o
Example 5.4
A fourth spectrophotometric method for the quantitative analysis of Pb2+
in blood yields an Ssamp of 0.712 for a 5.00 mL sample of blood. After spik-
ing the blood sample with 5.00 mL of a 1560-ppb Pb2+ external standard,
an Sspike of 1.546 is measured. What is the concentration of Pb2+ in the
original sample of blood?
Solution
To determine the concentration of Pb2+ in the original sample of blood,
we make appropriate substitutions into equation 5.11 and solve for CA.
0.712 = 1.546 Vo + Vstd = 5.000 mL + 5.00×10
–3
mL
CA
C A 55..005
00 mL + 1560 ppb 5.00 # 10 -3 mL
mL 5.005 mL = 5.005 mL
0.712 = 1.546
CA 0.9990C A + 1.558 ppb
Vo Vo
Concentration Vo Vstd
CA CA + C std
of Analyte Vo + Vs td Vo + Vs td
C A = 1.33 ppb
2+
The concentration of Pb in the original sample of blood is 1.33 ppb.
Example 5.5
Beginning with equation 5.8 show that the equations in Figure 5.7a for
the slope, the y-intercept, and the x-intercept are correct.
Solution
We begin by rewriting equation 5.8 as
S spike = k A C A Vo k A C std
V f + V f # Vstd
which is in the form of the equation for a straight-line
y = y-intercept + slope × x
where y is Sspike and x is Vstd. The slope of the line, therefore, is kACstd/Vf
and the y-intercept is kACAVo/Vf. The x-intercept is the value of x when y
is zero, or
0 = kAC A Vo k A Vstd
V f + V f # x-intercept
k A C A Vo V f
x-intercept =- =- CCA Vo
k A C std V f std
(a) 0.60
kACAVo
0.50 y-intercept =
Vf
0.40
Sspike 0.30 kACstd
slope =
Vf
0.20
0.10
0
-2.00 0 2.00 4.00 6.00
Figure 5.7 Shown at the top of the
Vstd (mL)
-C V figure is a set of six standard additions
x-intercept = A o
Cstd
for the determination of Mn2+. The
(b) 0.60
flask on the left is a 25.00 mL sample
kACAVo
0.50 y-intercept =
Vf
diluted to 50.00 mL with water. The
0.40
remaining flasks contain 25.00 mL of
sample and, from left-to-right, 1.00,
Sspike 0.30 slope = kA 2.00, 3.00, 4.00, and 5.00 mL spikes
0.20 of an external standard that is 100.6
mg/L Mn2+. Shown below are two
0.10
ways to plot the standard additions
0 calibration curve. The absorbance for
-4.00 -2.00 0 2.00 4.00 6.00 8.00 10.00 12.00
each standard addition, Sspike, is shown
-CAVo Vstd
x-intercept = Cstd × (mg/L) by the filled circles.
Vf Vf
Example 5.6
A fifth spectrophotometric method for the quantitative analysis of Pb2+
in blood uses a multiple-point standard addition based on equation 5.8.
The original blood sample has a volume of 1.00 mL and the standard used
for spiking the sample has a concentration of 1560 ppb Pb2+. All samples
were diluted to 5.00 mL before measuring the signal. A calibration curve
of Sspike versus Vstd has the following equation
S spike = 0.266 + 312 mL-1 # Vstd
What is the concentration of Pb2+ in the original sample of blood?
Solution
To find the x-intercept we set Sspike equal to zero.
0 = 0.266 + 312 mL-1 # Vstd
160 Analytical Chemistry 2.1
Solving for Vstd, we obtain a value of –8.526 × 10–4 mL for the x-intercept.
Substituting the x-intercept’s value into the equation from Figure 5.7a
- 8.526 # 10 -4 mL =- CCA Vo =- C A # 1.00 mL
std 1560 ppb
and solving for CA gives the concentration of Pb2+ in the blood sample as
1.33 ppb.
S IS = k IS C IS
where kA and kIS are the sensitivities for the analyte and the internal stan-
dard, respectively. Taking the ratio of the two signals gives the fundamental
equation for an internal standardization.
SA kACA CA
S IS = k IS C IS = K # C IS 5.12
Because K is a ratio of the analyte’s sensitivity and the internal standard’s
sensitivity, it is not necessary to determine independently values for either
kA or kIS.
Example 5.7
A sixth spectrophotometric method for the quantitative analysis of Pb2+
in blood uses Cu2+ as an internal standard. A standard that is 1.75 ppb
Pb2+ and 2.25 ppb Cu2+ yields a ratio of (SA/SIS)std of 2.37. A sample of
blood spiked with the same concentration of Cu2+ gives a signal ratio,
(SA/SIS)samp, of 1.80. What is the concentration of Pb2+ in the sample of
blood?
Solution
Equation 5.13 allows us to calculate the value of K using the data for the
standard
C A = CKIS # a SS A k =
2.25 ppb Cu 2+
# 1.80 = 1.33 ppb Pb 2+
IS samp ppb Cu 2+
3.05
ppb Pb 2+
Example 5.8
A seventh spectrophotometric method for the quantitative analysis of Pb2+
in blood gives a linear internal standards calibration curve for which
a SS A k = (2.11 ppb -1) # C A - 0.006
IS std
a SS A k + 0.006
= 2.80 + 0.006
IS samp
CA = = 1.33 Pb 2+
2.11 ppb -1 2.11 ppb -1
The concentration of Pb2+ in the sample of blood is 1.33 ppb.
You might wonder if it is possible to in-
clude an internal standard in the method
In some circumstances it is not possible to prepare the standards so of standard additions to correct for both
that each contains the same concentration of internal standard. This is the matrix effects and uncontrolled variations
case, for example, when we prepare samples by mass instead of volume. We between samples; well, the answer is yes
as described in the paper “Standard Dilu-
can still prepare a calibration curve, however, by plotting (SA/SIS)std versus tion Analysis,” the full reference for which
CA/CIS, giving a linear calibration curve with a slope of K. is Jones, W. B.; Donati, G. L.; Calloway,
C. P.; Jones, B. T. Anal. Chem. 2015, 87,
2321-2327.
5D Linear Regression and Calibration Curves
In a single-point external standardization we determine the value of kA
by measuring the signal for a single standard that contains a known con-
centration of analyte. Using this value of kA and our sample’s signal, we
then calculate the concentration of analyte in our sample (see Example
5.1). With only a single determination of kA, a quantitative analysis using
a single-point external standardization is straightforward.
A multiple-point standardization presents a more difficult problem.
Consider the data in Table 5.1 for a multiple-point external standardiza-
tion. What is our best estimate of the relationship between Sstd and Cstd? It
is tempting to treat this data as five separate single-point standardizations,
determining kA for each standard, and reporting the mean value for the
five trials. Despite it simplicity, this is not an appropriate way to treat a
multiple-point standardization.
So why is it inappropriate to calculate an average value for kA using
the data in Table 5.1? In a single-point standardization we assume that the
reagent blank (the first row in Table 5.1) corrects for all constant sources
of determinate error. If this is not the case, then the value of kA from a
single-point standardization has a constant determinate error. Table 5.2
demonstrates how an uncorrected constant error affects our determination
Table 5.2 Effect of a Constant Determinate Error on the Value of kA From a Single-
Point Standardization
Sstd kA = Sstd/ Cstd (Sstd)e kA = (Sstd)e/ Cstd
Cstd (without constant error) (actual) (with constant error) (apparent)
1.00 1.00 1.00 1.50 1.50
2.00 2.00 1.00 2.50 1.25
3.00 3.00 1.00 3.50 1.17
4.00 4.00 1.00 4.50 1.13
5.00 5.00 1.00 5.50 1.10
mean kA (true) = 1.00 mean kA (apparent) = 1.23
of kA. The first three columns show the concentration of analyte in a set of
standards, Cstd, the signal without any source of constant error, Sstd, and
the actual value of kA for five standards. As we expect, the value of kA is the
same for each standard. In the fourth column we add a constant determi-
nate error of +0.50 to the signals, (Sstd)e. The last column contains the cor-
responding apparent values of kA. Note that we obtain a different value of
kA for each standard and that each apparent kA is greater than the true value.
How do we find the best estimate for the relationship between the sig-
nal and the concentration of analyte in a multiple-point standardization?
Figure 5.8 shows the data in Table 5.1 plotted as a normal calibration curve.
Although the data certainly appear to fall along a straight line, the actual
calibration curve is not intuitively obvious. The process of determining the
best equation for the calibration curve is called linear regression.
60
50
40
Sstd 30
20
10
0
0.0 0.1 0.2 0.3 0.4 0.5
Cstd
Figure 5.8 Normal calibration curve data for the hypothetical multiple-point
external standardization in Table 5.1.
(1) that the difference between our experimental data and the calculated
regression line is the result of indeterminate errors that affect y,
(2) that indeterminate errors that affect y are normally distributed, and
(3) that the indeterminate errors in y are independent of the value of x.
Because we assume that the indeterminate errors are the same for all stan-
dards, each standard contributes equally in our estimate of the slope and
the y-intercept. For this reason the result is considered an unweighted
linear regression.
The second assumption generally is true because of the central limit the-
orem, which we considered in Chapter 4. The validity of the two remaining
assumptions is less obvious and you should evaluate them before you accept
the results of a linear regression. In particular the first assumption always is
suspect because there certainly is some indeterminate error in the measure-
ment of x. When we prepare a calibration curve, however, it is not unusual
to find that the uncertainty in the signal, Sstd, is significantly larger than the
uncertainty in the analyte’s concentration, Cstd. In such circumstances the
first assumption is usually reasonable.
nounce T
where b0 and b1 are estimates for the y-intercept and the slope, and Vy is the
If you are reading this aloud, you pro-
y as y-hat.
predicted value of y for any value of x. Because we assume that all uncer-
tainty is the result of indeterminate errors in y, the difference between y and
Vy for each value of x is the residual error, r, in our mathematical model.
ri = (y i - Vy i)
The reason for squaring the individual
residual errors is to prevent a positive re-
Figure 5.10 shows the residual errors for the three data points. The smaller
sidual error from canceling out a negative the total residual error, R, which we define as
R = / (y i - Vy i) 2
residual error. You have seen this before in n
the equations for the sample and popula- 5.16
tion standard deviations. You also can see i=1
from this equation why a linear regression the better the fit between the straight-line and the data. In a linear regres-
is sometimes called the method of least
squares. sion analysis, we seek values of b0 and b1 that give the smallest total residual
error.
ŷ3 ŷ = b0 + b1 x
y2
r3 = ( y3 − yˆ3 )
r2 = ( y 2 − yˆ2 )
y3
ŷ1 ŷ 2
r1 = ( y1 − yˆ1 )
y1
Figure 5.10 Illustration showing the evaluation of a linear regression in which we assume that all un-
certainty is the result of indeterminate errors in y. The points in blue, yi, are the original data and the
points in red, Vy i , are the predicted values from the regression equation, Vy = b 0 + b 1 x .The smaller
the total residual error (equation 5.16), the better the fit of the straight-line to the data.
Chapter 5 Standardizing Analytical Methods 167
b0 =n
i=1 i=1
Example 5.9
Using the data from Table 5.1, determine the relationship between Sstd and
Cstd using an unweighted linear regression.
Solution
We begin by setting up a table to help us organize the calculation.
xi yi xiyi xi2
0.000 0.00 0.000 0.000
Equations 5.17 and 5.18 are written in
0.100 12.36 1.236 0.010 terms of the general variables x and y. As
0.200 24.83 4.966 0.040 you work through this example, remem-
ber that x corresponds to Cstd, and that y
0.300 35.91 10.773 0.090 corresponds to Sstd.
0.400 48.79 19.516 0.160
0.500 60.42 30.210 0.250
Adding the values in each column gives
n n n n
/x i = 1.500 /y i = 182.31 /x y i i = 66.701 /x 2
i = 0.550
i=1 i=1 i=1 i=1
Substituting these values into equation 5.17 and equation 5.18, we find
that the slope and the y-intercept are
6 See, for example, Draper, N. R.; Smith, H. Applied Regression Analysis, 3rd ed.; Wiley: New
York, 1998.
168 Analytical Chemistry 2.1
/ (y - Vy )
(equation 5.19) and the standard devia- n
2
tion for a sample (equation 4.1)? i i
5.19
n
sr = n-2
i=1
60
50
40
Sstd 30
20
10
0
0.0 0.1 0.2 0.3 0.4 0.5
Cstd
Figure 5.11 Calibration curve for the data in Table 5.1 and Example 5.9.
Chapter 5 Standardizing Analytical Methods 169
sb = ns r2 = s 2r
n / x i2 - c / x i m / ^x - x h
n n n
1 2
2 5.20
i
i=1 i=1 i=1
n n
s 2r / x 2i s 2r / x 2i
sb = i=1
= i=1
5.21
n / x - c/ xi m n / ^ x i - x h2
0 n n 2 n
2
i
i=1 i=1 i=1
tercept from Example 5.9, and the squares of the residual error, _ y i - Vy i i .
2
member that x corresponds to Cstd, and
that y corresponds to Sstd.
Using the last standard as an example, we find that the predicted signal is
Vy 6 = b 0 + b 1 x 6 = 0.209 + (120.706 # 0.500) = 60.562
and that the square of the residual error is
_ y i - Vy i i = (60.42 - 60.562) 2 = 0.2016 . 0.202
2
The following table displays the results for all six solutions.
Vy i _ y i - Vy i i
2
xi yi
0.000 0.00 0.209 0.0437
0.100 12.36 12.280 0.0064
0.200 24.83 24.350 0.2304
0.300 35.91 36.421 0.2611
0.400 48.79 48.491 0.0894
0.500 60.42 60.562 0.0202
170 Analytical Chemistry 2.1
Adding together the data in the last column gives the numerator of equa-
tion 5.19 as 0.6512; thus, the standard deviation about the regression is
0.6512 = 0.4035
sr =
6-2
Next we calculate the standard deviations for the slope and the y-intercept
using equation 5.20 and equation 5.21. The values for the summation
terms are from in Example 5.9.
ns r2 6 # (0.4035) 2
sb = = = 0.965
n / x i2 - c / x i m
n n
1 2
(6 # 0.550) - (1.500) 2
i=1 i=1
n
s 2r / x 2i
(0.4035) 2 # 0.550
sb = i=1
= = 0.292
n / x i2 - c / x i m
n n
0 2
(6 # 0.550) - (1.500) 2
i=1 i=1
You can find values for t in Appendix 4. Finally, the 95% confidence intervals (a = 0.05, 4 degrees of freedom) for
the slope and y-intercept are
b 1 = b 1 ! ts b = 120.706 ! (2.78 # 0.965) = 120.7 ! 2.7
1
The standard deviation about the regression, sr, suggests that the signal, Sstd,
is precise to one decimal place. For this reason we report the slope and the
y-intercept to a single decimal place.
S samp - b 0
CA = 5.24
b1
What is less obvious is how to report a confidence interval for CA that
expresses the uncertainty in our analysis. To calculate a confidence interval
we need to know the standard deviation in the analyte’s concentration, s C , A Equation 5.25 is written in terms of a cali-
which is given by the following equation bration experiment. A more general form
of the equation, written in terms of x and
^ S samp - S std h2 y, is given here.
sC = sr 1 1
m+n+
(b 1) 2 / ^C std - C std h2
b1 n
5.25
^Y - y h
A
2
i
sr 1 1
^ x i - x h2
i=1 sx = + + n
b1 m n 2 /
(b 1)
where m is the number of replicate we use to establish the sample’s average i=1
signal, S samp , n is the number of calibration standards, S std is the average
signal for the calibration standards, and C std and C std are the individual and
i A close examination of equation 5.25
the mean concentrations for the calibration standards.7 Knowing the value should convince you that the uncertainty
in CA is smallest when the sample’s av-
of s C , the confidence interval for the analyte’s concentration is
A
erage signal, S samp , is equal to the aver-
n C = C A ! ts C
A A
age signal for the standards, S std . When
practical, you should plan your calibration
where nCA is the expected value of CA in the absence of determinate errors, curve so that Ssamp falls in the middle of
and with the value of t is based on the desired level of confidence and n–2 the calibration curve.
degrees of freedom.
Example 5.11
Three replicate analyses for a sample that contains an unknown concentra-
tion of analyte, yield values for Ssamp of 29.32, 29.16 and 29.51 (arbitrary
units). Using the results from Example 5.9 and Example 5.10, determine
the analyte’s concentration, CA, and its 95% confidence interval.
Solution
The average signal, S samp , is 29.33, which, using equation 5.24 and the
slope and the y-intercept from Example 5.9, gives the analyte’s concentra-
tion as
S samp - b 0
CA = = 29.33 - 0.209 = 0.241
b1 120.706
To calculate the standard deviation for the analyte’s concentration we must
determine the values for S std and for / ^C std - C std h2 . The former is just
i
the average signal for the calibration standards, which, using the data in
Table 5.1, is 30.385. Calculating / ^C std - C std h2 looks formidable, but
i
7 (a) Miller, J. N. Analyst 1991, 116, 3–14; (b) Sharaf, M. A.; Illman, D. L.; Kowalski, B. R. Che-
mometrics, Wiley-Interscience: New York, 1986, pp. 126-127; (c) Analytical Methods Commit-
tee “Uncertainties in concentrations estimated from calibration experiments,” AMC Technical
Brief, March 2006.
172 Analytical Chemistry 2.1
60
50
40
Sstd 30
the calibration standards. Using the data in Table 5.1 we find that s C is std
0.1871 and
s C = 0.4035 3
A
120.706 6 (120.706) 2 # 0.175
Finally, the 95% confidence interval for 4 degrees of freedom is
n C = C A ! ts C = 0.241 ! (2.78 # 0.0024) = 0.241 ! 0.007
A A
Figure 5.12 shows the calibration curve with curves showing the 95%
confidence interval for CA.
C A = x-intercept = - b 0
b1
and the standard deviation in CA is
^ S std h2
sC = sr 1+
(b 1) 2 / ^C std - C std h2
A
b1 n n
i
i=1
residual error
residual error
residual error
0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.1 0.2 0.3 0.4 0.5
Cstd Cstd Cstd
Figure 5.13 Plots of the residual error in the signal, Sstd, as a function of the concentration of analyte, Cstd, for an
unweighted straight-line regression model. The red line shows a residual error of zero. The distribution of the residual
errors in (a) indicates that the unweighted linear regression model is appropriate. The increase in the residual errors in
(b) for higher concentrations of analyte, suggests that a weighted straight-line regression is more appropriate. For (c),
the curved pattern to the residuals suggests that a straight-line model is inappropriate; linear regression using a quadratic
model might produce a better fit.
Practice Exercise 5.5
Using your results from Practice Exercise 5.4, construct a residual plot
and explain its significance.
Click here to review your answer to this exercise.
n n n
n / wi xi yi - / wi xi / wi yi
b1 = i=1 i=1 i=1
5.27
n / wi x - c/ wi xi m
n n 2
2
i
i=1 i=1
/ ^s h-2
n
5.28
yi
i=1
Example 5.12
Shown here are data for an external standardization in which sstd is the
standard deviation for three replicate determination of the signal.
Cstd (arbitrary units) Sstd (arbitrary units) sstd
0.000 0.00 0.02
0.100 12.36 0.02 This is the same data used in Example 5.9
0.200 24.83 0.07 with additional information about the
standard deviations in the signal.
0.300 35.91 0.13
0.400 48.79 0.22
0.500 60.42 0.33
As you work through this example, re-
member that x corresponds to Cstd, and
Determine the calibration curve’s equation using a weighted linear regres- that y corresponds to Sstd.
sion.
Solution
We begin by setting up a table to aid in calculating the weighting factors.
xi yi sy i ^s y h-2
i wi
0.000 0.00 0.02 2500.00 2.8339
As a check on your calculations, the sum
0.100 12.36 0.02 2500.00 2.8339 of the individual weights must equal the
0.200 24.83 0.07 204.08 0.2313 number of calibration standards, n. The
sum of the entries in the last column is
0.300 35.91 0.13 59.17 0.0671 6.0000, so all is well.
0.400 48.79 0.22 20.66 0.0234
0.500 60.42 0.33 9.18 0.0104
Adding together the values in the forth column gives
/ ^s h-2
n
yi
i=1
which we use to calculate the individual weights in the last column. After
we calculate the individual weights, we use a second table to aid in calculat-
ing the four summation terms in equation 5.26 and equation 5.27.
xi yi wi wi xi wi yi wi xi2 wi xi yi
0.000 0.00 2.8339 0.0000 0.0000 0.0000 0.0000
0.100 12.36 2.8339 0.2834 35.0270 0.0283 3.5027
0.200 24.83 0.2313 0.0463 5.7432 0.0093 1.1486
0.300 35.91 0.0671 0.0201 2.4096 0.0060 0.7229
0.400 48.79 0.0234 0.0094 1.1417 0.0037 0.4567
0.500 60.42 0.0104 0.0052 0.6284 0.0026 0.3142
Adding the values in the last four columns gives
176 Analytical Chemistry 2.1
n n
/w x i i = 0.3644 /w y i i = 44.9499
i=1 i=1
n n
/w x i
2
i = 0.0499 /w x yi i i = 6.1451
i=1 i=1
Substituting these values into the equation 5.26 and equation 5.27 gives
the estimated slope and estimated y-intercept as
(6 # 6.1451) - (0.3644 # 44.9499)
b1 = = 122.985
(6 # 0.0499) - (0.3644) 2
44.9499 - (122.985 # 0.3644)
b0 = = 0.0224
6
The calibration equation is
S std = 122.98 # C std + 0.02
Figure 5.14 shows the calibration curve for the weighted regression and the
calibration curve for the unweighted regression in Example 5.9. Although
the two calibration curves are very similar, there are slight differences in the
slope and in the y-intercept. Most notably, the y-intercept for the weighted
linear regression is closer to the expected value of zero. Because the stan-
dard deviation for the signal, Sstd, is smaller for smaller concentrations of
analyte, Cstd, a weighted linear regression gives more emphasis to these
standards, allowing for a better estimate of the y-intercept.
40
Sstd 30
20
10
0
0.0 0.1 0.2 0.3 0.4 0.5
Cstd
Figure 5.14 A comparison of the unweighted and the weighted normal calibra-
tion curves. See Example 5.9 for details of the unweighted linear regression and
Example 5.12 for details of the weighted linear regression.
Chapter 5 Standardizing Analytical Methods 177
Equations for calculating confidence intervals for the slope, the y-in-
tercept, and the concentration of analyte when using a weighted linear
regression are not as easy to define as for an unweighted linear regression.8
The confidence interval for the analyte’s concentration, however, is at its
optimum value when the analyte’s signal is near the weighted centroid, y c ,
of the calibration curve.
n
1 /w x
yc = n i i
i=1
The regression models in this chapter apply only to functions that con-
tain a single independent variable, such as a signal that depends upon the
analyte’s concentration. In the presence of an interferent, however, the signal
may depend on the concentrations of both the analyte and the interferent
S = k A C A + k I C I + S reag
Check out this chapter’s Additional Re- where kI is the interferent’s sensitivity and CI is the interferent’s concentra-
sources at the end of the textbook for
more information about linear regression
tion. Multivariate calibration curves are prepared using standards that con-
with errors in both variables, curvilinear tain known amounts of both the analyte and the interferent, and modeled
regression, and multivariate regression. using multivariate regression.11
Table 5.4 Equations and Resulting Concentrations of Analyte for Different Approaches
to Correcting for the Blank
Concentration of Analyte in...
Approach for Correcting The Signal Equation Sample 1 Sample 2 Sample 3
W A = S samp
CA = W
ignore calibration and reagent blank samp k A W samp 0.1707 0.1610 0.1552
W A = S samp - CB
CA = W
use calibration blank only samp k A W samp 0.1441 0.1409 0.1390
W A = S samp - RB
CA = W
use reagent blank only samp k A W samp 0.1494 0.1449 0.1422
W A = S samp - CB - RB
use both calibration and reagent blank C A = W samp k A W samp 0.1227 0.1248 0.1261
W A = S samp - TYB
CA = W
use total Youden blank samp k A W samp 0.1313 0.1313 0.1313
CA = concentration of analyte; WA = weight of analyte; Wsamp = weight of sample; kA = slope of calibration curve (0.075; see Table
5.3); CB = calibration blank (0.125; see Table 5.3); RB = reagent blank (0.100; see Table 5.3); TYB = total Youden blank (0.185; see
text)
In working up this data, the analytical chemists used at least four dif-
ferent approaches to correct the signals: (a) ignoring both the calibration
blank, CB, and the reagent blank, RB, which clearly is incorrect; (b) using
the calibration blank only; (c) using the reagent blank only; and (d) using
both the calibration blank and the reagent blank. The first four rows of
Table 5.4 shows the equations for calculating the analyte’s concentration
using each approach, along with the reported concentrations for the analyte
in each sample.
That all four methods give a different result for the analyte’s concentra-
tion underscores the importance of choosing a proper blank, but does not
tell us which blank is correct. Because all four methods fail to predict the
same concentration of analyte for each sample, none of these blank correc-
tions properly accounts for an underlying constant source of determinate
error.
To correct for a constant method error, a blank must account for sig-
nals from any reagents and solvents used in the analysis and any bias that
results from interactions between the analyte and the sample’s matrix. Both
the calibration blank and the reagent blank compensate for signals from Because we are considering a matrix effect
reagents and solvents. Any difference in their values is due to indeterminate of sorts, you might think that the method
of standard additions is one way to over-
errors in preparing and analyzing the standards. come this problem. Although the method
Unfortunately, neither a calibration blank nor a reagent blank can cor- of standard additions can compensate for
rect for a bias that results from an interaction between the analyte and the proportional determinate errors, it cannot
correct for a constant determinate error;
sample’s matrix. To be effective, the blank must include both the sample’s see Ellison, S. L. R.; Thompson, M. T.
matrix and the analyte and, consequently, it must be determined using the “Standard additions: myth and reality,”
sample itself. One approach is to measure the signal for samples of differ- Analyst, 2008, 133, 992–997.
180 Analytical Chemistry 2.1
ent size, and to determine the regression line for a plot of Ssamp versus the
amount of sample. The resulting y-intercept gives the signal in the absence
of sample, and is known as the total Youden blank.13 This is the true
blank correction. The regression line for the three samples in Table 5.3 is
Ssamp = 0.009844 × Wsamp + 0.185
giving a true blank correction of 0.185. As shown by the last row of Table
5.4, using this value to correct Ssamp gives identical values for the concentra-
tion of analyte in all three samples.
The use of the total Youden blank is not common in analytical work,
with most chemists relying on a calibration blank when using a calibra-
tion curve and a reagent blank when using a single-point standardization.
As long we can ignore any constant bias due to interactions between the
analyte and the sample’s matrix, which is often the case, the accuracy of an
analytical method will not suffer. It is a good idea, however, to check for
constant sources of error before relying on either a calibration blank or a
reagent blank.
5F.1 Excel
Let’s use Excel to fit the following straight-line model to the data in Ex-
ample 5.9.
y = b0 + b1 x
Enter the data into a spreadsheet, as shown in Figure 5.15. Depending
upon your needs, there are many ways that you can use Excel to complete
A B
a linear regression analysis. We will consider three approaches here.
1 Cstd Sstd
2 0.000 0.00 Use Excel’s Built-In Functions
3 0.100 12.36 If all you need are values for the slope, b1, and the y-intercept, b0, you can
4 0.200 24.83 use the following functions:
5 0.300 35.91 = intercept(known_y’s, known_x’s)
6 0.400 48.79
7 0.500 60.42 = slope(known_y’s, known_x’s)
Figure 5.15 Portion of a spread-
sheet containing data from Exam-
ple 5.9 (Cstd = Cstd; Sstd = Sstd).
13 Cardone, M. J. Anal. Chem. 1986, 58, 438–445.
Chapter 5 Standardizing Analytical Methods 181
where known_y’s is the range of cells that contain the signals (y), and
known_x’s is the range of cells that contain the concentrations (x). For ex-
ample, if you click on an empty cell and enter
= slope(B2:B7, A2:A7)
Excel returns exact calculation for the slope (120.705 714 3).
SUMMARY OUTPUT
Regression Statistics
Multiple R 0.99987244
R Square 0.9997449
Adjusted R Square 0.99968113
Standard Error 0.40329713
Observations 6
ANOVA
df SS MS F Significance F
Regression 1 2549.727156 2549.72716 15676.296 2.4405E-08
Residual 4 0.650594286 0.16264857
Total 5 2550.37775
Coefficients Standard Error t Stat P-value Lower 95% Upper 95% Lower 95.0% Upper 95.0%
Intercept 0.20857143 0.29188503 0.71456706 0.51436267 -0.60183133 1.01897419 -0.60183133 1.01897419
Cstd 120.705714 0.964064525 125.205016 2.4405E-08 118.029042 123.382387 118.029042 123.382387
Figure 5.16 Output from Excel’s Regression command in the Analysis ToolPak. See the text for a discussion of how to
interpret the information in these tables.
182 Analytical Chemistry 2.1
10 range from –1 to +1. The closer the correlation coefficient is to ±1, the bet-
r = 0.993
8 ter the model is at explaining the data. A correlation coefficient of 0 means
there is no relationship between x and y. In developing the calculations for
6
y
linear regression, we did not consider the correlation coefficient. There is
4 a reason for this. For most straight-line calibration curves the correlation
2
coefficient is very close to +1, typically 0.99 or better. There is a tendency,
however, to put too much faith in the correlation coefficient’s significance,
0
and to assume that an r greater than 0.99 means the linear regression model
0 2 4 6 8 10
x is appropriate. Figure 5.17 provides a useful counterexample. Although
Figure 5.17 Example of fitting a the regression line has a correlation coefficient of 0.993, the data clearly is
straight-line (in red) to curvilinear curvilinear. The take-home lesson here is simple: do not fall in love with
data (in blue). the correlation coefficient!
The second table in Figure 5.16 is entitled ANOVA, which stands for
analysis of variance. We will take a closer look at ANOVA in Chapter 14.
For now, it is sufficient to understand that this part of Excel’s summary
provides information on whether the linear regression model explains a
significant portion of the variation in the values of y. The value for F is the
result of an F-test of the following null and alternative hypotheses.
See Section 4F.2 and Section 4F.3 for a H0: the regression model does not explain the variation in y
review of the F-test.
HA: the regression model does explain the variation in y
The value in the column for Significance F is the probability for retaining
the null hypothesis. In this example, the probability is 2.5×10–6%, which
is strong evidence for accepting the regression model. As is the case with
the correlation coefficient, a small value for the probability is a likely out-
come for any calibration curve, even when the model is inappropriate. The
probability for retaining the null hypothesis for the data in Figure 5.17, for
example, is 9.0×10–7%.
The third table in Figure 5.16 provides a summary of the model itself.
The values for the model’s coefficients—the slope, b1, and the y-intercept,
b0—are identified as intercept and with your label for the x-axis data, which
in this example is Cstd. The standard deviations for the coefficients, sb0 and
sb1, are in the column labeled Standard error. The column t Stat and the
column P-value are for the following t-tests.
slope H0: b1 = 0, HA: b1 ≠ 0
See Section 4F.1 for a review of the t-test.
y-intercept H0: b0 = 0, HA: b0 ≠ 0
The results of these t-tests provide convincing evidence that the slope is not
zero, but there is no evidence that the y-intercept differs significantly from
zero. Also shown are the 95% confidence intervals for the slope and the
y-intercept (lower 95% and upper 95%).
Chapter 5 Standardizing Analytical Methods 183
A B C D E F
1 x y xy x^2 n= 6
2 0.000 0.00 =A2*B2 =A2^2 slope = =(F1*C8 - A8*B8)/(F1*D8-A8^2)
3 0.100 12.36 =A3*B3 =A3^2 y-int = =(B8-F2*A8)/F1
4 0.200 24.83 =A4*B4 =A4^2
5 0.300 35.91 =A5*B5 =A5^2
6 0.400 48.79 =A6*B6 =A6^2
7 0.500 60.42 =A7*B7 =A7^2
8
9 =sum(A2:A7) =sum(B2:B7) =sum(C2:C7) =sum(D2:D7) <--sums
Figure 5.18 Spreadsheet showing the formulas for calculating the slope and the y-intercept for the data in Example 5.9.
The shaded cells contain formulas that you must enter. Enter the formulas in cells C3 to C7, and cells D3 to D7. Next,
enter the formulas for cells A9 to D9. Finally, enter the formulas in cells F2 and F3. When you enter a formula, Excel
replaces it with the resulting calculation. The values in these cells should agree with the results in Example 5.9. You can
simplify the entering of formulas by copying and pasting. For example, enter the formula in cell C2. Select Edit: Copy,
click and drag your cursor over cells C3 to C7, and select Edit: Paste. Excel automatically updates the cell referencing.
70
60
50
40
y-axis
y-axis
30
20
10
0
Figure 5.19 Example of an Excel scatterplot showing
0 0.1 0.2 0.3 0.4 0.5 0.6
the data and a regression line. x-axis
Practice Exercise 5.6 concentration, CA, given the signal for a sample, Ssamp. Another limitation
is that Excel does not have a built-in function for a weighted linear regres-
Use Excel to complete the sion. You can, however, program a spreadsheet to handle these calculations.
regression analysis in Practice
Exercise 5.4. 5F.2 R
Click here to review your an- Let’s use R to fit the following straight-line model to the data in Example
swer to this exercise. 5.9.
y = b0 + b1 x
0.6
0.4
0.2
Residuals
0
0 0.1 0.2 0.3 0.4 0.5 0.6
Residuals
-0.2
-0.4
where y and x are the objects the objects our data. To access the results of
the regression analysis, we assign them to an object using the following
command
You can choose any name for the object
> model = lm(signal ~ conc) that contains the results of the regression
analysis.
where model is the name we assign to the object.
blue)and the regression line (in red). You can customize your plot
30
argument col allows you to select a color for the points or the line,
10
and the argument cex sets the size for the points. You can use the
command
0
> model=lm(signal~conc)
> summary(model)
Call:
lm(formula = signal ~ conc)
Residuals:
1 2 3 4 5 6
-0.20857 0.08086 0.48029 -0.51029 0.29914 -0.14143
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 0.2086 0.2919 0.715 0.514
conc 120.7057 0.9641 125.205 2.44e-08 ***
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
Figure 5.22 The summary of R’s regression analysis. See the
Residual standard error: 0.4033 on 4 degrees of freedom
text for a discussion of how to interpret the information in the Multiple R-Squared: 0.9997, Adjusted R-squared: 0.9997
output’s three sections. F-statistic: 1.568e+04 on 1 and 4 DF, p-value: 2.441e-08
The reason for including the argument The resulting output, shown in Figure 5.22, contains three sections.
which = 1 is not immediately obvious. The first section of R’s summary of the regression model lists the re-
When you use R’s plot command on an sidual errors. To examine a plot of the residual errors, use the command
object created by the lm command, the
default is to create four charts summa- > plot(model, which = 1)
rizing the model’s suitability. The first
of these charts is the residual plot; thus, which produces the result shown in Figure 5.23. Note that R plots the re-
which = 1 limits the output to this plot. siduals against the predicted (fitted) values of y instead of against the known
values of x. The choice of how to plot the residuals is not critical, as you can
see by comparing Figure 5.23 to Figure 5.20. The line in Figure 5.23 is a
smoothed fit of the residuals.
The second section of Figure 5.22 provides the model’s coefficients—
the slope, b1, and the y-intercept, b0—along with their respective standard
deviations (Std. Error). The column t value and the column Pr(>|t|) are for
the following t-tests.
Residuals vs Fitted
0.6
3
0.4
5
0.2
Residuals
0.0
-0.2
-0.4
4
-0.6
Fitted values
lm(signal ~ conc)
Chapter 5 Standardizing Analytical Methods 187
$`Standard Error`
[1] 0.002363588
$Confidence
[1] 0.006562373
$`Confidence Limits`
[1] 0.2346974 0.2478221
Figure 5.24 Output from R’s command for predicting the ana-
lyte’s concentration, CA, from the sample’s signal, Ssamp.
Call:
lm(formula = signal ~ conc, weights = w)
Residuals:
1 2 3 4 5 6
-2.223 2.571 3.676 -7.129 -1.413 -2.864
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 0.04446 0.08542 0.52 0.63
conc 122.64111 0.93590 131.04 2.03e-08 ***
---
Figure 5.25 The summary of R’s regression analysis for
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
a weighted linear regression. The types of information
Residual standard error: 4.639 on 4 degrees of freedom shown here is identical to that for the unweighted linear
Multiple R-Squared: 0.9998, Adjusted R-squared: 0.9997 regression in Figure 5.22.
F-statistic: 1.717e+04 on 1 and 4 DF, p-value: 2.034e-08
5G Key Terms
calibration curve external standard internal standard
method of standard
linear regression matrix matching
additions
multiple-point
normal calibration curve primary standard
standardization
reagent grade residual error secondary standard
single-point standard deviation about
serial dilution
standardization the regression
unweighted linear
total Youden blank weighted linear regression
regression
5H Chapter Summary
In a quantitative analysis we measure a signal, Stotal, and calculate the
amount of analyte, nA or CA, using one of the following equations.
S total = k A n A + S reag
S total = k A C A + S reag
To obtain an accurate result we must eliminate determinate errors that af-
fect the signal, Stotal, the method’s sensitivity, kA, and the signal due to the
reagents, Sreag.
To ensure that we accurately measure Stotal, we calibrate our equipment
and instruments. To calibrate a balance, for example, we use a standard
weight of known mass. The manufacturer of an instrument usually suggests
appropriate calibration standards and calibration methods.
To standardize an analytical method we determine its sensitivity. There
are several standardization strategies available to us, including external
standards, the method of standard addition, and internal standards. The
190 Analytical Chemistry 2.1
5I Problems
1. Suppose you use a serial dilution to prepare 100 mL each of a series of
standards with concentrations of 1.00×10–5, 1.00×10–4, 1.00×10–3,
and 1.00×10–2 M from a 0.100 M stock solution. Calculate the uncer-
tainty for each solution using a propagation of uncertainty, and com-
pare to the uncertainty if you prepare each solution as a single dilution
of the stock solution. You will find tolerances for different types of
volumetric glassware and digital pipets in Table 4.2 and Table 4.3. As-
sume that the uncertainty in the stock solution’s molarity is ±0.0002.
6. A standard sample contains 10.0 mg/L of analyte and 15.0 mg/L of in-
ternal standard. Analysis of the sample gives signals for the analyte and
the internal standard of 0.155 and 0.233 (arbitrary units), respectively.
Sufficient internal standard is added to a sample to make its concentra-
tion 15.0 mg/L. Analysis of the sample yields signals for the analyte
and the internal standard of 0.274 and 0.198, respectively. Report the
analyte’s concentration in the sample.
7. For each of the pair of calibration curves shown in Figure 5.26, select
the calibration curve that uses the more appropriate set of standards.
Briefly explain the reasons for your selections. The scales for the x-axis
and the y-axis are the same for each pair.
(a)
Signal
Signal
CA CA
(b)
Signal
Signal
CA CA
(c)
Signal
Signal
8. The following data are for a series of external standards of Cd2+ buffered
to a pH of 4.6.14
[Cd2+] (nM) 15.4 30.4 44.9 59.0 72.7 86.0
Sspike (nA) 4.8 11.4 18.2 26.6 32.3 37.7
(a) Use a linear regression analysis to determine the equation for the
calibration curve and report confidence intervals for the slope and
the y-intercept.
(b) Construct a plot of the residuals and comment on their significance.
At a pH of 3.7 the following data were recorded for the same set of
external standards.
[Cd2+] (nM) 15.4 30.4 44.9 59.0 72.7 86.0
Sspike (nA) 15.0 42.7 58.5 77.0 101 118
(c) How much more or less sensitive is this method at the lower pH?
(d) A single sample is buffered to a pH of 3.7 and analyzed for cadmium,
yielding a signal of 66.3 nA. Report the concentration of Cd2+ in
the sample and its 95% confidence interval.
(a) Determine the equation for the calibration curve using a linear
regression, and report confidence intervals for the slope and the y-
intercept. Average the replicate signals for each standard before you
complete the linear regression analysis.
(b) Based on your results explain why the authors concluded that the
internal standardization was inappropriate.
11. In Chapter 4 we used a paired t-test to compare two analytical methods Although this is a common approach for
comparing two analytical methods, it
that were used to analyze independently a series of samples of vari- does violate one of the requirements for
able composition. An alternative approach is to plot the results for one an unweighted linear regression—that in-
method versus the results for the other method. If the two methods determinate errors affect y only. Because
indeterminate errors affect both analytical
yield identical results, then the plot should have an expected slope, b1, methods, the result of an unweighted lin-
of 1.00 and an expected y-intercept, b0, of 0.0. We can use a t-test to ear regression is biased. More specifically,
compare the slope and the y-intercept from a linear regression to the ex- the regression underestimates the slope,
b1, and overestimates the y-intercept, b0.
pected values. The appropriate test statistic for the y-intercept is found We can minimize the effect of this bias by
by rearranging equation 5.23. placing the more precise analytical meth-
od on the x-axis, by using more samples
b0 - b0 b0 to increase the degrees of freedom, and
t exp = sb = s by using samples that uniformly cover the
0 b 0
range of concentrations.
Rearranging equation 5.22 gives the test statistic for the slope. For more information, see Miller, J. C.;
Miller, J. N. Statistics for Analytical Chem-
istry, 3rd ed. Ellis Horwood PTR Pren-
b1 - b1 1 - b1 tice-Hall: New York, 1993. Alternative
t exp = sb = sb
1 1 approaches are found in Hartman, C.;
Smeyers-Verbeke, J.; Penninckx, W.; Mas-
Reevaluate the data in problem 25 from Chapter 4 using the same sart, D. L. Anal. Chim. Acta 1997, 338,
19–40, and Zwanziger, H. W.; Sârbu, C.
significance level as in the original problem. Anal. Chem. 1998, 70, 1277–1280.
12. Consider the following three data sets, each of which gives values of y
for the same values of x.
Data Set 1 Data Set 2 Data Set 3 These three data sets are taken from Ans-
combe, F. J. “Graphs in Statistical Analy-
x y1 y2 y3 sis,” Amer. Statis. 1973, 27, 17-21.
10.00 8.04 9.14 7.46
8.00 6.95 8.14 6.77
13.00 7.58 8.74 12.74
9.00 8.81 8.77 7.11
11.00 8.33 9.26 7.81
14.00 9.96 8.10 8.84
6.00 7.24 6.13 6.08
4.00 4.26 3.10 5.39
12.00 10.84 9.13 8.15
7.00 4.82 7.26 6.42
5.00 5.68 4.74 5.73
194 Analytical Chemistry 2.1
(a) An unweighted linear regression analysis for the three data sets gives
nearly identical results. To three significant figures, each data set
has a slope of 0.500 and a y-intercept of 3.00. The standard devia-
tions in the slope and the y-intercept are 0.118 and 1.125 for each
data set. All three standard deviations about the regression are 1.24.
Based on these results for a linear regression analysis, comment on
the similarity of the data sets.
(b) Complete a linear regression analysis for each data set and verify
that the results from part (a) are correct. Construct a residual plot
for each data set. Do these plots change your conclusion from part
(a)? Explain.
(c) Plot each data set along with the regression line and comment on
your results.
(d) Data set 3 appears to contain an outlier. Remove the apparent out-
lier and reanalyze the data using a linear regression. Comment on
your result.
(e) Briefly comment on the importance of visually examining your
data.
rewriting it as
0 = kAC A Vo
& Vstd 0
V f + k A # C std V f
which is in the form of the linear equation
y = y-intercept + slope × x
where y is Sspike and x is Cstd × Vstd/Vf. The slope of the line, therefore,
is kA, and the y-intercept is kACAVo/Vf. The x-intercept is the value of x
when y is zero, or
0 = kAC V f + k A # " x-intercept ,
A Vo
k A C A Vo V f
x-intercept =- =- CVA Vo
kA f
n n
/x y i i = 4.110 # 10 -3 /x 2
i = 1.378 # 10 -4
i=1 i=1
When we substitute these values into equation 5.17 and equation 5.18,
we find that the slope and the y-intercept are
6 # (4.110 # 10 -3) - (2.371 # 10 -2) # (0.710)
b1 = = 29.57
6 # (1.378 # 10 -4) - (2.371 # 10 -2) 2
0.710 - 29.57 # (2.371 # 10 -2)
b0 = = 0.0015
6
and that the regression equation is
Sstd = 29.57 × Cstd + 0.0015
To calculate the 95% confidence intervals, we first need to determine
the standard deviation about the regression. The following table helps us
organize the calculation.
xi yi Vy i (y i - Vy i) 2
0.000 0.00 0.0015 2.250×10–6
1.55×10–3 0.050 0.0473 7.110×10–6
3.16×10–3 0.093 0.0949 3.768×10–6
Chapter 5 Standardizing Analytical Methods 197
n = 3.880 # 10 -3 M ! 0.13 # 10 -3 M
Click here to return to the chapter.
xi yi Vy i y i - Vy i
0.000 0.00 0.0015 –0.0015
0.010 1.55×10 –3 0.050 0.0473 0.0027
3.16×10–3 0.093 0.0949 –0.0019
4.74×10–3 0.143 0.1417 0.0013
residual error
0.000
6.34×10–3 0.188 0.1890 –0.0010
7.92×10–3 0.236 0.2357 0.0003
-0.010
Figure 5.27 shows a plot of the resulting residual errors. The residual er-
0.000 0.002 0.004 0.006 0.008 rors appear random, although they do alternate in sign, and that do not
CA show any significant dependence on the analyte’s concentration. Taken
Figure 5.27 Plot of the residual errors for together, these observations suggest that our regression model is appro-
the data in Practice Exercise 5.5. priate.
Click here to return to the chapter
SUMMARY OUTPUT
Regression Statistics
Multiple R 0.99979366
R Square 0.99958737
Adjusted R Square
0.99948421
Standard Error 0.00199602
Observations 6
ANOVA
df SS MS F Significance F
Regression 1 0.0386054 0.0386054 9689.9103 6.3858E-08
Residual 4 1.5936E-05 3.9841E-06
Total 5 0.03862133
Coefficients Standard Error t Stat P-value Lower 95% Upper 95% Lower 95.0% Upper 95.0%
Intercept 0.00139272 0.00144059 0.96677158 0.38840479 -0.00260699 0.00539242 -0.00260699 0.00539242
Cstd 29.5927329 0.30062507 98.437342 6.3858E-08 28.7580639 30.4274019 28.7580639 30.4274019
Figure 5.28 Excel’s summary of the regression results for Practice Exercise 5.6.
Chapter 5 Standardizing Analytical Methods 199
(0.114 - 0.1183) 2
s C = 1.996 # 10 -3
29.59
1+1+
3
A
6 (29.59) 2 # (4.408 # 10 -5)
= 4.772 # 10 -5
and the 95% confidence interval is
n = C A ! ts C = 3.80 # 10 -3 ! " 2.78 # (4.772 # 10 -5) ,
A
n = 3.80 # 10 -3 M ! 0.13 # 10 -3 M
Click here to return to the chapter
Call:
lm(formula = signal ~ conc)
Residuals:
1 2 3 4 5 6
-0.0013927 0.0027385 -0.0019058 0.0013377 -0.0010106 0.0002328
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 0.001393 0.001441 0.967 0.388
conc 29.592733 0.300625 98.437 6.39e-08 ***
---
Signif. codes: 0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1
$`Standard Error`
[1] 4.771723e-05
$Confidence
[1] 0.0001324843
Figure 5.29 R session for completing
$`Confidence Limits`
[1] 0.003672750 0.003937719 Practice Exercise 5.7.
200 Analytical Chemistry 2.1