Basic Business Statistics: Concepts and Applications
Chapter 1. Defining and Collecting Data
1.46 Visit the official website of a statistical software package that your class uses. Identify three
features that would help a business analyst summarize, visualize, or clean data.
1.47 A bank surveys 220 of its 12,000 small-business customers about satisfaction with online
banking. Of the 220 respondents, 154 say they are satisfied. Identify the population,
sample, one parameter, and the corresponding statistic.
1.48 A news organization reports a poll of 1,500 adults about the most important economic
issue. State one categorical variable and one numerical variable that could be collected in
such a poll.
1.49 A consulting firm surveys 1,200 CEOs about risks related to data security. Give one possible
coverage error, one possible nonresponse error, and one possible measurement error in this
survey.
1.50 A community survey records age, employment status, annual household income, number
of household vehicles, and home ownership status. Classify each variable as categorical or
numerical; for numerical variables, indicate discrete or continuous.
1.51 Design two survey questions for an alumni association: one that produces a categorical
response and one that produces a numerical response.
1.52 A university has 9,600 employees. A random sample of 2,400 employees is asked about
awareness of retirement-planning rules. Among them, 1,020 report awareness of at least
one rule. Identify the population, sample, parameter, and statistic.
1.53 A marketer plans a survey about social-media use. For each item, state whether the variable
is categorical or numerical: main platform used, number of visits per day, minutes spent
per visit, device used most often, and whether the respondent follows brands online.
1
Chapter 2. Organizing and Visualizing Variables
2.88 A campus bookstore estimates the following revenue shares for a textbook: publisher 64%,
bookstore operations 21%, author royalty 12%, and shipping/other 3%. Construct a per-
centage summary table and list the categories in Pareto order.
2.89 The table gives movie releases by source type in one year.
Source type Number
Original 310
Book or story 120
Real event 90
Sequel or spin-off 55
Game or toy 25
Compute the percentage distribution and identify the largest category.
2.90 A survey of 500 business marketers gives: centralized team 210, small team 145, no formal
team 85, outsourced 60. Prepare a relative-frequency distribution and name a suitable
chart.
2.91 A restaurant records entree orders: beef 155, chicken 110, fish 95, pasta 70, vegetarian 45,
and other 25. Construct the percentage distribution and state the top two categories.
2.92 Dessert orders by gender are shown below.
Gender Dessert No dessert
Female 38 62
Male 30 70
Compute row percentages and compare dessert ordering by gender.
2.93 A retailer has sales counts by region and channel.
Region Online Store
North 80 120
South 95 105
West 110 90
Compute the column totals and the grand total.
2.94 For the waiting times 12, 15, 17, 18, 22, 24, 26, 29, 31, 33, 35, 38, 41, 44, 49, construct a stem-
and-leaf display using tens as stems.
2.95 Using the waiting times in Problem 2.94, form a frequency distribution using classes 10–19,
20–29, 30–39, and 40–49.
2
2.96 For the frequency distribution in Problem 2.95, compute the cumulative frequency and
cumulative percentage distribution.
2.97 Monthly sales revenue, in thousands of dollars, is 44, 47, 49, 53, 56, 59, 62, 66, 68, 71, 75, 79.
Which graph should be used to display the pattern over time? Describe the main pattern.
2.98 Advertising spending X and sales Y for six stores are (2, 41), (3, 45), (4, 50), (5, 53), (6, 58), (7, 62).
Which graph displays the relationship? Describe the association.
2.99 A chart compares two product defect rates of 4.8% and 5.2% using a vertical axis from
4.7% to 5.3%. Explain the potential pitfall.
2.100 An A/B test for a website button gives the table below.
Button Clicked Did not click
Old button 420 1580
New button 510 1490
Compute the click-through percentage for each button.
2.101 A customer survey asks respondents to rate service as poor, fair, good, very good, or
excellent. What type of variable is this? Name a suitable graph.
2.102 A store records daily customers for 10 days: 93, 101, 98, 110, 117, 115, 121, 125, 130, 136.
Construct a frequency distribution using classes 90–99, 100–109, 110–119, 120–129, 130–
139.
2.103 For the data in Problem 2.102, which chart is more appropriate for showing the distribution:
a bar chart or histogram? Explain.
2.104 A dashboard must summarize region, product category, sales dollars, and number of com-
plaints. Suggest two visual summaries that combine more than one variable.
2.105 A dataset contains 1,000 transactions. Write a filtering rule to display only online purchases
above $100 made by returning customers.
2.106 The counts of customer ratings are: 1 star 8, 2 stars 12, 3 stars 45, 4 stars 90, 5 stars 145.
Compute the percentage with at least 4 stars.
2.107 Class project: each student records commute mode and commute time. Identify one cate-
gorical variable, one numerical variable, and a graph for each.
2.108 Class project: cross-classify students by commute mode (car, bus, walk) and year (first,
second, third, fourth). What table and graph would you use?
2.109 Refer to the A/B test in Problem 2.100. If the goal is to compare click and no-click
percentages for the two buttons, what graph should be used and what conclusion follows?
3
Chapter 3. Numerical Descriptive Measures
3.63 For the salary data, in thousands of dollars, 82, 91, 78, 110, 95, 88, 104, 99, compute the
mean, median, range, sample variance, and sample standard deviation.
3.64 Application approval times, in days, are 12, 15, 14, 19, 11, 13, 17, 16, 21, 18. Find the five-
number summary and interquartile range.
3.65 Customer satisfaction ratings are 7, 8, 6, 9, 5, 8, 7, 6, 10, 9, 8, 7. Compute the mean, median,
mode, and sample standard deviation.
3.66 Annual returns, in percent, are −3, 4, 6, 2, 9, −1, 5. Compute the arithmetic mean and
geometric mean return.
3.67 A test has mean 60 and standard deviation 8. Compute the Z scores for scores 72 and 55.
3.68 A bell-shaped distribution has mean 50 and standard deviation 5. Use the empirical rule to
describe the intervals containing approximately 68%, 95%, and 99.7% of the observations.
3.69 A dataset has mean 100 and standard deviation 15. Use Chebyshev’s theorem for the
interval 55 to 145.
3.70 For five stores, advertising spending X and sales Y are (2, 50), (4, 55), (5, 58), (7, 62), (8, 67).
Compute the sample covariance and correlation coefficient.
3.71 Department A has mean processing time 24 minutes and standard deviation 3 minutes.
Department B has mean 40 minutes and standard deviation 4 minutes. Compare relative
variability using the coefficient of variation.
3.72 In a sample of household donations, the mean is $86 and the median is $45. What does
this suggest about the shape of the distribution?
3.73 A training program has scores and frequencies: score 60 (4 employees), 70 (6), 80 (12), 90
(8). Compute the weighted mean.
3.74 For the data 18, 19, 20, 21, 22, 22, 23, 24, 25, 40, use the 1.5IQR rule to determine whether
there is an outlier.
3.75 Explain the difference between skewness and kurtosis in descriptive statistics.
3.76 An applicant scores 83 on Exam A, which has mean 75 and standard deviation 10, and 72
on Exam B, which has mean 65 and standard deviation 5. On which exam is the applicant
relatively stronger?
3.77 Commuting times, in minutes, are 31, 28, 45, 52, 39, 41, 36, 33, 47, 58, 29, 40. Compute the
mean and sample standard deviation, and state whether the mean or median is more
resistant to extremes.
4
3.78 Credit scores for two cities are: City A 710, 695, 720, 740, 705 and City B 660, 690, 675, 700, 685.
Compare their means and standard deviations.
3.79 Two assets have annual returns: Asset 1 4, 6, −2, 8, 3 and Asset 2 3, 5, 0, 7, 4. Compute the
covariance and correlation.
3.80 For six beers, alcohol percentage and calories are (4.2, 95), (4.5, 110), (5.0, 140), (5.2, 145), (6.0, 180), (6.5, 210
Compute the correlation coefficient.
5
Chapter 4. Basic Probability
4.62 In a survey of 1,000 investors, 620 prefer hybrid advice. Among 500 millennials, 330
prefer hybrid advice; among 500 baby boomers, 290 prefer hybrid advice. Construct the
contingency table and compute P (hybrid).
4.63 Using the table in Problem 4.62, compute P (hybrid | millennial) and determine whether
group and advice preference appear independent.
4.64 A website-builder survey classifies 300 firms by business size and whether they plan a
redesign.
Size Redesign No redesign
Small 72 78
Large 96 54
Compute P (large and redesign) and P (redesign | large).
4.65 A marketing email is opened by 30% of customers. If opened, the probability of purchase
is 0.18; if not opened, the probability of purchase is 0.04. What is the overall probability
of purchase?
4.66 A fraud screen flags 95% of fraudulent transactions and 6% of legitimate transactions.
Fraud occurs in 2% of transactions. If a transaction is flagged, find the probability that it
is fraudulent.
6
Chapter 5. Discrete Probability Distributions
5.33 An event insurer charges a $3, 500 premium for a contest that pays $1, 000, 000 with prob-
ability 0.002. Compute the insurer’s expected profit.
5.34 Suppose the probability that the stock market rises in a year is 0.62. Assuming indepen-
dence, find the probability that it rises in exactly 4 of the next 5 years and in none of the
next 5 years.
5.35 Seventeen percent of a population is smartphone-dependent. For a sample of 10, find the
probability that exactly 3, at least 3, and at most 6 are smartphone-dependent.
5.36 An indicator is correct in 11 or more of 14 years. Compute P (X ≥ 11) for X ∼ Bin(14, 0.50)
and for X ∼ Bin(14, 0.75).
5.37 Suppose 75% of medical bills contain an error. For a sample of 10 bills, compute P (X = 0),
P (X = 5), P (X > 5), the mean, and the standard deviation.
5.38 Repeat Problem 5.37 if the error rate is reduced to 40%.
5.39 For a sample of 10 social log-ins, compute the probability that more than 5 use Facebook
when p = 0.45, more than 5 use Google when p = 0.26, and none use Facebook when
p = 0.45.
5.40 A bank complaint category has probability 0.46. In a sample of 20 complaints, compute
the probability of exactly 10 and at least 12 in that category.
5.41 In the same complaint system, suppose 24% of complaints are referred to investigation. In
a sample of 20, compute the probability of no referrals and at most 3 referrals.
5.42 An indicator is correct in 37 or more of 42 years. Compute P (X ≥ 37) for p = 0.50, 0.70,
and 0.90.
5.43 A spurious indicator is correct in 38 of 50 years. If it has no predictive value, so that
p = 0.50, compute P (X ≥ 38).
5.44 The number of service calls arriving in 10 minutes follows a Poisson distribution with mean
3.2. Compute P (X = 0), P (X ≤ 2), P (X > 5), the mean, and the standard deviation.
7
Chapter 6. The Normal Distribution and Other Continuous Dis-
tributions
6.33 Ball-bearing diameters are normally distributed with mean 0.753 inch and standard devi-
ation 0.004 inch. Find the probability that a bearing is between 0.750 and 0.760, above
0.760, and below 0.740.
6.34 Bottle fills are normally distributed with mean 2.00 liters and standard deviation 0.05 liter.
Find P (X < 1.90) and P (1.95 < X < 2.05).
6.35 Online spending before visiting a store is normally distributed with mean $250 and standard
deviation $40. Compute P (X < 210), P (270 < X < 300), and the 90th percentile.
6.36 Assume instead that spending is uniformly distributed between $200 and $300. Find P (X <
210), P (270 < X < 300), and the mean and standard deviation.
6.37 A data set has mean 142, median 141, and a normal probability plot that is nearly linear
except for one high value. What does this suggest about normality?
6.38 A job-aptitude score is normally distributed with mean 70 and standard deviation 8. What
score cuts off the top 10%?
6.39 A call duration is normally distributed with mean 6.5 minutes and standard deviation 1.2
minutes. Find P (5 < X < 8) and P (X > 9).
6.40 Annual stock returns are modeled as normal with mean 0.08 and standard deviation 0.18.
Find the probability of a loss greater than 20%.
6.41 Monthly intern pay is normally distributed with mean $3, 500 and standard deviation $400.
Find the probability that a randomly selected intern earns between $3, 000 and $4, 200.
6.42 Explain how a normal probability plot can be used to evaluate whether a numerical variable
is approximately normally distributed.
8
Chapter 7. Sampling Distributions
7.26 For the bearing process in Problem 6.33, a random sample of 25 bearings is selected. Find
P (X̄ > 0.755).
7.27 Bottle fills have mean 2.00 liters and standard deviation 0.05. For n = 36, find P (1.99 <
X̄ < 2.01).
7.28 The amount of juice from an orange has mean 4.8 ounces and standard deviation 0.8 ounce.
For n = 64, find P (X̄ > 5.0).
7.29 In Problem 7.28, suppose the mean is 5.0 ounces. Find the symmetric interval around the
mean that contains 70% of sample means for n = 64.
7.30 If the population proportion favoring a new package is 0.42 and n = 200, approximate
P (p̂ > 0.50).
7.31 State two reasons why the sampling distribution of the sample mean may be approximately
normal even if the population is not normal.
7.32 A population has N = 1, 000, standard deviation σ = 20, and a sample size n = 100 is
selected without replacement. Compute the finite-population corrected standard error of
X̄.
7.33 In a class project, each student tosses a coin 10 times and records the sample proportion
of heads. What are the mean and standard error of the sampling distribution of p̂ if the
coin is fair?
7.34 The number of cars waiting in a line has mean 6 and standard deviation 3. For samples of
n = 36 time periods, find P (X̄ < 5).
7.35 City credit scores have mean 690 and standard deviation 40. For a random sample of 64
cities, find the probability that the sample mean exceeds 700.
9
Chapter 8. Confidence Interval Estimation
8.54 A survey of 1,000 Internet users finds device ownership counts: PC/laptop 840, smartphone
910, tablet 500, and smart watch 100. Construct 95% confidence intervals for the four
population proportions.
8.55 A grocery survey of 731 smartphone owners finds that 358 access coupons, 355 look up
recipes, 234 read reviews, and 154 locate in-store items. Construct 95% confidence intervals
for each population proportion.
8.56 A sample of 40 residents gives mean weekly viewing time 51 hours and standard deviation
3.5 hours. Also, 32 have HD service. Construct 95% confidence intervals for the mean
viewing time and the HD proportion.
8.57 A study of 50 firms finds mean time spent on a ransomware incident 42 hours with s = 8
hours. Thirteen firms lost customers. Construct a 99% CI for the mean time and a 95%
CI for the customer-loss proportion.
8.58 A sample of 25 managers has mean absenteeism 6.2 days and s = 7.3 days; 13 cite stress.
Construct 95% CIs for the mean and proportion. Also find the sample sizes needed for
estimating a mean within 1.5 days with σ = 8 at 95%, and a proportion within 0.075 at
90% with no prior estimate.
8.59 For 100 organizations, mean turnover cost is $12, 500 with s = $1, 000; 30 have talent-
development programs. Construct 95% CIs for the mean cost and the population propor-
tion.
8.60 A process has known σ = 12. For n = 64 and x̄ = 74, construct a 95% CI for µ.
8.61 How large a sample is needed to estimate a population mean within E = 2.5 units at 95%
confidence if σ = 15?
8.62 How large a sample is needed to estimate a population proportion within 0.04 at 95%
confidence if no previous estimate is available?
8.63 Explain what happens to the margin of error when the sample size is quadrupled, all else
equal.
8.64 For n = 18, x̄ = 32, and s = 6, construct a 90% CI for the mean.
8.65 In a sample of 400 customers, 260 are satisfied. Construct a 99% CI for the population
satisfaction proportion.
8.66 For 100 call-center contacts, mean answer time is 5.2 minutes with s = 1.1; 62 calls are
resolved on the first contact. Construct 95% CIs for the mean and proportion.
8.67 For 170 shingle samples, the mean granule loss is 0.35 gram with s = 0.05 gram. Construct
a 95% CI for the population mean granule loss.
10
8.68 Compare a 90%, 95%, and 99% confidence interval computed from the same sample. Which
is widest and why?
8.69 Using the call-center results in Problem 8.66, write a one-sentence managerial interpretation
of the interval for the first-contact resolution proportion.
11
Chapter 9. Fundamentals of Hypothesis Testing: One-Sample Tests
9.69 A web-design test has H0 : p = 0.50 versus H1 : p ̸= 0.50 and p-value 0.20. At α = 0.05,
state the decision and interpretation.
9.70 A model predicts bankruptcy for 28 of 400 firms. Test H0 : p = 0.05 versus H1 : p > 0.05
at α = 0.05.
9.71 A filling machine is set to 500 grams. A sample of 64 packages has x̄ = 503 grams. Assume
σ = 18 grams. Test H0 : µ = 500 versus H1 : µ ̸= 500 at α = 0.05.
9.72 A sample of 25 service times has x̄ = 71 seconds and s = 9 seconds. Test H0 : µ = 68
versus H1 : µ > 68 at α = 0.05.
9.73 A store claims the mean checkout time is less than 15 minutes. A sample of 36 customers
has x̄ = 14.2 and s = 2.1. Test the claim at α = 0.05.
9.74 A bank branch wants mean waiting time below 5 minutes. In a sample of 40 visits, x̄ = 4.7
and s = 1.2. Test at α = 0.05.
9.75 A hypothesis test yields p-value 0.047. State the decision at α = 0.05 and at α = 0.01.
9.76 In a sample of 250 customers, 142 prefer a new package. Test H0 : p = 0.52 versus
H1 : p ̸= 0.52 at α = 0.05.
9.77 For 170 shingle samples, x̄ = 0.35 gram and s = 0.05. Test whether mean granule loss is
less than 0.36 gram at α = 0.05.
9.78 A product team expects more than 60% satisfaction. In a sample of 300, 183 are satisfied.
Test H0 : p = 0.60 versus H1 : p > 0.60 at α = 0.05.
9.79 Summarize the results of Problems 9.76–9.78 in terms of statistical significance and practical
significance.
12
Chapter 10. Two-Sample Tests
10.58 Two job titles have salary summaries: group 1 n = 109, x̄ = 98445, s = 24120; group 2
n = 39, x̄ = 79749, s = 28086. Test equality of variances, then test whether group 1 has a
larger mean salary.
10.59 Private and public colleges have average student debt summaries: private n = 100, x̄ =
32500, s = 6200; public n = 100, x̄ = 27800, s = 5400. Test whether the mean debts differ
at α = 0.05.
10.60 Eight employees are timed before and after a training program. Before: 12, 15, 14, 18, 20, 16, 17, 19;
after: 10, 13, 13, 15, 16, 14, 15, 16. Use a paired t test for whether mean time decreased.
10.61 In two independent samples, 78 of 200 customers in region A and 55 of 180 in region B
purchase a warranty. Test whether the proportions differ at α = 0.05.
10.62 Two production lines have sample standard deviations 5.6 and 3.8 based on sample sizes
30 and 28. Test equality of variances at α = 0.05.
10.63 For the warranty data in Problem 10.61, construct a 95% confidence interval for pA − pB
using the unpooled standard error.
10.64 Two call centers have summaries: center A n = 35, x̄ = 48.2, s = 6.1; center B n = 32,
x̄ = 44.9, s = 7.8. Test whether center A has a larger mean service score.
10.65 For Problem 10.64, compute Cohen’s standardized mean difference using the pooled stan-
dard deviation.
10.66 Two call centers have mean answer times: A n = 100, x̄ = 5.2, s = 1.1; B n = 90, x̄ = 5.6,
s = 1.3. Test whether A has a smaller mean answer time.
10.67 Two shingle types have granule-loss summaries: type A n = 170, x̄ = 0.35, s = 0.05; type
B n = 160, x̄ = 0.38, s = 0.06. Test whether the means differ.
10.68 A paired experiment measures quality scores before and after a process change. Before:
64, 70, 75, 68, 72, 77, 69, 73, 71, 76; after: 67, 72, 78, 69, 75, 80, 71, 76, 73, 78. Test whether the
mean score increased.
10.69 Among 250 younger users, 74 use a mobile wallet; among 300 older users, 92 use one. Test
whether the proportions differ.
10.70 Write a short report summarizing Problems 10.67 and 10.68 for a production manager.
13
Chapter 11. Analysis of Variance
11.36 A parachute manufacturer tests fabric strength from three suppliers. Supplier A: 82, 85, 84, 86, 83;
B: 79, 78, 80, 77, 81; C: 88, 90, 87, 91, 89. Use one-way ANOVA at α = 0.05.
11.37 Medical-wire tensile ratios are measured for three reduction angles. Angle 1: 14.2, 14.5, 14.1, 14.4;
Angle 2: 13.8, 13.9, 14.0, 13.7; Angle 3: 15.1, 15.3, 14.9, 15.2. Perform one-way ANOVA.
11.38 Yarn breaking strengths are recorded at three air-jet pressures. 30 psi: 25.5, 24.9, 26.1, 24.7, 24.2, 23.6;
40 psi: 24.8, 23.7, 24.4, 23.6, 23.3, 21.4; 50 psi: 23.2, 23.7, 22.7, 22.6, 22.8, 24.9. Test equality
of mean strengths.
11.39 For the yarn experiment, suppose a second factor, side of nozzle, is added. Explain what
questions a two-way ANOVA can answer that the one-way ANOVA in Problem 11.38
cannot.
11.40 A hotel compares room-service delivery differences by breakfast type. Continental observa-
tions are 1.2, 2.1, 3.3, 4.4, 3.4, 5.3, 2.2, 1.0, 5.4, 1.4, −2.5, 3.0, −0.2, 1.2, 1.2, 0.7, −1.3, 0.2, −0.5, 3.8
and American observations are 4.4, 1.1, 4.8, 7.1, 6.7, 5.6, 9.5, 4.1, 7.9, 9.4, 6.0, 2.3, 4.2, 3.8, 5.5, 1.8, 5.1, 4.2, 4.9,
Test the effect of breakfast type.
11.41 Repeat Problem 11.40 if the second set of Continental observations is changed to −0.5, 5.0, 1.8, 3.2, 3.2, 2.7, 0
11.42 A pet-food team studies two factors: piece size (fine/current) and fill height (low/current),
with 20 cans per combination. State the null hypotheses for a two-way ANOVA and the
managerial meaning of an interaction.
14
Chapter 12. Chi-Square and Nonparametric Tests
12.56 A pizza survey gives this table.
Gender Pizza Hut Other
Female 4 13
Male 6 12
Use a chi-square test for a difference in proportions at α = 0.05.
12.57 A social-media tool survey cross-classifies tool use (rows: Tool A, Tool B, Tool C) by
business focus (columns: B2B, B2C): (310, 355), (420, 275), (250, 300). Test independence.
12.58 Digital-transformation progress is classified as early, middle, or advanced, and industry
sector as service or manufacturing. Counts are early (48, 52), middle (60, 40), advanced
(35, 65). Test independence at 0.05.
12.59 Customer trust scores for two advertising formats are: Format A 12, 15, 17, 19, 21, 22 and
Format B 18, 20, 23, 25, 27, 29. Use the Wilcoxon rank sum test to compare distributions.
15
Chapter 13. Simple Linear Regression
13.73 For the data relating Tomatometer rating X to opening weekend receipts Y , with obser-
vations (16, 7.8, 61, 6.1, 80, 28.4, 76, 11.3, 93, 24.8, 52, 13.3, 15, 2.9, 61, 4.0, 20, 5.1, 38, 1.9), fit
the least-squares line, predict Y at X = 55, and report r2 .
13.74 For the data relating number of cases delivered X to delivery time Y , with observations
(52, 32.1, 64, 34.8, 73, 36.2, 85, 37.8, 95, 37.8, 103, 39.7, 116, 38.5, 121, 41.9, 143, 44.2, 157, 47.1, 161, 43.0, 184, 4
fit the least-squares line, predict Y at X = 150, and report r2 .
13.75 For the data relating diameter at breast height X to redwood height Y , with obser-
vations (12, 160, 18, 195, 20, 210, 24, 235, 25, 240, 27, 255, 30, 270, 33, 285, 35, 295, 38, 310), fit
the least-squares line, predict Y at X = 25, and report r2 .
13.76 For the data relating living space X to asking price Y , with observations (1200, 280, 1450, 315, 1600, 340, 1800
fit the least-squares line, predict Y at X = 2000, and report r2 .
13.77 For the data relating commute distance X to commute time Y , with observations (3, 12, 5, 17, 7, 20, 8, 23, 12,
fit the least-squares line, predict Y at X = 10, and report r2 .
13.78 For the data relating credit score X to loan rate Y , with observations (620, 8.5, 650, 7.9, 680, 7.2, 700, 6.8, 720
fit the least-squares line, predict Y at X = 700, and report r2 .
13.79 For the data relating efficiency ratio X to ROATCE Y , with observations (45, 18, 50, 16, 55, 14, 60, 13, 65, 11,
fit the least-squares line, predict Y at X = 60, and report r2 .
13.80 For the data relating alcohol percentage X to calories Y , with observations (4.2, 95, 4.5, 110, 5.0, 140, 5.2, 145
fit the least-squares line, predict Y at X = 5.5, and report r2 .
13.81 For the data relating advertising X to sales Y , with observations (2, 41, 3, 45, 4, 50, 5, 53, 6, 58, 7, 62, 8, 67),
fit the least-squares line, predict Y at X = 6, and report r2 .
13.82 For the data relating training hours X to productivity Y , with observations (1, 50, 2, 53, 3, 58, 4, 60, 5, 64, 6, 6
fit the least-squares line, predict Y at X = 5, and report r2 .
13.83 For the data relating temperature X to energy demand Y , with observations (50, 110, 55, 105, 60, 100, 65, 98,
fit the least-squares line, predict Y at X = 72, and report r2 .
13.84 For the data relating store size X to monthly revenue Y , with observations (8, 90, 10, 110, 12, 120, 15, 150, 18,
fit the least-squares line, predict Y at X = 16, and report r2 .
13.85 For the data relating customer age X to monthly spending Y , with observations (22, 80, 25, 90, 30, 105, 35, 12
fit the least-squares line, predict Y at X = 42, and report r2 .
13.86 For the data relating CEO compensation X to stock return Y , with observations (4, −2, 5, 1, 6, 3, 8, 4, 10, 5, 12
fit the least-squares line, predict Y at X = 10, and report r2 .
16
13.87 For the data relating market return X to GE stock return Y , with observations (−10, −12, −5, −6, 0, 1, 5, 4, 1
fit the least-squares line, predict Y at X = 12, and report r2 .
17
Chapter 14. Introduction to Multiple Regression
14.71 Consider the regression equation Ŷ = −3.888 + 1.449X1 + 1.462X2 − 0.190X1 X2 . Compute
predictions for (X1 , X2 ) = (2, 2), (2, 7), (7, 2), (7, 7) and interpret the negative interaction.
14.72 A moving company predicts cost using cubic feet X1 and miles X2 . Data are (1200, 15, 280), (1500, 18, 330), (
Fit a multiple regression model.
14.73 Using the model from Problem 14.72, predict cost for a move with 2, 500 cubic feet and 26
miles.
14.74 A housing model uses living space X1 and number of bedrooms X2 to predict asking price.
Observations are (1200, 3, 280), (1450, 3, 315), (1600, 4, 340), (1800, 4, 380), (2100, 3, 430), (2400, 4, 500), (2600
Fit the model and report r2 .
14.75 For Problem 14.74, interpret the coefficient of living space.
14.76 For Problem 14.74, test the overall significance of the regression at α = 0.05.
14.77 A model for sales is Ŷ = 25+4.2X1 +1.8X2 , where X1 is advertising in thousands of dollars
and X2 is sales-force hours. Predict sales when X1 = 10 and X2 = 35.
14.78 Explain the difference between r2 and adjusted r2 in multiple regression.
14.79 A multiple regression has n = 40, k = 3 predictors, SSE = 480, and SST = 2400. Compute
r2 , adjusted r2 , and the overall F statistic.
14.80 A dummy-variable model is Ŷ = 60 + 5X + 12D, where D = 1 for premium customers and
0 otherwise. Interpret the coefficient of D.
14.81 A coffee shop models satisfaction using waiting time X1 , price rating X2 , and a drive-
through dummy variable D. Explain how an interaction term X1 D would be interpreted.
14.82 An extrusion process model is Ŷ = 2.40 + 0.015X1 − 0.020X2 + 0.006X3 , where Y is unit
density, X1 is temperature, X2 is speed, and X3 is pressure. Predict Y for X1 = 180,
X2 = 60, X3 = 120.
14.83 For the extrusion model in Problem 14.82, explain which coefficient suggests that increasing
a predictor decreases predicted density and how you would verify whether that effect is
statistically significant.
18