0% found this document useful (0 votes)
5 views23 pages

US Obesity and Data Analysis Insights

Uploaded by

赵雨枫
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views23 pages

US Obesity and Data Analysis Insights

Uploaded by

赵雨枫
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Homework section:

Chapter 1

7. According to the National Geographic obesity is rapidly becoming a serious problem in many parts of the world.
Suppose that a researcher hypothesizes that more than 25% of the US population is obese. To test the hypothesis, 1000
members of the US population will be used to collect data.

Solution
(a) The variable of interest is whether or not a person is obese
(b) Whether or not each member of the US population is obese or not obese
(c) Whether or not each member of the 1000 people sampled is obese or not obese
(d) Based on the sample of the 1000 people one could agree with or disagree with the hypothesis that more than 25% of the US
population is obese.
6. Classify each of the following quantitative variables as discrete or continuous:
(a) The number of times each week that a person goes for physical exercise. (a) Discrete: this is a counting variable
(b) The monthly wages (in dollars) of federal employees. (b) Continuous: money is measured as an amount
(c) The amount of snow falling per week at a ski area. (c) Continuous: snowfall is measured as an amount
(d) The number of California voters in a random sample of 100 voters who prefer raising taxes to cutting back on public services.
(d) Discrete: this is a counting variable

5. Classify the following data as quantitative or qualitative for the students in a statistics class:
(a) Ethnic classification of the student. (a) Qualitative: ethnicity is a category and is not measured in an amount
(b) Weight of the student. (b) Quantitative: weight is measured as an amount (pounds or kilograms usually)
(c) Number of textbooks used by the student during the current term. (c) Quantitative: textbooks are counted
(d) Number of hours of physical exercise each student gets per week\. (d) Quantitative: time is measured in hours, minutes, etc.
(e) Telephone number of the student. (e) Qualitative: telephone number is not measured or counted
(f) Student ID (or social security number). (f) Qualitative: student IDs are not measured or counted

6.

(a) What percent of the properties in the sample sold for more than $1 million?. (a) 11 + 7 + 5 = 23 and 23/45 = 0.5111. So,
approximately 51% of the properties sold for more than $1 million.
(b) If property taxes are assessed at 1.5% of the reselling price per year how many homes in the sample had property taxes of less
than $12,000 per year? (b) $12,000 / 0.015 = $800,000, so we need to count up the number of properties that sold for less than
$800,000. 3 + 10 + 4 = 17. So, 17 homes in the sample had property taxes of less than $12,000 per year.
(c) How would you describe the sample of home selling prices? (c) This graph does not fall nicely into one of our set phrases for
describing data. However, this data set does appear to be bimodal with several properties being sold between $400,000 to
$600,000 and several other properties being sold between $1 million and $1.2 million.

2. Hybrid cars use two types of engine/motor technologies: gasoline and electric. The electric motor, that takes some of the
load off the gasoline engine, is charged every time the brakes are applied (the energy created from the braking process is
transferred to charging the battery). As a result of the dual engine design, hybrid cars get considerably better gas mileage
than the standard gasoline engine vehicles. The graph below represents the best mileage ratings (in miles per gallon or
mpg) of several randomly selected hybrid vehicles on the road today (including experimental hybrid trucks and sport
utility vehicles).

Stem and Leaf graph of the best mileage ratings (in miles per gallon) of 45 hybrid vehicles
(a) n = 45 cars
(b) 26/45 = 0.5777. S. So, approximately 58% of the vehicles had a mileage rating between 40 mpg and 59 mpg (inclusive).
(c) 1 + 2 = 3 with a mileage rating of less than 20 miles per gallon.
(d) The distribution of mileage ratings is skewed to the LEFT.

Chapter 2

3. As information accessibility increases due to the internet, identity theft has been on a steady increase over the years. The data
below represents the percent of the population who fell victim to some form of identity theft over a seven-year period. Calculate
the MEDIAN and the MEAN percent of identity theft victims.

Identity Theft Victims for Years 2009 to 2015

Year 2009 2010 2011 2012 2013 2014 2015

% Identity Theft Victims 2% 4% 5% 7% 10% 15% 18%

Solution
MEAN=¯x=Σxn=617=8.714%
MEDIAN =7%

1. What does Chebyshev's Rule tell us about the amount of data that falls in the following intervals?

(a) The interval within 2.5 standard deviations to either side of the mean. (a) 1−1k2=1−12.52=0.84. So, at least 84% of the data
will fall in the interval ¯¯¯x±2.5s
(b) The interval within 1 standard deviation to either side of the mean. (b) 1−1k2=1−112=1−1=0. In order for Chebyshev’s Rule
to be useful, the value of k must be larger than 1.

6. At a local university it is known that student’s grade point averages (GPA’s) follow a mound shaped and symmetric
distribution with a mean of 2.89 and a standard deviation of 0.37.
(a) Since the data set is mound-shaped and symmetric, the Empirical Rule applies. So, we know that approximately 95% of the
data lies within 2 standard deviations of the mean.
¯¯¯x±2s⇒2.89±2(0.37)⇒(2.15,3.63)
(b) Find x by noting that since 68% of a mound-shaped and symmetric distribution lies within 1 standard deviation of the mean.
This leaves 16% for both tails of the distribution. Then, we need our unknown value of x to lie exactly 1 standard deviation above
the mean.
x=the sample mean+one standard deviation=2.89+(1)(0.37)=3.26

5. A building developer has the opportunity to purchase several acres of land that is divided up into several lots for
building single-family homes. She will buy the land provided that less than 30% of the lots will cost her more than $50,000
per lot. She only knows that the average sale price of all of the lots is $35,000 and the standard deviation of the sale price
is $8,000. Should she buy the lots?
Number line representing Chebychev’s Rule.

If less than 30% of the data is above 50,000, he will buy. We compute k: ¯¯¯x+ks=50,000⇒35,000+k(8000)=50,000⇒k=1.88=
So, 1k2=11.882=0.28. Then, according to Chebyshev's Rule 28% or fewer of the lots of land will cost more than $50,000.
Therefore, the developer should buy the land.
3. It is known that a particular hybrid vehicle has a maximum mileage rating of 70 mpg and a minimum mileage rating of
25 mpg. Estimate the standard deviation of mileage ratings for this hybrid vehicle.
Recall that: σ ≈Range6=70−256=7.5 mpg. So, the standard deviation is approximately 7.5 mpg.

Chapter 3

5. In a large midwestern university, a survey was taken that compared the class of an undergraduate to whether or not
they favor the US government s approach to containment of the spread of nuclear weapons around the world. The
following data is a summary of the survey taken by questioning 400 undergraduates at the university, 100 from each class.

University Opinion/Nuclear Weapons Data Table

Freshman Sophomore Junior Senior Totals

Favor 54 41 32 23 150

Oppose 46 59 68 77 250

Totals 100 100 100 100 400

If a student in the survey is selected at random, then find the probability that the student

(a) favors the US government's approach to nuclear weapons. (a) P(favor)=150/400=3/8=0.375


(b) is a freshman, given that the student favors the US government's approach. (b) P(freshman | favor)=54/150=9/25=0.36,
because there are 150 in the reduced sample space
(c) is a senior, given that the student favors the US government's approach. (c) P(senior | favor)=23/150=0.1533, because there
are 150 in the reduced sample space
(d) favors the US government's approach, given that the student is a junior or a senior. (d) P(favor | junior or
senior)=(32+23)/200=55/200=11/40=0.275

Chapter 4

3. If a binomial distribution has n=10, p=0.15 and x = the number of successes, then fill in the ? mark for each of the
following binomial probabilities:

(a) P(3)=10C?(.15)3 (.85)7. (a) 3


(b) P(7)=10C7 (.15)? (.85)3. (b) 7
(c) P(0)=10C0 (?)0 (?)10. (c) 0.15, 0.85
(d) P(9)=10C9 (.15)? (.85)?. (d) 9, 1

9. In a recent year billings and exports increased more than 10% compared to the previous year. Last year, planes with
piston engines accounted for 70% of all sales. If 100 planes were randomly selected from the sales of all planes last year,
then what is the probability that 75 or more of them had piston engines?
This is a binomial distribution.
n=100, p=0.70, x=number of planes sold that have piston engines
P(x≥75)=0.1631

5. A large University system, which currently uses a standard grading system of A, B, C, D or F for their students, is
analyzing the possibility of changing to a plus-minus grading policy. In addition to the computer system, which stores and
manages all the student records, the faculty and the students would also be affected by this possible change. One of the
concerns that students have, is that changing this policy might lower the overall GPA (grade point average) of the
students. One statement from students makes the conjecture that only 15% of all the students in the system are
supportive of the proposed change. If a random sample of 260 students in this University system gave the results that 60
students are supportive of the change, 150 are not supportive and the remaining 50 students are indifferent to the change.
Based on the data do you think that the student's conjecture that 15% of all the students in the system are supportive of
the proposed change could be correct?
This is a binomial distribution.
n=260, p=0.15
x=number of students who are supportive of the proposed change in the grading policy
μ =np=(260)(0.15)=39, σ =npq=√np(1−p)=√(260)(0.15)(0.85)=5.76, z−score=x− μ σ =60−395.76=3.65
So, based on the students’ conjecture we obtained a z−score more than 3 standard deviations above the mean. This is an
inconsistency as it is very unlikely for data to be more than 3 standard deviations above the mean. Since our data is from a
random sample, we would conclude that the students’ conjecture that only 15% of all the students in the system are supportive of
the proposed change in the grading policy is, in fact, not correct. Furthermore, based on the data we conclude that the percentage
of students in favor of the change is higher than 15%. Another way to think about this is that 60/260 is approximately 23%,
which is more than 3 standard deviations above the mean.

Chapter 5

6. Weights of newborn babies in the United States are normally distributed with a mean of 3420 grams and a standard
deviation of 495 grams. A newborn weighing less than 2200 grams is considered to be at risk, because the mortality rate
for this group is at least 1%. What percentage of newborn babies are not in the "at-risk" category?
Let X=the weight of a newborn baby, z=(X−mean)/stdev=(2200−3420)/495=−2.46
Then, P(baby not at−risk)=P(X>2200)=P(z>−2.46)=0.9931

10. Flat screen monitors manufactured by an offshore OEM supplier have life spans which are normally distributed with
a mean life of 13,500 hours and a standard deviation of 1,200 hours. Suppose that a flat screen monitor is selected at
random from this offshore OEM supplier.
(a) Find the probability that the life span of the monitor will be more than 15,000 hours. Let X=the life span of a randomly
selected monitor from this offshore OEM supplier
(a) P(X>15,000)=P(z>1.25)=0.1056
(b) Find the probability that the life span of the monitor will be between 12,500 hours and 14,000 hours.
(b) P(12,500<X<14,000)=P(−0.83<z<0.42)=0.4595
(c) Suppose that a potential buyer needs to be fairly certain (say 99%) that the life span of the flat screen monitors from this
offshore OEM supplier will have a life span of at least 10,000 hours. Can we be certain that these monitors will meet the
requirement? (c)P(the 10,000 minimum hour life span requirement is met) =P(X≥10,000)=P(z≥−2.92)=0.9982
Since the probability of a life span of at least 10,000 hours is more than 0.99, the requirement is met.

9. A statistics instructor has decided to give the top 15% of all students in her class A's based on total points in the course.
Suppose that the total points in the course have a mean of 750 with a standard deviation of 85 and the total points are
approximately normally distributed. What number of total points will provide a student an A in the course?
Let X=the total points required in the course to obtain an A, X= μ +z σ =750 points+( 1.04)(85 points)=838.4 points
So, you need to obtain 838.4 points or more in the course to obtain an A.

4. Due to increasing fuel prices, a new fuel saving device was developed that, when mounted to the intake manifold on
commercial diesel engines, increases mileage significantly. The device has a limited life and therefore must be replaced at
certain intervals. The manufacturer claims that the device has an average life of 750 hours. Assume the lives of all of these
devices has a normal distribution, with a mean of 750 hours and a standard deviation of 58 hours.
(a) If one of these devices is selected at random, what is the probability that the lifetime of this device will exceed 760 hours?
(a) This is a question about X and not ¯¯¯x.
z=x− μ σ =760−75058=0.17
P(X>760)=P(z>0.17)=0.4325
(b) If a random sample of 36 of these devices what is the probability that the sample mean will exceed 760 hours?
σ¯¯¯x= σ√n=58√36=9.6667
z=(¯¯¯x−μ¯x)σ¯x=(760−750)9.6667=1.03
P(¯¯¯x>760)=P(z>1.03)=0.1515
Chapter 6

14. Ferguson Corporation has two furniture stores in the city in different locations. The company’s quality control
department wanted to check if the customers are equally satisfied with the service provided at these two stores. A random
sample of 25 customers selected from Store Location I produced a mean satisfaction index of 7.6 (on a scale from 1 to 10,
1 being the lowest and 10 being the highest) with a standard deviation of 0.75. Another random sample of 28 customers
selected from Store Location II produced a mean satisfaction index of 8.1 with a standard deviation of 0.59. Assume that
the customer satisfaction index for each supermarket has a normal distribution with the same population standard
deviation. Construct and interpret a 99% confidence interval for the difference between the mean satisfaction indexes for
all customers for the two supermarkets.
Let μ 1=the mean satisfaction index for Store Location 1.
Let μ 2=the mean satisfaction index for Store Location 2.
We have independent random samples from the two populations, which are normally distributed with equal population
variances, which lead us to the formula:
¯¯¯x1−¯¯¯x2±t√s2p(1n1+1n2) with df=n1+n2−2
s2p=(n1−1)s21+(n2−1)s22n1+n2−2
s2p=(25−1)(0.75)2+(28−1)(0.59)225+28−2=0.4490
Then, df =25+28−2=51 and
(7.6−8.1)±2.68√0.4490(125+128)
(−0.994, −0.006)
We are 99% confident that the difference between the mean satisfaction index of Store Location 1 and Store Location 2 lies
between -0.994 and -0.006. This confidence interval does show that the satisfaction index of Store Location 1 is lower than the
satisfaction index of Store Location 2.

Chapter 8

16. Complete all four steps of the hypothesis testing procedure.


The Los Angeles Times reported that 32% of all adult Americans have attended at least one year of college. Suppose that a
random sample of 200 adults in the western United States included 82 with one or more years of college. At a 0.05 level of
significance, do the data support the assertion that a greater proportion of westerners (as compared to the United States as a
whole) have attended at least one year of college?
Step #1
H0:p=0.32
HA:p>0.32
p=the true proportion of Westerners who have attended at least one year of college
Step #2 z∗=2.73
Step #3 p−value=P(z>2.73)=0.0032
Step #4 At a 5% level of significance, the sample evidence is sufficient to show that a greater proportion of westerners (as
compared to the United States as a whole) have attended at least one year of college.

Calcium blockers are among several classes of medicines commonly prescribed to relieve high blood pressure. A study in
Denmark has found that calcium blockers may also be effective in reducing the risk of heart attacks. A total of 897 randomly
selected Danish patients, each recovering from a heart attack, were given a daily dose of the drug Verapamil, a calcium blocker.
After eighteen months of follow-up, 146 of these patients had recurring heart attacks. In a control group of 878 randomly selected
individuals, each of whom took placebos, 180 had a heart attack. At a 1% level of significance, do the data provide sufficient
evidence to infer that calcium blockers are effective in reducing the risk of heart attacks? Assume that the samples are
independent.
Step #1
H0:p1−p2=0
HA:p1−p2<0
p1=the true proportion of patients taking Verapamil who had recurring heart attacks
p2=the true proportion of patients taking a placebo ho had recurring heart attacks
Step #2 z∗=− 2.30
Step #3 p−value=P(z <−2.30)=0.0107
Step #4 At a 1% level of significance, the sample evidence is insufficient to show that calcium blockers are effective in reducing
the risk of heart attacks.

Chapter 9

4. Physicians have used the “diving reflex” to reduce abnormally rapid heartbeats in humans by briefly submerging the
patient’s face in cold water. The reflex, triggered by cold water temperatures, is an involuntary neural response that shuts
off circulation to the skin, muscles, and internal organs to divert extra oxygen-carrying blood to the heart, lungs and
brain. A research physician conducted an experiment to investigate the effects of various cold-water temperatures on the
pulse rate of small children. The data for seven 6-year-old children consisted of two measurements taken on each child:
x=the temperature of the water (°F), and y=the decrease in pulse rate (beats/minute). Use the following data to answer the
questions below. For this experiment the coldest water temperature used was 58°F, and the warmest water temperature
used was 70°F.
Correlation of TEMP and PULSE=− 0.970
The regression equation is:
PULSE=57.9−0.811(TEMP)
(a) State and interpret the correlation coefficient.
(b) State the least squares regression equation for the pulse rate data.
(c) If the water temperature is 61°F, predict the drop in pulse rate for a 6-year-old child.
(d) If the water temperature is 72°F, predict the drop in pulse rate for a 6-year-old child.
(e) Interpret the slope of the regression line.
Solution
(a) Since the correlation coefficient (r=−0.970) is close to -1, we can conclude that there is a strong negative linear relationship
between the temperature of the water and the decrease in pulse rate of 6-year-old children. As a result, a linear model is
appropriate.
(b) ^y =57.9−0.811x
x=temperature of the water
^y=predicted pulse rate of 6−year−olds
(c) ^y=57.9−0.811(61)=8.429 beats per minute
For a water temperature of 61°F, the model predicts a decrease of 8.429 beats/minute in the pulse rate of 6-year-olds.
(d) A water temperature of 72°F is beyond the scope of the data. As a result, our regression model is not appropriate for
prediction purposes.
(e) For every 1°F increase in the temperature of the water, the model predicts a decrease of 0.811 beats per minute in the pulse
rate of a 6-year-old child.

1. Infestation of crops by insects has long been of great concern to farmers and agricultural scientists. A study was
performed where a crop of strawberry plants was evaluated over time and the percentage of damaged plants was
recorded. Twelve observations were recorded where the researchers were hoping to be able to predict the percentage of
damaged crops from the age of the crop, in days. The data collection began on day 9 of the crop and ended on day 33. Use
the computer output below to answer the following questions.
Correlation of AGE and DAMAGE=0.955
The regression equation is:
DAMAGE=−20.3+3.34(AGE)
(a) Would the data be considered close enough to linear in order to use simple linear regression for prediction purposes? Explain
and defend your answer.
(b) Interpret the slope of the least squares regression equation.
(c) Predict the percentage of crop damage for a crop of strawberries that is 25 days old.
(d) Predict the percentage of crop damage for a crop of strawberries that is 47 days old.
Solution
(a) Since the correlation coefficient (r=0.955) is close to 1, we can conclude that there is a strong positive linear relationship
between the age of the strawberry crop (x) and the amount of damage done to the crop by insects. As a result, a linear model is
appropriate.
(b) For every 1 day increase in the age of the crop, the model predicts that the amount of insect damage done will increase by
3.34%.
(c) ^y=−20.3+3.34(25)=63.2 %
For a crop that is 25 days old, the model predicts that 63.2% of the crop will be damaged due to insects.
(d) Because 47 days is beyond the scope of the data, we cannot use our model for prediction purposes.

Sample Final:

I. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 1 THROUGH 3.

Among 100 marriage license applications, chosen at random in 2001, there were 8 in which the women were at least
one year older than the men, and among 30 marriage license applications, chosen at random in 2007, there were 6 in
which the women were at least one year older than the men. Assume you are interested in a 80% confidence interval
for the difference between the corresponding true proportions of marriage license applications in which the women
were at least one year older than the men.

1.

2. Choose from the following concerning the z-value or t-value that would be used in the 80% confidence interval for the
difference between the corresponding true proportions of marriage license applications in which the women were at least one
year older than

z = 1.28

3. If we assign group 1 to be the sample from 2001 and group 2 to be the sample from 2007, you may assume that the appropriate
confidence interval for the above problem to be (-0.220, -0.020). Based on this interval choose the most appropriate portion
below from the conclusion for the confidence interval.

We are 80% confident that the true proportion of marriage licenses in which the women were at least one year older than the men was greater in 2007 than in 2001.

Since the interval contains all negative numbers, we are 80% confident that p2 > p1

II. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 4 THROUGH 6.

In a study of heart surgery, one issue was the effect of drugs called beta-blockers on the pulse rate of patients during surgery. The available subjects were divided at random into two groups of 21
patients each. One group received a beta-blocker; the other a placebo. The surgical team recorded the pulse rate of each patient at a critical point during the operation. The treatment group had
mean 65.2 beats per minute. For the control group, the mean was 70.3. Below is Minitab output for the beta-blocker data. Assume equal population variances and assume the populations are
independent and normally distributed.

TWO SAMPLE CI FOR TREATMENT VS CONTROL

N MEAN
TREATMENT 21 65.20

CONTROL 21 70.30

95 PCT CI FOR MU TREATMENT - MU CONTROL: (-1.3, 5.47)

4. Choose the best interpretation below for the confidence interval in the Minitab output We are 95% confident that there is no
statistically significant difference between the true mean beats per minute during surgery for the two groups.

Since the 95% confidence interval for the difference in the means contains both positive and negative numbers, we
conclude that there is no statistically significant difference between the true mean beats per minute during surgery for
the two groups

5.

6. Pick the appropriate z or t value used in the above interval.

2.02

Since the confidence level is 95% and df = n1 + n2 - 2 = 21 +21 - 2 = 40, we find that the value is t = 2.02.

III. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 7 THROUGH 11.

A researcher wants to test whether the proportion of Foothill students who transfer to a four-year university is
different than the proportion of De Anza students who transfer to a four-year university.

7.

8. If the test statistic for the test is 𝑧∗=1.48 what would be the p-value for the test?

0.0694

0.1388

The alternative is "not equal" so we find the area in the right tail and double it to obtain the p-value.

Area is 0.0694

p-value = 2(Area) = 2(0.0694) = 0.1388


9. If the test statistic for the test is 𝑧∗=1.48 , then, at the 5% significance level, what would be your conclusion for the
test?

There is sufficient evidence to conclude HA is true.

There is insufficient evidence to conclude HA is true.

Since the p-value (0.1388) is greater than the significance level (0.05), there is insufficient evidence to conclude that
HA is true.

10. Choose the most appropriate conclusion for the hypothesis test.

At the 5% significance level, there is sufficient evidence to conclude that the true proportion of Foothill students who transfer to a four-year university is different than the true

proportion of DeAnza students who transfer to a four-year university.

At the 5% significance level, there is insufficient evidence to conclude that the true proportion of Foothill students who transfer to a four-year university is different than the true

proportion of DeAnza students who transfer to a four-year university.

This is the same as Problem 9 above. However, for this answer we are applying the correct phrasing using the context
of the problem.

11.

IV. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 12 THRU 15.

A major car manufacturer wants to test a new engine to see whether it meets new air pollution standards – the
population mean carbon emissions all engines of this type must be less than 20 parts per million(ppm). 21 engines are
randomly sampled for testing purposes and the mean and standard deviation of the emissions for those engines are
calculated to be 16.46 ppm and 9.49 ppm respectively. You are interested in whether the data supply sufficient
evidence to allow the manufacturer to conclude that this type of engine meets the pollution standard at a 10% level of
significance. Assume a normal distribution. Assume the p-value is 0.051.
12.

13.

14. Choose the most appropriate conclusion for the hypothesis test in problem 12.

At the 10% significance level, there is sufficient evidence to conclude that this type of engine meets the pollution standard.

Notice that p-value = P(t < -1.71) = 0.0514, using df = 20.

Then, the p-value is less than the significance level (0.10); so we conclude that there is sufficient evidence to conclude
that HA is true. So, this type of engine meets the pollution standard.

15.

V. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 16 THRU 18.

Suppose you record how long it takes you to get to school over many months and discover that the one-way travel
times, in minutes, are approximately normally distributed with a mean of 13.88 minutes and standard deviation of 4
minutes. If your first class is at 10:30 am, what is the latest time you should leave everyday so that you are on time at
least 93.7% of the time?
16.

17. Select the most appropriate item that pertains to the problem.

z = 1.53

Using the beoga app we find the z value that has 0.937 area below it.

z = 1.53

18.

19.
20. In the next three births, what is the probability that at least two are girls?

0.3333

0.4805

Note: p = P(a girl) = 1 - P(a boy) = 1 - 0.513 = 0.487

This is a binomial distribution where x = the number of girls, n = 3 and p = 0.487.

P(at least 2 girls) = P(x >= 2) = 0.4805 by using the beoga app.

VII. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 21 THROUGH 24.

With the recent growth in the number of U.S. households wired for cable television, major advertisers are now
beginning to show an increased interest in cable advertising. Since its inception, cable television has offered
advertisers relatively inexpensive rates and selective target audiences. For purposes of comparison, A.J. Bush and J.H.
Leigh conducted a content analysis of television commercials on the three major networks, ABC, CBS, and NBC, and
three of the more popular, well established cable networks, CNN, ESPN, and TBS. One of the variables measured was
the number of advertisements shown during prime-time (7:00 - 10:00 P.M.) and late-night (10:00 P.M. - 12:00 A.M.)
viewing segments on two weekday evenings during November 1981. The results are summarized in the table below.

Number of Advertisements

Prime-time Late-night TOTALS

MAJOR ABC 94 69 163

NETWORKS CBS 90 84 174

NBC 86 94 180

CABLE TBS 66 47 113

NETWORKS CNN 95 60 155

ESPN 87 45 132

TOTALS 518 399 917

21. Suppose an advertisement is selected at random from advertisements shown on the two weekday evenings during
November 1981. What is the probability that the advertisement selected was shown during a late-night segment?

163/399

399/917

Late-nite total = 399

P(ad is shown Late-nite) = 399/917

22. Suppose an advertisement is selected at random from advertisements shown on the two weekday evenings during
November 1981. What is the probability that the advertisement selected was shown on TBS during a Prime-time
segment?
113/119

113/917

66/917

There are 66 in the cell that corresponds to "TBS" and "Prime Time"

P(on TBS during Prime-time) = 66/917

23.

24.

VIII. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 25 THROUGH 27.

Electric power plants that use water for cooling their condensers sometimes discharge heated water into rivers, lakes,
or oceans. It is known that water heated above certain temperatures has a detrimental effect on the plant and animal
life in the water. Suppose it is known that the increased temperature of the heated water discharged by a certain
power plant in any given day has a normal distribution with a mean of 2.8 degrees Celsius and a standard deviation of
0.6 degrees Celsius. Suppose a sample of 35 days was collected. You are interested in the probability that the mean
increase in temperature of the discharged water for the sample exceeds 2.6 degrees Celsius.

25.
26.

27.

IX. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 28 THRU 30.

The albatross is known to have the longest wingspan of any living bird, typically ranging from 8.25 feet to 11.5 feet.
The wing span for the albatross is approximately normally distributed. A random sample of 10 albatross wingspans is
taken and the wingspans are recorded.

28.

29. Pick the appropriate z or t value used in the above interval.

1.64

1.83

For a small sample we use a t value corresponding to a 90% confidence level with df = 10 - 1 = 9.

t = 1.83
30.

X. USE THE INFORMATION BELOW TO ANSWER PROBLEMS 31 THRU 34

Acme Properties, Inc. specializes in custom home sales in the Evergreen Estates. A random sample of nine currently
listed custom homes provided information on size and price. The size data are in hundreds of square feet, rounded to
the nearest hundred; the price data are in thousands of dollars, rounded to the nearest thousand. The smallest house
in the sample was 22 hundred square feet and had a price of 195 thousand dollars. The largest house in the sample
was 40 hundred square feet and had a price of 475 thousand dollars.

Correlation of SIZE and PRICE = 0.930

The regression equation is PRICE = - 141 + 14.4 SIZE

31.

32. Predict the price of a custom home in the Evergreen Estates that is 47 hundred square feet.

$535,800

817,800

$475,000

47 hundred square feet is beyond the scope of the data

The value of x = 47 is beyond the scope of the data as the largest value is x = 40. So, we are not allowed to predict y
for this value of x.
33.

34. Interpret the correlation coefficient.

There is a weak negative linear relationship between square footage and sale price.

There is a strong negative linear relationship between square footage and sale price.

There is a strong positive linear relationship between square footage and sale price.

Since r is close to 1 (r = 0.930), there is a strong positive linear relationship between x and y.
XI. USE THE INFORMATION BELOW TO ANSWER PROBLEM 35

35. Which of the following two variables are both quantitative and continuous?

1. The number of ships passing under the Golden Gate Bridge in a day.

2. The amount of gasoline purchased at a gas station.

3. The time it takes to complete a test on an exam.

4. The type of surgery performed in an outpatient clinic.

2 and 4

2 and 3

Recall, quantitative and continuous means the data is numerical where any number on some interval is possible. As a
result, options 2 and 3 are correct.
36.

37. Choose the correct statement below (only one is correct).

The mean of the distribution is approximately $65,000

Approximately 15% of the salaries are $80,000 or higher

Note that:

The mean of the distribution is close to $55,000 not $65,000. The range of the distribution is close to $90,000 not
$45,000. Approximately 50% of the salaries are below $55,000 not $30,000.

By process of elimination and inspection of the graph, we conclude that "approximately 15% of the salaries are
$80,000 or higher".
• Approximately 68% of the data falls within one standard deviation ( σ\sigmaσ) of the mean (μ\muμ).
• Approximately 95% of the data falls within two standard deviations ( 2σ2\sigma2σ) of the mean (μ\muμ).
• Approximately 99.7% of the data falls within three standard deviations ( 3σ3\sigma3σ) of the mean (μ\muμ).

P(∣X−μ∣≥kσ)≤k21

Where:

• μ\muμ is the mean of the data set.


• σ\sigmaσ is the standard deviation.
• kkk is any number greater than 1.
• PPP denotes the probability.
• For k=2k = 2k=2, at least 122=14=0.25\frac{1}{2^2} = \frac{1}{4} = 0.25221 =41 =0.25 of the data lies within
2 standard deviations of the mean, meaning at least 75% of the data lies within this range.
• For k=3k = 3k=3, at least 132=19≈0.1111\frac{1}{3^2} = \frac{1}{9} \approx 0.1111321 =91 ≈0.1111 of the
data lies within 3 standard deviations of the mean, meaning at least 88.89% of the data lies within this range.
.

You might also like