0% found this document useful (0 votes)
2 views10 pages

Section5 1text

Chapter 5 discusses discrete probability distributions, focusing on how random variables can be classified as discrete or continuous based on their values. It explains the concept of discrete probability distributions, providing examples such as household sizes from the 2010 U.S. Census, and outlines methods for calculating the mean and standard deviation using technology. The chapter also introduces the rare event rule for inferential statistics, helping to determine if an event is unusual based on its probability.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views10 pages

Section5 1text

Chapter 5 discusses discrete probability distributions, focusing on how random variables can be classified as discrete or continuous based on their values. It explains the concept of discrete probability distributions, providing examples such as household sizes from the 2010 U.S. Census, and outlines methods for calculating the mean and standard deviation using technology. The chapter also introduces the rare event rule for inferential statistics, helping to determine if an event is unusual based on its probability.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 5: Discrete Probability Distributions

Section 5.1: Basics of Probability Distributions


In the first chapter we talked about different types of data that resulted from collecting some sort
of variable information from each individual in a sample (or sometimes a population):
qualitative and quantitative. No matter what type of data have been collected, we need a way to
translate that information into numeric information. For qualitative data, that usually entails
counting how many are in a particular category or by computing the percentage in a particular
category. For quantitative data, the data are already numeric, but many times we are interested
in other numeric summaries of that data as well such as the mean.

Random variable – a numerical count or measure of an outcome of a probability experiment, so


its value is determined by chance. Random variables are usually denoted by capital letters such
as X. We usually use the abbreviation rv to stand for random variable.

This numerical measure can be discrete or continuous.


Discrete random variable – a random variable that can only take on particular values in a
certain interval of numbers.
Continuous random variable – a random variable the can take on any value in a certain interval
of numbers.

Discrete random variables usually arise from counting while continuous random variables
usually arise from measuring.

Examples of each:
How tall is a plant given a new fertilizer?
rv X = the height of a randomly selected plant given the new fertilizer
This random variable is continuous since height is something you measure.
How many fleas are on a prairie dog?
rv X = the number of fleas on a randomly selected prairie dog
This random variable is discrete since you count the number of fleas.

Now suppose you put all the values of the random variable together with the probability that that
value of the random variable would occur. You could then have a distribution like before, but
now it is called a probability distribution since it involves probabilities. In this chapter and the
next we are going to focus on discrete probability distributions.

Discrete probability distribution – a table, graph or formula that gives all of the values possible
for the discrete random variable along with their corresponding probabilities.
For discrete probability distributions, 0  P  x   1 and  P  x   1
Example #5.1.1: Probability Distribution
The 2010 U.S. Census found the chance of a household being a certain size. The data is
in table #5.1.1 ("Households by age," 2013).

Table #5.1.1: Household Size from U.S. Census of 2010


Size of household 1 2 3 4 5 6 7 or more
Probability 26.7% 33.6% 15.8% 13.7% 6.3% 2.4% 1.5%

Solution:
In this case, the random variable is X = number of people in a household. This is a
discrete random variable, since you are counting the number of people in a household.
This is a discrete probability distribution since you have each x value and the
probabilities that go with it, all of the probabilities are between zero and one, and the sum
of all of the probabilities is 100% which is 1.

You can give a probability distribution in table form (as in table #5.1.1) or as a graph. The graph
looks like a histogram. A discrete probability distribution is basically a relative frequency
distribution based on a very large sample.

Example #5.1.2: Graphing a Probability Distribution


The 2010 U.S. Census found the chance of a household being a certain size. The data is
in the table ("Households by age," 2013). Draw a histogram of the probability
distribution.

Table #5.1.2: Household Size from U.S. Census of 2010


Size of 7 or
household 1 2 3 4 5 6 more
Probability 26.7% 33.6% 15.8% 13.7% 6.3% 2.4% 1.5%

Solution:
State random variable:
rv X = number of people in a household
You draw a histogram, where the x values are on the horizontal axis and are the x values
of the classes (for the 7 or more category, just call it 7). The probabilities are on the
vertical axis.
Graph #5.1.1: Histogram of Household Size from U.S. Census of 2010

Notice that this distribution of U.S. household size in 2010 is skewed right.
In problems involving a probability distribution function (pdf), you consider the probability
distribution the population even though the pdf in most cases comes from repeating an
experiment many times. This is because you are using the data from repeated experiments to
estimate the true probability (remember the Law of Large Numbers from the last chapter). Since
a pdf is basically a population, we use the population parameter symbols for the mean and
standard deviation. Note: the mean can be thought of as the expected value. It is the value you
expect to get on average if the trials were repeated infinite number of times. The mean or
expected value does not need to be a whole number, even if the possible values of the random
variable X are whole numbers.

For a discrete probability distribution function,


The mean or expected value is x   x  P  x 

The standard deviation is  x     x     P  x 


x
2

where x = the value of the random variable and P(x) = the probability corresponding to that
particular x value.
NOTE: Formulas are included here but you will always be using technology to compute these.

TECHNOLOGY: MEAN AND STANDARD DEVIATION OF DISCRETE PDF


Using your TI84:
 First push STAT 1 and enter the values of X into L1 and the corresponding probabilities
into L2 (or any other two lists)
 Push STAT  1 to open 1-Var Stats.
 You need to make your input screen look like one of the screens below (depending on
which operating system your calculator has). Push 2nd 1 to type L1 and push 2nd 2 to type
L2.

or
 Then push ENTER or highlight “Calculate” and push ENTER

Example #5.1.3: Calculating Mean and Standard Deviation for a Discrete Probability
Distribution
The 2010 U.S. Census found the chance of a household being a certain size. The data is
in the table ("Households by age," 2013). Calculate the mean and standard deviation.

Table #5.1.3: Household Size from U.S. Census of 2010


Size of 7 or
household 1 2 3 4 5 6 more
Probability 26.7% 33.6% 15.8% 13.7% 6.3% 2.4% 1.5%
Solution:
State random variable:
rv X = number of people in a household

Using the TI84:


o First push STAT 1 and enter the data into two lists:

o Push STAT  1 to open 1-Var Stats and enter the lists names like below:

or
o Highlight “Calculate” and push ENTER you will see the output below.

Figure #5.1.1: TI-84 Output

The calculator assumes you have sample data and does not display the correct symbol
for the mean when using probability distributions. The correct notation here would
be  x  2.525 people and  x  1.422 people.
This means that you expect a household in the U.S. to have an average of 2.525 people
in it with a standard deviation of 1.422 people.

Example #5.1.4: Calculating the Expected Value


In the Arizona lottery called Pick 3, a player pays $1 and then picks a three-digit number.
If those three numbers are picked in that specific order the person wins $500. What are
the expected winnings in this game?
Solution:
To find the expected winnings, you need to first create the probability distribution. In
this case, the random variable X = winnings. If you pick the right numbers in the right
order, then you win $500, but you paid $1 to play, so you actually win $499. If you
didn’t pick the right numbers, you lose the $1, the x value is $1. You also need the
probability of winning and losing. Since you are picking a three-digit number, and for
each digit there are 10 numbers you can pick with each independent of the others, you
can use the multiplication rule. To win, you have to pick the right numbers in the right
order. The first digit, you pick 1 number out of 10, the second digit you pick 1 number
out of 10, and the third digit you pick 1 number out of 10. The probability of picking the
1 1 1 1
right number in the right order is * *   0.001 . The probability of losing
10 10 10 1000
1 999
(not winning) would be 1    0.999 . Putting this information into a table
1000 1000
will help to calculate the expected value.

Table #5.1.6: Finding Expected Value


Win or lose x P(x)
Win $499 0.001
Lose $1 0.999

Now, for this one the table is small enough and we only have to compute the mean so we
could use the formula and get the following:
x  $499  0.001  $1 0.999  $  0.50
In the long run, you will expect to lose $0.50 per game on average. Notice that you can
never play the lottery and lose 50 cents: you either win $499 or you lose $1. But if you
were to play this lottery many, many times, and record your winnings each time, the list
of winnings if you add it up and divide by how many times you played would tend to get
closer to $  0.50 the more times you played.

Sometimes with games we are interested in whether a game is fair or not. A game that is
fair is one where the expected winnings are $0. Since the expected winnings for the
lottery are not $0, this game is not fair. Since you lose money, Arizona makes money,
which is why they have the lottery.

The reason probability is studied in statistics is to help in making decisions in inferential


statistics. To understand how that is done, the concept of a rare event is needed.

Rare (unusual) event – an event that has a small chance of happening. For now, we are going
to define an event as rare (unusual) if its probability is less than or equal to 0.05

This leads us to the rare event rule that we will be using for inferential statistics.
Rare Event Rule for Inferential Statistics
If, under a given assumption, the probability of a particular observed event is extremely small,
then you can conclude that the assumption is probably not correct.

As an example, suppose that you assume a die is fair and you roll it 100 times and you roll a 6 on
sixty of those rolls. The chance of getting at least sixty rolls with a 6 out of 100 rolls of a fair die
is only 0.0284. This is less than 0.05, so it would be reasonable to believe that the assumption
about the die being fair is not true.

Determining if an event is unusual


If you are looking at a value of x for a discrete random variable, and the P(the variable has a
value of x or more)  0.05, then you can consider the x an unusually high value. Another way
to think of this is if the probability of getting such a high value is less than 0.05, then the event of
getting the value x is unusual.

Similarly, if the P(the variable has a value of x or less)  0.05, then you can consider the x an
unusually low value. Another way to think of this is if the probability of getting a value as small
as x is less than 0.05, then the event of getting the value x is considered unusual.

Why is it "x or more" or "x or less" instead of just "x" when you are determining if an event is
unusual? Consider this example: you and your friend go out to lunch every day. Instead of each
paying for their own lunch, you decide to flip a coin, and the loser pays for both. Your friend
seems to be winning more often than you'd expect, so you want to determine if this is unusual
before you decide to change how you pay for lunch (or accuse your friend of cheating). The
process for how to calculate these probabilities will be presented in the next section on the
binomial distribution. If your friend won 6 out of 10 lunches, the probability of that happening
turns out to be about 20.5%, not unusual. The probability of winning 6 or more is about
37.7%. But what happens if your friend won 501 out of 1,000 lunches? That doesn't seem so
unlikely! The probability of winning 501 or more lunches is about 47.8%, and that is consistent
with your hunch that this isn't so unusual. But the probability of winning exactly 501 lunches is
much less, only about 2.5%. That is why the probability of getting exactly that value is not the
right question to ask: you should ask the probability of getting that value or more (or that value
or less on the other side).
The value 0.05 will be explained later, and it is not the only value you can use.

Below are the common wordings for inequality symbols that you should make sure you know.
is at least  is greater than or equal to 
is at most  is less than or equal to 
is less than < exceeds >
is greater than > does not exceed 
is a maximum of  is below <
is a minimum of  is above >
is no less than  is no greater than 
is between < is between inclusive 
Example #5.1.5: Is the Event Unusual
The 2010 U.S. Census found the chance of a household being a certain size. The data is
in the table ("Households by age," 2013).

Table #5.1.7: Household Size from U.S. Census of 2010


Size of 7 or
household 1 2 3 4 5 6 more
Probability 26.7% 33.6% 15.8% 13.7% 6.3% 2.4% 1.5%
Solution:
State random variable:
rv X = number of people in a household

a.) Is it unusual for a household to have six people in the family?

Solution:
To determine this, you need to look at two probabilities and compare each of them to
0.05. However, you cannot just look at the probability of six people. You need to
look at the probability of x being six or more people or the probability of x being six
or less people.
P  X  6  P  X  1  P  X  2  P  X  3  P  X  4  P  X  5  P  X  6
P  X  6  26.7%  33.6%  15.8%  13.7%  6.3%  2.4%
P  X  6  98.5%
Since this probability is more than 5%, six is not an unusually low value.

P  X  6  P  X  6  P  X  7 
P  X  6  2.4%  1.5%
P  X  6  3.9%
Since this probability is less than 5%, six is an unusually high value.

It is unusual for a household to have six people in the family.


b.) If you did come upon many families that had six people in the family, what would
you think?

Solution:
Since it is unusual for a family to have six people in it, then you may think that either
the size of families is increasing from what it was or that you are in a location where
families are larger than in other locations.

c.) Is it unusual for a household to have four people in the family?

Solution:
To determine this, you need to look at probabilities. Again, look at the probability of
x being four or more or the probability of x being four or less.
P  X  4   P  X  4   P  X  5  P  X  6   P  X  7 
P  X  4  13.7%  6.3%  2.4%  1.5%
P  X  4  23.9%
Since this probability is more than 5%, four is not an unusually high value.
P  X  4  P  X  1  P  X  2  P  X  3  P  X  4
P  X  4  26.7%  33.6%  15.8%  13.7%
P  X  4  89.8%
Since this probability is more than 5%, four is not an unusually low value.

Thus, four is not an unusual size of a family.

d.) If you did come upon a family that has four people in it, what would you think?

Solution:
Since it is not unusual for a family to have four members, then you would not think
anything is amiss.
Section 5.1: Homework
1.) Eyeglassomatic manufactures eyeglasses for different retailers. The number of days it
takes to fix defects in a pair of eyeglasses and the probability that it will take that number
of days are in the table.
Table #5.1.8: Number of Days to Fix Defects
Number of days Probabilities
1 24.9%
2 10.8%
3 9.1%
4 12.3%
5 13.3%
6 11.4%
7 7.0%
8 4.6%
9 1.9%
10 1.3%
11 1.0%
12 0.8%
13 0.6%
14 0.4%
15 0.2%
16 0.2%
17 0.1%
18 0.1%
a.) State the random variable.
b.) Draw a histogram of the number of days to fix defects
c.) Find the mean number of days to fix defects.
d.) Find the standard deviation for the number of days to fix defects.
e.) Find probability that it will take at least 16 days to fix the defect.
f.) Is it unusual for it to take 16 days to fix a defect on a pair of eyeglasses?
g.) If it does take 16 days for a pair of eyeglasses to be repaired, what would you think?

2.) Suppose you have an experiment where you flip a coin three times. You then count the
number of heads.
a.) State the random variable.
b.) Write the probability distribution for the number of heads.
c.) Draw a histogram for the number of heads.
d.) Find the mean number of heads.
e.) Find the standard deviation for the number of heads.
f.) Find the probability of having two or more number of heads.
g.) Is it unusual to flip two heads?
3.) The Ohio lottery has a game called Pick 4 where a player pays $1 and picks a four-digit
number. If the four numbers come up in the order you picked, then you win $2,500.
What are your expected winnings?

4.) An LG Dishwasher, which costs $800, has a 20% chance of needing to be replaced in the
first 2 years of purchase. A two-year extended warrantee costs $112.10 on a dishwasher.
What is the expected value of the extended warranty assuming the dishwasher is replaced
in the first 2 years?

You might also like