0% found this document useful (0 votes)
7 views3 pages

Module 1 Review Questions on Probability

Uploaded by

Dawsar Abdelhadi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views3 pages

Module 1 Review Questions on Probability

Uploaded by

Dawsar Abdelhadi
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Review Questions – Module 1

Key: * = this is good practice for the module quiz. ** = Useful for understanding, but slightly
more difficult or open-ended. At the very least, skim the provided answers, but stretch your brain
a bit by trying them yourself first.

* Question 1: Suppose we have a discrete random variable X that can take the values 0, 1, 2,
3, 4, or 5. The table shown below is called a discrete probability distribution: it lists the values X
can take (x, along the top) and the probability that X takes each of those values (P(X=x), along
the bottom). For example, P(X=0)=0.01, and P(X=5)=0.50.

x 0 1 2 3 4 5
P(X=x) 0.01 0.15 0.05 0.25 0.04 0.50

Common question: Wait, I thought a probability was an area under a curve…why can X take a
specific value here? For example, why isn’t P(X=5) equal to 0?

Answer: Our random variable here is discrete, not continuous. A discrete random variable
takes countably many possible values. Here, this random variable can take any of 6 possible
values (0, 1, 2, 3, 4, or 5). That means it can only be 0, 1, 2, 3, 4, or 5, no other value is
possible. The total probability that X takes any of those values is 1 (that is,
0.01+0.15+0.05+0.25+0.04+0.50=1). There is more information on discrete random variables in
the required reading for Module 1. We’ll talk more about discrete random variables later in the
course.

a) Find the probability that X is greater than 3.

b) Find the probability that X is less than or equal to 2.

c) Find the probability that X is between 1 and 4, inclusive.

d) Find the probability that X is between 1 and 4, exclusive.

** Question 2: State whether each of these statements is true or false and explain your
reasoning.

a) A z-score is a percentile.

b) If X ~ N(54, 3), then P(X < 50) = P(X > 58).

c) A normally distributed random variable can never take on negative values.

d) If the standard deviation of a normally distributed random variable is 0, then the variance of
the random variable is large.
* Question 3: Let X ~ N(-10, 5)

a) What is P(X < -5)?

b) What is P(X > - 2)?

c) What is P(-2.5 < X < 2.5)?

d) What is P(-3 < X < 1)?

e) What is P(X = -10)?

Question 4: Let X ~ N(40, 4)

* a) What is the mean of this distribution?

* b) What is the median of this distribution? Interpret this value.

* c) Suppose P(X < x) = 0.5. Find x.

** d) Suppose P(X < x) = 0.975. Find x.

** e) Suppose P(X > x) = 0.90. Find x.

PLEASE NOTE if you struggle with parts d) and e) of question 4, don’t worry – as indicated in
the key at the top of this document, they are not representative of module quiz questions, so
there’s time to think further before this concept is assessed.

Question 5: Suppose X ~ N(43, 0.5). Suppose we sample 16 values from this distribution.

* a) What is P( X< 42)?

* b) What is P( X< 42)?

* c) What is P( X> 43.25)?

* d) What is P( X> 43.25)?

* e) What is P(43.4< X <44.1)?

* f) What is P(43.4< X <44.1)?


Question 6: Suppose heights of manufactured units of product are well modelled by a normal
distribution with a mean of 100 cm and standard deviation of 0.04 cm.

* a) Suppose we sample 1 unit at random from the production line. What is the probability that
its height is less than 99.9 cm?

* b) Suppose we sample 10 units at random from the production line. What is the probability that
their average height is less than 99.9 cm?

** c) Why was the sampling distribution of the sample mean normal in part b), even though we
were sampling fewer than 31 units?

* d) Suppose we sample 10 units at random from the production line. What is the probability that
their average height is between 99.99 cm and 100.01 cm?

** e) Suppose the engineering design specification indicate that at least 95% of units of product
must have a height between 99.95 cm and 100.05 cm. Are the engineering design
specifications currently being met? Why or why not?

** Question 7: Suppose X is a discrete random variable. Suppose we sample 9 values from this
distribution.

a) Suppose we are interested in finding the probability that the sample mean takes a value less
than 8. Is it appropriate to use the normal distribution for the sampling distribution of the sample
mean here? Why or why not?

b) What would you need to know in order to calculate P( X< 8)?.

** Question 8: Suppose X is a discrete random variable that has mean 75 and standard
deviation 5.

a) Suppose we sample 100 values from X. What is the sampling distribution of the sample
mean? Why can we state the sampling distribution of the sample mean here, even though X is
clearly not normally distributed?

b) Suppose we sample 10 values from X. Is it possible to declare the sampling distribution of the
sample mean, based only on the information given in this question? Why or why not?

c) Is it possible to calculate P(X = 1) based on the information given in this question? Why or
why not?

Common questions

Powered by AI

Z-scores standardize individual data points by considering how many standard deviations the value is from the mean. In a normal distribution, z-scores can be directly linked to percentile ranks as they determine the relative position of a data point within the distribution. This allows for consistent and meaningful comparison across different data sets and is crucial for hypothesis testing and establishing statistical significance.

A standard deviation of zero implies that all values of the random variable are identical, meaning there is no deviation among data points. As variance is calculated by squaring the standard deviation, it too will be zero, representing no dispersion in the dataset, making the statement that variance is large false.

In continuous random variables, the probability of choosing an exact value is zero because there are infinitely many possible values over any interval, which leads to a probability density rather than a probability mass. Conversely, with discrete random variables, specific values can be isolated, enabling a non-zero probability assigned directly to each value within the probability mass function.

The normal distribution assumes continuous data and a sufficiently large sample size to approximate a bell curve accurately. For discrete variables, especially with a small sample size, the normal approximation may not reflect the true distribution shape. Alternatives include using the binomial distribution or employing non-parametric methods that do not presume any specific distribution.

As sample size increases, the likelihood of observing outliers in the dataset may seem higher due to greater coverage of possible values. However, the precision in estimating probability distribution parameters improves, which often indicates that true outliers are less prevalent. Additionally, larger sample sizes provide more reliable statistical power to detect genuine outliers rather than random variation.

When calculating probabilities involving the sample mean of a normally distributed variable, the mean of the distribution remains the same, but the standard deviation becomes the standard error, calculated as the population standard deviation divided by the square root of the sample size. This adjustment reflects reduced variability expected in the sample mean compared to individual observations, making probabilities based on the sample mean more precise.

If engineering standards necessitate a strict height range, the alignment of the normal distribution parameters—mean and standard deviation—within this range is critical. If not properly aligned, a significant portion of units will fall outside prescribed limits, leading to non-compliance with these standards. Adjustments might involve re-calibrating production processes to ensure the mean and variance produce units within compliance requirements, leveraging statistical process control techniques without sacrificing efficiency.

The probability of a normally distributed random variable taking on negative values primarily depends on the mean and standard deviation. If the normal distribution has a mean much greater than zero and a small standard deviation, the probability of negative values is extremely low, though not zero. However, when considering a model that inherently assumes positive outcomes (e.g., height), negative values might indicate model unsuitability or a need to redefine context or transform data.

The Central Limit Theorem (CLT) is crucial as it states that the sampling distribution of the sample mean will approximate a normal distribution as the sample size becomes large, regardless of the shape of the original population distribution. This is particularly important with non-normally distributed discrete variables, as it allows for the use of normal probability models in inferential statistics, facilitating hypothesis testing and confidence interval construction.

A discrete random variable takes countably many possible values, meaning it can take on specific, isolated points rather than a range of values. The probability distribution for a discrete random variable assigns probabilities to each of these specific values, and the sum of these probabilities is 1. Unlike continuous variables that require an integral to find the area under a curve, discrete variables have a probability mass function that directly gives the probability of each individual value.

You might also like