0% found this document useful (0 votes)
5 views14 pages

Sample Distributions and Statistics Exercises

The document contains exercises on sample distributions and the Central Limit Theorem, focusing on calculations of population averages, standard deviations, expected values, and standard errors for various sample sizes. It also discusses the appropriateness of using normal distribution for approximating sampling distributions and includes applications related to Alzheimer's disease, professor salaries, potassium requirements, and car battery lifespans. The exercises require statistical analysis and probability calculations based on given data and distributions.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views14 pages

Sample Distributions and Statistics Exercises

The document contains exercises on sample distributions and the Central Limit Theorem, focusing on calculations of population averages, standard deviations, expected values, and standard errors for various sample sizes. It also discusses the appropriateness of using normal distribution for approximating sampling distributions and includes applications related to Alzheimer's disease, professor salaries, potassium requirements, and car battery lifespans. The exercises require statistical analysis and probability calculations based on given data and distributions.

Translated by

ScribdTranslations
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

UNIVERSITY OF THE VALLEY OF MEXICO

Inferential Statistics

ACTIVITY 1
Exercises On Sample Distributions

Individual

Professor

Date

25 de julio del 2022


EJERCICIOS SOBRE DISTRIBUCIONES MUESTRALES

• Based on the material consulted in the unit, solve the exercises that are proposed.
about the following topics:
sample distributions
Central Limit Theorem (CLT)

• Basic Techniques 1.

1. A population consists of five numbers: 2, 3, 6, 8, 11. Consider all the samples.


possible size two combinations that can be extracted with replacement from this population.
Find:
a) The average of the population
2 + 3 + 6 + 8 + 11 30
∑ = =
5 5

b) The population standard deviation


√ (2- 6 + ) 2 3- 6 (+ 6- )62+ 8- 2 6(
( 6 + )11- )2 ( )2
=
5
√ 6 + 9 + 0 + 4 + 25
=
5
√ 54
=
5
= 10.8

= .
c) The expected value of the sample mean

Samples Sample Mean

(2,3) 2.5

(2,6) 4

(2,8) 5

(2,11) 6.5

(3,6) 4.5

(3,8) 5.5

(3,11) 7

(6,8) 7

(6,11) 8.5

(8,11) 9.5

∑? 60

The expected value of the mean of a probability distribution, and the mean
the expected value of the sample means is:
µ? =∑?/?
µ? = 60/10
µ? = 6
d) The standard deviation (standard error) of the sample mean
=

= 3.2863
√2
= .

Data 2 3 6 8 11
Media 6
Standard Deviation 3.28
Expected Value of the Sample
2.2 2.3 2.6 2.8 2.11
3.2 3.3 3.6 3.8 3.11
6.2 6.3 6.6 6.8 6.11
8.2 8.3 8.6 8.8 8.11
11.2 11.3 11.6 11.8 11.11

The mean of the Sample Distribution is equal to the sum of all the means.
show them 150/25= 6.

The standard deviation of the sampling distribution of the means, that is, the error
typical of tights.

The variance of the sampling distribution of means is obtained by subtracting the value
from the average = 6 of each number of (1), squaring each difference,
adding the 25 obtained numbers and dividing by 25 we get the result
final.

2. Random samples of size n were selected from populations with the means and
variances given here. Find the mean and standard deviation of the distribution of
sampling of the sample mean X in each case:
a)n = 36, µ = 10, σ29
= ?
=10

=

√9
=
√ 36
= .
b)n= 100,µ= 5,σ2= 4
= ?
=5

=

√4
=
√ 100
= .
c)n= 8,µ= 120, σ2equals 4
= ?
? =120

=

√1 1
= =
√8 2.82
= .

If the sampled populations are normal, what is the sampling distribution of X


for items a, b, and c?

If the sampled population is normal, then the sampling distribution of X is also


It will be normal regardless of the size of the sample that is chosen.

According to the Central Limit Theorem, if the sampled populations are not
normal, what can be said about the sampling distribution of X for items a,
b y c?

If random samples of 'n' observations are obtained from a non-normal population


with finite mean 'm' and standard deviation 's' then when 'n' is large,
sampling distribution of the mean "m" and standard deviation, the approximation is
becomes more accurate as 'n' becomes large.

3. A random sample of n observations is selected from a population with


standard deviation 1. Calculate the standard error of the mean (SE) for the following
values den.
1
a)n= 1 = =
√ √1
1
b)n= 2 = = .
√ √2
1
c)n= 4 = = .
√ √4
1
d)n= 9 = = .
√ √9
1
e)n= 16 = = .
√ √ 16
1
f)n= 25 = = .
√ √ 25
1
g)n= 100 = = .
√ √ 100

a b c d e f g
Data n=1 n=2 n=4 n=9 n=16 n=25 n=100
Results 1 0.7 0.5 0.33 0.25 0.2 0.1
4. Random samples of size n were selected from binomial populations with
population parameters p given here. Find the mean and standard deviation of
the sampling distribution of the sample proportion p̂ in each case:
a) n = 100, p = 0.3

R( 1− )
=√

0.3(1 - 0.3)
=√ = .
100

b) n = 400, p = 0.1

R( 1− )
=√

P( 1- 0.1 )
=√ = .
400
c) n = 250, p = 0.6

R( 1− )
=√

R( 1-0.6 )
=√ = .
250

5. Is it appropriate to use the normal distribution to approximate the sampling distribution?


of P in the following circumstances?
a) n = 50, p = 0.05
If it is appropriate, because the condition n≥30 is met

b) n = 75, p = 0.1
It is adequate because the condition n≥30 is met.

c) n = 250, p= 0.99
It is appropriate because the condition n≥30 is met.
• Applications

1. Alzheimer's disease. The duration of Alzheimer's disease from the beginning


the time from symptoms to death ranges from 3 to 20 years; the average is 8 years with a
standard deviation of 4 years. The manager of a large medical center selects randomly,
from the central database, the medical records of 30 Alzheimer's patients already
deceased and note the duration of the illness for each unit in the sample. Find the
approximate probabilities for the following events:
If the duration of the illness in years is a random variable with a normal distribution,
the arithmetic mean also has a normal distribution, thus using a standardization and the
The table of accumulated probabilities allows us to answer the following items.

The duration of Alzheimer's disease from the onset of symptoms until the
deaths have a normal distribution with:

Mean = m = 8 years and standard deviation σ = 4 years

To find possibilities associated with this distribution, a probability table is used.


accumulated calculated as areas under the standard normal curve (z).

The average duration is less than 7 years


We want to find the probability that the mean is less than 7, then we have
there is a probability of 0.0853 that the disease on average lasts 7 years
or less.

̅
(?−?) ′(?−?)
?< ? ( ) = {? < ?
⁄ ? ⁄ ?
√ √

(7-8)
?< 4 ( ) P(? < -1.3693)

√ 30

(-1)
?< 4⁄ ( ) = 0.5 − 0.4147
5.4772

(-1)
?< ( )= .
0.7303

? < -1.3693
b) The average duration exceeds 7 years
The probability that the mean is greater than 7 given that the table shows
cumulative probabilities is necessary to work with the complementary event for
obtain the result of the distribution.

( ¿̅−¿ ) ′(?−?)
¿< ? ( ) = {? < ?
⁄ ? ⁄ ?
√ √
̅
(7-8)
?< 4 ( ) P(? < -1.3693)

√ 30
(-1)
?< 4⁄ ( ) = 0.4147+ 0.5
5.4772
(-1)
?< ( )= .
0.7303

? < -1.3693

c) The average duration is no more than one year from the population mean. =8
When working with intervals, probabilities are obtained by differences of
the cumulative probabilities at the left tail of the extremes of that
interval, then the probability that the average duration of the illness
this between 7 and 9 years is 0.8294.

(7 < z < 9

P(-1.3693 < z < 1.3693)

( ) 0.9147-0.0853

( )= .

Beginnings Range of Years n p σ Z Table Z P %


A 03 to 20 30 8 4 -1.37 ( <-1.37) 0.0853 8.53%
B -1.37 (x>7)=1-p(x<7)=1-0.0853 0.9147 91.47%
P(7<x<9)=p(-1.37<z<1.37)
C -1.37 0.8294 82.94%
=0.9147-0.0853

Graph the standard error of the mean (SE) against the sample size n and connect the points
with a smooth curve. What is the effect of increasing the sample size on the error?
standard?
Deviation: 4

N Standard Error
30 0.73029674
60 0.51639778
1000 0.12649111
50000 0.01788854
200000 0.00894427

2. Salaries of professors. Suppose that professors at a university in the U.S.A. -with rank
as a teacher in public institutions that offer two-year academic programs,
they earn an average of $71,802 per year, with a standard deviation of $4,000
dollars. In an exercise to verify this salary level, a random sample was selected
of 60 teachers from a database of the academic staff of all the institutions
public institutions offering two-year programs in the U.S.

a) Describe the sampling distribution of the sample mean X.


It would be a Normal Distribution

b) What limits would be expected for the sample mean,


probability 0.95?
Alpha = 1-0.95 (significance level)
Critical value alpha/2 = 0.025
Value in tables 1.96 (from center to extremes)
Confidence interval:
)
(2 * sigma
( μ) 95%= ±

1.96 ∗ 4000
( μ) 95%= 71.802 ±
√ 60
( μ) 95%= 71.802 ± 1012.139648

c) Calculate the probability that the sample mean x is greater than $73,000.
= 73,000

=
Value of Tables = 0.6141 ( )
( < 73000 ) = 0.6141
( < 73000 ) = 1 − 0.61410.3859 (38.59%)

d) If a random sample actually produced a sample mean of $73,000,


Would you consider this uncommon? What conclusion would you draw?
The data that is somewhat distant from the average given that it is
in the maximum upper values above 72814.

3. Potassium Requirement. The normal daily requirement of potassium in humans


it is in the range of 2,000 to 6,000 milligrams (mg), with larger amounts
necessary during the hot summer months. The amount of potassium in different
food varies but measurements indicate that the banana contains a high level of potassium,
with approximately 422 mg in a medium-sized banana. Assume that the
The distribution of potassium in bananas is normally distributed, with a mean equal to 422 mg.
and a standard deviation of 13 mg per banana. You eat n = 3 bananas a day and T is the
total number of milligrams of potassium received from them.

The amount of potassium in different foods varies, but measurements indicate that the
banana contains a high level of potassium, with approximately 422 mg and deviation
standard of 13 mg per banana.

You eat 3 bananas a day and 'T' is the total number of milligrams of potassium you receive from
they.

a) Find the mean and the standard deviation of T.


T = total number of milligrams of potassium.
M= 422mg
S = 13mg
b) Find the probability that your daily potassium intake from the three bananas exceeds
of 1,300 mg. (Suggestion: Notice that T is the sum of three random variables X1, X2, X3)
X X, and where X1 is the amount of potassium in banana 1, etc.
Find that the probability of your daily potassium intake from three bananas
exceeds 1,300mg.

( ̂− )
( 434 = ) ( < )

4. Duration of car batteries. A car battery manufacturer claims that


the distribution of the duration time (lifespan) of the batteries of your best brand
it has an average 54 months and a standard deviation = 6 months. Assume that a group
Consumers decide to verify the claim and for that purpose buy a sample of 50.
batteries and subjects them to testing to measure their lifespan.

a) Assuming that the manufacturer's claim is true, describe the distribution of


sampling of the sample mean when n = 50 batteries.
The distribution would be a random sampling distribution, which tends to be a
normal distribution. (Add a random sample distribution to that).

b) Assuming that the manufacturer's claim is true, what is the probability of


that the sample of 50 batteries has a lifespan of 52 months or less?
( ̂− )
( < 52 )= ( < )
()

( 52 - 54 )
( < 52 )= ( < ) equals negative 2.3570
6
( )
√ 50

P(X<52) = 0.00939 = It translates to a 0.9% probability of obtaining a battery


have a lifespan of

5. Body temperature. Suppose that the body temperature of healthy individuals


distributes approximately normally with a mean of 37.0 C and a standard deviation of 0.4 C.

a) If 130 healthy people are randomly selected, what is the probability of


should the average temperature for these individuals be 36.80 or lower?
X= Temperatura corporal promedio
M= 37° C
S= 0.4 C
N= 130

( 36.80 − 37 )
( < 36.80 )= ( < ) = −5.70008 = 0%
0.4
√ 130

b) Would you consider an average temperature of 36.80 as unlikely to occur,


if the true average temperature of healthy people is 37 C?
It is a little unlikely

6. Cost of an apartment. The average cost of an apartment in the Cedar development


Lakes is $62,000 USD with a standard deviation of $4,200 USD.

a) What is the probability that an apartment in this development costs at least


$65,000 USD?
65,000 − 62,000
( > ) = 0.7142
4,200
( 65,000 = )0.7142 = 0.23755

b) The probability that the average cost of a sample of two apartments is


at least $65,000 USD is greater than or less than the probability that a
How much does that apartment cost? How much does it differ?
The amounts differ so much because they deviate from the average.

7. Tossing a coin. A fair coin is tossed n = 80 times. Let p ̂ be the proportion


sample of faces (soles). Find P p (0.44 0.61)

(
1− )
̅= √

0.5 ( 1- 0.5 )
̅= √ = 0.0559
80
P (0.44 < p < 0.61)
P (Z1 < p < Z2)
̅ −
=
0.44 − 0.5
Z1 = = 1.0733 Values in Table = 0.1423
0.0559
0.61 − 0.5
Z2 = 1.9677 0.02455
0.0559

Z1-Z2 = 0.11775
[Link] tools. It has been found that 2% of the tools produced
Certain machines have a defect. What is the probability that in 400 of them
tools

a) Do 3% or more have any defects?


N>30 If Complies
Np = 0.02*400 = 8 It is met
N (1-p) = 400 (1-0.02) = 392 If It Is Met

0.02 (1-0.02 )
̅= √ = 7X10-3
400
0.03 − 0.02
= = 0.9222It(is a 1 Values in
) Table = 0.0888
0.007

b) Do 2% or less have any defect?


0.02 − 0.02
= = 0 Values in Table = 0.500
0.007

Conclusion:

In statistics, a parameter is a number that summarizes a large amount of data that


can be derived from the study of a statistical variable. The calculation of this number is
well defined, usually through an arithmetic formula obtained from data of
the population; In statistics, a statistic is a quantitative measure derived from a
data set from a sample, with the aim of estimating or inferring characteristics of a
population or statistical model.

In statistics, the sampling distribution is what results from considering all the samples.
possibilities that can be taken from a population. Its study allows to calculate the
probability of approaching the parameter given a single sample
population.

The standard error is the standard deviation of the sampling distribution of a statistic.
sample. The term also refers to an estimate of the standard deviation,
derivative of a particular sample used to compute the estimate.
References

▪ Mathematical Statistics with


Applications (7th ed.). Mexico, Mexico: Cengage Learning.
▪ Devore, J. L. (2016). Probabilidad y Estadistica para Ingenieria y Ciencias (9 ed.). Cengage
Learning. Retrieved from[Link]
▪ McClave, J., & Sincich, T. (2014). Statistics (12 ed.). Harlow: Pearson.
▪ Mendenhall, W. I., Beaver, R. J., & Beaver, B. M. (2015). Introduction to Probability and
Statistics (14th ed.). Mexico City: CENGAGE Learning.
▪ Sweeney, D. J., Anderson, D. R., & Williams, T. (2011). Estadistica para Negocios y Economia
(11th ed.). Cengage Learning. Retrieved from[Link]

* * *

You might also like