0% found this document useful (0 votes)
9 views3 pages

Statistical Inference for Distributions

The document consists of a series of statistical problems related to estimation, hypothesis testing, and distributions for various probability models including normal, uniform, Poisson, and binomial distributions. It includes tasks such as finding maximum likelihood estimators, complete and sufficient statistics, UMVUEs, and testing hypotheses with specific significance levels. Each problem requires detailed calculations and justifications for the statistical properties and methods applied.

Uploaded by

Kaiho Daniel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
9 views3 pages

Statistical Inference for Distributions

The document consists of a series of statistical problems related to estimation, hypothesis testing, and distributions for various probability models including normal, uniform, Poisson, and binomial distributions. It includes tasks such as finding maximum likelihood estimators, complete and sufficient statistics, UMVUEs, and testing hypotheses with specific significance levels. Each problem requires detailed calculations and justifications for the statistical properties and methods applied.

Uploaded by

Kaiho Daniel
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

1. (27 marks + Bonus: 6 marks) Let (X1 , ...

, Xn ) be a random sample from N (µ, σ 2 ) with


an unknown µ ∈ R and a known σ 2 > 0. Define g(µ) = ecµ with a fixed c =
̸ 0.

(a) (1 mark) Does (X1 , . . . , Xn ) have a multivariate normal distribution? Why? No mark
will be given for an answer of “Yes” or “No”.

(b) (2 marks) Find the distribution of X̄ via the moment generating function. Write down
all intermediate steps.

(c) (2 marks) Does the random vector of (X1 , X̄) have a bivariate normal distribution?
Why? No mark will be given for an answer of “Yes” or “No”.

(d) (2 marks) Find the complete and sufficient statistic for µ.

(e) (5 marks) Find the maximum likelihood estimator of g(µ). Is it unbiased? Why?

(f) (4 marks) Find the limiting distribution of the maximum likelihood estimator for g(µ)
by Central Limit Theorem and Delta method.

(g) (3 marks) Find the Cramer-Rao lower bound for the variance of any unbiased estimator
of g(µ).

(h) (3 marks) Find the UMVUE of g(µ).

(i) (5 marks) Find the variance of the UMVUE of g(µ) and show that it is larger than its
Cramer-Rao lower bound but the ratio of the variance of the UMVUE over the Cramer-
Rao lower bound converges to 1 as n → ∞.
Hint: Taylor series of a real function f (x) that is infinitely differentiable at a real number
a is the power series
∑∞
f (n) (a)
(x − a)n
n=0
n!

(j) (Bonus: 6 marks) Find the UMVUE of P (X1 ≤ a) with a fixed a ∈ R.

2. (11 marks) Suppose that (X1 , . . . , Xn ) are independently and identically distributed according
to the uniform distribution U [0, θ] .

(a) (2 marks) Find the maximum likelihood estimator of θ.

(b) (4 marks) Find the sufficient statistic for θ and then prove that it is complete.

(c) (5 marks) Find the UMVUE of eθ .

1
3. (27 marks) Suppose that {Xi : i = 1, 2, . . . , n} is a random sample from a Poisson distribution
with an unknown mean parameter λ.

(a) (2 marks) Find the complete and sufficient statistic for λ.

(b) (2 marks) Derive the distribution of the complete and sufficient statistic for λ. Write
down all intermediate steps.

(c) (5 marks) Find the UMVUE of P (X1 = 1).

(d) (2 marks) Can the Cramer-Rao lower bound be achieved by any unbiased estimator of
P (X1 = 1)? Why? No mark will be given for an answer of “Yes” or “No”.

(e) (i) (4 marks) Find the UMP test at significance level α for testing H0 : λ ≤ λ0 against
λ > λ0 . Using the Central Limit Theorem to approximate the distribution.

(ii) (3 marks) Consider the specific case H0 : λ ≤ 1 against λ > 1. Using the Cen-
tral Limit Theorem to determine the sample size n so that a UMP test satisfies
Pr(Reject H0 |λ = 1) = 0.05 and Pr(Reject H0 |λ = 1.5) = 0.9.

In addition to {Xi : i = 1, 2, . . . , n}, {Yj : j = 1, 2, . . . , m} is also a random sample from a


population having a Poisson distribution with different mean parameter µ. We are interested
in testing the hypothesis H0 : λ = µ versus H1 : λ ̸= µ.

(f) (2 marks) Find the MLE of the unknown parameter in Θ0 .

(g) (2 marks) Find the MLEs of unknown parameters of λ and µ in Θ.

(h) (3 marks) Derive the likelihood ratio statistic.

(i) (2 marks) Using the asymptotic distribution of the likelihood ratio statistic to construct
a test procedure of rejecting H0 at a significance level α.

4. (20 marks)

(a) (12 marks) Assume that Xi ∼ Binomial(ni , Pi ) for i = 1, 2, 3. We are interested in


testing the hypothesis H0 : P1 = P2 = P3 = P versus H1 : They are not all equal.
i. (1 mark) Write down the MLE of P in Θ0 .

ii. (1 mark) Write down the MLEs of P1 , P2 and P3 in Θ.

iii. (4 marks) Find the likelihood ratio statistic and then derive the large-sample like-
lihood ratio test at significance level α.

2
iv. (2 marks) Find the Pearson’s goodness of fit test statistic and state the critical
region for this test at significance level α.

v. (4 marks) It was found that the numbers of patients with hypertension were 21, 36
and 30 among 69 of nonsmokers, 62 of moderate smokers and 49 of heavy smokers,
respectively. Test whether or not the proportions of patients with hypertension for
these three groups were the same at the significance level of 0.05. State clearly the
hypothesis statements, value of test statistic, critical value and your conclusion for
each test.

A. Use the large-sample likelihood ratio test.

B. Use the Pearson’s goodness of fit test.

(b) (8 marks) Let X = (X1 , ...Xm ) ∼ multinomial (n, θ1 , ..., θm ). Consider testing H0 :
θ1 = θ2 versus H1 : θ1 ̸= θ2 .
i. (1 mark) Write down the MLEs of unknown parameters in Θ0 .

ii. (1 mark) Write down the MLEs of unknown parameters in Θ.

iii. (4 marks) Show that the Pearson’s goodness of fit test statistic can be written as

(X1 − X2 )2
.
X1 + X2
State the critical region of this test at significance level α.

iv. (2 marks) Consider the following table:

After
Agree Disagree

Agree X1 X2
Before
Disagree X3 X4

Let X = (X1 , ...X4 ) ∼ multinomial (n, θ1 , ..., θ4 ). Using the result in (iii), test the
proportion of people who change from agree to disagree is the same as the proportion
of people who change from disagree to agree at the significance level of 0.05 when
X1 = 34, X2 = 19, X3 = 6 and X4 = 16. State clearly value of test statistic, critical
value and your conclusion.

Common questions

Powered by AI

Pearson’s goodness-of-fit test statistic, often expressed as X^2 = ∑(O-E)^2/E, approximates the distribution under the null hypothesis to identify deviations. The likelihood ratio test (LRT) converges in large samples to a chi-square distribution similar to X^2, offering an alternate method to evaluate fit. Both test deviations from the model were attributed to proportions, with LRT focusing on likelihood comparison and Pearson’s on observed versus expected count discrepancies.

The unique property of a sufficient statistic, such as the sum ∑Xi for a Poisson distribution parameter λ, is it captures all available information about the parameter in the sample. This means that conditioned on this statistic, the sample data adds no further insight into λ, thereby enabling simplified inferences directly from the sufficient statistic itself, as seen when developing minimal variance unbiased estimators.

The Central Limit Theorem states that the distribution of a standardized sum of independent random variables approaches a normal distribution as the sample size grows. The Delta method uses this principle to apply a smooth, differentiable transformation to an estimate, enabling derivation of its asymptotic distribution. For an MLE g(θ), the Delta method helps extend the CLT result to non-linear functions, resulting in g(θ̂) being approximately normal with derived variance.

The moment generating function is useful for finding the distribution of the sample mean in a normal distribution because it provides a way to determine the distribution of sums (or averages) of random variables. The MGF of a normal distribution is M(t) = exp(µt + 0.5σ^2t^2), and for the average of n independent normals, the MGF is M(t/n)^n. This demonstrates that the sample mean is distributed normally with mean µ and variance σ^2/n.

The variance of an UMVUE might be larger than the Cramer-Rao lower bound because the lower bound applies to unbiased estimators in general, not necessarily those derived as UMVUEs in finite samples. However, as sample size n increases, the ratio often converges to 1 due to asymptotic efficiency, wherein estimators become closer to achieving the lower bound due to the Law of Large Numbers and central limit considerations.

The maximum likelihood estimator of θ from a uniform distribution U[0, θ] is given by making the sample maximum, X(n), the MLE of θ. This results directly from considering the likelihood function, L(θ), which is proportional to 1/θ^n for the n observations, subject to θ being greater than or equal to the observed maximum. Thus the MLE is the largest observed value, X(n)

The Cramer-Rao lower bound provides a theoretical lower limit on the variance of any unbiased estimator of a parameter. For an unbiased estimator θ̂ of a parameter θ, if I(θ) is the Fisher information, Var(θ̂) ≥ 1/I(θ). This indicates the best variance that can be achieved by any unbiased estimator. Applying it involves calculating the Fisher information and using it to assess the efficiency of any estimator compared to this benchmark.

A random sample (X1, ..., Xn) from a normal distribution with parameters N(µ, σ^2) has a multivariate normal distribution because each Xi is independently distributed as N(µ, σ^2). Therefore, by definition, any linear combination of these independent normal variables will also be normally distributed, satisfying the criteria for a multivariate normal distribution.

To construct a likelihood ratio test for comparing multiple proportions in a binomial distribution, derive the likelihood under the null hypothesis H0 (same proportions) and alternative H1 (differing proportions). Calculate the likelihood ratio statistic, Λ = L(under H0)/L(under H1), and use it to form a test statistic, often a chi-square with degrees of freedom equal to the difference in parameters under H0 and H1. Compare this to a critical value to decide on rejection at a given significance level.

To find the UMVUE of a function like g(µ) = e^(cµ), identify a complete, sufficient statistic for the parameter, and use the Lehmann-Scheffé theorem. For normal samples, the sample mean is complete and sufficient for µ. Then, finding or deriving the unbiased estimator of g(µ) using the Rao-Blackwell theorem ensures the UMVUE.

You might also like