Statistical Inference
Statistical Inference
STATISTICAL INFERENCE
We know that statistical data is nothing but a random sample of observations drawn from a
population described by a random variable whose probability distribution is unknown or partly
unknown and we try to know about the properties of the population on the basis of knowledge of the
properties of the sample. This inductive process of going from known sample to the unknown
population is called ‘Statistical Inference ‘
Formally, let x be a random variable describing the population under investigation. Suppose X has
þ.m.ƒ 𝑓𝑜(𝑥) = 𝑃(𝑥 = 𝑥) or þ d ƒ 𝑓𝑜(𝑥) which depend on some unknown parameter 𝜃 (single or vector
valued) that may have any value in a set Ω (called the parameters space). We assume that the
functional form of 𝑓𝑜(𝑥) is known but not the parameter 𝜃(except that 𝜃 ∈ Ω). For example, the family
of distributions {𝑓𝜃(𝑥), 𝜃 ∈ Ω} may be the family of Poisson distribution {𝑃(𝜆), 𝜆 ≥ 0} or normal
distribution {𝑁(𝜇, 𝜎2),−∞ <𝜇 < ∞, 𝜎 ≥ 0}
POINT ESTIMATION
Definition: A random sample of size ‘n’ from the distribution of X is a set of independent and
identically distributed random variables {𝑥1, 𝑥2, … , 𝑥𝑛} each of which has the same distribution as that
of X. The probability of the sample is given by
Definition: A statistic T = T (x1,x2,…, xn) is any function of the sample values, which does not depend
on the unknown parameter 𝜃. Evidently, T is a random variable which has its own probability
distribution (called the ‘ Sampling distribution’ of T)
If we use the statistic T to estimate the unknown parameter 𝜃, it is called the estimator (or point
estimators) of 𝜃 and the value of T obtained from a given sample is its ‘estimate’
Then,
(I)E(̅𝑋)̅ =𝜇
Prof: We have
= 𝑛𝜎2 − 𝑛𝜎2/𝑛
= (𝑛 − 1)𝜎2
PROPERTIES OF ESTIMATORS
UNBIASEDNESS:
𝐸𝑥𝑎𝑚𝑝𝑙𝑒. If (𝑥1, 𝑥2, … , 𝑥𝑛) is a random sample from any population with mean 𝜇 and variance 𝜎2, the
sample mean 𝑥̅ is an unbiased estimator of 𝜇 but the sample variance 𝑆2 is not an unbiased estimator
of 𝜎2.
𝐸𝑥. if (𝑥1, 𝑥2, … 𝑥𝑛) is a random sample from 𝑎 normal distribution 𝑁(𝜇, 𝐼) show that 𝑇 =
is an unbiased estimator of 𝜇2 ,
Soln.
𝑬𝒙𝒂𝒎𝒑𝒍𝒆: Let (𝑥1, 𝑥2, … 𝑥𝑛) be a random sample of observation from a Bernoulli distribution ƒ𝜃(𝑥) =
Soln: We know that 𝐸(𝑥𝑖) = 𝜃 and 𝑉(𝑥𝑖) = 𝜃(1 − 𝜃) so that 𝐸(𝑌) = 𝑛 𝜃 and 𝑉(𝑌) = 𝑛𝜃(1 − 𝜃)
Now
= 𝑛𝜃(1 − 𝜃) + 𝑛2𝜃2 − 𝑛𝜃
= 𝑛(𝑛 − 1)𝜃2
E(T)=
Example: Show that the mean 𝑥̅ of a random sample of size 𝑛 from the exponential distribution
Example: Let (𝑥1, 𝑥2, … 𝑥𝑛) to a random sample from a normal distribution with mean 0 and variance
𝜃 (0< 𝜃 < ∞) so that is an unbiased estimator of 𝜃 and has variance 2𝜃2/n Sohm we know
that
4
Also
Example Let (𝑥1, 𝑥2, … 𝑥𝑛) be a random sample from the rectangular distribution 𝑅(0, 𝜃) having
þ, 𝑑, 𝑓
𝑑. 𝑓. of Yn is-
𝐹𝓎(𝓎) = 𝑃(𝑌𝑛 ≤ 𝓎)
= 𝑃(𝑥𝑖 ⪕ 𝓎, 𝑥𝑛 ⪕ 𝓎)
= [𝑃(𝑥 ⪕ 𝓎)]𝑛
þ, 𝑑, 𝑓 𝑜𝑓 𝑌𝑛 is-
5
Hence,
Or
𝐹𝑌𝑖(𝓎) = 𝑃{𝑌𝑖 ≤ 𝓎}
þ, 𝑑, 𝑓 of 𝑌𝑖is
Hence,
Example: Let ((𝑥1, 𝑥2, … 𝑥𝑛) be a random variable from the Rectangular distribution 𝑅(𝜃, 2𝜃) having
þ, 𝑑, 𝑓
6
Show that
Example: Let 𝑦1, 𝑦2,𝑦3 be the order statistics of a random sample of size 3 from 𝑎 uniform
*If 𝑦1,𝑦2,… , 𝑦𝑛 are two unbiased estimator with variance and correlation coeff. P between than
the linear combination which is unbiased and has minimum variance is.
*If 𝑦1,𝑦2,… , 𝑦𝑛 are ind ept unbiased estimators if 𝜃 with variance the linear
combination with minimum variance is
Where
Example Let ‘T’ be an unbiased estimator of . Does it imply that and , are unbiased for
respectively?
Soln :
7
Also,
Sohn Let
Evidently, or
Then
Minimising we get
Or
Remarks: An unbiased estimator may not exist. Let x be a random variable with Bernoulli
distribution.
Let X be a random variable having Poisson distribution 𝑃(𝑥) and suppose we want estimator 𝓰(𝜆)
=ℯ3𝜆. Consider a sample of one observation and the estimator T= . Then E(T)= ℯ−3𝜆 so that T is an
unbiased estimator of ℯ−3𝜆 but T(x)= (-2) X for x even and T(x) < 0 for 𝑥 odd, which is absurd since
ℯ−3𝜆 is always positive.
(𝑖𝑖𝑖) Instead of the parameter 𝜃 we may be interested in estimating a function 𝓰(𝜃). 𝓰(𝜃) is said to
be ‘estimable’ if there exists an estimator T Such that E(T)= 𝓰(𝜃), 𝜃 ∈ 𝛺.
Minimum Variance Unbiased (MVU) estimators : The class of unbiased estimators may, in general,
be quite large and we would like to choose the best estimator from this class. Among two
8
estimators of 𝜃 which are both unbiased , we would choose the one with smaller variance. The
reason for doing this rests on the interpretation of variance as a measure of concentration about the
mean. Thus, if T is unbiased for 𝜃, then by Chebyshev’s inequality-
Therefore, the smaller 𝑉𝑎𝑟(𝑇) is, the larger the lower bound of the probability of concentration of T
about 𝜃 becomes. Consequently, within the restricted class of unbiased estimators we would choose
the estimator with the smallest variance.
(UMVU) estimator of 𝜃 (or an estimator for 𝓰(𝜃) if it is unbiased and has the smallest variance
within the class of unbiased estimators of 𝜃 (or 𝓰(𝜃),) of all 𝜃 ∈ 𝛺. That is if T is any other unbiased
estimator of 𝜃, then-
Suppose we decide to restrict ourselves to the class of all unbiased estimators with finite variance.
The problem arises as to how we find an UMVU estimator, if such an estimator exists. For this we
would first determine a lower bound for the variances of all estimators (in the class of unbiased
estimators under consideration) and then would try to determine an unbiased estimator whose
variance is equal to this lower bound. The lower bound for the variances will be given by the Cramer-
Rao inequality for which we assume the following regularity conditions:
May be differentiated under the integral sign where T(X1, Xn) is any unbiased estimator of 𝜃
Cramer-Rao inequality: Let (X1,…, Xn) be a random sample of n observations on X with þ. 𝑑. 𝑓 𝑓(𝑥;
𝜃) and suppose the above regularity conditions hold. If T is any unbiased estimator of 𝜃, then-
Proof: We have
9
Or
Or
But
Where
And
10
=1
Proof:
I𝓯𝓯
𝑅(𝑇, 𝑍) = 1
𝑇 = 𝑎𝜃 + 𝑏𝜃𝑧
But 𝐸(𝑇) = 𝑎𝜃 = 𝜃
𝑖. 𝑒 T= 𝜃 + 𝑏𝜃𝑧
CRB
Therefore, we have
(𝑖𝑖) If 𝓰(𝜃) is an estimable function for which an unbiased estimator is T (𝑖. 𝑒. 𝐸(𝑇) = ℊ(𝜃)) then
C.R Inequality becomes-
(𝑖𝑣) If an unbiased estimator exists which is such that its variance is equal to the lower bound
CRB= then it will be UMVUE.
.
(𝑣) If there is no unbiased estimator whose variance equals the C R B it does not mean that
UMVUE will not exist. Such estimators can be found (if these exists ) by other methods.
(𝑣𝑖) In case of distributions not satisfying the regularity conditions (e.g.: Rectangular distribution)
UMVU estimators, if these exists can be found by other methods. For such cases UMVU estimator may
have variance less than CRB.
Example: Let (𝑥1, . . . 𝑥𝑛) be a random sample from a Bernoulli distribution 𝑓(𝑥; 𝜃) = 𝜃𝑥(1 −
So that
By CR inequality we have C R B
12
Soln:
So that
Example: Let (𝑥1, . . , 𝑥𝑛) be a random sample from a normal distribution 𝑁( 𝜃 , 𝜎2) where variance 𝜎
is known show that 𝑥̅ is UMVUE of 𝜃.
Soln:
13
Or
Example Let 𝑥1, . . , 𝑥𝑛 be a random sample from a normal distribution 𝑁(𝜇, 𝜃) where 𝜇 is known and
𝜃 is that variance to be estimated. Show that is UMVUE of 𝜃
(𝑥−𝜇)2
Soln:
Or
Consider the estimator for which E(S2)= 𝜃 and V(Sso that S2 is UMVUE of
𝜃
𝑛
Example An UMVU estimator is unique , in the sense that if T O and TI are both UMVU estimator
then TO = TI almost surely (𝑖. 𝑒 𝑃(𝑇𝑂 ≠ 𝑇𝐼) = 0)
By definition, 𝑉(𝑇) ≥ 𝑉(𝑇𝑂). It follows that 𝜌 ≥ 𝐼. Therefore 𝜌 =I so that, for every 𝜃, 𝑇𝑂 and 𝑇𝐼 are
linearly related, 𝑖. 𝑒.
𝑇𝑂 = 𝑎 + 𝑏𝑇𝐼
Where 𝑎, 𝑏 are amstants (may depend on 𝜃) and b≥ 0 . 𝑇aking expectation and variance we get
𝑇0 = 𝑇
CONSISTENCY
Definition: A sequence of estimator {𝑇𝑛}.𝑛 = 1,2, … of a parameter 𝜃 is said to be consistent if, as n→∞
𝑇𝑛 → 𝑝 𝜃 for each fixed 𝜃 𝜖 𝛺 that is , for any 𝜖(> 0)
𝑇𝑛 𝑐𝑜𝑛𝑣𝑒𝑟𝑔𝑒𝑠 𝑡𝑜 𝜃 𝑖𝑛 𝑝𝑟𝑜𝑏𝑎𝑏𝑙𝑖𝑡𝑦
Or 𝑃{|𝑇𝑛 − 𝜃| > 𝜖} → 0
Or 𝑃{|𝑇𝑛 − 𝜃| ≤ 𝜖} → 1
𝑎𝑠 𝑛 → ∞
Remarks:
(𝒊) For increase in sample size a consistent estimator will become more and more close to 𝜃
(iv) We will show later that if {𝑇𝑛} is a sequence of estimators such that 𝐸(𝑇𝑛) → 𝜃 and 𝑉(𝑇𝑛) → 0
and 𝑛 → ∞ then {𝑇𝑛} is consistent.
Examples:
1. Let (𝑥1, … 𝑥𝑛) be a random sample from any distribution with finite mean 𝜃.
Then it follows from LLN that 𝑥̅ so that ̅𝑥̅→̅ is consistent for 𝜃. If the
distribution has finite variance
15
(𝜎2, 𝑠𝑎𝑦) 𝑉(𝑥̅) = 𝜎2⁄𝑛 → 0 so that it follows from Remark (IV) that 𝑥̅ is consistent .it can be shown
Let (𝑥1, … 𝑥𝑛) be a random sample from rectangular distribution. 𝑅(𝑂, 𝜃) and let 𝑌𝑖 = 𝑚𝑖𝑛(𝑥1, … 𝑥𝑛)
consider the estimator 𝑇 = (𝑛 + 1)𝑌1. This is unbiased . Now for a any 𝐸(> 0),
𝜖 𝜖
(𝑒𝜃 − 𝑒−𝜃)
𝑛→∞
𝑃{[𝑇 − 𝜃]𝜖} + 1
By remark (iv) above 𝑠2 + 𝑠′2 are both constant for is biased and 𝑠′2 is unbiased.
1 𝑥𝑥þ−1(𝑥 ≥ 𝜃, 𝜃 > 0)
þ 𝑘𝑛𝑜𝑤𝑛 f(x, 𝜃) = 𝜃þΓ(þ) 𝑒𝜃
Soln:
𝑛þ
𝐸(𝑇𝑛) = 𝜃𝑛 → 𝜃
And 𝑉(𝑇𝑛) → 0
𝑃{|𝑇𝑛 − 𝜃| ≤ 𝜖1} → 1
As n→∞
Also , since 𝓰(𝜃) is a continuous function , given 𝜖(> 0)we can choose 𝜖1(> 𝑜)such that
Therefore ,
𝑃{|ℊ(Tn) − ℊ(𝜃)| ≤ 𝜖} → 1
(ii) If {Tn} is consistent for 𝜃(R and non-negative) then √𝑇𝑛 is consistent for √𝜃. Proof For
(iii) If {Tn} is consistent for , then {𝑇𝑛 ± 𝑇′𝑛} is consistent for 𝜃 + 𝜃′.
As n→∞.
(iv)if 𝑇𝑛 and are consistent for 𝜃 and 𝜃’ respectively , is consistent for 𝜃𝜃′. Proof:
we can write
EFFICIENCY:
If 𝑇1 and 𝑇2 are two unbiased estimators of a parameter 𝜃 , each having finite variance 𝑇1 is said to be
more efficient then 𝑇2 if 𝑉(𝑇1) >𝑉(𝑇2). The (relative) efficient of 𝑇1 relative to 𝑇2 is defined by
It is used to judge the efficiency of an unbiased estimator by comparing its variance with the
Cramer- Rao lower bound (C R B) .
18
Definition: Assume that the regularity condition of CR inequality hold (we call it a regular situation)
for family{𝑓(𝑥, 𝜃), 𝜃 ∈ 𝛺}. An unbiased estimator T* of 𝜃 is called most efficient if 𝑉(𝑇∗) equals the
CRB. In this situation, the ‘efficiency’ of any other unbiased estimator T of 𝜃 is defined by
(𝑎) regular situation when there is no unbiased estimator whose variance equals the CRB but
an UMVUE exists and maybe found by other methods.
(b)Non-regular situations when an UMVUE exists and may be found by other methods
(ii)The UMVUE is ‘most efficient‘ estimator in the examples considered earlier all UMVUE, whose
variances equalled CRB are most efficient
Example Consider 𝑎, 𝑟, 𝑠(𝑥1, … 𝑥𝑛) from a normal distribution𝑁(𝜇, 𝜃) where mean 𝜇 is known and
variance 𝜃(0 < 𝜃 < ∞ ) is to be estimated
We has seen that is UMVUE of 𝜃 for which the variance is equal to CRB and
Asymptotic efficiency: As different from the above definition of efficiency we may define efficiency
Let us confine ourselves to consistent estimators which are asymptotically normally distributed.
Among this class, the estimator with the minimum asymptotic variance is called the ‘most efficient
estimator’. It is also called best asymptotically normal (BAN) or consistent asymptotically normal
efficient (CANE) estimator it we denote by avar(T*) the asymptotic variance of a BAN estimator T*
then the efficiency of any other estimator T (within the class of asymptotically normal estimators) is
defined by
Example: Let (𝑥1, … , 𝑥𝑛) be a random sample from a normal distribution 𝑁(𝜇, 𝜎), Consider the ‘most
efficient estimator 𝑥̅ and another estimator 𝑥̅me. It can be show that both are CAN estimator. We have
19
And
Example: let T1, T2 be two unbiased estimators of 𝜃, having the same variance. Show that the
correlation coefficient ρ between T1, T2 cannot be smaller than 2e-1, where e is the efficiency of each
estimator,
Its variance is
Example: let 𝑇𝑜 be an UMVME (or most efficient estimator) where 𝑇1 an unbiased with efficiency ‘e’.
If 𝜌 is the correction coefficient between 𝑇𝑜 and𝑇1, then show that . Soln: we have
𝑒 = 𝑉(𝑇𝑜)/𝑉(𝑇1)
Or 𝑉(𝑇1) = 𝑉(𝑇𝑜)/𝑒
(Which the linear combination of 𝑇𝑜, 𝑇1 with minimum variance) then T is also unbiased, having
variance
Or
20
Since are both non-negative 𝑉(𝑇) ≤ 𝑉(𝑇𝑜) but since 𝑇𝑜 is UMVUE, 𝑉(𝑇) 𝑉(𝑇𝑜).
therefore 𝑉(𝑇) = 𝑉(𝑇𝑜) , and 𝜌 = √𝑒
SUFFICIENCY CRITERION:
A preliminary choice among statistics for estimating 𝜃 , before having for a UMVUE as BAN estimator,
can be made on the basic of another enter on suggested by R.A fisher. This is called ‘sufficiency’
criterion.
Definition: let (𝒙𝟏,… , 𝒙𝒏) be a random sample from the distribution of X having þ, 𝑑, 𝑓 𝑓(𝑥, 𝜃) 𝜃 𝜖
𝛺.A statistic 𝑇 = 𝑇(𝒙𝟏,… , 𝒙𝒏) is defined to be sufficient statistic if and only if the conditional
distribution of (𝒙𝟏,… , 𝒙𝒏) given T=t does not depend on 𝜃, for any value t.
[Note: In such a case if we know the value of the sufficient statistic T, then the sample values are not
needed to tell us anything more about 𝜃].
Also the conditional distribution of any other statistic T (which is not for 𝛺 tray) given T is
independent of 𝜃.
A necessary and sufficient condition for T to be sufficient for 𝜃 is that the joint þ. 𝑑, 𝑓 of (𝒙𝟏,… , 𝒙𝒏)
should be of the form
Where the first term on 𝑟, ℎ, 𝑠., depends on T and 𝜃 and the second them is independent of 𝜃. 𝑇his is
known as Nyman’s Factorisation Theorem which provides a simple method of judging whether a
statistic T is sufficient
Remark: Any one to one function of a sufficient statistic is also a sufficient statistic
Example: Consider n Bernoulli trials with probability of success P. The associated Bernoulli random
variables (𝑥1, … , 𝑥𝑛) have common distribution given by
Where
Where
Hence.
Example: let (𝑥1, … , 𝑥𝑛)be a random sample from a Normal population 𝑁(𝜇, 𝜎). Case
Where
Which shows that an jointly sufficient for [𝜇, 𝜎] Similarly,[ 𝑥̅, ∑(𝑥𝑖, 𝑥)2 / n-1]are also
sufficient for [𝜇, 𝜎],
Example let (𝑥1, … , 𝑥𝑛) be a random sample from a gamer distribution having þ, 𝑑, 𝑓
22
We have
We can write
We can write
Case III : Both 𝜃 and þ are unknown it is seen that are jointly sufficient for (𝜃, þ)
Example: let (𝑥𝑖, 𝑥𝑛, ) be a random sample from the experiential distribution
Example let (𝑥1, … , 𝑥𝑛)be a random sample from the distribution with þ, 𝑑, 𝑓
𝑓(𝑥, 𝜃) = 𝜃𝑥𝜃−1, 𝜃 ≤ 𝑥 ≤ 1 We
have
We have
For no single statistics T it is possible to express the above in the form ℊ[T, θ]𝒽(𝑥𝑖, 𝑥𝑛, ) . Hence there
exists no statistic T which taken alone is sufficient for θ. However the whole set (𝑥1, … , 𝑥𝑛) or the set
of order statistics (𝑥(1), … , 𝑥(𝑛))is jointly sufficient for θ
23
Example let (𝑥1, … , 𝑥𝑛) be a random sample from the Rectangular distribution 𝑅(0, 𝜃) having þ, 𝑑, 𝑓.
We have
But
Where 𝑋(1) and 𝑥(𝑛) are the minimum and maximum of sample values(𝑥1, … , 𝑥𝑛) Therefore,
we can write
Example : If x has þ, 𝑑, 𝑓
Example Let (𝑥1, … , 𝑥𝑛) be a random sample from the rectangular distribution 𝑅(𝜃1, 𝜃2) having
þ, 𝑑, 𝑓
The
Where We
can write
24
= 𝓰[𝑥(𝑖),𝑥(𝑛),𝜃1𝜃2]𝓱(𝑥𝑖, 𝑥𝑛)
Where
Example: let ((𝑥1, … , 𝑥𝑛)) be 𝑎, 𝑟, 𝑠 from the rectangular distribution R (-𝜃, 𝜃).
Then
Example: [𝑥(1), … , 𝑥(𝑛)] are jointly sufficient for and 𝑅(𝜃, 𝜃 + 1) Example:
METHHODS OF ESTIMATION:
For important methods of obtaining estimators are (I) methods of moments,(II) methods of
maximum likelihood (III)method of minimum χ2 and (IV) method of least squares.
(I)Method of moments
Whose solution is say , where is the estimate of Those are the method of
moments estimators of the parameters.
Example let
The equation
Let
∑𝑛𝑖 𝑥 𝑖2 ∑𝑛𝑖 (𝑥 𝑖 −𝑥 )2
𝜎=√ 𝑛
− 𝑥 2 =√ 𝑛
The equation
Remark: (I) the method of moments estimators are not uniquely defined. We may equate the central
(II) These estimator are not, in general, consistent and efficient but will be so only if the parent
(III) When population moments do not exist auchy population) this method of estimation is
inapplicable.
We want to know from which distribution . for what value of is the likelihood largest for this set
of observations. In other words we want to find the value of , denoted by which maximizes
say
In many cases it would be more convenient to deal with log , rather then , since log is
maximized for the some value of as . For obtaining we find the value of for which
27
We must however, check that this provides the absolute maximum. It the derivate dose not exists at 𝜃
= 𝜃 or equation (1) is not solvable this method of solving (1) will fail.
And
Or
Or 𝑒 = ∑𝑛𝑖 𝑥𝑖/𝑛=𝑥̅
𝑚. ℓ. 𝑒 of 𝜃 is 𝜃 = 𝑥̅
𝑛𝑥
𝑛𝑥
Then
And log !
Or
Example: Let 𝑥1, . . , 𝑥𝑛) be 𝑎, 𝑟, 𝑠 from the truncated Binomial distribution having þ, 𝑑, 𝑓
28
Then
And
−2𝑛𝜃(1 − 𝜃)2] = 𝜃
Or Or
𝑛 2
Then
And
Or
𝑚. ℓ. 𝑒 Of 𝜇 = 𝑥̅
29
Then
And
Or Equating
to zero we get
2 ∑𝑛𝑖 (𝑥 𝑖 −𝜇 𝜃 )
𝑚. ℓ. 𝑒 Of 𝜎 is 𝜎=√ 𝑛
𝑛 2
Then
And
And
∑𝑛𝑖 (𝑥 𝑖 −𝑥 )2
√
Equating to zero both the derivatives and solving the equations we get 𝜇 = 𝑥̅ and 𝜎 = 𝑛
𝑚. ℓ. 𝑒 are 𝜇 = 𝑥̅ and
Then And
30
Then
If we differentiate and equate to zero we get which does not yield any
result. Now is maximized by choosing the maximum value of subject to the condition
Example: X has
Then
And log 𝐿(𝜃) = 𝑛𝑙𝑜𝑔 𝜃 + (𝜃 − 1)∑𝑛𝑖 𝑙𝑜𝑔𝑥𝑖
Or
Then
0 ≤ 𝑥(1) ≤ ⋯ ≤ 𝑥(𝑛) ≤ 𝜃
𝑚. ℓ. 𝑒 of 𝜃 is𝜃 = 𝑥(𝑖)
Then
𝑚. ℓ. 𝑒 of 𝜃 is 𝜃 = −𝑥(𝑖)
Example: Let(𝑥1, . . . 𝑥𝑛) be 𝑎, 𝑟, 𝑠 from the regular distribution 𝑅(𝜃1, 𝜃2) having þ, 𝑑, 𝑓
Then
In maximized when (𝜃2 − 𝜃1) is minimum 𝑖, 𝑒𝜃1 is maximum and 𝜃2 is minimum subject to the
condition
𝜃1 ⪕ 𝑥(𝑖) ⪕. ⪕ 𝑥(𝑛) ⪕ 𝜃2
We have to take 𝜃2 = 𝑥(𝑛) and 𝜃1 = 𝑥(𝑖) so that 𝑚. ℓ. 𝑒 𝑜𝑓𝜃1and 𝜃2 are 𝜃1 = 𝑥(𝑖) and𝜃2 = 𝑥(𝑛) Example:
𝜃 − 𝑐 ⪁ 𝑥(𝑖) ⪁ ⋯ ⪕ 𝑥(𝑛) ⪕ 𝜃 + 𝑐
This shows that any statistics which lies in between 𝑥(𝑛) − 𝑐 and the
Example 12 It x has 𝑅(𝜃, 𝜃 + 𝐼),any statistics which lies between 𝑥(𝑛) − 1 and 𝑥(𝑖) is a 𝑚. 𝑙. 𝑒 if 𝜃
Example 13 Let(𝑥1, . . . 𝑥𝑛) be 𝑎, 𝑟, 𝑠 from the regular distribution 𝑅(𝜃, 2𝜃) having þ, 𝑑, 𝑓
Then
𝑖. 𝑒 𝜃 ⪁ 𝑥(1) … … (𝑖)
Since
Then
Then
And
Which is maximized when is the sample median.
of is
And
34
Or
Or
Of is
We have
(c)The sequence of estimators has the smallest asymptotic variance among all consistent,
(b)
(b)
43
44
45
W1 W2 W3
Let
0.01024 1.00000
0.23040 0.98976
0.34560 0.68256
0.25420 0.33696
0.07776 0.07776
47
48
49
: MP test is given by W=
50
51
52
H0 : f
Let us take a single observation. The MP test of H0 Vs H1 has the critical region
C}
²
Or C’
H
H1: f1(x) = 1 ; 0<x<1
53
Let us take a single observation. The MP test of H0 VS H1has the critical region given by
Where
We see that C
Hence MP or region is
f ϴ
We want to test
H0: ϴ = ϴ0 Vs
H1:ϴ = ϴ1(>ϴ0)
We have
L(ϴ) =
Now,
𝐿(ϴ0)
This shows that 𝐿(ϴ1) is an increasing function of 𝑥(𝑛) and, therefore
{ 𝑥(𝑛) ≥k}
P { 𝑥(𝑛) ≥k/𝛳0} = α
𝑛 1
We have
Remark: the above test is UMP for H0: ϴ=ϴ0 against H1:ϴ>ϴ0
As we have remarked, UMP test may not always exist. Therefore we for their restrict the class of
tests by considering unbiased tests (defined below) and then try to obtain UMP test in the class of
unbiased tests. If such a test exists we call it uniformly not powerful unbiased test (UMPU test)
54
Definition Suppose we are testing a sample hypothesis Hϴ: 𝜃 = 𝜃0 against a conqurite alternative
Remark: Suppose 𝜃 = 𝜃1 is one of the alternative value of 𝜃. If the test is not unbiased it may happen
that 𝑃𝑜(𝑇) <∝= 𝑃00(𝑇) which means that the probability of rejecting 𝐻𝑜 when it is false is less then
the probability if rejecting 𝐻𝑜 when it is true if the test is unbiased it will not happen. Theorem
A MP test or UMP test is unbiased.
Prof Let T be a MP (or UMP) test of size ∝. Consider another test T which rejects the null hypothesis
HO: 𝜃 = 𝜃0 with probability ∝ irrespective of the sample outcome. We may just toss a coin for which
the probability of is ∝ and decide to reject the null hypothesis Hϴ if we get ∝ , irrespective if the
sample values obtained. Then
𝑃𝑇{𝑅𝑒𝑗𝑒𝑐𝑡𝐻𝑜/HO is 𝑡𝑟𝑢𝑒} =∝
So that the size of the test T=∝. Also the power of test T is also∝, since
𝑃𝑇{𝑅𝑒𝑗𝑒𝑐𝑡𝐻𝑜/HO is 𝑓𝑎𝑙𝑠𝑒 } =∝
Or 𝑃𝑇 (𝜃) ⩾∝ for 𝜃 ≠ 𝜃0
Remark: It may be shown that the following tests are UMPU for two sided alternative 𝐻𝑖 ∶ 𝜃 ≠ 𝜃0 in
example 1,2 and 3
Now we consider a produce for constructing tests that has some intuitive appeal and that .
Frequently, though not always, leads to UMP or UMPU test. Also the produce leads to test that have
decided large sample properties
Suppose we are given a sample (𝑥1, … , 𝑥𝑛) from a distribution with þ, 𝑑, 𝑓 𝑓(𝑥, 𝜃 ) (where 𝜃 may be
a vector) and we deice to test the null hypothesis 𝐻𝑜 ∶ 𝜃 𝜖 𝑤(⊂ 𝛺) against the alternative hypothesis
𝐻𝑖 ∶ 𝜃 𝜖 𝑤(⊂ 𝛺) where 𝛺 is the parameter space,
max 𝐿(𝜃)
Where denotes the maximum of the likelihood function when 𝜃 is restricted to values in 𝜃 𝜔
w and max 𝐿(𝜃) denotes the maximum of the likelihood for when 𝜃 takes all possible values in𝛺
Obviously, 0 ≤ 𝜆 ≤ 1 and λ is also to 1 of the sample shows that 𝜃 lies actually in 𝛚. Definition
𝑤 = {𝜆 ⪕ 𝜆𝑜}
Remark (1) For testing a simple hypothesis against a simple alternative likelihood ratio test is
equivalent to the test given by the Neyman –Pearson lemma.
(ii) if a sufficient statistics exists the L.R test is a function of the sufficient statistics.
Example: (1) Let X be a r.v. having a normal distribution 𝑁(𝜇, 𝜎) where 𝜎 (=𝜎𝑜) is known
Against 𝐻1: 𝜇 ≠ 𝜇𝑜
Then
− ∑𝑛(𝑥𝑖−𝜇0)2/2𝜎20
Or
𝑒2𝜎20[∑(𝑥𝑖−𝑥̅) −∑(𝑥𝑖−𝜇0)² ≤ 𝜆0
Or
56
or
since there exists other UMP tests for 𝐻1:𝜇 > 𝜇0and
𝐻𝑂:𝜇 = 𝜇𝑜
Against 𝐻𝑖: 𝜇 ≠ 𝜇0
Therefore, we have
And
Or
2
Or
Or
Or ’’
Where ’
It is know that has t distribution on under There fore the values of can be
found from the size condition
Where Y~
Against
Under , the of is
In general, of is and of is
Then we have
And
Or
Or
But it is know that has distribution on (n-i) using the tables and size
condition we can get the values of and
58
(3a) suppose in example 3 the value of is know. Then the L.R cr region because
Where /n
We want to test
Against
Then we get
Also
Because of is
Where
Remark (i) if one take we shall get the L.R critical region as in both case of one –
sided alternation the L.R test are UMP test.
(2) Since has gamma distribution we can find the value of by using size condition
We want to test
Where it is assumed that 𝜎1 = 𝜎2(= 𝜎𝑢𝑛𝑘𝑜𝑤𝑛 ) we that the like hood function
59
And
Also and
Therefore
And
Therefore
Or
Or Or
2 2
Where
The cr region can be within as
Where 𝛾~𝑡𝑛1+𝑛2−2
(6)Let (𝑋𝐼,. 𝑋𝑛𝐼) be 𝑎, 𝑟, 𝑠 from N (𝜇, 𝜎𝑖) and(𝛾1, . 𝛾𝑛2 )𝑛 N (𝜇2, 𝜎2) where two samples (and two
distributions) are indecent
We want to test
Against
So that
𝑛 +𝑛2
1 − 12
So that max 𝐿(𝜇1 , 𝜇2 , 𝜎1 , 𝜎2 ) = 𝑛 1 +𝑛 2 𝑒
𝑛 +𝑛 2
(2𝜋) 1 2(𝑠 ) 2
𝑛1 𝑛2
Or
Or
Where
Setting 𝑔(𝑓)𝑓𝑜𝑟the 𝐿. 𝐻. 𝑆 of (i) we have 𝓰(o)=o and 𝓰(f)→o∞. Furthermore 𝓰(f) attains its
maximum for , it is impressing between o and f may and derision in (f may, ∞).
Therefore 𝓰(f)⪕ 𝜆𝑜 if and only if 𝑓 < 𝓀1 or 𝑓 > the LR cr region can be within as {𝐹 < 𝓀1 𝑜𝑟 𝐹 > 𝓀2}
Where
Hence 𝓀1, 𝓀2 can be obtained from the size condition 𝑃{𝑓 > 𝓀1 𝑜𝑟 𝐹 < 𝓀2} =∝whese F~𝐹𝑛1−1,𝑛2−1
61
;x≥0
=0 ;𝑥<0
(∝> 0, 𝛽 > 0) We
have , 𝑡<𝛽
𝐸(𝑋) =∝/𝛽
𝑉(𝑋) =∝/𝛽2
𝐸(𝑋) = 1/𝛽
𝑉(𝑋) = 1/𝛽2
=0 , x<0
The upper percent point of the F(m ,n, distribution is Fm ,n, where
Note that
Use of and distribution in testing problem
Use of distribution (i) Testing the variance of a of a distribution: Given a sample of
size n from a normal distribution where is unknown, we would like to test
against alternative > or < or the tests are summarised n follows know
HO:
Ho:
Ho:
know
HO: ,
Ho:
Ho:
Where
and the expected frequencies under the 𝐻𝑜 be 𝑒1,𝑒2, … … . 𝑒𝓀 (∑𝑛𝑖 𝑒𝑖 = 𝑛) where 𝑒𝑖 = 𝑛þ𝑖
calculate
Them, for large sample, 𝑥2has 𝑥2(𝓀 − 𝑖) the test of 𝐻𝑜 has the cr. region
(3) Testing goodness of fit: given a sample (𝑥1, . . 𝑥𝑛) of Observation on 𝑎. 𝑟. 𝑣 X arranged in the
form of a frequencies distribution having 𝓀 classes AI, … . . A𝓀 we would like to test the
hypothesis that distribution of X has a specified from with þ, 𝑑, 𝑓(𝑜𝑟 þ, 𝑚, 𝑓)𝑓𝑜(𝑥, 𝜃) the
parameter 𝜃 be a simple one or a vector (𝑜𝑖,… . . 𝜃ℯ)
Then, for large sample, 𝑥2has 𝑥2(𝓀 − 𝑖) the test of 𝐻𝑜 has the cr. Region
Note if 𝑟(𝑜𝑓 ℓ) parameters in 𝜃 are estimated from the sample then χ2 has χ2(𝓀 − 𝑟 − 𝑖)if any
expected frequency is lass then 5 we pool this class with the adjoining class and denote by𝓀 the
effective number ƪ classes after paroling
And 𝑒𝑖𝑗 = expected a=(𝑖𝑡ℎ 𝑟𝑜𝑤 𝑡𝑜𝑡𝑎𝑙 𝑥𝑗𝑡ℎ𝑐𝑜𝑙𝑢𝑚 𝑡𝑜𝑡𝑎𝑙)n “ “ “under 𝐻𝑜 Calculate
Where n=total frequency. Then has on the test of has the cr. Region
Us : all correlation coefficients are not equal we use the friskers z-trans function of correlation
And
Uses if t-distribution:
be two
samples from in dept normal populations and respectively let be as
usual and let
We would like to test 𝐻𝑂: 𝜇1 = 𝜇2 against alternative 𝜇1 < 𝜇2 or 𝜇1 ≠ 𝜇2 the test are summarised as
follows:
Case I
𝐻𝑖: 𝜇1 > 𝜇2
𝐻𝑖: 𝜇1 ≠ 𝜇2
65
Uses of F-distribution:
Let two samples of sizes 𝑛1 and 𝑛2 be given from two independent normal population 𝑁(𝜇1, 𝜎1) and
𝑁(𝜇2, 𝜎2), respectively .Let be the two sample variance. We would like to test the null
hypothesis 𝐻𝑜:𝜎1 = 𝜎2 against𝐻𝑖: 𝜎1 ≠ 𝜎2 The test are cr follows:
Reject 𝐻𝑜 if either
Or
Reject 𝐻𝑜 if either
Or
(2) Testing the multiple correlation coefficient: Given a sample of size or from a bivariate
normal population (𝑥1, 𝑥2, 𝑥3) with multiple correlation coefficient 𝑅1(23) of 𝑥1𝑜𝑟(𝑥2, 𝑥3) we
would like to test the null hypotheses 𝐻𝑂𝑅1(23) = 0 let the sample multiple correlation coefficient
be 𝑅1(23). The test is to reject 𝐻𝑂 at level ∝ if
(3) Testing the equality of means of 𝓀 normal distribution (𝓴 > 𝟐)[see left page]
Where r is a sample correlation coefficient Though the population correlation coefficient P may be
widely different from zero, the new statistics z may be amounted to be normally distributed even
when n is as small as 10 it has hen show that z has approximate mean
(ii)For testing 𝐻𝑜 ∶ 𝑝1 = 𝑝2 against 𝐻𝑖 ∶ 𝑝1 ≠ 𝑝2 involving two populations, let 𝑟1, 𝑟2 be the sample
correlation coefficient for two independent sample of size 𝑛1, 𝑛2 from the two populations and let
𝑧1,𝑧2be there transformed values,𝑖, 𝑒
(iii)Let 𝑟1, 𝑟2 … . . 𝑟𝓀 be sample correlation coefficient for 𝓀 sample of sizes 𝑛1, 𝑛2 … 𝑛𝓀 drown from
𝓀 independent vicariate normal population with correlation coefficients 𝑝1𝑝2 … . . 𝑝𝓀. Let 𝑧1, 𝑧𝓀 be
the transformed values and let
Large sample tests so for we have considered tests of hypothesis which contain assumptions
regarding the population are satisfied .Now we consider some approximate test which are valid only
for sufficiently large samples, but they have wide applicability and hold for all populations satisfying
certain general conditions rather than being valid for some particular populations only (e.g. normal )
67
(ii)Testing the equality of two population proportions: Let 𝑝1, 𝑝2 be two population proportions and þ1,
þ2 be the two sample proportions dream from there indecent population the test of 𝐻𝑜: , 𝑃1,𝑃2 is to
reject 𝐻𝑜at level ∝if
Where
(iii)Testing for a st. deviation: let s be the st. Deviation of a sample of observation of size drown from a
population with st. Deviation 𝜎(x) the test of 𝐻𝑂: 𝜎 = 𝜎𝑜 is to reject 𝐻𝑂 at level ∝if
(iv)Testing for equality of two population st. Deviation Let 𝓈1, 𝓈2 be the st. Deviation of two sample of
sprees 𝑛1, 𝑛2 from two independent population with st. Deviation 𝜎1, 𝜎2 Let
Definition:- For a random sample (𝑥1, … , 𝑥𝑛) from the distribution of a 𝑟. 𝑣. 𝑥 having þ, 𝑑, 𝑓 𝑓(𝑥, 𝜃) Let
𝐿1𝐿1(𝑥1, … , 𝑥𝑛)and 𝐿2(𝑥1, … , 𝑥𝑛) be two statistics such that 𝐿1 ≤ 𝐿2. The interval [𝐿1, 𝐿2] is a confidence
interval for 𝜃 with. Confidence coefficient 1−∝ (0 <∝< 1) if 𝑃𝜃[𝐿1 ≤ 𝜃 ≤ 𝐿2] = 1−∝ for all 𝜃 𝜖 𝛺 𝐿1 and
𝐿2 are called the lower and upper confidence limits, respectively at least one of them should not be a
constant.
Interval Estimation
Estimation of a parameter by a sample value is known as point estimation. An alternation produce is to
give an interval within which the parameter may be supposed to lie with high probability. This is called
interval estimation and the interval is called the confidence for the parameter
Suppose 𝑎, 𝑟, 𝑣 x has Normal distribution 𝑁(𝜇, 𝜎) with unknown mean 𝜇 and known st. Deviation𝜎. Let
(𝑥𝑖, … , 𝑥𝑛) be the values of a random sample of size or from then distribution .We know that the sample
Or, equivalently,
68
This shows that, in respected sampling the probability is 0.95 that the interval
Will include 𝜇, We say that above is a confidence interval for 𝜇 with confidence coefficient ,95. The two
end points are known as 95% confidence limits for𝜇.
Let us now consider the general problem Let 𝑎, 𝑟, 𝑣 x has distribution depending on an unknown
parameter 𝜃 which is to be estimated. Suppose Z is a statistics (usually it is a function of a sufficient
statistics if it exists) which is a function of 𝜃 but whose distribution does not depend on𝜃. Such a
statistics z is called a ploetal function Let 𝜆1 and 𝜆2 be two numbers such that
The above inequality can be solved such that it assumes the from
For all 𝜃 where 𝜃1and 𝜃2are random variables which do not depend on𝜃.Finally, if we astute the sample
value [𝜃1((𝑥1, … , 𝑥𝑛)), 𝜃2((𝑥1, … , 𝑥𝑛))] becomes a confidence interval for 𝜃 with desired confidence
coefficient 1−∝.
Remark: the numbers 𝜆1, 𝜆2 may be chosen in several ways, giving rise to several confidence intervals.
We usually choose confidence intervals of shortest length.
Which has 𝑁(𝑂, 𝐼) distribution For a specified∝ let 𝑁∝/2 be the % critical value of 𝑁(𝑜, 1)then
Or
So that
where
Or
So that
69
Or
Where
Or
Or
Or
confidence interval of
Let Z
Then
Or
Or
So that
One many chose a confidence region for using the two relations
Where
But it is difficult to find the probability of the sample to full in the shaded region (confidence region)
Alternatively, using the independence of and we chose the cofidence region by the help of relation
71
A,d Since
are indept
Obtain the boundaris of the confidence .region without difficully this is shown by the shaded region
below
Where
Or
Or
So that
(II) For two sample we can similerly find a confidence interval for as follows:
Where
So that
(iii) Let x be having mean , variance and we want a confidence interval for For that
approximately .
Or
Then
Is a confidence for
(iv) For two sample we an similerly find a cofidence interval for as follows:
Where
So that
(v) Let have a bivanate normal distribution with coefficient P and me want to find a confidence
region for P.
and
size n
Then So
that
Or
So that
Gives a confidence interval for ξ.From this we can earily obtain the corrponding confidence
interval for P.
73
NON-PARAMETRIC INFERENCE
In all problems of statictics inference considered so fan we assumed that the distribution of the random
variable breing sampled is know n except for some parameters . in pratice however the functional
from in the distribution is seldom if ever , known if is therefore desivable to devise some produres
that are free from this assumption concering distribution such produres are commonly refered to as
distribution free or non-parametric methods the term distribution free refers to the fact that no
assumptions are made about the underlying distribution execpt that the distribution function being
sampled is absolutely continuous or purely discrete. The term non-parametric refers to the factors that
there are no parameters involved in the traditional sense of the parameter used so for.
We will consider only the inferential problem of testing of hypothesis and dercribe a few
non-parametrictests
Single- sample problems : (a)The problem of fit : the problem of fit is to test the hypothesis that a
sample of obsevations (𝑥𝑖, 𝑥𝑛) is from some specified distribution against the alternative that it is from
some other [Link] we have to test
(i)Chi- square test: Let there be 𝓀 categories and let þ𝑖 be the probality of a random obsevation from
𝐹𝑜(𝑥) to fall in the 𝑖𝑡ℎ category (𝑖 = 1,2, … . 𝑛).For a sample of size n, Let 𝑜𝑖 be the obsevarved freqnecy
in the 𝑖𝑡ℎ category and let 𝑒𝑖 = 𝑛þ𝑖 be the expected frequency in the 𝑖𝑡ℎ category under 𝐻𝑜. To test 𝐻𝑜
we use the chi-square statics
The larger the value of 𝑥2 the more likely it is that the 𝑜𝑖,𝑠 did not come from 𝐹𝑜(𝑥). The 𝑥2 −statistic
for large samples has a 𝑥2 distribution on (𝓀 − 1)d.f .Thus an approximate level ∝ test is provided by
rejecting 𝐻𝑜 if
𝑥2 > 𝑥𝓀2−1∝,
(ii)Kolmogoror – Smironv one sample test : For the sample (𝑥𝑖,… 𝑥𝑛)let the empirical distribution
function 𝐹𝑛(x) be given by
𝑜 𝑥(𝑖)
𝑖𝑓 𝑥 <
𝐹𝑛(𝑥) {𝓀⁄𝑛 𝑖𝑓 𝑥(𝓀) ⪕ 𝑥 < 𝑖𝑓 𝑥(𝓀−𝑖)
𝑖 𝑥⩾ 𝑥(𝑛)
(𝓀 = 1,2, … 𝑛, −1) whese 𝑥(1),𝑥(2), … . 𝑥(𝑛) are the order statistic , Evidently ,
For testing 𝐻𝑂: 𝐹(𝑥) = 𝐹𝑜(𝑥) against the two sided alternative 𝐻𝑖: 𝐹(𝑥) ≠ 𝐹𝑜(𝑥) we use the Kolmogoror –
Smironv statictic
It can be shown that the K-S statistic 𝐷𝑛 is completely distribution free for any continouns distribution
𝐹𝑜(𝑥)
𝐷𝑛 > 𝐷𝑛,∝
Remark1:For testing𝐻𝑂: 𝐹(𝑥) = 𝐹𝑜(𝑥) against one-sided alternatives 𝐻1: 𝐹(𝑥) > 𝐹𝑜(𝑥) or 𝐻2:𝐹(𝑥) <
𝐹𝑜(𝑥) based on one-sided K.S statistics and 𝐷𝑛− are also available
Remark 2: For small sample 𝑥2 −test is not available but K.S test can be applied. For discrete distibution
[Link] is not availible but 𝑥2 −test can be appled K.S test is more powerful then 𝑥2 −test.
(B) The problem of Location: Let (𝑥𝑖, … . 𝑥𝑛) be a radom sample from a distribution 𝐹(𝑥) with unknown
median ξ ,where 𝐹(𝑥) is assumed to be continus in the neigbourhood of ξ. By definition of median
.We would like to test the hypothesis (x) If n>25, normal appronimution may be used
We take
Sign Test: We from the n differences (𝑥𝑖 − 𝜉𝑂)𝐼 = 1,2 … … … 𝑛 and find out the number, R,of position
differences (differences having postive signs ) 𝑖, 𝑒 when (𝑥𝑖 − 𝜉𝑂) > 𝑜.
1
If 𝐻𝑂 is true, and R has a Biomial distribution with paramer2 .
We
may use an exect test of𝐻𝑂 based on the Biomial Distribution. In the case of one-sided alternative
𝐻𝑖: 𝜉 > 𝜉𝑜
The sample will have an excess of positive signs and in the case of
𝐻𝑖: 𝜉 > 𝜉𝑜
The critical values 𝑅1∝, 𝑅2∝, 𝑅∝/2, 𝑅∝/2 are calculate from tables of Biomaial distribution
Rajred –sample signs test: Here we assume that we have a random sample of n pains (𝑥𝑛, 𝑥𝑛) giving the
the differences
𝐷𝑖 = 𝑥𝑖 − 𝑦𝑖 ,𝑖 = 1, … 𝑛
We have , now a single sample 𝐷𝐼, … . . 𝐷𝑛 and we can test 𝐻𝑜: ξ = 𝜉𝑜 which can be taken to be oby the
sign test descrited above.
Remark the above two sign test s are , repectively aralogoun to single sample 𝑡 − 𝑡𝑒𝑠𝑡 and paired ttest
for testing location of a normal distribution ,
Two sample problems : let (𝑥𝑖,… … 𝑥𝑛) and (𝛾𝑖, … … 𝛾𝑛) be independent random sample s from two
absolutely continous distribution 𝐹𝑥(𝑥) and 𝐹𝛾(𝓎) , respectively
Run test(Wald –Wolfowitz): we assarge the m, x’s and n 𝛾′𝑠 in increasing order of size
and count the numbers of runs .if 𝐻𝑜 is true the (m+n) values will be well mixed up and
we expect that R, the total number of runs , will be relatively large. But R will be small if the samples come
from differernt popaltions 𝑖, 𝑒 𝐻𝑜 is false in the extreme case , if all the value of y are greater than all the
value of x, or vice – vera , there will be only two runs
𝑅 ⪕ 𝑅∝
𝑃(𝑅 ⪕ 𝑅∝/𝐻𝑜) ⪕∝
And
Tables of critical values of R based on above have been given by swed and Eisenhant For
large m,n(both greater then 10), Ris asymptohcally Normally distributed with
And
Median it test: We arrange the x’s and y’s in asscending order of size and find the median M of the
contied sample let
If V is large it is reasomable to conclude that the actual median of x is smaller than the median of Y
𝑖, 𝑒 𝐻𝑜: 𝐹𝑥(𝑥) = 𝐹𝑌(𝑥) is respected
On the other hand , if V is too small it is reamable to condude that the actual median of X is greater
than the median of y 𝑖. 𝑒 𝐻𝑜: 𝐹𝑥(𝑥) = 𝐹𝑦(𝑥)is respected in fovoues of 𝐻𝑖:𝐹𝑥(𝑥) < 𝐹𝑌(𝑥)
For the two sided alternative , we use the two sided test .
𝑚 𝑛
And
76
𝑚 𝑛
Wilcoxon- Mann –Whitney U test: This is the most widely used two- sample non-parametric test and is
a useful alternative to the t-test assumotions.
The test is like the run test based on the pattern of 𝑚, 𝑥′𝑠 and 𝑛, 𝑦′𝑠 arranged in ascending order of size .
The Main- Whitney U statistic is defined as the number of times as X preades 𝑎 𝑌 In the combined sample
of size 𝑚 + 𝑛. We define
And write
Note that is the number of 𝑦𝑗′𝑠 that are larger than 𝑥𝑖 and hence U is the number of values of
𝑥𝑖, … … … 𝑥𝑛 that are smaller than each of 𝑦𝑖,… … … , 𝑦𝑛. For example , suppose the contined sample
when ordered is as follows :
Then U=7, becouse there are three values of X<𝑌1, two values of X<𝑌2 and two values of X<𝑌3
It is obseved that U=0 if all the are larger than all and U=mn of all the are smaller than all
the 𝑦𝑖′𝑠. Thus 𝑜 ⪕ 𝑈 ⪕ 𝑚𝑛. If U is large the values of y tend to be larger than X (Y is stochastically larger
than X) and this supposts the alternative 𝐹𝑥(𝑥) > 𝐹𝛾(𝑥). Similarly, if U is small, the values of Y tend to be
smaller than X and this supposts the alternative 𝐹𝑥(𝑥) > 𝐹𝛾(𝑥).
𝐻𝑜 𝐻𝑖 𝑅𝑒𝑗𝑒𝑐𝑡𝐻𝑜 𝑖𝑓
And
The tables of distribution of U for small samples are given by table and Mann-Whitney. For large samples
U has asymptotic normal distribution,𝑖, 𝑒
APPENDIX
Therom: suppose Xis a continuous 𝑟, 𝑢 with þ, 𝑑, 𝑓 𝑓𝑥(𝑥). Set 𝑥 = {𝑥,𝑓𝑥(𝑥) > 𝑜}.Let
(ii)the derivative of 𝑥 = ℊ−1(𝑥) 𝜔. 𝑟. 𝑡 𝓎 is continous and non-zero for 𝓎 𝜖 𝑥, where ℊ−1(𝓎) is the
inverse for of 𝓎(𝑥) 𝑖, 𝑒 ℊ−1(𝓎) isthat 𝑥 for which ℊ(𝑥) = 𝓎 Then 𝛾 = ℊ(𝑥) is a cont. 𝑟, 𝑢 with þ, 𝑑,
𝑓.
(i)𝓎1, = ℊ1(𝑥1, 𝑥2) and 𝓎2, = ℊ2(𝑥1, 𝑥2) defines i:i transformation of x onto x.
(iii) The jacebian of transformation is non-zero for (𝓎1𝓎1)𝜖𝑥.Then the joint þ, 𝑑, 𝑓 of 𝛾1 = ℊ, (𝑥1,
𝑥2)and 𝛾2 = ℊ, (𝑥1, 𝑥2)is given by
Where
X2- distribution
Definition : A continous 𝑟, 𝑢, 𝑥 is said to have the X2- distribution on n degrees of freedom if its þ, 𝑑, 𝑓 is
given by
=𝑜 𝑥<𝑜
The 𝑚, ℊ, 𝑓 of x is given by
For 𝑛 ⪕ 2the þ, 𝑑, 𝑓 of 𝑥2(𝑛) steadily dencress as 𝑥 iscrese while for 𝑛 > 2 there is a uniqne maximum
at 𝑥 = 𝑛 − 2
Theorom : Let 𝑥1, 𝑥2 … … . . 𝑥𝑛 be n independent standand normal r,v,s 𝑖. 𝑒 𝑥𝑖~𝑁(𝑜, 1), 𝑖 = 1, … 𝑛 Then
has a X2- distribution on 𝑛, 𝑑, 𝑓.
78
Therom : Let 𝛾1, 𝛾2 … . 𝛾𝑛 be indepent 𝑟, 𝑢, 𝑠 with X2- distribution on 𝑛𝑖, … … . 𝑛𝓀 degrees of freedom
resp .
Then
Proof :the 𝑚, ℊ, 𝑓 Z
𝑀𝑍(1) = 𝐸𝑒𝑡𝑧
= (1 − 2𝑡)−(𝑛𝑖+..+𝑛𝓀)/2
Crollanj : Let (𝑥𝑖, … . . 𝑥𝑛)be a random simple from a Normal distributuion 𝑁(𝜇, 𝜎).Then has
𝑥2 distribution on 𝑛, 𝑑, 𝑓.
And be the sample mean and sample variance. Then has 𝑥2 distribution on
(𝑛 − 𝑖)𝑑, 𝑓.
Therom: For large 𝑛, √2𝑥2 can be shown to be approximately normally distributred with mean
and st-dearation unity.
Therom: Assume that y has distribution function 𝐹𝑌 which satifies some regularity conditions ad which
has r-unknown parameters 𝜃1, 𝜃2 … . 𝜃𝑟 and that (𝑦𝑖,. . 𝑦𝑛) is a random sample of [Link] 𝜃𝑖, 𝜃𝑟 be the 𝑚.
ℓ, 𝑒 of 𝜃′𝑠 .Suppose the sample is distribution in 𝓀 non-orerlapping intervals {𝐼𝐽} where
𝐼𝐽 = {𝓎: 𝑎𝑗−𝑖 < 𝑦 < 𝑎𝑗−𝑖},𝑗 = 1, … 𝓀(𝑎𝑜 = −∞𝑎𝓀 = ∞and . Let 𝑥𝑖,… . . 𝑥𝓀 be the number of sample values
falling in these inervals, respectively if me define
79
𝑝𝑗 = 𝑃{𝑌𝑓𝑎𝑙𝑙𝑠 𝑖𝑛𝐼𝐽}, 𝑗 = 1, … 𝓀
Where 𝜃𝑖, 𝜃𝓀 replace 𝜃𝑖, 𝜃𝓀 in 𝐹𝑦 ,then the distribution of the statistics Lerger is
appoximately distributed as 𝑥2on 𝓀 − 𝑟 − 𝑖 𝑑, 𝑓 as n gets
Students t-distribution
Which shows that it is a couchy distribution We will therefore, assume that 𝑛 > 𝑖
Remark:the þ, 𝑑, 𝑓 of t-distribution is symmctric about again. For large n the t-distribution tends to
Normal distribution. For small n hawever t-distribution deviates considerally from the normal in fact if
𝑇~𝑡(𝑛)and 𝑧~𝑁(𝑜, 𝑖)
2r<n
Therom : Let 𝑥~𝑁(𝑜, 1) and 𝑦~𝑥2(𝑛) and Let 𝑥and 𝑦 be independent .Then
80