ECONOMETRICS
Semester V Paper VIII 4 Credits 100 marks
Module I STATISTICAL INFERENCE
Concepts and Steps in Hypothesis Testing (Population , Sample,
Population Parameter, Sample Statistic, Null & Alternative Hypothesis,
Test of Significance, Critical Region, One- Tailed & Two-Tailed Tests,
Type I & Type II Errors)
Basic Statistical Methods for Hypothesis Testing - Standard Normal z
Distribution, t – distribution, Chi – square χ2 , F distribution
Population & Sample
Population refers to any group of people or objects
that form the subject of study in a particular survey and
are similar in one or more ways.
Target population is the collection of objects that
possess the information sought by the researcher and
about which inferences are to be made.
A Sample is a smaller part of a larger group. It is a
subset of the population, containing the characteristics of a
larger population.
• Samples are used in statistical testing when population
sizes is too large for the test to include all possible
members or observations.
Statistical Inference
Inductive - extension from particular to general
Researcher performs an experiment, obtains data and on the basis of
empirical evidence draws conclusions which go beyond the experiment but
are expected to be valid for the entire field of study.
Examining the sample to draw certain facts about the population based on
results found in the sample
11-4
Population Parameter & Sample Statistic
Population parameter refers to the numerical characteristics of the
population under investigation.
Quantities of population which has a distribution f(x) – MEAN ‘μ’
,VARIANCE ‘σ2’ , Median , Skewness …
Researcher interested only in some characteristics of population – draws
inference in probabilistic terms on the basis of sample statistics – 2 types of
judgements
• Estimation of parameter using an estimator, specific value of which is an
‘estimate’
• Testing some hypothesis about the population parameter based on apriori
assumptions – Accept / Reject hypothesis base d on evidence provided by
sample observations in the form of test statistics.
Population Parameter & Sample Statistic
Quantities of population which has a distribution f(x) – MEAN ‘μ’
,VARIANCE ‘σ2’ , Median , Skewness …
Sample statistic – Process of sampling ( selection of random samples from
the population) carried out for sample investigation with the objective of
obtaining information about the target population which may be finite or
infinite.
The characteristics of the sample from which conclusions are drawn are
called sample statistics.
Sample mean ‘ ͞xi’, standard deviation ‘s’ ….
HYPOTHESIS - STATISTICAL, NULL, ALTERNATIVE
HYPOTHESIS – logical predictive statements or guesses –
• statements about the probability distribution of the population
i) Maintained - qualitative statements – not for testing
ii) Statistical / Testable – tentative statement about the value of
population parameter – to test whether the statement is true or false –
verify based on sample observations drawn – to find how far the sample
differs from the hypothesized value of population parameter.
iii) Simple Hypothesis – population is specified completely
iv) Composite Hypothesis – population not specified completely
HYPOTHESIS - STATISTICAL, NULL, ALTERNATIVE
Statistical / Testable HYPOTHESIS - to test whether the statement is true
or false –– to find how far the sample differs from the hypothesized value of
population parameter.
TYPES for testing :
Null Hypothesis – Neutral statement / non-committal attitude of the
modeler – researcher makes a statement of indifference while testing the
hypothesis – Hypothesis of ‘NO DIFFERENCE’ - denoted as H0
Alternative Hypothesis – Specifies values that the modeler believes to be
true – Researcher’s Claim – Acceptance or rejection of H0 meaningful only
against a rival alternative hypothesis – denoted as HA or H1 .
One Tailed Test - HA or H1 Alternative Hypothesis shows positive or
negative bias on the part of the researcher.
Two Tailed Test - HA or H1 Alternative Hypothesis shows no bias on the
part of the researcher. Hypothesis of DIFFERENCE…
Alternative Hypothesis – Specifies values that the modeler believes to be true
– Researcher’s Claim– denoted as HA or H1 .
One Tailed Test - HA or H1 Alternative Hypothesis shows positive or
negative bias on the part of the researcher.
Two Tailed Test - HA or H1 Alternative Hypothesis shows no bias on the
part of the researcher. Hypothesis of DIFFERENCE…
LEVEL OF SIGNIFICANCE and CRITICAL REGION
LEVEL OF SIGNIFICANCE – l.o.s determines the confidence interval
within which a statistician accepts the null hypothesis
• The validity of a null hypothesis H0 against an alternative hypothesis H1 is
always tested at a certain level of significance.
•If confidence level is 95 % => statistician feels that he should accept
H0 such that he will not be taking a wrong decision in 95 out of 100
cases i.e. probability he’s not wrong is 0.95 OR probability he’s wrong is
0.05 – researcher confident he made the right decision in accepting true
HO
• Generally consider 95 % or 99% Level of confidence
• TEST PROCEDURE – tests of hypothesis – tests of significance –
decision rules to accept or reject the null hypothesis H0 – to determine
whether sample observations differ significantly from expected results.
Region of ACCEPTANCE of Null: the region of sample space in which if
the value of test statistic falls, we accept the null hypothesis H0 . The margin
that separates the region of acceptance of null from the region of rejection
depends on predetermined level of significance. (l.o.s) – The l.o.s is fixed
before drawing the sample so the result will not influence the decision – “Fail
to Reject Null”
CRITICAL REGION – Region of rejection of Null Hypothesis – the
region of sample space in which if the value of test statistic falls, we reject
the null hypothesis H0 .
TYPE I and TYPE II ERRORS
The decision to accept or reject the null hypothesis Ho is made on thebasis
of information given by sampleobservations. The conclusions drawn
however may not always be true for the population.
Following possibilities arise:
i. Null Hypothesis Ho is true, but the test procedure rejects it
ii. Null Hypothesis Ho is false, but the test procedure accepts it
iii. Null Hypothesis Ho is true and the test procedure accepts it
iv. Null Hypothesis Ho is false and the test procedure rejects it
Null ACCEPT HO REJECT HO
Hypothesis/DECISION
Ho is true NO Error / Correct TYPE I ERROR
decision
Ho is false TYPE II ERROR NO Error / Correct Decision
α and β
i. P(Rejecting Ho when Ho is true) = P (TYPE I ERROR) = α => l.o.s.
α => Size of the critical region
i. P(Accepting Ho when Ho is false) = P (TYPE II ERROR) = β
(1- β ) – Power of Test – The value of the power function of the test
hypothesis Ho against the alternative HA at a parameter point is called
the power of the test a that point.
• An ideal test keeps both errors α and β under control. But for a
sample of fixed size ‘n’ both are related and any attempt to reduce
one would increase the other – Both errors may be reduced by
increasing sample size but it may not always be possible
• TYPE I Error deemed more serious
p value
A p-value higher than 0.05 (> 0.05) is not statistically significant and indicates strong
evidence for the null hypothesis. This means we retain the null hypothesis and reject the
alternative hypothesis.
• One cannot accept the null hypothesis, we can only reject the null or fail to reject it.
A statistically significant result cannot prove that a research hypothesis is correct (as this
implies 100% certainty).
Instead, one may state that the results “provide support for”
or “give evidence for” our research hypothesis (as there is
still a slight probability that the results occurred by chance
and the null hypothesis was correct – e.g. less than 5%).
p value
The p value of the test is the probability of obtaining results as extreme as the
observed results of a statistical hypothesis test, assuming that the null
hypothesis is correct. If l.o.s α= 5 % or 0.05, then the decision rule is reject the
null hypothesis Ho if p-value < 0.05 - A p-value less than 0.05 (typically ≤
0.05) is statistically significant. It indicates strong evidence against the null
hypothesis, as there is less than a 5% probability the null is correct (and the
results are random). Therefore, we reject the null hypothesis, and accept the
alternative hypothesis.
POINT & INTERVAL ESTIMATION
• Estimation of population parameter using a sample estimator. A specific value of the
estimator is called an ‘estimate’
• Given a population of values of a continuous random variable X which follows Normal
probability distribution, estimating it’s parameter MEAN ‘μ’,then we use a sample
statistic μ̂ to obtain the estimate μo .The problem of point estimation is obtaining an
estimate that would represent the best guess about the value of the parameter.
• Interval Estimation is an estimation process where the parameters may be estimated
within a certain range with some given level of probability i.e. confidence level and the
interval is called confidence interval.
• Generally the confidence interval chosen is 95% or 99 %
Law of Large Numbers
If a population of variable X which is normally distributed with mean μ and
variance σ2 and if repeated samples of size n are drawn from this population
then the theoretical distribution of sample means ͞ xi will be normal with mean
μ and variance σ2 / n => As ‘n’ increases , s.d. of sampling distribution
decreases.
• Sample mean a reliable , unbiased estimate of population mean.
• When n becomes large , sample mean is expected to be closer to population
mean.
•If population distribution normal then sampling distribution is also normal…
Central Limit Theorem
If the size of the sample is large (n → ∞) the theoretical sampling distribution
of ͞ xi (i.e. sample mean) will be close to a normal curve regardless of the
shape of the distribution of the basic parent population.
• A good approximation generally obtained at n ≥ 30
NORMAL p.d.f. – A continuous random variable r.v. X is said to have a
normal distribution with parameters with mean μ and variance σ2 if it’s
probability distribution function is given by;
f(x) = -∞ < x ,μ < ∞ ; σ > 0
=0 o.w.
Properties of Normal p.d.f
Area properties of Normal p.d.f - Standard Normal Variate (S.N.V.)
Standard Normal Table
Confidence Interval & Confidence Limits – 95 % or 99%
TEST PROCEDURE - Design and Evaluation of Test
TESTS
1. LARGE SAMPLE
One Sample / Two Sample
1 Tailed / 2 tailed
α = 5 % / 1%
2) Students’ t – distribution – t test (small samples)
Dist. of sample statistic not approximated by Normal dist
1 sample (single mean) / 2 sample (Difference of means)
Chi-Square Distribution (TEST OF VARIANCE)
Definition: The Chi-Square Distribution, denoted as χ2 is related to the standard normal
distribution such as, if the independent normal variable, let’s say Z assumes the standard normal
distribution, then the square of this normal variable Z2 has the chi-square distribution
with ‘K’ degrees of freedom. Here, K is the sum of the independent squared normal variables.
The Sampling distribution of chi-square can be closely approximated by a continuous normal
curve as long as the sample size remains large.
PROPERTIES
[Link] chi-square distribution is a continuous probability distribution with the values ranging
from 0 to ∞ (infinity) in the positive direction. The χ2 can never assume negative values.
[Link] shape of the chi-square distribution depends on the number of degrees of freedom ‘ν’.
When ‘ν’ is small, the shape of the curve tends to be skewed to the right, and as the ‘ν’ gets
larger, the shape becomes more symmetrical and can be approximated by the normal
distribution.
3. χ2 distribution depends on the degrees of distribution as its shape
F-Distribution (DIFFERENCE OF VARIANCES)
Definition: The F-distribution depends on the degrees of freedom and is usually defined as the
ratio of variances of two populations normally distributed and therefore it is also called
as Variance Ratio Distribution.
Properties of F-Distribution
1) The F-distribution is positively skewed and with the increase in the degrees of freedom ν1 and
ν2, its skewness decreases.
2) The value of the F-distribution is always positive, or zero since the variances are the square of
the deviations and hence cannot assume negative values. Its value lies between 0 and ∞.
3) The shape of the F-distribution depends on its parameters ν1 and ν2 degrees of freedom.