Estimation Theory
Laura Sánchez, Angie Madrid, Cristina Martínez and Libardo Sánchez
University of Magdalena
Faculty of Engineering
Statistics II
April 2021
University of Magdalena - Faculty of Engineering 0
Estimation Theory:
Estimating is establishing conclusions about population characteristics based on results
show them, this term indicates that based on what is observed in a sample (a summary
statistical with the measures we know from Descriptive) this result is generalized
sampled from the total population, so the estimate is the generalized value to the
population. It consists of the search for the value of the population parameters in question.
study.
As it is an estimation, there is some error. Even if the estimator has all the
optimal properties. No matter how small it is, there will always be an error.
Thus, to obtain estimates adapted to that reality, intervals are created for
confidence. That is, ranges within which those estimated values lie with a certain degree of
trust. The degree of trust (reliability) can be modified. The greater it is
The greater the degree of confidence, the larger the interval will be. However, the less error there is...
initial estimate, the narrower the confidence interval will be.
To estimate what will happen regarding something (or what is happening, or what happened), despite
being a clearly statistical element, it is deeply rooted in our daily lives.
Within this, we also make estimates within a range of possibilities.
example: "I think I will finish the statistics assignment in about 5-6 days." What we do in
The field of data analysis is to apply technical nuances to this habit.
Estimation model:
The model commonly used consists of the following elements:
Parameter space: It is an unobservable space, whose elements are the
possible parameters that data generation depends on.
Observation space: It is the space whose elements are empirical data or
measures that will be used in the estimation.
- Probabilistic transition rule: Statistical distribution of the observations,
depending on the parameter or parameters.
University of Magdalena – Faculty of Engineering 1
- Estimator: Function of empirical data, which is used to generate the measure or
estimate.
Point estimate
A point estimate is when a single value taken from the sample is used to estimate.
the unknown parameter of the population (mean execution time of an algorithm,
average height of women in a population, difference in the average result between two
medical treatments, proportion of people who improve with a medical treatment... At
The value used is called an estimator, and these are the common estimators:
When we do inferential statistics, we use the sample values to try to
obtain the population values and this is particularly used in point estimation
Population parameter Sample parameter
μ : population average X́: sample mean
σ:population standard deviation S : sample standard deviation
p : population proportion ṕ : sample proportion
When we do inferential statistics, we use the sample values to try to
obtain the population values and this is particularly used in point estimation
We use the point estimator to find the estimator of the population parameter.
the values of mean, deviation, and proportion have an emphasis since we will not be able to
calculate exactly unless we could measure all the elements of the population,
something that in many cases is difficult; for this reason, in the case of the first one, the average is calculated
sample and that value I obtain is nothing more than the estimator of the population parameter
just like what happens with the second, but in the case of proportion I am not talking about number but about
percentage.
University of Magdalena - Faculty of Engineering 2
Population parameter estimator Point estimator
n
1
^ X́ = ∑ xi
ni=1
σ^ S=
√ ∑ ( x−I x́)2
n−1
x
^p ṕ=
n
There is a problem related to the use of point estimates, and that is that some
Estimates will be closer to the parameter being calculated than others. Without
embargo, we do not know how close our only point estimate of the parameter is
true. In a given situation, we can even consider extremely
it's unlikely that the point estimate is exactly the same as the parameter, but we are not
in a position to say how much we have erred.
Theory of interval estimation
In general, to try to resolve the issues of point estimates (such as the
mentioned earlier), we built an interval estimation of the parameter of
interest, in such a way that we can establish a degree of confidence in that interval
Include within its boundaries the parameter that is being estimated.
The confidence interval is determined by two values within which we assert
what is the true parameter with a certain probability. They are some limits or margin of
variability that we give to the estimated value, in order to affirm, under a criterion of
probability that the true value will not exceed them. It is an expression of the type [θ1, θ2] or θ1
≤ θ ≤ θ2, where θ is the parameter to be estimated. This interval contains the estimated parameter.
with a certain certainty or level of confidence.
In interval estimation, the following concepts are used:
University of Magdalena - Faculty of Engineering 3
Parameter variability: If it is not known, an approximation can be obtained in
the data or in a pilot study. There are also methods to calculate the size of the
shows that disregard this aspect. It is commonly used as a measure of this
variability the population standard deviation is denoted by σ.
Estimation error: It is a measure of its accuracy that corresponds to the
amplitude of the confidence interval. The more precision is desired in the estimation
of a parameter, the confidence interval should be narrower and therefore,
the smaller the error, and more subjects should be included in the studied sample. We will call
to this precision E, according to the formula E = θ2 - θ1.
Level of confidence: It is the probability that the true value of the parameter
estimated in the population is situated in the confidence interval obtained. The level of
confidence is denoted by (1-α), although it is usually expressed with
percentage (1-α)·100%. It is common to take a confidence level of 95% or a
99%, which correspond to α values of 0.05 and 0.01, respectively.
Alpha value: Also called the level of significance. It is the probability (in percentage)
one) of failing in our estimation, that is, the difference between certainty (1) and the
confidence level (1-α). For example, in an estimation with a confidence level
At 95%, the value α is (100-95) / 100 = 0.05.
Critical value: It is represented by Zα/2. It is the value of the abscissa at a certain
distribution that leaves to its right an area equal to α/2, where 1-α is the level of
confidence. Typically, critical values are tabulated or can be calculated in
function of the population distribution. For example, for a distribution
normal, with mean 0 and standard deviation 1, the critical value for α = 0.05 would be calculated
in the following way: that value is sought in the distribution table (or the most
approximately), under the "Area" column; it is observed that it corresponds to -0.64.
Then Zα/2 = 0.64. If the mean or standard deviation of the normal distribution does not
matching those in the table, the variable change t=(X-μ)/σ can be made for
its calculation.
Construction of Confidence Intervals for the mean:
University of Magdalena - Faculty of Engineering 4
A confidence interval is an estimation technique used ininference
statisticthat allows to limit one or several pairs of values, within which there
you will find thepoint estimatesearched (with a certain probability).
A confidence interval will allow us to calculate two values around a mean.
sample (one upper and one lower). These values will delimit arankwithin which,
With a certain probability, the population parameter will be located.
Confidence interval = mean ± margin of error
Example of confidence interval for the mean, assuming normality and
known the standard deviation
The pivot statistic used for the calculation would be the following:
The resulting interval would be the following:
We see how in the interval to the left and right of the inequality we have the bound.
lower and upper respectively. Therefore, the expression tells us that the probability of
that the population mean lies between those values is 1-alpha (confidence level).
Let's take a better look at the above with a solved exercise as an example.
University of Magdalena - Faculty of Engineering 5
It is desired to estimate the average time that a runner takes to complete a marathon.
Ten marathons have been timed, and an average of 4 hours has been obtained.
a standard deviation of 33 minutes (0.55 hours). It is desired to obtain a 95% interval of
confianza.
To obtain the interval, we would just need to substitute the data into the formula of the
interval.
The confidence interval would be the part of the distribution that is shaded in blue.
The 2 values bounded by this would correspond to the 2 red lines.
The central line that divides the distribution into 2 would be the true population value.
It is important to highlight that in this case, since the density function of the distribution
N(0,1) gives us the cumulative probability (from the left up to the critical value), we have
to find the value that leaves us to the left 0.975% (this is 1.96).
University of Magdalena - Faculty of Engineering 6
Construction of intervals for the proportion:
The confidence interval for the population proportion is centered on
sample proportion; with its upper and lower limits
where /2 is the critical value corresponding to the confidence level 1- of the
standard normal distribution and It is the typical proportion error.
Example of confidence interval for the population proportion:
A sample of 100 voters, chosen at random from all in a district, indicated
that 55% of them were in favor of a certain candidate. Find the interval.
99.73% confidence for the proportion of all voters in the district that
they were in favor of the candidate.
Magdalena University - Faculty of Engineering 7