INFERENTIAL STATISTICS
One of the most important applications of statistics is making estimations about an
entire population based on the information from a small sample. This process is
known as statistical inference.
A population is a group of all distinct individuals or objects that you want to draw
conclusions about. The number of individuals/objects in a population is called
population size.
In statistics, we commonly use a sample that is a small subset of a larger set of data
for making inferences about the large set. Here the larger set is the population out of
which sample is drawn.
Sampling is a technique of selecting a small group (subset) of population for
estimating the characteristics, without having to investigate every individual.
1. Probability Sampling: Randomization (choosing something at random) is used in
this sampling method to ensure that every member of the population has an equal
chance of being included in the selected sample.
● Simple Random Sampling: a simple random sample is a subset of individuals
chosen from a larger set in which a subset of individuals are chosen
randomly, all with the same probability.
● Stratified Sampling: stratified sampling is a method of sampling from a
population which can be partitioned into subpopulations.
● Systematic Sampling: one-dimensional systematic sampling is a statistical
method involving the selection of elements from an ordered sampling frame.
● Every element is selected by following some rule.
● Cluster Sampling: cluster sampling is a sampling plan used when mutually
homogeneous yet internally heterogeneous groupings are evident in a
statistical population.
2. Non-Probability Sampling: Randomization is not used in non-probability sampling.
The result of this method can be biased, making it difficult for all the elements of
population to be included in the sample equally.
● Convenience Sampling:Convenience sampling is a type of non-probability
sampling that involves the sample being drawn from that part of the
population that is close to hand.
● Quota Sampling: Quota sampling is a method for selecting survey participants
that is a non-probabilistic version of stratified sampling.
● Purposive Sampling: Purposive sampling refers to intentionally selecting
participants based on their characteristics, knowledge, experiences, or some
other criteria. Resume selection is purposive sampling
● Snowball Sampling: snowball sampling is a nonprobability sampling technique
where existing study subjects recruit future subjects from among their
acquaintances. Thus the sample group is said to grow like a rolling snowball.
Representative Sample: A sample which accurately represents, reflects, or
matches with some of the features of your population.
Unrepresentative Sample: When the statistic does not represent the
population parameter, it is called an unrepresentative sample.
Unbiased Sampling: If every individual or the elements in the population has
an equal chance to be part of the selected sample, then the sampling process
is called unbiased.
Biased Sampling: If a sampling process systematically favours certain
outcomes over others, it is said to be biased. Convenience sampling type is
one of the biased sampling.
The difference between a population parameter and a sample statistic is
known as a sampling error.
Parameter and statistic both are related yet distinct measures. Parameter
refers to the whole population, while statistic refers to part of the population.
Statistical significance is a measure of reliability of findings which establishes
that when a finding is significant, it simply means we are confident that it is
real and the sample was framed wisely.
The sampling distribution of a statistic is the distribution of all possible values
taken by the statistic when all possible samples of a fixed size n are taken
from the population.
Central limit theorem (CLT) implies that the distribution of a sample leads to
become a normal distribution (bell curve shaped) as the sample size becomes
larger, considering that all the sizes of samples are identical, whatever be the
shape of the population distribution.
A sample size of 30 or more is considered to be sufficient to hold CLT and as
the sample size becomes large the prediction of characteristics of population
becomes more accurate.
We use Confidence Interval (CI) to express the precision and uncertainty of a
sampling process. A confidence level, a statistic and a margin of error are the
three components of it.
The margin of error describes the accuracy of a sampling method, while the
confidence level explains its uncertainty.
Standard error of the mean (SEM) measures how much dispersion there is
likely to be in a sample's mean compared to the population mean i.e it
measures the standard deviation of sampling distribution about the
mean.
The number of independent pieces of information on which an approximation
is based is known as the degrees of freedom.
The t distribution table values are critical values of the t distribution. The
column header are the t distribution probabilities (alpha) whereas the row
depicts the degrees of freedom (df). This can be used for both one-sided
(lower and upper) and two-sided tests using the appropriate value of .
“t” that we evaluated using formula is known as test statistics
|t| < = t alpha(given in the question), then Ho is accepted
Otherwise rejected
|t| > t alpha(given in question) then reject Ho
Otherwise Ho