DEFINITIONS
1. Parameter: A summary description of a fixed characteristics of measure of the target
population denoting the true value that would be obtained if a census rather than a sample were
undertaken.
- How many women prefer Pantene shampoo in Indore market?
- A census among the Indore women population would give clear data denoting true value.
- The money spent on an average by a woman in the Indore market is INR. 125, but it may
not be known.
2. Statistic: A summary description of a characteristic or measure of the sample; the sample
statistic is used as an estimate of the population parameter.
- Sample of the women population in Indore market would lead to sample statistic, which
can be used to estimate the total number of women in Indore preferring Pantene
shampoo.
3. Statistical inference: The process of generalizing the sample results to the population results.
- The process is called inferential statistics.
4. Confidence interval: The range into which the true population parameter will fall, assuming a
given level of confidence.
Confidence level Confidence interval
95% INR.120-130
99% INR.122-128
99% INR. 0-500 (It is not useful as such estimate does not give a good
estimate)
- Confidence level has to be higher and range of confidence interval has to be smaller.
5. Confidence level: The probability that a confidence interval will include the population
parameter.
6. Precision level: The desired size o the estimating interval; the maximum permissible
difference between the sample statistic and the population parameter.
- Practically useful confidence interval is the precision level, i.e. INR.122-128
7. Sampling distribution: The distribution of the values of a sample statistic computed for each
possible sample that could be drawn from the target population under a specified sampling plan.
8. Normal distribution: A basis for classical statistical inference that is bell shaped and
symmetrical in appearance. Its measures of central tendency are all identical.
9. Measure of central tendency: It is a single value that attempts to describe a set of data by
identifying the central position within that set of data.
- The mean, median and mode are all valid measures of central tendency, but under
different conditions, some become more appropriate to use than others.
10. Mean: The mean is equal to the sum of all the values in the data set divided by the number of
values in the data set.
So, if we have n values in a data set and they have values x1, x2, ⋯xn, the sample mean, usually
denoted by x¯ (pronounced "x bar"), is:
x¯= (x1+x2+⋯+xn) / n
This formula is usually written using the Greek capitol letter, ∑, pronounced "sigma", which
means "sum of...": x¯ = ∑x/n
The mean is essentially a model of your data set. It is the value that is most commonly used.
The mean is not often one of the actual values that you have observed in your data set.
It minimises error in the prediction of any one value in your data set. (Spending INR.0-1000, e.g.
mean – INR.350)
It is not to be used when the data contains unusually small or large numbers. (If women are
spending INR.10,000 per month on shampoo with gold particles).
11. Median: The median is the middle score for a set of data that has been arranged in order of
magnitude.
The median is less affected by outliers.
Arrange the data in ascending order and then find the middle number in case of total odd
numbers.
In case total even numbers find the average of the middle two numbers.
Advantages
It is very simple to understand and easy to calculate.
In some cases, it is obtained simply by inspection.
Median lies at the middle part of the series and hence it is not affected by the extreme
values. (e.g. some women spending exorbitant money on shampoo with gold particles)
It is a special average used in qualitative phenomena like intelligence or beauty which are
not quantified but ranks are given. Thus we can locate the person whose intelligence or
beauty is the average. (Ordinal level data)
In grouped frequency distribution it can be graphically located by drawing ogives.
It is specially useful in open-ended distributions since the position rather than the value of
item that matters in median.
Disadvantages
In simple series, the item values have to be arranged. If the series contains large number
of items, then the process becomes tedious.
It is a less representative average because it does not depend on all the items in the series.
(e.g. capabilities of MBA students in India… students of IIMs don’t get to influence the
results even though their capabilities are far higher than small institute MBA students.)
It is not capable of further algebraic treatment. For example, we can not find a combined
median of two or more groups if the median of different groups are given.
12. Mode: It is the most commonly observed value in a set of data.
- For the normal distribution, the mode is also the same value as the mean and median.
- In many cases, the modal value will differ from the average value in the data.
Advantages
It is easy to understand and simple to calculate.
It is not affected by extremely large or small values.
It can be located just by inspection in ungrouped data and discrete frequency distribution.
(SKUs of shampoo are e.g., of discrete frequency distribution; consumption ml per day is
going to be continuous frequency distribution.)
It can be useful for qualitative data. (e.g., Statements by different women to opt for
Pantene shampoo)
Disadvantages
It is not based on all the values.
It is not capable of further mathematical treatment.
Standard error: The standard deviation of the sampling distribution of the mean or proportion.
z value: The number of standard errors a point is away from the mean.