Unit-II
Statistical Inference
Indira College of Engineering Management, Pune
Need of statistics in Data Science and Big
Data Analytics -
• In data science, statistics is at the core of sophisticated machine
learning algorithms, capturing and translating data patterns into
actionable evidence.
• Data scientists use statistics to gather, review, analyze, and draw
conclusions from data, as well as apply quantified mathematical
models to appropriate variables.
2
Importance of statistics in Data Science and
Big Data Analytics -
• Identify the importance of features by using various statistical tests.
• Finding the relationship between features to eliminate the possibility of
duplicate features.
• Converting the features into the required format.
• Normalizing and scaling the data. This step also involves the identification
of the distribution of data and the nature of data.
• Taking the data for further processing by using required adjustment in the
data.
• After processing the data identify the right mathematical approach/model.
• Once the results are obtained the results are verified on the different
accuracy measurement scales.
3
Measures of Central Tendency -
• A measure of central tendency (also referred to as measures of centre
or central location) is a summary measure that attempts to
describe a whole set of data with a single value that represents
the middle or centre of its distribution.
• There are some measures of central tendency: Mean, Mode, Median
and Mid Range.
4
Measures of Central Tendency -
1. Mean - The mean represents the average value of the dataset.
Examples – 1) 2,2,5,6,7,8
2) 10, 20,30,40,50,60
3) 2,2,5,6,7,8
5
Measures of Central Tendency -
2. Median – Median is the middle value of the dataset in which the
dataset is arranged in the ascending order or in descending order.
6
Measures of Central Tendency -
• Median –
• Examples –
• 23,21,18,16,15,13,12,10,9,7,6,5,2,1
• 2,5,7,2,6,8,9
• 6,7,4,7,7,6,4,6,5
• 4.5,4.1,2.0,4.6,4.2
• 75p,99p,89p,79p & 85p
7
Measures of Central Tendency -
3. Mode – The mode represents the frequently occurring value in the
dataset.
Examples –
5,4,2,3,2,1,5,4,5
1,3,3,3,5,6,6,9,9,9
54,67,32,54,72,98,32,33,21,32,67
82,75,78,82,71 & 82
8
Measures of Central Tendency -
4. Mid Range – The midrange is used to identify a measure of centre.
To calculate the midrange of a set of numbers:
Step1: Sort out your numbers(either into ascending or descending
order)
Step2: Find the maximum & minimum numbers.
Step3: Use the midrange formula, M=(max+min)/2
9
Measures of Central Tendency -
4. Mid Range –
Examples –
100,30,17,620,77,900,12,470,4
2,5,3,4,5
70,80,55,65,67,87,90,81,69,71
10
Measures of Dispersion -
• Dispersion in statistics refers to the measures of the variability of
data or terms.
• In simple terms, it shows how squeezed or scattered the variable is.
• The types of absolute measures of dispersion are:
1. Range
2. Variance
3. Mean Deviation
4. Standard Deviation
11
Measures of Dispersion -
1. Range – It is simply the difference between the maximum value
and the minimum value given in a data set.
• Examples:
• 1, 3, 5, 6, 7 => Range = 7 -1= 6
• 3, 4, 1, 5, 9, 7, 2
• 90, 72, 95, 10, 26, 12, 55, 8, 92
• 75.2, 10.1, 88.5, 90.2, 67.5, 12.5, 70.5
12
Measures of Dispersion -
2. Variance - Deduct the mean from each data in the set, square each of
them and add each square and finally divide them by the total no of values in
the data set to get the variance.
- Population vs. sample variance
• Different formulas are used for calculating variance depending on whether you
have data from a whole population or a sample.
13
Measures of Dispersion -
• Population variance-
14
Measures of Dispersion -
Sample variance -
15
Measures of Dispersion -
Steps for calculating the variance
• Step 1: Find the mean
• Step 2: Find each score’s deviation from the mean
• Step 3: Square each deviation from the mean
• Step 4: Find the sum of squares
• Step 5: Divide the sum of squares by n – 1 or N
16
Measures of Dispersion -
Example –
1.
17
Measures of Dispersion -
Step 1: Find the mean
• To find the mean, add up all the scores, then divide them by the
number of scores.
18
Measures of Dispersion -
Step 2: Find each score’s deviation from the mean
• Subtract the mean from each score to get the deviations from the mean.
• Since x̅ = 50, take away 50 from each score.
19
Measures of Dispersion -
Step 3: Square each deviation from the mean
• Multiply each deviation from the mean by itself. This will result in
positive numbers.
20
Measures of Dispersion -
Step 4: Find the sum of squares
• Add up all of the squared deviations. This is called the sum of
squares.
21
Measures of Dispersion -
Step 5: Divide the sum of squares by n – 1 or N
• Divide the sum of the squares by n – 1 (for a sample) or N (for a
population).
• Since we’re working with a sample, we’ll use n – 1, where n = 6.
22
Measures of Dispersion -
Examples –
Find variance
2) 105, 100, 102, 95, 100, 98 107
3) 4, 8, 11, 17, 20, 24, 32
4) 3, 8, 13, 18, 23
5) 5,6,8,9,10,11,14
23
Measures of Dispersion -
3. Standard Deviation - The square root of the variance is known as
the standard deviation i.e. S.D. = √var.
Examples –
• 6, 7, 10, 12, 13, 4, 8, 12
• 6, 8, 10, 12, 14, 16, 18, 20, 22, 24
24
Measures of Dispersion -
4. Mean Deviation –
• The mean deviation is defined as a statistical measure that is used to
calculate the average deviation from the mean value of the given data set.
• The mean deviation of the data values can be easily calculated using the
below procedure.
• Step 1: Find the mean value for the given data values
• Step 2: Now, subtract the mean value from each of the data values given (Note:
Ignore the minus symbol)
• Step 3: Now, find the mean of those values obtained in step 2.
Mean Deviation = [Σ |X – µ|]/N
25
Measures of Dispersion -
4. Mean Deviation –
Mean Deviation = [Σ |X – µ|]/N
Here,
• Σ represents the addition of values
• X represents each value in the data set
• µ represents the mean of the data set
• N represents the number of data values
• | | represents the absolute value, which ignores the “-” symbol
26
Measures of Dispersion -
4. Mean Deviation –
Mean Deviation for Frequency Distribution –
• To present the data in the more compressed form we group it and mention
the frequency distribution of each such group. These groups are known as
class intervals.
• Grouping of data is possible in two ways:
1. Discrete Frequency Distribution
2. Continuous Frequency Distribution
27
Measures of Dispersion -
1. Mean Deviation for Discrete Distribution Frequency –
• As the name itself suggests, by discrete we mean distinct or non-
continuous. In such a distribution the frequency (number of observations)
given in the set of data is discrete in nature.
• If the data set consists of values x1,x2, x3………xn each occurring with a
frequency of f1, f2… fn respectively then such a representation of data is
known as the discrete distribution of frequency.
28
Measures of Dispersion -
1. Mean Deviation for Discrete Distribution Frequency –
To calculate the mean deviation for grouped data and particularly for discrete
distribution data the following steps are followed:
• Step I: The measure of central tendency about which mean deviation is to
be found out is calculated. Let this measure be a.
• If this measure is mean then it is calculated as,
29
Measures of Dispersion -
1. Mean Deviation for Discrete Distribution Frequency –
• Step II: Calculate the absolute deviation of each observation from the
measure of central tendency calculated in step (I)
• Step III: The mean absolute deviation around the measure of central
tendency is then calculated by using the formula
1. If the central tendency is mean then,
2. In case of median,
30
Measures of Dispersion -
Mean Deviation Examples -
1. Determine the mean deviation for the data values 5, 3,7, 8, 4, 9.
31
Measures of Dispersion -
Mean Deviation Examples -
1. Determine the mean deviation for the data values 5, 3,7, 8, 4, 9.
Solution – Step 1: find the mean for the given data:
• Mean, µ = ( 5+3+7+8+4+9)/6
• µ = 36/6
• µ=6
• Therefore, the mean value is 6.
32
Measures of Dispersion -
Mean Deviation Examples -
1. Determine the mean deviation for the data values 5, 3,7, 8, 4, 9.
Step 2: Now, subtract each mean from the data value, and ignore the minus symbol if any
• (Ignore”-”)
5–6=1
3–6=3
7–6=1
8–6=2
4–6=2
9–6=3
Now, the obtained data set is 1, 3, 1, 2, 2, 3.
33
Measures of Dispersion -
Mean Deviation Examples -
1. Determine the mean deviation for the data values 5, 3,7, 8, 4, 9.
Step 3: Finally, find the mean value for the obtained data set
Therefore, the mean deviation is
= (1+3 + 1+ 2+ 2+3) /6
= 12/6
=2
Hence, the mean deviation for 5, 3,7, 8, 4, 9 is 2.
34
Measures of Dispersion -
Mean Deviation Examples -
2. In a foreign language class, there are 4 languages, and the frequencies of
students learning the language and the frequency of lectures per week are
given as:
35
Measures of Dispersion -
Mean Deviation Examples -
36
Measures of Dispersion -
2. Mean Deviation of Grouped Data –
• In frequency distribution of continuous type, the class intervals or groups
are arranged so that there are no gaps between the classes and each class
in the table has its respective frequency.
37
Measures of Dispersion -
2. Mean Deviation of Grouped Data –
• The following table represents the age group of employees working in a
certain company.
38
Measures of Dispersion -
2. Mean Deviation of Grouped Data –
• Steps to Calculate Mean Deviation of Continuous Frequency Distribution:
Step i) Assume that the frequency in each class is centered at the mid-point.
The mean is calculated for these mid-points.
Considering the above example, the midpoints are given as follows:
39
Measures of Dispersion -
2. Mean Deviation of Grouped Data –
• Steps to Calculate Mean Deviation of Continuous Frequency Distribution:
The mean is calculated by the formula
Step ii) The mean absolute deviation about the mean is given by:
40
Measures of Dispersion -
2. Mean Deviation of Grouped Data –
• Steps to Calculate Mean Deviation of Continuous Frequency Distribution:
41
Measures of Dispersion -
4. Mean Deviation about Median –
Mean Deviation(M) = [Σ |X – M|]/N
• Example – 6,7,10,12,13,4,8,12
• Mean Deviation about median for grouped data-
Mean Deviation(M) = ∑fi|xi−M|∑fi
∑fi=N
42
Conditional Probability -
• Conditional probability is defined as the likelihood of an event or outcome
occurring, based on the occurrence of a previous event or outcome.
• Conditional probability is calculated by multiplying the probability of the
preceding event by the updated probability of the succeeding, or conditional,
event.
43
Conditional Probability -
Example 2: In a group of 100 computer buyers, 40 bought CPU, 30
purchased monitor, and 20 purchased CPU and monitors. If a
computer buyer chose at random and bought a CPU, what is the
probability they also bought a Monitor?
44
Conditional Probability -
Solution: As per the first event, 40 out of 100 bought CPU,
• So, P(A) = 40% or 0.4
• Now, according to the question, 20 buyers purchased both CPU and monitors.
So, this is the intersection of the happening of two events. Hence,
• P(A∩B) = 20% or 0.2
• By the formula of conditional probability we know;
• P(B|A) = P(A∩B)/P(A)
• P(B|A) = 0.2/0.4 = 2/4 = ½ = 0.5
• The probability that a buyer bought a monitor, given that they purchased a CPU,
is 50%.
45
Bayes Theorem -
• Bayes' Theorem states that the conditional probability of an event,
based on the occurrence of another event, is equal to the likelihood
of the second event given the first event multiplied by the probability
of the first event.
• Bayes' Theorem calculates the conditional probability of an event,
based on the values of specific related known probabilities.
46
Bayes Theorem -
• Bayes’ Theorem has two types of probabilities :
1. Prior Probability [P(A)]
2. Posterior Probability [P(A|B)]
Here,
B – B is a data tuple.
A – A is some Hypothesis.
47
Bayes Theorem -
• Example –
P(King | Face) = P(Face | King).P(King)
P(Face)
48
Bayes Theorem -
1. Prior Probability - Prior Probability is the probability of occurring an event
before the collection of new data. It is the best logical evaluation of the probability
of an outcome which is based on the present knowledge of the event before the
inspection is performed.
2. Posterior Probability - When new data or information is collected then the
Prior Probability of an event will be revised to produce a more accurate measure of
a possible outcome. This revised probability becomes the Posterior Probability and
is calculated using Bayes’ theorem. So, the Posterior Probability is the probability
of an event B occurring given that event A has occurred.
49
Basics and need of hypothesis and hypothesis
testing -
Hypothesis –
“Hypothesis is an idea/Guess that you test using data. It helps us check if
something we believe is true or not”.
• Hypothesis is the beginning of the study that translates research questions
into predictions that might or might not be true.
• Use of word “Support”, not “Prove”.
• Powerful tool in research process(helps researcher to relate theory to
observation).
50
Basics and need of hypothesis and hypothesis
testing -
Hypothesis –
• A hypothesis is a tentative generalization, the validity of which remains to
be tested. In its most elementary stages, the hypothesis may be any hunch,
guess imaginative idea or Intuition whatsoever which becomes the basis
of action or Investigation.
• Hypothesis is a shrewd guess or inference that is formulated and
provisionally adopted to explain observed facts or conditions and to guide
in further investigation.
51
Basics and need of hypothesis and hypothesis
testing -
Two main types of hypotheses:
1. Null Hypothesis(H₀): Nothing changes or there is no effect ( e.g.
Discounts do not increase sales).
2. Alternative Hypothesis(H₁ or Ha): There is a change or effect ( e.g.
Discounts increase sales).
52
Basics and need of hypothesis and hypothesis
testing -
Characteristics of Hypothesis -
• The hypothesis should be clear and precise to consider it to be reliable.
• If the hypothesis is a relational hypothesis, then it should be stating the
relationship between variables.
• The hypothesis must be specific and should have scope for conducting more
tests.
• The way of explanation of the hypothesis must be very simple and it should also
be understood that the simplicity of the hypothesis is not related to its
significance.
53
Pearson Correlation -
• The Pearson correlation measures the strength of the linear relationship
between two variables.
• It has a value between -1 to 1, with a value of -1 meaning a total negative
linear correlation, 0 being no correlation, and + 1 meaning a total positive
correlation.
• Example – The price and demand of a product is in correlation.
54
Pearson Correlation -
• Examples –
1. X Y 2. Age(X) Weight(Y)
1 2 40 78
2 4 21 70
3 7 25 60
4 9 31 55
5 12 38 80
6 14 47 66
55
Pearson Correlation -
• Examples –
3. Stock(X) Stock(Y) 4. X Y
45 9 17 94
50 8 13 73
53 8 12 59
58 7 15 80
60 5 16 93
14 85
16 66
16 79
18 77
19 91
56
Pearson Correlation -
• Examples –
5. X Y
6 45
12 47
13 39
17 58
22 68
25 76
27 75
29 74
30 78
32 81
57
Pearson Correlation -
• Examples –
5. Solution -
58
Degree of Freedom(df)-
• Degree of freedom (df) refers to the number of independent values or
observations in a dataset that can vary while still satisfying given constraints. It
is essential in hypothesis testing, confidence intervals, and many statistical
calculations.
• Examples in Different Statistical Tests:
1. One-Sample t-Test:
df = n – 1
where, n = sample size
2. Two-Sample t-Test (Independent Samples):
df=(n1−1)+(n2−1)
Where n1 and n2 are the sample sizes of both groups.
3. Chi-Square Test:
df=(Rows−1)×(Columns−1)
59
Chi-Square test -
• A chi-square test is a statistical test that is used to compare observed and
expected results.
60
Chi-Square test -
1. Agree Undecided Married Total
Male 10 13 11 34
Female 15 12 14 41
Total 25 25 25 75
61
Chi-Square test -
2.
62
Chi-Square test -
3.
Low Medium High Total
For 213 203 182
Against 138 110 154
Total
63
T-test -
• A t-test is a statistical test that is used to compare the means of two groups.
• It is often used in hypothesis testing to determine whether a process or treatment
actually has an effect on the population of interest, or whether two groups are
different from one another.
• A t-test can only be used when comparing the means of two.
64
T-test -
65
T-test -
.
66
T-test -
.
67
X1 X2
15.2 15.9
T-test- 15.3 15.9
16.0 15.2
• Examples – 15.8 16.6
5. 15.6 15.2
14.9 15.8
15.0 15.8
15.4 16.2
15.6 15.6
15.7 15.6
15.5 15.8
15.2 15.5
15.5 15.5
15.1 15.5
15.3 14.9
15.0 15.9
68