Statistics & Probability
Typed Notes — Measures of Dispersion, Correlation, Regression, Probability Theory
Compiled and typed from handwritten class notes (15 May 2026 – 1 Jul 2026). Organised by topic; original worked
examples are preserved with cleaned-up steps.
1. Measures of Dispersion — Range & Coefficient of Range
2. Quartile Deviation (Q.D.) & Coefficient of Q.D.
3. Mean Deviation (M.D.) & Coefficient of M.D.
4. Mean, Median, Mode of Grouped Data + Empirical Relationship
5. Variance & Standard Deviation — Natural Numbers Example
6. Karl Pearson's Correlation, Regression Analysis (Bivariate)
7. Spearman's Rank Correlation Coefficient (with & without ties)
8. Probability — Basic Terminology & Axioms
9. Theorem of Total Probability & Addition Theorem
10. Bayes' Theorem
11. Worked Probability Problems
1. Measures of Dispersion
Dispersion measures how the observations are deviated, in an overall sense, from a central (usually mean)
value.
Absolute Measures Relative Measures
Range (R) Coefficient of Range
Quartile Deviation (Q.D.) Coefficient of Q.D.
Mean Deviation (M.D.) Coefficient of M.D.
Standard Deviation (S.D.) Coefficient of S.D. (Variation)
Range
R = Maximum value − Minimum value
Example: Data set: 3, 7, 14, 11, 2, 4, 8, 13
Maximum = 14, Minimum = 2
R = 14 − 2 = 12
Coefficient of Range
Coefficient of R = (Max value − Min value) / (Max value + Min value) × 100
Using the same data:
Coefficient of R = 12 / (14 + 2) × 100 = 12/16 × 100 = 75%
2. Quartile Deviation
Q.D. = (Q3 − Q1) / 2
Q3 = 3rd Quartile = value of the variable such that CF< at 3N/4 covers 75% of the observations.
Q = 1st Quartile = value of the variable such that CF< at N/4 covers 25% of the observations.
1
Q3 = l1 + [(3N/4 − F) / f3] × C
where: l = lower class-boundary of the Q class | 3N/4 = three-fourths of total frequency | F = CF<
1 3
preceding the Q class | f = ordinary frequency of the Q class | C = common class width.
3 3 3
Q1 is found the same way, replacing 3N/4 with N/4 and locating the Q1 class.
Worked Example
Class F CF
10-20 5 5
20-30 7 12
30-40 18 (f■) 30
40-50 31 (f■) 61 (F■)
50-60 24 (f■) 85
60-70 12 97
70-80 3 100
Step 1: 3N/4 = (3/4) × 100 = 75 → Q3 class = 50–60 class (CF just exceeding 75 is 85).
(Note: original notes marked the 40–50 class as the reference row for the Q3 lookup boundary.)
Q3 = l1 + [(3N/4 − F)/f3] × C
Coefficient of Quartile Deviation
Co-eff. of Q.D. = (Q3 − Q1) / (Q3 + Q1) × 100
Alternative form seen in notes: Co-eff. of Q.D. = Q.D. / Q2 × 100 (Q2 = median).
3. Mean Deviation (M.D.)
Mean Deviation can be measured about the Mean, the Median, the Mode, or any arbitrary constant A.
Co-eff. of M.D./Mean = (M.D. about Mean / Mean) × 100
Co-eff. of M.D./Median = (M.D. about Median / Median) × 100
Co-eff. of M.D./Mode = (M.D. about Mode / Mode) × 100
Co-eff. of M.D./A = (M.D. about A / A) × 100 (A = arbitrary constant)
Worked Example
Given: M.D. about Mean = 17.14, Mean = 45
Co-eff. of M.D./Mean = (17.14 / 45) × 100 = 38.08%
4. Mean, Median & Mode of Grouped Data
Mode of Grouped Observations
Mo = l1 + [d1 / (d1 + d2)] × C
where d = f − f , d =f −f , l = lower boundary of modal class, f = frequency of modal class.
1 m m−1 2 m m+1 1 m
Worked Example (7 class intervals, N = 100)
Class Freq
10-20 5
20-30 7
30-40 18 (d■)
40-50 (Modal) 31
50-60 24 (d■)
60-70 12
70-80 3
Mo = 40 + (31−18) / [(31−18)+(31−24)] × 10
= 40 + 13/(13+7) × 10 = 40 + 13/20 × 10 = 40 + 6.5 = 46.5
Empirical Relationship Among Mean, Median & Mode
If a data-set has moderate skewness (i.e. skewness magnitude < 3), then:
Mean − Mode = 3 (Mean − Median)
Equivalently: M − Mo = 3(M − Mc), where M = Mean, Mc = Median, Mo = Mode.
Worked Example
Given Mode Mo = 46.5, Median Mc = 46.45 (from a similar grouped data-set):
M − 46.5 = 3(M − 46.45)
M − 46.5 = 3M − 139.35 ⇒ 3M − M = 139.35 − 46.5 = 92.85
2M = 92.85 ⇒ M = 46.425
A second version in the notes uses Mode = 53.33 and Median = 51.67, giving M − 53.33 = 3(M − 51.67), ⇒ 2M =
101.66 ⇒ M = 50.835.
Advantages / Disadvantages of Mean, Median, Mode
For a raw data set such as 2, 5, 6, 8, 9, 11, 11, 13, 15, 18, 21, 21, 21 — Mean, Median and Mode would
each summarise the data differently; the choice depends on the presence of outliers, the level of
measurement, and whether the distribution is skewed. (See separate write-up if a detailed comparison table
is required.)
Second Worked Problem — Mean, Median, Mode & S.D. from a “candidates
obtaining marks X or higher” table
This is a cumulative frequency (CF≥) “more than” type distribution. n = 140 candidates.
X CF (≥) F■ (ordinary) F■ CF≤
10 140 7 0 0
20 133 15 7 7
30 118 18 15 22
40 100 25 18 40
50 (M→) 75 30 25 65
60 45 20 30 (f■) 95
70 25 16 20 115
80 9 7 16 131
90 2 2 7 138
100 0 0 2 160
N/2 = 140/2 = 70, which falls in the 50 class using CF≤.
Mc = l1 + [(N/2 − F2)/f2] × C
By simple interpolation: (M − 50)/(60−50) = (70−65)/(95−65)
c
(Mc − 50)/10 = 5/30 ⇒ Mc = 50 + 5/3 = 50 + 1.67 = 51.67 (approx.)
Mo = l1 + d1/(d1+d2) × C = 50 + (30−25)/[(30−25)+(30−20)] × 10
= 50 + 5/(5+10) × 10 = 50 + 10/3 = 50 + 3.34 = 53.34
5. Variance & Standard Deviation — 10 Natural Numbers
Problem: Find the variance and S.D. of the first 10 natural numbers having the same value as their
corresponding frequencies (i.e. x = f = 1,2,3,...,10).
Var(x) = σx2 = (1/N) ∑ fi(xi − x■)2 = ∑fx2/N − (∑fx/N)2
Here N = ∑f = 1+2+...+10 = 55
∑fx = 12+22+...+102 = n(n+1)(2n+1)/6 = 10·11·21/6 = 385
∑fx2 = 13+23+...+103 = [n(n+1)/2]2 = (5×11)2 = 552 = 3025
σx2 = 3025/55 − (385/55)2 = 55 − 72 = 55 − 49 = 6
Variance = 6
S.D. = σx = √Var(x) = √6 = 2.45
S.D. = 2.45
Coefficient of S.D.
Co-eff. of S.D. = (S.D./Mean) × 100 = (2.45/7) × 100 = 245/7 = 35%
Mean = ∑fx/∑f = 385/55 = 7.
6. Bivariate Regression Analysis
Regression → to move back (estimate one variable from another using a fitted historical relationship).
Progression → to move ahead.
y = f(x) — y is the dependent variable, x is the independent variable.
Regression line of y on x
y − y■ = byx (x − x■)
byx = Cov(x,y) / σx2
Regression line of x on y
x − x■ = bxy (y − y■)
bxy = Cov(x,y) / σy2 = r · σx/σy
bxy is called the “ratio of S.D. of x to y” scaled by r.
Supporting formulas
Cov(x,y) = ∑xy/n − (∑x/n)(∑y/n)
Var(x) = σx2 = ∑x2/n − (∑x/n)2
rxy = Cov(x,y) / (σx·σy)
Cov(x,y) = rxy·σx·σy
Worked Example (n = 10 paired observations, x = y = 1..10)
∑x = ∑y = 55, ∑xy = 220, x■ = y■ = 5.5, n = 10
Cov(x,y) = 220/10 − (5.5)(5.5) = 22 − 30.25 = −8.25
σx2 = 385/10 − (5.5)2 = 38.5 − 30.25 = 8.25 ⇒ σx = √8.25
σy2 = 8.25 (by symmetry) ⇒ σy = √8.25
rxy = −8.25 / (√8.25 × √8.25) = −8.25/8.25 = −1
r■■ = −1 (perfect negative correlation)
7. Spearman's Rank Correlation Coefficient
Without repeated ranks (ties)
ρxy = 1 − 6∑d2 / (n3 − n)
where d = R1 − R2 (difference of ranks) and n = number of observations.
Worked Example — Marks in Statistics & Accountancy, ranked, n = 4:
Student Statistics Accountancy R■ R■ d = R■−R■ d²
S 40 35 2 3 −1 1
R 38 40 3 1 2 4
N 42 34 1 4 −3 9
P 28 36 4 2 2 4
∑d² = 18
ρxy = 1 − 6(18)/(43−4) = 1 − 108/60 = 1 − 1.8 = −0.8
Rank correlation = −0.8 (strong negative correlation between the two subjects)
With repeated ranks (ties) — corrected formula
ρ = 1 − 6[∑d2 + ∑m(m2−1)/12] / (n3 − n)
m = number of times a rank repeats (a correction term m(m²−1)/12 is added for each group of tied ranks).
Worked Example — Two judges A & B rank 10 contestants in a musical contest; several ties occur and the
corrected Spearman coefficient is required, taking all 4 ties into consideration.
[Link] Judge A Judge B R■ R■ d = R■−R■ d²
1 10 4 1 8 −7 49
2 6 5 4 6 −2 4
3 5 8 6 2.5 3.5 12.25
4 6 4 4 8 −4 16
5 3 7 8 4 4 16
6 2 10 9.5 1 8.5 72.25
7 4 8 7 2.5 4.5 20.25
8 6 1 4 10 −6 36
9 7 6 2 5 −3 9
10 2 2 9.5 8 1.5 2.25
∑d² = 237 (before tie correction)
Ties in Judge A's ranks: value 6 repeats 3 times (rank 4), value 2 repeats 2 times (rank 9.5) → correction terms of
the form m(m²−1)/12 are added for each tied group, and similarly for Judge B's ties.
Applying the full formula with all tie-corrections (4 tied groups across both judges) gives the corrected
Spearman's rank correlation coefficient for this data set.
8. Probability — Basic Terminology & Axioms
Term Meaning
Random Experiment
An experiment with an uncertain outcome, e.g. coin tossing, card drawing, dice throwing, drawing from a bag.
Events & Outcome Events are user-defined and denoted A, B, C, D... — the face that turns up.
Mutually Exclusive
A &
(ME) events
B are ME if occurrence of A guarantees non-occurrence of B; A & B cannot occur simultaneously.
Mutually ExhaustiveThe
events
set of all possible outcomes that a random experiment can generate, which are ME to each other.
Independent events Occurrence of one event does not affect the probability of the other.
Key probability rules
3. P(A ∪ B) = P(A+B) = probability of occurrence of at least one of the events, i.e. either A or B or both.
4. P(A ∩ B) = P(AB) = probability of occurrence of the joint (compound) event — A and B together /
simultaneously.
5. P(A/B) = probability of occurrence of A, given that B has already taken place (conditional probability).
Conversely: P(B/A) = probability of occurrence of B, given that A has already occurred.
Range of probability
0 ≤ P(A) ≤ 1
P(A) = 0 0 < P(A) < 1 P(A) = 1
Impossible event — Sure event
Theorem 2 — Bounds on the correlation coefficient
−1 ≤ rxy ≤ +1 ⇒ |rxy| ≤ 1
r = −1 −1 < r < 0 r=0 0<r<1 r = +1
Perfect negative correlation Negative correlation No linear correlation Positive correlation Perfect positive correlation
9. Addition Theorem & Theorem of Total Probability
Addition Theorem (Additive model)
Disjoint (mutually exclusive) sets:
P(A ∪ B) = P(A) + P(B)
When A and B are not mutually exclusive (joint set exists):
P(A ∪ B) = P(A) + P(B) − P(A ∩ B)
Theorem of Total Probability
Statement: If two events A & B are mutually exclusive, then the probability of occurrence of at least one of
them is given by P(A ∪ B) = P(A) + P(B) (disjoint sets).
More generally, for the union of several events:
P(E■ ∪ E■ ∪ E■ ∪ E■) = P(E■)+P(E■)+P(E■)+P(E■) = ∑P(Ei)
Numerical check (from notes): 18/35 + 4/35 + 12/35 + 1/35 = 35/35 = 1 (certain event).
10. Bayes' Theorem
(On reverse conditional probability)
Statement: An event A occurs if any of the mutually exclusive events B1, B2, ..., Bn occurs. If the
probabilities of B1,...,Bn are known, and the conditional probabilities P(A/B1), P(A/B2), ..., P(A/Bn) are known,
then the probability of (Bi/A) is given by:
P(Bi/A) = [P(Bi)·P(A/Bi)] / ∑j=1n [P(Bj)·P(A/Bj)]
Worked Example — Three machines M■, M■, M■
Three automatic machines M1, M2, M3 produce 30%, 20%, 50% of total output respectively, with defective
rates 3%, 2%, 5%. An item is selected at random and found defective. Find the probability it came from M3.
P(M1)=0.3, P(M2)=0.2, P(M3)=0.5
P(D/M1)=0.03, P(D/M2)=0.02, P(D/M3)=0.05
P(M3/D) = [0.5×0.05] / [0.3×0.03 + 0.2×0.02 + 0.5×0.05]
= 0.025 / (0.009+0.004+0.025) = 0.025/0.038 = 0.6578 ≈ 0.66
P(M■/D) ≈ 0.66
11. Worked Probability Problems
11.1 Contractor and two contracts (P■ and E)
P(getting plumbing contract) P(P ) = 3/4. P(not getting electrical contract) P(Ec) = 1/3. P(at least one
1
contract) = P(P ∪E) = 4/5. Find P(both contracts).
1
P(P■)=3/4, P(P■■)=1−3/4=1/4
P(E■)=1/3, P(E)=1−1/3=2/3
P(P■∪E)=P(P■)+P(E)−P(P■∩E)
4/5 = 3/4 + 2/3 − P(P■∩E)
P(P■∩E) = 3/4+2/3−4/5 = (45+40−48)/60 = 37/60
P(both contracts) = 37/60
11.2 Three horses in a race
Prob. of A winning is twice that of B, and thrice that of C. Find each probability.
Let P(C) = p. Then P(B) = 3p (since A is thrice C's, and A is twice B's ⇒ B = 3p), P(A) = 6p.
S = {p, 3p, 6p}, p+3p+6p = 1 ⇒ 10p = 1 ⇒ p = 1/10
P(A) = 6/10, P(B) = 3/10, P(C) = 1/10
11.3 Two friends speak the truth
A speaks truth in 70% of cases, B speaks false in 40% of cases. Find the % of cases in which they are likely
to contradict each other.
P(A■)=0.7, P(A■)=0.3; P(B■)=0.4, P(B■)=0.6
T F F T
Contradiction occurs when exactly one of them tells the truth (the ME events (A ∩B ) or (A ∩B ) occurs):
P[(A■∩B■) + (A■∩B■)] = P(A■)P(B■) + P(A■)P(B■)
= 0.7×0.4 + 0.3×0.6 = 0.28+0.18 = 0.46
They contradict each other in 46% of the cases (so agree in 54% of cases).
Check — probability they agree: P(A■∩B■)+P(A■∩B■) = 0.7×0.6+0.3×0.4 = 0.42+0.12 = 0.54 = 54%. ✓
11.4 Picnic weather problem (Theorem of Total Probability)
P(rain) = 90%. If it rains, P(good picnic) = 0.3. Otherwise P(good picnic) = 0.8. Find P(picnic will be good).
P(R)=0.9, P(R■)=0.1
P(G■/R)=0.3, P(G■/R■)=0.8
Picnic is good if either of the two ME events (G ∩R) or (G ∩Rc) occurs:
p p
P(G■) = P(R)·P(G■/R) + P(R■)·P(G■/R■)
= 0.9×0.3 + 0.1×0.8 = 0.27+0.08 = 0.35
P(good picnic) = 0.35 = 35%
11.5 Drawing 3 balls (4 White, 3 Black) — combinatorial probability
A bag has 4 White (W) and 3 Black (B) balls; 3 balls are drawn at random. Find the probability of each
composition:
nCr = n! / [r!(n−r)!], total ways = ■C■ = 35
Event Combination Probability
Exactly 2 White & 1 Black ■C■ × ³C■ / ■C■ 18/35
3 White & 0 Black ■C■ × ³C■ / ■C■ 4/35
1 White & 2 Black ■C■ × ³C■ / ■C■ 12/35
0 White & 3 Black ■C■ × ³C■ / ■C■ 1/35
Sanity check: 18/35 + 4/35 + 12/35 + 1/35 = 35/35 = 1 (all possible compositions accounted for).