0% found this document useful (0 votes)
6 views10 pages

Answer Key 67004 Statistical Analysis

The document is an answer key for a statistical analysis exam, detailing properties of correlation, median, skewness, price index numbers, and variance. It includes calculations for mode, standard deviation, time series forecasting, quartile deviation, and rank correlation. Additionally, it covers methods for fitting a straight line using the least squares method and provides examples and formulas for various statistical concepts.

Uploaded by

deepikabarlota
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views10 pages

Answer Key 67004 Statistical Analysis

The document is an answer key for a statistical analysis exam, detailing properties of correlation, median, skewness, price index numbers, and variance. It includes calculations for mode, standard deviation, time series forecasting, quartile deviation, and rank correlation. Additionally, it covers methods for fitting a straight line using the least squares method and provides examples and formulas for various statistical concepts.

Uploaded by

deepikabarlota
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

ANSWER KEY

Paper / Subject Code: 67004 | Statistical Analysis


Time: 2½ hrs | Total Marks: 75

Q1: Attempt ALL (Short Notes) [10 Marks]


A. Properties of Correlation
1. Correlation lies between -1 and +1 (i.e., -1 ≤ r ≤ +1).
2. It is a pure number — no units.
3. r is symmetric: r(X,Y) = r(Y,X).
4. Independent of change of origin and scale.
5. +1 = perfect positive, -1 = perfect negative, 0 = no linear correlation.
6. r only measures LINEAR relationship; it does not imply causation.

B. Properties of Median
1. It is the middle value of an ordered dataset — divides distribution into two equal halves.
2. Not affected by extreme values (outliers) — hence preferred for skewed data.
3. Can be calculated for open-end class intervals.
4. It can be located graphically using an Ogive.
5. For a symmetrical distribution: Mean = Median = Mode.
6. It may not be an actual value in the data (it is interpolated for grouped data).

C. Positive vs Negative Skewness


Feature Positive Skewness Negative Skewness
Shape of tail Right (longer right tail) Left (longer left tail)
Relation of averages Mean > Median > Mode Mean < Median < Mode
Sk value Sk > 0 Sk < 0
Example Income distribution Exam scores when most score high

D. Laspeyres & Fisher's Price Index Numbers


Laspeyres' Price Index: Uses BASE YEAR quantities (Q0) as weights.
L = (ΣP1Q0 / ΣP0Q0) × 100
It tends to OVERSTATE price increase since it uses old consumption patterns.
Fisher's Ideal Price Index: Geometric mean of Laspeyres and Paasche.
F = √(L × Pa) where Pa = (ΣP1Q1 / ΣP0Q1) × 100
Called 'Ideal' because it satisfies both the Time Reversal Test and Factor Reversal Test.

E. Properties of Variance
1. Variance = σ² (square of Standard Deviation). Always non-negative.
2. Variance = 0 means all values are identical.
3. Not independent of change of scale; it changes when data is multiplied by a constant.
4. Variance is additive for independent random variables: Var(X+Y) = Var(X) + Var(Y).
5. It gives more weight to extreme deviations (squares them).
Var = Σ(X - X̅ )² / N (ungrouped) | Σf(X - X̅ )² / Σf (grouped)
Q2: Attempt Any THREE [15 Marks]
A. Mode from Frequency Distribution (Histogram question — locate mode)
Data:
CI 50-100 100-150 150-200 200-250 250-300 300-350 350-400 400-450
f 12 18 26 14 10 9 3 17

▶ Highest frequency = 26, so Modal Class = 150-200


▶ L = 150, f1 = 26, f0 = 18 (preceding), f2 = 14 (succeeding), i = 50
Mode = L + [(f1 - f0) / (2f1 - f0 - f2)] × i
Mode = 150 + [(26 - 18) / (2×26 - 18 - 14)] × 50
Mode = 150 + [8 / (52 - 32)] × 50 = 150 + (8/20) × 50
Mode = 150 + 20 = 170
✅ Mode = 170

B. Standard Deviation
Data:
CI 20-40 40-60 60-80 80-100 100-120 120-140 140-160 160-180
f 11 10 8 7 4 10 3 2

Using Step Deviation Method: A (Assumed Mean) = 90, i = 20

CI Mid (X) f d=(X-A)/i fd fd²


20-40 30 11 -3 -33 99
40-60 50 10 -2 -20 40
60-80 70 8 -1 -8 8
80-100 90 7 0 0 0
100-120 110 4 1 4 4
120-140 130 10 2 20 40
140-160 150 3 3 9 27
160-180 170 2 4 8 32
Total 55 -20 250

▶ N = Σf = 55, Σfd = -20, Σfd² = 250


σ = i × √[(Σfd²/N) - (Σfd/N)²]
σ = 20 × √[(250/55) - (-20/55)²]
σ = 20 × √[4.5455 - (-0.3636)²]
σ = 20 × √[4.5455 - 0.1322]
σ = 20 × √4.4133 = 20 × 2.1008
✅ Standard Deviation (σ) = 42.02

C. Uses and Components of Time Series in Business Forecasting


TIME SERIES is a set of data recorded at successive time intervals. It helps businesses forecast future trends.
USES IN BUSINESS FORECASTING:
1. Demand Forecasting: Predicting future sales to plan production and inventory.
2. Financial Planning: Budget preparation and profit projections.
3. Economic Policy: Government uses it to study GDP, inflation, unemployment trends.
4. Stock Market Analysis: Identifying price trends and seasonal patterns.
5. Supply Chain Management: Anticipate seasonal demand to avoid shortages.
COMPONENTS OF TIME SERIES:
1. Secular Trend (T): Long-term upward/downward movement. E.g. population growth, rising sales.
2. Seasonal Variations (S): Regular periodic changes within a year. E.g. umbrella sales in monsoon, AC sales in summer.
3. Cyclical Variations (C): Wave-like fluctuations over many years, linked to business cycles (boom, recession, recovery).
4. Irregular / Random Variations (I): Unpredictable, erratic changes due to unforeseen events. E.g. floods, wars,
pandemics.
Model: Y = T × S × C × I (Multiplicative) or Y = T + S + C + I (Additive)

D. Quartile Deviation
Data:
CI 100-150 150-200 200-250 250-300 300-350 350-400 400-450 450-500
f 9 11 12 8 12 23 11 8

CI f Cumulative Frequency (CF)


100-150 9 9
150-200 11 20
200-250 12 32
250-300 8 40
300-350 12 52
350-400 23 75
400-450 11 86
450-500 8 94
Total (N) 94

▶ N = 94
▶ Q1 class: N/4 = 23.5 → falls in 150-200 (CF before = 9, f = 11, L = 150)
Q1 = 150 + [(23.5 - 9) / 11] × 50 = 150 + (14.5/11) × 50 = 150 + 65.91
= 215.91
▶ Q3 class: 3N/4 = 70.5 → falls in 350-400 (CF before = 52, f = 23, L = 350)
Q3 = 350 + [(70.5 - 52) / 23] × 50 = 350 + (18.5/23) × 50 = 350 + 40.22
= 390.22
Quartile Deviation (QD) = (Q3 - Q1) / 2 = (390.22 - 215.91) / 2 =
174.31 / 2
✅ QD = 87.155
Coefficient of QD = (Q3 - Q1)/(Q3 + Q1) = 174.31 / 606.13 = 0.2876
✅ Coefficient of QD = 0.288 (approx)

E. Index Numbers — Explanation


An Index Number is a statistical measure that shows relative changes in a variable or group of variables compared to a
base period (= 100).
Types: (1) Price Index Numbers (2) Quantity Index Numbers (3) Value Index Numbers
Uses: Measuring inflation (CPI), industrial production (IIP), cost of living, comparing economic performance.
Base Period: The reference period against which comparisons are made. Its index = 100.
Characteristics: Expressed as percentages, always relative measures, aggregative in nature.
Q3: Attempt Any TWO [20 Marks]
A. Laspeyres, Paasche, and Fisher's Price Index Numbers
Items P0 Q0 P1 Q1 P0Q0 P1Q0 P0Q1 P1Q1
A 100 12 122 13 1200 1464 1300 1586
B 123 13 156 12 1599 2028 1476 1872
C 167 17 143 19 2839 2431 3173 2717
D 125 18 129 14 2250 2322 1750 1806
E 160 12 165 10 1920 1980 1600 1650
Σ 9808 10225 9299 9631

▶ ΣP0Q0 = 9808, ΣP1Q0 = 10225, ΣP0Q1 = 9299, ΣP1Q1 = 9631


Laspeyres' Index (L) = (ΣP1Q0 / ΣP0Q0) × 100 = (10225 / 9808) × 100
✅ Laspeyres' Index = 104.25
Paasche's Index (Pa) = (ΣP1Q1 / ΣP0Q1) × 100 = (9631 / 9299) × 100
✅ Paasche's Index = 103.57
Fisher's Ideal Index = √(L × Pa) = √(104.25 × 103.57) = √10796.9
✅ Fisher's Ideal Index = 103.91

B. Rank Correlation (Spearman's ρ)


Data:
X 11 4 8 11 17 5 13
Y 16 5 16 7 16 8 12

X Y Rank (Rx) Rank (Ry) D = Rx-Ry D²


11 16 4.5 (tie) 6 (tie) -1.5 2.25
4 5 1 1 0 0
8 16 3 6 (tie) -3 9
11 7 4.5 (tie) 2 2.5 6.25
17 16 7 6 (tie) 1 1
5 8 2 3 -1 1
13 12 6 4 2 4
ΣD² 23.5

▶ X: value 11 appears 2 times → ranks 4,5 → avg rank = 4.5 | CF_X = (2³-2)/12 = 0.5
▶ Y: value 16 appears 3 times → ranks 5,6,7 → avg rank = 6 | CF_Y = (3³-3)/12 = 24/12 = 2
▶ N = 7, ΣD² = 23.5
ρ = 1 - [6(ΣD² + CF_X + CF_Y)] / [N(N² - 1)]
ρ = 1 - [6 × (23.5 + 0.5 + 2)] / [7 × (49 - 1)]
ρ = 1 - [6 × 26] / [7 × 48]
ρ = 1 - 156 / 336 = 1 - 0.4643
✅ Spearman's Rank Correlation (ρ) = 0.536 (Moderate Positive Correlation)

C. P50 (Median) and D4 for given data


CI 0-100 100-200 200-300 300-400 400-500 500-600 600-700 700-800
f 12 6 5 18 23 15 12 10

CI f CF
0-100 12 12
100-200 6 18
200-300 5 23
300-400 18 41
400-500 23 64
500-600 15 79
600-700 12 91
700-800 10 101
Total 101

▶ N = 101
▶ P50 = Median: N/2 = 50.5 → falls in 400-500 (CF before = 41, f = 23, L = 400, i = 100)
P50 = 400 + [(50.5 - 41) / 23] × 100 = 400 + (9.5/23) × 100 = 400 +
41.30
✅ P50 = 441.30
▶ D4 (4th Decile): 4N/10 = 4×101/10 = 40.4 → falls in 300-400 (CF before = 23, f = 18, L = 300, i = 100)
D4 = 300 + [(40.4 - 23) / 18] × 100 = 300 + (17.4/18) × 100 = 300 +
96.67
✅ D4 = 396.67
Q4: Attempt Any THREE [30 Marks]
a. Least Square Method — Fit Straight Line
Year 2005 2006 2007 2008 2009 2010 2011
Production 67 56 78 43 12 19 23
(Y)

▶ Origin: 2008 (middle year), i = 1 year


Year Y X (coded) XY X²
2005 67 -3 -201 9
2006 56 -2 -112 4
2007 78 -1 -78 1
2008 43 0 0 0
2009 12 1 12 1
2010 19 2 38 4
2011 23 3 69 9
Σ 298 0 -272 28

▶ N = 7, ΣY = 298, ΣXY = -272, ΣX² = 28, ΣX = 0


a = ΣY / N = 298 / 7 = 42.57
b = ΣXY / ΣX² = -272 / 28 = -9.71
✅ Trend Equation: Yc = 42.57 - 9.71X (Origin: 2008)
▶ Negative b means production is declining by 9.71 units per year.

b. Regression Equations
X 12 6 8 4 11 8 10
Y 10 7 21 16 10 5 7

X Y X² Y² XY
12 10 144 100 120
6 7 36 49 42
8 21 64 441 168
4 16 16 256 64
11 10 121 100 110
8 5 64 25 40
10 7 100 49 70
59 (ΣX) 76 (ΣY) 545 (ΣX²) 1020 (ΣY²) 614 (ΣXY)

▶ N = 7, X̅ = 59/7 = 8.43, Y̅ = 76/7 = 10.86


byx = [NΣXY - ΣXΣY] / [NΣX² - (ΣX)²]
byx = [7×614 - 59×76] / [7×545 - 59²] = [4298 - 4484] / [3815 - 3481] =
-186/334
byx = -0.557
bxy = [NΣXY - ΣXΣY] / [NΣY² - (ΣY)²]
bxy = [4298 - 4484] / [7×1020 - 76²] = -186 / [7140 - 5776] = -186/1364
bxy = -0.136
Regression Line of Y on X:
Y - Y̅ = byx(X - X̅ )
Y - 10.86 = -0.557(X - 8.43)
✅ Y = 15.56 - 0.557X ...(Regression of Y on X)
Regression Line of X on Y:
X - X̅ = bxy(Y - Y̅ )
X - 8.43 = -0.136(Y - 10.86)
✅ X = 9.907 - 0.136Y ...(Regression of X on Y)
r = √(byx × bxy) = √(0.557 × 0.136) = √0.0757 = -0.275 (negative
since both b are negative)

c. Combined Standard Deviation


Group N Mean (X̅) SD (σ)
1 50 45 6
2 30 48 5

▶ Step 1: Combined Mean


X̅ _combined = (N1X̅ 1 + N2X̅ 2) / (N1 + N2) = (50×45 + 30×48) / (50+30)
= (2250 + 1440) / 80 = 3690 / 80 = 46.125
▶ Step 2: Deviations from combined mean
d1 = X̅ 1 - X̅ = 45 - 46.125 = -1.125
d2 = X̅ 2 - X̅ = 48 - 46.125 = 1.875
▶ Step 3: Combined Standard Deviation
σc = √[(N1(σ1² + d1²) + N2(σ2² + d2²)) / (N1 + N2)]
σc = √[(50(36 + 1.2656) + 30(25 + 3.5156)) / 80]
σc = √[(50 × 37.2656 + 30 × 28.5156) / 80]
σc = √[(1863.28 + 855.47) / 80] = √[2718.75/80] = √33.984
✅ Combined Standard Deviation = 5.83 (approx)

d. Karl Pearson's Coefficient of Correlation


X 11 34 6 11 5 4 11
Y 12 8 21 8 9 11 10
X Y X² Y² XY
11 12 121 144 132
34 8 1156 64 272
6 21 36 441 126
11 8 121 64 88
5 9 25 81 45
4 11 16 121 44
11 10 121 100 110
82 (ΣX) 79 (ΣY) 1596 (ΣX²) 1015 (ΣY²) 817 (ΣXY)

▶ N = 7, ΣX = 82, ΣY = 79, ΣX² = 1596, ΣY² = 1015, ΣXY = 817


r = [NΣXY - ΣXΣY] / √{[NΣX² - (ΣX)²][NΣY² - (ΣY)²]}
r = [7×817 - 82×79] / √{[7×1596 - 82²][7×1015 - 79²]}
r = [5719 - 6478] / √{[11172 - 6724][7105 - 6241]}
r = -759 / √{4448 × 864}
r = -759 / √{3842832} = -759 / 1960.3
✅ Karl Pearson's r = -0.387 (Low Negative Correlation)

e. Regression — Explanation
Regression analysis is a statistical method that establishes the mathematical relationship between a dependent variable
and one or more independent variables, enabling prediction.
TWO Lines of Regression exist:
1. Regression of Y on X: Used to predict Y (dependent) for a given value of X (independent).
Y - Y̅ = byx(X - X̅ ) where byx = r(σy/σx)
2. Regression of X on Y: Used to predict X for a given value of Y.
X - X̅ = bxy(Y - Y̅ ) where bxy = r(σx/σy)
Important Properties:
• Both regression lines pass through the point (X̅, Y̅).
• The correlation coefficient r = √(byx × bxy).
• Both regression coefficients must carry the same sign.
• If r = 0, the lines are perpendicular; if r = ±1, both lines coincide.

--- End of Answer Key ---

You might also like