Paper II: Statistical Analysis for
Business Research
COMPREHENSIVE DOCTORAL STUDY GUIDE & REFERENCE
MATRIX
Bharathiar University — Ph.D. Part I Examinations
Curriculum Alignment Matrix: 40% Theory & 60% Advanced Problem Workflows
Examination Blueprint Note: As per university regulations, the question paper covers exactly 40%
theoretical frameworks and 60% numerical problem derivation and analysis.
UNIT I: BUSINESS STATISTICS & DESCRIPTIVE
MEASURES
1. Foundations of Business Statistics
Meaning and Definition
Business Statistics represents an applied branch of advanced data science involving the systematic
collection, organization, presentation, analysis, and mathematical interpretation of quantitative datasets to
mitigate operational risk and optimize corporate decision-making structures under conditions of
macroeconomic or systemic uncertainty.
Bharathiar University Ph.D. Part I Examination • Paper II Page 1
Classical Formulations:
• Croxton and Cowden: "Statistics may be defined as the collection, presentation, analysis, and
interpretation of numerical data."
• Horace Secrist: "By statistics we mean aggregates of facts affected to a marked extent by a
multiplicity of causes, numerically expressed, enumerated or estimated according to reasonable
standards of accuracy, collected in a systematic manner for a predetermined purpose and placed
in relation to each other."
Scope and Operational Functions in Corporate Research
• Market Analysis and Consumer Demographics: Econometric modeling of purchasing elasticities,
brand migration patterns, and market share distribution across diverse geographical landscapes.
• Statistical Quality Control (SQC): Deploying Shewhart control charts, Six Sigma tracking, and
acceptance sampling frameworks to minimize variance within continuous manufacturing lines.
• Financial Engineering & Risk Forecasting: Utilizing stochastic modeling, capital asset pricing
structures, and volatility forecasting metrics to optimize equity portfolios.
• Strategic Human Capital Management: Quantitative processing of employee lifecycle metrics,
actuarial turnover vectors, and objective performance distribution structures.
Epistemological Advantages and Structural Limitations
Core Operational Advantages Inherent Methodological Limitations
Aggregate Exclusive Paradigm: Statistical
Definiteness of Expression: Converts complex,
techniques are formulated exclusively for aggregate
vague corporate qualitative realities into objective,
datasets; they cannot derive independent properties
mathematically clear numeric values.
for isolated observations.
Condensation of Mass Datasets: Compresses Qualitative Blindspots: Cannot directly process
millions of data rows into actionable, descriptive non-quantified qualitative attributes (e.g., brand
parameters like central trends and dispersion loyalty, corporate integrity) without proxy
metrics. measurement scaling.
Probabilistic Indeterminism: Statistical
Comparative Benchmarking: Enables clear cross-
conclusions are valid only on average and
sectional and longitudinal comparative analysis
macroscopically; they do not dictate exact
across competitive industries.
deterministic individual outcomes.
Bharathiar University Ph.D. Part I Examination • Paper II Page 2
2. Foundations of Data & Sampling Architecture
Taxonomy of Variables and Data Modalities
Data forms the foundational empirical basis of business research. It is categorized broadly into Primary
Data (investigator-collected via surveys or experimental tracking) and Secondary Data (leveraging
compiled financial sheets or census repositories).
• Qualitative (Categorical) Variables: Nominal or ordinal scales that describe non-numeric traits (e.g.,
corporate structure type, geographical region).
• Quantitative (Numerical) Variables: Expressed in distinct numeric amounts. These are further
divided into:
◦ Discrete Variables: Countable integers with strict step intervals (e.g., daily transaction counts).
◦ Continuous Variables: Infinite fractional values within a measurement range (e.g., quarterly revenue
growth rate).
Mathematical Concept of the Random Variable
A random variable is a real-valued mathematical function mapped from the sample space of a random
experiment to a real number line. A Discrete Random Variable possesses a countable probability mass
function (PMF), whereas a Continuous Random Variable is modeled via an integrating probability
density function (PDF).
Sampling Topologies and Rigorous Execution Techniques
In empirical research, the complete universe of items under study is the Population (N), while the
scientifically extracted subset is the Sample (n).
Probability Sampling Frameworks (Rigorous/Unbiased)
1. Simple Random Sampling (SRS): Every element in the population frame has an identical,
independent probability of selection. Executed via lottery or pseudorandom number generators.
2. Stratified Random Sampling: Used when the population is highly heterogeneous. The population is
split into mutually exclusive, exhaustive subsets called strata based on control variables (e.g., firm
asset sizes). Independent simple random samples are then extracted from each stratum.
3. Systematic Sampling: Elements are drawn at fixed periodic intervals from the population list. The
sampling interval is calculated as k = N / n, starting with a random integer between 1 and k.
4. Cluster Sampling: The population is divided into multi-element, naturally occurring geographic or
administrative units called clusters. A random selection of entire clusters is audited, which lowers data
collection costs over widespread areas.
Bharathiar University Ph.D. Part I Examination • Paper II Page 3
Non-Probability Sampling Frameworks (Subjective/Purposive)
• Convenience Sampling: Selection is based purely on ease of access. This approach limits
generalizability and can introduce selection bias.
• Judgmental / Purposive Sampling: Handpicked by the researcher using expert criteria to focus on
specific, illustrative case profiles.
• Quota Sampling: Non-random stratification where interviewers fill pre-set demographic quotas based
on fixed population characteristics.
• Snowball Sampling: Relies on chain-referrals where initial small groups recruit additional hidden
subjects from their professional network.
3. Mathematical Measures of Central Tendency
Arithmetic Mean (X̄)
The mathematical center of a distribution, representing the sum of all values divided by the sample size. It
acts as the balance point of the dataset.
Continuous Series Step-Deviation Method: X̄ = A + ( (Σ fd) / N ) × c where d = (X - A) / c
Inviolable Mathematical Properties of the Mean:
1. The algebraic sum of individual deviations from the true arithmetic mean is always zero: Σ(X - X̄) = 0.
2. The sum of squared deviations from the arithmetic mean is minimized: Σ(X - X̄)2 < Σ(X - A)2 for any
value A ≠ X̄.
3. Combined Mean Formula: X̄ = (n1X̄1 + n2X̄2) / (n1 + n2)
12
Median (M)
The positional value that divides an ordered dataset into two equal halves. It is highly robust against data
outliers.
Continuous Interpolation Formula: M = L + [ ( (N/2) - cf ) / f ] × c
Where L is the lower limit of the median class, cf is the cumulative frequency of the preceding class, f is
the actual frequency of the median class, and c is the width of the class interval.
4. Mathematical Measures of Dispersion
Mean Deviation (MD)
The average of the absolute deviations of data points around a central measure (Mean or Median),
ignoring algebraic signs.
Bharathiar University Ph.D. Part I Examination • Paper II Page 4
MDX̄ = (Σ f|X - X̄|) / N
Standard Deviation (σ) and Variance (σ2)
The standard deviation is the positive square root of the mean of squared deviations from the arithmetic
mean. It provides a reliable, standard measure of risk and volatility.
Continuous Step-Deviation Formula: σ = √[ (Σ fd2 / N) - (Σ fd / N)2 ] × c
Coefficient of Variation (CV)
A relative measure of dispersion that expresses the standard deviation as a percentage of the mean. It is
used to compare the stability and consistency of different datasets regardless of scale.
CV = (σ / X̄) × 100
Bharathiar University Ph.D. Part I Examination • Paper II Page 5
NUMERICAL APPLICATION BLUEPRINT — UNIT I
Problem: The following continuous frequency table details the monthly sales performance of 50
regional distribution hubs. Calculate the Arithmetic Mean, Standard Deviation, and Coefficient of
Variation.
Sales (in ₹ Lakhs) Number of Hubs (f)
10 - 20 5
20 - 30 12
30 - 40 20
40 - 50 10
50 - 60 3
Comprehensive Solution Matrix: Let Assumed Mean A = 35, Class width c = 10.
Class f Midpoint (X) d = (X-35)/10 fd d2 fd2
10 - 20 5 15 -2 -10 4 20
20 - 30 12 25 -1 -12 1 12
30 - 40 20 35 0 0 0 0
40 - 50 10 45 1 10 1 10
50 - 60 3 55 2 6 4 12
Total N=50 - - Σfd = -6 - Σfd2 = 54
1. Mean Derivation:
X̄ = A + (Σ fd / N) × c = 35 + (-6 / 50) × 10 = 35 - 1.2 = 33.8 ext{ Lakhs}
2. Standard Deviation Derivation:
σ = √[ (54 / 50) - (-6 / 50)2 ] × 10 = √[ 1.08 - 0.0144 ] × 10 = √[ 1.0656 ] × 10 = 1.03227 × 10 = 10.32 ext{
Lakhs}
3. Relative Volatility (CV):
CV = (10.32 / 33.8) × 100 = 30.53\%
Bharathiar University Ph.D. Part I Examination • Paper II Page 6
UNIT II: CORRELATION & REGRESSION ANALYSIS
1. Advanced Correlation Architecture
Correlation analysis explores the linear relationships between variables, measuring the strength and
direction of their association without assuming a causal link.
Karl Pearson Product-Moment Coefficient of Correlation (r)
Measures the linear relationship between two continuous variables. The value of r is constrained within
the range [-1, 1].
Mathematical Formulation: r = [ nΣXY - (ΣX)(ΣY) ] / √[ [nΣX2 - (ΣX)2][nΣY2 - (ΣY)2] ]
Spearman’s Rank Correlation Coefficient (R)
Used for ordinal rankings or when continuous data violates normal distribution assumptions. It measures
monotonic relationships.
Tied/Repeated Ranks Formula: R = 1 - [ 6 ( ΣD2 + Σ (mi3 - mi) / 12 ) ] / [ n(n2 - 1) ]
Where D = R1 - R2 is the difference between paired ranks, and mi is the number of observations tied within
the i-th tie group.
Partial and Multiple Correlation
• Multiple Correlation (R ): Measures the joint linear relationship between a primary dependent
1.23
variable (X1) and a combination of multiple independent predictors (X2, X3).
R1.23 = √[ (r122 + r132 - 2r12r13r23) / (1 - r232) ]
12.3): Isolates the clean linear relationship between X1 and X2 by statistically
• Partial Correlation (r
removing the confounding effects of a shared variable X3.
r12.3 = [ r12 - r13r23 ] / √[ (1 - r132)(1 - r232) ]
Autocorrelation & Durbin-Watson Testing Framework
Autocorrelation occurs when errors in a time series are correlated over time, violating the assumption of
independent errors in classical linear regression.
Durbin-Watson Test Statistic: d = [ Σt=2T (et - et-1)2 ] / [ Σt=1T et2 ]
Bharathiar University Ph.D. Part I Examination • Paper II Page 7
The statistic ranges from 0 to 4. A value of d ≈ 2 indicates no autocorrelation. Values near 0 indicate
positive autocorrelation, while values near 4 point to negative autocorrelation.
2. Predictive Regression Models
Simple Linear Regression Framework
Models the linear relationship between an independent predictor (X) and a dependent criterion variable (Y)
using the Ordinary Least Squares (OLS) method.
Y = a + bX
The parameters are computed by solving the following normal equations:
ΣY = na + bΣX
ΣXY = aΣX + bΣX2
Multiple Linear Regression Architecture
Extends simple regression by including two or more independent predictors to explain variation in the
dependent variable.
Y = a + b1X1 + b2X2 + ... + bkXk + ε
Use of Dummy Variables (Categorical Coding)
Dummy variables use binary coding (0 or 1) to incorporate qualitative attributes into regression models. To
avoid perfect multicollinearity, known as the Dummy Variable Trap, a categorical attribute with m
categories must be represented by exactly m - 1 dummy variables.
Bharathiar University Ph.D. Part I Examination • Paper II Page 8
NUMERICAL APPLICATION BLUEPRINT — UNIT II
Problem: Track the relationship between marketing spend (X in ₹ Thousands) and quarterly
revenue performance (Y in ₹ Lakhs) across 6 regional corporate offices. Compute Karl Pearson’s r
and derive the OLS regression equation of Y on X.
Office Index Spend (X) Revenue (Y) X2 Y2 XY
1 2 4 4 16 8
2 4 7 16 49 28
3 6 9 36 81 54
4 8 12 64 144 96
5 10 15 100 225 150
6 12 17 144 289 204
Total ΣX=42 ΣY=64 ΣX2=364 ΣY2=724 ΣXY=540
1. Pearson Correlation Coefficient Derivation:
r = [ 6(540) - (42)(64) ] / √[ [6(364) - 422][6(724) - 642] ]
r = [ 3240 - 2688 ] / √[ [2184 - 1764][4344 - 4096] ] = 552 / √[ 420 × 248 ] = 552 / 322.74 = 0.971
2. OLS Regression Parameters Derivation:
Slope b = [ nΣXY - (ΣX)(ΣY) ] / [ nΣX2 - (ΣX)2 ] = 552 / 420 = 1.314
Means: X̄ = 42/6 = 7; Ȳ = 64/6 = 10.667
Intercept a = Ȳ - bX̄ = 10.667 - (1.314 × 7) = 1.469
Final Regression Equation: Ŷ = 1.469 + 1.314X
Interpretation: For every additional unit increase in marketing spend, revenue increases by 1.314
units.
Bharathiar University Ph.D. Part I Examination • Paper II Page 9
UNIT III: HYPOTHESIS TESTING (LARGE & SMALL
SAMPLES)
1. Foundations of Statistical Inference
Hypothesis testing is a structured inferential framework used to assess whether sample statistics provide
sufficient evidence to reject a specified null hypothesis about a population parameter.
Key Concepts in Statistical Inference:
• Null Hypothesis (H ): A statement of no effect, no difference, or status quo. It is assumed true
0
until proven otherwise.
• Alternative Hypothesis (H ): The statement that contradicts H , representing the effect or
1 0
relationship the researcher aims to demonstrate.
• Type I Error (α): Rejecting the null hypothesis when it is true. This probability is the level of
significance.
• Type II Error (β): Failing to reject the null hypothesis when it is false. The power of the test is
defined as 1 - β.
2. Z-Test Frameworks (Large Samples: n ≥ 30)
Z-tests rely on the standard normal distribution, supported by the Central Limit Theorem for larger sample
sizes.
Testing a Single Population Mean
Z = (X̄ - μ0) / (σ / √n)
Testing the Difference Between Two Independent Means
Z = (X̄1 - X̄2) / √[ (σ12 / n1) + (σ22 / n2) ]
Testing a Single Population Proportion
Z = (p - P0) / √[ (P0Q0) / n ] where Q0 = 1 - P0
Bharathiar University Ph.D. Part I Examination • Paper II Page 10
Testing the Difference Between Two Proportions
Z = (p1 - p2) / √[ P̂Q̂ ( (1/n1) + (1/n2) ) ]
Where the pooled proportion estimator is: P̂ = (n1p1 + n2p2) / (n1 + n2) and Q̂ = 1 - P̂.
3. Student's t-Test Frameworks (Small Samples: n < 30)
Used when sample sizes are small and the population standard deviation (σ) is unknown, assuming the
underlying population is normally distributed.
t-Test for a Single Mean
t = (X̄ - μ0) / (s / √n) with ν = n - 1 degrees of freedom.
Where s is the sample standard deviation: s = √[ Σ(X - X̄)2 / (n - 1) ].
Independent Two-Sample t-Test
t = (X̄1 - X̄2) / (sp √[ (1/n1) + (1/n2) ]) with ν = n1 + n2 - 2 df.
Where the pooled standard deviation is: sp = √[ [ (n1-1)s12 + (n2-1)s22 ] / (n1 + n2 - 2) ].
Paired t-Test for Matched Observations
t = d̄ / (sd / √n) with ν = n - 1 df.
Where d represents the differences between paired scores, d̄ is the mean difference, and sd is the
standard deviation of those differences.
Bharathiar University Ph.D. Part I Examination • Paper II Page 11
NUMERICAL APPLICATION BLUEPRINT — UNIT III
Problem: A logistics provider claims that a route optimization update reduces package delivery
times below the previous average of 42 minutes. A random sample of n = 25 deliveries records a
mean time of X̄ = 39.8 minutes with a sample standard deviation of s = 4.1 minutes. Test this claim at
the 5% level of significance.
Step-by-Step Solution:
• 1. Hypotheses: H : μ ≥ 42 minutes; H : μ < 42 minutes (Left-Tailed Test).
0 1
• 2. Test Selection: Small sample (n = 25 < 30), population variance unknown. Use the single-
sample t-test.
• 3. Critical Value: For degrees of freedom ν = 25 - 1 = 24 at a one-tailed α = 0.05, the critical
value is -1.711. Reject H0 if tcalc ≤ -1.711.
• 4. Test Statistic Computation:
t = (39.8 - 42) / (4.1 / √25) = -2.2 / (4.1 / 5) = -2.2 / 0.82 = -2.683
• 5. Conclusion: Since t
calc = -2.683 < -1.711, it falls into the rejection region. We reject the null
hypothesis. The data supports the claim that the update reduces delivery times.
Bharathiar University Ph.D. Part I Examination • Paper II Page 12
UNIT IV: CHI-SQUARE, F-TEST, AND ANOVA
1. Chi-Square (χ2) Distribution Testing
The Chi-Square test is a non-parametric method used to evaluate categorical frequency data based on
differences between observed and expected frequencies.
χ2 = Σ [ (O - E)2 / E ]
Independence of Attributes Test
Evaluates whether two categorical variables are independent within a contingency table.
Expected Cell Frequency: Eij = (Row Total × Column Total) / N
The degrees of freedom are calculated as: ν = (r - 1)(c - 1).
Methodological Constraints
The total frequency N should ideally exceed 50. If any expected cell frequency in a 2 × 2 table is less than
5, apply Yates' Correction for Continuity: χ2Yates = Σ [ (|O - E| - 0.5)2 / E ].
2. Variance Comparison: The F-Test
The F-test is used to compare the variances of two independent populations to determine if they differ
significantly.
F = s12 / s22 where s12 > s22
The degrees of freedom are ν1 = n1 - 1 and ν2 = n2 - 1.
3. Analysis of Variance (ANOVA) Frameworks
ANOVA splits total dataset variance into distinct components to test for significant differences across
multiple group means.
One-Way ANOVA Classification
Analyzes a single factor across multiple independent groups, partitioning total variation into between-
group and within-group components.
• Correction Factor: CF = T2 / N
Bharathiar University Ph.D. Part I Examination • Paper II Page 13
• Total Sum of Squares: SST = ΣΣX 2 - CF
ij
• Sum of Squares Between Groups: SSB = Σ(T 2 / n ) - CF
j j
• Sum of Squares Within Groups (Error): SSW = SST - SSB
Source of Sum of Squares Degrees of Mean Square Calculated F-
Variance (SS) Freedom (ν) (MS) Ratio
MSB = SSB / (k -
Between Groups SSB k-1 F = MSB / MSW
1)
Within Groups MSW = SSW / (N
SSW N-k
(Error) - k)
Total Variation SST N-1
Bharathiar University Ph.D. Part I Examination • Paper II Page 14
NUMERICAL APPLICATION BLUEPRINT — UNIT IV
Problem: Evaluate productivity scores across three distinct workplace training modules (A, B, C)
using a sample of 12 employees. Perform a One-Way ANOVA at the 5% level of significance.
Module A Module B Module C
5 9 12
7 10 11
4 8 9
8 9 12
Step-by-Step Solution:
• Group Totals: T = 24, T = 36, T = 44. Grand Total T = 104. Total observations N = 12, groups k
A B C
= 3.
• Sum of all squared elements: ΣΣX2 = 52+72+...+122 = 970.
• Correction Factor: CF = 1042 / 12 = 901.33.
• Total Variation: SST = 970 - 901.33 = 68.67.
• Between-Group Variance: SSB = [ (242/4) + (362/4) + (442/4) ] - 901.33 = [ 144 + 324 + 484 ] -
901.33 = 50.67.
• Error Variance: SSW = SST - SSB = 68.67 - 50.67 = 18.00.
Source SS ν MS Calculated F
Between Treatments 50.67 2 25.335 12.67
Error Within Groups 18.00 9 2.000 -
Inference: The critical table value for F(2,9) is 4.26. Since 12.67 > 4.26, we reject the null
hypothesis. There is a statistically significant difference in productivity outcomes between the
training modules.
Bharathiar University Ph.D. Part I Examination • Paper II Page 15
UNIT V: NON-PARAMETRIC STATISTICS &
MULTIVARIATE ANALYSIS
1. Non-Parametric Frameworks
Non-parametric methods are distribution-free alternatives used when data scales are ordinal or nominal, or
when the assumption of a normally distributed population is violated.
The Sign Test
A simple method used to test hypotheses about a population median by analyzing the direction (plus or
minus signs) of deviations from the hypothesized value.
Runs Test for Randomness
Evaluates whether a sequence of binary events occurs in a random order by counting the total number of
runs (r).
Large Sample Expected Mean: μr = [ (2n1n2) / (n1 + n2) ] + 1
Mann-Whitney U Test
The non-parametric counterpart to the independent two-sample t-test. It converts data values into
continuous pooled ranks to test whether two independent samples come from the same distribution.
Kruskal-Wallis H Test
The non-parametric alternative to the One-Way ANOVA. It ranks pooled data across multiple groups to
determine if they share a common median.
H = [ 12 / (N(N+1)) ] × Σ [ Rj2 / nj ] - 3(N + 1)
The statistic follows a Chi-Square distribution with ν = k - 1 degrees of freedom.
2. Time Series Modeling
A time series tracks data points over uniform time intervals to understand patterns and forecast future
values based on past performance.
Bharathiar University Ph.D. Part I Examination • Paper II Page 16
Structural Decomposition Components:
• Secular Trend (T): Consistent, long-term smooth directional movement over an extended period.
• Seasonal Variation (S): Periodic patterns that repeat within a single annual cycle (e.g., holiday
retail surges).
• Cyclical Variation (C): Long-term wave-like oscillations around the trend line, driven by broader
business cycles.
• Irregular Residuals (I): Unpredictable, random shocks or noise (e.g., policy shifts or natural
disruptions).
Mathematical Formulation Models
• Additive Specification: Y = T + S + C + I
t t t t t
• Multiplicative Specification: Y = T × S × C × I
t t t t t
Method of Least Squares for Linear Trends
Fits a linear trend line by minimizing the sum of squared errors. When the time variable is centered so that
ΣX = 0, the parameters are simplified as:
a = ΣY / n and b = ΣXY / ΣX2
3. Overview of Multivariate Analysis Techniques
• Principal Component Analysis (PCA): A data reduction technique that projects correlated variables
into a smaller set of uncorrelated factors called Principal Components, retaining the maximum possible
variance.
• Factor Analysis: Identifies underlying latent variables or factors from groups of highly correlated
observed measures to uncover internal structure.
• Discriminant Analysis: Predicts group membership for a categorical dependent variable based on a
linear combination of continuous predictor variables.
• Cluster Analysis: An unsupervised classification method that groups individuals or items into clusters
to maximize within-group similarity and between-group differences.
• Path Analysis: An extension of multiple regression that maps direct and indirect causal pathways
across an integrated network of variables.
Bharathiar University Ph.D. Part I Examination • Paper II Page 17
NUMERICAL APPLICATION BLUEPRINT — UNIT V
Problem: Fit a Least Squares linear trend line to the annual revenue data of a regional enterprise
(Y in ₹ Crores) over a 5-year period, and project the estimated revenue for the upcoming year
(2027).
Year Label 2021 2022 2023 2024 2025
Revenue (Y) 12 14 19 23 27
Step-by-Step Solution: Centering the timeline around the midpoint year (2023) ensures ΣX = 0.
Year Y Time Code (X) XY X2
2021 12 -2 -24 4
2022 14 -1 -14 1
2023 19 0 0 0
2024 23 1 23 1
2025 27 2 54 4
Total ΣY = 95 ΣX = 0 ΣXY = 39 ΣX2 = 10
Parameter Calculation:
a = ΣY / n = 95 / 5 = 19
b = ΣXY / ΣX2 = 39 / 10 = 3.9
Trend Line Equation: Yt = 19 + 3.9X (with origin set at 2023).
Forecast for 2027: The time code for 2027 is X = 2027 - 2023 = 4.
Y2027 = 19 + 3.9(4) = 19 + 15.6 = 34.6 ext{ Crores}.
Bharathiar University Ph.D. Part I Examination • Paper II Page 18