Understanding the Wishart Distribution
Understanding the Wishart Distribution
The Wishart distribution is significant in multivariate statistics as it extends the concept of the chi-square distribution to the multivariate setting, dealing with symmetric positive definite matrices. It is often used as a conjugate prior for the inverse of normal covariance matrices in Bayesian computation and when symmetric positive definite matrices are the primary focus, such as in diffusion tensor studies . The Wishart distribution is related to Hotelling's T² statistic because the statistic is derived from data that can be considered under the Wishart distribution. Specifically, if S follows a Wishart distribution, T² measures the deviations of a sample mean from a population mean, scaled by the inverse sampled covariance matrix .
The inversion lemma, which provides a method to invert partitioned matrices, aids in demonstrating Theorem 3 by facilitating the computation of inverses required in deriving the distribution properties of a'Σ⁻¹a/a'M⁻¹a for M ∼ Wp(n, Σ). It allows the analyst to express complicated matrix inversions in terms of simpler sub-matrix inversions, thus making it feasible to derive and prove distributional results for expressions involving inverted Wishart-distributed matrices .
The matrix additivity property of the Wishart distribution, where the sum of independent Wishart-distributed matrices remains Wishart-distributed, is significant in statistical applications because it allows for the accumulation of information across independent samples. This property is essential in multivariate analysis and computational methods, such as Monte Carlo simulations or Bayesian updating, where sample matrices are aggregated to form a single, larger Wishart-distributed matrix, permitting straightforward statistical inference .
The independence between the sample mean vector and the sample covariance matrix for multivariate normal samples is established through the decomposition of the total sum of squares. Specifically, by showing that the matrix of de-meaned vectors 31i=1 (2Xiden - X̄)(2Xiden - X̄)' splits into independent components Pn i=2 YiY'i, it demonstrates the independence between Y1, representing the sample mean, and Yi for i ≥ 2, representing the components contributing to the sample covariance matrix .
The corollary states that if M follows a Wishart distribution Wp(n, Σ) and a is a non-zero vector in Rp such that a'Σa ≠ 0, then the statistic a'Ma/a'Σa follows a chi-square distribution with n degrees of freedom. This connection highlights the role of the Wishart distribution as a multivariate generalization of the chi-square distribution .
The law of large numbers for matrices following the Wishart distribution implies that as the number of degrees of freedom n increases, the scaled Wishart matrix Mn/n converges in probability to the population covariance matrix Σ. This result ensures that for large samples, the sample covariance matrix provides a good approximation of the true covariance matrix .
The density function of the Wishart distribution is particularly used in Bayesian computation, where it serves as a conjugate prior for the inverse of normal covariance matrices, and in studies focusing on symmetric positive definite matrices, such as diffusion tensor imaging. Outside these contexts, the exact form of the density function is rarely used due to its complexity and because analyses involving the Wishart distribution typically focus on its properties rather than its explicit density .
Theorem 6 implies that for i.i.d. samples from a multivariate normal distribution, the sample mean vector and the sample covariance matrix are independent. Additionally, it shows that the scaled sample mean follows a multivariate normal distribution Np(0, Σ), and the adjusted sample covariance (n-1)S follows a Wishart distribution Wp(n-1, Σ). This independence is crucial for statistical inference and simplifies multivariate hypothesis testing .
Hotelling's T² statistic functions as a multivariate analogue to the Student's t-statistic by testing hypotheses about the mean vector of a multivariate normal distribution when the covariance matrix is unknown. It is used to determine if there is a significant difference between the sample means of different groups, similar to how the t-test is used for univariate analysis .
The condition m > p-1 is necessary because it ensures the invertibility of the sample covariance matrix S, which is required for calculating Hotelling's T² statistic. Without this condition, the covariance matrix would be singular, and its inverse would not exist, making it impossible to compute the statistic and derive its relationship to the F-distribution as stated in Theorem 5 .