STEP III Statistics: Key Concepts
STEP III Statistics: Key Concepts
Differentiating the probability generating function G_X(t) with respect to t and then substituting t = 1 gives the expected value E(X). This technique exploits the fact that the differentiation details the weights of probabilities at successively higher powers of t, which correspond to the potential outcomes of the random variable .
The expectation is derived by considering the infinite series formed by E(X) = p + 2qp + 3q^2p + ..., where each term is weighted by the number of trials and previous outcomes' probabilities. Recognizing the pattern in the derivatives of a geometric series allows us to simplify E(X) to 1/p, which matches the expected outcome for the geometric distribution .
If two random variables X and Y are independent, then their covariance Cov(X, Y) equals 0. This is because independence implies no linear relationship between the variables, which results in zero covariance .
As the sample size n increases, the sample variance Var(¯X) decreases, because it is inversely proportional to n (σ²/n). This reflects the increasing precision of the sample mean as an estimate of the population mean with larger samples .
Calculating the variance using a probability generating function involves differentiating it twice to obtain E(X^2) and once to get E(X). The variance is then given by subtracting the square of the first derivative from the second derivative evaluated at t=1, using algebraic manipulation and properties of power series .
The product moment correlation coefficient, ρ, quantifies the strength and direction of a linear relationship between two variables. It is computed as the covariance of the variables divided by the product of their standard deviations, controlling for scale and allowing ρ to be bounded between -1 and 1 .
To calculate variance using moment generating functions, differentiate the M.G.F. with respect to t twice and evaluate at t = 0. This gives E(X^2). The variance Var(X) is then found by subtracting the square of the first derivative evaluated at t=0, i.e., M'(0)^2, from the second derivative M''(0).
The Central Limit Theorem states that irrespective of the population distribution, the distribution of the sample mean approaches a normal distribution as the sample size n becomes large. This is due to the averaging effect and the convergence in distribution towards a Gaussian distribution .
For dependent random variables X and Y, the variance of a linear combination Var(aX ± bY) includes an additional term 2abCov(X, Y) to account for the covariance between X and Y. This term adjusts the variance calculation when considering potential correlation or dependence between the variables .
The probability density function (P.D.F.) of a random variable can be derived by differentiating its cumulative distribution function (C.D.F.) with respect to y. This conversion is grounded in the Fundamental Theorem of Calculus, where the derivative of an integral (C.D.F.) with respect to its upper limit gives the integrand (P.D.F.).