Joint and Marginal Distributions in Statistics
Joint and Marginal Distributions in Statistics
prob models that involve more than one rv → multivariate models
•
∈
• : constants
•
∈
∈
∈
mpmf of
Mathematical Statistics 4-2
∞ ∞
•
∞ ∞
∞
∞
mpdf of
∞ ∞
•
∞ ∞
(ⅱ) Consider ≥ 1
≥
1
≡
Mathematical Statistics 4-3
Consider ≥
≥
1
≥
• joint cdf
≤ ≤
continuous case : ∞ ∞
[Def. 4.2.1] discrete bivariate random vector with jpmf
and mpmf
the conditional pmf of given that
for any st
the conditional pmf of given that
for any st
Mathematical Statistics 4-4
.
•
For
For
Mathematical Statistics 4-5
∴
≤
→ ∣ ∼ two parameter exponential dist.,
where the location parameter
the scale parameter
∞
→
∞
⇒
exp exp
exp
∴ ∼ ⇒ Thm. 4.2.14
⋯
∴ ∼
By Thm. 4.2.12,
∴ ∼
: the Jacobian of the transformation
Mathematical Statistics 4-10
Let and
→ →
∣∣
∴ the jpdf of and :
⋅
∴ the mpdf of :
let → ie
∴ ∼ ∵
Mathematical Statistics 4-11
∣ ∣
∴ ⋅∣ ∣
exp ⋅
⋅
a ft of alone a ft of alone
⇒ and are indep. by Lemma 4.2.7
Also, ∼
∼
(Note) Sums and differences of indep. normal r.v.s are indep. normal
r.v.s regardless of the means of and , so long as
.
(p.f)
Assume and are continuous r.v.s.
Mathematical Statistics 4-12
∈ ∈
→ by p
∈ ⋅ ∈
∴ By Lemma 4.2.7, and are indep.
∞
use exponential
∞ ∞ → Cauchy dist
wrt wrt
wrt
∵ ∼
∵ ∼
∞
∞
use
geometric dist
∼
⇒ Three-stage hierarchy in Ex. 4.4.5 is equivalent to the two-stage
hierarchy ∼
∼
Mathematical Statistics 4-16
∴
∣
∴ Similarly can be calculated.
Taking expectation yields
⋆
Also,
∣
(ⅰ)
∼
∴
(ⅱ)
≤ ≤
∴
(ⅲ)
⇒
Mathematical Statistics 4-20
[Thm. 4.5.6] and : any two r.v.s, and : any two constants
⇒
If and are indep., then .
(p.f) Omit.
∼ , ∼ : indep.
Let
∵
let ∣ ∣
→
∴
∴ ← much stronger relationship than Ex. 4.5.4
∼ , ∼ : indep.
Let
→
Since ∼ and since and
are indep., .
∴
→
exp
∞ ∞
Nice properties:
a ∼
b ∼
c
d ∼
where : constant
Mathematical Statistics 4-23
• For any ⊂ ,
∈
∈
if discrete random vector
⋯ ⋯ ⋯
⋯
if conti random vector
•
∈
∞ ∞
∞
⋯
∞
• ⋯
⋯ ∈
∞ ∞
⋯ ∞ ∞
⋯
⋯
• ⋯ ∣ ⋯
⋯
(ⅰ)
⋯
Mathematical Statistics 4-24
(ⅱ)
∞ ∞
(ⅲ) ∞ ∞
∞ ∞
∞ ∞
∣
Mathematical Statistics 4-25
[Def. 4.6.2] positive integer ⋯ ≤ ≤
,
where nonnegative integer and
• multinomial coefficient
⋯
the number of ways that objects can be divided
into groups with in the first group, in the
second group, …, and in the th group
⋯
⋯ ∈
⋯
⋯ ∈
⋯ ⋅
⋯
∵
⋯ ∈
⋯ ⋯
≡ ∵ ⋯
⋯
∴ ∼
∴ Similarly ∼ ⋯
⋯ ⋯
∼ dist
with trials and cell probs ⋯
: one-dimensional ∀
→ ⋯ are mutually indep. r.v.s
[Corollary 4.6.9] ⋯ are mutually indep. r.v.s with mgfs
⋯ and ⋯ are fixed constants.
Let ⋯ . Then the mgf of is
∑
⋯
Mathematical Statistics 4-28
∑
(p.f)
∑
⋯
∑
⋯
∼
∑
exp ⋯ exp
∑
exp ∑
∴ ∼
one to one on ⋯ ⋅⋅⋅
⋮ ⋮ ⋮ ⋮
where
⋮ ⋮ ⋮ ⋮
⋅⋅⋅
⇒
⇒
∼
∴ are mutually indep.
Mathematical Statistics 4-30
4.7 Inequality
[Lemma 4.7.1]
⇒ ≥ with equality iff
(p.f) Omit.
≤
→
≤
⇒
≥
(ⅱ) is convex if ″ ≥ ∀
is concave if ″ ≤ ∀
(ⅲ) If is concave, then ≤ .
log
(ⅰ) log
log log ≤ log log
∴ ≤
log
≥ log log
(ⅱ) log log log
log
∴ log ≥ log ⇒ ≤
∵
⋅ ∞ ⋅ ∞
≥ ⋅ ∞ ⋅ ∞
∵ is nondecreasing and
⋅ ∞ ≤
The method of transformations can simplify the study of Poisson distributed variables by reducing complex problems to simpler forms. For instance, when two independent Poisson variables X~Poisson(θ) and Y~Poisson(λ) are added, the resulting sum Z = X + Y is also Poisson distributed with parameter θ + λ . This property allows for the study of compound Poisson processes through relatively straightforward transformations, such as using moment-generating functions or leveraging properties of sums to easily infer about aggregates of Poisson-distributed events, thus simplifying analysis and solution derivation.
Holder's Inequality generalizes Cauchy-Schwarz Inequality by extending it to transformations and variables with different p-norms. While Cauchy-Schwarz is a specific case for p = q = 2, stating that |E[XY]| ≤ (E[X^2])^(1/2)(E[Y^2])^(1/2), Holder's Inequality covers a broader range by stating |E[XY]| <= (E[|X|^p])^(1/p) * (E[|Y|^q])^(1/q), where 1/p + 1/q = 1. This allows it to handle a wider range of mathematical situations involving sums of products, facilitating analysis under more general conditions.
Understanding the conditional distribution of random variables is crucial because it provides insights into how one variable behaves given specific information about another variable. This knowledge is important for predictive modeling and decision making in various fields such as finance, machine learning, and insurance where multi-variate dependencies must be modeled accurately . For instance, the conditional probability density function f(Y|X=x) helps in determining the probability distribution of Y given that X takes on a certain value, which can significantly refine predictions and analyses.
Mixture distributions are utilized in statistical modeling to represent populations that exhibit heterogeneity or are composed of distinct sub-populations, each having its own probability distribution. For example, combining a binomial distribution with a Poisson distribution yields a compound distribution used to model data where the mean of the binomial varies with a Poisson distribution . In practical terms, mixture models can be applied in scenarios like market segmentation, where the overall distribution of consumer preferences is seen as a composition of different underlying preference distributions.
Calculating the variance of composite random variables is fundamental to understanding the variability and reliability of statistical predictions. For two independent random variables X and Y, the variance of their sum is additive, Var(X+Y) = Var(X) + Var(Y). This additive property allows statisticians to account for sources of variation and gauge the uncertainty in composite measures, such as the collective risk in portfolio returns or aggregated error in sensor data fusion, which is integral to precision in forecasting, risk management, and decision-making models.
When two independent normal random variables X and Y are considered, their sum U = X + Y and difference V = X - Y are also independent if X and Y have the same variance . This is a remarkable property of normal distributions which does not extend to other distributions. Mathematically, this implies that the transformation maintaining the joint normality leads to a diagonal covariance matrix for (U, V), which simplifies the joint analysis, eases the calculation of joint probabilities, and enhances decomposition into independent components for further statistical inferences.
The correlation coefficient, denoted as ρ, measures the linear relationship between two random variables X and Y. It provides information about the direction and strength of the relationship. A positive correlation indicates that as one variable increases, the other tends to increase, whereas a negative correlation suggests that as one variable increases, the other tends to decrease. Values close to -1 or 1 indicate a strong linear relationship, while values near 0 indicate a weak linear relationship . However, the correlation does not imply causation and does not measure non-linear relationships.
If X and Y are independent, the joint PDF f(x, y) can be expressed as the product of their marginal PDFs, i.e., f(x, y) = f_X(x)f_Y(y) for all values of x and y . This means that knowing the value of one variable does not provide any information about the other. Independently, their conditional PDFs are also equal to the marginal PDFs, e.g., f(y|x) = f_Y(y). This property is essential for simplification in statistical computations involving independent random variables.
Jensen's Inequality states that for a convex function g, E[g(X)] ≥ g(E[X]). This inequality shows that the expected value of a convex function of a random variable is at least the convex function of the expected value of that variable. It provides essential insights into the behavior of random variables and expectations under convex transformations, indicating that convex transformations tend to amplify variability, while in a practical scenario, it can be used to derive bounds or estimate the likelihood of certain events when dealing with convex loss functions or utility functions.
Bivariate transformations are significant in statistical analysis because they allow the study of relationships between two random variables by transforming them into a new set of variables that might be easier to analyze or interpret . For instance, transformation can help in finding the distribution of a sum or a product of variables, which is a common requirement in many statistical problems. The probability distribution of the transformed variables is fully determined by the original joint distribution, enabling new insights and simplifications in handling complex data forms.