0% found this document useful (0 votes)
7 views7 pages

Data Science Statistics Formulas

This document provides a collection of important formulas for exam preparation, including various probability distributions such as Bernoulli, Binomial, Geometric, Negative Binomial, Poisson, and Normal distributions. It specifies how these formulas will be made available during the exam, either in printed form for face-to-face exams or digitally for online exams. Additionally, it warns that the collection does not include calculation rules or other mathematical laws.

Uploaded by

jolina256
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views7 pages

Data Science Statistics Formulas

This document provides a collection of important formulas for exam preparation, including various probability distributions such as Bernoulli, Binomial, Geometric, Negative Binomial, Poisson, and Normal distributions. It specifies how these formulas will be made available during the exam, either in printed form for face-to-face exams or digitally for online exams. Additionally, it warns that the collection does not include calculation rules or other mathematical laws.

Uploaded by

jolina256
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PRÜFUNGSAMT

[Link]

DLBDSSPDS01-01
Formula Collection + Tabs

Dear Students,
The following collection contains some formulas that are important for the exam.
Please note that this PDF document is only for exam preparation. During the exam, this collection of formulas and tables
will also be made available to you:
− Face-to-face exam: The formula collection and the tables will be printed out for you as part of the exam and must be
handed in again together with the remaining exam pages.
− Online exam: The formula collection and the tables are stored for you on the platform of the proctor provider and you
can download them there before activating the exam.

DLBDSSPDS01_Formula Collection & Tables for [Link]

Warning: This formula collection only contains the formulas shown in the script. Calculation rules or other
important mathematical laws are not listed here.

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]

Discreet distributions
Bernoulli-distribution. For a Binomial(𝑛𝑛, 𝑝𝑝)-distributed random variable X applies:
1 − 𝑝𝑝, 𝑘𝑘 = 0,
Probability function 𝑓𝑓𝑋𝑋 (𝑘𝑘) = 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)1−𝑘𝑘 = �
𝑝𝑝, 𝑘𝑘 = 1,
Expected value 𝐸𝐸[𝑋𝑋] = 𝑝𝑝
Variance Var[𝑋𝑋] = 𝑝𝑝(1 − 𝑝𝑝)
𝑋𝑋 − 𝜇𝜇 3 1 − 2𝑝𝑝
Skewness 𝐸𝐸 �� � �=
𝜎𝜎 �𝑝𝑝(1 − 𝑝𝑝)
𝑋𝑋 − 𝜇𝜇 4 1 − 6𝑝𝑝(1 − 𝑝𝑝)
Kurtosis 𝐸𝐸 �� � �−3 =
𝜎𝜎 𝑝𝑝(1 − 𝑝𝑝)
Moment generating
𝑀𝑀𝑋𝑋 (𝑡𝑡) = 1 − 𝑝𝑝 + 𝑝𝑝e𝑡𝑡 , 𝑡𝑡 ∈ ℝ
functions (mgf)

Binomial distribution. For a Binomial(𝑛𝑛, 𝑝𝑝)- distributed random variable X applies:


𝑛𝑛
Probability function 𝑓𝑓𝑋𝑋 (𝑘𝑘) = � � 𝑝𝑝𝑘𝑘 (1 − 𝑝𝑝)𝑛𝑛−𝑘𝑘 , 𝑘𝑘 ∈ {0, 1, … , 𝑛𝑛}
𝑘𝑘
Expected value 𝐸𝐸[𝑋𝑋] = 𝑛𝑛𝑛𝑛
Variance Var[𝑋𝑋] = 𝑛𝑛𝑛𝑛(1 − 𝑝𝑝)
𝑋𝑋 − 𝜇𝜇 3 1 − 2𝑝𝑝
Skewness 𝐸𝐸 �� � �=
𝜎𝜎 �𝑛𝑛𝑛𝑛(1 − 𝑝𝑝)
𝑋𝑋 − 𝜇𝜇 4 1 − 6𝑝𝑝(1 − 𝑝𝑝)
Kurtosis 𝐸𝐸 �� � �−3 =
𝜎𝜎 𝑛𝑛𝑛𝑛(1 − 𝑝𝑝)
Moment generating
𝑀𝑀𝑋𝑋 (𝑡𝑡) = (1 − 𝑝𝑝 + 𝑝𝑝e𝑡𝑡 )𝑛𝑛 , 𝑡𝑡 ∈ ℝ
functions (mgf)

Geometric distribution. For a Geometric(𝑝𝑝)- distributed random variable X applies:


Probability function 𝑓𝑓𝑋𝑋 (𝑘𝑘) = 𝑝𝑝(1 − 𝑝𝑝)𝑘𝑘 , 𝑘𝑘 ∈ ℕ0
1 − 𝑝𝑝
Expected value 𝐸𝐸[𝑋𝑋] =
𝑝𝑝
1 − 𝑝𝑝
Variance Var[𝑋𝑋] = 2
𝑝𝑝
𝑋𝑋 − 𝜇𝜇 3 2 − 𝑝𝑝
Skewness 𝐸𝐸 �� � �=
𝜎𝜎 �1 − 𝑝𝑝
𝑋𝑋 − 𝜇𝜇 4 𝑝𝑝2
Kurtosis 𝐸𝐸 �� � �−3 = 6+
𝜎𝜎 1 − 𝑝𝑝
𝑝𝑝
Moment generating , 𝑡𝑡 < − ln(1 − 𝑝𝑝)
𝑀𝑀𝑋𝑋 (𝑡𝑡) = �1 − (1 − 𝑝𝑝)e𝑡𝑡
functions (mgf)
∞, 𝑡𝑡 ≥ − ln(1 − 𝑝𝑝)

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]

Negative Binomial distribution. For a NegBinomial(𝑛𝑛, 𝑝𝑝)- distributed


random variable X applies:
𝑛𝑛 + 𝑘𝑘 − 1 𝑛𝑛
Probability function 𝑓𝑓𝑋𝑋 (𝑘𝑘) = � � 𝑝𝑝 (1 − 𝑝𝑝)𝑘𝑘 , 𝑘𝑘 ∈ ℕ0
𝑛𝑛 − 1
1 − 𝑝𝑝
Expected value 𝐸𝐸[𝑋𝑋] = 𝑛𝑛
𝑝𝑝
1 − 𝑝𝑝
Variance Var[𝑋𝑋] = 𝑛𝑛 2
𝑝𝑝
𝑋𝑋 − 𝜇𝜇 3 2 − 𝑝𝑝
Skewness 𝐸𝐸 �� � �=
𝜎𝜎 �𝑛𝑛(1 − 𝑝𝑝)
𝑋𝑋 − 𝜇𝜇 4 6 𝑝𝑝2
Kurtosis 𝐸𝐸 �� � �−3 = +
𝜎𝜎 𝑛𝑛 𝑛𝑛(1 − 𝑝𝑝)
𝑝𝑝 𝑛𝑛
Moment generating � � , 𝑡𝑡 < − ln(1 − 𝑝𝑝)
𝑀𝑀𝑋𝑋 (𝑡𝑡) = � 1 − (1 − 𝑝𝑝)e𝑡𝑡
functions (mgf)
∞, 𝑡𝑡 ≥ − ln(1 − 𝑝𝑝)

Poisson-Distribution. For a Poisson(𝜆𝜆)- distributed random variable X applies:


𝜆𝜆𝑘𝑘 −𝜆𝜆
Probability function 𝑓𝑓𝑋𝑋 (𝑘𝑘) = 𝑒𝑒 , 𝑘𝑘 ∈ ℕ0
𝑘𝑘!
Expected value 𝐸𝐸[𝑋𝑋] = 𝜆𝜆
Variance Var[𝑋𝑋] = 𝜆𝜆
𝑋𝑋 − 𝜇𝜇 3 1
Skewness 𝐸𝐸 �� � �=
𝜎𝜎 √𝜆𝜆
𝑋𝑋 − 𝜇𝜇 4 1
Kurtosis 𝐸𝐸 �� � �−3 =
𝜎𝜎 𝜆𝜆
Moment generating �𝜆𝜆�e𝑡𝑡 −1��
functions (mgf) 𝑀𝑀𝑋𝑋 (𝑡𝑡) = e , 𝑡𝑡 ∈ ℝ

Multivariate hypergeometric distribution. For a MultiHypergeom(𝑛𝑛, 𝑛𝑛1 , … , 𝑛𝑛𝑘𝑘 )-


distributed random variable (𝑋𝑋1 , … , 𝑋𝑋𝑘𝑘 ) applies:
𝑛𝑛 𝑛𝑛
⎧� 𝑥𝑥1 � ⋅ … ⋅ �𝑥𝑥𝑘𝑘 �
⎪ 1 𝑘𝑘
, 𝑥𝑥1 + ⋯ + 𝑥𝑥𝑘𝑘 = 𝑛𝑛
Probability function 𝑓𝑓(𝑋𝑋1 ,…,𝑋𝑋𝑘𝑘 ) (𝑥𝑥1 , … , 𝑥𝑥𝑘𝑘 ) = 𝑛𝑛1 + ⋯ + 𝑛𝑛𝑘𝑘
⎨� 𝑛𝑛


⎩ 0, Otherwise

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]

Continues Distributions
1 dimensional continues uniform distribution. For a Uniform([𝑎𝑎, 𝑏𝑏])-distributed
random variable X applies
1
, 𝑥𝑥 ∈ [𝑎𝑎, 𝑏𝑏]
Probability function 𝑓𝑓𝑋𝑋 (𝑥𝑥) = �𝑏𝑏 − 𝑎𝑎
0, 𝑥𝑥 ∉ [𝑎𝑎, 𝑏𝑏]
1
Expected value 𝐸𝐸[𝑋𝑋] = (𝑎𝑎 + 𝑏𝑏)
2
1
Variance Var[𝑋𝑋] = (𝑏𝑏 − 𝑎𝑎)2
12
𝑋𝑋 − 𝜇𝜇 3
Skewness 𝐸𝐸 �� � �=0
𝜎𝜎
𝑋𝑋 − 𝜇𝜇 4 6
Kurtosis 𝐸𝐸 �� � �−3 = −
𝜎𝜎 5
𝑒𝑒 𝑏𝑏𝑏𝑏 − 𝑒𝑒 𝑎𝑎𝑎𝑎
Moment generating
𝑀𝑀𝑋𝑋 (𝑡𝑡) = � (𝑏𝑏 − 𝑎𝑎)𝑡𝑡 , 𝑡𝑡 ≠ 0
functions (mgf)
1, 𝑡𝑡 = 0

1 Dimensional Normal Distribution. For a Normal(𝜇𝜇, 𝜎𝜎 2 )-distributed random


variable X applies
1 (𝑥𝑥−𝜇𝜇)2
Probability density �− �
function 𝑓𝑓𝑋𝑋 (𝑥𝑥) = 𝑒𝑒 2𝜎𝜎2 , 𝑥𝑥 ∈ ℝ
√2𝜋𝜋𝜎𝜎 2
Expected value 𝐸𝐸[𝑋𝑋] = 𝜇𝜇
Variance Var[𝑋𝑋] = 𝜎𝜎 2
𝑋𝑋 − 𝜇𝜇 3
Skewness 𝐸𝐸 �� � �=0
𝜎𝜎
𝑋𝑋 − 𝜇𝜇 4
Kurtosis 𝐸𝐸 �� � �−3 = 0
𝜎𝜎
Moment generating 1 2 2
functions (mgf) 𝑀𝑀𝑋𝑋 (𝑡𝑡) = 𝑒𝑒 𝜇𝜇𝜇𝜇+2𝜎𝜎𝑡𝑡
, 𝑡𝑡 ∈ ℝ

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]
Exponential distribution. For an Exponential(𝜆𝜆)-distributed random variable X
applies
Probability density −𝜆𝜆𝜆𝜆
𝑓𝑓𝑋𝑋 (𝑥𝑥) = �𝜆𝜆𝑒𝑒 𝑥𝑥 > 0
function 0, 𝑥𝑥 ≤ 0
Cumulative −𝜆𝜆𝜆𝜆
𝐹𝐹𝑋𝑋 (𝑥𝑥) = �1 − 𝑒𝑒 , 𝑥𝑥 > 0
distribution function 0, 𝑥𝑥 ≤ 0
1
Expected value 𝐸𝐸[𝑋𝑋] =
𝜆𝜆
1
Variance Var[𝑋𝑋] = 2
𝜆𝜆
𝑋𝑋 − 𝜇𝜇 3
Skewness 𝐸𝐸 �� � �=2
𝜎𝜎
𝑋𝑋 − 𝜇𝜇 4
Kurtosis 𝐸𝐸 �� � �−3 = 6
𝜎𝜎
𝜆𝜆
Moment generating
𝑀𝑀𝑋𝑋 (𝑡𝑡) = �𝜆𝜆 − 𝑡𝑡 , 𝑡𝑡 < 𝜆𝜆
functions (mgf)
∞, 𝑡𝑡 ≥ 𝜆𝜆

Student’s T distribution. For a T (𝜈𝜈)- distributed random variable X applies


𝜈𝜈 + 1 −
𝜈𝜈+1
Probability density Γ� 2 � 𝑥𝑥 2 2
𝑓𝑓𝑋𝑋 (𝑥𝑥) = 𝜈𝜈 �1 + � 𝑥𝑥 ∈ ℝ
function 𝜈𝜈
√𝜋𝜋𝜋𝜋Γ �2�

Γ(𝑥𝑥) = ∫0 𝑡𝑡 𝑥𝑥−1 𝑒𝑒 −𝑡𝑡 𝑑𝑑𝑑𝑑 𝑥𝑥 > 0
Gamma function
Γ(𝑛𝑛 + 1) = 𝑛𝑛! 𝑛𝑛 ∈ ℕ0

Gamma distribution. For a Gamma (𝛼𝛼, 𝛽𝛽)- distributed random variable X applies
𝛽𝛽 𝛼𝛼 𝛼𝛼−1 −𝛽𝛽𝛽𝛽
Probability density 𝑥𝑥 e 𝑥𝑥 > 0
𝑓𝑓𝑋𝑋 (𝑥𝑥) = �Γ(𝛼𝛼)
function
0, 𝑂𝑂𝑂𝑂ℎ𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒
𝛼𝛼
Expected value 𝐸𝐸[𝑋𝑋] =
𝛽𝛽
𝛼𝛼
Variance Var[𝑋𝑋] = 2
𝛽𝛽

Beta distribution. For a Beta(𝛼𝛼, 𝛽𝛽)- distributed random variable X applies


Γ(𝛼𝛼 + 𝛽𝛽) 𝛼𝛼−1
Probability density 𝑥𝑥 (1 − 𝑥𝑥)𝛽𝛽−1 0 ≤ 𝑥𝑥 ≤ 1
𝑓𝑓𝑋𝑋 (𝑥𝑥) = �Γ(𝛼𝛼)Γ(𝛽𝛽)
function
0, 𝑂𝑂𝑂𝑂ℎ𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒
𝛼𝛼
Expected value 𝐸𝐸[𝑋𝑋] =
𝛼𝛼 + 𝛽𝛽
𝛼𝛼𝛼𝛼
Variance Var[𝑋𝑋] = 2
(𝛼𝛼 + 𝛽𝛽) (𝛼𝛼 + 𝛽𝛽 + 1)

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]
Weibull distribution. For a Weibull(𝑘𝑘, 𝜆𝜆)- distributed random variable X applies
𝑘𝑘 𝑥𝑥 𝑘𝑘−1 −�𝑥𝑥�𝑘𝑘
Probability density
𝑓𝑓𝑋𝑋 (𝑥𝑥) = �𝜆𝜆 �𝜆𝜆 � 𝑒𝑒 𝜆𝜆 , 𝑥𝑥 ≥ 0
function
0, 𝑂𝑂𝑂𝑂ℎ𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒𝑒
𝑥𝑥 𝑘𝑘
Cumulative −� �
𝐹𝐹𝑋𝑋 (𝑥𝑥) = �1 − 𝑒𝑒 𝜆𝜆 , 𝑥𝑥 > 0
distribution function 0, 𝑥𝑥 ≤ 0

𝜇𝜇1
Multivariate Normal distribution. For a Normal ��𝜇𝜇 � , Σ�-distributed random
2
variable (𝑋𝑋, 𝑌𝑌) applies:
T
1 𝑥𝑥 𝜇𝜇1 𝑥𝑥 𝜇𝜇1
1 − ��𝑦𝑦�−�𝜇𝜇 �� Σ−1 ��𝑦𝑦�−�𝜇𝜇 ��
2
Probability density 𝑓𝑓𝑋𝑋,𝑌𝑌 (𝑥𝑥, 𝑦𝑦) = 𝑒𝑒 2 2
,
function �(2𝜋𝜋)𝑛𝑛 det(Σ)
𝑥𝑥, 𝑦𝑦 ∈ ℝ

Markov’s inequality
𝐸𝐸[𝑋𝑋]
𝑃𝑃(𝑋𝑋 ≥ 𝑡𝑡) ≤ .
𝑡𝑡
Chebyshev’s inequality
σ2
𝑃𝑃(|𝑋𝑋 − μ| ≥ 𝑡𝑡) ≤ 2 .
𝑡𝑡
Hoeffding’s inequality
2𝑛𝑛𝑡𝑡 2
𝑃𝑃(𝑋𝑋� − μ ≥ 𝑡𝑡) ≤ 𝑒𝑒𝑒𝑒𝑒𝑒 �− �
(𝑏𝑏 − 𝑎𝑎)2

Misprints and errors are reserved page 1 from 6 Date 03/17/2023


PRÜFUNGSAMT
[Link]

Misprints and errors are reserved page 1 from 6 Date 03/17/2023

You might also like