Mimo
Mimo
Communication
MIMO Wireless Communication
(ELL8415)
Anupama Rajoriya
Assistant Professor
Department of Electrical Engineering
Motivation for MIMO communication
Multiple transmit and/or receive antennas
o Diversity gain: Transmit same signal over multiple channels→ boosts the SNR
o Spatial multiplexing: Provides transmission of multiple parallel streams
o Beamforming- serving users in a particular direction→ Interference reduction
Course Objectives
30
30 % 30 %
25
20
20 % 20 %
15
10
0
Assignments/Quizes Midsem Term project Endsem
A Wireless communication link
• Additive white Gaussian noise (AWGN)
Transmitter Receiver
𝑥 𝑦
o Received signal:
𝑦 =𝑥+𝑛
− 𝑃 + 𝑛, 𝑖𝑓 𝑥 = − 𝑃
o𝑦= ൝
𝑃 + 𝑛, 𝑖𝑓 𝑥 = 𝑃
Transmitter Receiver
𝑥 𝑦
o Capacity: maximum number of bits per second that can be transmitted reliably over a wireless channel
𝑃
𝐶 = log 2 1 + 2
𝜎
𝑃
o Here is commonly known as SNR - Signal to Noise Ratio
𝜎2
𝑃 𝑃 𝑃 Property of Logarithm
o When 𝜎2 ≪ 1 , 𝐶 = log 2 1 + ≈ log 2 𝑒→ Low SNR regime
𝜎2 𝜎2
o
𝑃
When 𝜎2 ≫ 1 , 𝐶 = log 2 1 +
𝑃
≈
𝑃
log 2 𝜎2 → High SNR regime 𝑥 log 2 𝑒, if 𝑥 ≪ 1
𝜎2 log 2 1 + 𝑥 = ቊ
log 2 𝑥, if 𝑥 ≫ 1.
Transmitter Receiver
𝑥 𝑦
o Capacity: maximum number of bits per second that can be transmitted reliably over a wireless channel
𝑃
𝐶 = log 2 1 + 2
𝜎
𝑃
o Here is commonly known as SNR - Signal to Noise Ratio
𝜎2
𝑃 𝑃 𝑃 Property of Logarithm
o When 𝜎2 ≪ 1 , 𝐶 = log 2 1 + ≈ log 2 𝑒→ Low SNR regime
𝜎2 𝜎2
o
𝑃
When 𝜎2 ≫ 1 , 𝐶 = log 2 1 +
𝑃
≈
𝑃
log 2 𝜎2 → High SNR regime 𝑥 log 2 𝑒, if 𝑥 ≪ 1
𝜎2 log 2 1 + 𝑥 = ቊ
log 2 𝑥, if 𝑥 ≫ 1.
Single input and multiple output (SIMO) Multiple input and single output (MISO)
o One transmit antenna, two receive antennas o Two transmit antennas, one receive antenna
𝑦1 𝑥1
Transmitter Receiver Transmitter
𝑥 Receiver
𝑦2 𝑥2 𝑦
Single input and multiple output (SIMO)
• Additive white Gaussian noise (AWGN)
𝑦1
o Received signal at antenna 1:
𝑦1 = 𝑥 + 𝑛1 Transmitter Receiver
o Received signal at antenna 2: 𝑥
𝑦2 = 𝑥 + 𝑛2 𝑦2
o 𝑛1 and 𝑛2 are independently random variables with distribution CN(0, 𝜎 2 )
o Added received signal
𝑦 = 𝑦1 + 𝑦2 = 2𝑥 + 𝑛1 + 𝑛2
2𝑃
o SNR calculation : 𝜎2
o Capacity
2𝑃
𝐶 = log 2 1 + 2
𝜎
2𝑃 2𝑃
o At low SNR, 𝐶 = log 2 1 + 𝜎2 ≈ 𝜎2 log 2 𝑒→ low SNR regime
At low SNR, capacity increases linearly with number of receive antennas→ Boosting the SNR
Multiple input and single output (MISO)
• Additive white Gaussian noise (AWGN)
o One symbol transmitted over two antennas 𝑥1
o Signal transmitted over antenna 1:
Transmitter
𝑥1 = 𝑥/ 2 Receiver
o Signal transmitted over antenna 2: 𝑥2 𝑦
𝑥2 = 𝑥/ 2
P P
o Total transmit power: 2 + 2 = 𝑃 (same as single antenna case)
o Received signal:
𝑥 𝑥
𝑦 = 𝑥1 + 𝑥2 + 𝑛 = + + 𝑛 = 2𝑥 + 𝑛
2 2
2𝑃
o SNR calculation : 𝜎2
o Capacity
2𝑃
𝐶 = log 2 1 + 2
𝜎
2𝑃 2𝑃
o At low SNR, 𝐶 = log 2 1 + 𝜎2 ≈ 𝜎2 log 2 𝑒→ Low SNR regime
At low SNR, capacity increases linearly with number of receive antennas→ Boosting the SNR
Motivation for MIMO communication?
• Fading channel
Channel
Transmitter Receiver
𝑥 𝑦
o SNR calculation :
ℎ1 2 + ℎ2 2 2 𝑃 ℎ1 2 + ℎ2 2 𝑃
= Example
ℎ1 2 𝜎 2 + ℎ2 2 𝜎 2 𝜎2
o Capacity For ℎ1 = 1, ℎ2 = −0.5,
2 2
ℎ1 + ℎ2 𝑃 1.25𝑃
𝐶 = log 2 1 + SNR = 𝜎2
𝜎2
Matched filter used here maximizes the SNR
Single input and multiple output (SIMO)
• Fading channel
ℎ1
𝑦1
o Received signal at antenna 1:
𝑦1 = ℎ1 𝑥 + 𝑛1 ℎ2
Transmitter Receiver
o Received signal at antenna 2: 𝑥
𝑦2 = ℎ2 𝑥 + 𝑛2 𝑦2
o 𝑛1 and 𝑛2 are independently random variables with distribution CN(0, 𝜎 2 )
o Weigh and Add the two signals to strengthen the SNR Receive beamforming/combining
o Lets use Matched filtering
𝑦 = ℎ1∗ 𝑦1 + ℎ2∗ 𝑦2 = ℎ1 2 + ℎ2 2 𝑥 + ℎ1∗ 𝑛1 + ℎ2∗ 𝑛2
o SNR calculation :
ℎ1 2 + ℎ2 2 2 𝑃 ℎ1 2 + ℎ2 2 𝑃
= Example
ℎ1 2 𝜎 2 + ℎ2 2 𝜎 2 𝜎2
o Capacity For ℎ1 = 1, ℎ2 = −0.5,
2 2
ℎ1 + ℎ2 𝑃 1.25𝑃
𝐶 = log 2 1 + SNR = 𝜎2
𝜎2
Matched filter used here maximizes the SNR
Multiple input and single output (MISO)
• Fading channel ℎ1
o One symbol transmitted over two antennas 𝑥1 ℎ2
o Channel is assumed to be unknown at the transmitter
o Signal transmitted over antenna 1: Transmitter
Receiver
𝑥1 = 𝑥/ 2 𝑥2 𝑦
o Signal transmitted over antenna 2:
𝑥2 = 𝑥/ 2
o Received signal:
𝑥 𝑥 𝑥
𝑦 = ℎ1 𝑥1 + ℎ2 𝑥2 + 𝑛 = ℎ1 + ℎ2 + 𝑛 = ℎ1 + ℎ2 +𝑛
2 2 2
o SNR calculation :
| ℎ1 + ℎ2 |2 𝑃
Example
2𝜎 2
For ℎ1 = 1, ℎ2 = −0.5,
0.125𝑃
SNR = 𝜎2
o Received signal:
2
𝑥 2
𝑥 2 2
𝑥
𝑦 = ℎ1 𝑥1 + ℎ2 𝑥2 + 𝑛 = ℎ1 + ℎ2 +𝑛= ℎ1 + ℎ2 +𝑛
𝑐 𝑐 𝑐
Example
o SNR calculation :
ℎ1 2 + ℎ2 2 𝑃
For ℎ1 = 1, ℎ2 = −0.5,
𝜎2 1.25𝑃
SNR = 𝜎2
Multiple input and single output (MISO)
• Fading channel ℎ1
o Assume the channels to be known at the transmitter 𝑥1 ℎ2
o Signal transmitted over antenna 1:
ℎ1∗ 𝑥 Transmitter
Receiver
𝑥1 =
𝑐 𝑥2 𝑦
o Signal transmitted over antenna 2:
ℎ2∗ 𝑥
Transmit beamforming 𝑥2 = 𝑐
ℎ1 2 + ℎ2 2 𝑃
o Transmit power (should be P): =𝑃
𝑐
⟹𝑐 = ℎ1 2 + ℎ2 2
o Received signal:
2
𝑥 2
𝑥 2 2
𝑥
𝑦 = ℎ1 𝑥1 + ℎ2 𝑥2 + 𝑛 = ℎ1 + ℎ2 +𝑛= ℎ1 + ℎ2 +𝑛
𝑐 𝑐 𝑐
Example
o SNR calculation :
ℎ1 2 + ℎ2 2 𝑃
For ℎ1 = 1, ℎ2 = −0.5,
𝜎2 1.25𝑃
SNR = 𝜎2
Multiple input and single output (MISO)
• Fading channel ℎ1
o Assume the channels to be known at the transmitter 𝑥1 ℎ2
o Signal transmitted over antenna 1:
ℎ1∗ 𝑥 Transmitter
Receiver
𝑥1 =
𝑐 𝑥2 𝑦
o Signal transmitted over antenna 2:
ℎ2∗ 𝑥
Transmit beamforming 𝑥2 = 𝑐
ℎ1 2 + ℎ2 2 𝑃
o Transmit power (should be P): =𝑃 Multiple receive antennas Multiple transmit antennas
𝑐
⟹𝑐 = ℎ1 2 + ℎ2 2
Combining Beamforming
o Received signal:
2
𝑥 2
𝑥 2 2
𝑥
𝑦 = ℎ1 𝑥1 + ℎ2 𝑥2 + 𝑛 = ℎ1 + ℎ2 +𝑛= ℎ1 + ℎ2 +𝑛
𝑐 𝑐 𝑐
Example
o SNR calculation :
ℎ1 2 + ℎ2 2 𝑃
For ℎ1 = 1, ℎ2 = −0.5,
𝜎2 1.25𝑃
SNR = 𝜎2
Multiple input and multiple output (MIMO)
ℎ11
• Additive white Gaussian noise (AWGN) ℎ12
o Two symbols transmitted over two antennas ℎ21 𝑦1
𝑥1
o Signals transmitted: 𝑥1 and 𝑥2 ℎ22
o Received signal at antenna 1: Transmitter Receiver
𝑦1 = ℎ11 𝑥1 + ℎ12 𝑥2 + 𝑛1 𝑥2 𝑦2
o Received signal at antenna 2:
𝑦2 = ℎ21 𝑥1 + ℎ22 𝑥2 + 𝑛2
𝒚 = 𝑯𝒙 + 𝒏
2×1
2×1 2×2
Motivation for MIMO communication
Multiple transmit and/or receive antennas
o Diversity gain: Transmit same signal over multiple channels→ boosts the SNR
A signal transmitted over two independently fading channels is more likely to be correctly detected than the
same signal transmitted over a single fading channel
Motivation for MIMO communication
Multiple transmit and/or receive antennas
o Diversity gain: Transmit same signal over multiple channels→ boosts the SNR
A signal transmitted over two independently fading channels is more likely to be correctly detected than the
same signal transmitted over a single fading channel
o Spatial multiplexing: Provides transmission of multiple parallel streams
ℎ11
ℎ12
ℎ21 𝑦1
𝑥1
ℎ22
Transmitter Receiver
𝑥2 𝑦2
Motivation for MIMO communication
Multiple transmit and/or receive antennas
o Diversity gain: Transmit same signal over multiple channels→ boosts the SNR
A signal transmitted over two independently fading channels is more likely to be correctly detected than the
same signal transmitted over a single fading channel
o Spatial multiplexing: Provides transmission of multiple parallel streams
ℎ11
ℎ12
ℎ21 𝑦1
𝑥1
ℎ22
Transmitter Receiver
𝑥2 𝑦2
Lecture 2
Reference:
Elements of Information Theory
Thomas M. Cover, Joy A. Thomas
Information theory
o Information theory- compressibility, storage of data and reliable communication
o Every channel has a characterizing quantity (capacity), such that, for the transmission rates below it,
the error probability could be made arbitrarily small.
o In other words, it is the maximum achievable rate with reliability (very low error probability)
o Information required to describe a quantity ≡ uncertainty
o If a quantity is certain, no information
o Appropriate to model information and data communication using probability theory
o Familiarity with probability theory is expected
Example
1, with probabilit𝑦 p
o Entropy of a Bernoulli random variable 𝑋 = ቊ
0, with probabilit𝑦 (1 − p)
𝐻 𝑋 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2(1 − 𝑝)
o X is output of a fair coin toss: 𝑝 = 1/2
1 1 1 1
o Entropy: 𝐻 𝑋 = − log 2 − log 2 = 1 bit
2 2 2 2
Entropy
o Let 𝑋 ∈ 𝒳be a discrete random variable
o With probability mass function (pmf) :
𝑝 𝑥 = Pr 𝑋 = 𝑥 , 𝑋 ∈ 𝒳
𝐻 𝑋
𝑥∈𝒳
o Non-negative: 𝐻 𝑋 ≥ 0
o larger entropy ⇒ 𝑋 is more unpredictable
o Zero entropy- 𝑋 is deterministic (no uncertainty)
o Information (no of bits) required to describe the quantity
o Function of probabilities p(x), not of X
Example
1, with probabilit𝑦 p
o Entropy of a Bernoulli random variable 𝑋 = ቊ
0, with probabilit𝑦 (1 − p)
𝐻 𝑋 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2(1 − 𝑝)
o X is output of a fair coin toss: 𝑝 = 1/2
1 1 1 1
o Entropy: 𝐻 𝑋 = − log 2 − log 2 = 1 bit
2 2 2 2
Joint entropy and conditional entropy
o For a pair of random variables (𝑋, 𝑌) with the joint distribution 𝑝(𝑥, 𝑦), joint entropy is defined as
𝐻 𝑋, 𝑌 = − 𝑝 𝑥, 𝑦 log 2 𝑝 𝑥, 𝑦 = −𝔼 log 2 𝑝(𝑥, 𝑦)
𝑥∈𝒳 𝑦∈𝒴
o Conditional entropy of random variable 𝑌 given 𝑋 is defined as
𝐻 𝑌|𝑋 = − 𝑝 𝑥 𝐻(𝑌|𝑋 = 𝑥) = −𝔼𝑝 𝑥 𝐻(𝑌|𝑋 = 𝑥)
𝑥∈𝒳
𝑝 𝑥 𝑝 𝑥
𝐷(𝑝||𝑞) = 𝑝 𝑥 log = 𝔼𝑝 log .
𝑞 𝑥 𝑞 𝑥
𝑥∈𝒳
0 𝑝
o Convention used: 0 log 𝑞 = 0 and p log 0 = ∞
o measure of the inefficiency of assuming that the distribution is q when the true distribution is p
o if true distribution 𝑝 𝑥 is known, code can be constructed with average description length H(p)
o if a distribution 𝑞 𝑥 is used, we would need H p + 𝐷(𝑝||𝑞) bits on the average
o Always non-negative and 𝐷(𝑝| 𝑞 = 0, if and only 𝑝 = 𝑞.
o Not symmetric: 𝐷(𝑝| 𝑞 ≠ 𝐷(𝑞||𝑝)
Example
o Consider two distributions 𝑝 and 𝑞 on a random variable 𝑋 ∈ {0,1}
o Let 𝑝(0) = 1 − 𝑟, 𝑝 1 = 𝑟 and 𝑞 0 = 1 − 𝑠, 𝑞 1 = 𝑠
o Relative entropy:
𝑝 0 𝑝 1 1−𝑟 𝑟
D(p| q = 𝑝 0 log 2 𝑞 + 𝑝 1 log 2 𝑞 = 1 − 𝑟 log 2 + 𝑟 log 2 𝑠
0 1 1−𝑠
1−𝑠 𝑠
D(q| p = 1 − 𝑠 log 2 + 𝑠 log 2
1−𝑟 𝑟
Relative Entropy
Example
o Consider two distributions 𝑝 and 𝑞 on a random variable 𝑋 ∈ {0,1}
o Let 𝑝(0) = 1 − 𝑟, 𝑝 1 = 𝑟 and 𝑞 0 = 1 − 𝑠, 𝑞 1 = 𝑠
o Relative entropy:
𝑝 0 𝑝 1 1−𝑟 𝑟
D(p| q = 𝑝 0 log 2 𝑞 + 𝑝 1 log 2 𝑞 = 1 − 𝑟 log 2 + 𝑟 log 2 𝑠
0 1 1−𝑠
1−𝑠 𝑠
D(q| p = 1 − 𝑠 log 2 + 𝑠 log 2
1−𝑟 𝑟
o If 𝑟 = 𝑠, ⇒ D(p| q = D(q| p = 0
1 1
o Let 𝑟 = 2 and s = 4
1 1
1 2 1 2
D(p| q = log 2 3 + log 2 1 = 0.2075 bits
2 2
4 4
3 1
3 4 1 4
D(q| p = log 2 1 + log 2 1 = 0.187 bits
4 4
2 2
Mutual information
o measure of the amount of information that one random variable contains about another random variable
o For X and Y, the mutual information 𝐼 𝑋; 𝑌 is defined as
𝑝 𝑥, 𝑦
𝐼 𝑋; 𝑌 = 𝑝 𝑥, 𝑦 log
𝑝 𝑥 𝑝(𝑦)
𝑥∈𝒳 𝑦∈𝒴
= 𝐷(𝑝(𝑥, 𝑦)||𝑝 𝑥 𝑝(𝑦))
𝑝 𝑥, 𝑦
= 𝔼𝑝 𝑥,𝑦 log
𝑝 𝑥 𝑝(𝑦)
o Quantifies the reduction in the uncertainty of one random variable due to the knowledge of the other
o Mutual information is symmetric: 𝐼(𝑋; 𝑌) = 𝐼(𝑌; 𝑋)
Mutual information and entropy
o Rewriting the mutual information expression
𝑝 𝑥, 𝑦
𝐼 𝑋; 𝑌 = 𝑝 𝑥, 𝑦 log
𝑝 𝑥 𝑝(𝑦)
𝑥∈𝒳 𝑦∈𝒴
𝑝 𝑥|𝑦
= 𝑝 𝑥, 𝑦 log
𝑝 𝑥
𝑥∈𝒳 𝑦∈𝒴
𝑓 𝑝𝑖 𝑥𝑖 ≤ 𝑝𝑖 𝑓 𝑥𝑖
𝑖=1 𝑖=1
Jensen’s inequality
𝑘−1 𝑘−1
𝑓 𝑝𝑖 𝑥𝑖 ≤ 𝑝𝑖 𝑓 𝑥𝑖 (1)
𝑖=1 𝑖=1
o For distribution with k mass points and probabilities 𝑞1 , 𝑞2 , . . 𝑞𝑘
𝑘 𝑘−1
𝑞𝑖 𝑓 𝑥𝑖 = 𝑞𝑘 𝑓 𝑥𝑘 + 𝑞𝑖 𝑓(𝑥𝑖 )
𝑖=1 𝑖=1
𝑞𝑖 𝑞𝑖
o Lets define 𝑝𝑖 = σ𝑘−1 = . Replace 𝑞𝑖 with 𝑝𝑖 for 𝑖 = 1, . . (𝑘 − 1) in the RHS to get
𝑗=1 𝑞𝑗 (1−𝑞𝑘 )
𝑘 𝑘−1
𝑞𝑖 𝑓 𝑥𝑖 = 𝑞𝑘 𝑓 𝑥𝑘 + (1 − 𝑞𝑘 ) 𝑝𝑖 𝑓(𝑥𝑖 )
𝑖=1 𝑖=1
𝑘−1
By using Eq (1)
≥ 𝑞𝑘 𝑓 𝑥𝑘 + (1 − 𝑞𝑘 )𝑓 𝑝𝑖 𝑥𝑖
𝑖=1
𝑘−1 𝑘
≥ 𝑓 𝑞𝑘 𝑥𝑘 + 1 − 𝑞𝑘 𝑝𝑖 𝑥𝑖 = 𝑓 𝑞𝑖 𝑥𝑖 Convexity definition
𝑖=1 𝑖=1
𝑘 𝑘
⇒ 𝑞𝑖 𝑓 𝑥𝑖 ≥ 𝑓 𝑞𝑖 𝑥𝑖 ⇒𝑓 𝔼𝑋 ≤ 𝔼[𝑓(𝑋)]
𝑖=1 𝑖=1
Information inequality
• Relative entropy
= −𝐻 𝑋 + log 𝒳 𝑝 𝑥
𝑥∈𝒳
𝐷(𝑝| 𝑢 = −𝐻 𝑋 + log 𝒳 ≥ 0
⇒ 𝐻 𝑋 ≤ log |𝒳|
Lecture 3
Reference:
Elements of Information Theory
Thomas M. Cover, Joy A. Thomas
Recap
o Some basic definitions for discrete random variables:
▪ Entropy,
▪ Joint and conditional entropies,
▪ Relative entropy,
▪ Mutual information
In today’s class
o For continuous random variables:
▪ Differential entropy
▪ Joint and conditional entropies,
▪ Relative entropy
▪ Mutual information
o Entropy of Gaussian random variables and random vectors
More on Relative entropy
o Kullback-Leibler divergence (KL divergence)
o Not a mathematical distance metric (no symmetry, does not satisfy triangular inequality)
o Signifies the “information loss”
o if true distribution 𝑝 𝑥 is known: average description length 𝐻(𝑝)
o if a distribution 𝑞 𝑥 is used, we would need H p + 𝐷(𝑝||𝑞) bits on the average
o Simple difference between distributions will not have similar meaning
o Example: Lets assume the case where 𝑞 𝑋 = 𝑥 = 0 when the true distribution 𝑝 𝑋 = 𝑥 > 0
Relative entropy/KL divergence becomes ∞→ signifying complete failure of the assumption
o Simple difference will give just a number, which does not properly represent the catastrophic failure
𝑝(𝑋 = 𝑥)
4
𝑞(𝑋 = 𝑥)
1 1
1 2 1 2
4 4
0 1 2 𝑥 0 1 2 𝑥
Differential entropy
o Let 𝑋 be a continuous random variable with cumulative distribution function (CDF)
𝐹 𝑥 = Pr(𝑋 ≤ 𝑥)
o Its probability distribution function (pdf):
𝑓 𝑥 = 𝐹′(𝑥)
o Set where 𝑓 𝑥 > 0 is called the support set of 𝑋
Example
When a<1, ℎ 𝑋 < 0
o Differential entropy of a uniformly distributed random variable 𝑋 ∈ [0, 𝑎]
1
⇒ Differential entropy can be negative
if 0 ≤ 𝑥 ≤ 𝑎,
o Pdf: 𝑓 𝑥 = ൝𝑎
0 otherwise.
𝑎
𝑎 1 1
ℎ 𝑋 = −∫0 𝑓 𝑥 log 𝑓 𝑥 𝑑𝑥 = − න log 𝑑𝑥 = log 𝑎
0 𝑎 𝑎
Joint and conditional differential entropy
o For pair of random variables 𝑋 and 𝑌 with joint pdf 𝑓(𝑥, 𝑦), joint differential entropy is defined as
ℎ 𝑋, 𝑌 = − න න 𝑓 𝑥, 𝑦 log 𝑓 𝑥, 𝑦 𝑑𝑦 𝑑𝑥
𝑆𝑋 𝑆𝑌
o 𝑆𝑋 : support set of 𝑋 and 𝑆𝑌 : support set of 𝑌
o The conditional differential entropy is defined as
ℎ 𝑋 𝑌 = − න න 𝑓 𝑥, 𝑦 log 𝑓 𝑥 𝑦 𝑑𝑥 𝑑𝑦
𝑆𝑋 𝑆𝑌
𝑓(𝑥, 𝑦)
= − න න 𝑓 𝑥, 𝑦 log 𝑑𝑥 𝑑𝑦
𝑆𝑋 𝑆𝑌 𝑓 𝑦
= − න න 𝑓 𝑥, 𝑦 log 𝑓(𝑥, 𝑦) 𝑑𝑥 𝑑𝑦 + න න 𝑓 𝑥, 𝑦 log 𝑓 𝑦 𝑑𝑥 𝑑𝑦
𝑆𝑋 𝑆𝑌 𝑆𝑋 𝑆𝑌
= ℎ 𝑋, 𝑌 + න 𝑓 𝑦 log 𝑓 𝑦 𝑑𝑦
𝑆𝑌
= ℎ 𝑋, 𝑌 − ℎ(𝑌)
Relative entropy and Mutual information
o Relative Entropy: KL-divergence between two densities 𝑓 𝑥 and 𝑔(𝑥) on a random variable 𝑋
𝑓 𝑥
𝐷(𝑓| 𝑔 = න 𝑓 𝑥 log 𝑑𝑥
𝑆 𝑔 𝑥
o Mutual Information:
𝐼 𝑋; 𝑌 = 𝐷(𝑓(𝑥, 𝑦)||𝑓 𝑥 𝑓(𝑦))
𝑓(𝑥, 𝑦)
= න න 𝑓 𝑥, 𝑦 log 𝑑𝑥 𝑑𝑦
𝑆 𝑋 𝑆𝑌 𝑓(𝑥)𝑓 𝑦
𝑓 𝑥 𝑓(𝑦|𝑥)
= න න 𝑓 𝑥, 𝑦 log 𝑑𝑥 𝑑𝑦
𝑆 𝑋 𝑆𝑌 𝑓(𝑥)𝑓 𝑦
= − න න 𝑓 𝑥, 𝑦 log 𝑓(𝑦) 𝑑𝑥 𝑑𝑦 + න න 𝑓 𝑥, 𝑦 log 𝑓 𝑦|𝑥 𝑑𝑥 𝑑𝑦
𝑆𝑋 𝑆𝑌 𝑆𝑋 𝑆𝑌
= − න 𝑓 𝑦 log 𝑓 𝑦 𝑑𝑦 − ℎ(𝑌|𝑋)
𝑆𝑌
= ℎ 𝑌 − ℎ(𝑌|𝑋)
= ℎ 𝑋 − ℎ(𝑋|𝑌)
= ℎ 𝑋 + ℎ(𝑌) − ℎ(𝑋, 𝑌) As ℎ 𝑌 𝑋 = ℎ 𝑋, 𝑌 − ℎ(𝑌)
Properties
o Relative Entropy: 𝐷(𝑓| 𝑔 ≥ 0
o Proof:
𝑓 𝑥 𝑔 𝑥
−𝐷(𝑓| 𝑔 = − න 𝑓 𝑥 log 𝑑𝑥 = න 𝑓 𝑥 log 𝑑𝑥
𝑆 𝑔 𝑥 𝑆 𝑓 𝑥
𝑔 𝑥 𝑔 𝑥
= 𝔼 log ≤ log 𝔼
𝑓 𝑥 𝑓 𝑥
𝑔 𝑥
= log න 𝑓 𝑥 𝑑𝑥 = log 1 = 0
𝑆 𝑓 𝑥
⇒ 𝐷(𝑓| 𝑔 ≥ 0
o Mutual Information: 𝐼 𝑋; 𝑌 ≥ 0. Equality holds iff X and Y are independent
o Conditioning reduces entropy: ℎ 𝑋 𝑌 ≤ ℎ(𝑋). Equality holds iff X and Y are independent
For Gaussian random variables
Differential entropy of a Gaussian random variable:
o Let 𝑋 be a zero mean Gaussian random variable with variance 𝜎 2 : 𝑋 ∼ 𝑁(0, 𝜎 2 )
o Its pdf
1 𝑥2
− 2
𝑓 𝑥 = e 2𝜎
2𝜋𝜎 2
o Differential entropy:
∞
ℎ 𝑋 = − න 𝑓 𝑥 log 𝑓 𝑥 𝑑𝑥
−∞
∞
1 𝑥2
−
= − න 𝑓 𝑥 log e 2𝜎2 𝑑𝑥
−∞ 2𝜋𝜎 2
∞ 2∞
𝑥
=න 𝑓 𝑥 log 2𝜋𝜎 2 𝑑𝑥 + න 2 𝑓 𝑥 log 𝑒 𝑑𝑥
−∞ −∞ 2𝜎
∞
2
log 𝑒 ∞ 2
= log 2𝜋𝜎 න 𝑓 𝑥 𝑑𝑥 + 2 න 𝑥 𝑓 𝑥 𝑑𝑥
−∞ 2𝜎 −∞
1 log 𝑒
= log 2𝜋𝜎 2 × 1 + 2 𝔼 𝑋2
2 2𝜎
1 log 𝑒
= log 2𝜋𝜎 2 × 1 + 2 × 𝜎2
2 2𝜎
1 1 1
= log 2𝜋𝜎 + log 𝑒 = log 2𝜋𝑒𝜎 2
2
2 2 2
For Gaussian random variables
Differential entropy of a Gaussian random vector:
o Let 𝑿 = [𝑋1 , 𝑋2 , … , 𝑋𝑛 ] ∈ ℝ𝑛×1be a random vector with pdf
1 1
−2 𝒙−𝝁 𝑇 𝚺 −1 (𝒙−𝝁)
𝑓 𝒙 = 1e
𝑛
2𝜋 𝚺 2
o Here 𝝁 ∈ ℝ𝑛×1 : mean of the vector 𝑿 and 𝚺 ∈ ℝ𝑛×𝑛 : covariance, defined as 𝚺 = 𝔼 (𝒙 − 𝝁) 𝒙 − 𝝁 𝑇
∞ ∞
𝑛 1 1
=න 𝑓 𝒙 log 2𝜋 𝚺 2 𝒙 − 𝝁 𝑇 𝚺 −1 (𝒙 − 𝝁)𝑓 𝒙 log 𝑒 𝑑𝒙
𝑑𝒙 + න
−∞ −∞ 2
∞
𝑛 1 log 𝑒 ∞
= log 2𝜋 𝚺 2 න 𝑓 𝒙 𝑑𝒙 + න 𝒙 − 𝝁 𝑇 𝚺 −1 (𝒙 − 𝝁)𝑓 𝒙 𝑑𝒙
−∞ 2 −∞
1 𝑛
log 𝑒
= log 2𝜋 𝚺 × 1 + 𝔼 𝒙 − 𝝁 𝑇 𝚺 −1 (𝒙 − 𝝁)
2 2
For Gaussian random variables
Differential entropy of a Gaussian random vector:
Proof:
o Let 𝑔(𝒙) be any pdf on 𝑿 such that 𝔼𝑔 𝑿 = 𝟎 and 𝔼𝑔 𝑿𝑿𝑇 = 𝚺
o Let 𝑓 𝒙 be the pdf when 𝑿 ∼ 𝑁(𝟎, 𝚺), i.e., 𝑓 𝒙 = 𝑁(𝟎, 𝚺) or
1 1
−2𝒙𝑇 𝚺 −1 𝒙
𝑓 𝒙 = 1e
𝑛
2𝜋 𝚺 2
o Relative entropy:
𝑔 𝒙
𝐷(𝑔| 𝑓 = න 𝑔 𝒙 log 𝑑𝒙
𝑓 𝒙
= න 𝑔 𝒙 log 𝑔 𝒙 𝑑𝒙 − න𝑔 𝒙 log 𝑓 𝒙 𝑑𝒙
1 1
−2𝒙𝑇 𝚺 −1 𝒙
= −ℎ 𝑿 − න 𝑔 𝒙 log 1e 𝑑𝒙
𝑛
2𝜋 𝚺 2
Gaussian random vector has maximum entropy
Let the random vector 𝑿 ∈ ℝ𝑛×1 has zero mean and covariance matrix 𝚺 = 𝔼 𝑿𝑿𝑇 ∈ ℝ𝑛×𝑛 . Then
1
• ℎ 𝑿 ≤ 2 log 2𝜋𝑒 𝑛 |𝚺|
• Equality holds if and only if 𝑿 ∼ 𝑁(𝟎, 𝚺)
𝑛 1 1 𝑇 −1
𝐷(𝑔| 𝑓 = −ℎ 𝑿 + log 2𝜋 𝚺 2 න𝑔 𝒙 𝑑𝒙 + න 𝒙 𝚺 𝒙 𝑔 𝒙 log 𝑒 𝑑𝒙
2
𝑛 1 1
= −ℎ 𝑿 + log 2𝜋 𝚺 2 × 1 + log 𝑒 𝔼𝑔 𝒙𝑇 𝚺 −1 𝒙
2
1 1
= −ℎ 𝑿 + log 2𝜋 𝑛 𝚺 + log 𝑒 Tr 𝔼𝑔 𝒙𝑇 𝚺 −1 𝒙
2 2
1 1
= −ℎ 𝑿 + log 2𝜋 𝑛 𝚺 + log 𝑒 𝔼𝑔 Tr 𝒙𝑇 𝚺 −1 𝒙
2 2
1 1
= −ℎ 𝑿 + log 2𝜋 𝑛 𝚺 + log 𝑒 Tr 𝚺 −1 𝔼𝑔 𝒙𝒙𝑇
2 2
1 𝑛
= −ℎ 𝑿 + log 2𝜋 𝑛 𝚺 + log 𝑒
2 2
1
= −ℎ 𝑿 + log 2𝜋𝑒 𝑛 𝚺 ≥0 Equality holds when 𝑓 𝒙 = 𝑔(𝒙)
2
1 ⇒ 𝑔 𝒙 = 𝑁(𝟎, 𝚺)
⇒ ℎ 𝑿 ≤ log 2𝜋𝑒 𝑛 𝚺
2
Motivation for MIMO communication
Multiple transmit and/or receive antennas
o Diversity gain: Transmit same signal over multiple channels→ boosts the SNR
A signal transmitted over two independently fading channels is more likely to be correctly detected than the
same signal transmitted over a single fading channel
o Spatial multiplexing: Provides transmission of multiple parallel streams
o Beamforming- serving users in a particular direction→ Interference reduction
ℎ11
ℎ12
ℎ21 𝑦1
𝑥1
ℎ22
Transmitter Receiver
𝑥2 𝑦2
Lecture 4
Reference:
Elements of Information Theory
Thomas M. Cover, Joy A. Thomas
Recap
o For continuous random variables:
▪ Differential entropy
▪ Joint and conditional entropies,
▪ Relative entropy
▪ Mutual information
o Entropy of Gaussian random variables and random vectors
In today’s class
o Definition of capacity
o Examples: Channel capacity
Capacity
Input 𝑋 Output 𝑌
Channel
o Capacity of a channel is the maximum of the mutual information between the input and output over all
distributions on the input that satisfy the power constraint:
𝐶 = max𝑝 𝑥 𝐼 𝑋; 𝑌
o Such that 𝔼 𝑋 2 ≤ 𝑃
o Noiseless ⇒ 𝑌 = 𝑋
o When 0 is transmitted 0 is received and vice versa
o Mutual information: 𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌
o Lets calculate 𝑝 𝑥 𝑦
𝑝 𝑋=1𝑌 =1 =1
𝑝 𝑋=1𝑌 =0 =0
𝑝 𝑋=0𝑌 =0 =1
𝑝 𝑋=0𝑌 =1 =0
1 if 𝑥 = 𝑦,
𝑝 𝑋=𝑥𝑌=𝑦 =ቊ
0 if 𝑥 ≠ 𝑦
o Conditional entropy:
𝐻 𝑋 𝑌 = − 𝑝 𝑥 𝑦 log 𝑝 𝑥 𝑦 = −1 × log 1 = 0
𝑥 𝑦
Examples
Noiseless binary channel
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
Examples
Noiseless binary channel
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
Examples
Noiseless binary channel
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
𝐼 𝑋; 𝑌 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2 (1 − 𝑝)
Examples
Noiseless binary channel
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
𝐼 𝑋; 𝑌 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2 (1 − 𝑝)
o Capacity
𝐶 = max 𝐼 𝑋; 𝑌 = max 𝐻 𝑋
Examples
Noiseless binary channel
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
𝐼 𝑋; 𝑌 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2 (1 − 𝑝)
o Capacity
𝐶 = max 𝐼 𝑋; 𝑌 = max 𝐻 𝑋
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
𝐼 𝑋; 𝑌 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2 (1 − 𝑝)
o Capacity
𝐶 = max 𝐼 𝑋; 𝑌 = max 𝐻 𝑋
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑋 − 𝐻 𝑋 𝑌 = 𝐻 𝑋
𝐼 𝑋; 𝑌 = −𝑝 log 2 𝑝 − (1 − 𝑝) log 2 (1 − 𝑝)
o Capacity
𝐶 = max 𝐼 𝑋; 𝑌 = max 𝐻 𝑋
we can send 1 bit per transmission over this channel with no errors
Examples
Noisy channel: Binary symmetric channel 1−𝑝
𝑥=0 𝑦=0
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
𝑝
o Mutual information: 𝐼 𝑋; 𝑌 = 𝐻 𝑌 − 𝐻 𝑌 𝑋
o Lets calculate 𝐻 𝑌 𝑋 = 𝑥 first
𝑝
𝐻 𝑌 𝑋 = 𝑥 = − 𝑝 𝑦 𝑋 = 𝑥 log 𝑝 𝑦 𝑋 = 𝑥 1−𝑝
𝑦 𝑥=1 𝑦=1
𝐻 𝑌 𝑋 = 0 = − 𝑝 𝑦 𝑋 = 0 log 𝑝(𝑦|𝑋 = 0)
𝑦
= −𝑝 log 𝑝 − 1 − 𝑝 log 1 − 𝑝 = 𝐻(𝑝)
o Similarly: 𝐻 𝑌 𝑋 = 1 = 𝐻(𝑝)
o Conditional entropy: 𝐻 𝑌 𝑋 = σ𝑥 𝑝 𝑥 𝐻 𝑌 𝑋 = 𝑥
o Mutual information:
𝐼 𝑋; 𝑌 = 𝐻 𝑌 − 𝐻 𝑌 𝑋
= 𝐻 𝑌 − 𝑝 𝑥 𝐻 𝑌 𝑋 = 𝑥
𝑥
=𝐻 𝑌 −𝐻 𝑝
≤ 1 − 𝐻(𝑝) As 𝐻 𝑌 = 1 is the maximum value of entropy
Examples
Noisy channel: Binary symmetric channel 1−𝑝
𝑥=0 𝑦=0
Input Output
Channel
𝑋 ∈ {0,1} 𝑌 ∈ {0,1}
𝑝
o Mutual information: 𝐼 𝑋; 𝑌 = 𝐻 𝑌 − 𝐻 𝑌 𝑋
o Lets calculate 𝐻 𝑌 𝑋 = 𝑥 first
𝑝
𝐻 𝑌 𝑋 = 𝑥 = − 𝑝 𝑦 𝑋 = 𝑥 log 𝑝 𝑦 𝑋 = 𝑥 1−𝑝
𝑦 𝑥=1 𝑦=1
𝐻 𝑌 𝑋 = 0 = − 𝑝 𝑦 𝑋 = 0 log 𝑝(𝑦|𝑋 = 0)
𝑦
= −𝑝 log 𝑝 − 1 − 𝑝 log 1 − 𝑝 = 𝐻(𝑝) Capacity of a binary symmetric channel:
o Similarly: 𝐻 𝑌 𝑋 = 1 = 𝐻(𝑝) C = 1 − 𝐻(𝑝)
o Conditional entropy: 𝐻 𝑌 𝑋 = σ𝑥 𝑝 𝑥 𝐻 𝑌 𝑋 = 𝑥
o Mutual information: It is achieved when 𝑋 is uniformly
𝐼 𝑋; 𝑌 = 𝐻 𝑌 − 𝐻 𝑌 𝑋 distributed, i.e., 𝑝 = 1/2
= 𝐻 𝑌 − 𝑝 𝑥 𝐻 𝑌 𝑋 = 𝑥
𝑥
=𝐻 𝑌 −𝐻 𝑝
≤ 1 − 𝐻(𝑝) As 𝐻 𝑌 = 1 is the maximum value of entropy
𝑍: noise
Examples
Noisy Gaussian channel:
Input Output
Channel
𝑋 𝑌
o Input output relation: 𝑋 𝑌
𝑌 =𝑋+𝑍
o The noise follows 𝑍 ∼ 𝑁(0, 𝜎 2 )
o Noise 𝑍 is independent of 𝑋
o Capacity with specified power constraint
max 𝐼(𝑋; 𝑌)
𝑝 𝑥
o Such that 𝔼 𝑋2≤𝑃
o Mutual information calculation
𝐼 𝑋; 𝑌 =ℎ 𝑌 −ℎ 𝑌 𝑋
=ℎ 𝑌 −ℎ 𝑋+𝑍 𝑋
=ℎ 𝑌 − ℎ(𝑍|𝑋) Using the result ℎ 𝑋 + 𝑍 𝑋 = ℎ(𝑍|𝑋)
=ℎ 𝑌 − ℎ(𝑍) Since 𝑍 is independent of 𝑋
o Since 𝑍 ∼ 𝑁(0, 𝜎 2 ), its entropy becomes:
1
ℎ 𝑍 = log 2𝜋𝑒𝜎 2
2
𝑍: noise
Examples
Noisy Gaussian channel:
Input Output
Channel
𝑋 𝑌
1 𝑋 𝑌
o Also ℎ 𝑌 ≤ 2 log 2𝜋𝑒𝔼[𝑌 2 ]: Gaussian distribution has maximum enetropy
o Lets calculate 𝔼[𝑌 2 ]
𝔼 𝑌2 = 𝔼 𝑋 + 𝑍 2
= 𝔼 𝑋 2 + 𝑍 2 + 2𝑋𝑍
= 𝑃 + 𝜎2 + 0
1 1
o We now have ℎ 𝑌 ≤ 2 log 2𝜋𝑒(𝑃 + 𝜎 2 ) and ℎ 𝑍 = 2 log 2𝜋𝑒𝜎 2
o This gives us
𝐼 𝑋; 𝑌 = ℎ 𝑌 − ℎ(𝑍)
1 1
≤ log 2𝜋𝑒 𝑃 + 𝜎 − log 2𝜋𝑒𝜎 2
2
2 2
1 𝑃
= log 1 + 2
2 𝜎
o Equality holds/maximum is achieved when 𝑌 is Gaussian
o As 𝑍 is Gaussian, if 𝑋 ∼ 𝑁(0, 𝑃), 𝑌 becomes Gaussian
𝑍: noise
Examples
Noisy Gaussian channel:
Input Output
Channel
𝑋 𝑌
1 𝑋 𝑌
o Also ℎ 𝑌 ≤ 2 log 2𝜋𝑒𝔼[𝑌 2 ]: Gaussian distribution has maximum enetropy
o Lets calculate 𝔼[𝑌 2 ]
𝔼 𝑌2 = 𝔼 𝑋 + 𝑍 2
= 𝔼 𝑋 2 + 𝑍 2 + 2𝑋𝑍
= 𝑃 + 𝜎2 + 0
1 1
o We now have ℎ 𝑌 ≤ 2 log 2𝜋𝑒(𝑃 + 𝜎 2 ) and ℎ 𝑍 = 2 log 2𝜋𝑒𝜎 2
o This gives us
𝐼 𝑋; 𝑌 = ℎ 𝑌 − ℎ(𝑍)
1 1
≤ log 2𝜋𝑒 𝑃 + 𝜎 − log 2𝜋𝑒𝜎 2
2
2 2
1 𝑃 Capacity of an AWGN channel:
= log 1 + 2 1 𝑃
2 𝜎 C = log 1 + 2
o Equality holds/maximum is achieved when 𝑌 is Gaussian 2 𝜎
o As 𝑍 is Gaussian, if 𝑋 ∼ 𝑁(0, 𝑃), 𝑌 becomes Gaussian
It is achieved when 𝑋 is Gaussian
Channel capacity
Input 𝑋 Output 𝑌
Channel
o Objective of data transmission: receiver should be able to decode the transmitted data with small
probability of error
o In practice, we cannot always identify a subset of the inputs to send information without error
o The idea is to send a sequence of data over multiple such channels (use codes with long block length)
o This allows us to send information at a rate C bits per transmission with an arbitrarily low probability of
error→ channel capacity theorem
Some definitions and notations
o Discrete channel: consists of input alphabet 𝒳, output alphabet 𝒴 and pmf 𝑝(𝑦|𝑥)
o Example: binary symmetric channel
Theorem 8.7.1 in
“Elements of Information Theory”,
Thoman M cover, Joy A Thomas
Rate R is said to be achievable for a Gaussian channel with a power constraint P if there exists a
sequence of 𝑀, 𝑛 codes with codewords satisfying the power constraint such that the maximal
probability of error 𝜆(𝑛) → 0.
Gaussian channel
o We derived the capacity of a single random variable→ considered single transmission
𝑊 𝑃
C= log 1 + bits /sec
2 𝑊𝜎 2
1 𝑃
C = log 1 + 2 Bits/sec/Hz
2 𝜎
o Capacity per complex dimensions
𝑃
C = log 1 + 2 Bits/sec/Hz
𝜎
Lecture 5
Reference:
Chapter 5
Fundamentals of Wireless communication
David Tse, Pramod Vishwanath
Recap
o Definition of capacity
o Examples: Channel capacity
In today’s class
o Capacity of AWGN channel
• SISO system
• SIMO system
• MISO system
Gaussian channel
o Capacity per complex dimensions
C = log 1 + 𝑆𝑁𝑅 Bits/sec/Hz
o Capacity
ℎ 2𝑃
𝐶 = log 2 1+ 2
𝜎
o Will talk about fading channels subsequently
SIMO Gaussian channel ℎ1
𝑦1
o Consider single-antenna transmitter and 𝐿 −antenna receiver .
o Received signal at the 𝑙-th antenna : Transmitter .
ℎ𝐿 . Receiver
𝑦𝑙 = ℎ𝑙 𝑥 + 𝑛𝑙 𝑥
o For all 𝑙 = 1, 2, … , 𝐿
o Here ℎ𝑙 is constant channel gain from transmitter to 𝑙-th antennas at the receiver 𝑦𝐿
o 𝑛𝑙 are independently random variables with distribution CN(0, 𝜎 2 )
o Vector representation of the channel model:
𝒚 = 𝒉𝑥 + 𝒏
o Here 𝒚 = 𝑦1 , 𝑦2 , … , 𝑦𝐿 𝑇 , 𝒉 = ℎ1 , ℎ2 , … , ℎ𝐿 𝑇 and 𝒏 = 𝑛1 , 𝑛2 , … , 𝑛𝐿 𝑇
o SNR:
𝒘𝐻 𝒉 2 𝑃
SNR = 2
𝜎 𝒘 𝟐
o By applying Cauchy-Schwartz inequality ( 𝒖𝑇 𝒗 ≤ ‖𝒖‖‖𝒗‖) in the numerator, we get
𝒘𝐻 𝒉 2 𝑃 𝒉 𝟐𝑃
SNR = 2 2 ≤
𝜎 𝑤 𝜎2
o Equality is met when 𝒘 = 𝑐𝒉, for some constant 𝑐
o 𝒘 = 𝒉→ This receive beamforming is called maximal ratio combining
SIMO Gaussian channel ℎ1
𝑦1
.
Transmitter .
ℎ𝐿 .
o Overall system after combining becomes 𝑥 Receiver
෨ + 𝑛
𝑦 = ℎ𝑥 𝑦𝐿
o Here 𝑦 = 𝒉𝐻 𝒚, ℎ෨ = 𝒉 𝟐 and 𝑛 = 𝒘𝐻 𝒏
o SNR becomes
𝒉 𝟐𝑃
SNR =
𝜎2
o And the capacity becomes
‖𝒉‖2 𝑃
𝐶 = log 2 1 +
𝜎2
o 𝐿 receive antennas provide beamforming/power gain
ℎ1
MISO Gaussian channel 𝑥1
.
.
. Receiver
Transmitter ℎ𝐿
o Consider a system with 𝐿-antenna transmitter 𝑦
o And single-antenna receiver
o Transmit single symbol 𝑥 over multiple antennas 𝑥𝐿
o Multiply the symbol 𝑥 with an 𝐿-length vector 𝒘→ 𝒙 = 𝒘𝑥
o Here 𝒘 = 𝑤1 , 𝑤2 , … , 𝑤𝐿 𝑇 and 𝒙
= 𝑥1 , 𝑥2 , … , 𝑥𝐿 𝑇 = 𝑤1 𝑥, 𝑤2 𝑥, … , 𝑤𝐿 𝑥 𝑇
o Received signal
𝑦 = 𝒉𝑇 𝒙
+𝑛
= 𝒉𝑇 𝒘𝑥 + ณ 𝑛
desired signal noise
o Noise 𝑛 ∼ 𝐶𝑁(0, 𝜎 2 )
o Channels gains are constant and known at both transmitter and receiver
o Calculate the SNR and maximize it
o Signal power 𝔼 𝒉𝑇 𝒘𝑥 2 = 𝒉𝑇 𝒘 2 𝑃
o Noise power 𝔼 𝑛 2 = 𝜎 2 ℎ1
o Apply Cauchy Schwartz inequality to get
𝑇 2 𝑥1
𝒉 𝒘 𝑃 .
SNR = .
𝜎2 . Receiver
ℎ 2 𝑤 2𝑃 Transmitter ℎ𝐿
≤ 𝑦
𝜎2
o Equality holds when 𝒘 = 𝑐𝒉
1 𝑥𝐿
o To make the transmit power constrained to 𝑃, 𝑐 = h
MISO Gaussian channel
Problem: Design 𝒘 such that the SNR is maximized OR capacity is achieved
ℎ
𝑤=
ℎ
o This transmit beamforming is called maximal ratio transmission
o Capacity
‖𝒉‖2 𝑃
𝐶 = log 2 1 +
𝜎2
o 𝐿 transmit antennas also provide beamforming/power gain
ℎ1
𝑥1
.
.
. Receiver
Transmitter ℎ𝐿 𝑦
𝑥𝐿
Lecture 6
Reference:
Chapter 5
Fundamentals of Wireless communication
David Tse, Pramod Vishwanath
Recap
o Capacity of AWGN channel
• SISO system
• SIMO system
• MISO system
In today’s class
o Fading channels
• Slow fading
• Fast fading
o Outage capacity for fading channels
• Slow fading SISO channels
SIMO Gaussian channel MISO Gaussian channel
ℎ1
ℎ1
𝑦1
. 𝑥1
. .
Transmitter ℎ𝐿 .
. Receiver
𝑥 Receiver Transmitter .
ℎ𝐿 𝑦
𝑦𝐿
𝑥𝐿
𝒉
o 𝒘= → maximal ratio transmission
o 𝒘 = 𝒉→ maximal ratio combining 𝒉
o Capacity of SIMO Gaussian channel o Capacity of MISO Gaussian channel
‖𝒉‖2 𝑃 ‖𝒉‖2 𝑃
𝐶 = log 2 1 + 𝐶 = log 2 1 +
𝜎2 𝜎2
o 𝐿 receive antennas provide beamforming/power gain o 𝐿 transmit antennas also provide
beamforming/power gain
Fading channels
o Channel varies over time
o Linear time variant channel
o SISO system model at instant 𝑚:
𝑦 𝑚 =ℎ 𝑚 𝑥 𝑚 +𝑛 𝑚
o Two kinds of fading:
o Slow fading
o fast fading
o Defined based on the relation between coherence time and symbol duration
o Coherence time is the time duration over which the channel can be considered approximately constant
o If the symbol duration is larger than the coherence time, the symbol will see the changing channel over its
transmission- fading of the channel is fast
o Coherence time ≫ symbol duration → Slow fading
o Coherence time ≪ symbol duration → Fast fading
o At 𝜖 = 0.01, 𝐶𝜖 is 14% of 𝐶AWGN , unlike the SISO case where it was just 1%.
Why are more antennas helping?
o Plot shows the PDF of 𝒉 2 for different values of 𝐿
o Notice that as 𝐿 increases the pdf around 𝒉 2 = 0 decreases
o It is less probable to have all the channels bad when 𝐿 is high
Lecture 7
Reference:
Chapter 5
Fundamentals of Wireless communication
David Tse, Pramod Vishwanath
Recap
o Fading channels
• Slow fading
• Fast fading
o Outage capacity for fading channels
• Slow fading SISO channels
• Slow fading SIMO channels
In today’s class
o Outage capacity for fading channels
• Slow fading MISO channels
o Capacity of fast fading SISO, SIMO and MISO channels
Slow fading SISO and SIMO channels
o Transmitter can send the data at a rate 𝑅 > 𝐶
o If the channel is strong→ communication can happen
o If the channel is weak → Outage event occurs
o Outage probability when operating at a rate 𝑅
𝑃out 𝑅 = Pr 𝑅 > 𝐶
o Outage capacity: Largest transmission rate such that any rate 𝑅 < 𝐶𝜖 follows 𝑃out 𝑅 < 𝜖
𝑃out 𝐶𝜖 = 𝜖
o Here ℎ is the random quantity, not known at the transmitter
o Here 𝑥1 and 𝑥2 have power constraint 𝑃/2 for total power constraint to be 𝑃
o Channel is assumed to be constant over two symbol durations
Alamouti codes for 2x1 MISO channel
o System model:
𝑥1 −𝑥2∗
𝑦1 𝑦 2 = ℎ1 ℎ2 + 𝑛1 𝑛2
𝑥2 𝑥1∗
o It can rewritten as
𝑦1 ℎ1 ℎ2 𝑥1 𝑛[1]
= ∗ ∗ ต +
𝑦2 ∗ ℎ2 − ℎ1 𝑥2 𝑛2∗
𝒚 𝑯 𝒙 𝒏
⇒ 𝒚 = 𝑯𝒙 + 𝒏
ℎ1 2 + ℎ2 2 𝟐 𝑃 ℎ1 2
+ ℎ2 2
𝑃 𝒉 2 𝑆𝑁𝑅
SNR = = =
ℎ1 2 + ℎ2 2 2𝜎 2 2𝜎 2 2
𝒉 2
o Maximum achievable rate or the capacity thus becomes: log 1 + SNR
2
o Outage probability for Alamouti scheme:
‖𝒉‖2
𝑃out 𝑅 = Pr log 1 + SNR < 𝑅
2
o When the transmitter has the channel knowledge, the outage probability:
𝑃out 𝑅 = Pr log 1 + ‖𝒉‖2 SNR < 𝑅
o Power loss of factor 2, but same diversity gain
o Alamouti scheme radiates energy in isotropic manner: all the signals have same energy
o Outage probability for 𝐿 transmit antennas:
‖𝒉‖2
𝑃out 𝑅 = Pr log 1 + SNR < 𝑅
𝐿
Fast fading channels
o Coherence time ≪ symbol duration → Fast fading
o SISO system model at instant 𝑚:
𝑦 𝑚 =ℎ 𝑚 𝑥 𝑚 +𝑛 𝑚
log 1 + ℎ𝑙 2 SNR
𝑙=1
o Average rate
𝐿
1
log 1 + ℎ𝑙 2 SNR
𝐿
𝑙=1
o Outage probability
𝐿
1
𝑃out 𝑅 = Pr log 1 + ℎ𝑙 2 SNR < 𝑅
𝐿
𝑙=1
Fast fading channels
o What if the number of blocks in the block fading model goes to infinity 𝐿 → ∞
o The capacity, using law of large number reduces to
𝐿
1
log 1 + ℎ𝑙 2 SNR → 𝔼 log 1 + ℎ 2 SNR
𝐿
𝑙=1
o The randomness due to channel has been removed by taking the expectation
o Capacity for a SISO fast fading channel
𝐶𝐹𝐹 = 𝔼 log 1 + ℎ 2 SNR
o Transmitter needs to know the statistics of the channel, and not the exact channel!
o Similarly the capacity of fast fading SIMO channel (L receive antennas) and MISO channel (L transmit antennas)
𝐶𝐹𝐹 = 𝔼 log 1 + ‖𝒉‖2 SNR
o Whole codeword needs to be of length 𝐿𝑇𝑆 where 𝐿 and 𝑇𝑆 both are large
Fast fading channels
o AWGN capacity for SISO channel 𝐶AWGN = log(1 + SNR)
o Here we assume 𝔼 ℎ 2
=1
o At low SNR:
𝔼 log 1 + ℎ 2 SNR ≈𝔼 ℎ 2
SNR log 𝑒 = 𝐶AWGN
o At high SNR:
𝔼 log 1 + ℎ 2 SNR ≈ 𝔼 log ℎ 2 SNR = 𝐶AWGN + 𝔼[log ℎ 2 ]
o Constant difference term can be further bounded as 𝔼[log ℎ 2 ] ≤ log 𝔼 ℎ 2 = 0
Fast fading channels
Reference:
Chapter 7
Fundamentals of Wireless communication
David Tse, Pramod Vishwanath
Recap
o Fading channels
• Slow fading
• Fast fading
o Outage capacity for slow fading SISO, SIMO and MISO channels
o Capacity of fast fading SISO, SIMO and MISO channels
In today’s class
o MIMO channel representation
o MIMO channel characteristics
o Capacity MIMO channels
▪ Deterministic channel
▪ Full CSI
▪ CSI at receiver only
Summary: capacity of wireless channels
Deterministic channel with channel information at transmitter and receiver
MISO 𝑦 = 𝒉𝑇 𝒙
+𝑛 𝐶 = log(1 + ‖𝒉‖2 SNR)
Summary: capacity of wireless channels
Slow fading channel without channel information at transmitter
MISO 𝑦 = 𝒉𝑇 𝒙
+𝑛 ‖𝒉‖2 𝐶𝜖
Pr 𝑅 > log 1 + SNR SNR
2 = log 1 + × 𝐹 −1 1 − 𝜖
2
Summary: capacity of wireless channels
Fast fading channel without channel information at transmitter
MISO 𝑦 = 𝒉𝑇 𝒙
+𝑛 𝐶 = 𝔼𝒉 log(1 + ‖𝒉‖2 SNR)
MIMO channels
o Multiple antennas at the Tx or Rx give power and diversity gain
o Can we increase the capacity linearly with number of antennas?
o Multiple Tx and Rx antennas help us achieve that
MIMO channels
ℎ11
o Consider a system with 𝑛𝑡 -antenna transmiiter and ℎ𝑛𝑟 1
𝑦1
𝑛𝑟 -antenna receiver 𝑥1
o Let the data transmitted over 𝑛𝑡 antennas is given . .
by: 𝒙 = 𝑾𝑡 𝒙 . .
ℎ1𝑛𝑡
. Receiver
o Here, 𝒙 ∈ ℂ𝑛𝑡 ×1 data to be transmitted and 𝑾𝑡 ∈ Transmitter .
ℎ𝑛𝑟 𝑛𝑡
ℂ𝑛𝑡×𝑛𝑡 is the beamforming matrix
o Transmit signal:
𝑇 𝑥𝑛𝑡 𝑦𝑛𝑟
= 𝑥1 , 𝑥2 , … , 𝑥𝑛𝑡
𝒙
𝑛𝑡
o Here each 𝑥𝑖 = σ𝑗=1 𝑤𝑖𝑗 𝑥𝑗
o Receive signal at antenna 1:
𝑦1 = ℎ11 𝑥1 + ℎ12 𝑥2 + ⋯ + ℎ1𝑛𝑡 𝑥𝑛𝑡 + 𝑛1
o Receive signal at antenna 2:
𝑦2 = ℎ21 𝑥1 + ℎ22 𝑥2 + ⋯ + ℎ2𝑛𝑡 𝑥𝑛𝑡 + 𝑛2
o Receive signal at antenna 𝑛𝑟 :
𝑦𝑛𝑟 = ℎ𝑛𝑟 1 𝑥1 + ℎ𝑛𝑟 2 𝑥2 + ⋯ + ℎ𝑛𝑟 𝑛𝑡 𝑥𝑛𝑡 + 𝑛𝑛𝑟
MIMO channels
o Complete receive signal in vector representation
𝑦1 ℎ11 𝑥1 + ℎ12 𝑥2 + ⋯ + ℎ1𝑛𝑡 𝑥𝑛𝑡 𝑛1
𝑦2 ℎ21 𝑥1 + ℎ22 𝑥2 + ⋯ + ℎ2𝑛𝑡 𝑥𝑛𝑡 𝑛2
… = …
+ …
𝑦𝑛𝑟 ℎ𝑛𝑟 1 𝑥1 + ℎ𝑛𝑟 2 𝑥2 + ⋯ + ℎ𝑛𝑟 𝑛𝑡 𝑥𝑛𝑡 𝑛𝑛𝑟
o Can be re-written as
𝑦1 ℎ11 ℎ12 ⋯ ℎ1𝑛𝑡 𝑥1 𝑛1
𝑦2 ℎ21 ℎ22 ⋯ ℎ2𝑛𝑡 𝑥2 𝑛2
… = ⋮ ⋮ ⋱ ⋮ … + …
𝑦𝑛𝑟 ℎ𝑛 1 ℎ𝑛 2 ⋯ ℎ𝑛 𝑛 𝑥𝑛𝑡 𝑛𝑛𝑡
𝑟 𝑟 𝑟 𝑡
𝒚
𝒙 𝒏
𝑯
⇒ 𝒚 = 𝑯
𝒙+𝒏
o Here, 𝒚, 𝒏 ∈ ℂ𝑛𝑟 ×1 , 𝒙
∈ ℂ𝑛𝑡×1 and 𝑯 ∈ ℂ𝑛𝑟 ×𝑛𝑡
o Noise 𝒏 ∼ 𝒞𝒩(𝟎, 𝚺), where 𝚺 = 𝔼 𝒏𝒏𝐻 = 𝜎 2 𝑰 ∈ ℂ𝑛𝑟 ×𝑛𝑟 is the noise covariance matrix
o At the receiver end, combining is also performed using 𝑾𝒓 ∈ ℂ𝑛𝑟 ×𝑛𝑟
𝑛𝑟 𝑛𝑟
o Also𝜆1 ≥ 𝜆2 ≥ ⋯ ≥ 𝜆𝑛min × 𝑛𝑟 × (𝑛𝑡 −𝑛𝑟 )
MIMO channels
o For the channel
ℎ11 ℎ12 ⋯ ℎ1𝑛𝑡
ℎ21 ℎ22 ⋯ ℎ2𝑛𝑡
𝑯=
⋮ ⋮ ⋱ ⋮
ℎ𝑛𝑟 1 ℎ𝑛𝑟 2 ⋯ ℎ𝑛𝑟 𝑛𝑡
o SVD decomposition: 𝑯 = 𝑼𝚲𝐕 H
o very useful tool if the channel is known at both transmitter and receiver
MIMO channels
o For the channel
ℎ11 ℎ12 ⋯ ℎ1𝑛𝑡
ℎ21 ℎ22 ⋯ ℎ2𝑛𝑡
𝑯=
⋮ ⋮ ⋱ ⋮
ℎ𝑛𝑟 1 ℎ𝑛𝑟 2 ⋯ ℎ𝑛𝑟 𝑛𝑡
o SVD decomposition: 𝑯 = 𝑼𝚲𝐕 H
o very useful tool if the channel is known at both transmitter and receiver
𝑦1
𝑥1
. . .
. . .
. Receiver
Transmitter . .
𝑥𝑛𝑡 𝑦𝑛𝑟
MIMO channels
o For the channel
ℎ11 ℎ12 ⋯ ℎ1𝑛𝑡
ℎ21 ℎ22 ⋯ ℎ2𝑛𝑡
𝑯=
⋮ ⋮ ⋱ ⋮
ℎ𝑛𝑟 1 ℎ𝑛𝑟 2 ⋯ ℎ𝑛𝑟 𝑛𝑡
o SVD decomposition: 𝑯 = 𝑼𝚲𝐕 H
o very useful tool if the channel is known at both transmitter and receiver
𝑦1
𝑥1
. . .
. . .
. Receiver
Transmitter . .
𝑥𝑛𝑡 𝑦𝑛𝑟
MIMO Gaussian channels
o We consider the channel deterministic in this case
o Assume it to be known at both transmitter and receiver (full CSI)
o As there are only at max 𝑛min distinct paths, we send the data with 𝑛min non-zeros
o The data vector 𝒙 = 𝑥1 , 𝑥2 , … , 𝑥𝑛𝑚𝑖𝑛 , 0, … , 0 ∈ ℂ𝑛𝑡 ×1 ⇒ (𝑛𝑡 − 𝑛min ) zeros
o Recall the data transmitted over 𝑛𝑡 antennas is given by: 𝒙 = 𝑾𝑡 𝒙
o Here, 𝑾𝑡 ∈ ℂ 𝑡 𝑡 is the beamforming/precoding matrix
𝑛 ×𝑛
o Received signal:
𝒚 = 𝑯 𝒙+𝒏
o Use the SVD decomposition of 𝑯
𝒚 = 𝑼𝚲𝐕 H 𝒙 + 𝒏 = 𝑼𝚲𝐕 H 𝑾𝑡 𝒙 + 𝒏
Reference:
Chapter 7
Fundamentals of Wireless communication
David Tse, Pramod Vishwanath
Recap
o MIMO channel representation
o MIMO channel characteristics
In today’s class
o Capacity MIMO channels
▪ Deterministic channel
▪ Full CSI
▪ CSI at receiver only
ℎ11
. .
o Consider a system with 𝑛𝑡 -antenna transmiiter and . ℎ1𝑛𝑡 .
𝑛𝑟 -antenna receiver . Receiver
Transmitter .
ℎ𝑛𝑟𝑛𝑡
o Let the data transmitted over 𝑛𝑡 antennas is given
by: 𝒙 = 𝑾𝑡 𝒙
o Here, 𝒙 ∈ ℂ𝑛𝑡 ×1 data to be transmitted and 𝑾𝑡 ∈ 𝑥𝑛𝑡 𝑦𝑛𝑟
ℂ𝑛𝑡×𝑛𝑡 is the beamforming matrix
o Complete receive signal in vector representation
𝑦1 ℎ11 ℎ12 ⋯ ℎ1𝑛𝑡 𝑥1 𝑛1
𝑦2 ℎ21 ℎ22 ⋯ ℎ2𝑛𝑡 𝑥2 𝑛2
… = ⋮ ⋮ ⋱ ⋮ … + …
𝑦𝑛𝑟 ℎ𝑛 𝑟 1 ℎ𝑛 𝑟 2 ⋯ ℎ𝑛 𝑟 𝑛 𝑡 𝑥𝑛𝑡 𝑛𝑛 𝑡
𝒚
𝒙 𝒏
𝑯
⇒ 𝒚 = 𝑯
𝒙+𝒏
o Here, 𝒚, 𝒏 ∈ ℂ ,𝒙
𝑛𝑟 ×1
∈ℂ and 𝑯 ∈ ℂ
𝑛𝑡 ×1 𝑛𝑟 ×𝑛𝑡
o Noise 𝒏 ∼ 𝒞𝒩(𝟎, 𝚺), where 𝚺 = 𝔼 𝒏𝒏𝐻 = 𝜎 2 𝑰 ∈ ℂ𝑛𝑟 ×𝑛𝑟 is the noise covariance matrix
o Received signals are combined using 𝑾𝑟 ∈ ℂ𝑛𝑟×𝑛𝑟 to get
𝑾𝐻 𝐻
𝑟 𝒚 = 𝑾𝑟 𝑯𝑾𝑡 𝒙 + 𝑾𝑟 𝒏
𝐻
MIMO channels
o SVD decomposition of the channel 𝑯
𝑯 = 𝑼𝚲𝐕 H
o It is defined for any matrix
o Matrices 𝑼 ∈ ℂ𝑛𝑟×𝑛𝑟 and 𝑽 ∈ ℂ𝑛𝑡 ×𝑛𝑡 are unitary matrices, i.e., 𝑼𝐻 𝑼 = 𝑼𝑼𝐻 = 𝑰𝑛𝑟 and 𝑽𝐻 𝑽 = 𝑽𝑽𝐻 = 𝑰𝑛𝑡
𝑛 ×𝑛
o Matrix 𝚲 ∈ ℝ+𝑟 𝑡 is a diagonal matrix, with ordered singular values of 𝑯 in its diagonals
o 𝑯 has 𝑛min = min 𝑛𝑟 , 𝑛𝑡 positive singular values
⇒𝒚 = 𝑾𝐻 H
𝑟 𝑼𝚲𝐕 𝑾𝑡 𝒙 + 𝒏
o By taking 𝑾𝑡 = 𝑽 and 𝑾𝒓 = 𝑼, we get
= 𝑼𝐻 𝑼𝚲𝐕 H 𝐕𝒙 + 𝒏
𝒚
⇒𝒚 = 𝚲𝒙 + 𝒏
o We already saw that 𝔼 𝒙 𝟐 = 𝔼 𝒙 𝟐 and 𝔼[ 𝐻 ] = 𝔼[𝒏𝒏𝐻 ] = 𝜎 2 𝑰
𝒏𝒏
MIMO Gaussian channels
o Combined received signal
= 𝚲𝒙 + 𝒏
𝒚
o Interestingly, this can be written as multiple (𝑛min to be precise) SISO channels
𝑦𝑖 = 𝜆𝑖 𝑥𝑖 + 𝑛 𝑖
o Here the index 𝑖 = 1, 2, … , 𝑛min
o Capacity of each of these channels is
𝜆2𝑖 𝑃𝑖
𝐶𝑖 = log 1 + 2
𝜎
o We used the fact that 𝔼 𝑥𝑖 = 𝑃𝑖
2
𝑛
o For constrained transmit power, we need to have σ𝑖 min 𝑃𝑖 ≤ 𝑃
o Recall the capacity C is achieved when 𝑥𝑖 is Gaussian distributed
o Total capacity: Sum of capacities of all these channels
𝑛min
𝜆2𝑖 𝑃𝑖
𝐶sum = log 1 + 2
𝜎
𝑖
o Still there is a scope to maximize the rate: optimal power allocation
o Optimal power allocation: Find 𝑃𝑖 for all 𝑖 such that 𝐶sum is maximized
MIMO Gaussian channels
o Optimal power allocation: Find 𝑃𝑖 for all 𝑖 such that 𝐶sum is maximized
max 𝐶sum
𝑃𝑖
𝑛min
𝜆2𝑖 𝑃𝑖
⇒ max log 1 + 2
𝑃1 ,𝑃2 ,…,𝑃𝑛𝑚𝑖𝑛 𝜎
𝑖
𝑛
Subject to σ𝑖 min 𝑃𝑖 = 𝑃
−3 2
1 0 52 0 0 0 1 0
𝑯= 2 3 0 0 13 0 1 0 0
13 0 0 13 0 2 0 0 1
0
o Singular values are square root of eigenvalues of 𝑯𝐻 𝑯, its eigenvectors constitute 𝑽 and each column of 𝑼 follows
𝟏
𝒖𝑖 = 𝑯𝒗𝑖
𝜆𝑖
𝜆2𝑖 𝑃𝑖
o Capacity: 𝐶sum = σ𝑛𝑖 min log 1+
𝜎2
𝜆12 𝑃1 𝜆22 𝑃2 𝜆23 𝑃3
𝐶sum = log 1 + 2 + log 1 + 2 + log 1 + 2
𝜎 𝜎 𝜎
52𝑃1 13𝑃2 4𝑃3
= log 1 + + log 1 + + log 1 +
2 2 2
o Such that 𝑃1 + 𝑃2 + 𝑃3 = 0.75
MIMO Gaussian channels
o Example: SVD-based precoder and combiner, and Waterfilling power allocation:
o Lets determine 𝜇
𝑛min
1 𝜎2
− =𝑃
𝜇 𝜆2𝑖
𝑖=1
1 2 1 2 1 2
⇒ − + − + − = 0.75
𝜇 52 𝜇 13 𝜇 4
1 25
⇒ = = 0.4808
𝜇 52
o First allocate the power to the weakest channel– channel 3
1 2 1
𝑃3 = − =− <0
𝜇 4 52
⇒ 𝑃3 = 0
o No power allocated to channel 3
o Recalculate the value of 𝜇
1 2 1 2
− + − = 0.75
𝜇 52 𝜇 13
1 49
⇒ = = 0.4712
𝜇 104
MIMO Gaussian channels
o Example: SVD-based precoder and combiner, and Waterfilling power allocation:
o Lets now allocate the power to the current weakest channel– channel 2
1 2 33
𝑃2 = − = = 0.3174 > 0
𝜇 13 104
⇒ 𝑃2 = 0.3174
o Similarly, the power of the strongest channel- channel 1 becomes
1 2 45
𝑃1 = − = = 0.4327 > 0
𝜇 52 104
⇒ 𝑃1 = 0.4327
MIMO Gaussian channels
o Example: SVD-based precoder and combiner, and Waterfilling power allocation:
MIMO Gaussian channels
o Example: SVD-based precoder and combiner, and Waterfilling power allocation:
+
1 𝜎2
−
𝜇 𝜆12
𝜎2
𝜆12
Waterfilling algorithm Unfair to weaker channels
o Set 𝑘 = 𝑛min
1. Calculate 𝜇 according to the following equation
𝑘
1 𝜎2
− 2 =𝑃
𝜇 𝜆𝑖
𝑖=1
2. Start with the channel with lowest SNR (i.e., lowest singular value)
3. Allocate power to the 𝑘-th channel using the value of 𝜇 obtained as follows
+
1 𝜎2
𝑃𝑘 = −
𝜇 𝜆2𝑘
Degrees of freedom
MIMO channels
o Capacity linearly increases with number of antennas→ degrees of freedom