0% found this document useful (0 votes)
11 views26 pages

22+1-Lane 6.4Gb/s RX Core in 65nm CMOS

The document presents a 22+1 lane source-synchronous RX core designed for high-speed serial links in server systems, achieving data rates of 3.2 to 6.4 Gb/s. It details the architecture, including a cleanup PLL, pulsed CDR for power reduction, and specifications such as power efficiency of 4.5mW/Gbps. The implementation utilizes 65nm CMOS technology and emphasizes low power consumption and area efficiency while maintaining high supply noise immunity.

Uploaded by

Saqib Shah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views26 pages

22+1-Lane 6.4Gb/s RX Core in 65nm CMOS

The document presents a 22+1 lane source-synchronous RX core designed for high-speed serial links in server systems, achieving data rates of 3.2 to 6.4 Gb/s. It details the architecture, including a cleanup PLL, pulsed CDR for power reduction, and specifications such as power efficiency of 4.5mW/Gbps. The implementation utilizes 65nm CMOS technology and emphasizes low power consumption and area efficiency while maintaining high supply noise immunity.

Uploaded by

Saqib Shah
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

A Multi-Standard 6.

4Gb/s 22+1-Lane
Source-Synchronous Link RX Core
with Optional Cleanup PLL
in 65nm CMOS

R. Reutemann 1, M. Ruegg 1, F. Keyser 2,


J. Bergkvist 2, D. Dreps 2, T. Toifl 3, M. Schmatz 3

Miromico 1 / IBM STG 2 / IBM Zurich Research Laboratory 3


Outline
▪ High speed serial links in server systems
▪ Source synchronous RX core overview
▪ Clocking
• Cleanup PLL and polyphase filter paths
▪ CDR operation
• Pulsed CDR mode for power reduction
▪ Implementation, key results

2
Application Area Overview
High-speed serial links in server systems:
• Many (10 – 20+) lanes per bus
• May have different north/south width  simplex cores
• Example standards: Intel QPI, HT3, proprietary
• Mostly source synchronous links for on-board busses

3
Application Area
Important parameters:
• Throughput (Gb/s per pin)
• Power (mW per Gb/s)
• Limited die area (µm2 per Gb/s, fit beneath C4 balls)
• Minimum latency (memory links, SMP links)
• Fairly good channels ...
• Less than 20dB attenuation at ½ baud rate
• Open eye standards (with FFE, but no DFE)
• … for current generation, will change
as speeds increase

4
RX PHY Core Overview
RX PHY core presented:
• 3.2 – 6.4 Gb/s per data lane
• Half-rate source synchronous clock
• 23 and 17-lane cores (22 or 16 data/spare lanes plus clock)
• RX data lanes implemented in pairs for circuit sharing

5
Clock Lane Macro
▪ Generate I/Q phases from half-rate link clock input
 parallel PLL and polyphase filter (PPF) path
▪ Half-rate I/Q clock generation from
100MHz ref clock for test modes
▪ Central reference/bias

6
Dual-Lane Data Channel Macro

7
Input Stage and Preamp
▪ Selectable (gnd/diff.) tuned termination
▪ T-coils for return loss and bandwidth
▪ Series switches for BIST and offset tuning
▪ Selectable 0 – 6dB peaking
▪ Global and per-latch offset correction

8
Test Support
▪ Macro BIST for high-speed RX path:
• PLL in multiplying mode provides half-rate clock
• Aux phase rotator drives generator/serializer

▪ Higher level test modes:


• LSSD, ASST support (clocking, boundary)
• (AC-) JTAG receiver, high-speed transparent mode
9
BIST/Cleanup PLL Architecture

10
PLL Core Implementation

11
PLL Settings and Results
▪ BIST mode (multiply by 16..32): loop BW ≈ 10MHz
▪ Cleanup mode (no multiplication, 1:1):
• PFD at half speed: loop BW 30..120MHz
• PFD at full speed: loop BW 50..250MHz
▪ Total PLL area: 0.032mm2
▪ Total PLL power: 41mW
▪ RMS jitter,
measured at test point (1/4 rate, single-ended):
• refclk 700 fs
• fbclk (BW 50 MHz) 800 fs
• fbclk (BW 90 MHz) 700 fs
• fbclk (BW 250 MHz) 680 fs

12
Excursion: Source Sync Link Background

Full CDR Link: Source Synchronous Link:


▪ clock recovered from ▪ clock shipped along
data at RX with data
▪ RX clock vs. data jitter ▪ RX clock vs. data jitter
mostly uncorrelated highly correlated
13
Source Sync Link Jitter Model

Simplified Data-Clock Jitter Model:


▪ no PLL:
D - C = NT(f) · | 1 - e-j 2π f ∆T |2 + NX(f)
▪ with PLL:
D - C = NT(f) · | 1 - e-j 2π f ∆T · HP(f) |2 + NX(f) · |HP(f)|2 + NP(f)

14
Crosstalk-Induced Jitter
-30 0 3
NRZ Data Spectrum
without PLL 
-35 -10 2.5
FEXT
-40 -20 2

rms jitter [ps]


dBV

1.5

dBV
-45 -30
FEXT*Data
-50 -40 1

-55 -50 0.5


FEXT*Data*PLL
-60 -60 0
0 2 4 6 0 500 1000
f [GHz] (linear scale) PLL bandwidth [MHz]

Crosstalk induced clock jitter component (example):


▪ 8 aggressors (data channels) on clock as victim
▪ cross-talk seen as additional jitter component on RX clock
▪ RX PLL removes a significant portion of this component

15
In-System PLL and PPF Results
System 1:
at ½ baud rate:
- attenuation 13dB
- signal/Xtalk 33dB

System 2:
at ½ baud rate:
- attenuation 12dB
- signal/Xtalk 21dB

16
System 1 Breakdown: The Culprit
▪ Performance breakdown for low loop-bandwidth case in
system 1 traced back to a reference clock buffer in this
system
▪ This buffer injects significant DJ at multiples of 16MHz
into the TX reference clock path
▪ At low loop-BW settings, RX PLL filters out some of this
fully correlated jitter, resulting in additional clock-vs-data
jitter at the
RX latch

17
CDR Loop Architecture
▪ E/L aggregation  quasi linear phase detector
▪ Adjustable data to edge sampling point offset
▪ Selectable CDR loop bandwidth (1 .. 600 kHz)

18
“Pulsed” CDR for Power Reduction

Pulsed CDR power reduction vs. fully-on CDR mode


(at 6.4Gb/s, per lane, edge path duty cycle 10%):
• Clock path: -2.3mW (27%)
• Data path, demux, CDR: -3.1mW (21%)
• Total savings: 0.85mW per Gb/s (15%)
19
Pulsed CDR: Subsampling Effects
▪ Jitter amplification due
to subsampling
▪ Mitigation:
• Reduce CDR loop
bandwidth (averaging)
• Dithering to trade-off
notch width vs. depth

20
Pulsed CDR Results
Sinusoidal jitter tolerance,
bit-true simulation model results and measured data:

21
RX Core Power Consumption
Power consumption (22+1 lanes, 6.4Gb/s, polyphase filter,
pulsed CDR mode with 10% on time):
• Total power: 635mW = 4.5mW/Gb/s
• Clock lane and distribution: 80mW = 0.6mW/Gb/s
• Per data lane: 25mW = 3.9mW/Gb/s

8%
5% Clock Lane
27%
Clock Distribution
20% IQ Restore, Phase Rotators
Data Lane Front-End
10% Demux, CDR C4
Logic, Support
30%
22
22+1 Lane Core Implementation
▪ 65nm bulk CMOS
technology
▪ Core area:
• 22+1 lane core 2.96 mm2
• clock lane 0.16 mm2
• dual data lane 0.17 mm2
▪ Development versions
characterized on link
testsite
▪ Two variants of core set
used in chips for different
server lines, links fully
operational
23
Selected Specifications and Results
Application Area Differential and Ground Terminated Link
Protocols, Including Industry Standard
and IBM Proprietary Memory Links
Data Rates 3.2, 4.8 – 6.4Gb/s
Selectable Termination 42.5Ω to Ground, 100Ω Differential
Return Loss @ 3.2GHz ≥ 10dB Differential Mode,
≥ 9dB Common Mode
Input-Referred RX Noise 1mV,rms
Process 65nm Bulk CMOS
Power: 22+1 Lane RX Core 635mW (4.5mW/Gbps)
Supply 1.0V Core, 1.05V I/O, 2.0V AVDD
ESD Robustness 2kV HBM, 500V CDM, 200V MM
Testability At-Speed RX BIST , LSSD/ASST
Support, AC-JTAG, Transparent Mode
24
Summary
▪ Product-level 22+1 lane RX core supporting
data rates of 3.2 – 6.4 Gb/s
▪ CML preamp and clock distribution with CMOS
latches and demux for low power and area at
high supply noise immunity
▪ Parallel poly-phase filter and cleanup PLL based
clock paths allow selection of optimum bandwidth
for system and direct comparison
▪ Optional pulsed CDR operation for lower power
▪ Total power efficiency of 4.5mW/Gbps with
nominal hardware in functional mode
25
Contact & References
▪ Contacts:
• Miromico AG, Robert Reutemann
(see [Link] )
• IBM Research - Zurich, Dr. Thomas Toifl
(see [Link] )

▪ References:
• ISSCC 2010, Session 8.3:
[Link]
• JSSC December 2010:
[Link]
• see references there for additional material

26

You might also like