Conception des systèmes sur puce
Conception des systèmes sur puce
Tronc Commun
• les systèmes sur puce (SoC)
De l'algorithme au système sur puce o architecture, principaux composants, bus
o outils de conception système, compilation logicielle
Méthodologies, applications et perspectives o métriques (performance, énergie, coût)
• les nouvelles architectures des DSP et FPGA
Olivier Sentieys
saurez modéliser un algorithme (signal) par un graphe
IRISA • métriques, transformations et optimisation
ENSSAT - Université de Rennes 1 saurez concevoir un composant ou un processeur spécialisé
ISE
depuis l'algorithme (notion de synthèse d’architecture)
saurez concevoir et optimiser du code sur une architecture
spécialisée
EII3/M2R - 2
Silicon Atom
Software!
Mask set is few M$US
EII3/M2R - 7 EII3/M2R - 8
[ITRS2002]
[© R. Rutenbar, CMU]
EII3/M2R - 11 EII3/M2R - 12
EII3/M2R - 15 EII3/M2R - 16
Ère post PC
2. Évolutions des applications
EII3/M2R - 18
Log Complexity
3G Cellular generations
EDGE/GPRS
10km
3GPP-LTE
GSM
UMTS
2.5G
1km 3G 4G Bit/nJ
Processor' Performance
WiMax
802.16a
Moore’s Law
100m
DECT
Bluetooth 802.11n/b 2G
10m ZigBee WLAN
802.11g/a
0
ISDN/ADSL ATM, SONET, … Battery Capacity
1G
Data Rate
10kbs 2Mbs 100Mbs
1982 1992 2002 2012 Time
EII3/M2R - 19 EII3/M2R - 20
Interface
MPEG4 TDMA Turbo/ Smart
MP3/AC3 W-CDMA Viterbi Antennas
Internet access Codes
500mW @ 6 GOPS
Image 12 GIPS/W @ 6 GOPS
Multiple Channel Demodul. RF
Demult. Access Decoder Filter
Equalizer
Voice
Avec les processeurs actuels
Source
Decoder 30 Kg ou 10 minutes !!!
EII3/M2R - 23 EII3/M2R - 24
EII3/M2R - 25 EII3/M2R - 26
WWW
Services
Monitoring et contrôle
(W)LAN
Identification et sécurité Moteur : Gestion du moteur, Boîte de vitesses automatique, Contrôle d’embrayage, 4WD
Température Châssis : ABS/ASR/DSC, Suspension, 4WS.
Réseaux multimédia
Réseaux de données
?? Sécurité : Air Bag, Prétensionneur, Système anti-collisions, Croisière.
Sécurité : Alarmes diverses, Fermeture avec ou sans clés
Agrément : Vitres, Sièges, Miroir, Chauffage, ...
Wifi, ZigBee, UWB
Instrumentation : Affichage, Navigation, GPS, Audio, Téléphone, CAN.
EII3/M2R - 27 EII3/M2R - 28
Engine Power
• Volumes importants, bas coûts, haute fiabilité, peu de Speed Electronicxs
maintenance, haute qualité, temps de mise sur le marché
court, contraintes physiques importantes (poids, taille). Actuators
• Real Time
• DSP + MCU
EII3/M2R - 29 EII3/M2R - 30
EII3/M2R - 31 EII3/M2R - 32
DSP core
quality recognition
24 sqmm enhancement
RAM
Multiplier A digital
down
image
decoder
speech • Slow processing
conv coder
D decoder
IP
Analog DSP core Memory
EII3/M2R - 35 EII3/M2R - 36 On-chip bus
Old GSM
EII3/M2R - 39 EII3/M2R - 40
EII3/M2R - 41 EII3/M2R - 42
Ex. 2: Network Processor IXP1200 Intel Ex. 3: Set Top Box STb STMicro.
®
EII3/M2R - 43 EII3/M2R - 44
SPDIF
2x16 2x16
Hub metadata read from stream
DDR2
256
DDR2
256
• Host CPU is performing playback control only:
EII3/M2R - 45 Mbytes Mbytes EII3/M2R - 46 o navigation, parsing, streaming,…
STB Product
(65nm LP 7ML) Block Diagram
Clkgen A DDR2 150Mtransistors
886 pads 50µm stag.
SATA Codec
73 initiators+96 targets
115 propagated clocks
dsp
(19 for interconnect)
MTP Content:
36 soft IPs
2 hard blocks
HDMI 16 analog IPs
19 IOLIBs
Clkgen Audio 29 internal blocks/glues
RFDAC B C DAC VideoDAC 140 memory cuts
Host: 500DMIPs
32 32
4 x TS Input
Or 1394 out USB Disk Drives 16
Peripherals
EII3/M2R - 50
EII3/M2R - 51 EII3/M2R - 52
A/D
Interconnect Bus
D/A IP
Modèles P=>S
I talk PCI
réutilisables S=>P µP RAM ROM
Paramètres OK, let’s and AMBA
+ Interfaces talk PCI DSP1 DSP2 CPU
DMA DSP
AMBA
Exemples
OK, let’s • TI's OMAP, Philips' Nexperia, Intel's PCA (Personal
talk VCI Wrapper Wrapper Wrapper Internet Communications Architecture), Infineon'
VCI
Bluetooth, Mgold (3G), ...
1981
1983
1985
1987
1989
1991
1993
1995
1997
1999
2001
2003
2005
2007
2009
[SIA 97]
EII3/M2R - 58
Processeur
+ +
Image Image ARM
Memory Codeur Vidéo Memory
Motion Motion
Estimation Estimation
Constraints
Simulation, Verification
Time
Hardware / Software
Cost • Cohérences des descriptions à tous les niveaux
Power Partitionning d'abstraction
Test
Programmable Hardware
Reliability Processors Accelerators • SystemC ?
Software Hardware • Preuve d'une spécification de bas niveau, par rapport à la
Algorithm i Algorithm j
spécification initiale
C code VHDL/C code
RTL/HLS Exploration de différents modèles et découpages H/S
correspondant aux spécifications initiales
Software Hardware Library • Notion de partitionnement H/S
Compilation Synthesis
Co-simulation et co-vérification
EII3/M2R - 63 DSP IP EII3/M2R - 64
Synthèse architecturale
5. Evolution des méthodologies
Détail en cours option ISE
ENTITY fir IS!
! PORT (xn:IN INTEGER; yn:OUT INTEGER);!
END fir;! Simulation système
ARCHITECTURE behavioral OF fir IS!
Blocks
SystemC
& interfaces Matlab block Matlab block Matlab block
block
description
1b 3b 3b
4b
Test bench & verification (verification : both blocks (verification : both blocks
Test bench & verification
Process definition should give same results) should give same results)
Process definition
[Brodersen 2001]
EII3/M2R - 73 EII3/M2R - 74
C Abstraction levels
C
MCU O
DSP
• AL = Algorithm
M O
M o Prior to HW/SW partition
• TLM = Transaction-Level Model
o After HW/SW partition, models bit-true behavior, register bank, data
C transfers, system synchronisation; no timing needed.
O RAM C
Bus Bus O IP 2 • T-TLM = Timed TLM
model M model M o TLM + timing annotation, refined communication model
• BCA = Bus Cycle Accurate
o Models state at each clock edge
o e.g. Instruction Set Simulator (ISS) of a microprocessor
IP 1
C
O
SOC example • RT= Register Transfer
M • MCU : Microcontroler Unit o Synthesisable model
• DSP : Digital Signal Proc.
• IP : Hardware Block
[Courtesy of F. Rocheteau] EII3/M2R - 76
C C C C
MCU
ISS O MCU O
M O DSP M O DSP
M M
C C C C
Communication analysis Bus O RAM
TLM Bus O RTL
IP 2 Bus O RAM Bus O IP 2
TLM M BCA M
- Bus sizing model model M model model M
- Cache analysis
C Manage complexity C
- Early performance analysis Focus on functionality
Emulator
IP 1 O IP 1 O Simplified communication protocols
M - Mixed abstraction levels M
- Heterogeneous environment Throughput (no pagination, address generation)
Frequency
Size
[Courtesy of F. Rocheteau]
EII3/M2R - 81 EII3/M2R - 82
BCA
Memoire_Body
BCA
SYSTEM
EII3/M2R - 83 EII3/M2R - 84
Cœurs de processeur
6. Solutions architecturales Processeurs enfouis sur un SOC
Délivré sous licence, modulaire, bloc IP
1. Cœurs de processeur Caractérisation d’un cœur
• foundry-captive, licenciable (code RTL)
Contenu du cœur
Processeurs RISC • cœur (+ mémoire (+ périphériques ))
Processeurs configurables Exemples
Processeurs DSP • Infineon Carmel, Infineon TriCore
• ARM
• DSP Group OAK/PINE
e.g. ARM, TI, Xtensa, ST, … • ST D950, ST Lx
• TI C64x, C55x
EII3/M2R - 90
EII3/M2R - 91 EII3/M2R - 92
Performance Characteristics
ARM920T ARM920T ARM922T ARM922T
0.18µ 0.13µ 0.18µ 0.13µ
EII3/M2R - 93 EII3/M2R - 94
EII3/M2R - 95 EII3/M2R - 96
Texas Instruments
Very Long Instruction Word TMS 320C6x Series - VelociTI ‘C6200 CPU
Caractéristiques MPY
ADD
SHL
SUB
ADD
LDW
SUB
LDW
STW
B
STW
MVK
ADDK
NOP
B
NOP
Fetch
• Plusieurs instructions par cycle, empaquetées dans une MPY MPY ADD ADD STW STW ADDK NOP
32x8=256 bits
"super-instruction" large
• Architecture plus régulière, plus orthogonale, plus proche Dispatch Unit
du RISC
• Jeu de registres uniforme, plus large
L:ALU
Exemples
Functional Functional Functional Functional Functional Functional Functional Functional
Unit Unit Unit Unit Unit Unit Unit Unit
ST200
6. Solutions architecturales
21mm2
4.3. Architectures reconfigurables
fly" Voice
Architectures Source
Decoder
• FPGA, Data Path (DP), Processeurs
Granularité Coprocesseur
• Grain fin (porte logique, LUT), grain moyen (DP) Reconfigurable
temps
Reconfiguration
• Statique ou dynamique Processeur Processeur
1 CLB=4 SLICES
EII3/M2R - 118 EII3/M2R - 119
Virtex II Pro
DART DPR
Contrôleur de tâche
• Haute performance SB
Ctrl
DPR DPR
cluster1
maison
cluster2 • Faible consommation DPR
DPR DPR
o Peu d’instructions
Contrôleur mémoire
DPR DPR SB
DPR DPR Ctrl o Peu d’accès mémoire Ctrl
DPR DPR DPR
Ctrl
E/S Mem
DPR DPR SB
Mem
L1
maison
cluster3
DPR Cluster4
DPR L1
config. FPGA
DPR
Collaboration ENSSAT/STMicroelectronics L1
EII3/M2R - 130 EII3/M2R - 131
Segmented Network
$D1 $D2 $D3 $D4 Ctrl RDP3
Data.
802.11a (Channel Est.)
.. .. .. .. .. .. .. .. .. .. .. .. .. .. CB
RDP4
Mem.
.. .. .. .. .. .. .. .. .. .. .. .. .. ..
Ctrl
DMA
RDP5
.. .. .. .. .. .. .. .. .. .. .. .. .. .. Config
Mem.
FPGA
RDP6
STMicroelectronics
.. .. .. .. .. .. .. .. .. .. .. .. .. .. CEA LIST/LETI
.. .. .. .. .. .. .. .. .. .. .. .. .. ..
Loop Managment
RAC TX
Fresh Circuit (CEA)
4G mobile terminals ARM9
AHB RAM
Technology: ST 0.13µ
CPU core: ARM946
EST ETH
4.8 Mgates
RX
Chip area = 80 mm2
Package: TBGA 420
Core power supply: 1.2 V
DART
EII3/M2R - 137
Complex block FIR filter FIR filter that operates on a block of Modem channel equalisation Intel Pentium 200 MHz 23
complex data Analog Device ADSP 2106x 60 MHz 17
Real single-sample FIR filter FIR filter that operates on a single Speech processing, general filtering
Texas Instruments TMS320C67xx 167 MHz 65
sample of real data
Least-mean-square adaptive FIR filter LMS adaptive FIR filter that operates on Channel equalisation, servo control, linear predictive Texas Instruments TMS320C3x 80 MHz 9
Processeurs
ARM7 TDMI/picolo 70 MHz 14
Vector dot product Sum of the pointwise multiplication of Convolution, correlation, matrix multiplication,
two vectors multidimensional signal processing ARM7 TDMI 80 MHz 7
Vector add Pointwise addition of two vectors Graphics, combining audio signals or images, vector Motorola DSP 566xx 60 MHz 15
producing a third vector search Motorola DSP 563xx 100 MHz 25
Vector maximum Discovery of the value and location of a Error-control coding, algorithms using block floating-
Lucent technologies DSP 16xxx 100MHz 37
vector's maximum value point arithmetic
Convolutional encoder Application of convolutional forward North American digital cellular telephone equipment Lucent technologies DSP 16xx 120MHz 22
error-correction code to a block of bits (IS-54 standard) Analog device ADSP21xx 75 MHz 19
Finite-state machine (FSM) A contrived series of control operations Control operations appear in nearly all digital signal Texas Instruments TMS320C80 60 MHz 26
(test, branch, push, pop) and bit processing applications
Texas Instruments TMS320C54x 100 MHz 25
manipulations
256-point, radix-2, in-place fast Fourier FFt conversion of a normal time-domain Radar, MPEG audio compression, spectral analysis Texas Instruments TMS320C62x 200 MHz 99
Watt-microsecondes
Texas Instruments TMS320C44 Virgule Fixe
14
Texas Instruments TMS320C31 Virgule Flottante • 200 MHz, 1.8V
12
Texas Instruments TMS320C209
NEC µPD77015 10
DSP16210
Motorola DSP56166
8 • 100 MHz, 3.3V
Motorola DSP56002
IBM MDSP2780
6 ZSP16401
Lucent Technologies DSP 3207 4 • 200 MHz, 2.5V
Lucent Technologies DSP32C
01
60
49
10
1
1
Analog Devices ADSP-2171
C6701
20
70
64
11
C5
62
C6
C6
P1
-2
P1
0 10000 20000 30000 40000 50000 60000 70000
• 167 MHz, 1.8V
SP
ZS
DS
AD
Cost-Time Product (µs$)
EII3/M2R - 140 EII3/M2R - 141
mW/MIPS
I286 $200
Processeur MIPS MHz Vdd Pmoy MIPS/Watt I386 $300
1000
DSP1 $150 Pentium $500
CoolRisc 14 14 3V 2.8 mW 5000
Pentium MMX $700
TMS C54x 30-200 30-200 1.8V 460 mW 1500-3000 100
DSP32C $250
TMS C6x 1600 200 2.5V 2W 800 10 DSP16A $15
StrongARM 200 200 1.5V 420 mW 500 DSP16210 <$10
1 DSP1600 $10
a 21164 1200 300 3V 60 W 20
1980 1985 1990 1995 2000
a 21364 4000 1000 1.5V 100 W 40
EII3/M2R - 144 EII3/M2R - 145 [Ackland ISLPD98]
Energy Efficiency
Flexibility Processor/Logic 10-50 MOPS/mW
10
Run-time C B A C E E ASIPs
Flexibility DSPs 2 V DSP: 3 MOPS/mW
1
Top Speed D E B C A A SA110
Embedded Processors 0.4 MIPS/mW
0.1
Energy C D C B A A Flexibility (Coverage)
Efficiency
EII3/M2R - 146 EII3/M2R - 147
WCDMA WCDMA
Emetteur WCDMA Récepteur WCDMA
Filtre y’(n)
e’(n) x’(n)
DATA(n) x(n) Filtre y(n) s(n) Démodulation RIF
Modulation
RIF Rake Receiver DATA’(n)
WCDMA WCDMA
DSP C6x DSP C54x ASIC
Etude d’une solution architecturale pour le système émetteur-
récepteur FIR
• Pour les 3 blocs de filtre RIF complexe, Rake Receiver et modulation-
démodulation, indiquer :
o pour un ASIC, le nombre d’opérateurs nécessaires ;
o pour chaque DSP, le nombre de processeurs nécessaires, indiquer si la
contrainte de temps peut être respectée ; MOD
o pour les 3 architectures la puissance moyenne pour l’exécution du bloc.
Complexité des Fréquence Consommation
opérations Puissance
calcul
1 ASIC 0.18u ⊗ : 4ns ⊗ : 130pJ/opération RAKE
⊕ : 3 ns ⊕ : 40pJ/opération