Embedded Processors II Overview
Embedded Processors II Overview
YOAB/YEAB
ZOAB/ZEAB
DAAU: Data Address Arithmetic Unit
XOAB/XEAB
PCU: Program Control Unit
OFU: Operand Fetch Unit
XODB/XEDB
YODB/YEDB
ZODB/ZEDB
CBU: Computation and Bit –Manipulation Unit
IDU: Instruction Decode Unit
BFO: Bit Field Operation DAAU
User- Pointers, SW stack
Defined Alternative Bank of Registers
Registers Bit Reversal
CBU
Barrel MUL0 MUL1 BI
EXP Shifter NMI
INT0
INT1
BFO ALU IDU PCU INT2
VINT
BMU CU IVA
OFU
y0 x0 x1 y1
Bus Align
sign &
zero ext. y x x y
MUL0 MUL1
Barrel p0 p1
EXP SV Shifter
Scaling Shifter Scaling Shifter
const const
ALU
Saturation Unit
a0 a1 b0 b1
Saturation Unit
General-purpose
General-purpose processors
processors areare
not
not fast
fast enough
enough for
for data-intensive
data-intensive
applications,
applications, don’t
don’t have
have enough
enough I/O
I/O or
or
compute
compute bandwidth
bandwidth
RTL
RTL –– often
often aa good
good choice:
choice:
•• High
High performance
performance due due
to
to parallelism
parallelism
•• Large
Large number
number of of wires
wires
General ROM
A/D in
in // out
out of
of the
the block
block
Purpose
•• Languages
Languages and and tools
tools
32b CPU RAM familiar
familiar to to many
many
But
But … …
I/O
•• Slow
Slow to to design
design andand verify
verify
Hardwired •• Inflexible
Inflexible after
after tapeout
tapeout
Logic •• High
PHY High re-spin
re-spin risk
risk and
and cost
cost
•• Slows
Slows timetime to
to market
market
General
General Control RAM A/D
Control A/D Silicon RISC
RAM
RISC scaling
Controller Image Video Video
Logic Logic Logic
Data Processing: I/O I/O
Image, video, audio, End-product Audio Packet
packet processing, Logic Logic
complexity
security, or DSP
Hardwired Logic PHY Security DSP
PHY
Logic Logic
I/O
Audio Packet
Processor Processor
Security DSP
Processor Processor PHY
Performance and Power Savings
Dedicated
Hardware
Reconfigurable
Hardware
Application
Specific I/S
Processors
Domain-
Specific
Processors
General-
Purpose
Processors
Flexibility
Software
Extensions Extensions
Base Base Base
1. 2. 3. 4. 5.
Processor
Hardware Design:
RTL, scripts, test
bench
ALU OCD
High-level Tensilica
specification: Processor
Tensilica Instruction Generator
Extension (TIE)
Customized
Software Tools:
Compiler,
Design processor in one hour debuggers,
simulators, RTOS
Processor Controls Instruction Fetch / PC Instruction RAM
Trace TRACE Port Instruction ROM
JTAG JTAG Tap Control Extended Instruction Instruction Instruction Instruction
Align, Decode, Decode/Dispatch MMU Cache
On Chip Debug
Dispatch
Exception Support
User User Base Register External Interface
Exception Handling Defined Defined File
Registers Register Register Base ALU Xtensa
Data Address Files Files
Processor
Watch Registers MAC 16 DSP
Interface PIF
Instruction Address User User MUL 16/32 Control
Watch Registers Defined Defined
Floating Point
Interrupt Control Execution Execution Write
Interrupts Units and Units
Timers User Defined Buffer
Interfaces
Execution Unit
User Defined
Queues and Wires Vectra Data Data
Vectra DSP Vectra DSP
DSP MMU Cache
Used Defined Data Data
Data ROMs
Load/Store Units Load/Store
Base ISA Feature
Unit Data RAMs
Configurable Function
Xtensa
Optional Function
Local
Optional & Configurable Memory
User Defined Features (TIE) Interface
2.0
2.0 0.5 0.473 0.14
1.2 0.3
0.08
1.0 0.3 0.23
0.06
0.8 0.2
300000
243968
35
Code Bytes
20
15
100000
70328 10
59648 57149
5
0
0
optimized out-of-box optimized out-of-box
Trimedia TM1300 optimized Trimedia TM1300 out-of-box Xtensa Xtensa Trimedia TM1300 Trimedia TM1300
ARM1020E Xtensa optimized
Xtensa out-of-box core area code area
100
Xtensa Telecom
Xtensa Network
Xtensa Consumer
Xtensa Telecom OPT
EEMBCPerformance/MHz
50 iffe
x re
d
en nc
er e
gy
1
0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8
Evaluates millions
of possible Tuned Software
extensions: Tools
• SIMD operations
• operator fusion
• parallel execution
Option: manually refine
configuration
status
Mechanism
reg
CPU
data
reg
status
(8 bit)
CPU xmit/
8251
rcv
data serial
(8 bit) port
while (TRUE) {
/* read */
while (peek(IN_STATUS) == 0);
achar = (char)peek(IN_DATA);
/* write */
poke(OUT_DATA,achar);
while (peek(OUT_STATUS) != 0);
}
intr request
status
Mechanism
reg
intr ack
PC
IR
CPU
data/address data
reg
device 1 device 2 device n
interrupt
acknowledge
L1 L2 .. Ln
CPU
Vector
Device
interrupt interrupt
request acknowledge
CPU
Interrupt
vector handler 0
table head handler 1
handler 2
handler 3
asel
RDATA
ARM Memory
WDATA system
csel bsel
CPDIN CPDOUT
Coprocessor
data
address cache
SRAM
CPU ~MB main
Controller
~10ns
cache
register memory
CMOS DRAM
<KB address ~GB
~1ns ~100ns
data
data
logical physical
address memory address
main
CPU management
memory
unit
page 1
page 2
segment 1
memory
segment 2
segment base address logical address
linear/physical address
page offset
page i base
concatenate
page offset
page
descriptor
page descriptor
flat tree
Translation table
1st index 2nd index offset
base register
descriptor concatenate
1st level table
descriptor concatenate
2nd level table
physical address
accelerator
CPU
memory
I/O
cost
performance
Data input Data output
Accelerated
computation
M1 M2
P1 5 5 ?
P2 5 6
P3 - 5
Example process execution times
P1 P2
d1=2 d2=4
P3
Task graph
M1 M2
P1 5 5
P2 5 6
P3 - 5
M1 P2 P1 P2
Time = 14
d1=2 d2=4
M2 P1 P3 P3
network d2
5 10 15 20
time
M1 M2
P1 5 5
P2 5 6
P3 - 5
M1 P1 P1 P2
Time = 12
d1=2 d2=4
M2 P2 P3
P3
network d1
5 10 15 20
time
Function
Function units
units Memories
Memories I/O
I/O ports
ports
Controller
… …
interconnect
Registers
Registers
interconnect
Interconnect
Interconnect Controller
Controller
inport rom
outport ram_1 ram_2 alu_1 alu_2
acu_1 acu_2 acu_3
DesignDE DEvelop
µA
set target
#include …
.tmp
run / profile main()
{
001010
create_resource
save load … = dct();
ISS 110011
010110
}
LIFETIME
LIFETIME
110111
LOAD
LOAD
DEvelop
+ +
C Source 2 3 Analysis feedback
Description INPUT
0 * +
1 2
* + +
1 *
4
1 2 3
2 * +
3 6
* *
4 5 3 * +
8 7
+ +
7 6 4 *
5
*
8 *
1
5 * +
9 10 Micro-
OUTPUT
code
Group 1
Group 3 Group 4
6
Speedup
0
3Des AES Blowfish Md5 Rc4 SHA
10
7
Speedup
0
0 2 4 6 8 10
CFU Cost Budget (Relative to 32-bit ripple-carry adder)
Input 1 Input 2
+ Input 3
RF RF RF RF + ^
… 0x1B 0x5 Input 3
Control Memory | ^
Output Output
CFU Reg. File and
OptimoDE = 5.5 mm2 in 0.13 µ 2% Interconnect
ARM 926EJ = 5.0 mm in 0.13 µ
2
9%
Control
Memory
11%
RF RF RF RF
…
Baseline
Design
78%
Control Memory
Input 1 0x8
0xFF
>>
Input 2 |
0xFF
>>
+
Input 2 |,&
Output
+,-
Output
Speedup
3d
0
1
2
3
4
5
6
7
3d es
es -3d
-B es
lo
3d wfi
e sh
3d s-R
es c4
3des
3d -AE
B e S
Bl low s-S
ow fis ha
fis h-3
h d
Bl -Blo es
ow w
Bl fis fish
ow h-
R
Bl fish c4
ow -A
fis ES
Blowfish
R h-S
R c4- ha
c4 3d
-B es
lo
R wfis
c4 h
R -Rc
c4 4
-
R AE
Rc4
c4 S
AE -S
AE S- ha
S- 3d
Bl es
o
AE wfi
S sh
AE -R
S - c4
AES
AE AE
S S
Key: application run – application designed for
Sh -Sh
Sh a- a
a- 3de
Bl s
o
Sh wfis
a h
Sh -Rc
Base
Sha
a- 4
Sh AE
a- S
Generalized
Sh
a