0% found this document useful (0 votes)
10 views68 pages

Embedded Processors II Overview

Uploaded by

김동주
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views68 pages

Embedded Processors II Overview

Uploaded by

김동주
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Y space X space Z space

YOAB/YEAB

ZOAB/ZEAB
DAAU: Data Address Arithmetic Unit

XOAB/XEAB
PCU: Program Control Unit
OFU: Operand Fetch Unit

XODB/XEDB

YODB/YEDB

ZODB/ZEDB
CBU: Computation and Bit –Manipulation Unit
IDU: Instruction Decode Unit
BFO: Bit Field Operation DAAU
User- Pointers, SW stack
Defined Alternative Bank of Registers
Registers Bit Reversal

Status and OFU


Mode
Registers

CBU
Barrel MUL0 MUL1 BI
EXP Shifter NMI
INT0
INT1
BFO ALU IDU PCU INT2
VINT
BMU CU IVA

B0 Accumulator A0 Accumulator PAB


B1 Accumulator A1 Accumulator PDB
External
Xspace Yspace Zspace
Register

OFU

DAAU PCU CBU IDU


YDB
XDB/ZDB

y0 x0 x1 y1
Bus Align
sign &
zero ext. y x x y
MUL0 MUL1
Barrel p0 p1
EXP SV Shifter
Scaling Shifter Scaling Shifter
const const

ALU

Saturation Unit

a0 a1 b0 b1

Saturation Unit
General-purpose
General-purpose processors
processors areare
not
not fast
fast enough
enough for
for data-intensive
data-intensive
applications,
applications, don’t
don’t have
have enough
enough I/O
I/O or
or
compute
compute bandwidth
bandwidth
RTL
RTL –– often
often aa good
good choice:
choice:
•• High
High performance
performance due due
to
to parallelism
parallelism
•• Large
Large number
number of of wires
wires
General ROM
A/D in
in // out
out of
of the
the block
block
Purpose
•• Languages
Languages and and tools
tools
32b CPU RAM familiar
familiar to to many
many
But
But … …
I/O
•• Slow
Slow to to design
design andand verify
verify
Hardwired •• Inflexible
Inflexible after
after tapeout
tapeout
Logic •• High
PHY High re-spin
re-spin risk
risk and
and cost
cost
•• Slows
Slows timetime to
to market
market
General
General Control RAM A/D
Control A/D Silicon RISC
RAM
RISC scaling
Controller Image Video Video
Logic Logic Logic
Data Processing: I/O I/O
Image, video, audio, End-product Audio Packet
packet processing, Logic Logic
complexity
security, or DSP
Hardwired Logic PHY Security DSP
PHY
Logic Logic

Traditional Application-Specific IC System-On-Chip (SOC)


(ASIC)
General
Control RAM A/D
Processor

Image Video Video


Processor Processor Processor

I/O

Audio Packet
Processor Processor

Security DSP
Processor Processor PHY
Performance and Power Savings

Dedicated
Hardware
Reconfigurable
Hardware
Application
Specific I/S
Processors
Domain-
Specific
Processors
General-
Purpose
Processors

Flexibility
Software
Extensions Extensions
Base Base Base

1. 2. 3. 4. 5.
Processor
Hardware Design:
RTL, scripts, test
bench
ALU OCD

DSP Cache Timer

Register File FPU

High-level Tensilica
specification: Processor
Tensilica Instruction Generator
Extension (TIE)
Customized
Software Tools:
Compiler,
Design processor in one hour debuggers,
simulators, RTOS
Processor Controls Instruction Fetch / PC Instruction RAM
Trace TRACE Port Instruction ROM
JTAG JTAG Tap Control Extended Instruction Instruction Instruction Instruction
Align, Decode, Decode/Dispatch MMU Cache
On Chip Debug
Dispatch
Exception Support
User User Base Register External Interface
Exception Handling Defined Defined File
Registers Register Register Base ALU Xtensa
Data Address Files Files
Processor
Watch Registers MAC 16 DSP
Interface PIF
Instruction Address User User MUL 16/32 Control
Watch Registers Defined Defined
Floating Point
Interrupt Control Execution Execution Write
Interrupts Units and Units
Timers User Defined Buffer
Interfaces
Execution Unit
User Defined
Queues and Wires Vectra Data Data
Vectra DSP Vectra DSP
DSP MMU Cache
Used Defined Data Data
Data ROMs
Load/Store Units Load/Store
Base ISA Feature
Unit Data RAMs
Configurable Function
Xtensa
Optional Function
Local
Optional & Configurable Memory
User Defined Features (TIE) Interface
2.0
2.0 0.5 0.473 0.14

1.8 0.5 0.123


0.12
1.6 0.4

1.4 0.4 0.10

1.2 0.3
0.08
1.0 0.3 0.23
0.06
0.8 0.2

0.6 0.2 0.04


0.03
0.4 0.1
0.018 0.017
0.02 0.016
0.2 0.1 0.03 0.023 0.01
0.087 0.080 0.059 0.058 0.039 0.017 0.016 0.013 0.011
0.0 0.0 0.00
Optimized ConsumerMarks/MHz Optimized TeleMarks/MHz Optimized NetMarks/MHz
Xtensa optimized Xtensa optimized Xtensa optimized
Xtensa out-of-box TI C6203 optimized Xtensa out-of-box
MIPS64 20Kc Xtensa out-of-box MIPS64 20Kc
ARM1020E TI C6203 out-of-box ARM1020E
MIPS64b (NEC VR5000) MIPS64 20Kc MIPS64b (NEC VR5000)
MIPS32b (NEC VR4122) MIPS64b (NEC VR5000) MIPS32b (NEC VR4122)
ARM1020E
MIPS32b (NEC VR4122)
Source: EEMBC © 2004, Tensilica, Inc.
EEMBC ConsumerMark
400000 387328

300000

243968
35
Code Bytes

core + code area mm 2 (0.13µ)


30
200000
25

20

15
100000
70328 10
59648 57149
5

0
0
optimized out-of-box optimized out-of-box
Trimedia TM1300 optimized Trimedia TM1300 out-of-box Xtensa Xtensa Trimedia TM1300 Trimedia TM1300
ARM1020E Xtensa optimized
Xtensa out-of-box core area code area
100

Xtensa Telecom
Xtensa Network
Xtensa Consumer
Xtensa Telecom OPT
EEMBCPerformance/MHz

Xtensa Network OPT


(ARM1026EJ-S = 1)

Xtensa Consumer OPT


ARM1026EJ-S Telecom
ARM1026EJ-S Network
ARM1026EJ-S Consumer
10

50 iffe
x re
d
en nc
er e
gy
1
0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8

Processor Power Dissipation


(mW/MHz including cache, 130nm, 1V)
Designer selects Processor
Original Run XPRES Xtensa
“best” Hardware
C/C++ Code Compiler Processor
configuration Generator ALU OCD

DSP Cache Timer


int
int main(
main( )) Register File FPU
{{
int
int i;
i;
short
short c[100];
c[100];
for (i=0;i<N/2;i++)
for (i=0;i<N/2;i++)
{{

Evaluates millions
of possible Tuned Software
extensions: Tools
• SIMD operations
• operator fusion
• parallel execution
Option: manually refine
configuration
status

Mechanism
reg
CPU
data
reg
status
(8 bit)
CPU xmit/
8251
rcv
data serial
(8 bit) port
while (TRUE) {
/* read */
while (peek(IN_STATUS) == 0);
achar = (char)peek(IN_DATA);
/* write */
poke(OUT_DATA,achar);
while (peek(OUT_STATUS) != 0);
}
intr request
status

Mechanism
reg
intr ack

PC
IR
CPU

data/address data
reg
device 1 device 2 device n

interrupt
acknowledge

L1 L2 .. Ln
CPU
Vector
Device

interrupt interrupt
request acknowledge

CPU

Interrupt
vector handler 0
table head handler 1
handler 2
handler 3
asel

RDATA
ARM Memory
WDATA system

csel bsel
CPDIN CPDOUT

Coprocessor
data
address cache
SRAM
CPU ~MB main
Controller

~10ns
cache

register memory
CMOS DRAM
<KB address ~GB
~1ns ~100ns
data
data
logical physical
address memory address
main
CPU management
memory
unit
page 1
page 2
segment 1

memory

segment 2
segment base address logical address

segment lower bound range range


segment upper bound check error

linear/physical address
page offset

page i base

concatenate

page offset
page
descriptor
page descriptor

flat tree
Translation table
1st index 2nd index offset
base register

descriptor concatenate
1st level table

descriptor concatenate
2nd level table

physical address
accelerator

CPU
memory

I/O
cost

performance
Data input Data output
Accelerated
computation

Execution time on CPU


P1 P2
M1 M2
d1 d2
P3

Task graph Hardware platform

M1 M2

P1 5 5 ?
P2 5 6

P3 - 5
Example process execution times
P1 P2

d1=2 d2=4

P3

Task graph
M1 M2

P1 5 5

P2 5 6

P3 - 5

M1 P2 P1 P2
Time = 14
d1=2 d2=4
M2 P1 P3 P3

network d2

5 10 15 20
time
M1 M2

P1 5 5

P2 5 6

P3 - 5

M1 P1 P1 P2
Time = 12
d1=2 d2=4
M2 P2 P3
P3

network d1

5 10 15 20

time
Function
Function units
units Memories
Memories I/O
I/O ports
ports

Controller
… …

interconnect
Registers
Registers

regs regs regs … regs


µC

interconnect
Interconnect
Interconnect Controller
Controller
inport rom
outport ram_1 ram_2 alu_1 alu_2
acu_1 acu_2 acu_3

imm_1 imm_2 control

• 2 Custom FUs, 2 RAM, 1 ROM, 3 ACU, 2 I/O ports


• Designer responsible for creating custom units manually
µA Definition µA Evaluation
OptimoDE Librarian
OptimoDE User µA
µA µA µA
Library µA µA
User
Library .inc
.c1
.inc .inc
.c2
.inc .inc
.c3
.inc
Library Library
xxxxx
xx
dct
instantiate xxxxx

DesignDE DEvelop
µA
set target
#include …
.tmp
run / profile main()
{

001010
create_resource
save load … = dct();
ISS 110011
010110
}
LIFETIME
LIFETIME
110111
LOAD
LOAD
DEvelop
+ +
C Source 2 3 Analysis feedback
Description INPUT
0 * +
1 2

* + +
1 *
4
1 2 3
2 * +
3 6
* *
4 5 3 * +
8 7
+ +
7 6 4 *
5
*
8 *
1
5 * +
9 10 Micro-
OUTPUT
code

analysis/opti bind schedule

Syntax checks Match architecture Optimize code and


Dataflow analysis and dataflow graph register use
Group 1 Group 2 Group 2

Group 1

Cost: 2 Adders Cost: 0.5 Adders


Gain: 10,000 Cycles Gain: 1,000 Cycles

Group 3 Group 4

Cost: 1 Adder Cost: 1 Adder


Gain: 1,500 Cycles Gain:
Group 3 2,500 Cycles
Group 4
9
ARM 926EJ (200 MHz)
8 OptimoDE (333 MHz, 5 Issue: 3 ALU, 1 Mem, 1 Brn)
OptimoDE + 1 CFU (333 MHz)
7

6
Speedup

0
3Des AES Blowfish Md5 Rc4 SHA
10

7
Speedup

0
0 2 4 6 8 10
CFU Cost Budget (Relative to 32-bit ripple-carry adder)
Input 1 Input 2

+ Input 3

+ Input 4 Input 2 Input 1


ALU 1 ALU 2 ACU SRAM CFU

RF RF RF RF + ^
… 0x1B 0x5 Input 3

>> << &

Control Memory | ^

Output Output
CFU Reg. File and
OptimoDE = 5.5 mm2 in 0.13 µ 2% Interconnect
ARM 926EJ = 5.0 mm in 0.13 µ
2
9%
Control
Memory
11%

ALU 1 ALU 2 ACU SRAM CFU

RF RF RF RF

Baseline
Design
78%

Control Memory
Input 1 0x8

0xFF
>>

Input 1 0x8, 0x4

Input 2 |

0xFF
>>

+
Input 2 |,&
Output

+,-

Output
Speedup
3d

0
1
2
3
4
5
6
7
3d es
es -3d
-B es
lo
3d wfi
e sh
3d s-R
es c4
3des

3d -AE
B e S
Bl low s-S
ow fis ha
fis h-3
h d
Bl -Blo es
ow w
Bl fis fish
ow h-
R
Bl fish c4
ow -A
fis ES
Blowfish

R h-S
R c4- ha
c4 3d
-B es
lo
R wfis
c4 h
R -Rc
c4 4
-
R AE
Rc4

c4 S
AE -S
AE S- ha
S- 3d
Bl es
o
AE wfi
S sh
AE -R
S - c4
AES

AE AE
S S
Key: application run – application designed for

Sh -Sh
Sh a- a
a- 3de
Bl s
o
Sh wfis
a h
Sh -Rc
Base
Sha

a- 4
Sh AE
a- S
Generalized

Sh
a

You might also like