Chapter 7: Microarchitecture
Advanced
Microarchitecture
Advanced Microarchitecture
• Deep Pipelining
• Microoperations
• Branch Prediction
• Superscalar Processors
• Out of Order Processors
• Register Renaming
• SIMD
• Multithreading
• Multiprocessors
136 Digital Design & Computer Architecture Microarchitecture
Deep Pipelining
• 1020 stages typical
• Number of stages limited by:
– Pipeline hazards
– Sequencing overhead
– Power 300
– Cost 250
Time (ps)
200
Tc
150
Instruction
Time
5 6 7 8 9 10 11 12
N: # of pipeline stages
137 Digital Design & Computer Architecture Microarchitecture
Microoperations
• Decompose complex instructions into series of simple
instructions called micro-operations (microops or µops)
• At runtime, complex instructions are decoded into one or
more microops
• Used heavily in CISC (complex instruction set computer)
architectures (e.g., x86)
Complex Op Microop Sequence
lw s1, 0(s2), postincr 4 lw s1, 0(s2)
addi s2, s2, 4
Without μops, would need 2nd write port on the register file
138 Digital Design & Computer Architecture Microarchitecture
Branch Prediction
• Guess whether branch will be taken
– Backward branches are usually taken (loops)
– Consider history to improve guess
• Good prediction reduces fraction of branches
requiring a flush
139 Digital Design & Computer Architecture Microarchitecture
Branch Prediction
• Ideal pipelined processor: CPI = 1
• Branch misprediction increases CPI
• Static branch prediction:
– Check direction of branch (forward or backward)
– If backward, predict taken
– Else, predict not taken
• Dynamic branch prediction:
– Keep history of last several hundred (or thousand)
branches in branch target buffer, record:
• Branch destination
• Whether branch was taken
140 Digital Design & Computer Architecture Microarchitecture
Dynamic Branch Prediction
• 1bit branch predictor
• 2bit branch predictor
141 Digital Design & Computer Architecture Microarchitecture
Branch Prediction Example
addi s1, zero, 0 # s1 = sum
addi s0, zero, 0 # s0 = i
addi t0, zero, 10 # t0 = 10
For: # for (i=0; i<10; i=i+1)
bge s0, t0, Done
add s1, s1, s0 # sum = sum + i
addi s0, s0, 1 # i = i + 1
j For
Done:
142 Digital Design & Computer Architecture Microarchitecture
1Bit Branch Predictor
• Remembers whether branch was taken the
last time and does the same thing
• Mispredicts first and last branch of loop
addi s1, zero, 0 # s1 = sum
addi s0, zero, 0 # s0 = i
addi t0, zero, 10 # t0 = 10
For: # for (i=0; i<10; i=i+1)
bge s0, t0, Done
add s1, s1, s0 # sum = sum + i
addi s0, s0, 1 # i = i + 1
j For
Done:
143 Digital Design & Computer Architecture Microarchitecture
2Bit Branch Predictor
Strongly Weakly Weakly Strongly taken
Taken taken Taken taken Not Taken taken Not Taken
predict predict predict predict
taken taken taken taken taken not taken taken not taken
addi s1, zero, 0 # s1 = sum
addi s0, zero, 0 # s0 = i
addi t0, zero, 10 # t0 = 10
For: # for (i=0; i<10; i=i+1)
bge s0, t0, Done
add s1, s1, s0 # sum = sum + i
addi s0, s0, 1 # i = i + 1
j For
Done:
Only mispredicts last branch of loop
144 Digital Design & Computer Architecture Microarchitecture
Chapter 7: Microarchitecture
Superscalar & Out of
Order Processors
Superscalar Processors
• Multiple copies of datapath execute multiple
instructions at once
• Dependencies make it tricky to issue multiple
instructions at once
CLK CLK CLK CLK
CLK
PC RD A1
A A2
A3 RD1
RD4
ALUs
A4 A1 RD1
Instruction A5 Register A2 RD2
A6 File RD2
Memory RD5 Data
WD3 Memory
WD6
WD1
WD2
146 Digital Design & Computer Architecture Microarchitecture
Superscalar Example
Ideal IPC: 2
Actual IPC: 2
1 2 3 4 5 6 7 8
Time (cycles)
s0
lw s7
lw s7, 40(s0) 40 +
RF t1 DM RF
IM
add s8
add s8, t1, t2 t2 +
s1
sub s9
sub s9, s1, s3 s3 -
RF s3 DM RF
IM
and s10
and s10, s3, t4 t4 &
s1
or s11
or s11, s1, t5 t5 |
RF s2 DM RF
IM
sw s5
sw s5, 80(s2) 80
147 Digital Design & Computer Architecture Microarchitecture
Superscalar with Dependencies
Ideal IPC: 2
Actual IPC: 6/5 = 1.2
1 2 3 4 5 6 7 8
Time (cycles)
s0
lw s8
lw s8, 40(s0) 40 +
RF t5 DM RF
IM
or s11
or s11, t5, t6 t6 |
RAW
s11
sw s7
sw s7, 80(s11) 80 +
RF DM RF
IM
two cycle latency
RAW
between load and
use of s8 s8
add s9
add s9, s8, t1 t1 +
RF t2 DM RF
WAR IM
sub s8
sub s8, t2, t3 t3 -
RAW
s4
and s10
and s10, s4, s8 s8 &
RF DM RF
IM
148 Digital Design & Computer Architecture Microarchitecture
Out of Order (OOO) Processor
• Looks ahead across multiple instructions
• Issues as many instructions as possible at once
• Issues instructions out of order (as long as no
dependencies)
• Dependencies:
– RAW (read after write): one instruction writes, later
instruction reads a register
– WAR (write after read): one instruction reads, later
instruction writes a register
– WAW (write after write): one instruction writes, later
instruction writes a register
149 Digital Design & Computer Architecture Microarchitecture
Out of Order (OOO) Processor
• Instruction level parallelism (ILP): number
of instruction that can be issued
simultaneously (average < 3)
• Scoreboard: table that keeps track of:
– Instructions waiting to issue
– Available functional units
– Dependencies
150 Digital Design & Computer Architecture Microarchitecture
Out of Order Processor Example
Ideal IPC: 2
Actual IPC: 6/4 = 1.5
1 2 3 4 5 6 7 8
Time (cycles)
s0
lw s8
lw s8, 40(s0) 40 +
RF t5 DM RF
IM
or s11
or s11, t5, t6 t6 |
RAW
s11
sw s7
sw s7, 80(s11) 80 +
RF DM RF
IM
two cycle latency
RAW
between load and
use of R8 s8
add s9
add s9, s8, t1 t1 +
RF t2 DM RF
WAR IM
sub s8
sub s8, t2, t3 t3 -
RAW
s4
and s10
and s10, s4, s8 s8 &
RF DM RF
IM
151 Digital Design & Computer Architecture Microarchitecture
Register Renaming
Ideal IPC: 2
Actual IPC: 6/3 = 2
1 2 3 4 5 6 7
Time (cycles)
s0
lw s8
lw s8, 40(s0) 40 +
RF t2 DM RF
IM
sub r0
sub r0, t2, t3 t3 -
2-cycle RAW RAW s4
and s10
and s10, s4, r0 r0 &
RF t5 DM RF
IM
or s11
or s11, t5, t6 t6 |
RAW s8
s9
add
add s9, s8, t1 t1 +
RF s11 DM RF
IM
sw s7
sw s7, 80(s11) 80 +
152 Digital Design & Computer Architecture Microarchitecture
SIMD
• Single Instruction Multiple Data (SIMD)
– Single instruction acts on multiple pieces of data at once
– Common application: graphics
– Can apply to short arithmetic operations (also called
packed arithmetic)
• For example, add eight 8bit elements
63 56 55 48 47 40 39 32 31 24 23 16 15 8 7 0 Bit position
a7 a6 a5 a4 a3 a2 a1 a0 D0
+ b7 b6 b5 b4 b3 b2 b1 b0 D1
a7 + b7 a6 + b6 a5 + b5 a4 + b4 a3 + b3 a2 + b2 a1 + b1 a0 + b0 D2
153 Digital Design & Computer Architecture Microarchitecture
Chapter 7: Microarchitecture
Multithreading &
Multiprocessors
Advanced Architecture Techniques
• Multithreading
– Wordprocessor: thread for typing, spell checking,
printing
• Multiprocessors
– Multiple processors (cores) on a single chip
155 Digital Design & Computer Architecture Microarchitecture
Threading: Definitions
• Process: program running on a computer
– Multiple processes can run at once: e.g., surfing
Web, playing music, writing a paper
• Thread: part of a program
– Each process has multiple threads: e.g., a word
processor may have threads for typing, spell
checking, printing
156 Digital Design & Computer Architecture Microarchitecture
Threads in a Conventional Processor
Singlecore system:
• One thread runs at once
• When one thread stalls (for example, waiting
for memory):
– Architectural state of that thread stored
– Architectural state of waiting thread loaded into
processor and it runs
– Called context switching
• Appears to user like all threads running
simultaneously
157 Digital Design & Computer Architecture Microarchitecture
Multithreading
• Multiple copies of architectural state
• Multiple threads active at once:
– When one thread stalls, another runs immediately
– If one thread can’t keep all execution units busy,
another thread can use them
• Does not increase instructionlevel parallelism
(ILP) of single thread, but increases
throughput
Intel calls this “hyperthreading”
158 Digital Design & Computer Architecture Microarchitecture
Multiprocessors
• Multiple processors (cores) with a method of
communication between them
• Types:
– Homogeneous: multiple cores with shared main
memory
– Heterogeneous: separate cores for different tasks (for
example, DSP and CPU in cell phone)
– Clusters: each core has own memory system
159 Digital Design & Computer Architecture Microarchitecture
About these Notes
Digital Design and Computer Architecture Lecture Notes
© 2021 Sarah Harris and David Harris
These notes may be used and modified for educational and/or
noncommercial purposes so long as the source is attributed.
160 Digital Design & Computer Architecture Microarchitecture