Computer Architecture
Lecture 1
Review of CS 161
What is a von-Neumann computer? =>
The Stored Program Concept Sequential
Execution of a program instructions in
binary for storing in memory
Design of an Instruction Set - RISC Vs.
CISC, Examples
Design of CPU Datapath
CPU Control Design (Hardwire vs.
Microprogramming)
Memory Design (Main memory, Cache
Memory, Virtual memory)
Input-Output
MIPS ISA
1. all MIPS instructions are same length
simplifies fetch and decode (steps 1,2)
Intel 80x86 and IBM 360/370 instructions
are variable length, 1-17 bytes
2. few instruction formats in MIPS
source register fields are same place in
all instructions
can read two registers and decode
instruction in the same cycle
Explicit Load/Store instructions for
memory-register operations
Review of CPU Datapath Design
Instruction operation consists of 5 parts,
namely, Fetch (IF), Decode (ID), Execute
(EX), Memory (DM), and Write-back (WB)
stages
Single Cycle Design Big Cycle CPI = 1
Problems: (1) Low frequency meaning less
number of instructions executed per cycle (2)
All instns take same one big cycle
Multicycle Design Small Cycle CPI < 5
Break the datapath to several stages, each
taking one cycle. Frequency is increased and
some instns can finish earlier. Problems: Need
extra registers to separate the stages and
control must ensure that right control signals
Review: Datapath for MIPS
Stage 5
PC
Instruction
Memory
(Imem)
Registers
Stage 1
Stage 2
ALU
Stage 3
Use datapath figure to represent stages
IFtch Dcd Exec Mem WB
Reg
ALU
IM
DM
Reg
Data
Memory
(Dmem)
Stage 4
Pipelined Execution IPC= 1
Time
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
Program Flow
IFtch Dcd Exec Mem WB
To simplify pipeline, every
instruction takes same number of
steps, called stages
Graphical Pipeline Representation
Time (clock cycles)
ALU
I
n
IM Reg
DM Reg
Load
s
IM Reg
DM Reg
t Add
r.
IM Reg
DM Reg
Store
O
IM Reg
DM Reg
Sub
r
IM Reg
DM Reg
d Or
e
r
(right half highlighted means read, left half write)
ALU
ALU
ALU
ALU
time
Reg
Reg
IM
DM
10
12
14
Reg
DM
Reg
Reg
ALU
IM
ALU
IM
ALU
Reg
IM
DM
ALU
IM
ALU
Reg
DM
Reg
Reg
IM
ALU
Reg
Example: Single-cycle vs. Pipelined
DM
Reg
16
18
20
Advanced Architectural Concepts
Can we achieve CPI < 1?
(i.e., can we have IPC > 1?)
State-of-the-Art Microprocessor
Superscalar execution or
Instruction Level Parallelism (ILP)
Deeper Pipeline => Dynamic Branch
Prediction => Speculation => Recovery
Out-of-order Execution => Instruction
Window and Prefetch => Reorder
Buffers
VLIW Ex: Intel/HP Titanium
Instruction Level Parallelism (ILP) IPC > 1
Time
IFtch Dcd Exec Mem WB
IFetchDcd Exec Mem WB
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
IFtch Dcd Exec Mem WB
Program Flow
ILP = 2
EX: Pentium, SPARC, MIPS 10000, IBM Power
PC
Very Large Instruction Word (VLIW) IPC > 1
Time
IFtch Dcd Exec Mem WB
Exec
IFtch Dcd Exec Mem WB
Exec
IFtch Dcd Exec Mem WB
Exec
Program Flow
EX: Itanium