Module 5
Basic Processing Unit
Computer Organization, Fifth Edition
Carl Hamacher, Zvonko Vranesic, Safwat Zaky
Presented By:
Prof. Sunanda,
Assistant Professor,
CSE, DSCE
Outline
• Some Fundamental concepts
• Execution of a Complete Instruction
• Multiple bus organization
• Hard-wired control
• Micro programmed control
[Link], Assistant Professor, Dept of CSE, DSCE
Fundamental Concepts
• Processor fetches one instruction at a time and
perform the operation specified.
• Instructions are fetched from successive memory
locations until a branch or a jump instruction is
encountered.
• Processor keeps track of the address of the memory
location containing the next instruction to be fetched
using Program Counter (PC).
• Instruction Register (IR)
[Link], Assistant Professor, Dept of CSE, DSCE
Executing an Instruction
• Fetch the contents of the memory location pointed to by
the PC. The contents of this location are loaded into the IR
(fetch phase).
IR ← [[PC]]
• Assuming that the memory is byte addressable, increment
the contents of the PC by 4 (fetch phase).
PC ← [PC] + 4
• Carry out the actions specified by the instruction in the IR
(execution phase).
[Link], Assistant Professor, Dept of CSE, DSCE
Processor Organization – Single Bus Structure
[Link], Assistant Professor, Dept of CSE, DSCE
Executing an Instruction
• Transfera word of data from one processor register to
another or to the ALU.
• Perform an arithmetic or a logic operation and store the
result in a processor register.
• Fetch the contents of a given memory location and load
them into a processor register.
• Store a word of data from a processor register into a given
memory location.
[Link], Assistant Professor, Dept of CSE, DSCE
Register Transfers
Internal processor
bus
Riin
Ri
Riout
Y in
Constant 4
Select MUX
A B
ALU
Z in
Z out
[Link], Assistant Professor, Dept of CSE, DSCE
Figure 7.2. Input and output gating for the registers in Figure 7.1.
Performing an Arithmetic or
Logic Operation
• The ALU is a combinational circuit that has no internal
storage.
• ALU gets the two operands from MUX and bus. The
result is temporarily stored in register Z.
• What is the sequence of operations to add the contents of
register R1 to those of R2 and store the result in R3?
1. R1out, Yin
2. R2out, SelectY, Add, Zin
3. Zout, R3inAssistant Professor, Dept of CSE, DSCE
[Link],
Fetching a Word from Memory
• Address into MAR; issue Read operation; data into MDR.
[Link], Assistant Professor, Dept of CSE, DSCE
Fetching a Word from Memory
• The response time of each memory access varies (cache
miss, memory-mapped I/O,…).
• To accommodate this, the processor waits until it
receives an indication that the requested operation has
been completed (Memory-Function-Completed, MFC).
• Move (R1), R2
[Link], Assistant Professor, Dept of CSE, DSCE
Timing Diagram – Memory Read Operation
[Link], Assistant Professor, Dept of CSE, DSCE
Execution of a Complete Instruction
• Add (R3), R1
• Fetch the instruction
• Fetch the first operand (the contents of the memory
location pointed to by R3)
• Perform the addition
• Load the result into R1
[Link], Assistant Professor, Dept of CSE, DSCE
Execution of a Complete Instruction
Add (R3), R1
[Link], Assistant Professor, Dept of CSE, DSCE
Execution of Branch Instructions
• A branch instruction replaces the contents of PC with the
branch target address, which is usually obtained by adding
an offset X given in the branch instruction.
• The offset X is usually the difference between the branch
target address and the address immediately following the
branch instruction.
• Conditional branch
[Link], Assistant Professor, Dept of CSE, DSCE
Execution of Branch Instructions
[Link], Assistant Professor, Dept of CSE, DSCE
Multiple-Bus Organization
• While the steps still have to be performed by sequentially
changing from one step to another, the number of
independent steps can be clubbed into single step in a
multiple bus organization.
• Improvement in performance in a multiple bus
organization
[Link], Assistant Professor, Dept of CSE, DSCE
Multiple-Bus Organization
[Link], Assistant Professor, Dept of CSE, DSCE
Multiple-Bus Organization
• Add R4, R5, R6
[Link], Assistant Professor, Dept of CSE, DSCE
Traditional Pipeline Concept
•Laundry Example
•Ann, Brian, Cathy, Dave
each have one load of clothes
to wash, dry, and fold
A B C D
•Washer takes 30 minutes
•Dryer takes 40 minutes
•“Folder” takes 20 minutes
[Link], Assistant Professor, Dept of CSE, DSCE
Traditional Pipeline Concept
6 PM 7 8 9 10 11 Midnight
Time
30 40 20 30 40 20 30 40 20 30 40 20
A
• Sequential laundry takes 6
hours for 4 loads
B
• If they learned pipelining, how
long would laundry take?
D
[Link], Assistant Professor, Dept of CSE, DSCE
Traditional Pipeline Concept
6 PM 7 8 9 10 11 Midnight
T Time
a
30 40 40 40 40 20
s
k A
•Pipelined laundry takes 3.5
O B hours for 4 loads
r
d C
e
r D
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Concept
• Pipelining doesn’t help latency
6 PM 7 8 9 of single task, it helps
throughput of entire workload
Time
T • Pipeline rate limited by slowest
a 30 40 40 40 40 20 pipeline stage
s
A • Multiple tasks operating
k
simultaneously using different
resources
O B
r • Potential speedup = Number
d pipe stages
C
e • Unbalanced lengths of pipe
r stages reduces speedup
D
[Link], Assistant Professor, Dept of CSE, DSCE • Time to “fill” pipeline and time
to “drain” it reduces speedup
Use the Idea of Pipelining in a
Computer
Fetch + Execution
Tim
I1 I2 I3 e
Tim
Clock 1 2 3 4 e
F E F E F E cycle
1 1 2 2 3 3 Instruction
I1 F1 E1
(a) Sequential execution
I2 F2 E2
Interstage
buffer B
1 I3 F3 E3
Instructio E ecutio
n fetc x nuni
(c) Pipelined
huni t execution
t
Figure 8.1. Basic idea of instruction
(b) Hardware organization
pipelining.
[Link], Assistant Professor, Dept of CSE, DSCE
Use the Idea of Pipelining in a
Computer
Fetch + Decode
+ Execution + Write
Textbook page: 457
[Link], Assistant Professor, Dept of CSE, DSCE
Use the Idea of Pipelining in a
Computer
• Fetch(F)- read the instruction from the memory
• Decode(D)- Decode the instruction and fetch the source operand
• Execute(E)- perform the operation specified by the instruction
• Write(W)- store the result in the destination location
[Link], Assistant Professor, Dept of CSE, DSCE
Role of Cache Memory
• Each pipeline stage is expected to complete in one
clock cycle.
• The clock period should be long enough to let the
slowest pipeline stage to complete.
• Faster stages can only wait for the slowest one to
complete.
• Since main memory is very slow compared to the
execution, if each instruction needs to be fetched from
main memory, pipeline is almost useless.
• Fortunately, we have cache.
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Performance
• The potential increase in performance resulting from
pipelining is proportional to the number of pipeline stages.
• However, this increase would be achieved only if all
pipeline stages require the same time to complete, and
there is no interruption throughout program execution.
• Unfortunately, this is not true.
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Performance
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Performance
• The previous pipeline is said to have been stalled for two clock
cycles.
• Any condition that causes a pipeline to stall is called a hazard.
• Data hazard – any condition in which either the source or the
destination operands of an instruction are not available at the time
expected in the pipeline. So some operation has to be delayed, and
the pipeline stalls.
• Instruction (control) hazard – a delay in the availability of an
instruction causes the pipeline to stall.
• Structural hazard – the situation when two instructions require the
use of a given hardware resource at the same time.
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Performance
Instruction
hazard
Idle periods
– stalls
(bubbles)
[Link], Assistant Professor, Dept of CSE, DSCE
Pipeline Performance
Structural Load X(R1), R2
hazard
[Link], Assistant Professor, Dept of CSE, DSCE
Model Questions
1. Explain with neat diagram, the basic organization of a microprogrammed
control.
2. Describe the three bus organization of the data path and describe in detail.
3. Write control sequence for the instruction Add R1, R2, R3.
4. Explain a complete processor with a neat diagram
5. Write and explain the control sequences for the execution of the following
instruction: Add (R3), R1.
6. Bring out the differences between micro programmed controls and hard
wired control.
7. Explain Field coded micro instructions with a neat diagram.
8. Explain multiple bus organization and its advantages.
9. Explain the role of cache memory in pipelining.
10. Explain pipelining performance.
[Link], Assistant Professor, Dept of CSE, DSCE