Unit I Computer Evolution and Performance
Unit I Computer Evolution and Performance
➢ What is a computer?
Computer is an advanced electronic device that takes raw data as an input from the user and
processes it under the control of a set of instructions (called program), produces a result(output),
and saves it for future use. This tutorial explains the foundational concepts of computer hardware,
software, operating systems, peripherals, etc. along with how to get the most value and impact
from computer technology.
Functionalities of a Computer
There are three basic functionalities of a Computer System and they are
1. Input
2. Process
3. Output
But if we look at it in a very broad sense, any digital computer carries out the following five
functions:
Step 2 - Stores the data/instructions in its memory and uses them as required.
• Accuracy − Computers can perform very complex computations accurately in a very short
period of time. If a user inputs the correct input to the computer, it gives accurate results
that can be used in decision-making.
• Storage − Computers can store large amounts of data permanently. The data is saved in
files, which can be accessed at any time; these files are saved for a long time period until a
user deletes them.
• Diligently − A computer can do the assigned task diligently. A computer can work for hours
without getting tired. Hence, it can do thousands of complex computations with the same
accuracy.
• No I.Q. − A computer does not have its own I.Q.; it carries out the predetermined tasks and
does not take its own decisions.
• No Feelings − A computer does not have emotions. It works as per the given instructions
by users.
Disadvantages of Computers
• Health Issues − Working long hours on computers leads to health issues. Student's playing
games and accessing related applications for long periods of time cause serious health
problems.
• Spread of Pornography − The growing trend of the internet has spread pornography. In
today's time, pornography is a big threat to society and the youth.
• Virus and hacking attacks − Viruses are unwanted programmes that enter computers
through networks or the internet. These programmes may steal information or damage
computers. Sometimes these lock the application programmes of the computer to affect its
working.
• No IQ − Computers cannot make their own decisions. Its functioning depends on human
interventions.
• Negative effect on the environment − The increasing use of computers and automated
devices has posed a major threat to the environment.
• Crashed Networks − Hackers may destroy the network, which affects the overall working
of the existing system. In todays time, most of the data is on servers, so destroying the
network may be a serious threat to communication.
3000 This is the earliest computing device. This was used to do basic arithmetic calculations. In
Abacus
BCE this computer, beads were moved along rods to represent numbers.
Mechanical
computer (Step Mechanical devices and ideas that were important precursors to the development of
Reckoner, Turk's computers and automation were introduced. It uses mechanical components, such as gears,
18th ce
Head, Difference levers, and switches, to perform calculations and process information. The popular
ntury
Engine, mechanical devices developed during 18th century were Step Reckoner, Turk's Head,
Analytical
Difference Engine, Analytical Engine.
Engine).
Electromechanic
al
Computers Devi
ces like the Z3,
Mark I
ENIAC (1945)
Stored-Program
Computers
(1940s-1950s) Computers developed during 19th century were crucial in shaping the
concepts and ideas that eventually led to the creation of the computers we use
Transistors and today. Most of the devices were based on the combination of mechanical and
Integrated
Circuits (1950s- electrical switches to perform computation. The 19th century was the time
19th ce 1960s) where invention of computing devices was more and more. The size of
ntury computer was reduced and the devices with large storage and high
Minicomputers
computations were introduced. Interconnectivity with multiple devices and
(1960s-1970s)
data sharing, remote accessing were recorded as the features of the computers
Microprocessors which makes it popular in the world and make the computer as most
and Personal demandable computing device in the world.
Computers
(1970s-1980s)
Graphical User
Interfaces (1980s-
1990s)
Internet and
World Wide Web
(1990s)
Laptops, 20th century is the time where computer technology is the next level. Portable and light
Smartphones, weighted high computing devices are introduced and in trend. Cloud Computing
and Tablets technology makes the internet as a more useful platform to keep the data centralise in
20th ce (2000s- terms of accessing and its computation on server. Hence, cloud computing involves
ntury Present) delivering various services over the internet, such as computing power, storage, and
applications. The 20th century saw significant developments in the field of artificial
Cloud intelligence (AI) also. AI technologies began to be integrated into various applications,
Computing such as speech recognition, image processing, natural language processing, and robotics.
and AI (Most These developments set the stage for the further evolution of AI in the 21st century, where
demandable in the focus shifted toward more data-driven approaches.
cutting edge
technology)
The history of computers is marked by a continuous cycle of innovation, with each generation
building upon the achievements of the previous one. This overview provides just a glimpse into
the rich and complex evolution of computing technology.
➢ Generations of Computers
Generation in computer terminology is a change in technology a computer is/was being used.
Initially, the generation term was used to distinguish between varying hardware technologies.
Nowadays, generation includes both hardware and software, which together make up an entire
computer system.
There are five computer generations known till date. Each generation has been discussed in
detail along with their time period and characteristics. In the following table, approximate dates
against each generation has been mentioned, which are normally accepted.
• Very costly
• Huge size
• Need of AC
• Non-portable
• EDVAC
• UNIVAC
• IBM-701
• IBM-750
2. Second Generation Computers
The period of second generation was from 1959-1965. In this generation, transistors were used
that were cheaper, consumed less power, more compact in size, more reliable and faster than the
first-generation machines made of vacuum tubes. In this generation, magnetic cores were used as
the primary memory and magnetic tape and magnetic disks as secondary storage devices.
In this generation, assembly language and high-level programming languages like FORTRAN,
COBOL were used. The computers used batch processing and multiprogramming operating
system.
• Use of transistors
• AC required
• IBM 1620
• IBM 7094
• CDC 1604
• CDC 3600
• UNIVAC 1108
3. Third Generation Computers
The period of third generation was from 1965-1971. The computers of third generation used
Integrated Circuits (ICs) in place of transistors. A single IC has many transistors, resistors, and
capacitors along with the associated circuitry.
The IC was invented by Jack Kilby. This development made computers smaller in size, reliable,
and efficient. In this generation remote processing, time-sharing, multi-programming operating
system were used. High-level languages (FORTRAN-II TO IV, COBOL, PASCAL PL/1,BASIC,
ALGOL-68 etc.) were used during this generation.
• IC used
• Faster
• Lesser maintenance
• Costly
• AC required
• Honeywell-6000 series
• IBM-370/168
• TDC-316
4. Fourth Generation Computers
The period of fourth generation was from 1971-1980. Computers of fourth generation used
Very Large Scale Integrated (VLSI) circuits. VLSI circuits having about 5000 transistors and
other circuit elements with their associated circuits on a single chip made it possible to have
microcomputers of fourth generation. Fourth generation computers became more powerful,
compact, reliable, and affordable. As a result, it gave rise to Personal Computer (PC) revolution.
In this generation, time sharing, real time networks, distributed operating system were used. All
the high-level languages like C,C++, DBASE etc., were used in this generation.
The main features of fourth generation are:
• Very cheap
• Portable and reliable
• Use of PCs
• Pipeline processing
• No AC required
• DEC 10
• STAR 1000
• PDP 11
• CRAY-1(Super Computer)
• CRAY-X-MP(Super Computer
5. Fifth Generation Computers
The period of fifth generation is 1980-till date. In the fifth generation, VLSI technology became
ULSI (Ultra Large Scale Integration) technology, resulting in the production of microprocessor
This generation is based on parallel processing hardware and AI (Artificial Intelligence) software.
AI is an emerging branch in computer science, which interprets the means and method of making
computers think like human beings. All the high-level languages like C and C++, Java, .Net etc.,
are used in this generation.
• ULSI technology
• Desktop
• Laptop
• Notebook
• Ultrabook
• Chromebook
➢ Von Neumann Architecture
Von Neumann architecture is a computer design model proposed by John von
Neumann in 1945. It is based on the stored-program concept, where both
program instructions and data are stored in the same memory. The architecture
consists of a Central Processing Unit (CPU), main memory, input unit,
and output unit. Instructions are executed sequentially using the fetch-
decode-execute cycle. A major limitation of this architecture is the Von
Neumann bottleneck, caused by the use of a single bus for both data and
instructions.
Core Principle: Stored-Program Concept
The key idea of Von Neumann architecture is that:
• Instructions and data are represented in binary form
• Both are stored in a single shared memory
• Instructions can be modified like data during execution
This enables flexibility, conditional branching, and complex program control.
2. Main Memory
• Stores both program instructions and data
• Memory locations are identified using unique addresses
• Accessed sequentially during program execution
3. Input Unit
• Accepts data and instructions from external devices
• Converts input into machine-readable binary form
4. Output Unit
• Produces results in human-readable form
• Converts binary data into output signals
Solutions:
• Cache Memory: Small, fast memory close to CPU to store frequently used
data and instructions.
• Pipelining & Prefetching: Fetch instructions ahead of execution to hide
memory latency.
• Harvard Architecture: Separate memory and buses for instructions and data
to allow simultaneous access.
• Wider or Multiple Memory Channels: Increase data transfer per cycle.
Conclusion:
The Von Neumann bottleneck occurs because the CPU and memory cannot
communicate simultaneously for instructions and data, making memory
access the limiting factor in system performance. Modern computers reduce
this bottleneck using cache, pipelining, and hybrid architectures.
Characteristics
• Single memory for instructions and data
• Sequential instruction execution
• Simple and cost-effective design
• Flexible program storage
Advantages
• Simple hardware implementation
• Lower cost
• Easy program modification
• Suitable for general-purpose computing
Disadvantages
• Limited performance due to memory bottleneck
• Slower execution compared to Harvard architecture
• Inefficient for high-speed parallel processing
➢ Harvard Architecture
Harvard architecture is a computer design model where program
instructions and data are stored in separate memory units that are accessed
through independent buses. This separation allows the processor to fetch
instructions and access data simultaneously, which helps avoid the bottleneck
present in traditional Von Neumann systems.
• Eliminates the Von Neumann bottleneck.
• Faster and predictable performance (suitable for real-time systems).
• Parallel access to both instructions and data.
Working Principle
In Harvard Architecture, fetching an instruction from instruction memory and
reading/writing data from/to data memory happen at the same time without
waiting for one to finish. Separate buses prevent the bottleneck that occurs
when data and instructions share a path. For example, while an instruction is
being executed, the next instruction can be fetched simultaneously, speeding
up processing.
There is common bus for data and Separate buses are used for
instruction transfer. transferring data and instruction.
CPU can not access instructions CPU can access instructions and
and read/write at the same time. read/write at the same time.
Here are several factors that can impact the performance of a computer system,
including:
• Processor speed: The speed of the processor, measured in GHz (gigahertz), determines
how quickly the computer can execute instructions and process data.
• Memory: The amount and speed of the memory, including RAM (random access memory)
and cache memory, can impact how quickly data can be accessed and processed by the
computer.
• Storage: The speed and capacity of the storage devices, including hard drives and solid-
state drives (SSDs), can impact the speed at which data can be stored and retrieved.
• I/O devices: The speed and efficiency of input/output devices, such as keyboards , mice,
and displays, can impact the overall performance of the system.
• Software optimization: The efficiency of the software running on the system, including
operating systems and applications, can impact how quickly tasks can be completed.
• Designing for performance focuses on reducing execution time and increasing system
efficiency
• Achieved by optimizing processor speed, memory system, and architecture
• Requires a balanced design approach.
• It Depends on Microprocessor Speed, Performance Balance, Improvements in Chip
Organization and Architecture
1. Microprocessor Speed
Definition
• Microprocessor speed indicates how fast a processor executes instructions.
Key Points
• Depends on:
o Clock frequency
o Instruction execution efficiency
• Higher clock speed allows more operations per second
• Performance is also affected by Cycles Per Instruction (CPI)
• Clock speed alone does not guarantee high performance
Performance Formula
CPU Execution Time = Instruction Count × CPI × Clock Cycle Time
OR
Instruction Count × CPI
CPU Time =
Clock Rate
2. Performance Balance
Definition
• Performance balance ensures that CPU, memory, and I/O operate efficiently
together without creating bottlenecks.
Key Points
• Avoids bottlenecks in the system
• Fast CPU with slow memory leads to poor performance
• All components must be upgraded proportionally
Techniques to Achieve Performance Balance
• Memory hierarchy
o Registers
o Cache
o Main memory
• High-speed buses
• Parallel data transfer
• Efficient I/O handling
Advantages
• Efficient utilization of hardware resources
• Reduced idle time of CPU
• Improved overall system performance
Applications
• Embedded systems
• Real-time systems
• Database servers
• Network servers
Evolution Summary
• Increase in word size from 4-bit to 64-bit
• Growth in processing speed and efficiency
• Improved memory addressing capabilities
➢ Performance Assessment
2. CPI
Total Clock Cycles
CPI =
Total Instructions
4. Overall Performance
1
Performance =
Execution Time
Applications
• High-Performance Computing (HPC)
o Scientific simulations, weather modeling, AI/ML workloads.
• Embedded and Real-Time Systems
o Automotive, medical devices, robotics.
• Servers and Cloud Computing
o Database servers, web servers, virtualization platforms.
• Software Optimization
o Compiler and algorithm performance evaluation.
• System Design
o Helps architects design efficient pipelines, caches, and memory hierarchies.
The processing required for a single instruction is called an instruction cycle. The two steps are
referred to as the fetch cycle and the execute cycle. Program execution halts only if the machine is
turned off, some sort of unrecoverable error occurs, or a program instruction that halts the
computer is encountered.
At the beginning of each instruction cycle, the processor fetches an instruction from memory. In a
typical processor, a register called the program counter (PC) holds the address of the instruction to
be fetched next. The instruction fetch and execute cycle (or fetch-decode-execute) is the core
operational process of a CPU, which continuously fetches machine-level instructions from
memory, decodes them into control signals, and executes the operations to run programs. This
cycle repeats billions of times per second until the system shuts down.
The instruction cycle (or fetch-decode-execute cycle) is the fundamental process a CPU uses to
run programs, involving fetching an instruction from memory, decoding it to understand the
operation, and then executing it, repeating until halted. Key stages are Fetch (get instruction using
Program Counter (PC) into Instruction Register (IR)), Decode (interpret opcode and
operands), Execute (perform operation), and sometimes Write-back/Store (save result)
. A diagram shows these steps as a flowchart, moving from fetching to decoding, then executing,
with the cycle looping back to fetch the next instruction
.
Stages of the Instruction Cycle
1. Fetch Cycle:
• The address in the Program Counter (PC) is sent to the Memory Address
Register (MAR).
• A read signal is sent to memory.
• The instruction at that address is read into the Memory Data Register (MDR).
• The instruction from MDR moves to the Instruction Register (IR).
• The PC is incremented to point to the next instruction.
2. Decode Cycle:
• The control unit decodes the opcode (the operation part) in the IR.
• It identifies the instruction type (memory, register, I/O) and operands
(data/address).
• If it's a memory reference with indirect addressing, the effective address is fetched
from memory.
3. Execute Cycle:
• The control unit generates signals to perform the operation (e.g., addition in
the Arithmetic Logic Unit (ALU)).
• Operands are fetched if needed.
4. Write-back/Store Cycle (Optional):
• The result of the execution is written back to a register or memory.
Key Components Involved:
• Program Counter (PC): Holds the address of the next instruction.
• Memory Address Register (MAR): Stores the address in memory that the CPU wants to
access.
• Memory Data Register (MDR): Stores the data being transferred to or from memory.
• Instruction Register (IR): Holds the current instruction being decoded and executed.
• Control Unit (CU): Generates signals to manage the flow of data and control the
processor.
• Arithmetic Logic Unit (ALU): Performs calculation and logical operations.
This process is synchronized by the system clock, where faster clock speeds allow for more cycles
per second, enhancing performance.
Interconnection structure
A computer system consists of three main modules:
• CPU (Central Processing Unit)
• Main Memory
• Input / Output (I/O) Modules
These components must communicate continuously to execute programs. The collection of
communication paths that connect CPU, memory, and I/O modules is known as the
interconnection structure. The design of this structure determines how efficiently information is
exchanged within the computer system.
Definition
Interconnection Structure is the arrangement of communication pathways that allow data,
instructions, and control signals to flow between CPU, memory, and I/O devices.
1. Address Bus
A collection of wires used to identify particular location in main memory is called Address Bus. Or
in other words, the information used to describe the memory locations travels along the address
bus.
• The address bus transports memory addresses which the processor wants to access in
order to read or write data..
• The address bus is unidirectional.
• The size of address bus determines how many unique memory locations can be
addressed.
Example:
• A system with 4-bit address bus can address 24 = 16 Bytes of memory.
• A system with 16-bit address bus can address 216 = 64 KB of memory
• A system with 20-bit address bus can address 220 = 1 MB of memory.
2. Data Bus
A collection of wires through which data is transmitted from one part of a computer to another
is called Data Bus. It can be thought of as a highway on which data travels within a computer.
• The main objective of data bus is transfer of the data between microprocessor to input/
output devices or memory.
• The data bus transfers instructions coming from or going to the processor.
• The data bus is bidirectional because the data can flow in either direction from CPU to
memory(or input/output device) or from memory to the CPU.
• The size (width) of bus determines how much data can be transmitted at one time.
Example:
• A 16-bit bus can transmit 16 bits of data at a time.
• 32-bit bus can transmit 32 bits at a time.
3. Control Bus
The connections that carry control information between the CPU and other devices within the
computer is called Control Bus. The control bus transports orders and synchronization signal
coming from the control unit and travelling to all other hardware components
• The main objective of control bus is all signals controller carried from processor to other
hardware device.
• The Control bus is bidirectional because the data can flow in either direction from CPU to
memory(or input/output device) or from memory to the CPU.
• It also transmits response signals from the hardware.
Comparison Between System Buses
The below table shows the comparison between the three buses as below :
Address Bus
Carries memory addresses; Identifies where data should go
(Unidirectional)
Data Bus (Bidirectional) Carries actual data, moves data between components
• 1. Arithmetic Operations
• ALU performs arithmetic operations to carry out basic mathematical calculations on binary
numbers.
3. Shift Operations
• ALU can shift bits left or right to aid in fast arithmetic and bit manipulation.
• Types include logical shift, arithmetic shift, and rotate operations.
• Often used for multiplication/division by powers of two and bit-level data formatting.
4. Comparison Operations
• ALU supports comparison by evaluating two operands and setting appropriate flags.
• It checks for conditions like equality, greater than, or less than.
• Useful in branching and conditional instructions based on flag values.
Binary arithmetic is the core operational element in every digital device, ranging from microcontrollers in
household appliances to multi-core processors in data centres. A full adder circuit lies at the heart of binary
arithmetic. It’s a combinational logic network that adds three input bits and produces a two-bit result—sum
and carry. Full adders are very simple in concept. The design of a full adder and its integration into larger
arithmetic units involves important trade-offs in terms of:
• speed
• Power
• Area
• Scalability
Therefore, engineers working with digital logic must understand how a full adder works, how it differs from
a half adder, how to build it using basic gates or hardware description languages, and how modern research
pushes its performance limits.
This article discusses a holistic view of the full adder circuit. We will discuss the theory of binary addition
and the differences between half- and full-adders. Later, we will cover several implementation strategies—
from gate-level schematics and transistor-level realisations to high-level hardware design flows for FPGAs
and ASICs. The discussion extends to multi-bit architectures, such as ripple-carry and carry-look-ahead
adders, and highlights commercial ICs as well as low-power innovations.
Binary Addition - Why Do We Need a Full Adder?
Digital systems represent numbers in base-2, where each bit can be 0 or 1. It means that a carry may be
generated during bit-by-bit addition.
The simplest adder is a half adder, which adds two single-bit inputs but cannot handle a carry-in from a
previous addition. Its output comprises a sum and a carry bit. A full adder was developed to tackle the carry
problem in multi-bit addition.
Half adder vs full adder
The main differences between a half adder and a full adder are summarised in the table below. Each keyword
entry is kept short to fit within narrow columns. In a half adder, the sum is the XOR of A and B and the
carry is the AND of A and B.
Fig 1: Comparison of a half adder and full adder circuit
However, if a carry arrives from a previous stage, the half adder cannot process it.
A full adder solves this problem by accepting an additional carry-in bit. This allows it to add three bits—
two operands and a carry—and produce a sum and carry-out, making it suitable for cascading in multi-bit
adders.
Feature Half adder Full adder
Inputs A, B A, B, Carry-in
Carry handling None Adds incoming carry
Outputs Sum and carry Sum and carry-out
Complexity Simple More complex due to extra input
Typical use Building block of full adders Multi-bit addition, digital processors
Logic description of a full adder
A full adder is a combinational circuit that adds two binary digits and a carry bit and generates a sum bit
and a carry bit. Internally, one XOR gate, three AND gates, and one OR gate connect to realise the circuit.
The operation is straightforward:
• Inputs: A, B and Cin
• Sum output (S): A ⊕ B ⊕ Cin
• Carry output (Cout): A·B + A·Cin + B·Cin
The truth table for the full adder, reproduced below, shows all eight input combinations and their
corresponding outputs:
A B C_in Sum C_out
0 0 0 0 0
0 0 1 1 0
0 1 0 1 0
0 1 1 0 1
1 0 0 1 0
1 0 1 0 1
1 1 0 0 1
1 1 1 1 1
We can summarize the truth table with the following Boolean expressions:
• S=A⊕B⊕Cin
• Cout=AB+ACin+BCin
Using these expressions, designers can derive logic gate implementations or optimise the circuit using
Karnaugh maps and Boolean algebra.
Implementing a Full Adder
Using half adders and basic gates
One intuitive way to build a full adder is to combine two half adders with an OR gate. When two half adder
circuits are connected, the first adds inputs A and B, and its sum is fed into a second half adder along with
the carry-in.
The two carry outputs from these half adders are ORed to produce the final carry. This modular approach
explains the relationship between half and full adders but also simplifies testing and debugging when
designing in hardware description languages (HDLs).
When implementing this structure with logic gates, each half adder uses an XOR gate for the sum and an
AND gate for the carry. The resulting full adder consists of two XOR gates, two AND gates and one OR
gate.
Universal gate implementations
Many teaching labs require building circuits from universal gates to demonstrate gate equivalence. A full
adder can be implemented using only NAND gates or only NOR gates.
A NAND-only design utilises nine NAND gates. It features two half adder equivalents and an extra NAND
to combine carries.
Fig 2: A Full Adder using NAND Gates only
A NOR-only design is the same as the NAND implementation, but designed with NOR gates.
Basic Idea
Booth’s algorithm examines two bits at a time:
• Current least significant bit of multiplier (Q₀)
• An extra bit Q₋₁ (previous bit)
Based on these bits, the algorithm decides whether to add, subtract, or do nothing.
Registers Used
Register Purpose
A Accumulator
Q Multiplier
M Multiplicand
Q₋₁ Extra bit
Count Number of bits
Algorithm Steps
1. Initialize registers
o A=0
o Q = Multiplier
o M = Multiplicand
o Q₋₁ = 0
o Count = number of bits
2. Check Q₀ and Q₋₁.
3. Perform operation according to Booth rule.
4. Perform Arithmetic Right Shift (A, Q, Q₋₁).
5. Decrease count.
6. Repeat until count = 0.
7. Final result stored in (A, Q).
Example: Multiply 7 × 3 using Booth Algorithm
Binary values:
M = 0111 (7)
Q = 0011 (3)
A = 0000
Q₋₁ = 0
Count = 4
Step Table
Step A Q Q₋₁ Operation
Initial 0000 0011 0 Start
1 1001 0011 0 A=A−M
Shift 1100 1001 1 Shift
2 1100 1001 1 No operation
Shift 1110 0100 1 Shift
3 0101 0100 1 A=A+M
Shift 0010 1010 0 Shift
4 0010 1010 0 No operation
Shift 0001 0101 0 Final
Final Result:
AQ = 00010101
Decimal value:
21
So,
7 × 3 = 21
Registers Used
Register Meaning
A Accumulator
M Multiplicand
Q Multiplier
Q₋₁ Extra bit (initially 0)
Count Number of bits
Booth’s Decision Rules
Q₀ Q₋₁ Operation
0 0 No operation
1 1 No operation
0 1 A=A+ M
1 0 A=A− M
After the operation → perform Arithmetic Right Shift (A, Q, Q₋₁).
Step-by-Step Table
Step A Q Q₋₁ Operation Comment
Initial 0000 0011 0 — Registers initialized
1 1011 0011 0 A=A− M Q₀=1, Q₋₁=0 → subtract M
Shift 1101 1001 1 Arithmetic shift Shift A,Q,Q₋₁ right
2 1101 1001 1 No operation Q₀=1, Q₋₁=1
Shift 1110 1100 1 Shift Arithmetic shift
3 0011 1100 1 A=A+ M Q₀=0, Q₋₁=1 → add M
Shift 0001 1110 0 Shift Arithmetic shift
4 0001 1110 0 No operation Q₀=0, Q₋₁=0
Shift 0000 1111 0 Final shift End of iterations
Final Result
AQ = 00001111
Decimal value:
1111₂ = 15₁₀
So,
5 × 3 = 15