Organização e Arquitetura de Computadores
Organização e Arquitetura de Computadores
[Link]
2
PCS 3612 - 2025 © CBM
Introdução
Seção 2.1 do livro-texto
3
PCS 3612 - 2025 © CBM
Conjunto de instruções
4
PCS 3612 - 2025 © CBM
Conjunto de instruções RISC-V
5
PCS 3612 - 2025 © CBM
Conceito de Programa Armazenado
e Arquiteturas Associadas
6
PCS 3612 - 2025 © CBM
Conceito de Programa Armazenado
e Arquiteturas Associadas
● Referências:
○ Arquitetura e organização de computadores, William Stallings, Editora Pearson, 8a edição,
2010.
■ Capítulo 1 e 2.
○ First Draft of a Report on the EDVAC. John von Neumann, distributed by Herman Goldstine.
1945.
■ Arquivo pdf disponível no Ae.
○ Livro texto: Seção 1.13 Historical Perspectives & Further Reading
■ Disponível online em
[Link]
○ Livro texto: Seção 2.24 Historical Perspectives & Further Reading
■ [Link]
7
PCS 3612 - 2025 © CBM
Organização e Arquitetura
8
PCS 3612 - 2025 © CBM
ENIAC (Electronic Numerical Integrator
and Computer)
9
PCS 3612 - 2025 © CBM
Ideia
10
PCS 3612 - 2025 © CBM
Princípios e Arquitetura de Von Neumann
● Simplicidade: Operações
elementares sobre operandos
elementares
● Linearidade e Uniformidade:
Memória linear e uniforme
● Sequencialidade e Centralidade:
Processamento sequencial e
controle centralizado
● Unicidade: Programa e dados
armazenados numa mesma
memória
12
PCS 3612 - 2025 © CBM
Interpretador de instruções
13
PCS 3612 - 2025 © CBM
Arquitetura Harvard
14
PCS 3612 - 2025 © CBM Fonte: [Link]
Arquitetura Harvard
15
PCS 3612 - 2025 © CBM
Organização
16
PCS 3612 - 2025 © CBM
Como fica o código C = A + B nas diferentes
organizações?
17
PCS 3612 - 2025 © CBM
Comparando as organizações
18
PCS 3612 - 2025 © CBM
Evolução
19
PCS 3612 - 2025 © CBM
CISC (Complex Instruction Set Computer)
20
PCS 3612 - 2025 © CBM
RISC (Reduced Instruction Set Computer)
21
PCS 3612 - 2025 © CBM
CISC x RISC?
22
PCS 3612 - 2025 © CBM
Operações do hardware do computador
Seção 2.2 do livro-texto
23
PCS 3612 - 2025 © CBM
Operações Aritméticas
● Add e sub
○ Três operandos: duas origens e um destino
add a, b, c → a = b + c
24
PCS 3612 - 2025 © CBM
Exemplo
● Código em C:
f = (g + h) - (i + j);
25
PCS 3612 - 2025 © CBM
Operandos do hardware do computador
Seção 2.3 do livro-texto
26
PCS 3612 - 2025 © CBM
Registradores como operandos
27
PCS 3612 - 2025 © CBM
Registradores RISC-V
28
PCS 3612 - 2025 © CBM
Exemplo
Código em C:
f = (g + h) - (i + j);
29
PCS 3612 - 2025 © CBM
Operandos em memória
30
PCS 3612 - 2025 © CBM
Exemplo
● Código em C:
A[12] = h + A[8];
h em x21,
endereço base de A em x22
● Código compilado RISC-V:
31
PCS 3612 - 2025 © CBM
Registradores e Memória
32
PCS 3612 - 2025 © CBM
Constantes ou operandos imediatos
33
PCS 3612 - 2025 © CBM
Números com sinal e sem sinal
Seção 2.4 do livro-texto
34
PCS 3612 - 2025 © CBM
Extensão de sinal
35
PCS 3612 - 2025 © CBM
Representando instruções no computador
Seção 2.5 do livro-texto
36
PCS 3612 - 2025 © CBM
Representando instruções
37
PCS 3612 - 2025 © CBM
RISC-V: Instrução formato R
38
PCS 3612 - 2025 © CBM
Exemplo de instrução formato R
39
PCS 3612 - 2025 © CBM
RISC-V: Instrução formato I
40
PCS 3612 - 2025 © CBM
RISC-V: Instrução formato S
41
PCS 3612 - 2025 © CBM
Programas armazenados
42
PCS 3612 - 2025 © CBM
Operações lógicas
Seção 2.6 do livro-texto
43
PCS 3612 - 2025 © CBM
Operações lógicas
FIGURE 2.8 C and Java logical operators and their corresponding RISC-V instructions.
One way to implement NOT is to use XOR with one operand being all ones (FFFF FFFF FFFF FFFFhex). 44
Shift ● AND
○ útil para mascarar bits em uma palavra
● Instrução formato I ● OR
● imediato: quantas posições deslocar ○ útil para incluir bits em uma palavra
● Shift left logical ● XOR
○ Desloca à esquerda e preenche com bits 0 ○ útil para diferenciar bits
○ slli por i bits equivale a multiplicar por 2i
● Shift right logical
○ Desloca à direita e preenche com bits 0
○ srli por i bits equivale a dividir por 2i
(somente números sem sinal!)
45
PCS 3612 - 2025 © CBM
Instruções para tomada de decisões
Seção 2.7 do livro-texto
46
PCS 3612 - 2025 © CBM
Operações condicionais
47
PCS 3612 - 2025 © CBM
Compilando if
48
PCS 3612 - 2025 © CBM
Compilando while (loop)
Exit: …
49
PCS 3612 - 2025 © CBM
Blocos básicos
50
PCS 3612 - 2025 © CBM
Mais operações condicionais
51
PCS 3612 - 2025 © CBM
Signed vs. Unsigned
52
PCS 3612 - 2025 © CBM
53
PCS 3612 - 2025 © CBM
RISC-V (continuação)
54
PCS 3612 - 2025 © CBM
Suporte a procedimentos no hardware do
computador
Seção 2.8 do livro-texto
55
PCS 3612 - 2025 © CBM
Procedimento
● Ou função
● Forma de implementar abstração no software
● É uma sub-rotina armazenada que realiza uma tarefa com base nos
parâmetros que lhe são passados
56
PCS 3612 - 2025 © CBM
Passos necessários
57
PCS 3612 - 2025 © CBM
Instruções para chamada
58
PCS 3612 - 2025 © CBM
Uso de registradores
59
PCS 3612 - 2025 © CBM
Exemplo - procedimento folha
● Código em C:
int leaf_example (int g, int h, int i, int j)
{
int f;
f = (g + h) - (i + j);
return f;
}
● Argumentos g, … , j em x10, …, x13
● f em x20
● temporários x5, x6
● Precisa salvar x5, x6, x20 na pilha
60
PCS 3612 - 2025 © CBM
Exemplo - procedimento folha
61
PCS 3612 - 2025 © CBM
Dados na pilha
FIGURE 2.10 The values of the stack pointer and the stack (a) before, (b) during, and © after the procedure call.
The stack pointer always points to the “top” of the stack, or the last word in the stack in this drawing.
62
PCS 3612 - 2025 © CBM
Procedimentos aninhados
63
PCS 3612 - 2025 © CBM
Exemplo - Procedimento aninhado
Código RISC-V:
Código em C: fact:
int fact (int n) addi sp,sp,-8 Salva endereço de retorno e n na pilha
{ sw x1, 4(sp)
sw x10, 0(sp)
if (n < 1) return f; addi x5, x10, -1 x5 = n -1
else return n * fact(n - 1); bge x5, x0, L1 if n>=1, vá para L1
} addi x10, x0, 1 senão, valor de retorno é 1
addi sp, sp, 8 Libera a pilha, não restaura valores
jalr x0, 0(x1) Retorna
Argumento em x10
L1: addi x10,x10,-1 n = n -1
Resultado em x10 jal x1,fact chama fact(n-1)
addi x6,x10,0 move resultado de fact(n-1) para x6
lw x10,0(sp) restaura n de quem chamou
lw x1,4(sp) restaura endereço de quem chamou
addi sp,sp,8 libera pilha (pop)
mul x10,x10,x6 retorna n*fact(n-1)
jalr x0,0(x1) retorna 64
PCS 3612 - 2025 © CBM
O que deve ser preservado?
FIGURE 2.11 What is and what is not preserved across a procedure call. If the software relies on the global
pointer register, discussed in the following subsections, it is also preserved.
65
PCS 3612 - 2025 © CBM
Alocação da pilha
frame pointer (fp): um valor indicando o local dos registradores salvos e as variáveis
locais para um determinado procedimento.
FIGURE 2.12. Ilustração da alocação de pilha (a) antes, (b) durante e (c) após a chamada de um procedimento.
O frame pointer ($fp) aponta para a primeira palavra do frame, normalmente um registrador de argumento salvo, e o stack pointer ($sp) aponta para o topo da pilha.
A pilha é ajustada de modo a criar espaço para todos os registradores salvos e quaisquer variáveis locais residentes na memória. Como o stack pointer pode mudar
durante a execução do programa, é mais fácil para os programadores referenciarem variáveis por meio do frame pointer estável, embora isso também pudesse ser
feito por meio do stack pointer e um pouco de aritmética de endereços. Se não houver variáveis locais na pilha dentro de um procedimento, o compilador ganhará
tempo não atribuindo um endereço ao frame pointer, e depois, restaurando-o. Quando um frame pointer é usado, ele é inicializado usando o endereço que está no
$sp em uma chamada, e o $sp é restaurado usando o valor do $fp.
66
PCS 3612 - 2025 © CBM
Alocação de memória
● Text
○ código do programa
● Static data
○ variáveis globais
○ ex. variáveis estáticas em C, constantes,
vetores de tamanho definido
○ x3 (global pointer)
● Dynamic data
○ heap
○ ex. malloc em C, new em Java
● Stack
67
PCS 3612 - 2025 © CBM
Comunicando-se com as pessoas
Seção 2.9 do livro-texto
68
PCS 3612 - 2025 © CBM
Caracteres
FIGURE 2.15 ASCII representation of characters. Note that upper- and lowercase letters differ by exactly 32; this observation
can lead to shortcuts in checking or changing upper- and lowercase. Values not shown include formatting characters. For example,
8 represents a backspace, 9 represents a tab character, and 13 represents a carriage return. Another useful value is 0 for null, the
value the programming language C uses to mark the end of a string.
69
PCS 3612 - 2025 © CBM
RISC-V: Operações load/store
70
PCS 3612 - 2025 © CBM
Exemplo: copiando string
71
PCS 3612 - 2025 © CBM
Endereçamento no RISC-V para
operandos imediatos e endereços
Seção 2.10 do livro-texto
72
PCS 3612 - 2025 © CBM
Constantes de 32 bits
73
PCS 3612 - 2025 © CBM
Endereço de desvios condicionais
74
PCS 3612 - 2025 © CBM
Endereço de saltos (jump)
● Alvo de jump and link (jal) usa imediato de 20 bits para maior alcance
● Formato UJ:
75
PCS 3612 - 2025 © CBM
RISC-V: Modos de endereçamento
76
PCS 3612 - 2025 © CBM
RISc-V: formatos de instrução
FIGURE 2.19 Four RISC-V instruction formats. Figure 4.14.6 reveals the missing RISC-V formats for conditional branch (SB) and unconditional jumps (UJ),
whose formats match the lengths of the fields in the S and U types, but the bits are swirled around. The rationale for SB and UJ makes more sense once you
have an understanding of hardware given in Chapter 4, as SB and UJ simplify the hardware but give the assembler a little more to do.
77
PCS 3612 - 2025 © CBM
Sincronização
Seção 2.11 do livro-texto
78
PCS 3612 - 2025 © CBM
Sincronização
79
PCS 3612 - 2025 © CBM
Sincronização no RISC-V
80
PCS 3612 - 2025 © CBM
Exemplo
81
PCS 3612 - 2025 © CBM
Traduzindo e iniciando um programa
Seção 2.12 do livro-texto
82
PCS 3612 - 2025 © CBM
Hierarquia de tradução para C
Estático
FIGURE 2.20 A translation hierarchy for C. A high-level language program is first compiled into an assembly language program and then assembled into an
object module in machine language. The linker combines multiple modules with library routines to resolve all references. The loader then places the machine code
into the proper memory locations for execution by the processor. To speed up the translation process, some steps are skipped or combined. Some compilers
produce object modules directly, and some systems use linking loaders that perform the last two steps. To identify the type of file, UNIX follows a suffix convention
for files: C source files are named x.c, assembly files are x.s, object files are named x.o, statically linked library routines are x.a, dynamically linked library routes
are [Link], and executable files by default are called [Link]. MS-DOS uses the suffixes .C, .ASM, .OBJ, .LIB, .DLL, and .EXE to the same effect.
83
PCS 3612 - 2025 © CBM
Produzindo um objeto
84
PCS 3612 - 2025 © CBM
Ligando objetos (estático)
85
PCS 3612 - 2025 © CBM
Carregando um programa
86
PCS 3612 - 2025 © CBM
Ligações dinâmicas
87
PCS 3612 - 2025 © CBM
DLLs e lazy procedure
89
PCS 3612 - 2025 © CBM
Um exemplo de ordenação em C para juntar
tudo
Seção 2.13 do livro-texto
90
PCS 3612 - 2025 © CBM
Exemplo
91
PCS 3612 - 2025 © CBM
Função swap (folha)
Função em C
swap:
slli x6,x11,2 // reg x6 = k * 4
add x6,x10,x6 // reg x6 = v + (k * 4)
lw x5,0(x6) // reg x5 (temp) = v[k] Código RISC-V
lw x7,4(x6) // reg x7 = v[k + 1]
sw x7,0(x6) // v[k] = reg x7
sw x5,4(x6) // v[k+1] = reg x5 (temp)
jalr x0,0(x1) // return to calling routine
92
PCS 3612 - 2025 © CBM
Função sort
Função em C
Código RISC-V
93
PCS 3612 - 2025 © CBM
Função sort: corpo
94
PCS 3612 - 2025 © CBM
Função sort: procedimento
95
PCS 3612 - 2025 © CBM
sort - procedimento completo
96
PCS 3612 - 2025 © CBM
Desempenho e otimizações do compilador
FIGURE 2.26 Comparing performance, instruction count, and CPI using compiler optimization for Bubble Sort.
The programs sorted 100,000 32-bit words with the array initialized to random values. These programs were run on a Pentium 4 with a clock rate of
3.06 GHz and a 533 MHz system bus with 2 GB of PC2100 DDR SDRAM. It used Linux version 2.4.20.
97
PCS 3612 - 2025 © CBM
Desempenho: linguagem e algoritmos
FIGURE 2.27 Performance of two sort algorithms in C and Java using interpretation and optimizing compilers relative to unoptimized
C version.
The last column shows the advantage in performance of Quicksort over Bubble Sort for each language and execution option.
These programs were run on the same system as in Figure 2.29. The JVM is Sun version 1.3.1, and the JIT is Sun Hotspot version 1.3.1.
● Bubble sort:
○ código em C não otimizado é 8,3 vezes mais rápido do que código em Java interpretado
○ Compilador JIT torna execução em Java 2,1 vezes mais rápida que código em C não
otimizado e desempenho melhor que C otimizado
● Quicksort:
○ diferença de algoritmo predomina, baixo tempo de execução para código compilado
98
PCS 3612 - 2025 © CBM
Então…
99
PCS 3612 - 2025 © CBM
Conteúdo das seções 2.14 (Arrays versus
ponteiros) e 2.15 (Compiling C and Interpreting
Java) não fazem parte do escopo da disciplina
100
PCS 3612 - 2025 © CBM
Arrays versus ponteiros
101
PCS 3612 - 2025 © CBM
Vida real: instruções MIPS
Seção 2.13 do livro-texto
102
PCS 3612 - 2025 © CBM
MIPS
● Projetado na década de 80
● RISC-V possui projeto similar
○ acesso a memória via load/store
○ 32 registradores, sendo um fixo em zero
○ instruções com 32 bits de largura
○ ambos possuem bne e beq
● Principal diferença: instrução de desvio
○ MIPS: somente bne e beq
○ Para outras comparações: usar comparação
■ slt ou sltu
● Conjunto de instruções completo do MIPS é maior do que RISC-V
103
PCS 3612 - 2025 © CBM
Vida real: instruções ARMv7 (32 bits)
Seção 2.17 do livro-texto
104
PCS 3612 - 2025 © CBM
Formato de instruções ARM, RISC-V e MIPS
105
PCS 3612 - 2025 © CBM
ARMv7 (32 bits)
106
PCS 3612 - 2025 © CBM
Instruções equivalentes
FIGURE 2.31 ARM register–register and data transfer instructions equivalent to the RISC-V core.
Dashes mean the operation is not available in that architecture or not synthesized in a few instructions. If there are several choices of
instructions equivalent to the RISC-V core, they are separated by commas. ARM includes shifts as part of every data operation instruction, so a
shift with superscript 1 is just a variation of a move instruction, such as lsr1. Note that ARM has no divide instruction.
107
PCS 3612 - 2025 © CBM
Modos de endereçamento
108
PCS 3612 - 2025 © CBM
Vida real: instruções ARMv7 (64 bits)
Seção 2.18 do livro-texto
109
PCS 3612 - 2025 © CBM
ARMv8 (64 bits)
110
PCS 3612 - 2025 © CBM
Vida real: Instruções x86
Seção 2.18 do livro-texto
111
PCS 3612 - 2025 © CBM
Instruções x86
112
PCS 3612 - 2025 © CBM
Instruções x86
113
PCS 3612 - 2025 © CBM
Instruções x86: mais evolução
114
PCS 3612 - 2025 © CBM
Instruções x86: e ainda mais...
115
PCS 3612 - 2025 © CBM
Registradores do 386
116
PCS 3612 - 2025 © CBM
Instruções com registradores
FIGURE 2.35 Instruction types for the arithmetic, logical, and data transfer instructions.
The x86 allows the combinations shown. The only restriction is the absence of a memory-memory mode. Immediates may be 8, 16, or 32 bits in
length; a register is any one of the 14 major registers in Figure 2.33 (not EIP or EFLAGS).
117
PCS 3612 - 2025 © CBM
Modos de endereçamento
FIGURE 2.36 x86 32-bit addressing modes with register restrictions and the equivalent RISC-V code. The Base plus Scaled Index
addressing mode, not found in RISC-V or MIPS, is included to avoid the multiplies by 4 (scale factor of 2) to turn an index in a register into a byte
address (see Figures 2.26 and 2.28). A scale factor of 1 is used for 16-bit data, and a scale factor of 2 for 32-bit data. A scale factor of 0 means
the address is not scaled. If the displacement is longer than 12 bits in the second or fourth modes, then the RISC-V equivalent mode would need
more instructions, usually a lui to load bits 12 through 31 of the displacement, followed by an add to sum these bits with the base register. (Intel
gives two different names to what is called Based addressing mode—Based and Indexed—but they are essentially identical and we combine
them here.)
118
PCS 3612 - 2025 © CBM
Algumas instruções
119
PCS 3612 - 2025 © CBM
Formato de instrução
120
PCS 3612 - 2025 © CBM
Codificando instruções
FIGURE 2.40 The encoding of the first address specifier of the x86: mod, reg, r/m. The first four columns show the encoding of the 3-bit reg
field, which depends on the w bit from the opcode and whether the machine is in 16-bit mode (8086) or 32-bit mode (80386). The remaining
columns explain the mod and r/m fields. The meaning of the 3-bit r/m field depends on the value in the 2-bit mod field and the address size.
Basically, the registers used in the address calculation are listed in the sixth and seventh columns, under mod = 0, with mod = 1 adding an 8-bit
displacement and mod = 2 adding a 16-bit or 32-bit displacement, depending on the address mode. The exceptions are 1) r/m = 6 when mod = 1
or mod = 2 in 16-bit mode selects BP plus the displacement; 2) r/m = 5 when mod = 1 or mod = 2 in 32-bit mode selects EBP plus displacement;
and 3) r/m = 4 in 32-bit mode when mod does not equal 3, where (sib) means use the scaled index mode shown in Figure 2.39. When mod = 3,
the r/m field indicates a register, using the same encoding as the reg field combined with the w bit.
121
PCS 3612 - 2025 © CBM
Implementando IA-32
122
PCS 3612 - 2025 © CBM
Demais Instruções RISC-V
Seção 2.20 do livro texto
123
PCS 3612 - 2025 © CBM
Demais instruções RISC-V
FIGURE 2.41 The remaining five instructions in the base RISC-V instruction set architecture.
124
PCS 3612 - 2025 © CBM
RISC-V: arquitetura base e extensões
FIGURE 2.42 The RISC-V instruction set architecture is divided into the base ISA, named I, and five standard
extensions, M, A, F, D, and C. RISC-V International is developing many other optional instruction extensions. Unlike most
architectures, the RISC-V software stack only assumes the base architecture (I), with other extensions optional that are only
issued by the compiler if the processor includes those options.
125
PCS 3612 - 2025 © CBM
Multiplicação de matrizes em Python
Seção 2.21 do livro texto
126
PCS 3612 - 2025 © CBM
Multiplicação de matrizes em Python
● Código em Python
○ sem utilizar bibliotecas otimizadas, pois o objetivo é mostrar as diferenças de linguagens entre
Python e C
127
PCS 3612 - 2025 © CBM
Escrevendo em C…
FIGURE 2.43 C version of a double-precision matrix multiply, widely known as DGEMM for Doubleprecision GEneral Matrix Multiply (GEMM).
128
PCS 3612 - 2025 © CBM
x86 assembly para o corpo
FIGURE 2.44 The x86 assembly language for the body of the nested loops generated by compiling the unoptimized C code in Figure
2.43 using gcc with -O3 optimization flags.
129
PCS 3612 - 2025 © CBM
Desempenho
FIGURE 2.45 Speed of DGEMM in Figure 2.43 over the Python program in Section 1.10 as we increase optimization levels.
130
PCS 3612 - 2025 © CBM
Falácias e armadilhas
Seção 2.22 do livro texto
131
PCS 3612 - 2025 © CBM
Falácias e armadilhas
132
PCS 3612 - 2025 © CBM
Falácias e armadilhas
134
PCS 3612 - 2025 © CBM
Comentários finais
Seção 2.23 do livro texto
135
PCS 3612 - 2025 © CBM
Comentários finais
● Princípios de projeto
○ simplicidade favorece a regularidade
○ menor é mais rápido
○ um bom projeto exige bons compromissos
● Abstrações: camadas de software/hardware
○ Compilador, assembler, hardware
● RISC-V: exemplo típico de conjunto de instruções RISC
○ diferente do x86
○ similar MIPS e ARM v8
136
PCS 3612 - 2025 © CBM
RISC-V: Conjunto de instruções
& classes de instruções
FIGURE 2.48 RISC-V instruction classes, examples, correspondence to high-level program language constructs, and
percentage of RISC-V instructions executed by category for the average integer and floating point SPEC CPU2006
benchmarks. Figure 3.24 in Chapter 3 shows average percentage of the individual RISC-V instructions executed.
137
PCS 3612 - 2025 © CBM
PCS3612: Organização e Arquitetura de Computadores I
[Link]