0 ratings 0% found this document useful (0 votes) 7 views 50 pages Microprocessor Design Changes
The document discusses the evolution of processor design over the past 50 years, highlighting key shifts such as the RISC revolution and the transition to Chip Multithreading (CMT). It emphasizes that change is constant in technology, driven by the need for innovation and efficiency, and outlines the historical patterns of experimentation and consolidation in the industry. The talk also addresses the importance of Moore's Law and advancements in processor architectures as critical factors influencing future developments in computing.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content,
claim it here .
Available Formats
Download as PDF or read online on Scribd
Go to previous items Go to next items
BUCH em eee f
Laceleces tea t+)
As the title says, this talk is on the changes that have occurred and continue to occur
in processor design. In order to cover what is now over 50 years of computing
history in general, including 30+ years of microprocessor design in specific, I'll
keep things at a a fairly high level. So, although there are a fair number of technical
concepts involved that may not be familiar to everybody here — terms like “MIPS”
and “RISC” and “IPC” and so on — we'll be able to explain these concepts as we 20
along. If you feel you're getting lost at any point, just stop me and I'll be happy to
clarify not only the specific idea in question, but how it fits into the bigger picture
of the Changing Parameters of Processor Design.
Like all worthwhile stories, this one has a moral, in fact two morals. The first moral
is simply that things do change. 1 know that seems pretty simplistic, but change is
by nature unsettling, so there's always a lot of resistance to it — especially from
those that benefit from the way things are. Hence, whatever the point you happen to
be at, there always are some ready to declare that, finally, change is now officially
over and the status is going to remain quo from now on. So the first basic point to
this talk is that those kinds of statements aren't true — things not only always can
change, but they do change and will continue to change on a fairly regular basis.
The second basic point is that, although I think there's always an element of fashion
in any kind of change, change in microprocessor design is not simply a matter of
fashion. There's also a real sense in which progress has been and will continue to
be made. Indeed, at bottom, what drives change in processor design is the periodic
need to consider what is really quite a radical alternative to the way processors have
been designed hitherto, in order to continue to make progress.| The Shifting Landscape
‘The Path of Progress
| Technology Drivers
The RISC Revolution
The Irrelevance of Itanium
The Commodity Shift
The Coming Transition to CMT
The Relevance of SPARC
The Performance Trap & How to Escape It
Space-ffficient Performance
Time-Efficient Performance
Building CMT Processors
Requirements
Benefits
So here's the outline of what we'll be covering. First, we'll look at an overview of
the history of computing, and set the major design points we need to consider. As
you will see, viewed at a very high level, there is a sort of pattern to change that can
help us understand the shifts that have occurred in the computing landscape. We'll
then look at the way in which these periodic changes add up to overall progress,
and the major technical drivers behind the progress that has been made.
We'll then run through the last two major design shifts in some detail, the RISC
revolution of the 1980s, and what I call the Commodity Shift of the 1990s. Since I
know some people were and perhaps still are expecting Intel's initiative around
their new Itanium processor to be the next big shift in the processor landscape, I
also want to take a few minutes to talk about Itanium and why I believe it's destined
to end up as a footnote in computing history, rather than as the next main chapter.
That will set the stage for what I believe is going to be the next main chapter,
namely Chip MultThreading, CMT, or what Sun calls “throughput computing.”
Since processor progress is, at bottom, about processor performance, that will
require that we spend some time talking about what now is impeding our ability to
reach higher levels of performance, and how CMT overcomes those problems with
designs that make efficient use of both a processor's physical resources and its
compute cycles — what I call space- and time-efficient computing.
Finally, we'll spend a little time looking at what's needed to build CMT-style
processors, and the benefits that we expect to result from these sorts of designs._— eu |
The Shifting Landscape
Just when you think it's all over - technology takes a new turn!
System Design
| companies Innovation
Semiconductor
companies
[Hi diversification (experimentation) Economies
Consolidation electing the winners) ___ OF Scale
Here's the big picture that we'll be covering. For present purposes, I've
conceptualized change in processor design as a periodic shift in the balance
of power between system companies and semiconductor companies, with
the respective forte of each side in this ongoing struggle for dominance
being innovative new designs by system companies, pitted against the
massive economies of scale that semiconductor companies can leverage.
To the extent this picture is accurate, one thought here is that, in the end,
both of these opposed forces really are necessary in order to keep moving,
ahead. That is, without periodic infusions of design innovation, economies
of scale ultimately turn flat and sterile. Whereas, innovative new designs
eventually must find their way into volume markets, or become
increasingly irrelevant.
Within each phase shift, there also seems to be a consistent two-part
pattern. First, there's a period of experimentation, characterized by a lot of
different approaches and players. Second, there's a period of consolidation,
which one or two dominant designs emerge. (That's the point where
there's always a strong temptation, at least on the part of the winners, to
declare the whole game is over, and from here on it's going to be just a
whole lot more of essentially the same thing.)Just when you think it's all over - technology takes a new turn!
ime
by System Di
companies Innovation
yor
Eeompebles
Semiconductor
companies
Economies
Diversification experimentation)
Consolidation (selecting the winners) of Scale
These are some sample names associated with the first era of computing
When the commercial computing industry got started back in the early
1950s, of course, there were no semiconductor companies as yet (indeed
semiconductor technology was still in its infancy, experimenting with the
germanium-based transistor). So the early computers were strictly a
systems play, with many companies trying their hand at some really very
interesting design variations, before IBM emerged as the dominant player
with their 360/370 line of machines.
Although I'm not showing it here, mainly due to lack of space, there was a
very important design shift from mainframes to minicomputers, during the
initial technological period when CPUs were still build out of discrete
components. This lead to the rise of a whole new group of companies in
the 1960s, supplementing and in some cases supplanting the players of the
1950s: the Digitals and Primes and Data Generals of the computing
landscape.Just when you think it's all over - technology takes a new turn!
ooh
we System Design
orem —- GOMpanies Innovation
yor
fre Semiconductor
es" _ companies a
conomies
Diversification (experimentation)
consolidation (electing the wimers) __ Of Seale
The development of the microprocessor in the 1970s had to wait on the
emergence of the semiconductor industry in the 1960s, and an entirely new
class of company, created to develop devices around the new possibilities
inherent in the integrated circuit or IC. The economics of these companies,
driven by the need to find volume markets for relatively inexpensive
silicon-based chip components, were very different from the economics of
even a minicomputer maker.
The first processor designs to emerge from semiconductor companies,
termed microprocessors in part to distinguish them from minicomputers,
existed for many years under the radar of established system companies,
with little apparent relevance to even 16-bit minicomputers, let alone to the
far more sophisticated supercomputer and mainframe designs created by
leading computer architects like Seymour Cray and Gene Amdahl. By the
middle of the 1980s, however, two 32-bit descendants of these original 4-
and 8-bit designs had come to dominate the major growth sectors of the
computing landscape: the Intel 80386 in the home and business PC, and the
Motorola 68000 in scientific and engineering workstations.Just when you think it's all over - technology takes a new turn!
ve Po
System 3, Design
companies ®* Innovation
ion
oven
ars
Tooele,
"Ss Semiconductor
«<"" companies
Diversification (experimentation)
consolidation electing the winers) __ Of Scale
Economies
But just when it seemed as if the mantle of leadership in processor design
was firmly settled on the shoulders of the semiconductor industry (albeit,
some doubt remained about which semiconductor company eventually
‘would emerge on top), the landscape shifted again. The driver of this next
major change was a new design point that emerged commercially in the
mid-1980s, called “RISC.” This is the first shift I want to look at in some
detail, so we won't discuss it here other than to point out that it enabled
system companies to wrest the initiative in processor design away from the
semiconductor companies.
Indeed, one of the most striking aspects of the “RISC revolution” is that,
while several semiconductor companies made serious efforts to introduce
competing RISC designs of their own — including Intel with the i860,
Motorola with the 88000, and AMD with the 29000 — not one of the
semiconductor-sponsored design initiatives survived. Indeed, in the high-
performance arena, only RISC designs that either were originally created
by a systems company (like the HP PA-RISC, the Sun SPARC, the DEC
Alpha, and the IBM POWER) or quickly acquired a systems “owner” (like
the Fairchild/Intergraph Clipper and the MIPS/SGI MIPS) survived for any
length of time.System
companies
era
‘isu Semiconductor
«companies
Diversification (experimentation)
Economies
of Scale
‘Through the RISC revolution, we have enough separation from past events
to enjoy the advantage of historical perspective. The next shift, however,
brings us to events that are still in the process of unfolding. So we move
here from retrodiction to prediction, from simple historical review to
forecasting events that still are in the process of playing themselves out.
The advantage of a pattern like the one laid out here is that it can help us
gain perspective on current events as they happen, without needing to wait
for the judgment of history to decide on probable outcomes, or to
distinguish likely winners from mere glamorous pretenders.
But even armed with a pattern, there's no gainsaying the point that
predicting the future remains a highly inexact science. Nonetheless, even
though I can't tell you what will happen over the course of the next several
years with anything like the same confidence I can speak about what has
happened to date, what I can do is present one possible future scenario for
your consideration, a scenario that I think not only has great potential to
revitalize computing once again, but that also makes considerable sense in
light of what has happened to date in the history of processor design, and
which seems to me to be perfectly aligned with the essential forces needed
to drive the industry into the future.Discrete Logie
Integrated circuits (cs)
A Protteof ve isruptve technology”
(radical reduction in cost with signifcat Ita
loss of neon}
Progress driven by Moore's Law » New Design Poin
Microprocessors
The previous slides make the point that processor design does change on a periodic
basis, with the initiative shifting back and forth between semiconductor companies
and system companies. Further, just when you finally get comfortable with the way
things are, suddenly everything changes. What the previous slides don't show is the
fact that the shifting landscape of processor design is not a reflection of some
underlying base instability, akin to periodic geological or political upheavals, but is
instead a function of progress, where the path of progress is measured by steadily
lower costs combined with an overall increase in functionality (as measured by a
variety of indicators, from larger address sizes to lower power consumption).
The one true “disruptive technology” in this chronicle of progress is the development
of the IC, which led to the creation of the microprocessor itself. The early micros, of
course, provided far less functionality than contemporary system processors, but with
continued development not only caught up with minis and even mainframes, but
eventually surpassed them — at a far lower cost for the same level of ability.
Subsequent changes have not been “disruptive” in this strict sense, but rather have
been new design points that lowered the effective cost of computing for the same
level of functionality (or what is the same thing, provided more functionality for the
same level of cost). Since we will be looking at each of the new design points
starting with RISC in some detail, I won't dwell further on this summary, other than
to emphasize that all real shifts in the processor landscape are marked by sudden and
significant movements along this “path of progress.”prog ereantiog after 3 year ail ontrack
tnamicroprocesor forthe 3B
double evry 2 years “lz
Since thay were tvented
al
| “The doubling will stow
down. You really get bit
| by the fact that materials
are made of atoms.”
“Gordon Moore, 7/5/2002
Hopefully, that's sufficient history to get us started. Before turning to the
first of the new processor design points I want to consider in some detail,
the RISC revolution, we need to look at two of the basic technical drivers of
movement along the path of progress.
‘The first driver is Moore's Law, which has governed progress in the
semiconductor industry since the invention of the planar transistor by
Fairchild back in 1959 — the key development that made it possible to
“print” ICs much like photographs. Moore's Law is a simple doubling
algorithm that says how many transistors will fit on a semiconductor chip
by any given year. In the initial 1965 version of his Law, Gordon Moore,
who at that time was a Director of R&D at Fairchild, said the number of
transistors per chip would double every year, starting from 1 in 1959. In
1975, Moore revised his Law downward to a doubling every 2 years.
This latter curve has been maintained by the semiconductor industry for an
astonishing period now, over 30 years, and looks to remain on track until
sometime after 2010 (before the doubling period stretches out for a second
time). The moral here for us is that we're headed to processors that will
contain 2-4 billion transistors by the end of the decade. With transistor
budgets now poised to go over the 1B milestone, entirely new kinds of
processor designs are becoming feasible for the first time. That's the first
key technology driver to keep in mind as we move forward.The more instructions a processor executes in a glen unit of time, the higher |
its performance |
MIPS or Millions of Instructions executed Per Second is the best known and |
most basic performance metric A MIPS rating fora processors simple to calculate | |
slven jst two numbers: |
MIPS = avg. instrs. executed per clock (IPC) < clock rate (in MHz)
{3,2 processor that on average, executes exactly 2 instructions every clock oe, and whose
ioc runs ats Giz (= 2000 Mit), would execute 200,000,000 htrucons in second, for 8
performance rating of 2,000 (2.1000) MIPS
2,000 MIPS = 2.0 IPC X 000 MH
‘Note: MIPS ial rete measure of perermace fo roctsrs wh the sae nsructon et eter OA.
[isnot aac measure as fret ance he yal etc ove a ny sotore ees ert on
‘he gpeatinsracion anchor 8h Tass eaimate complaint by OSC poorest oops PES as
{hat the aber MPs ratings toutes RSC pacer wer ace, sces lon ene eset
{ecompined ies actual wer than lon compen CSE nseucon The eeu eats shite Meare
‘iraatneprtcance, om te tine kates tw rent procs io eect ih sue tamer tarot
{oie mer tatesewe cferem oceans to eacte he sine pay foun fh
{hesitate nh Fe benchnak si, rrciadie whch ot eb MIS tgs wth
SPecnumber othe nvr bse maar fave paceso’ pata
The second key technology that continues to advance relentlessly, besides
semiconductors, is processor architectures. We will be looking at several
different types of processor architectures in some detail, but first we need
to be clear on the basic goal behind most architectural advances, namely,
their ability to enable higher levels of performance.
The simplest way to think about processor performance is in terms of the
number of instructions a processor can execute in a given unit of time, say
a second. This is a function of just two basic parameters.
The first parameter that controls performance is how many instructions a
processor can execute, on average, in each clock cycle. This is typically
abbreviated as “IPC” for average “Instructions Per Clock” executed.
The second parameter that controls performance is how fast the clock ticks.
Processor clock frequency is typically given in megahertz (MHz) or
gigahertz (GHz), ice, as either millions or billions of clock cycles a second.
Given an IPC number for a processor and its clock frequency in MHz,
multiplying the two numbers together will produce a MIPS rating for the
processor, specifying how many Millions of Instructions Per Second that
processor can execute.Osun
From the definition of performance (in terms
of MIPS), it follows that there are just two
possible ways to improve processor
performance (MIPS rating)
Increase IPC 0 |
Increase clock rate
Clock Rate |
While it might seem that processor performance ought to be a lot more
complicated that just MIPS, at the top level, that’s all there is to the subject.
When processor architects think about how to design a faster processor,
everything they think about boils down to one or the other of these two
fundamental parameters. Either they need to figure out a way to execute
more instructions, on average, in each clock cycle, or they need to figure
out a way to make the clock tick faster. All the complexity of processor
design lies in the details of achieving one or the other of these two goals.
Moore's Law and the Performance Vector are really all we need by way of
background. We're ready now to turn to the first basic shift in the way
microprocessors were designed, the so-called “RISC revolution.”a es |
The RISC Revolution |
| The Mlcroprocessor Era
The first new design point
1971. 1 microprocessor CPU
1985 2" com |
abit
36D spit [RISC sin
|
|
| aot
* eo Ce |
Radically simplify Instruction Set Architectures (ISAs)
rect execution of instructions (no microcode)
‘Microarchitecture improvements (pipelining, etc.)
Higher levels of onvchip integration
Improvements in memory latency (caches) & bandwidth
| less design complexity, fewer transistors
| (smaller design teams/faster design cycles, smaller die)
Functionality
The revolutionary thing about RISC, which is an acronym for "Reduced
Instruction Set Computing,” is that it fundamentally changed the
price/performance point for microprocessors. As the RISC acronym suggests,
the basic idea was to greatly simplify both the number and kinds of instructions
that a processor needed to execute, focusing specifically on the sorts of
relatively simple operations that effectively can be implemented in hardware
(ike adding or comparing two binary numbers). This might seem like a pretty
simple notion, but it stood in sharp contrast to the reigning design style of the
early 1980s, dubbed “CISC” (by RISC proponents), for “Complex Instruction
Set Computing.” CISC processors, like the Intel x86, the Motorola 68K, the
DEC Vax, the IBM 360/370, etc., had 100s of instructions, many of which
involved a lengthy sequence that started with fetching one or more values from
memory, included operating on them in some way, and ended by returning the
result to memory. The CISC idea was that the more powerful the instruction set,
the more potent the processor. The RISC idea that higher performance would
result, instead, from simpler instructions was really very radical.
Building simpler processors had a lot of other benefits, too. Simplifying the
core compute engine made transistors available for on-chip integration of other
key functional units, like MMUs and FPUs and caches, which also carried
significant performance benefits. And, since the designs were less complex, they
could be created more quickly using fewer engineers.| Why RI RISC Was Revolutionary _
cise RISC
i386 R2000
Year 1985 1985
Technology | 1.5. cmos | 2.0 CMOS
Die Size 103 mm2_| 80 mm2
Price
Transistors 115,000
Cache . | 32 KB ex.
sszcrca | 144¢PGA | An order of magnitude
Power aa ra improvement in p/p
Performance | 2.2 Speceo | 10.1 Specso | Performance
|
|
BH Better (hot) i comparable Wl Worse (cold) |
Here's a telling comparison between two processors, both introduced in 1985, the
Intel 80386, and one of the first commercial RISC microprocessors, the MIPS
R2000. The RISC design required just over 100,000 transistors to implement, vs.
nearly 300,000 transistors needed to build the 386. Hence, it was possible to build
the R2000 in an older, more cost-effective semiconductor technology, and still end
up with a smaller, cheaper die. Further, the R2000 was able to use half the external
cache of the 386 and still generate almost 5 times the performance when operating
at exactly the same frequency, in a very similar package, using no more power.
Multiply the cost savings times the performance gain, and RISC offered an order of
‘magnitude improvement (or more) in the price/performance ratio of
microprocessors. No wonder it was revolutionary! And no wonder that Intel
started to have doubts about the ability of their x86 line to stay competitive with
the new RISC alternatives that soon started to populate the processor landscape.
RISC was more than just a new design point for microprocessors. It also was a
liberation. Because RISC designs were simple, deciding to create a RISC
processor was, relatively speaking, a low-cost project. Even comparatively small
system companies, like Sun (in 1984), which previously had to rely on
semiconductor companies for CPUs, were able to take control of their own
processor destiny, driving designs to meet their system-level goals rather than the
lowest-common-denominator requirements of some volume chip market.RISC Performance
Basic Performance Equation per Task
{factoring indifference in RISC vs. CISC ISA)
. Instructions in Task
Task Time = we X Cycle Time
| Total number of instructions in a task ~ 2X worse than CISC
| fined by SA stare (omer, 05)
| ‘Average number of instructions executed Per Clock cycle 10X better than CISC
focus of micoarhitectua improvements wiser ut, 000 econ, ber ace
{lock cycle time - equal to or better than CISC
fe ecg basi ine age, agely eerie by semianducter technology
In Theory, RISC at least 5X faster than CISC
| Gain driven by improvements in IPC
Initial Goal - single-cycle execution (IPC =2)
Long-Term Goal - continually increasing IPC (2, 3, 4 ....)
Result is wider 8 wider superscalar designs with more outoforder execution, etc.
Here's the secret to the revolutionary performance levels achieved by RISC
processors, compared to their older CISC rivals. To be fair, I'm using a
more complicated formula that counts not simple MIPS, but factors in the
difference between CISC and RISC instruction sets. How long it takes to
accomplish a particular task is a question of the total number of instructions
a machine has to execute to get the job done, divided by the average number
of instructions the machine executes each clock cycle, times the clock cycle
time. [Plug in some numbers and give an example of using this formula]
Since RISC instruction sets are simpler than CISC instruction sets overall, it
is necessary to execute more RISC instructions than CISC instructions to
accomplish the same amount of work. This is typically about a factor of 2
to 1. So RISC machines start off with a performance handicap, but make up
for it by their ability to efficiently dispatch instructions. The original Vax,
for example, needed about 10 clock cycles to execute an instruction,
whereas the early RISC machines all set “single cycle execution” as their
goal. Assuming no difference in clock rate between a RISC and CISC
processor, that meant RISC designs were about 5 times faster overall that
CISC designs. Which is pretty much what we saw in the previous example
comparing the 386 with the R2000. After reaching the goal of “single cycle
execution,” RISC processors moved in the direction of increasingly wide
and complex superscalar designs, in an effort to continue to drive up IPC.A Limit to RISC Performance |
Sew weap
ase
The Superscaler Barrier
Problem: The complexity of dynamic superscalar
‘out of order instruction issue increases nondinearly
‘with the number of instructions issued each cycle,
with more and more impact on cycle time
"
F
2
=
i:
i:
Bo
cL
a
gs
;
| Time
By the mid 1990s, RISC processors had accomplished a great deal. They
had progressed from 32-bit to 64-bit architectures, they were able to issue 4
or more instructions every clock cycle, and they had incorporated all the more
practical innovations that had been proposed for higher performance,
including out-of-order execution, register renaming, sophisticated branch
prediction algorithms, set-associative caches, and both hardware and software
prefetch. There were two unintended consequences of these architectural
refinements, however.
The first was to effectively eliminate any meaningful distinction between
RISC and CISC along the complexity scale. The 300% difference between a
100,000 transistor RISC processor and a 300,000 transistor CISC processor
was highly significant, The 5% difference between a 3.3M transistor RISC
processor and a 3.5M transistor CISC processor was totally insignificant.
‘The second was to make RISC processors victims of their own success. Each
step up in performance required progressively more design effort, with
steadily diminishing results. In particular, the drive to higher and higher
levels of IPC tended to run out of steam around the 6-issue mark. At this
point, the detrimental effect on clock rate of a still wider-issue design tended
to outweigh any performance gain from the resulting improvement in IPC.complexities of parallel issue to software
Simple inorder hardware design
Unlimited IPC scaling
(to the extent new parallelizing compilers
‘can deliver independent parallel instructions)
'No impact on cycle time ~
(clock speeds can improve as fast as semiconductor
technology delivers shorter transistor gate lengths) |
The Superscaler Barrier
=
a
2
3a
:
£3
z>
ge
Bs
EB:
£3
As the limit to the ability of RISC processors to continue to push out along the
parameter of constantly increasing IPC became clear in the early 1990s, processor
architects began to cast about for new architectural ideas that could avoid the
superscalar barrier. For better or worse, there was a candidate ready and waiting
just off-stage, in VLIW or Very Long Instruction Word architectures, VLIW
theory was a close contemporary of RISC theory. Indeed, the two most notable
‘VLIW start-ups of the 1980s were both founded in 1984, the same year MIPS was
founded to pursue RISC architectures and Sun started work on SPARC.
Although there are a number of parallels between basic RISC vs. VLIW concepts,
including the notion of pushing complexity out of hardware and into compilers,
the notable difference for present purposes is that VLIW is a strategy specifically
intended to reach very high levels of ILP. In theory, the only limit to ILP in
VLIW architectures is the ability of compilers to detect and organize parallel
instructions in code. VLIW hardware is designed to allow arbitrary increases in
ILP, to the limit of current transistor budgets to accommodate parallel execution
units, with absolutely no impact on clock cycle time.
As the superscalar barrier loomed closer and closer, it was perhaps only natural
therefore that VLIW would come to the fore as a solution to this problem. The
most surprising thing was the alliance created to promote this solution, between
HP, one of the RISC pioneers, and Intel, the most prominent critic of RISC.an ie ss oa =
Osun
Three EPIC Truths
4 1. Evenin theory, offers no performance discontinuity - rather, 2
| _ way of extending the RISC performance curve (based on steadily
increasing IPC) past the limits of dynamic superscalar issue.
2. In actual practice, provides no performance breakthrough ~ rather,
EPIC processors aiso stall at a peak issue rate of 6 instructions/clock
{due to lack of ILP in typical code).
3. Ineality, no design simplification ~ rather, new complexities
in the EPic architecture outweigh any gains from
eliminating dynamic outoforder superscalar issue.
5
4
3
2
a
a0
3
The EPIC Era
}- The RISC Era The Superscaler Barrier
3
$
é
zg
3
z
5
It is now almost exactly a decade since EPIC was first announced by Intel and
HP, in June of 1994. Which is to say, our assessment of RISC vs. EPIC no
longer needs to rely on a practical understanding of the shortcomings of RISC
vs. the theoretical advantages of VLIW architectures. With two generations of
Itanium already shipping, and a third version of the processor scheduled to
begin shipping this summer, we now enjoy a practical understanding of EPIC,
too. Three major points have become manifest.
First, although the point to EPIC was to break through the superscalar barrier
that threatened to throttle RISC designs at the 6-issue mark, EPIC, unlike
RISC, did not promise any sort of sudden performance discontinuity, but rather
was a strategy for smoothly extending the curve of rising IPC (past 6).
Second, the superscalar barrier is not the problem to reaching higher levels of
IPC. For most programs, the IPC curve breaks before the superscalar barrier is
reached, due to a paucity of parallel instructions. In fact, there isn't enough ILP
in most code to fully sustain a 4-issue machine, let alone a 6-issue machine.
Third, EPIC processors do not succeed at reversing the trend toward more and
more complex hardware. If anything, Itanium processors are more complicated
to design and implement than other types of processors, not less so.= | - @Sun |
Why EPIC Isn't the Next Shift
cisc RISC cise ePIC
i386 R2000_ Xeon DP __itanium 2
Year 1985 1985 2002 2002
Technology | [Link] | 20ucmos | otgucmos | o.1eucmos
Die Size 103 mm2 | _ 80 mm2 145 mm2 421 mma
Transistors | 275,000 | 115,000 | 55,000,000 | 221,000,000
Cache caKpext. | s2kBon. | 532KBe 3.286 MBs
Package | tszeraa | wacrca | ronae |... OOGRe
Power 5s aw 130W
1 GHz
}1040 SPECint2K| 810 SPECIN2K
10.1 Spec89 T1048 SPECtp2K |1431 SPECIp2K
Tl comparable Wore (old)
To clearly see why Itanium isn't going to revolutionize the processor market the way
RISC processors changed the rules for processor design, let's make exactly the same
comparison between Itanium 2 and the reigning high-end x86 design of 2002, as we
did between the R2000 and the reigning high-end x86 design of 1985. As this table
shows, not only does Itanium 2 fail to provide an order of magnitude improvement in
price/performance over Xeon, you're hard pressed to find anything that Itanium does.
better than Xeon. The factors that affect the price of systems — die size, package, and
power consumption — are all substantially worse for Itanium than Xeon. The Itanium
design is not simpler than Xeon, but far more complicated. As a consequence, Xeon
outclocks Itanium by a factor of 3. For this reason, even outfitted with a cache only
1/6 the size of anium 2, it outperforms Itanium 2 on integer code by 25%.
To be fair, Itanium 2 is built with an older, cheaper process technology than Xeon,
albeit that's largely a reflection of the longer design cycles required for Itanium
processors, and is more than offset by Itanium's much larger transistor count and
associated giant die size. The one area where Itanium does excel is in its floating-
point performance. But even here, the appeal of Itanium is not to the thrifty, those
looking to pay less and get more, but rather to those in search of the last measure of
floating-point performance, and desperate enough to purchase it at any price.
In short, Itanium entirely lacks the distinctive signature of an important development
in processor history: namely, the ability to cause a sudden and significant forward
movement along the path of progress (towards higher functionality/lower cost).The Commodity Shift
The Microprocessor Era
The second new design point
1971. microprocessors * Pentium Pro reaches RISC performance | |
1985 cave (wth RISC micoareiecture),
AMO & ne iat sete 10 year egal
isc ‘ope overs done competion
1995
‘commodity Volume &
Processors Compabity
Price Perormante—— SSmpetion
Power consumption OMPeUt
Leverage dominant code base with specialized hardware
designs to achieve a key advantage ina volume market
1 Lower cost (AMD, Cyrix/NSC)
Lower power (Transmeta) |
| Higher specialized performance (Nexgen, MicroUnity)
L |
Although, in the end, Itanium turned out not to be the bright morning of a new era of
processor design, but rather a false dawn, there was a significant shift in the
processor landscape that happened in the mid-90s. Ironically, this involved Intel's
other processor, the one they intended to replace with Itanium,
As mentioned, by the mid-1990s, 64-bit, superscalar, out-of-order, etc. RISC designs
required huge transistor budgets (>3M) to implement. At this level of complexit
any meaningful distinction between RISC and CISC disappeared, x86 designers
fought back by the simple expedient of adopting RISC architectures wholesale,
creating RISC execution engines, and building front ends that converted complex x86
instructions into 2, 3, ot more RISC operations. True, this meant adding an extra
layer of complexity on top of the RISC engine in an x86 design, but one more
complexity, among so many others, was barely noticeable
Tve picked 1995 as the point when the balance of power shifted away from RISC and
over to the x86 architecture. 1995 was the year Pentium Pro briefly took the
performance leadership crown, demonstrating the viability of the x86 architecture,
and also the year that Intel finally terminated its legal efforts to prevent other
companies from building 32-bit x86 processors. The commodity shift was more
subtle than the RISC revolution, both because the x86 was old not new, and because
the shift was based not on a radical leap forward in computing performance, but
rather a reduction in the cost of computing based on volume and clone competition.Fr
i. ]
Sun
Intel's Unintended Triumph |
Intel Launches 1A-64, it's 3" attempt
to replace the x86 with something less
accidental, more controllable
“ EPIC” era falls to dawn because, at a
hhigher price point, the same (or less)
performance always loses
Intel dooms ttanium by demonstrating that
the x86 can reach RISC performance curve
“Commodity” era dawns because, at
the same (or higher) performance point,
lower cost always wins
RISC comes to dominate 32/64bit
| 1995 ‘embedded markets, where volume
ae without compatibility is possible
The last decade has seen a proliferation of companies convinced they can carve out
some piece of a huge commodity market, by building an x86 clone processor that
offers a special advantage: if not in terms of still lower cost, then in terms of a radical
reduction in power consumption, or a performance edge in some special computing,
niche. I think everybody has gotten comfortable with the idea that the x86 is really a
commodity processor, capable of being produced by many different companies, with
Intel holding a favored place only by virtue of their size and the associated economies
of scale they enjoy, their ability to address multiple design points, and their
aggressive pursuit of very high frequencies.
Itis worth noting, however, that this position, namely, the premier x86 producer
among many other actual and potential producers, is precisely the position Intel never
intended to occupy. When Intel gave up its legal battle to outlaw clone competition,
it did so in the expectation that it soon could replace the x86 with an exclusive
proprietary design that would enable them to avoid running in the commodity race,
rather than constantly keep striving to stay ahead in it.
That Intel's intended end run around the commodity market failed should come as no
surprise. As we have seen, Itanium provides absolutely no price/performance
advantage, but rather operates at a considerable disadvantage in this respect. Whereas
the mid-1990s legal transformation of the x86 from an Intel proprietary design into a
commodity processor, combined with Pentium Pro's existence proof that x86
processors could offer very competitive performance, had exactly the opposite effect,
endowing the x86 with a clear price/performance advantage against the field.aka “Horizontal Model”
(based on 18M PC: 8/12/3983 to present) component based
distribution, support (ors iteaation&
‘istomization bel Gate, PQ, igh specialization)
ee 1m Two privileged
iscrete hardware technology level tech companies
1 Worzonal ech
Sime nate
money
(unt Subsumed)
1 Driven by volume,
market orees
1 Delivers generic
functionality at
lowest possible
cost
value shift
Here's a visualization of the commodity or “horizontal” model that has driven
the x86 processor to dominance. There are two principal levels to this model, a
technology level, composed of independent hardware and software suppliers,
and an assembly level that handles building, selling, and servicing systems. The
system assemblers — the Dells and Gateways and HPQs, etc. of the world — are
free to select whatever hardware and software technologies they like from
among the competing suppliers of technology ~ albeit, Intel holds a favored
Position among the hardware suppliers as does Microsoft among the software
suppliers. Still, neither Intel nor Microsoft is quite as dominant as it might like
to be or, indeed, once was, thanks to AMD and Linux, respectively.
This model is chiefly remarkable for its ability to drive out cost, though the
price paid for the unerting ability of this model to home in on progressively
lower cost points is a kind of homogenization at the system level, that tends to
make assembled products indistinguishable from each other.
There also is a certain built-in instability to the model, as Moore's Law in the
processor realm and some analog of that principle of ever-increasing size in the
software world, enable Intel and Microsoft to assume more and more of the
necessary system-level technology into the CPU and OS, respectively. Thi
winnowing of technology suppliers creates tension, since a monopolistic end-
point for either hardware or software technology is incompatible with the
market-driven, competition-based nature of the commodity model.Intel's Vision for Hardware Technology |
(riven by Moore's Law of everexpanding transistor budgets/CPU functionality)
All devices that use CPUs |
(price competition between many assemblers to drive maximum volumes)
Intel's Dilemma:
Can't get here without Itanium
(0 proprietary lockout)
‘can't get here with Itanium
(no commodity model)
sole technology supplier
(after the last value shift)
But that's not to say a company can't dream about becoming a monopolistic
supplier. Certainly Intel does. This is their vision for the end point of the
hardware technology curve, after Moore's Law has enabled them to
incorporate all the (technically significant) hardware functionality needed to
build a system into their CPU chip. Note that the “horizontal” part of the
model now applies only its top assembly level, where Intel certainly hopes
to see many different assemblers, locked in a fierce price competition that
drives the maximum volumes for their CPUs. But the technology layer has
now taken on a decidedly ... vertical appearance.
The problem Intel faces in reaching this ultimate endpoint is that it’s flatly
inconsistent with the commodity model. Even after the system-on-a-chip
becomes standard for the most complex systems (as, by Moore's Law, it
eventually must), as long as the x86 reigns, market forces ensure that there
will be alternate suppliers of this commodity part.
Which is to say, in order to reach the goal envisaged here, Itanium not only
must become successful at the high-end of the processor market, it
ultimately must replace the x86 in the volume end of the market. But, as we
have seen, Itanium can't replace the x86 in volume use exactly because it
doesn't fit the commodity model, and hence can't reach the sort of
price/performance point that can be achieved only through that model.With Billions & Billions of
| transistors per chip ... .. all system functionality
(except large DRAM) gets
integrated into the CPU.
Intel's challenges aren't confined to the dilemma that the same commodity
model that has elevated them to their current lofty position as a World
Power simultaneously prevents them from ruling unchallenged in their
domain. Moore's Law is an independent source of worry. On the one hand,
it enables Intel continually to do more for less, providing an endless wealth
of new opportunities over time. On the other hand, it places them on an
endless treadmill, where doubling the world’s appetite for transistors every
two years amounts to nothing more (Moore?) than running in place.
For purposes of this discussion, the importance of Moore's Law is that,
quite apart from its specific implications for Intel (whether for good or ill),
it continually redefines the boundary between the CPU and the rest of the
system. In effect, Moore's Law means that the capacity of the CPU keeps
expanding, encompassing more and more of the system until, in the end,
only the CPU remains (along with some memory chips).
As the pendulum of Moore's Law swings further and further in the direction
of the system on a chip, however, it also creates more and more tension
with the component-based commodity model. The commodity model is
built on multiple suppliers, characterized by low integration and high
specialization. But a single chip cannot be parceled out across multiple
vendors. It is, by definition, an integrated whole, rather than an assemblage
of independent parts collected from interchangeable suppliers.The Real Endpoint of Moore's Law
With Billions & Billions of
transistors per chip ... ««. CPU design becomes system
design and system design
becomes CPU design!
system companies
(with CPU design expertise)
‘There wil be no end othe troubles of computing or of computer users themselves, til system
designers become cip designers inthis worl, oil those we now cll hip designers cally and
‘truly become system designers, and computing power andvison thus come int te same hands”
{with apologies to Plato's Republi)
In short, Moore's Law continually redefines the design landscape. It's real
endpoint is not Intel's vision of a single, monopolistic supplier of all
hardware technology, but a world in which CPU design has become system
design, and system design, conversely, is CPU design
I think it's fair to ask if this endpoint is a world friendly to chip companies,
all of which were created under the assumptions of a horizontal model, and
whose business is premised on the notion of supplying specialized
components to volume markets? Or is it rather a world friendly to
companies with the depth of expertise needed to create systems, and the
breadth of experience required to integrate many different functional
elements into a working whole?
In truth, it seems to me a world that favors a third kind of company: one
that combines the expertise of a systems company with the specialized
knowledge needed to design chips. Fortunately, there are a few companies
that fit that description, descendants of the system-side RISC era with the
strength of conviction and nimbleness of product to survive through the
semiconductor-side commodity era. Sometimes winning is just a matter of
surviving long enough to show up in the right place at the right time,Why SPARC!
|
So why does Sun design its own CPUs?
(not because it wants to be in the CPU business but rather)
Because Sun is a systems company
And, in the end*,
system design is CPU design
The altemative is to sell and service systems
designed by some other System/CPU company
(Don't confuse semiconductor fabrication
technology with CPU design technology)
If the foregoing analysis is right, the wheel of change is now poised to turn again,
moving us another notch further along the path of progress. Specifically, a design
point that lets us put not just hundreds of millions, but shortly billions and billions
of transistors on a single chip, provides system companies with the opportunity to
take advantage of another new design point that, like RISC, promises to combine a
radical advance in performance with a radical reduction in cost, At Sun we've
taken to calling this new design point Chip MultiThreading or CMT for short.
Unlike RISC, however, CMT is not strictly a compute-engine vision, but rather fits
into an overall system-level vision that we call “throughput computing.” And
that's really what I want to talk about today, though it's taken us awhile to get here.
But let me pause before turning to talk specifically about CMT and throughput
computing to answer the question most often asked about Sun's persistence in
building computer systems around its own SPARC processor, namely, why do you
still do it? RISC hasn't been truly relevant since, well, sometime around 1995,
Did Sun somehow miss the whole commodity shift?
The answer to that question, of course, is no, we didn’t miss the commodity shift.
In fact, today Sun builds systems around the commodity x86 processor and the
commodity Linux operating system. But for a systems company (as opposed to a
systems assembler) that can't be all there is. And for a company to stay in the
systems business, it ultimately has to stay in the CPU business because, in the end,
the two are one and the same thing.m systembased
= No privileged
companies
Cluster Tools Sa ae
Brey acs
Sen)
Bestof Breed Solaris OE
Highly Multithreaded SPARC CPU functionality for most
efficient execution of
target task(s)
“The Network is the Computer”
Open Systems |
|
At the system level, here's a snapshot of Sun's overall vision. Admittedly, this is
not a generic vision of some commodity system. It is a vision specific to Sun, and
incorporates Sun's ambitions to drive technology in given directions for the ultimate
benefit of its SPARC/Solaris customers. To the extent the goals underlying this
vision are successfully realized, both Sun and its customers will benefit.
Although I'm going to be focusing on just one element of the big picture shown
here, the highly multithreaded SPARC CPU that provides the hardware foundation
for this future Sun system, it’s important to understand from the outset that the CPU
strategy I'll be talking about is not a stand-alone proposition. It makes sense only
within the context of the Solaris OE and the robust threading model it provides, and
the inherently multithreaded network applications that Sun systems are designed to
run, In short, CMT is not a “one size fits all” kind of technology
Nonetheless, while the above picture is intended to deliver precisely the
coordinated functionality required for the most efficient execution of the tasks
targeted by Sun systems, absolutely nothing prevents other system companies from
designing CMT-style processors that fit into their own system vision, and that
provide the coordinated functionality needed to solve other types of system-level
problems. There are no privileged system-level visions, just the different visions of
different companies. And may the best vision, backed by the most effective
implementation, win.‘The Microprocessor Era
The third new design pe
1971 s*ICbased CPUs
Processor
1985, 1" commercial RISC CPUS
1995 2" RiSCcompettive 6
Commodity
64-bit Processors 2005 1" cmt cpus
Chip Efficiency
Price/Petrmance—~) Muaithreading
112th Lee race Power consumption|
ar enip atreadng
Throughput (LP)
Strategy: Reinvent the (stalled) IPC curve by switching from ILP
to TIP, focusing on tasks per time (not time per task)
‘Advantages: Much more efficent deployment of transistor resources
‘Much higher utilization of on-chip resources
‘Much less idle time
‘True practical IPC scaleability (to large numbers)
So here's what I think is going to be the next big wave of innovation in
CPU design, CMT. 2005 is the year that Sun will roll out its first radical
CMT design, marking the beginning of the CMT shift. Note that these
shifts, from RISC to commodity to CMT, seem to come on 10-year
boundaries, With just three data points that well may be only a
coincidence, but I'm planning to keep a weather eye out as we get nearer to
2015, in case there's some deeper cause behind the timing of this pattern.
Ina nutshell, the basic idea behind CMT is to revitalize the drive to higher
levels of performance by creating designs that use both on-chip resources
and processor clock cycles far more efficiently than any designs yet
produced. The key here is to invert the performance focus, to switch from
looking for ways to minimize the time needed to execute a given task, and
look instead for ways to maximize the number of tasks that can be executed
ina given time.
The advantage of this inversion of the performance problem is that it
enables the creation of processors that make much more efficient use of
available transistor budgets, i.e., that can use far more of their available
resources at the same time, and that suffer many fewer idle clock cycles.
Further, it lets designs be created that allow almost unlimited IPC scaling,
as Moore's Law enables the creation of bigger and bigger machines. The
result will be systems that offer unprecedented levels of throughput.The Performance Trap
Relatively small gains in CPU performance result
|
|
|
in relatively large losses in chip efficiency
(more on-chip resources sitting idle more of the time) |
Idle Resources
‘The more aggressive the CPU in attempting to drive up IPC,
‘the worse the problem
Idle Time
‘The higher the CPU clock rate,
the worse the problem
Perhaps the simplest way to understand this next quite radical departure in
processor design is to start from what I call the performance trap. If you think
back to the performance vector slide I showed earlier, recall that there are
only two basic parameters that control how quickly a processor can execute
instructions, namely, the average number of instructions it can execute in each
clock cycle, known as IPC or Instructions Per Clock, and how fast the
processor clock ticks, i.e., the processor's operating frequency.
The performance trap is simply that everything you do to drive up
performance has the unfortunate side-effect of making the processor less
efficient. Further, as processors attempt to reach ever higher levels of
performance, the problem rapidly gets worse — each new performance gain
tends to be smaller, yet must be purchased at a still larger loss of efficiency.
More specifically, everything a designer does to push up the IPC figure of
merit results in more on-chip resources sitting idle at any given time. And
everything a designer does to push up the CPU clock rate results in more
clock cycles in which the processor does nothing at all.
Why do we care how efficient or inefficient the processor is? For two
reasons. First, inefficiency is not merely wasteful, it's expensive. Second, if
we could somehow begin using all our on-chip resources and all our CPU
clock cycles efficiently, the potential performance improvement is staggering.Osun
Tesnurce cage ares oval due w ache mines
Frebssed on amount of vate n opi coe
‘Gap conceptual numbers co nt reser mesure les
Wide superscalar designs achieve
higher performance by using more
resources less efficiently
Today's widest Gissue
designs are operating at
a relatively low efficiency
aise ten 7a
Instructions Per Clock
(performance)
Let's start with the issue of idle resources. If we graph IPC against resource
usage, the resulting efficiency curve is going to look something like this.
Ignoring the problem of stalls for the moment, a processor that is resourced to
issue just a single instruction in a clock cycle is operating at perfect efficiency,
in the sense that it is always doing everything it is designed to do.
‘As you resource a processor to issue additional instructions in the same clock
cycle, efficiency necessarily drops off, because you have to design on the “best
case” assumption, that the additional slot will be filled every cycle, but in fact,
the “best case” will not always be realized. Thus, if you design a machine
with two issue slots, on some cycles the second slot will not be filled, and the
associated resources hence will be left idle. Add a third slot, and the
associated execution resources will be left idle even more often. And so on:
each new issue slot carries its own burden of added resources to support the
additional issue, but the extra resources will pay off less and less often for each
added slot. (Assuming what seems only reasonable, that the chance of issuing
nn instructions is always higher than the chance of issuing n+1 instructions.)
Today's widest designs try to issue 6 instructions every single clock cycle.
However, in fact, in many programs, they do well to average 2 instructions per
clock. Which is to say, these designs operate at very low levels of efficiency.
This “efficiency crises” for high-performance processors is now serious
enough that steps already are being taken to remedy it.HyperThreading
Js an attempt to solve the inherent
inefficiency of wide-issue machines
“comeing hema
Se re,
eam eaeenennae
Site
Make mor fet use of valable
crt resoores by runing two
treads nparal haf the Ihe
pertirend
a8
‘rade some pershvead
performance ptentat
gains abt to un
‘ore threads in parallel,
(patil efficiency)
Resource Usage %
45 6 7 8 9 0 1 WB
Instructions Per Clock
(performance)
“ypertrendng stl ame fx Syme MuitTivea ig or SMT
The term you may have seen in this context is “hyperthreading,” which is
Intel's name for an idea that comes out of some academic research done at
the University of Washington and elsewhere under the term “Symmetric
MultiThreading” or SMT. The basic idea behind SMT is fairly
straightforward. Build 6-issue machines (for example) that have the ability
to divide themselves into two virtual 3-wide machines. Then, in code where
6 instruction issue is a very distant hope indeed, put your idle resources to
work running a second thread in parallel with the first thread, on the built-in
second virtual core.
Like everything else in life, there are some trade-offs to this strategy.
Obviously, you sacrifice some portion of your potential single-thread
performance, though the sacrifice is smali on the operative assumption that
opportunities to issue more than 3 instructions at a time are rare. (If they
weren't rare, the 6-wide machine wouldn't be so inefficient.) And while you
do retreat back up the efficiency curve to a more productive use of the
available on-chip resources, a hyperthreaded machine is still far from
perfectly efficient, since three-wide (and even two-wide) issue will fail some
of the time.
So, yes, hyperthreading can be an improvement. But it’s far from being a
complete cure to the inefficiency plaguing today's wide-issue designs.Maximizing E Efficiency
If efficiency is the goal, there's a far
Ne solution than Hyperthreading
Build a single-issue core
It's the only perfect solution!
(ot course, you have to trade tvead performance)
{spatial efficiency)
SGP UOoOo OS
Resource Usage %
45 6 7
j Instructions Per Clock
(performance)
If I may be allowed an analogy, hyperthreading as a solution to wide-issue
inefficiency is a little like solving the pain associated with hitting oneself over
the head by putting on a helmet. Yes, it does succeed in reducing the felt
pain. But the fact remains that you're only in pain because you keep hitting
yourself over the head, So there's a much simpler and more direct way to
eliminate this pain than putting on a helmet — stop hitting yourself.
Similarly, the only reason hyperthreading makes any sense at all is because
you've gone and build a tremendously inefficient wide-issue machine, and
now you're trying to figure out some way to make it less inefficient. But
there's a simpler solution. Don't build a wide-issue machine. Problem solved.
In fact, if maximum efficiency (rather than maximum single-thread
performance) is the goal, the very best solution is to build a simple, 1-issue
machine. Because, as soon as you add a second issue slot, efficiency is going
to drop off dramatically. Indeed, where efficiency is the question, a single-
issue machine is the only absolutely perfect answer.
Of course, there's a trade-off. If you want to maximize efficiency, you have to
surrender some performance. Conversely, if you want to maximize
performance, you have to surrender efficiency. The good news about trading
in the direction of efficiency rather than performance, is that this trade
typically sacrifices mere pounds of performance, compared to the tons of
efficiency that tend to get thrown away when more performance is the goal.Le a osun
Space-Efficient Performance
TM
How to avoid the” | 8 xissue cores
idle resources
performance trap
Drive up performance while ‘Two 84PC designs
‘maintalingeficlency by
SPplyng transistors not Both require about
needed for wide sme the same size chip
to building motple
‘issue corer
‘once
Resource Usage %
(patil efficiency)
ges
endleve perfomance 1 Bissue core
forfarigher throughput
4 5 6 7 8 9
Instructions Per Clock
| (performance)
So here's the secret to building a processor that is able to use all its resources
efficiently, and that also takes advantage of the very large transistor budgets
that are currently available to CPU designers. Namely, rather than chase
Instruction Level Parallelism down the curve of vanishing efficiency by
building a massively-resourced, 8-issue, single-core processor, switch playing
fields. Instead of going after ILP, go after Thread Level Parallelism by using
the available transistors to build 8 single-issue cores on a die.
Both processors can work on 8 separate instructions every single clock cycle.
Both processors require about the same size (cost) die to build. The difference
is, every single core on the 8-issue TLP machine is operating at maximum
efficiency, supporting only the minimum set of resources that it both needs and
can use all the time. Whereas, the 8-issue ILP machine is operating far below
optimal efficiency, with most of its resources sitting idle most of the time.
‘True, any given thread will finish executing sooner on the 8-issue ILP
machine. But the 8-issue TLP machine will finish executing far more threads
in any given unit of time. Which is to say, if we change our definition of
performance from minimizing the amount of time necessary to execute a given
unit of threads, to maximizing the number of threads that can be executed in a
given unit of time, the TLP design is far higher performance than the ILP
machine (corresponding to its far more efficient use of resources).Efficiency,
Higher
Single Thread Single Thread
Performance Performance
active J idle
logic Bi logic
Here's a graphic that shows the difference between the alternative design points
that we've been discussing. On the left, we have a simple scalar design. As
mentioned, this design provides the most efficiency but the least performance.
To boost performance (defined as the ability to execute the same thread in less
time), designs moved from scalar to superscalar in the early 1990s, settling on 4-
issue machines as a sort of “sweet spot.” Although these machines “peak” at 4
IPC, their average IPC is certain to be lower ~ 2 IPC is probably on the generous
side. Which means half the peak issue width amounts to “dark logic” overhead.
Let’s say we now push on to an 8-issue design, in pursuit of still higher
performance. This does bring more active logic to bear on the computational
problem (say the average IPC now goes up to 3). But the toll in efficiency is high,
ice., the ratio of idle “dark” to active “light” logic gets even worse. Indeed, the
inefficiency of designs this wide is so obvious that it calls out for remedy. Enter
hyperthreading, which backs away from maximum single-thread performance in
order to run two threads in parallel, restoring the overall efficiency of a 4-issue
design (along with its somewhat lower level of single-thread performance).
But there is a far more efficient solution than hyperthreading: an 8-issue machine
designed not to run | thread very quickly, or even two threads somewhat less
quickly, but 8 threads at the standard scalar rate of execution. If efficiency is a
goal, and performance is defined to mean maximum throughput, this is by far the
best alternative among these possible design points._— @5un ]
More Idle Time
(ei sage asus no sais oe oak fremuces or sauce cots
‘ph eoncepta, mbes 9 ra rein messed lus
Higher frequency designs achieve
higher performance by using more
clock cycles less efficiently
gases
designs are operating at
relatively low efficiency
/
Today's highest frequency
cycle Usage %
&
(temporal efficiency)
°
‘© 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800,
Clock Rate
(performance)
The curse of idle resources, though, is only half the performance trap. The
other half of the trap is the curse of idle clock cycles.
Here's the shape of the problem. Recall that the other half of the MIPS
performance equation (besides IPC) is clock rate or processor MHz.
(actually GHz these days for high-performance machines). But, sure
enough, as clock rate goes up, processor efficiency ~ measured here as the
percentage of total clock cycles during which the processor is
accomplishing useful work (vs. idling along, accomplishing nothing of any
use) — goes down.
Don't stop me if you've seen this curve before. It's exactly the same
progressive decline in useful clock cycles, as the overall processor clock
rate is driven up, that we saw for productively employed resources, as the
overall processor IPC is driven up. And again, we've now reached a crises
point. Today's high frequency designs are throwing more and more clock
cycles at problems in the interest of ever higher performance, but with less
and less effect. Indeed, compared to, say, a 1 GHz design, a 3 GHz design
is shockingly inefficient in its use of available clock cycles,
We've seen there is way around the trap of ever more idle resources. Is
there also some way to avoid the trap of more and more idle cycles?—— Sun
| The Basic Problem
Typical High-frequency RISC Pipeline (Integer Ops)
Miss
Every Li Miss stalls the pipeline while the required instruction
or data item is fetched
Before we turn to consider possible solutions, let's take a minute to
understand the root cause of the problem here. This is a picture of a typical
high-frequency pipeline. Note that although it takes only a single cycle
(stage #8) to execute a typical integer operation (in keeping with the RISC
design precept of “single cycle execution”), there's quite a bit of processing
that has to go on before an operation can be executed, as well as some
processing that has to go on afterwards. But as long as the pipeline stays
full, i.e., every stage is engaged working on an instruction — the machine
operates at peak efficiency, with no wasted clock cycles.
All modern pipelined processors are designed to keep up a steady
execution flow, as long as both instructions and data are no further away
than the on-chip Level One (L1) caches. Indeed, the need to access the L1
caches in step with the pace at which instructions flow through the
pipeline, is a primary reason these caches have to remain relatively small.
The problem is, the L1 caches are small. For most real programs, they
hold only a fraction of the total instructions and data required to
completely execute a given thread. And every time there's a “miss” in
cither the L1 instruction or data cache ~ the next instruction or data item
that's needed for the computation is not in the L1 cache — the pipeline has
to stall while the missing item is fetched from some more distant location.Effect of increasing CPU Frequencies on Memory Latencies)
[fe be all memories st othe te >
The higher the clock frequency, the longer the stall on
every Li miss. nesmrey nmbere cc resp wating
The length of the stall depends on how far it's necessary to go to get the
missing item. If the item is on the processor chip in an L2 or possibly L3
cache, typically the stall lasts only a few clock cycles. If it's necessary to g0
off chip to a local high-speed cache, the stall might be a few 10s of cycles.
If the item is not in cache but has to be fetched all the way from main
memory, then the stall is likely to take 100s of CPU clock cycles. And
should the item have to be retrieved from disk, well, measured by CPU
clock cycles, that's a lengthy trip indeed. 5 ms., the seek time of a fast hard
disk drive, may be far less than a heartbeat, but it is 5 million clock ticks for
a 1 GHz processor.
The basic problem with escalating CPU clock frequencies is that the
memory hierarchy (including any large on-chip caches, external cache,
memory, and disk) does not scale up in speed at the same pace as the CPU
pipeline itself. Which is to say, every time the CPU pipeline clock gets
faster, in effect the entire memory hierarchy shifts to the right on this slide.
One of the consequences of making the CPU clock tick faster is the
unfortunate fact that more ticks will occur during an interval in which the
processor pipeline is stalled, waiting on a memory access.= ncross the industry, The gap only gets worse with time
today’s chips are largely
able to execute code faster
than we can feed them with
Instructions and data,
‘There no longer are
performance bottlenecks in
‘the floating point multipler
or. Integer unit... Ina
CPU Frequency
‘Memory Speeds
study using TPC... three
‘ut of every four CPU cycles
retired zero instructions;
Operational Frequency
Time
“1 expect that over the coming decade memory
subsystem design will be the only important design
issue for microprocessors.” - Richard L. sites (Alpha architect)
ets fon th ama Sp oper p99
This truth has been a dominant factor in processor performance for the last
decade, as these quotes from 1996 attest. The Alpha, the last important
high-end RISC processor to be introduced, was designed around a “speed-
demon” or clock-oriented performance strategy (rather than an ILP-based or
“brainiac” performance strategy), and so was one of the first designs to fully
appreciate the magnitude of the problem of idle clock cycles. As Richard
Sites here testifies, measurements show that in many programs as many as 3
out of every 4 clock cycles accomplish no useful work, but are spent
waiting for memory accesses to complete. Back when the Alpha was the
clear leader in clock frequency, it was sometimes said jokingly that Alpha
“waited faster” than any other processor. Which, in fact, was literally true,
if the measure of “fast waiting” was the number of CPU clock cycles that
expired uselessly during a memory access.
As the graph illustrates, this is not a problem that can be fixed by waiting
for things to get better sometime down the road. To the contrary, the
problem only gets worse with each passing year, as CPU clock frequencies
continue to rise at a much faster rate than memory speeds increase.| Computing Faster Is Little Help
decreasing compute times are of very modest value il emo
without corresponding improvements in memory latencies [7] soy trey
Ce aad
Here's an illustration of the “performance trap” encountered in trying to reach
higher performance levels by means of higher clock rates, Doubling the
processor clock rate cuts actual processing time in half. That's the good news.
‘The bad news is, unless memory is somehow made faster, all doubling the CPU
clock rate means for a memory stall is that twice as many ticks get wasted
waiting for the next instruction or data item to be returned.
To see the magnitude of the problem here, let's plug some sample numbers into
this illustration. Say we spend 100 clock ticks computing before we get an L1
cache miss, and then have to wait 300 clock ticks for the needed item to be
returned from memory. Obviously, the machine is not terribly efficient in its
use of clock cycles under these assumptions. In the small sample of computing
time shown, the machine spends 400 cycles doing actual computing and 900
cycles doing absolutely nothing but waiting while memory is accessed.
Now let's double the CPU clock rate. The good news is, the 400 cycles spent
computing now happen in half the time previously required. But all the higher
clock rate means for a memory access is that the machine now spends 1800
CPU cycles waiting on memory. True, the 1800 cycles don't take any longer
than the 900 cycles did formerly, and the 400 compute cycles happen in half the
time, so there is a modest overall speedup of about 15%. But there is a steep
price paid in efficiency for this 15% performance gain, since while there still
are only 400 compute cycles, there are now 1800 (not just 900) idle cycles.osx
Vertical Threading
|
| 3 oter4s is a way to solve the inherent
| inefficiency of high frequency
machines
Bea 2 22
Keep core busy at all times by
switching between tasks
It's the only perfect solution!
Cycle Usage %
(temporal efficiency)
st
0
3
© 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800
| Clock Rate
(oerformance)
_ |
Fortunately, there is a way to solve the inefficiency inherent in moving to
higher clock frequencies. Rather than follow the curve of single-thread
performance down in efficiency, increase the number of candidate threads a
core might operate on in a clock cycle. Then, when one thread stalls on a
memory access, rather than sitting though a period of enforced idleness,
waiting until the stalled thread can start up again, simply switch off to
another thread. When that thread stalls, too, switch again. And soon.
With proper planning and a little luck, few or no cycles will be wasted
waiting on memory accesses for stalled threads.= @sun
Time-Efficient Performance
|
| ‘ni wo tc esa
TLP 4 threads per virtual core
3
How to avoid Dieu performance whe
" mata fie
the idle time eideane stomata
the Gteaon tne of whcnes
performance thread is able to run (isn't stalled
trap ‘waiting for memory)
Stalled 75%, sti 25% fees on rable singl end
ILP 1 thread per virtual core
90
80
7
Cy
50
0
ee
se
ge
oe
#2
E
terteatng (SMT, arotal heading) dige sgle
tdssoe psa cre to fr more an ers"
Seseseidt tempter sale
° Pj ij ij jj ya
© 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800
Clock Rate
(performance)
Here's alittle more detailed picture of how to avoid the idle time
performance trap through vertical threading. Suppose each virtual (or
physical) core in a processor is able to maintain concurrent state for, say, 4
separate threads of execution. On the assumption that any given thread is
stalled up to 75% of the time waiting on memory accesses, three of the four
threads might be stalled at the same time. But one of the four threads ought
to be able to execute. When that thread stalls ~ as it will, probably sooner
than later — the memory access for one of the three stalled threads may have
completed, and the processor can resume that thread of execution. When it
stalls, another of the candidate threads might be ready to run, and so on,
While this scheme is not foolproof — it's possible that even with 4 candidate
threads available to a core, there will be clock cycles where all threads are
waiting and none are runable ~ it's obviously going to come far closer to
perfectly efficient use of a processor's clock cycles than a single-threaded
core. It will approach perfectly efficiently use of processor time, even if it
does not perfectly achieve this goal.
Note that hyperthreading (simultaneous multithreading) is distinct from
vertical threading. Addressing the inefficiency of wide-issue machines
(resource use) does nothing to solve the problem of each simultaneous
multithread idling for many clock cycles, waiting on memory.Time Usage Comparison
Single-Threaded Core
Th eee
(SRENRENENNEN Core compute time
corel tne
SS renoyancytme } Hm
4X Vertically-Threaded Core
Thread 2
Here's a more explicit comparison of the way in which a single-threaded core, and
a 4X vertically-threaded core, spend their respective clock cycles.
The single-threaded core follows the classic compute-stall, compute-stall model
shown at the top. In any given stretch of processing time, the larger part of the
total time expended is spent waiting on memory accesses, with core idle time
equal to the time taken by the memory accesses.
The 4X vertically-threaded core follows exactly the same compute-stall pattern
for each of its threads. No thread can run when missing a required instruction or
data item; every thread must wait while such items are fetched. But rather than
wait for a given thread to be able to resume execution, a vertically threaded core
moves on to a new thread of execution that is not currently stalled, and keeps
operating. When the new thread stalls, the core shifts again to a third runable
thread, and when that third thread stalls, it still continues operating on a fourth
thread. By the time the fourth thread stalls, with any luck at all, one of the first
three threads will have completed its memory access, and execution can continue
on that thread without missing so much as a single beat. And so on.
The bottom line is the core stays constantly busy, gainfully employed at executing
one thread or another all the time, with no idle time at all. True, there is a lot of
memory access time needed to keep this multi-threaded core fed, but all the cycles
required to access memory are overlapped with other access and compute cycles,
and so effectively “disappear” in terms of their impact on performance.Now let’s put it all together. Here's a picture of what happens when you
combine eight resource-efficient single-issue cores on a single processor
die, and enable each of these cores to run in a highly time-efficient way,
through 4X vertical threading. As this picture indicates, the prospect is a bit
overwhelming. You end up with a processor that is able to keep 32 threads
in play simultaneously.
True, a maximum of just 8 threads — one per core — will be executing
actively on any given clock cycle. But the other 24 threads are also all
“live”= its just that they happen to be stalled at the moment, waiting on a
memory access to complete so they can resume execution. So, in a very
real sense, this processor is running 32 threads all at the same time, each
thread just as fast as it is able to run,
The secret behind this processor's astonishing ability to run 32 threads at
once is its near-perfect efficiency. By virtue of its 8 single-issue cores, the
whole processor die is put to use, without the overhead burden of many
under-utilized execution resources that plague wide-issue designs. By
Virtue of its vertically-threaded design, each of its 8 cores operate with few
or no idle clock cycles spent waiting on memory accesses, without the
‘overhead burden of many idle clock cycles that plague both single-threaded
and hyperthreaded designs.“Hyperthreaded” Processor CMT Processor
use some on-chip resources
‘OHT Processors
‘recess
a a |
use all on-chip resources
= core compute time
ES core idle time
=== memory latency
some of the time
Admittedly, by using its “excess” fetch, issue, and execute capacity to run two
threads in parallel, a “hyperthreaded” processor does succeed in rectifying some of
the inefficiency of a very wide-issue design. But this retreat to two narrower virtual
cores still falls far short of perfect efficiency, in terms of its ability to keep all the
available on-chip resources gainfully employed at the same time. And
hyperthreading does absolutely nothing to address the fact that each parallel thread
will stall up to 75% of the time, while necessary memory accesses take place.
Compare this picture to a CMT processor, where virtually all of the available on-chip
resources are put to use by running 8 threads on eight single-issue cores (rather than
just 2 threads on 2 multi-issue cores). And since each single-issue core is vertically
threaded, all of their respective clock cycles can be spent computing (rather than
mostly waiting on memory). The bottom line is, rather than keeping some of its
resources busy some of the time, the CMT processor is able to apply all of its
resources to computing all of the time.
For designs of similar complexity, this difference in efficiency translates directly into
a difference in performance. If the CMT processor brings, on average, twice the on-
chip resources to bear on a compute problem, and four times as many clock cycles, it
will outperform the hyperthreaded design by a factor of 8-to-1. Compared to a
single-threaded design, the difference would an astonishing 16-to-1 — which is to say,
just about the factor of 15X improvement that we've cited for our first CMT blade
processor, compared to today's four-issue, single-thread blade processor.Why CMT Is a Radical Processing Advance
per thread
percore peak avg at same MHz
processor cores threads © IPC IPC Efficiency Throughput
“Niagara” 8 Pat 100%
POWERS, 40%
Itanium 3 33%
SPARCE4 VI 33%
Xeon 33%
Opteron 93 33%
Performance Comparison: dependent on
me wp
CMT architectures high none
‘ficiency = avg IPC / peak ie 7 rea
‘roghpats freeads mg PC ee en |
How radical an idea is CMT? Here's a comparative look at the processor
landscape circa 2005, based on what I've just told you about Sun's
forthcoming CMT “Niagara” processor, and what other vendors have said
about their processor plans for 2005. IBM and Intel both claim they will have
a dual-core CMP design, with hyperthreading capability for each core. Other
design points for 2005 are dual-core CMP designs without hyperthreading,
and single-core designs with and without hyperthreading. That's pretty much
the stated range of prospects.
Ive listed some fairly generous average IPC figures for each of the wide-issue
designs on this list. If we figure efficiency as the ratio between average and
peak IPC, no wide-issue design is very efficient. If we calculate throughput
performance by multiplying the number of cores per die, times the number of
threads per core, times the average IPC per thread, the dual core/dual thread
designs appear to top out at a throughput of about 8. On the same two metrics,
“Niagara” scores a flat 100% efficiency rating, and offers throughput of 32.
‘That's why CMT is a radical processing advance. Note the conditions under
which this comparison holds valid. The CMT design requires 32 parallel
threads to reach its throughput number, but depends not at all on ILP. The
other designs have a relatively modest dependence on TLP, since they can use
no more than 4 threads at a time, but must attempt to compensate for their low
TLP by applying more ILP to each thread.CMT - Keys to Success
Simple, Scalable (RISC) 64-bit Processor Core
| ~ Designs require a huge address space
| — SPARC is an ideal candidate
| Processor Design Expertise
| ~ Sun has been designing processors since 1984
“Throughput Computing” Design Expertise
~ Sun commercialized SMP technology in the 1990s
- CMT achieves at the CPU level what SMP does at the system level.
Highty Threaded Operating System
~ Solaris is the industry leader in effective thread support |
Highly Threaded Applications
— Intrinsic to most networking computing applications
— Not typical of most desktop applications
‘What's needed to succeed with a CMT-style processor design? Or, to put it another
way, if CMT designs are so great, why isn’t everybody building one? The answer is
that it takes a fairly unique combination of circumstances to be able to bring a
successful CMT design to market, starting with a simple, scalable 64-bit processor
core. SPARC is ideal here, while Itanium, for example, is a poor base (8 scalar
SPARC cores will fit on a die that uses fewer transistors than the next iteration of
the single-core Itanium 2).
Second, it takes processor design expertise. Some system companies, bowing to the
pressures of the commodity processor cycle, have abandoned CPU design on the
grounds that it is no longer possible for them to add value in this area. Sun, by
contrast, has continued to invest in its processor design capability.
But simply having both an appropriate CPU design base and CPU design capability
aren't enough. CMT designs translate to chip-scale what SMP accomplished at a
system level. To build these kinds of chips requires an understanding of MP design
issues at the system level. It takes a server company to build a “server on a chip.”
Even the right processor core, processor design expertise, and a system-level
understanding of SMP issues are not sufficient for CMT success. The whole point
to these designs is to switch the performance focus from ILP to TLP, from time per
unit task to tasks per unit time. To empower these designs, it takes an OS like
Solaris, with a robust multi-threading model, and applications with a high degree of
TLP, like those typical of network computing. All of which Sun uniquely has.Huge increase in
performance of Performance
threaded applications
Quantum reduction eae
in the cost of —Fewer servers
Network Computing -Less floor space
without ; Reduced power consumption
eee ~Lower air conditioning
Simplified administration
and maintenance
Major consolidation
of components and ;
interconnections Reli
‘So what benefits might one of Sun's customers expect to see from Sun's CMT
initiative? In a nutshell, benefits in proportion to the size of the advance in
computing technology.
‘An increase in single-processor performance anywhere from 4 to 8 times the
best performance available on any competing hyperthreaded ILP-style
processor in the same timeframe. A quantum-leap reduction in the cost of
computing, based on the need to have far fewer servers to handle the same
workload, housed in much smaller boxes, occupying much less valuable real
estate, consuming far less expensive electricity, needing much less
investment in cooling capacity, and offering substantial savings in the costs
of maintaining and administering systems.
As a bonus, the new systems should be far more reliable than any comparable
SMP system built today, since collapsing the entire server down onto a chip
also eliminates all the separate components and interconnections in today's
SMP systems, together with their additive chances of failure.
But from the customer's viewpoint, the best part about this whole prospect is
that it does not disrupt their existing compute environment at all. Sun's new
SPARC-based CMT processors will drop seamlessly into a new generation of
SPARC/Solaris systems that continue to run the same applications in exactly
the same way as before. Except far faster. In far less space. With far less
‘operating expense. More reliably than ever before.CMT - Sun Benefits
+ Provides competitive advantage and
differentiation
+ Increases revenue and market share
through superior value proposition
* Enables new markets and services
* Validates Sun's long-term
SPARC/Solaris strategy
What will this initiative do for Sun? Again, the benefits should be proportional to
the size of the advance being made in computing technology. To date, Sun is the
only company proclaiming CMT as its direction for future processors. Once other
companies have had a chance to think about the benefits of this strategy, that
differential is likely to go away, but we expect to maintain a leadership position
with CMT processors, just as SPARC was a leader throughout the RISC era.
As just explained, Sun's CMT systems should provide Sun customers with a
superior value proposition, enabling Sun to win new customers and increase both
its revenue and market share.
Further, we expect CMT (like RISC) to enable entirely new markets and services,
with a great expansion in the opportunities for “server-on-a-chip” technology.
Suppose you could put an entire E10K on your desktop. Suppose you could carry
one around in your pocket. Would that be interesting? What new things might
you be able to do in that sort of a world?
Last, but certainly not least, CMT validates the importance of continuing to invest
in key technologies, even during a commodity shift, when it might seem that there
is no way to add value through technology, and wise companies should focus
instead on achieving some form of non-technical advantage, whether that be
lower costs, better service, greater ease of use, or whatever. To the contrary, as
Jong as technology has limits, there will be value to devising new technologies
that allow the existing limits to be broken.Ue Osun =)
Innovation Pays |
* Throughput Computing fundamentally |
changes the price/performance ratio of |
Network Computing
+ Enables new markets and services
+ Sun is uniquely positioned to deliver
| Throughput Computing
(64-bit SPARC + Solaris + Threaded Apps)
The fact is that processor technology is not yet a story whose end can be
foretold, let alone announced. It still is the tale of a on-going journey, not
a story about the computer industry reaching some final destination already
visible on the horizon. Though the path of progress always tends toward
more functionality and lower costs, the specific parameters of processor
design are fluid and change from time to time. We already have witnessed
three very distinct design points just since the start of the microprocessor
era, from CISC to RISC to commodity, and a fourth shift to CMT is now
looming straight ahead.
Tn contrast to those who want to proclaim processor design a finished
subject, closed due to the exhaustion of possibilities, as long as there is no
obvious limit to the amount of progress that is possible towards higher
levels of functionality and lower levels of cost, there is no obvious end-
point to this process of change. Rather, quantum leaps forward will remain
possible as long as new ways of thinking about old problems are a
possibility. CMT is the next such leap. If this analysis is right, it will not
be the last. It's important to pay attention to CMT, because it will usher in a
new era of winners and losers, but don't expect the ending to be “... and
they lived happily ever after.” The longing for a stable end point is the
halimark of a fairy tale. The touchstone of reality is change.Well, that’s my story about the past, present, and future of processor design.
TI be happy to take any questions you might have, either about Sun's new
CMT initiative in processor design, or about any other points of interest
regarding the changing parameters of processor design over the years.
Obviously, to compress fifty years of design history into a 90 minute talk,
Ive had to simplify the big points and omit many interesting details, but I've
tried to do so without lapsing into either outright error or gross
exaggeration. Although perhaps the operative words here are “outright”
and “gross”, since some amount of distortion is probably inevitable. So I'd
be glad to try to clarify anything said I've said, or talk further about
anything that you might like to consider in a bit more depth.
Thanks for your interest and attention. The floor is now open for questions
and comments.