0% found this document useful (0 votes)
7 views50 pages

Microprocessor Design Changes

The document discusses the evolution of processor design over the past 50 years, highlighting key shifts such as the RISC revolution and the transition to Chip Multithreading (CMT). It emphasizes that change is constant in technology, driven by the need for innovation and efficiency, and outlines the historical patterns of experimentation and consolidation in the industry. The talk also addresses the importance of Moore's Law and advancements in processor architectures as critical factors influencing future developments in computing.

Uploaded by

harlan.mcghan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
0% found this document useful (0 votes)
7 views50 pages

Microprocessor Design Changes

The document discusses the evolution of processor design over the past 50 years, highlighting key shifts such as the RISC revolution and the transition to Chip Multithreading (CMT). It emphasizes that change is constant in technology, driven by the need for innovation and efficiency, and outlines the historical patterns of experimentation and consolidation in the industry. The talk also addresses the importance of Moore's Law and advancements in processor architectures as critical factors influencing future developments in computing.

Uploaded by

harlan.mcghan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF or read online on Scribd
BUCH em eee f Laceleces tea t+) As the title says, this talk is on the changes that have occurred and continue to occur in processor design. In order to cover what is now over 50 years of computing history in general, including 30+ years of microprocessor design in specific, I'll keep things at a a fairly high level. So, although there are a fair number of technical concepts involved that may not be familiar to everybody here — terms like “MIPS” and “RISC” and “IPC” and so on — we'll be able to explain these concepts as we 20 along. If you feel you're getting lost at any point, just stop me and I'll be happy to clarify not only the specific idea in question, but how it fits into the bigger picture of the Changing Parameters of Processor Design. Like all worthwhile stories, this one has a moral, in fact two morals. The first moral is simply that things do change. 1 know that seems pretty simplistic, but change is by nature unsettling, so there's always a lot of resistance to it — especially from those that benefit from the way things are. Hence, whatever the point you happen to be at, there always are some ready to declare that, finally, change is now officially over and the status is going to remain quo from now on. So the first basic point to this talk is that those kinds of statements aren't true — things not only always can change, but they do change and will continue to change on a fairly regular basis. The second basic point is that, although I think there's always an element of fashion in any kind of change, change in microprocessor design is not simply a matter of fashion. There's also a real sense in which progress has been and will continue to be made. Indeed, at bottom, what drives change in processor design is the periodic need to consider what is really quite a radical alternative to the way processors have been designed hitherto, in order to continue to make progress. | The Shifting Landscape ‘The Path of Progress | Technology Drivers The RISC Revolution The Irrelevance of Itanium The Commodity Shift The Coming Transition to CMT The Relevance of SPARC The Performance Trap & How to Escape It Space-ffficient Performance Time-Efficient Performance Building CMT Processors Requirements Benefits So here's the outline of what we'll be covering. First, we'll look at an overview of the history of computing, and set the major design points we need to consider. As you will see, viewed at a very high level, there is a sort of pattern to change that can help us understand the shifts that have occurred in the computing landscape. We'll then look at the way in which these periodic changes add up to overall progress, and the major technical drivers behind the progress that has been made. We'll then run through the last two major design shifts in some detail, the RISC revolution of the 1980s, and what I call the Commodity Shift of the 1990s. Since I know some people were and perhaps still are expecting Intel's initiative around their new Itanium processor to be the next big shift in the processor landscape, I also want to take a few minutes to talk about Itanium and why I believe it's destined to end up as a footnote in computing history, rather than as the next main chapter. That will set the stage for what I believe is going to be the next main chapter, namely Chip MultThreading, CMT, or what Sun calls “throughput computing.” Since processor progress is, at bottom, about processor performance, that will require that we spend some time talking about what now is impeding our ability to reach higher levels of performance, and how CMT overcomes those problems with designs that make efficient use of both a processor's physical resources and its compute cycles — what I call space- and time-efficient computing. Finally, we'll spend a little time looking at what's needed to build CMT-style processors, and the benefits that we expect to result from these sorts of designs. _— eu | The Shifting Landscape Just when you think it's all over - technology takes a new turn! System Design | companies Innovation Semiconductor companies [Hi diversification (experimentation) Economies Consolidation electing the winners) ___ OF Scale Here's the big picture that we'll be covering. For present purposes, I've conceptualized change in processor design as a periodic shift in the balance of power between system companies and semiconductor companies, with the respective forte of each side in this ongoing struggle for dominance being innovative new designs by system companies, pitted against the massive economies of scale that semiconductor companies can leverage. To the extent this picture is accurate, one thought here is that, in the end, both of these opposed forces really are necessary in order to keep moving, ahead. That is, without periodic infusions of design innovation, economies of scale ultimately turn flat and sterile. Whereas, innovative new designs eventually must find their way into volume markets, or become increasingly irrelevant. Within each phase shift, there also seems to be a consistent two-part pattern. First, there's a period of experimentation, characterized by a lot of different approaches and players. Second, there's a period of consolidation, which one or two dominant designs emerge. (That's the point where there's always a strong temptation, at least on the part of the winners, to declare the whole game is over, and from here on it's going to be just a whole lot more of essentially the same thing.) Just when you think it's all over - technology takes a new turn! ime by System Di companies Innovation yor Eeompebles Semiconductor companies Economies Diversification experimentation) Consolidation (selecting the winners) of Scale These are some sample names associated with the first era of computing When the commercial computing industry got started back in the early 1950s, of course, there were no semiconductor companies as yet (indeed semiconductor technology was still in its infancy, experimenting with the germanium-based transistor). So the early computers were strictly a systems play, with many companies trying their hand at some really very interesting design variations, before IBM emerged as the dominant player with their 360/370 line of machines. Although I'm not showing it here, mainly due to lack of space, there was a very important design shift from mainframes to minicomputers, during the initial technological period when CPUs were still build out of discrete components. This lead to the rise of a whole new group of companies in the 1960s, supplementing and in some cases supplanting the players of the 1950s: the Digitals and Primes and Data Generals of the computing landscape. Just when you think it's all over - technology takes a new turn! ooh we System Design orem —- GOMpanies Innovation yor fre Semiconductor es" _ companies a conomies Diversification (experimentation) consolidation (electing the wimers) __ Of Seale The development of the microprocessor in the 1970s had to wait on the emergence of the semiconductor industry in the 1960s, and an entirely new class of company, created to develop devices around the new possibilities inherent in the integrated circuit or IC. The economics of these companies, driven by the need to find volume markets for relatively inexpensive silicon-based chip components, were very different from the economics of even a minicomputer maker. The first processor designs to emerge from semiconductor companies, termed microprocessors in part to distinguish them from minicomputers, existed for many years under the radar of established system companies, with little apparent relevance to even 16-bit minicomputers, let alone to the far more sophisticated supercomputer and mainframe designs created by leading computer architects like Seymour Cray and Gene Amdahl. By the middle of the 1980s, however, two 32-bit descendants of these original 4- and 8-bit designs had come to dominate the major growth sectors of the computing landscape: the Intel 80386 in the home and business PC, and the Motorola 68000 in scientific and engineering workstations. Just when you think it's all over - technology takes a new turn! ve Po System 3, Design companies ®* Innovation ion oven ars Tooele, "Ss Semiconductor «<"" companies Diversification (experimentation) consolidation electing the winers) __ Of Scale Economies But just when it seemed as if the mantle of leadership in processor design was firmly settled on the shoulders of the semiconductor industry (albeit, some doubt remained about which semiconductor company eventually ‘would emerge on top), the landscape shifted again. The driver of this next major change was a new design point that emerged commercially in the mid-1980s, called “RISC.” This is the first shift I want to look at in some detail, so we won't discuss it here other than to point out that it enabled system companies to wrest the initiative in processor design away from the semiconductor companies. Indeed, one of the most striking aspects of the “RISC revolution” is that, while several semiconductor companies made serious efforts to introduce competing RISC designs of their own — including Intel with the i860, Motorola with the 88000, and AMD with the 29000 — not one of the semiconductor-sponsored design initiatives survived. Indeed, in the high- performance arena, only RISC designs that either were originally created by a systems company (like the HP PA-RISC, the Sun SPARC, the DEC Alpha, and the IBM POWER) or quickly acquired a systems “owner” (like the Fairchild/Intergraph Clipper and the MIPS/SGI MIPS) survived for any length of time. System companies era ‘isu Semiconductor «companies Diversification (experimentation) Economies of Scale ‘Through the RISC revolution, we have enough separation from past events to enjoy the advantage of historical perspective. The next shift, however, brings us to events that are still in the process of unfolding. So we move here from retrodiction to prediction, from simple historical review to forecasting events that still are in the process of playing themselves out. The advantage of a pattern like the one laid out here is that it can help us gain perspective on current events as they happen, without needing to wait for the judgment of history to decide on probable outcomes, or to distinguish likely winners from mere glamorous pretenders. But even armed with a pattern, there's no gainsaying the point that predicting the future remains a highly inexact science. Nonetheless, even though I can't tell you what will happen over the course of the next several years with anything like the same confidence I can speak about what has happened to date, what I can do is present one possible future scenario for your consideration, a scenario that I think not only has great potential to revitalize computing once again, but that also makes considerable sense in light of what has happened to date in the history of processor design, and which seems to me to be perfectly aligned with the essential forces needed to drive the industry into the future. Discrete Logie Integrated circuits (cs) A Protteof ve isruptve technology” (radical reduction in cost with signifcat Ita loss of neon} Progress driven by Moore's Law » New Design Poin Microprocessors The previous slides make the point that processor design does change on a periodic basis, with the initiative shifting back and forth between semiconductor companies and system companies. Further, just when you finally get comfortable with the way things are, suddenly everything changes. What the previous slides don't show is the fact that the shifting landscape of processor design is not a reflection of some underlying base instability, akin to periodic geological or political upheavals, but is instead a function of progress, where the path of progress is measured by steadily lower costs combined with an overall increase in functionality (as measured by a variety of indicators, from larger address sizes to lower power consumption). The one true “disruptive technology” in this chronicle of progress is the development of the IC, which led to the creation of the microprocessor itself. The early micros, of course, provided far less functionality than contemporary system processors, but with continued development not only caught up with minis and even mainframes, but eventually surpassed them — at a far lower cost for the same level of ability. Subsequent changes have not been “disruptive” in this strict sense, but rather have been new design points that lowered the effective cost of computing for the same level of functionality (or what is the same thing, provided more functionality for the same level of cost). Since we will be looking at each of the new design points starting with RISC in some detail, I won't dwell further on this summary, other than to emphasize that all real shifts in the processor landscape are marked by sudden and significant movements along this “path of progress.” prog ereantiog after 3 year ail ontrack tnamicroprocesor forthe 3B double evry 2 years “lz Since thay were tvented al | “The doubling will stow down. You really get bit | by the fact that materials are made of atoms.” “Gordon Moore, 7/5/2002 Hopefully, that's sufficient history to get us started. Before turning to the first of the new processor design points I want to consider in some detail, the RISC revolution, we need to look at two of the basic technical drivers of movement along the path of progress. ‘The first driver is Moore's Law, which has governed progress in the semiconductor industry since the invention of the planar transistor by Fairchild back in 1959 — the key development that made it possible to “print” ICs much like photographs. Moore's Law is a simple doubling algorithm that says how many transistors will fit on a semiconductor chip by any given year. In the initial 1965 version of his Law, Gordon Moore, who at that time was a Director of R&D at Fairchild, said the number of transistors per chip would double every year, starting from 1 in 1959. In 1975, Moore revised his Law downward to a doubling every 2 years. This latter curve has been maintained by the semiconductor industry for an astonishing period now, over 30 years, and looks to remain on track until sometime after 2010 (before the doubling period stretches out for a second time). The moral here for us is that we're headed to processors that will contain 2-4 billion transistors by the end of the decade. With transistor budgets now poised to go over the 1B milestone, entirely new kinds of processor designs are becoming feasible for the first time. That's the first key technology driver to keep in mind as we move forward. The more instructions a processor executes in a glen unit of time, the higher | its performance | MIPS or Millions of Instructions executed Per Second is the best known and | most basic performance metric A MIPS rating fora processors simple to calculate | | slven jst two numbers: | MIPS = avg. instrs. executed per clock (IPC) < clock rate (in MHz) {3,2 processor that on average, executes exactly 2 instructions every clock oe, and whose ioc runs ats Giz (= 2000 Mit), would execute 200,000,000 htrucons in second, for 8 performance rating of 2,000 (2.1000) MIPS 2,000 MIPS = 2.0 IPC X 000 MH ‘Note: MIPS ial rete measure of perermace fo roctsrs wh the sae nsructon et eter OA. [isnot aac measure as fret ance he yal etc ove a ny sotore ees ert on ‘he gpeatinsracion anchor 8h Tass eaimate complaint by OSC poorest oops PES as {hat the aber MPs ratings toutes RSC pacer wer ace, sces lon ene eset {ecompined ies actual wer than lon compen CSE nseucon The eeu eats shite Meare ‘iraatneprtcance, om te tine kates tw rent procs io eect ih sue tamer tarot {oie mer tatesewe cferem oceans to eacte he sine pay foun fh {hesitate nh Fe benchnak si, rrciadie whch ot eb MIS tgs wth SPecnumber othe nvr bse maar fave paceso’ pata The second key technology that continues to advance relentlessly, besides semiconductors, is processor architectures. We will be looking at several different types of processor architectures in some detail, but first we need to be clear on the basic goal behind most architectural advances, namely, their ability to enable higher levels of performance. The simplest way to think about processor performance is in terms of the number of instructions a processor can execute in a given unit of time, say a second. This is a function of just two basic parameters. The first parameter that controls performance is how many instructions a processor can execute, on average, in each clock cycle. This is typically abbreviated as “IPC” for average “Instructions Per Clock” executed. The second parameter that controls performance is how fast the clock ticks. Processor clock frequency is typically given in megahertz (MHz) or gigahertz (GHz), ice, as either millions or billions of clock cycles a second. Given an IPC number for a processor and its clock frequency in MHz, multiplying the two numbers together will produce a MIPS rating for the processor, specifying how many Millions of Instructions Per Second that processor can execute. Osun From the definition of performance (in terms of MIPS), it follows that there are just two possible ways to improve processor performance (MIPS rating) Increase IPC 0 | Increase clock rate Clock Rate | While it might seem that processor performance ought to be a lot more complicated that just MIPS, at the top level, that’s all there is to the subject. When processor architects think about how to design a faster processor, everything they think about boils down to one or the other of these two fundamental parameters. Either they need to figure out a way to execute more instructions, on average, in each clock cycle, or they need to figure out a way to make the clock tick faster. All the complexity of processor design lies in the details of achieving one or the other of these two goals. Moore's Law and the Performance Vector are really all we need by way of background. We're ready now to turn to the first basic shift in the way microprocessors were designed, the so-called “RISC revolution.” a es | The RISC Revolution | | The Mlcroprocessor Era The first new design point 1971. 1 microprocessor CPU 1985 2" com | abit 36D spit [RISC sin | | | aot * eo Ce | Radically simplify Instruction Set Architectures (ISAs) rect execution of instructions (no microcode) ‘Microarchitecture improvements (pipelining, etc.) Higher levels of onvchip integration Improvements in memory latency (caches) & bandwidth | less design complexity, fewer transistors | (smaller design teams/faster design cycles, smaller die) Functionality The revolutionary thing about RISC, which is an acronym for "Reduced Instruction Set Computing,” is that it fundamentally changed the price/performance point for microprocessors. As the RISC acronym suggests, the basic idea was to greatly simplify both the number and kinds of instructions that a processor needed to execute, focusing specifically on the sorts of relatively simple operations that effectively can be implemented in hardware (ike adding or comparing two binary numbers). This might seem like a pretty simple notion, but it stood in sharp contrast to the reigning design style of the early 1980s, dubbed “CISC” (by RISC proponents), for “Complex Instruction Set Computing.” CISC processors, like the Intel x86, the Motorola 68K, the DEC Vax, the IBM 360/370, etc., had 100s of instructions, many of which involved a lengthy sequence that started with fetching one or more values from memory, included operating on them in some way, and ended by returning the result to memory. The CISC idea was that the more powerful the instruction set, the more potent the processor. The RISC idea that higher performance would result, instead, from simpler instructions was really very radical. Building simpler processors had a lot of other benefits, too. Simplifying the core compute engine made transistors available for on-chip integration of other key functional units, like MMUs and FPUs and caches, which also carried significant performance benefits. And, since the designs were less complex, they could be created more quickly using fewer engineers. | Why RI RISC Was Revolutionary _ cise RISC i386 R2000 Year 1985 1985 Technology | 1.5. cmos | 2.0 CMOS Die Size 103 mm2_| 80 mm2 Price Transistors 115,000 Cache . | 32 KB ex. sszcrca | 144¢PGA | An order of magnitude Power aa ra improvement in p/p Performance | 2.2 Speceo | 10.1 Specso | Performance | | BH Better (hot) i comparable Wl Worse (cold) | Here's a telling comparison between two processors, both introduced in 1985, the Intel 80386, and one of the first commercial RISC microprocessors, the MIPS R2000. The RISC design required just over 100,000 transistors to implement, vs. nearly 300,000 transistors needed to build the 386. Hence, it was possible to build the R2000 in an older, more cost-effective semiconductor technology, and still end up with a smaller, cheaper die. Further, the R2000 was able to use half the external cache of the 386 and still generate almost 5 times the performance when operating at exactly the same frequency, in a very similar package, using no more power. Multiply the cost savings times the performance gain, and RISC offered an order of ‘magnitude improvement (or more) in the price/performance ratio of microprocessors. No wonder it was revolutionary! And no wonder that Intel started to have doubts about the ability of their x86 line to stay competitive with the new RISC alternatives that soon started to populate the processor landscape. RISC was more than just a new design point for microprocessors. It also was a liberation. Because RISC designs were simple, deciding to create a RISC processor was, relatively speaking, a low-cost project. Even comparatively small system companies, like Sun (in 1984), which previously had to rely on semiconductor companies for CPUs, were able to take control of their own processor destiny, driving designs to meet their system-level goals rather than the lowest-common-denominator requirements of some volume chip market. RISC Performance Basic Performance Equation per Task {factoring indifference in RISC vs. CISC ISA) . Instructions in Task Task Time = we X Cycle Time | Total number of instructions in a task ~ 2X worse than CISC | fined by SA stare (omer, 05) | ‘Average number of instructions executed Per Clock cycle 10X better than CISC focus of micoarhitectua improvements wiser ut, 000 econ, ber ace {lock cycle time - equal to or better than CISC fe ecg basi ine age, agely eerie by semianducter technology In Theory, RISC at least 5X faster than CISC | Gain driven by improvements in IPC Initial Goal - single-cycle execution (IPC =2) Long-Term Goal - continually increasing IPC (2, 3, 4 ....) Result is wider 8 wider superscalar designs with more outoforder execution, etc. Here's the secret to the revolutionary performance levels achieved by RISC processors, compared to their older CISC rivals. To be fair, I'm using a more complicated formula that counts not simple MIPS, but factors in the difference between CISC and RISC instruction sets. How long it takes to accomplish a particular task is a question of the total number of instructions a machine has to execute to get the job done, divided by the average number of instructions the machine executes each clock cycle, times the clock cycle time. [Plug in some numbers and give an example of using this formula] Since RISC instruction sets are simpler than CISC instruction sets overall, it is necessary to execute more RISC instructions than CISC instructions to accomplish the same amount of work. This is typically about a factor of 2 to 1. So RISC machines start off with a performance handicap, but make up for it by their ability to efficiently dispatch instructions. The original Vax, for example, needed about 10 clock cycles to execute an instruction, whereas the early RISC machines all set “single cycle execution” as their goal. Assuming no difference in clock rate between a RISC and CISC processor, that meant RISC designs were about 5 times faster overall that CISC designs. Which is pretty much what we saw in the previous example comparing the 386 with the R2000. After reaching the goal of “single cycle execution,” RISC processors moved in the direction of increasingly wide and complex superscalar designs, in an effort to continue to drive up IPC. A Limit to RISC Performance | Sew weap ase The Superscaler Barrier Problem: The complexity of dynamic superscalar ‘out of order instruction issue increases nondinearly ‘with the number of instructions issued each cycle, with more and more impact on cycle time " F 2 = i: i: Bo cL a gs ; | Time By the mid 1990s, RISC processors had accomplished a great deal. They had progressed from 32-bit to 64-bit architectures, they were able to issue 4 or more instructions every clock cycle, and they had incorporated all the more practical innovations that had been proposed for higher performance, including out-of-order execution, register renaming, sophisticated branch prediction algorithms, set-associative caches, and both hardware and software prefetch. There were two unintended consequences of these architectural refinements, however. The first was to effectively eliminate any meaningful distinction between RISC and CISC along the complexity scale. The 300% difference between a 100,000 transistor RISC processor and a 300,000 transistor CISC processor was highly significant, The 5% difference between a 3.3M transistor RISC processor and a 3.5M transistor CISC processor was totally insignificant. ‘The second was to make RISC processors victims of their own success. Each step up in performance required progressively more design effort, with steadily diminishing results. In particular, the drive to higher and higher levels of IPC tended to run out of steam around the 6-issue mark. At this point, the detrimental effect on clock rate of a still wider-issue design tended to outweigh any performance gain from the resulting improvement in IPC. complexities of parallel issue to software Simple inorder hardware design Unlimited IPC scaling (to the extent new parallelizing compilers ‘can deliver independent parallel instructions) 'No impact on cycle time ~ (clock speeds can improve as fast as semiconductor technology delivers shorter transistor gate lengths) | The Superscaler Barrier = a 2 3a : £3 z> ge Bs EB: £3 As the limit to the ability of RISC processors to continue to push out along the parameter of constantly increasing IPC became clear in the early 1990s, processor architects began to cast about for new architectural ideas that could avoid the superscalar barrier. For better or worse, there was a candidate ready and waiting just off-stage, in VLIW or Very Long Instruction Word architectures, VLIW theory was a close contemporary of RISC theory. Indeed, the two most notable ‘VLIW start-ups of the 1980s were both founded in 1984, the same year MIPS was founded to pursue RISC architectures and Sun started work on SPARC. Although there are a number of parallels between basic RISC vs. VLIW concepts, including the notion of pushing complexity out of hardware and into compilers, the notable difference for present purposes is that VLIW is a strategy specifically intended to reach very high levels of ILP. In theory, the only limit to ILP in VLIW architectures is the ability of compilers to detect and organize parallel instructions in code. VLIW hardware is designed to allow arbitrary increases in ILP, to the limit of current transistor budgets to accommodate parallel execution units, with absolutely no impact on clock cycle time. As the superscalar barrier loomed closer and closer, it was perhaps only natural therefore that VLIW would come to the fore as a solution to this problem. The most surprising thing was the alliance created to promote this solution, between HP, one of the RISC pioneers, and Intel, the most prominent critic of RISC. an ie ss oa = Osun Three EPIC Truths 4 1. Evenin theory, offers no performance discontinuity - rather, 2 | _ way of extending the RISC performance curve (based on steadily increasing IPC) past the limits of dynamic superscalar issue. 2. In actual practice, provides no performance breakthrough ~ rather, EPIC processors aiso stall at a peak issue rate of 6 instructions/clock {due to lack of ILP in typical code). 3. Ineality, no design simplification ~ rather, new complexities in the EPic architecture outweigh any gains from eliminating dynamic outoforder superscalar issue. 5 4 3 2 a a0 3 The EPIC Era }- The RISC Era The Superscaler Barrier 3 $ é zg 3 z 5 It is now almost exactly a decade since EPIC was first announced by Intel and HP, in June of 1994. Which is to say, our assessment of RISC vs. EPIC no longer needs to rely on a practical understanding of the shortcomings of RISC vs. the theoretical advantages of VLIW architectures. With two generations of Itanium already shipping, and a third version of the processor scheduled to begin shipping this summer, we now enjoy a practical understanding of EPIC, too. Three major points have become manifest. First, although the point to EPIC was to break through the superscalar barrier that threatened to throttle RISC designs at the 6-issue mark, EPIC, unlike RISC, did not promise any sort of sudden performance discontinuity, but rather was a strategy for smoothly extending the curve of rising IPC (past 6). Second, the superscalar barrier is not the problem to reaching higher levels of IPC. For most programs, the IPC curve breaks before the superscalar barrier is reached, due to a paucity of parallel instructions. In fact, there isn't enough ILP in most code to fully sustain a 4-issue machine, let alone a 6-issue machine. Third, EPIC processors do not succeed at reversing the trend toward more and more complex hardware. If anything, Itanium processors are more complicated to design and implement than other types of processors, not less so. = | - @Sun | Why EPIC Isn't the Next Shift cisc RISC cise ePIC i386 R2000_ Xeon DP __itanium 2 Year 1985 1985 2002 2002 Technology | [Link] | 20ucmos | otgucmos | o.1eucmos Die Size 103 mm2 | _ 80 mm2 145 mm2 421 mma Transistors | 275,000 | 115,000 | 55,000,000 | 221,000,000 Cache caKpext. | s2kBon. | 532KBe 3.286 MBs Package | tszeraa | wacrca | ronae |... OOGRe Power 5s aw 130W 1 GHz }1040 SPECint2K| 810 SPECIN2K 10.1 Spec89 T1048 SPECtp2K |1431 SPECIp2K Tl comparable Wore (old) To clearly see why Itanium isn't going to revolutionize the processor market the way RISC processors changed the rules for processor design, let's make exactly the same comparison between Itanium 2 and the reigning high-end x86 design of 2002, as we did between the R2000 and the reigning high-end x86 design of 1985. As this table shows, not only does Itanium 2 fail to provide an order of magnitude improvement in price/performance over Xeon, you're hard pressed to find anything that Itanium does. better than Xeon. The factors that affect the price of systems — die size, package, and power consumption — are all substantially worse for Itanium than Xeon. The Itanium design is not simpler than Xeon, but far more complicated. As a consequence, Xeon outclocks Itanium by a factor of 3. For this reason, even outfitted with a cache only 1/6 the size of anium 2, it outperforms Itanium 2 on integer code by 25%. To be fair, Itanium 2 is built with an older, cheaper process technology than Xeon, albeit that's largely a reflection of the longer design cycles required for Itanium processors, and is more than offset by Itanium's much larger transistor count and associated giant die size. The one area where Itanium does excel is in its floating- point performance. But even here, the appeal of Itanium is not to the thrifty, those looking to pay less and get more, but rather to those in search of the last measure of floating-point performance, and desperate enough to purchase it at any price. In short, Itanium entirely lacks the distinctive signature of an important development in processor history: namely, the ability to cause a sudden and significant forward movement along the path of progress (towards higher functionality/lower cost). The Commodity Shift The Microprocessor Era The second new design point 1971. microprocessors * Pentium Pro reaches RISC performance | | 1985 cave (wth RISC micoareiecture), AMO & ne iat sete 10 year egal isc ‘ope overs done competion 1995 ‘commodity Volume & Processors Compabity Price Perormante—— SSmpetion Power consumption OMPeUt Leverage dominant code base with specialized hardware designs to achieve a key advantage ina volume market 1 Lower cost (AMD, Cyrix/NSC) Lower power (Transmeta) | | Higher specialized performance (Nexgen, MicroUnity) L | Although, in the end, Itanium turned out not to be the bright morning of a new era of processor design, but rather a false dawn, there was a significant shift in the processor landscape that happened in the mid-90s. Ironically, this involved Intel's other processor, the one they intended to replace with Itanium, As mentioned, by the mid-1990s, 64-bit, superscalar, out-of-order, etc. RISC designs required huge transistor budgets (>3M) to implement. At this level of complexit any meaningful distinction between RISC and CISC disappeared, x86 designers fought back by the simple expedient of adopting RISC architectures wholesale, creating RISC execution engines, and building front ends that converted complex x86 instructions into 2, 3, ot more RISC operations. True, this meant adding an extra layer of complexity on top of the RISC engine in an x86 design, but one more complexity, among so many others, was barely noticeable Tve picked 1995 as the point when the balance of power shifted away from RISC and over to the x86 architecture. 1995 was the year Pentium Pro briefly took the performance leadership crown, demonstrating the viability of the x86 architecture, and also the year that Intel finally terminated its legal efforts to prevent other companies from building 32-bit x86 processors. The commodity shift was more subtle than the RISC revolution, both because the x86 was old not new, and because the shift was based not on a radical leap forward in computing performance, but rather a reduction in the cost of computing based on volume and clone competition. Fr i. ] Sun Intel's Unintended Triumph | Intel Launches 1A-64, it's 3" attempt to replace the x86 with something less accidental, more controllable “ EPIC” era falls to dawn because, at a hhigher price point, the same (or less) performance always loses Intel dooms ttanium by demonstrating that the x86 can reach RISC performance curve “Commodity” era dawns because, at the same (or higher) performance point, lower cost always wins RISC comes to dominate 32/64bit | 1995 ‘embedded markets, where volume ae without compatibility is possible The last decade has seen a proliferation of companies convinced they can carve out some piece of a huge commodity market, by building an x86 clone processor that offers a special advantage: if not in terms of still lower cost, then in terms of a radical reduction in power consumption, or a performance edge in some special computing, niche. I think everybody has gotten comfortable with the idea that the x86 is really a commodity processor, capable of being produced by many different companies, with Intel holding a favored place only by virtue of their size and the associated economies of scale they enjoy, their ability to address multiple design points, and their aggressive pursuit of very high frequencies. Itis worth noting, however, that this position, namely, the premier x86 producer among many other actual and potential producers, is precisely the position Intel never intended to occupy. When Intel gave up its legal battle to outlaw clone competition, it did so in the expectation that it soon could replace the x86 with an exclusive proprietary design that would enable them to avoid running in the commodity race, rather than constantly keep striving to stay ahead in it. That Intel's intended end run around the commodity market failed should come as no surprise. As we have seen, Itanium provides absolutely no price/performance advantage, but rather operates at a considerable disadvantage in this respect. Whereas the mid-1990s legal transformation of the x86 from an Intel proprietary design into a commodity processor, combined with Pentium Pro's existence proof that x86 processors could offer very competitive performance, had exactly the opposite effect, endowing the x86 with a clear price/performance advantage against the field. aka “Horizontal Model” (based on 18M PC: 8/12/3983 to present) component based distribution, support (ors iteaation& ‘istomization bel Gate, PQ, igh specialization) ee 1m Two privileged iscrete hardware technology level tech companies 1 Worzonal ech Sime nate money (unt Subsumed) 1 Driven by volume, market orees 1 Delivers generic functionality at lowest possible cost value shift Here's a visualization of the commodity or “horizontal” model that has driven the x86 processor to dominance. There are two principal levels to this model, a technology level, composed of independent hardware and software suppliers, and an assembly level that handles building, selling, and servicing systems. The system assemblers — the Dells and Gateways and HPQs, etc. of the world — are free to select whatever hardware and software technologies they like from among the competing suppliers of technology ~ albeit, Intel holds a favored Position among the hardware suppliers as does Microsoft among the software suppliers. Still, neither Intel nor Microsoft is quite as dominant as it might like to be or, indeed, once was, thanks to AMD and Linux, respectively. This model is chiefly remarkable for its ability to drive out cost, though the price paid for the unerting ability of this model to home in on progressively lower cost points is a kind of homogenization at the system level, that tends to make assembled products indistinguishable from each other. There also is a certain built-in instability to the model, as Moore's Law in the processor realm and some analog of that principle of ever-increasing size in the software world, enable Intel and Microsoft to assume more and more of the necessary system-level technology into the CPU and OS, respectively. Thi winnowing of technology suppliers creates tension, since a monopolistic end- point for either hardware or software technology is incompatible with the market-driven, competition-based nature of the commodity model. Intel's Vision for Hardware Technology | (riven by Moore's Law of everexpanding transistor budgets/CPU functionality) All devices that use CPUs | (price competition between many assemblers to drive maximum volumes) Intel's Dilemma: Can't get here without Itanium (0 proprietary lockout) ‘can't get here with Itanium (no commodity model) sole technology supplier (after the last value shift) But that's not to say a company can't dream about becoming a monopolistic supplier. Certainly Intel does. This is their vision for the end point of the hardware technology curve, after Moore's Law has enabled them to incorporate all the (technically significant) hardware functionality needed to build a system into their CPU chip. Note that the “horizontal” part of the model now applies only its top assembly level, where Intel certainly hopes to see many different assemblers, locked in a fierce price competition that drives the maximum volumes for their CPUs. But the technology layer has now taken on a decidedly ... vertical appearance. The problem Intel faces in reaching this ultimate endpoint is that it’s flatly inconsistent with the commodity model. Even after the system-on-a-chip becomes standard for the most complex systems (as, by Moore's Law, it eventually must), as long as the x86 reigns, market forces ensure that there will be alternate suppliers of this commodity part. Which is to say, in order to reach the goal envisaged here, Itanium not only must become successful at the high-end of the processor market, it ultimately must replace the x86 in the volume end of the market. But, as we have seen, Itanium can't replace the x86 in volume use exactly because it doesn't fit the commodity model, and hence can't reach the sort of price/performance point that can be achieved only through that model. With Billions & Billions of | transistors per chip ... .. all system functionality (except large DRAM) gets integrated into the CPU. Intel's challenges aren't confined to the dilemma that the same commodity model that has elevated them to their current lofty position as a World Power simultaneously prevents them from ruling unchallenged in their domain. Moore's Law is an independent source of worry. On the one hand, it enables Intel continually to do more for less, providing an endless wealth of new opportunities over time. On the other hand, it places them on an endless treadmill, where doubling the world’s appetite for transistors every two years amounts to nothing more (Moore?) than running in place. For purposes of this discussion, the importance of Moore's Law is that, quite apart from its specific implications for Intel (whether for good or ill), it continually redefines the boundary between the CPU and the rest of the system. In effect, Moore's Law means that the capacity of the CPU keeps expanding, encompassing more and more of the system until, in the end, only the CPU remains (along with some memory chips). As the pendulum of Moore's Law swings further and further in the direction of the system on a chip, however, it also creates more and more tension with the component-based commodity model. The commodity model is built on multiple suppliers, characterized by low integration and high specialization. But a single chip cannot be parceled out across multiple vendors. It is, by definition, an integrated whole, rather than an assemblage of independent parts collected from interchangeable suppliers. The Real Endpoint of Moore's Law With Billions & Billions of transistors per chip ... ««. CPU design becomes system design and system design becomes CPU design! system companies (with CPU design expertise) ‘There wil be no end othe troubles of computing or of computer users themselves, til system designers become cip designers inthis worl, oil those we now cll hip designers cally and ‘truly become system designers, and computing power andvison thus come int te same hands” {with apologies to Plato's Republi) In short, Moore's Law continually redefines the design landscape. It's real endpoint is not Intel's vision of a single, monopolistic supplier of all hardware technology, but a world in which CPU design has become system design, and system design, conversely, is CPU design I think it's fair to ask if this endpoint is a world friendly to chip companies, all of which were created under the assumptions of a horizontal model, and whose business is premised on the notion of supplying specialized components to volume markets? Or is it rather a world friendly to companies with the depth of expertise needed to create systems, and the breadth of experience required to integrate many different functional elements into a working whole? In truth, it seems to me a world that favors a third kind of company: one that combines the expertise of a systems company with the specialized knowledge needed to design chips. Fortunately, there are a few companies that fit that description, descendants of the system-side RISC era with the strength of conviction and nimbleness of product to survive through the semiconductor-side commodity era. Sometimes winning is just a matter of surviving long enough to show up in the right place at the right time, Why SPARC! | So why does Sun design its own CPUs? (not because it wants to be in the CPU business but rather) Because Sun is a systems company And, in the end*, system design is CPU design The altemative is to sell and service systems designed by some other System/CPU company (Don't confuse semiconductor fabrication technology with CPU design technology) If the foregoing analysis is right, the wheel of change is now poised to turn again, moving us another notch further along the path of progress. Specifically, a design point that lets us put not just hundreds of millions, but shortly billions and billions of transistors on a single chip, provides system companies with the opportunity to take advantage of another new design point that, like RISC, promises to combine a radical advance in performance with a radical reduction in cost, At Sun we've taken to calling this new design point Chip MultiThreading or CMT for short. Unlike RISC, however, CMT is not strictly a compute-engine vision, but rather fits into an overall system-level vision that we call “throughput computing.” And that's really what I want to talk about today, though it's taken us awhile to get here. But let me pause before turning to talk specifically about CMT and throughput computing to answer the question most often asked about Sun's persistence in building computer systems around its own SPARC processor, namely, why do you still do it? RISC hasn't been truly relevant since, well, sometime around 1995, Did Sun somehow miss the whole commodity shift? The answer to that question, of course, is no, we didn’t miss the commodity shift. In fact, today Sun builds systems around the commodity x86 processor and the commodity Linux operating system. But for a systems company (as opposed to a systems assembler) that can't be all there is. And for a company to stay in the systems business, it ultimately has to stay in the CPU business because, in the end, the two are one and the same thing. m systembased = No privileged companies Cluster Tools Sa ae Brey acs Sen) Bestof Breed Solaris OE Highly Multithreaded SPARC CPU functionality for most efficient execution of target task(s) “The Network is the Computer” Open Systems | | At the system level, here's a snapshot of Sun's overall vision. Admittedly, this is not a generic vision of some commodity system. It is a vision specific to Sun, and incorporates Sun's ambitions to drive technology in given directions for the ultimate benefit of its SPARC/Solaris customers. To the extent the goals underlying this vision are successfully realized, both Sun and its customers will benefit. Although I'm going to be focusing on just one element of the big picture shown here, the highly multithreaded SPARC CPU that provides the hardware foundation for this future Sun system, it’s important to understand from the outset that the CPU strategy I'll be talking about is not a stand-alone proposition. It makes sense only within the context of the Solaris OE and the robust threading model it provides, and the inherently multithreaded network applications that Sun systems are designed to run, In short, CMT is not a “one size fits all” kind of technology Nonetheless, while the above picture is intended to deliver precisely the coordinated functionality required for the most efficient execution of the tasks targeted by Sun systems, absolutely nothing prevents other system companies from designing CMT-style processors that fit into their own system vision, and that provide the coordinated functionality needed to solve other types of system-level problems. There are no privileged system-level visions, just the different visions of different companies. And may the best vision, backed by the most effective implementation, win. ‘The Microprocessor Era The third new design pe 1971 s*ICbased CPUs Processor 1985, 1" commercial RISC CPUS 1995 2" RiSCcompettive 6 Commodity 64-bit Processors 2005 1" cmt cpus Chip Efficiency Price/Petrmance—~) Muaithreading 112th Lee race Power consumption| ar enip atreadng Throughput (LP) Strategy: Reinvent the (stalled) IPC curve by switching from ILP to TIP, focusing on tasks per time (not time per task) ‘Advantages: Much more efficent deployment of transistor resources ‘Much higher utilization of on-chip resources ‘Much less idle time ‘True practical IPC scaleability (to large numbers) So here's what I think is going to be the next big wave of innovation in CPU design, CMT. 2005 is the year that Sun will roll out its first radical CMT design, marking the beginning of the CMT shift. Note that these shifts, from RISC to commodity to CMT, seem to come on 10-year boundaries, With just three data points that well may be only a coincidence, but I'm planning to keep a weather eye out as we get nearer to 2015, in case there's some deeper cause behind the timing of this pattern. Ina nutshell, the basic idea behind CMT is to revitalize the drive to higher levels of performance by creating designs that use both on-chip resources and processor clock cycles far more efficiently than any designs yet produced. The key here is to invert the performance focus, to switch from looking for ways to minimize the time needed to execute a given task, and look instead for ways to maximize the number of tasks that can be executed ina given time. The advantage of this inversion of the performance problem is that it enables the creation of processors that make much more efficient use of available transistor budgets, i.e., that can use far more of their available resources at the same time, and that suffer many fewer idle clock cycles. Further, it lets designs be created that allow almost unlimited IPC scaling, as Moore's Law enables the creation of bigger and bigger machines. The result will be systems that offer unprecedented levels of throughput. The Performance Trap Relatively small gains in CPU performance result | | | in relatively large losses in chip efficiency (more on-chip resources sitting idle more of the time) | Idle Resources ‘The more aggressive the CPU in attempting to drive up IPC, ‘the worse the problem Idle Time ‘The higher the CPU clock rate, the worse the problem Perhaps the simplest way to understand this next quite radical departure in processor design is to start from what I call the performance trap. If you think back to the performance vector slide I showed earlier, recall that there are only two basic parameters that control how quickly a processor can execute instructions, namely, the average number of instructions it can execute in each clock cycle, known as IPC or Instructions Per Clock, and how fast the processor clock ticks, i.e., the processor's operating frequency. The performance trap is simply that everything you do to drive up performance has the unfortunate side-effect of making the processor less efficient. Further, as processors attempt to reach ever higher levels of performance, the problem rapidly gets worse — each new performance gain tends to be smaller, yet must be purchased at a still larger loss of efficiency. More specifically, everything a designer does to push up the IPC figure of merit results in more on-chip resources sitting idle at any given time. And everything a designer does to push up the CPU clock rate results in more clock cycles in which the processor does nothing at all. Why do we care how efficient or inefficient the processor is? For two reasons. First, inefficiency is not merely wasteful, it's expensive. Second, if we could somehow begin using all our on-chip resources and all our CPU clock cycles efficiently, the potential performance improvement is staggering. Osun Tesnurce cage ares oval due w ache mines Frebssed on amount of vate n opi coe ‘Gap conceptual numbers co nt reser mesure les Wide superscalar designs achieve higher performance by using more resources less efficiently Today's widest Gissue designs are operating at a relatively low efficiency aise ten 7a Instructions Per Clock (performance) Let's start with the issue of idle resources. If we graph IPC against resource usage, the resulting efficiency curve is going to look something like this. Ignoring the problem of stalls for the moment, a processor that is resourced to issue just a single instruction in a clock cycle is operating at perfect efficiency, in the sense that it is always doing everything it is designed to do. ‘As you resource a processor to issue additional instructions in the same clock cycle, efficiency necessarily drops off, because you have to design on the “best case” assumption, that the additional slot will be filled every cycle, but in fact, the “best case” will not always be realized. Thus, if you design a machine with two issue slots, on some cycles the second slot will not be filled, and the associated resources hence will be left idle. Add a third slot, and the associated execution resources will be left idle even more often. And so on: each new issue slot carries its own burden of added resources to support the additional issue, but the extra resources will pay off less and less often for each added slot. (Assuming what seems only reasonable, that the chance of issuing nn instructions is always higher than the chance of issuing n+1 instructions.) Today's widest designs try to issue 6 instructions every single clock cycle. However, in fact, in many programs, they do well to average 2 instructions per clock. Which is to say, these designs operate at very low levels of efficiency. This “efficiency crises” for high-performance processors is now serious enough that steps already are being taken to remedy it. HyperThreading Js an attempt to solve the inherent inefficiency of wide-issue machines “comeing hema Se re, eam eaeenennae Site Make mor fet use of valable crt resoores by runing two treads nparal haf the Ihe pertirend a8 ‘rade some pershvead performance ptentat gains abt to un ‘ore threads in parallel, (patil efficiency) Resource Usage % 45 6 7 8 9 0 1 WB Instructions Per Clock (performance) “ypertrendng stl ame fx Syme MuitTivea ig or SMT The term you may have seen in this context is “hyperthreading,” which is Intel's name for an idea that comes out of some academic research done at the University of Washington and elsewhere under the term “Symmetric MultiThreading” or SMT. The basic idea behind SMT is fairly straightforward. Build 6-issue machines (for example) that have the ability to divide themselves into two virtual 3-wide machines. Then, in code where 6 instruction issue is a very distant hope indeed, put your idle resources to work running a second thread in parallel with the first thread, on the built-in second virtual core. Like everything else in life, there are some trade-offs to this strategy. Obviously, you sacrifice some portion of your potential single-thread performance, though the sacrifice is smali on the operative assumption that opportunities to issue more than 3 instructions at a time are rare. (If they weren't rare, the 6-wide machine wouldn't be so inefficient.) And while you do retreat back up the efficiency curve to a more productive use of the available on-chip resources, a hyperthreaded machine is still far from perfectly efficient, since three-wide (and even two-wide) issue will fail some of the time. So, yes, hyperthreading can be an improvement. But it’s far from being a complete cure to the inefficiency plaguing today's wide-issue designs. Maximizing E Efficiency If efficiency is the goal, there's a far Ne solution than Hyperthreading Build a single-issue core It's the only perfect solution! (ot course, you have to trade tvead performance) {spatial efficiency) SGP UOoOo OS Resource Usage % 45 6 7 j Instructions Per Clock (performance) If I may be allowed an analogy, hyperthreading as a solution to wide-issue inefficiency is a little like solving the pain associated with hitting oneself over the head by putting on a helmet. Yes, it does succeed in reducing the felt pain. But the fact remains that you're only in pain because you keep hitting yourself over the head, So there's a much simpler and more direct way to eliminate this pain than putting on a helmet — stop hitting yourself. Similarly, the only reason hyperthreading makes any sense at all is because you've gone and build a tremendously inefficient wide-issue machine, and now you're trying to figure out some way to make it less inefficient. But there's a simpler solution. Don't build a wide-issue machine. Problem solved. In fact, if maximum efficiency (rather than maximum single-thread performance) is the goal, the very best solution is to build a simple, 1-issue machine. Because, as soon as you add a second issue slot, efficiency is going to drop off dramatically. Indeed, where efficiency is the question, a single- issue machine is the only absolutely perfect answer. Of course, there's a trade-off. If you want to maximize efficiency, you have to surrender some performance. Conversely, if you want to maximize performance, you have to surrender efficiency. The good news about trading in the direction of efficiency rather than performance, is that this trade typically sacrifices mere pounds of performance, compared to the tons of efficiency that tend to get thrown away when more performance is the goal. Le a osun Space-Efficient Performance TM How to avoid the” | 8 xissue cores idle resources performance trap Drive up performance while ‘Two 84PC designs ‘maintalingeficlency by SPplyng transistors not Both require about needed for wide sme the same size chip to building motple ‘issue corer ‘once Resource Usage % (patil efficiency) ges endleve perfomance 1 Bissue core forfarigher throughput 4 5 6 7 8 9 Instructions Per Clock | (performance) So here's the secret to building a processor that is able to use all its resources efficiently, and that also takes advantage of the very large transistor budgets that are currently available to CPU designers. Namely, rather than chase Instruction Level Parallelism down the curve of vanishing efficiency by building a massively-resourced, 8-issue, single-core processor, switch playing fields. Instead of going after ILP, go after Thread Level Parallelism by using the available transistors to build 8 single-issue cores on a die. Both processors can work on 8 separate instructions every single clock cycle. Both processors require about the same size (cost) die to build. The difference is, every single core on the 8-issue TLP machine is operating at maximum efficiency, supporting only the minimum set of resources that it both needs and can use all the time. Whereas, the 8-issue ILP machine is operating far below optimal efficiency, with most of its resources sitting idle most of the time. ‘True, any given thread will finish executing sooner on the 8-issue ILP machine. But the 8-issue TLP machine will finish executing far more threads in any given unit of time. Which is to say, if we change our definition of performance from minimizing the amount of time necessary to execute a given unit of threads, to maximizing the number of threads that can be executed in a given unit of time, the TLP design is far higher performance than the ILP machine (corresponding to its far more efficient use of resources). Efficiency, Higher Single Thread Single Thread Performance Performance active J idle logic Bi logic Here's a graphic that shows the difference between the alternative design points that we've been discussing. On the left, we have a simple scalar design. As mentioned, this design provides the most efficiency but the least performance. To boost performance (defined as the ability to execute the same thread in less time), designs moved from scalar to superscalar in the early 1990s, settling on 4- issue machines as a sort of “sweet spot.” Although these machines “peak” at 4 IPC, their average IPC is certain to be lower ~ 2 IPC is probably on the generous side. Which means half the peak issue width amounts to “dark logic” overhead. Let’s say we now push on to an 8-issue design, in pursuit of still higher performance. This does bring more active logic to bear on the computational problem (say the average IPC now goes up to 3). But the toll in efficiency is high, ice., the ratio of idle “dark” to active “light” logic gets even worse. Indeed, the inefficiency of designs this wide is so obvious that it calls out for remedy. Enter hyperthreading, which backs away from maximum single-thread performance in order to run two threads in parallel, restoring the overall efficiency of a 4-issue design (along with its somewhat lower level of single-thread performance). But there is a far more efficient solution than hyperthreading: an 8-issue machine designed not to run | thread very quickly, or even two threads somewhat less quickly, but 8 threads at the standard scalar rate of execution. If efficiency is a goal, and performance is defined to mean maximum throughput, this is by far the best alternative among these possible design points. _— @5un ] More Idle Time (ei sage asus no sais oe oak fremuces or sauce cots ‘ph eoncepta, mbes 9 ra rein messed lus Higher frequency designs achieve higher performance by using more clock cycles less efficiently gases designs are operating at relatively low efficiency / Today's highest frequency cycle Usage % & (temporal efficiency) ° ‘© 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800, Clock Rate (performance) The curse of idle resources, though, is only half the performance trap. The other half of the trap is the curse of idle clock cycles. Here's the shape of the problem. Recall that the other half of the MIPS performance equation (besides IPC) is clock rate or processor MHz. (actually GHz these days for high-performance machines). But, sure enough, as clock rate goes up, processor efficiency ~ measured here as the percentage of total clock cycles during which the processor is accomplishing useful work (vs. idling along, accomplishing nothing of any use) — goes down. Don't stop me if you've seen this curve before. It's exactly the same progressive decline in useful clock cycles, as the overall processor clock rate is driven up, that we saw for productively employed resources, as the overall processor IPC is driven up. And again, we've now reached a crises point. Today's high frequency designs are throwing more and more clock cycles at problems in the interest of ever higher performance, but with less and less effect. Indeed, compared to, say, a 1 GHz design, a 3 GHz design is shockingly inefficient in its use of available clock cycles, We've seen there is way around the trap of ever more idle resources. Is there also some way to avoid the trap of more and more idle cycles? —— Sun | The Basic Problem Typical High-frequency RISC Pipeline (Integer Ops) Miss Every Li Miss stalls the pipeline while the required instruction or data item is fetched Before we turn to consider possible solutions, let's take a minute to understand the root cause of the problem here. This is a picture of a typical high-frequency pipeline. Note that although it takes only a single cycle (stage #8) to execute a typical integer operation (in keeping with the RISC design precept of “single cycle execution”), there's quite a bit of processing that has to go on before an operation can be executed, as well as some processing that has to go on afterwards. But as long as the pipeline stays full, i.e., every stage is engaged working on an instruction — the machine operates at peak efficiency, with no wasted clock cycles. All modern pipelined processors are designed to keep up a steady execution flow, as long as both instructions and data are no further away than the on-chip Level One (L1) caches. Indeed, the need to access the L1 caches in step with the pace at which instructions flow through the pipeline, is a primary reason these caches have to remain relatively small. The problem is, the L1 caches are small. For most real programs, they hold only a fraction of the total instructions and data required to completely execute a given thread. And every time there's a “miss” in cither the L1 instruction or data cache ~ the next instruction or data item that's needed for the computation is not in the L1 cache — the pipeline has to stall while the missing item is fetched from some more distant location. Effect of increasing CPU Frequencies on Memory Latencies) [fe be all memories st othe te > The higher the clock frequency, the longer the stall on every Li miss. nesmrey nmbere cc resp wating The length of the stall depends on how far it's necessary to go to get the missing item. If the item is on the processor chip in an L2 or possibly L3 cache, typically the stall lasts only a few clock cycles. If it's necessary to g0 off chip to a local high-speed cache, the stall might be a few 10s of cycles. If the item is not in cache but has to be fetched all the way from main memory, then the stall is likely to take 100s of CPU clock cycles. And should the item have to be retrieved from disk, well, measured by CPU clock cycles, that's a lengthy trip indeed. 5 ms., the seek time of a fast hard disk drive, may be far less than a heartbeat, but it is 5 million clock ticks for a 1 GHz processor. The basic problem with escalating CPU clock frequencies is that the memory hierarchy (including any large on-chip caches, external cache, memory, and disk) does not scale up in speed at the same pace as the CPU pipeline itself. Which is to say, every time the CPU pipeline clock gets faster, in effect the entire memory hierarchy shifts to the right on this slide. One of the consequences of making the CPU clock tick faster is the unfortunate fact that more ticks will occur during an interval in which the processor pipeline is stalled, waiting on a memory access. = ncross the industry, The gap only gets worse with time today’s chips are largely able to execute code faster than we can feed them with Instructions and data, ‘There no longer are performance bottlenecks in ‘the floating point multipler or. Integer unit... Ina CPU Frequency ‘Memory Speeds study using TPC... three ‘ut of every four CPU cycles retired zero instructions; Operational Frequency Time “1 expect that over the coming decade memory subsystem design will be the only important design issue for microprocessors.” - Richard L. sites (Alpha architect) ets fon th ama Sp oper p99 This truth has been a dominant factor in processor performance for the last decade, as these quotes from 1996 attest. The Alpha, the last important high-end RISC processor to be introduced, was designed around a “speed- demon” or clock-oriented performance strategy (rather than an ILP-based or “brainiac” performance strategy), and so was one of the first designs to fully appreciate the magnitude of the problem of idle clock cycles. As Richard Sites here testifies, measurements show that in many programs as many as 3 out of every 4 clock cycles accomplish no useful work, but are spent waiting for memory accesses to complete. Back when the Alpha was the clear leader in clock frequency, it was sometimes said jokingly that Alpha “waited faster” than any other processor. Which, in fact, was literally true, if the measure of “fast waiting” was the number of CPU clock cycles that expired uselessly during a memory access. As the graph illustrates, this is not a problem that can be fixed by waiting for things to get better sometime down the road. To the contrary, the problem only gets worse with each passing year, as CPU clock frequencies continue to rise at a much faster rate than memory speeds increase. | Computing Faster Is Little Help decreasing compute times are of very modest value il emo without corresponding improvements in memory latencies [7] soy trey Ce aad Here's an illustration of the “performance trap” encountered in trying to reach higher performance levels by means of higher clock rates, Doubling the processor clock rate cuts actual processing time in half. That's the good news. ‘The bad news is, unless memory is somehow made faster, all doubling the CPU clock rate means for a memory stall is that twice as many ticks get wasted waiting for the next instruction or data item to be returned. To see the magnitude of the problem here, let's plug some sample numbers into this illustration. Say we spend 100 clock ticks computing before we get an L1 cache miss, and then have to wait 300 clock ticks for the needed item to be returned from memory. Obviously, the machine is not terribly efficient in its use of clock cycles under these assumptions. In the small sample of computing time shown, the machine spends 400 cycles doing actual computing and 900 cycles doing absolutely nothing but waiting while memory is accessed. Now let's double the CPU clock rate. The good news is, the 400 cycles spent computing now happen in half the time previously required. But all the higher clock rate means for a memory access is that the machine now spends 1800 CPU cycles waiting on memory. True, the 1800 cycles don't take any longer than the 900 cycles did formerly, and the 400 compute cycles happen in half the time, so there is a modest overall speedup of about 15%. But there is a steep price paid in efficiency for this 15% performance gain, since while there still are only 400 compute cycles, there are now 1800 (not just 900) idle cycles. osx Vertical Threading | | 3 oter4s is a way to solve the inherent | inefficiency of high frequency machines Bea 2 22 Keep core busy at all times by switching between tasks It's the only perfect solution! Cycle Usage % (temporal efficiency) st 0 3 © 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800 | Clock Rate (oerformance) _ | Fortunately, there is a way to solve the inefficiency inherent in moving to higher clock frequencies. Rather than follow the curve of single-thread performance down in efficiency, increase the number of candidate threads a core might operate on in a clock cycle. Then, when one thread stalls on a memory access, rather than sitting though a period of enforced idleness, waiting until the stalled thread can start up again, simply switch off to another thread. When that thread stalls, too, switch again. And soon. With proper planning and a little luck, few or no cycles will be wasted waiting on memory accesses for stalled threads. = @sun Time-Efficient Performance | | ‘ni wo tc esa TLP 4 threads per virtual core 3 How to avoid Dieu performance whe " mata fie the idle time eideane stomata the Gteaon tne of whcnes performance thread is able to run (isn't stalled trap ‘waiting for memory) Stalled 75%, sti 25% fees on rable singl end ILP 1 thread per virtual core 90 80 7 Cy 50 0 ee se ge oe #2 E terteatng (SMT, arotal heading) dige sgle tdssoe psa cre to fr more an ers" Seseseidt tempter sale ° Pj ij ij jj ya © 400 800 1200 1600 2000 2400 2800 3200 3600 4000 4400 4800 Clock Rate (performance) Here's alittle more detailed picture of how to avoid the idle time performance trap through vertical threading. Suppose each virtual (or physical) core in a processor is able to maintain concurrent state for, say, 4 separate threads of execution. On the assumption that any given thread is stalled up to 75% of the time waiting on memory accesses, three of the four threads might be stalled at the same time. But one of the four threads ought to be able to execute. When that thread stalls ~ as it will, probably sooner than later — the memory access for one of the three stalled threads may have completed, and the processor can resume that thread of execution. When it stalls, another of the candidate threads might be ready to run, and so on, While this scheme is not foolproof — it's possible that even with 4 candidate threads available to a core, there will be clock cycles where all threads are waiting and none are runable ~ it's obviously going to come far closer to perfectly efficient use of a processor's clock cycles than a single-threaded core. It will approach perfectly efficiently use of processor time, even if it does not perfectly achieve this goal. Note that hyperthreading (simultaneous multithreading) is distinct from vertical threading. Addressing the inefficiency of wide-issue machines (resource use) does nothing to solve the problem of each simultaneous multithread idling for many clock cycles, waiting on memory. Time Usage Comparison Single-Threaded Core Th eee (SRENRENENNEN Core compute time corel tne SS renoyancytme } Hm 4X Vertically-Threaded Core Thread 2 Here's a more explicit comparison of the way in which a single-threaded core, and a 4X vertically-threaded core, spend their respective clock cycles. The single-threaded core follows the classic compute-stall, compute-stall model shown at the top. In any given stretch of processing time, the larger part of the total time expended is spent waiting on memory accesses, with core idle time equal to the time taken by the memory accesses. The 4X vertically-threaded core follows exactly the same compute-stall pattern for each of its threads. No thread can run when missing a required instruction or data item; every thread must wait while such items are fetched. But rather than wait for a given thread to be able to resume execution, a vertically threaded core moves on to a new thread of execution that is not currently stalled, and keeps operating. When the new thread stalls, the core shifts again to a third runable thread, and when that third thread stalls, it still continues operating on a fourth thread. By the time the fourth thread stalls, with any luck at all, one of the first three threads will have completed its memory access, and execution can continue on that thread without missing so much as a single beat. And so on. The bottom line is the core stays constantly busy, gainfully employed at executing one thread or another all the time, with no idle time at all. True, there is a lot of memory access time needed to keep this multi-threaded core fed, but all the cycles required to access memory are overlapped with other access and compute cycles, and so effectively “disappear” in terms of their impact on performance. Now let’s put it all together. Here's a picture of what happens when you combine eight resource-efficient single-issue cores on a single processor die, and enable each of these cores to run in a highly time-efficient way, through 4X vertical threading. As this picture indicates, the prospect is a bit overwhelming. You end up with a processor that is able to keep 32 threads in play simultaneously. True, a maximum of just 8 threads — one per core — will be executing actively on any given clock cycle. But the other 24 threads are also all “live”= its just that they happen to be stalled at the moment, waiting on a memory access to complete so they can resume execution. So, in a very real sense, this processor is running 32 threads all at the same time, each thread just as fast as it is able to run, The secret behind this processor's astonishing ability to run 32 threads at once is its near-perfect efficiency. By virtue of its 8 single-issue cores, the whole processor die is put to use, without the overhead burden of many under-utilized execution resources that plague wide-issue designs. By Virtue of its vertically-threaded design, each of its 8 cores operate with few or no idle clock cycles spent waiting on memory accesses, without the ‘overhead burden of many idle clock cycles that plague both single-threaded and hyperthreaded designs. “Hyperthreaded” Processor CMT Processor use some on-chip resources ‘OHT Processors ‘recess a a | use all on-chip resources = core compute time ES core idle time === memory latency some of the time Admittedly, by using its “excess” fetch, issue, and execute capacity to run two threads in parallel, a “hyperthreaded” processor does succeed in rectifying some of the inefficiency of a very wide-issue design. But this retreat to two narrower virtual cores still falls far short of perfect efficiency, in terms of its ability to keep all the available on-chip resources gainfully employed at the same time. And hyperthreading does absolutely nothing to address the fact that each parallel thread will stall up to 75% of the time, while necessary memory accesses take place. Compare this picture to a CMT processor, where virtually all of the available on-chip resources are put to use by running 8 threads on eight single-issue cores (rather than just 2 threads on 2 multi-issue cores). And since each single-issue core is vertically threaded, all of their respective clock cycles can be spent computing (rather than mostly waiting on memory). The bottom line is, rather than keeping some of its resources busy some of the time, the CMT processor is able to apply all of its resources to computing all of the time. For designs of similar complexity, this difference in efficiency translates directly into a difference in performance. If the CMT processor brings, on average, twice the on- chip resources to bear on a compute problem, and four times as many clock cycles, it will outperform the hyperthreaded design by a factor of 8-to-1. Compared to a single-threaded design, the difference would an astonishing 16-to-1 — which is to say, just about the factor of 15X improvement that we've cited for our first CMT blade processor, compared to today's four-issue, single-thread blade processor. Why CMT Is a Radical Processing Advance per thread percore peak avg at same MHz processor cores threads © IPC IPC Efficiency Throughput “Niagara” 8 Pat 100% POWERS, 40% Itanium 3 33% SPARCE4 VI 33% Xeon 33% Opteron 93 33% Performance Comparison: dependent on me wp CMT architectures high none ‘ficiency = avg IPC / peak ie 7 rea ‘roghpats freeads mg PC ee en | How radical an idea is CMT? Here's a comparative look at the processor landscape circa 2005, based on what I've just told you about Sun's forthcoming CMT “Niagara” processor, and what other vendors have said about their processor plans for 2005. IBM and Intel both claim they will have a dual-core CMP design, with hyperthreading capability for each core. Other design points for 2005 are dual-core CMP designs without hyperthreading, and single-core designs with and without hyperthreading. That's pretty much the stated range of prospects. Ive listed some fairly generous average IPC figures for each of the wide-issue designs on this list. If we figure efficiency as the ratio between average and peak IPC, no wide-issue design is very efficient. If we calculate throughput performance by multiplying the number of cores per die, times the number of threads per core, times the average IPC per thread, the dual core/dual thread designs appear to top out at a throughput of about 8. On the same two metrics, “Niagara” scores a flat 100% efficiency rating, and offers throughput of 32. ‘That's why CMT is a radical processing advance. Note the conditions under which this comparison holds valid. The CMT design requires 32 parallel threads to reach its throughput number, but depends not at all on ILP. The other designs have a relatively modest dependence on TLP, since they can use no more than 4 threads at a time, but must attempt to compensate for their low TLP by applying more ILP to each thread. CMT - Keys to Success Simple, Scalable (RISC) 64-bit Processor Core | ~ Designs require a huge address space | — SPARC is an ideal candidate | Processor Design Expertise | ~ Sun has been designing processors since 1984 “Throughput Computing” Design Expertise ~ Sun commercialized SMP technology in the 1990s - CMT achieves at the CPU level what SMP does at the system level. Highty Threaded Operating System ~ Solaris is the industry leader in effective thread support | Highly Threaded Applications — Intrinsic to most networking computing applications — Not typical of most desktop applications ‘What's needed to succeed with a CMT-style processor design? Or, to put it another way, if CMT designs are so great, why isn’t everybody building one? The answer is that it takes a fairly unique combination of circumstances to be able to bring a successful CMT design to market, starting with a simple, scalable 64-bit processor core. SPARC is ideal here, while Itanium, for example, is a poor base (8 scalar SPARC cores will fit on a die that uses fewer transistors than the next iteration of the single-core Itanium 2). Second, it takes processor design expertise. Some system companies, bowing to the pressures of the commodity processor cycle, have abandoned CPU design on the grounds that it is no longer possible for them to add value in this area. Sun, by contrast, has continued to invest in its processor design capability. But simply having both an appropriate CPU design base and CPU design capability aren't enough. CMT designs translate to chip-scale what SMP accomplished at a system level. To build these kinds of chips requires an understanding of MP design issues at the system level. It takes a server company to build a “server on a chip.” Even the right processor core, processor design expertise, and a system-level understanding of SMP issues are not sufficient for CMT success. The whole point to these designs is to switch the performance focus from ILP to TLP, from time per unit task to tasks per unit time. To empower these designs, it takes an OS like Solaris, with a robust multi-threading model, and applications with a high degree of TLP, like those typical of network computing. All of which Sun uniquely has. Huge increase in performance of Performance threaded applications Quantum reduction eae in the cost of —Fewer servers Network Computing -Less floor space without ; Reduced power consumption eee ~Lower air conditioning Simplified administration and maintenance Major consolidation of components and ; interconnections Reli ‘So what benefits might one of Sun's customers expect to see from Sun's CMT initiative? In a nutshell, benefits in proportion to the size of the advance in computing technology. ‘An increase in single-processor performance anywhere from 4 to 8 times the best performance available on any competing hyperthreaded ILP-style processor in the same timeframe. A quantum-leap reduction in the cost of computing, based on the need to have far fewer servers to handle the same workload, housed in much smaller boxes, occupying much less valuable real estate, consuming far less expensive electricity, needing much less investment in cooling capacity, and offering substantial savings in the costs of maintaining and administering systems. As a bonus, the new systems should be far more reliable than any comparable SMP system built today, since collapsing the entire server down onto a chip also eliminates all the separate components and interconnections in today's SMP systems, together with their additive chances of failure. But from the customer's viewpoint, the best part about this whole prospect is that it does not disrupt their existing compute environment at all. Sun's new SPARC-based CMT processors will drop seamlessly into a new generation of SPARC/Solaris systems that continue to run the same applications in exactly the same way as before. Except far faster. In far less space. With far less ‘operating expense. More reliably than ever before. CMT - Sun Benefits + Provides competitive advantage and differentiation + Increases revenue and market share through superior value proposition * Enables new markets and services * Validates Sun's long-term SPARC/Solaris strategy What will this initiative do for Sun? Again, the benefits should be proportional to the size of the advance being made in computing technology. To date, Sun is the only company proclaiming CMT as its direction for future processors. Once other companies have had a chance to think about the benefits of this strategy, that differential is likely to go away, but we expect to maintain a leadership position with CMT processors, just as SPARC was a leader throughout the RISC era. As just explained, Sun's CMT systems should provide Sun customers with a superior value proposition, enabling Sun to win new customers and increase both its revenue and market share. Further, we expect CMT (like RISC) to enable entirely new markets and services, with a great expansion in the opportunities for “server-on-a-chip” technology. Suppose you could put an entire E10K on your desktop. Suppose you could carry one around in your pocket. Would that be interesting? What new things might you be able to do in that sort of a world? Last, but certainly not least, CMT validates the importance of continuing to invest in key technologies, even during a commodity shift, when it might seem that there is no way to add value through technology, and wise companies should focus instead on achieving some form of non-technical advantage, whether that be lower costs, better service, greater ease of use, or whatever. To the contrary, as Jong as technology has limits, there will be value to devising new technologies that allow the existing limits to be broken. Ue Osun =) Innovation Pays | * Throughput Computing fundamentally | changes the price/performance ratio of | Network Computing + Enables new markets and services + Sun is uniquely positioned to deliver | Throughput Computing (64-bit SPARC + Solaris + Threaded Apps) The fact is that processor technology is not yet a story whose end can be foretold, let alone announced. It still is the tale of a on-going journey, not a story about the computer industry reaching some final destination already visible on the horizon. Though the path of progress always tends toward more functionality and lower costs, the specific parameters of processor design are fluid and change from time to time. We already have witnessed three very distinct design points just since the start of the microprocessor era, from CISC to RISC to commodity, and a fourth shift to CMT is now looming straight ahead. Tn contrast to those who want to proclaim processor design a finished subject, closed due to the exhaustion of possibilities, as long as there is no obvious limit to the amount of progress that is possible towards higher levels of functionality and lower levels of cost, there is no obvious end- point to this process of change. Rather, quantum leaps forward will remain possible as long as new ways of thinking about old problems are a possibility. CMT is the next such leap. If this analysis is right, it will not be the last. It's important to pay attention to CMT, because it will usher in a new era of winners and losers, but don't expect the ending to be “... and they lived happily ever after.” The longing for a stable end point is the halimark of a fairy tale. The touchstone of reality is change. Well, that’s my story about the past, present, and future of processor design. TI be happy to take any questions you might have, either about Sun's new CMT initiative in processor design, or about any other points of interest regarding the changing parameters of processor design over the years. Obviously, to compress fifty years of design history into a 90 minute talk, Ive had to simplify the big points and omit many interesting details, but I've tried to do so without lapsing into either outright error or gross exaggeration. Although perhaps the operative words here are “outright” and “gross”, since some amount of distortion is probably inevitable. So I'd be glad to try to clarify anything said I've said, or talk further about anything that you might like to consider in a bit more depth. Thanks for your interest and attention. The floor is now open for questions and comments.

You might also like