A personal record of understanding, deciphering, speculating and predicting the development of modern microarchitecture designs.

Sunday, November 04, 2012

AMD's ARM Strategy?

I'm quite baffled by the recent news that AMD is going to design server chips with licensed ARM(v8) cores compatible with their "freedom fabric" interconnect.

What I don't understand is which type of product are they after with such a combination? The freedom fabric is a storage fabric, more like a (scalable) south-bridge equivalent. It seems AMD will integrate the storage fabric into the (supposedly many-core) ARM-based chips? Since AMD is not licensing the freedom fabric to anyone else, the only "customer" of such a chip will be SeaMicro? Since they will use licensed ARM cores, their chips likely won't compete with others (Qualcomm, Samsung, nVidia, Broadcom, ...) who develop customized cores on power efficiency and performance. Does AMD seriously believe their processors can be more competitive based on storage technology?

It is hard for me to believe that with all the possibilities in ARM based products, AMD could only think of licensing ARM cores to sell their freedom fabric? I thought ARM isn't even a CPU design company with the same caliber as AMD! As someone who believe AMD took the right design approach with their Bulldozer/Trinity (and similarly Bobcat) designs, I even doubt the licensed ARM core could compete with their own next-generation low-power x86-64 cores.

Anyway this is IMHO yet another weird plan from AMD over the past few years. Although Rory Read assured everyone during the analyst day that AMD's problem wasn't the wrong plans , I had the feeling that it's precisely the wrong plans, or at least lack of a right/focused plan. They planned to make Llano based APU, Bobcat based APU, Bulldozer and Trinity, all in a narrow 2-year window, during which they will also revamp graphics core and spend $370M to buy storage technologies. If they were trying to do everything of course one of the things they tried might fare better than others, right? Now they are trying to make ARM based processors, but just for scalable storage. I really don't understand--what would that accomplish which their x86-64 expertise will not? If ARM really is lower-power than x86 at the same performance (which I seriously doubt), then hasn't AMD been wasting time & money in designing Bobcat, but instead should've licensed ARMv7 cores 2 years ago?

Their problem was not having the wrong plan? Ba!!!

Saturday, August 13, 2011

HP DV6z Quad Ed

Got an HP DV6z with AMD A8-3500M (Llano) recently. I'm extremely happy with the 1920x1080 display ($150 upgrade). It is not glossy, but very colorful and bright. If you get an HP DV6z with Llano, BE SURE to spend that $150 for a high resolution display!

dv6z Quad Ed
• steel gray
• Genuine Windows 7 Home Premium 64-bit
• AMD Quad-Core A8-3500M Accelerated Processor (2.4GHz/1.5GHz, 4MB L2 Cache)
• AMD Radeon(TM) Discrete-Class Graphics [HDMI, VGA]
• FREE Upgrade to 6GB DDR3 System Memory (2 Dimm)
• 640GB 7200RPM Hard Drive with HP ProtectSmart Hard Drive Protection
• 9 Cell Lithium Ion Battery
• 15.6" diagonal Full HD HP Anti-glare LED Display (1920 x 1080)
• FREE Upgrade to Blu-ray player & SuperMulti DVD burner
• HP TrueVision HD Webcam with Integrated Digital Microphone and HP SimplePass Fingerprint Reader
• 802.11b/g/n WLAN and Bluetooth(R)

Battery life  (9-cell) is about 5hr heavy use with Linux VM compiling in background. With idle/light use battery life could be up to 7hr. Not terrific but quite good enough for the 15.6" + 1080p screen.

It's relatively heavy, but on the other hand, it is fully loaded. I'd say it'll make a great desktop replacement, easily beating most laptop at comparable price ($800~$1000).

I haven't really optimized the laptop yet but here are few photos:


Thursday, June 09, 2011

Experience in Acer Iconia W500 Windows 7 Tablet

-
I recently got one of those Iconia W500 tablet from Acer, running the fully version Windows 7 Home Premium.

It's really a nice build. I had some doubt before when I read some "bad" reviews about this tablet on the Internet. Then I went to a BestBuy to take a look myself at the Iconia A500 tablet, which I believe uses the same build but Intel's Atom processors. Note that Iconia W500, which I eventually bought, is based on AMD's Fusion C-50 processor, which is a lot faster and supports the new DX11 graphics and OpenCL. I realized that if A500 looks and feels nicely, then W500 can only be better, since they are using the same tablet frame & accessories. Apparently some reviewers have sensory problem.

Anyway, it is a nice little tablet with detachable full-size keyboard.

Some reviews made it sound like difficult or even impossible to "fold" the tablet and keyboard. In fact it cannot be more easy---just fold it. The tablet part will detach ("slide") itself out of the keyboard base. I could even do this when the tablet is on with no problem.

The dual-core C-50 is noticeably faster than the single-core 1.6GHz Turion + X300 graphics (the laptop being replaced by this tablet). The C-50 also runs much cooler. Real usage battery life is ~5.5 hours idle, ~4.5 hours with light use (+wifi), and just over 3 hours web-based gaming. Not as good as iPad, but comparable to other Zacate based laptops with larger battery.

I don't find it very useful to take the keyboard on the road. But it's again very easy, and the keyboard is lighter than the slim Dell keyboard I use on my desktop.

It has two disadvantages as a tablet:
 - Windows 7 is NOT a good tablet OS.
 - It can use more motion sensors and actuators that are found in iPad and Xoom.

But perhaps it's not fair to compare W500  with iPad or Xoom. The Acer Iconia W500 is more of a professional tablet, for people to make presentations, take notes, communicate with colleagues and all that. The iPad and Xoom, on the other hand, are more of consumer toys. Not to say you can't play on the Iconia W500 (you can certainly play many great games better than any other tablet there), or you can't work on iPad or Xoom, but the general feeling of Windows 7 and IOS or Android are just different.

Thursday, April 28, 2011

The Battle between Netbook and AMD Fusion

-
In the 1st quarter of 2011, Microsoft's Windows revenue dropped 4% from last year, mainly due to 40% decline in netbook sales. CFO Peter Klein says that tablets "played a part" in this decline. It's no wonder that companies no longer make a big fuzz about netbooks like they did in 2008 when Intel Atom was first out. After two short years of coming to life, netbook is already conceived as slow, insufficient and uncool by the consumers, ready to be replaced by the next technology.

AMD Fusion is what I feel an exciting and revolutionary technology which combines a capable multi-core CPU and a general-purpose GPU into a cost and power efficient package. We have an HP dm1z based on the AMD E-350, and I have to say it is an excellent notebook. It's capable, thin-and-light, runs cool and has long battery life (7+ hours in actual usage).

I find it ridiculous, though, that HP dm1z is marketed as a netbook. It's not a netbook. Neither HP nor AMD should have defined it as a netbook. My wife plays many cool games on it, edits photos on it, views HD movies on it, and basically performs every task an advanced PC user would on it, many tasks cannot be satisfactorily performed on a netbook (she holds an M.S. degree in computer science). It's no surprise, since netbook, by definition, is only for the "net," not games nor computing in general. An Atom based 10" laptop with in-order cores may be a netbook; an ARM based 8" laptop with 2GB RAM may be a netbook. But these AMD Fusion laptops have full-grown notebook computing capabilities (good), with size and power consumption similar to a netbook (better!).

It is no wonder that AMD does not refer to their Fusion APUs as for netbook. But sadly it doesn't matter. When Intel released their CULV, people tried to define the slow single-core Celeron based laptop with 2GB RAM as "notebook;" but after AMD launched their Fusion APUs, with two 64-bit out-of-order cores at 1.6GHz accessing 4GB RAM, most people seem to become dense about the many distinctions from netbook. Seriously, if the AMD Fusion "netbook" plays cool games with DX11, runs Office suite and photo editors smoothly, plays HD movies and even runs virtual machines at close-to-native speed, then it is a full notebook. If it's thin and light, then it's an ultramobile notebook. The reasonable conclusion: Fusion APU is not another netbook chip, but a perfect replacement for those ultramobile processors that would otherwise cost you $1000 each.

I feel kind of sad for AMD, for they seem to live in a world which is mostly agnostic about how good their technologies are. But perhaps that is how people like me can by these powerful little notebooks at a great price?

Friday, April 15, 2011

AMD tapes out 28nm Wichita, Intel shows new Atom and peeks 32nm shrink

-
It was said a few days ago that AMD taped out Wichita, their 28nm shrink on the APU roadmap.

Currently, there are two types of APUs on the market. Zacate @18W targets ultraportable notebook (such as HP dm1z and MSI X370) and small form factor desktop, while Ontario @9W targets tablet, (fanless) set-top and embedded boxes. According to AMD, these APUs are very small (< 0.8 cm^2) compared to ordinary processors. Shrinking them from 40nm to 28nm would bring the die size to below 0.5cm^2. It was also said that TSMC will apply HKMG on their 28nm process, which should offer some significant power reduction or performance boost.

I think there are two ways that AMD can bring the APUs forward. Wichita (with 1--2 Bocat core) could have similar performance to Ontario/Zacate with lower power consumption for the tablet market. Krishna (2--4 Bobcat core) could have the same 9W/18W TDP with more cores and even higher performance than Zacate for future ultraportable notebook market. These ultraportable notebooks should prove themselves with flexibility and performance clearly above the tablets. Netbook, on the other hand, is perhaps a market to be gradually replaced by tablet. It would be a mistake, in my opinion, for AMD to invest in and build up products for the netbook market now.

Currently, AMD's APUs offer much better performance than Intel's dual-core Atom. It's no surprise that with the tape out of AMD's next APU, Intel is eager to announce the new Atom and disclose its 32nm shrink at IDF in Beijing this year.

Wednesday, April 13, 2011

Supid is as stupid does

Some people at the Internet forums are overly concerned with fanboism arguments such as asking for non-evident expertise or superficial credential when presented with technical discussions. To me such behavior represents degeneration of human intelligence because they didn't even attempt to read, think, and debate with their brains before they make judgment. To them, anyone who spoke of things that are out of their comprehension must be insane, and they wouldn't even voluntarily improve themselves with information that are novel to them.

Stupid is as stupid does. If something is wrong,  then no matter what credential or expertise its performer claims, it is still wrong. When person X corrects the spelling of person Y, what X needs is a dictionary, not an argument that "because you're in primary school but I am the President." Bottom line: to say something is wrong, one needs to explain why; and he needs to understand the thing himself and make sure his understanding is valid first.

As I have discussed previously in my blog articles, it continues to amaze me how people can consciously believe marketing crap about the mythical IPC or bandwidth "numbers" of specific processors or benchmarks. Really, stupid is as stupid does. It doesn't matter where those numbers come from, when we have hard evidence or sound analysis proving them being either wrong or irrelevant.

Sunday, April 10, 2011

First look at AMD Family 15h (Bulldozer) Software Optimization Guide

NOTE: If you only want to know whether AMD would K.O. Intel or the other way around, or if you believed technical discussions are nonsense while Internet rumors are gold, then please stay away. OTOH, if you like computer architecture and feel excited about state-of-the-art designs, please enjoy and let me know what you think (thanks)!


Updates --
* 4/13/2011 Updated with discussion on load-store unit and memory disambiguation.
* 4/12/2011 Updated with highlights on shared frontend and changes to other memory resource.


Prelude

AMD recently released the software optimization guide for its upcoming & most anticipated family 15h (Bulldozer) processors. In this article we take a high-level comparative look at the newly released document.

The new processor family features a revolutionary "cluster multi-threading" (CMT) architecture, where a processor consists of multiple modules, each being a cluster of two cores sharing the same instruction frontend, floating-point unit and level-2 cache. Newly supported ISA extensions include the 128-bit SSE4 and 128 & 256-bit AVX, XOP and FMA4.

Despite these major differences, the Bulldozer is fundamentally a continuation of the previous processor design from AMD. It is perhaps more useful to first mention some similarities between Bulldozer and the previous family 10h (K10) processors before going into detail of the differences:
  • The same (or very similar) macro-op and micro-op based instruction decode is utilized.
  • Similar register file superforwarding.
  • Same L1I cache, very similar L3 cache and system interconnect architecture are used.
  • Similar pick-pack instruction decode in a 32-byte window.
  • Loads and stores seem still performed in the load-store unit working as a backend to the integer core and FPU, rather than being scheduled directly in reservation stations.
  • The shared FPU design in Bulldozer has its root deep in the separated integer and FPU schedulers in K10.
  • Same or very similar microarchitecture for indirect branch (512-entry target array) and return address (24-entry return address stack) prediction.
That said, below we discuss some (not all!) of the major microarchitecture differences introduced in Bulldozer: shared frontend, execution pipelines, L1D and L2/L3 caches, and memory access resources. 


Highlights on the shared frontend:
  • Two 32-byte instruction fetch windows (one for each core? 1.6.4)
  • Fetch window tracking structure (to manage fetches for both cores? 2.6)
  • Hybrid (tournament) branch prediction with global and local branch predictors
  • 2-level BTB with 512+5120 entries, upped from 1-level 2048 entries
  • Instructions decoded from a 32-byte window or two 16-byte windows (for both cores? 2.7)
  • Introduce branch fusion
Instruction fetch and branching are greatly improved in Bulldozer. A more sophisticated conditional branch prediction is employed, utilizing a local predictor, a global predictor and a tournament selector. The branch target buffer (BTB) is increased to 2.5+ times larger.

Note that although a single frontend serves two cores, the same branch prediction information can be shared by both cores if they execute the same program. Even if the two cores run different programs, sharing the same instruction fetch and branch prediction resources can have benefit in latency hiding, especially for non-optimized and densely branching codes.

When stars align (instruction allocation is optimized and the code has pre-decode information), the frontend can decode up to 4 macro-ops from a 32-byte window per cycle for one core. Otherwise, a 16-byte window is scanned to find the boundaries for supposedly < 4 decodes per cycle. It is unclear whether in such cases one 16-byte window can be scanned for each core, thus still maintaining 32-byte decode (for both cores) per cycle. Note that it takes at least 2x time to scan a instruction window twice as large, but two instruction windows of same size can always be scanned concurrently by parallel resources, if available.

The branch fusion seems similar to Intel's macro-op fusion. It has limited applicability but would make Bulldozer more competitive for running Intel-optimized codes.


Highlights on the execution pipelines:
  • 4-way microarchitecture design
  • Integer core has two EX and two AGLU pipelines, plus an LSU (2.10.2)
  • Floating-point unit (FPU) has two FMAC and two IMMX pipelines (2.11)
Up to 4 macro-ops per clock cycle can be issued from the (shared) frontend to either of the two cores. Within each core, up to 4 macro-ops per clock cycle can be sent to an integer or the floating-point scheduler.

The integer scheduler can dispatch up to 4 micro-ops per cycle, one to each of the 4 pipelines. Almost all ALU operations are handled by the 2 EX pipelines, except some LEA instructions which also utilize AGU. Thus the integer core can execute only up to 2 x86 instructions per clock cycle, resulting in a maximum integer IPC of 2.0 (in units of x86 instructions). Note however this estimate does not include the computing throughput of the integer SIMD pipelines in the FPU.

The FPU scheduler can dispatch up to four 128-bit operations with the following combinations: (1) any of {FMUL, FADD, FMAC, FCVT, IMAC}; and (2) any of {FMUL, FADD, FMAC, Shuffle, Permute}; and (3) any of {AVX, MMX, ISSE}; and (4) any of {AVX, MMX, ISSE, FSTORE}.

From a layman's viewpoint, the shared FPU seems to offer only half the throughput of two K10 cores for independent FMUL and FADD operations. However, in previous Opteron, vectorized loads and stores also share the FMUL and FADD pipelines; in Bulldozer, vectorized loads are either "free" or handled by the IMMX pipelines. Note that when FPU is throughput bottleneck, each arithmetic operation should be paired with on average one load or store. A perhaps more significant overhead saving comes from various vectorized register moves which can now be dispatched concurrently to separate IMMX pipelines. Thus the shared FPU in Bulldozer is actually a very balanced design.


Changes to L1 data cache: (2.5.2)
  • Size reduced from 64kB to 16kB
  • Associativity increased from 2-way to 4-way
  • Number of banks increased from 8 to 16 banks
  • Load-to-use latency increased from 3 to 4 cycles
  • Access policy changed from write-back to write-through
The L1D cache seems to go through an almost complete overhaul in Bulldozer. In previous AMD Opteron the L1D cache is virtually indexed and physically tagged; this allows the cache size to be greater than (page_size)*(associativity) without the homonym and synonym problems. On the other hand, this also means every cache hit must be subject to TLB hit.

In Bulldozer, the L1D cache size is (page_size)*(associativity) = 4kB * 4 = 16kB. As such, it is possible that the L1D cache is now virtually tagged which would put the DTLB access out of the critical loop. While this limits the maximum cache size to 16kB, it can offer clock rate and power advantage.

Limiting the cache size, however, does not solve the synonym problem where two cores in a Bulldozer module map different virtual address to the same physical address. Inconsistency can occur when the two cores update contents in their (virtually tagged) data cache separately. This problem, however, can be solved by writing through to the physically tagged shared L2D cache.


Changes to L2 and L3 caches:
  • L2 cache is now a "mostly inclusive" cache (2.5.3)
  • L2 cache latency increases to 18 ~ 20 cycles from previous 12 (=9+3) cycles
  • L3 cache is logically partitioned into sub-caches each up to 2MB (2.5.4)
The "mostly inclusive" property of the L2 cache in Bulldozer is a direct consequence of the write-through policy of the L1D cache. Any cache line that has been modified in an L1D cache will also have a copy in the L2 cache. On the other hand, when there is L1D/L2 cache miss and L3 cache hit, a cache line is copied from L3 cache directly to L1D cache (same behavior as in K10), making the L2 cache not fully inclusive. Similar behavior applies to the memory prefetch instructions which copy cache lines directly to L1D. On the other hand, "cold" data are probably loaded to both L1D and L2 caches to take advantage of the sharing of L2 by both cores (different from K10), which could explain the "mostly" inclusive description to the L2 cache.

The L2 cache latency in K10 is 9 cycles beyond the (3-cycle) L1 cache access, or a total 12 cycles. In Bulldozer, the L2 cache latency is increased to 18 ~ 20 cycles; the greater value is probably for writes, or for L1D TLB miss. The increased latency shows Bulldozer core designed more as thinner and faster (higher clock rate) than wider and shorter (higher ILP).


On load-store unit and memory disambiguation:
  • 40-entry load queue and 24-entry store queue in LSU
The load-store unit (LSU) seems to be very similar to the one in K10. Both utilizes two queues, one primarily for pending loads and one exclusively for pending stores. There have been claims that Bulldozer offers better out-of-order loads to stores than K10. From the high-level point of view of the LSU, the only "major" difference is perhaps the use of virtual address for tagging the L1D cache in Bulldozer(?), but physical address in K10. Tagging L1D with virtual addresses may allow pending stores to retire sooner to L1D without being subject to any TLB miss latency, thus resolving store-to-load dependency faster. Otherwise, according to Section 6.3 of the software optimization guides, the same restrictions on store-to-load forwarding apply to both Bulldozer and K10.

There has been many claims (mostly from people outside of AMD?) that Bulldozer must offer some "memory disambiguation" similar to Core 2 or Nehalem. From the organization of Bulldozer's integer and load-store pipelines, which resemble K10 more than Core 2, AMD would have to use very different memory disambiguation mechanisms than Intel. The concept of memory disambiguation is actually simple: a memory access can be ambiguous when its target address is unknown. Once the address is known, then disambiguation (within the same process) can be performed by simply comparing the addresses.

Suppose there is a store to an address A specified by a memory reference M. If M is not in cache, then the store can be pending for a long time waiting for A (at address M) to come. During that time, all later (independent) loads are ambiguous because any of their addresses could be the same as A (which is yet unknown). Similarly, there can be memory access ambiguity for stores following a load from A, or stores following a store to A.

One disambiguation that can be done is to predict which of the later memroy accesses are to addresses that overlap with A. All those that are predicted not overlapping proceed speculatively, and have their results (and all those they affected) squashed if later A is found to overlap with their access addresses. Note, however, that such disambiguation cannot be performed by the LSU if the LSU receives load-store requests with known addresses. It seems to be the case in both K10 and Bulldozer where the LSU works as a backend to the reservation stations.

Is it worth it to allow ambiguous memory access requests to be sent speculatively to Bulldozer's LSU? I think it requires detailed analysis and simulation to know for sure. The software optimization guide does not tell us whether such a design is used in Bulldozer. (Note that a more "severe" type of memory disambiguation may be needed for Intel Nehalem where two processes can share the same LSU, where different virtual memory mapping can create extra memory reference ambiguity.)


Changes to other memory resources (hardware prefetch and write combining):
  • Hardware pretech to both L1 and L2 (prefetch instruction still to L1 only, 6.5)
  • Stride L1 prefetcher with up to 12 pretech patterns
  • "Region" L2 prefetcher for up to 4096 streams or patterns
  • 4KB 4-way WCC plus a (single?) 64-byte 4-entry WCB (?) WCB (A.5)
Due to the much smaller size of L1D in Bulldozer, it is reasonable to expect hardware prefetch to be less aggressive at L1D. Instead, part of the "aggressiveness" is transferred to the large and shared L2 cache. Although less aggressive, the prefetch mechanism is much more sophisticated, keeping multiple (12) prefetch patterns active at the same time.

A special design in Bulldozer is the addition of a 4KB 4-way associative write coalescing cache (WCC) for aggregating write-back (WB) memory writes (before committing them to L2?). This special "write cache" is inclusive with the L2 cache, and has its contents universally visible. It is unclear whether there is one WCC per core or one per module, although the former seems more plausible.

One of the design goals of WCC is probably to improve inter-core data transfer. Previously in K10, if core1 needs to send something to core2, the cache line containing the data must be (a) modified in core1's L1D, (b) evicted from core1's L1D to its L2, then (c) transferred from core1's L2 to core2's L1D. In Bulldozer, since every write to L1D also writes through to the WCC, steps (b) can be omitted and step (c) can be performed together with updating the L2 cache. Even less overhead is incurred if the data transfer occurs between two cores in the same module that share the L2 cache.

The WCC also acts as a write buffer for the write combining buffer (WCB) for streaming loads and write combine memory type. This can have other implications on the memory ordering requirement by the AMD64 execution model, which we will not touch upon here.

Bulldozer seems to have less write-combining resource per core for streaming stores and write combining memory type than K10. Performance "caveat" was mentioned for streaming store instructions in Section 6.5 of the software optimization guide, where writing >1 streams of data with streaming stores results in much less performance compared with K10. It appears, although unclear, that Bulldozer has a (single?) 64-byte 4-entry (sharing the 64 bytes? each having 64 bytes?) write combining buffer (per core?). K10 and even the later K8 revisions have 4 independent 64-byte WCBs per core. One explanation is that modern processors have more cores and thus fewer occasions to store multiple independent data streams per core. With only one stream of streaming stores, the performance in Bulldozer is still comparable to that in K10.

On the other hand, by beefing up the write-combining resource for write-back & temporal stores with the WCC, common memory writes are made much more efficient. Make the common case fast -- a rule of thumb in microarchitecture design!

~~

Wednesday, March 23, 2011

Tuesday, February 08, 2011

Battery life of HP ProBook 6455b w/ AMD Phenom II N620

Common perception has been that AMD based notebooks are power hungry and rarely get over 3hr battery life. That perception is wrong.

Evidence? The HP ProBook 6455b with 2.8 GHz dual-core Phenom II N620 and 6-cell battery has up to 3hr 50min battery life under normal usage (Internet surfing with wifi and 70% screen brightness), as shown in the following screenshots taken roughly 50 minutes apart:

 ===> 
Before start writing this blog article After writing this blog article

While 3.8hr battery life isn't stellar, keep in mind that this is a 14.1 notebook with 35W TDP 2.8 GHz dual-core CPU and DirectX 11 capable GPU. I believe not many sub-$1000 Core 2 Duo or Core i3 laptops reach that battery life, and they usually have much lower-end GPUs.

(The notebook claims to have over 5hr battery life, which is the case when the CPU is idle at 1.0 GHz, the screen dimmed to the lowest and wifi turned off. I find such battery life number pointless, though.)

Friday, December 31, 2010

AMD Bobcat Fusion APU -- A Big Deal?

AMD has been enthusiastic and optimistic about its upcoming Fusion Accelerated Processing Unit (APU) based on the Bobcat cores set for launch at next year's (really less than one week from now) International Consumer Electronics Show (CES) in Las Vegas. It even makes a supposedly humorous video on YouTube, showing its main competitor spying on and astonished by AMD's "Fusion technology".

Is the Fusion APU really a big deal and, if so, in what sense? Will it really revolutionize personal computing as claimed by AMD?


The Facts

We already know the performance bound of these Fusion APUs, straight from AMD: compared to current CPU designs, the Bobcat core will achieve 90% performance with 50% die area. So a 1.6GHz Bobcat core will have performance comparable to a 1.4GHz Turion, definitely not a stellar specification. In fact, the APU's performance has been previewed and shown to be comparable to Intel's CULV CPU + nVidia's ION GPU.

The more impressive part is perhaps that the APU has both the CPU and GPU sitting on the same die, sharing the same system interface and 18W power envolope. Thus from the performance perspective, APU is much better than Intel's Atom processor (which powers most of the current low-cost netbooks), while from the power and cost perspective, APU is much better than Intel CULV + nVidia ION. So the whole point of these Fusion APU is really not about better performance (in both processing speed and power), but to reach a "better" power-performance tradeoff, i.e., performance-per-watt.


The Advantage

But is this power-performance tradeoff the real "advantage" of the Fusion APU, that it is unreachable by other players? I highly doubt it. For example, if one combines Intel's Yonah and nvidia's ION2 and manufactures them on Intel or TSMC 32nm, the same level of performance-per-watt could very well be reached.

However, even if Intel and nVidia work together, such a product probably won't make money for Intel due to all those redesign efforts required and the erosion to Intel's existing products. So IMHO one critical "advantage" that AMD has with APUs is that AMD's current market share in low-power laptops is so small that it doesn't worry about cannibalization by releasing cheaper products. Intel OTOH doesn't want to replace their existing laptops with lower performance cheaper ones. Instead they designed Atom to target on the smartphone and tablet markets. They make sure there's significant performance gap between Atom and Core i3 so the two markets are well separated.


The Extra

Hardware is only part of the story. By combining CPU and GPU closely together, every laptop based on AMD's Fusion APU becomes DirectCompute and OpenCL capable. Such "universal" GPGPU availability makes GPGPU acceleration a viable choice for software developers, which in turn makes these Fusion APUs better products (since more programs will be optimized for the CPU+GPU package). OpenCL came along somewhere in 2008 is an industry standard that replaced the original ATI Stream. Kernel programming in OpenCL is also very similar to that in nVidia CUDA, making OpenCL a fine choice for developers who are looking for or already taking advantage of GPGPU.


The "Better" Product?

However, even with GPGPU acceleration, a 18W APU still won't achieve stellar performance. Do you really believe the 18W TDP can translate to personal supercomputer, artificial intelligence and immersive 3D interface? What would be more interesting instead is the Fusion APU with the Bulldozer CPU core and the "Southern Island" GPU. That plus OpenCL could really be revolutionary in terms of software acceleration. But that plan, first disclosed by AMD in 2007, had been delayed until at least 2012/2013. Instead, Bobcat-based low-end Fusion APUs came to fill the void for the next 1 to 2 years.

While the current Fusion APU is not in AMD's original plan, with some irony it is probably a "better" product than originally planned. Why? Because believe it or not, most laptop users really don't need higher CPU performance! Most people will be quite happy with a dual-core 1.6GHz computer which they use mostly for e-mail and web surfing. The good graphics offered by these APUs is just a sweetening plus.

So what AMD does with the Bobcat-based APU is to depress CPU+GPU prices and power budgets so laptop makers can give us better other stuff, such as longer battery life, better webcam, and faster WiFi/3G/4G. And although this Fusion APU will reduce CPU+GPU ASPs and will hurt high-end laptop sales, AMD has little to lose in those areas anyway. :-)

Thursday, September 02, 2010

The IPC Myths

While Instruction Per Cycle (IPC) is an important metric for program optimization, it has been misused in many contexts. Below are a few common examples:
  • IPC can be used to describes how good a CPU is.
  • IPC is roughly proportional to pipeline width of the CPU.
  • IPC of modern CPUs are high (>>1).
  • Amdahl's law says CPU with higher IPC will have higher single-threaded performance.
  • ... 


Myth #1: IPC described as a single value

A common problem of all the "statements" above is that they all refer to IPC as if it is some intrinsic property determined by the CPU microarchitecture. In fact, IPC is a property determined not just by the CPU, but more by the program from algorithm down to instruction scheduling. For example, it is very possible for a CPU1 to have higher IPC than CPU2 running program A, but lower IPC running program B.

Thus, saying "CPU1 has higher (or lower) IPC than CPU2" has to be inaccurate, especially when the two processors have different microarchitectures.

Myth #2: Higher IPC means better

Many people believe higher IPC means higher (single-thread) performance. This is as wrong as when people thought higher clock rate means higher performance. Still, many believe higher IPC is better because the CPU can run as fast with slower clock rate. This seems an over-reaction to the Pentium 4, which had very high clock rate but moderate performance compared to Athlon64/Opteron.

The problem with this type of thinking is that the relation between IPC and clock rate is really a tradeoff. Like any tradeoff relation, you don't get optimal results by sliding towards either edge. With microarchitecture and circuit-level advancements, both clock rate and/or IPC can be increased. Which one to improve should depend on the design and application of the processor, and it's definitely not always (not even usually) IPC.

Myth #3: IPC is proportional to CPU pipeline width

We see many arguments like below on the Internet--
  • Core 2 can issue up to 4 x86 instructions per cycle, so it should have an IPC close to 4.
  • Nehalem brings [this or that features] to circumvent the decode limit, so it's IPC is 25% or 33% higher than Core 2.
  • K10 (AMD Family 10h) can only decode 3 x86 instructions per cycle, so its IPC has "bottleneck" at the instruction decode.
None of these statements is correct. It's not that the conclusion of these statements are absolutely false, but that their reasoning does not hold water. The best we can say about them is that without profiling or cycle-accurate simulation, we simply don't know.

In the case of Core 2 and Nehalem, we actually know for sure that the statements above are false. IPC of Core 2 Duo running SPEC CPU2006 was measured in this paper. The values were between 0.4 to 1.8 among all sub-benchmarks, with average only around 1.0, no where near its 4-way decoder width.

If we compare actual SPECint measurements of Core 2 (22.6) with Nehalem (25.1 or 27.8), we see that Nehalem has 11% to 23% higher single-thread performance after taking into account potentially 20% turbo frequency. Thus Nehalem's IPC for SPECint is at most ~20% higher than Core 2, and most likely much less when exclude the turbo mode effect. In other words, if Core 2's IPC for SPECint sub-benchmarks were 0.4~1.8, then Nehalem's should be between 0.5~2.1. Both are far below what is implied by their 4-way pipelines or any sexy-sound marketing features.

Myth #4: Amdahl's law favors CPU designed for higher IPC

This is the strangest argument that I have seen on the Internet, because it is completely the opposite of truth. The main thing that Amdahl's law says is that performance improvement is intrinsically limited by the available parallelism in a program. In the context of single-threaded programs, this means that performance at the same clock rate is limited by the Instruction-Level Parallelism (ILP) available in the program.

Some people see that "limited by the ILP" part and immediately relate it to a CPU designed for higher IPC. The problem here is that, according to Amdahl's law, the ILP is limited by the program, not the CPU. In other words, if your program has low ILP, it will not run fast no matter how high an IPC the CPU was designed for. Thus in fact Amdahl's law favors a CPU designed for higher clock rate but lower IPC than the available ILP in the program.

Furthermore, the available ILP in a program is also a strong function of the window size and the branch prediction accuracy. Both are very difficult to increase in the uber-complex microarchitectures of modern CPUs. That is why features such as SIMD (SSE and AVX), SMT, and turbo frequency are used in Nehalem to improve single-thread processor performance. None of them increases IPC of the CPU.

Conclusion

IPC is very useful when one wants to optimize his program for a particular system. It is one of the most important metrics that profiling produces. But like any metric, generalizing its implication outside of its intended usage context is usually meaningless and even misleading.

Wednesday, May 19, 2010

GPGPU and its battle of nVidia vs ATI

GPGPU seems to be really taking off. I came across a new YouTube video showing IBM new mainstream server using nVidia Tesla graphics cards for compute intensive acceleration.

For the past few years, AMD/ATI enthusiasts have assumed that Radeon is AMD's crown jewels. The truth might be just the opposite.

At a workshop in a recent conference, an nvidia researcher compared CUDA and OpenCL. His argument (whether true or not) was simple: OpenCL is more a device driver level language. He "proved" it by showing the same program written in CUDA and in OpenCL side-by-side. The CUDA one took about 2 slides. The OpenCL about 10. If you are a researcher/programmer/engineer, which one will you use?

It may be true that ATI Evergreen gives higher performance per dollar for games, but nVidia Fermi seems to give better GPGPU performance on average. Evergreen has more parallelism and higher theoretical flops, but Fermi is easier to program and to get real speedup. Hundreds universities worldwide are teaching students how to optimize their programs for Fermi. This is a formidable rival and I don't share at all the optimism of many ATI enthusiasts.

I believe In a year or two we will see the market of GPGPU surpassing that of enthusiast graphics. Very few people in the world care about fps when playing games. On the other hand, everyone benefits from GPGPU. I feel that AMD/ATI is too conservative in pushing for GPGPU. Most of their laptop/desktop chipsets still use r700 or even r600 based IGP, which are very hard to get good GPGPU performance, if any at all. Every time I see a laptop with HD42xx IGP I feel disgusted. They are selling those 2-year-old stuff which doesn't let users take proper advantage of OpenCL. They sell them for cheap, but is it a good thing? Do they also want to sell r800 IGPs for cheap 2 years from now?

Then it's OpenCL which everyone's heard of but few is interested. In my humble opinion, AMD should hire a few programmers to fully integrate OpenCL into their Catalyst driver, so that every computer with an ATI GPU can use them after a simple driver install. In contrast, perhaps for fear of Microsoft or whatever reason, AMD/ATI want end users to manually install and upgrade every OpenCL release to match the installed Catalyst driver. If I'd go so much trouble then why don't I just install CUDA which more people are using anyway?

And if I'm going with CUDA, then I will not only buy Tesla for my workstation, but also GeForce for my desktop & laptop; there will be less incentive for me to buy an AMD CPU as well. That is really a good way to keep me away from being AMD's customer, isn't it?

Tuesday, April 27, 2010

Stating the facts or bad-mouthing his former employer?

Over at the MacRumors forum someone claimed to be a former AMD employee recently started to criticize AMD and people working in it. His posts can be seen in the following links: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 14, 15 (thanks to dm7000s at AMDZone for collecting these links).

In my opinion, it is both "interesting" and "fishy" to see someone do this to his former employer. On one hand, he may be revealing some real problems inside (parts of) the company which other people (either inside or outside AMD) wouldn't know or recognize. On the other hand, he may be right about things that he claims to know, but due to his bitterness wrong about the conclusions.

I believe in this case, it is the latter. In any rate, lets go through some of his points below and see, assuming these are all facts, how true or false they can be:

AMD has not been financially successful since "K8". This may not be due to any of AMD's problem, but Intel's monopoly tactics. One should ask why was AMD "financially successful" during the K8 days in the first place? Was it because only those who designed K8 knew that they were doing? Was it because they hand crafted every transistor? Or was it really because both Itanium and Netburst terribly sucked in real-world tests? I'd argue it's only the last.

AMD has been losing key employees. Losing employees is tough for any company. Yet, sometimes a company has to lose weight when it is evolving and before it can start growing again. The question is not whether someone did something grand. But whether he will do something grander. What would have been the grander next step after K8? Could AMD have beaten Intel by making an over-complicated "K9" with SMT and turbo mode and everything else? I'd argue with the required design and verification efforts, no "key employee" could have made this happen timely and cost effectively.

AMD is not hand-instantiating designs transistors anymore. Anyone (who is a electrical engineer) can hand craft transistors. It is at the end of the day primarily a labor intensive task. If you're a CTO and you expect your company to hand craft transistors better than your 10x oversized competitor, then you're not being realistic. You won't win. And with the "unfair" agreements between Intel and AMD prior to their 2009 settlement, AMD was simply forbidden to reap the same amount of profit by selling hand-crafted processors of higher performance.  I was informed by a kind reader, who seem to know what was going on inside AMD, that hand-crafting transistors is emphatically what AMD did not do for K8. Unlike Intel which does a lot of custom designs, AMD used standard cells for the ALUs and most other components. However, AMD did a lot of circuit placement and routing by hand, with a superb physical design (implementation) team. Somewhat related to the previous comment, I was also told that many in that team have left AMD over the past few years.

AMD did not make any new architecture after K8. Architecture is a flimsy thing. At their hearts K8 is no more than K7 plus extra 64-bit registers and integrated NB. The way instructions are broken down to macro-ops and micro-ops, the basic organization of the ROB and the separate INT and FP schedulers are all the same between K7 and K8. HyperTransport based NB, the exclusive L1/L2 cache and improve TLB gave K8 solid performance. But so are the improvements made to K10 like the shared L3, unganged memory, probe filter and greater scalability. K8's NB was designed to have up to 8P in a single system; few wanted that over the years. Today K10-based Magny Cours processors allow 48 cores with perhaps tighter inter-core communication. I bet Intel very much want to do the same.



In conclusion.... is that guy simply stating the facts, or is he bad-mouthing his former employer? Personally, I think what he did was immature and immoral, even if what he said were facts. I was told that there's been some political struggles inside AMD during the post-K8 years, and I also believe that such politics must've brought with it some waste of time and money as well as loss of talents. But still.... in my humble opinion, that's no good excuse for picking on your former exployee and starting a public brawl fight.

:mrgreen:

Thursday, April 24, 2008

Internet Misinformation - On K10 vs Core2 Bandwidths

The Internet is filled with all kinds of information. For the most part this is a very good thing- who don't like new things and thoughts? However, if not being careful enough such free information could easily become misinformation, especially regarding technical things where either the author doesn't really know what he is writing, or he actually knows but deliberately misleads his readers for marketing or financial reasons.

In recently years with the widespread of consumer-grade "benchmarks" we see many what I'd call "folklore comparison" of different PC platforms, most recently AMD's Barcelona (K10) and Intel's Xeon (Core2) microarchitectures. Specifically, many of these folklore comparisons made by on-line reviews give Intel-favoring misinformation to "justify" Core2's "theoretical ILP advantage" on paper. In this article I will look more closely at both sides of this argument: Is it or is it not justified to attribute Core2's better performance in some cases to the supposed "advantage" in its microarchitecture, or does such advantage actually exist?

Misinformation on L1 Data Cache Bandwidth

The first example starts with a "test" AnandTech performed regarding K10 vs Core2 L1 data cache bandwidth:

Lavalys Everest L1 Bandwidth

Read (MB/s)Write (MB/s)Copy (MB/s)Bytes/cycle (Read)Latency (ns)
Opteron 2350 2 GHz32117160822393516.061.5
Xeon 5160 3.047860477469547515.951
Xeon E5345 2.3337226371347426815.961.3
Opteron 2224 SE51127256014408015.980.9
Opteron 8218HE 2.6 GHz41541208013581515.981.1


From the values above it would appear that, per clock cycle, K10 can load as much data (16 bytes) but store only half (8 bytes) as Core2 can. As pointed out by scientia at AMDZone, these numbers do not seem correct. In fact, according to this presentation slide, AMD's Barcelona (K10) processors could theoretically perform two 128-bit (16-byte) loads per clock cycle, or twice the Core2's L1 data cache bandwidth. Why the contradiction?

Rebuttal and Explanation

What is shown here is a perfect example of misinformation coming out of such "folklore comparison" performed by AnandTech, using synthetic benchmark tools without really knowing what it is doing. A synthetic benchmark would underestimate Opteron's L1 bandwidth because, with sequential accesses and small strides, it stresses only one of the two ports of K10's L1 cache.

Recall that K10's L1 data cache (L1D) is 2-way set associative with two 128-bit ports. Internally each port is connected to one bus going to the Load-Store Unit (LSU). This arrangement is described both verbally (2nd paragraph, page 223, A.5.2) and graphically (Figure 11, page 230, A.13) in the Software Optimization Guide for AMD Family 10 Processors:


Optimally in every clock cycle two 128-bit words, one from each port, can be read from L1D to LSU and forwarded to the execution units. What happens with synthetic benchmarks is probably that, due to their fine-grain assembly-level "optimization," they could generate unrealistic codes favoring one microarchitecture instead of another. On K10, it is in fact possible to force data accesses to the same cache way or the same cache bank. Such accesses will only utilize one of the two available ports and result in half the optimal bandwidth.

In practice, do we always find data accesses to the same port? Of course not. Clearly the type of tests that AnandTech did reflect little if any realistic processor performance. They are at best echoes to uneducated folklore opinions, or worse practices of darn FUDs. On the other hand, we also won't find all data accesses to different ports, thus a theoretical calculation of K10's maximum L1D bandwidth (that it is twice as high as Core2's) is also overly optimistic and unrealistic. In reality, if 50% of cache accesses are spread to two ports, then in average it would take 3 cycles to read 4 words (one cycle to read 2 words, two cycles to read 1 word each). The average bandwidth would be 4/3 = 1.33 reads or writes per clock cycle. In other words, K10 would achieve about 67% (1.33/2) of its theoretical max L1D bandwidth, which would be 33% higher than Core 2 for reads and 33% lower for writes.

Proof and Conclusion

To prove that this theory is true, I wrote a program in C with gcc's SSE intrinsics to test the bandwidths myself. Skipping other I/O & maintenance parts, the kernel of the code looks like this:


The above code is compiled by gcc 3.4.4 with -O2 and -march=k8. The generated assembly is also included below for assurance that the code really does what it's supposed to do. (Note: When trying it on gcc 4.2 the compiler is smart enough to know that the loop doesn't actually do any useful work, and would not generate any SSE load instruction. In this case the code above achieves 1 iteration (skipping 8 loads) per cycle, limited by the conditional branch bubble.)


I have the program with three tests: SSE store, SSE load, and SSE PAND; only SSE PAND is shown above but the other two are very similar. When running on a Phenom 9500 @2.2GHz, the achieved bandwidths for store, load, and PAND are 28GB/s, 46GB/s, 46GB/s, respectively. This translates to about 1.3x 16B reads/cycle and 0.8x 16B writes/cycle. So without any special treatment, on the C-source level, I could already get 30% better L1D read bandwidth and 60% better L1D write bandwidths than AnandTech's "test" results; furthermore, the theoretical estimates that I offered in the previous section were actually fairly close.

Out of curiosity I took the same program to run on a Core 2-based 2.0GHz Xeon machine. It turns out that it only achieves ~85% max bandwidth there, i.e., less than 14B reads and writes per clock cycle. Thus per clock cycle, the L1D read bandwidth on Core 2 is only 2/3rd of that on K10, whereas the write is just 8% higher. Frankly I have no idea how Lavalys Everest does its benchmarking codes to generate vastly different results, but really anyone with a clear mind shouldn't care about how some synthetic binary runs, but what he can achieve on the platform on the source level, without dirtying his hands by assembly optimizations (or de-optimization for the non-Intel cases).

Saturday, September 22, 2007

AMD's latest x86 extension: SSE5 - Part 2

Series Index -

In this part we will compare Intel's SSSE3 and SSE4.x with AMD's SSE5. More specifically we will look at how one can (or cannot) use SSE5 to accomplish the same tasks performed by SSSE3 and SSE4.x. The pinnacle question we're trying to answer here is whether the SSE5 from AMD is strictly an extension to Intel's SSE4, or in some sense a replacement for SSSE3 and SSE4.x (which none of AMD's current processors - including Barcelona and Phenom - supports)?

Syntactical Similarity

The original 8086/8087 have one-byte opcode instructions (if we ignore the ModRM bits used for 8087 and a handful others such as bit rotations). One remaining opcode byte that was usefully unused turned out to be 0Fh; had it been used, it would've had the meaning of POP CS , which was not there because it would create some interesting program flow control problems. Using 0Fh as an escape byte followed by a second byte, a number of two-byte opcode instructions were added by 80{2|3|4}86, Pentium, MMX, 3DNow!, and SSE/2/3/4a.

After the addition of SSE4a from AMD, the free two-byte opcodes left are only the followings: 0F0{4,A,C}h, 0F2{4-7}h, 0F3{6-F}h, 0F7{A,B}h, and 0FA{6,7}h. Why is this important? Because these points in the two-byte opcode space are the only entries where the x86 ISA can be further extended. Obviously, the two-dozen or so entries are not enough for any large-scale extension.

In order to further extend the instruction set in a significant way, the opcode itself must be extended from two-byte to three-byte. This is where SSSE3/SSE4.x and SSE5 bear the most similarity: they all consist (mainly) of instructions with three opcode bytes. Intel carved out 0F38xxh and 0F3Axxh for SSSE3 and SSE4.x, whereas AMD took 0F24xxh, 0F25xxh, 0F7Axxh and 0F7Bxxh for SSE5.

Syntactical Differences

However, the syntactical similarity between Intel's and AMD's extensions pretty much ends right here. As we've seen in Part 1. of this series, SSE5 instruction encoding is regular and orthogonal: the 3rd opcode byte (Opcode3) always has 5 bits for opcode extension, 1 bit for operand ordering, and 2 bits for operand size.

On the other hand, the encoding of SSSE3 and SSE4.x instructions may well have been arbitrary for anyone outside Intel. For example, look at the following SSSE3 instructions:

PSIGNB - 0F380h 1000b ... PABSB - 0F381h 1100b
PSIGNW - 0F380h 1001b ... PABSW - 0F381h 1101b
PSIGND - 0F380h 1010b ... PABSD - 0F381h 1110b

It may seem from above that the right-most bits encode the operand size - 00b for byte, 01b for word, and 10b for dword. However, take anther look at the following SSSE3 instructions:

PSHUFB - 0F380h 0000b ... PMADDUBSW - 0F380h 0100b
PHADDW - 0F380h 0001b ...... PHSUBW - 0F380h 0101b
PHADDD - 0F380h 0010b ...... PHSUBD - 0F380h 0110b

For some (probably legitimate) reason, Intel designers decided not to include horizontal byte additions and subtractions; instead they (most "exceptionally") squeezed in a byte-shuffle instruction and a specialized multiply-add instructions. We see that 30-years later, people at Intel still design instructions exactly the same way like 30-years ago: doesn't make sense.

Even worse cases are seen in SSE4.x. The following example shows the encodings used for packed MAX and packed MIN instructions:

PMAXSB - 0F383h 1100b ... PMINSB - 0F383h 1000b
PMAXSD - 0F383h 1101b ... PMINSD - 0F383h 1001b
PMAXUW - 0F383h 1110b ... PMINUW - 0F383h 1010b
PMAXUD - 0F383h 1111b ... PMINUD - 0F383h 1011b

Note how the different operand types and operand sizes are squeezed cozily into consecutive opcode byte values without much sense. For some mystical reason, the unsigned word operations are put quite arbitrarily right next to the signed dword operations . But wait... what happens to P{MAX|MIN}SW and P{MAX|MIN}UB? Well, they already are SSE2 instructions with opcode 0FE{E|A}h and 0FD{E|A}h, respectively. As can be seen in this example, the irregularity of SSE4.x also inherits from the poor design of SSE2.

From software programmer's point of view, the irregularity really doesn't matter as long as the compiler can generate these opcodes automatically. But such extension irregularity is no circuit designer's love to implement. This is probably why Intel, assumed not incompetent, chose in such poor styles to design SSEx - to make it as difficult as possible for anyone else (most prominently AMD) to offer compatible decoding. In the end, not only Intel's competitors but also its customers suffer from the bad choices: had Intel designed the original SSE/SSE2 the same way as AMD does SSE5, we would've had a much more complete & efficient set of x86 SIMD instructions that makes sense! (Now, does Intel promote open & fair competition that benefits the consumers? Or does it aims nothing but to screw up its competitors, sometimes together with its customers?)

In any rate, as we've been above the encoding of SSE5 is different from SSSE3/SSE4.x and thus the former does not exclude the latter. In other words, it is possible for a processor to offer both SSE5 and SSSE3/SSE4.x (much like 3DNow! and MMX). What about their functionalities, then? Below we'll look at each SSSE3 and SSE4.x instruction and see how its functionalities can or cannot be accomplished by SSE5.

Functional Comparison to SSSE3

For SSSE3 instructions:
  • PHADDx/PHSUBx
    • Horizontally add/subtract word & dword in both source and destination sub-operands and pack them into destination.
    • Each PHADDx/PHSUBx in SSE5 operates on only one 128-bit packed source.
  • PMADDx
    • Multiply destination and source sub-operands, horizontally add the results, and store them back to destination.
    • PMADx in SSE5 offers more powerful multiply-add intrinsics
    • No byte-to-word multiply-add in SSE5, though.
  • PSHUFB
    • Shuffle bytes in destination according to source.
    • Special & weaker cases of the first-half of PPERM in SSE5.
  • PALIGNR
    • Shift concatenated destination & source bytes back into destination.
    • Special & weaker cases of the first-half of PPERM in SSE5.
  • PSIGNx
    • Retain, negate, or set zero sub-operands in destination if corresponding sub-operands in source is positive, negative, or zero, respectively.
    • No direct implementation in SSE5.
  • PABSx
    • Store the unsigned absolute values of source sub-operands into destination sub-operands.
    • No direct implementation in SSE5.
  • PMULHRSW
    • Multiply 16-bit sub-operands of destination and source and store the rounded high-order 16-bit results back to destination.
    • No direct implementation in SSE5.

It can be seen that most SSSE3 instructions are not directly implemented in SSE5, with possibly the exceptions of PSHUFB, PALIGNR, and PADDx/PSUBx. However, these latter SSSE3 instructions can still be useful as lower-latency, lower-instruction count shortcuts to the more generic & powerful SSE5 counterparts. Thus from this point of view, future AMD processors will probably still benefit from implementing SSSE3 together with SSE5.

Functional Comparison to SSE4.x

For SSE4.1 instructions:
  • PMULLD
    • Multiply 32-bit sub-operands of destination and source and store the low-order 32-bit results back to destination.
    • Can be done by two PMULDQ (SSE2) followed by a PPERM.
  • DPPS/DPPD
    • Horizontally dot-product single/double precision floating-point sub-operands in destination and source and selectively store results to destination sub-operand fields.
    • FMADx in SSE5 offer more powerful & flexible floating-point dot product intrinsics.
  • MOVNTDQA
    • Non-temporal dword load from WC memory type into an internal buffer of processor, without storing to the cache hierarchy.
    • Specific to Intel processor implementation.
    • PREFETCHNTA in Opteron & later works for the same purpose.
  • BLENDx and PBLENDx
    • Conditionally copy sub-operands from source into destination.
    • Special and weaker cases of PERMPx and PPERM in SSE5.
  • PMAXx and PMINx
    • Packed max and min operations of destination and source
    • Can be accomplished by a PCOMx followed by a PPERM in SSE5.
  • EXTRACTPS/PEXTRx
    • Extract sub-operands from an XMM register (source) to memory or a general-purpose register (destination).
    • Special and weaker case of PERMPx for memory destination.
    • No direct implementation for GPR destination in SSE5.
  • INSERTPS/PINSRx
    • Optionally copy sub-operands from source to destination.
    • Special and weaker case of PERMPx in SSE5.
  • PMOVx
    • Sign- or zero-extend source sub-operands to destination.
    • Special and weaker case of PPERM with a proper mux/logical argument.
  • PCMPEQQ
    • Packed compare-equal between destination and source and store results back to destination.
    • Special and weaker case of PCOMQ in SSE5.
  • MPSADBW
    • Compute "sum of absolute byte-difference" between one 4-byte group in source and eight 4-byte groups in destination and store the eight results back to destination
    • No direct implementation in SSE5.
  • PHMINPOSUW
    • Find the minimum word horizontally in source and put its value in DEST[15:0] and its index in DEST[18:16]
    • No direct implementation in SSE5.
  • PACKUSDW
    • Convert signed dword to unsigned word with saturation.
    • Complements PACKSSDW/PACKUSWB/PACKSSWB in SSE2
    • No direct implementation in SSE5.
  • PTEST, ROUNDx
    • Llogical zero test, packed precision rounding.
    • Copied directly to SSE5.

For SSE4.2 instructions:
  • PCMPGTQ
    • Packed compare for greater than
    • Special & weaker case of PCOMQ in SSE5.
  • String match, CRC32
    • No direct implementation in SSE5.
  • POPCNT
    • Copied directly from AMD's POPCNT.

A few evidences from above show that it's probably not very likely for a future AMD processor to implement SSE4.1 & SSE4.2 in addition to SSE5. First, some of the instructions are copied directly from SSE4.1 to SSE5 (TEST and ROUNDx); had AMD wanted to implement SSE4.1 before SSE5, it would've been unnecessary to copy these instructions. Second, those instructions in SSE4.x that do not have superior SSE5 counterparts are either extremely specialized (MPSADBW, PHMINPOSUW, string match & CRC32), or able to be accomplished more flexibly by two or less SSE5 instructions.

We can also see how Intel designers work very hard to squeeze functionalities into the poor syntax of SSE4.x, resulting in a poor extension design. One example is the BLENDx/PBLENDx instructions. Instead of using the proper SSE5-like 3-way syntax, the variable selector in SSE4.1 is set implicitly to XMM0, not only requiring additional register shuffling but also limiting the number of permutation types to only 1 at any moment.

Another example is the DPPS/DPPD instructions, where the dot-product is performed partially vertical and partially horizontal. To make these instructions useful the two source vectors must be arranged to alternate positions: (A0, B0), (A1, B1), (A2, B2), ... Not only such arrangement can be costly by itself, but also after the operation one of the arranged source vectors is destroyed (replaced by the dot-product result).

Concluding Part 2.

Comparing SSE5 with SSSE3/SSE4, it seems that after years of being dragged along by Intel's poor extension designs, AMD finally decides to make its own next step in a better way. As I've discussed above, it's probably more advantageous for AMD to implement SSSE3 together with SSE5, and less so to implement SSE4.1 & SSE4.2.

However, as we know the commercial software in general and benchmarks in particular, especially on the desktop enthusiast market, are heavily influenced by the bigger company, thus if it turns out SSE4.x are excessively used to benchmark processor performance then it is still possible for AMD to implement them in its future processors. But lets hope for all customers' sake this is not going to happen, and future x86 extension will follow more of AMD's SSE5 than Intel's SSE4.x.

Friday, September 21, 2007

AMD's latest x86 extension: SSE5 - Part 1

Series Index -

The SSE5 announcement made by AMD earlier this month is something big. In fact, in terms of instruction scope and architectural design, it is bigger SSE3, SSSE3, and SSE4 combined. If we think of AMD64 as completely revamping x86-based general-purpose computing (as generally conceived by the industry), then we can also think of SSE5 as completely revamping x86-based SIMD acceleration. In my opinion, the leaps made by AMD in both AMD64 and SSE5 firmly assert the company as the leader in x86 computing architectures, leaving Intel gasping far behind.

The SSE5 Superiority

There are a few things that make SSE5 a "superior" kind of SIMD (Single-Instruction Multiple-Data) instructions different from all the previous SSE{1-4}:
  • SSE5 is a generic SIMD extension that aims to accelerate not just multimedia but also HPC and security applications.
    • In contrast, previous SSEx, especially SSE3 and later, were designed specifically with media processing in mind.
    • The CRC and string match instructions of SSE4.2 are too specialized to be generally useful.
  • SSE5 instructions can operate on up to three distinct memory/register operands.
    • It allows true 3-operand operations, where the destination operand is different from any of the two source operands.
    • It allows 3-way 4-operand operations, where the destination operand is the same as one of the three source operands.
  • SSE5 includes powerful and generic Vector Conditional Moves (both integer and floating-point).
    • Only four instructions (mnemonics) are added: PCMOV for generic bits, PPERM for integer bytes/(d,q)words, PERMPD/PERMPS for single/double-precision floating points.
    • Powerful enough to move data from any part of the 128-bit source memory/register to any part of the 128-bit destination register, plus optional logical post-operations.
  • SSE5 includes both integer arithmetic & logic, and floating-point arithmetic & compare instructions.
    • For integer arithmetics, it includes both true vertical Multiply-Accumulate and flexible horizontal Adds/Subs.

An Analytical View of SSE5 Instruction Format

All above show one thing: SSE5 is a well-planned, thoroughly articulated, and carefully designed ISA extension. The amazing thing is that the designers at AMD accomplish all these by simply adding a single DREX byte in-between the SIB and Displacement bytes, as shown in the figure below (taken from page 2 of AMD's SSE5 documentation):
A question naturally arises: will the additional DREX byte further increase instruction lengths? Fortunately, not a single bit. According to the official document linked above, those SSE5 instructions that use the DREX byte can not only take 3 distinctive operands but also access all 16 XMM registers without the AMD64 REX prefix; in fact, the use of the DREX byte in an SSE5 instruction excludes the use of the REX prefix. SSE5 instruction lengths are just as long as needed and as short as it can be. (We will talk more about possible further extensions to AMD64 REX and SSE5 DREX in a later part.)

Another great merit of SSE5 instruction encoding is that it is simple and regular. Note the "Opcode3" byte in the above picture, the main byte that distinguishes among different SSE5 instructions: its encoding is astonishingly simple: 5 bits for opcode, 1 bit for operand ordering, and 2 bits for operand size. The result is an orthogonal instruction encoding - you only need to look at an opcode field by itself to know what it means. In contrast, the 3rd opcodes of Intel's SSSE3 and SSE4 instructions seem like picked by spoiled child to purposely screw up any implementation. (We will talk more about comparison between AMD's SSE5 and Intel's SSSE3/SSE4 in a later part.)

Types of SSE5 Instructions

There are several major types of instructions in SSE5:
  1. Various integer and floating-point multiply-accumulate (MAC) instructions.
  2. Vector conditional move (CMOV) and permutation (PERM) instructions.
  3. Vector compare and predicate generation instructions.
  4. Packed integer horizontal add and subtract.
  5. Vectorized rounding, precision control, and 16-bit FP conversion.
A single PTEST instructions in Type 3 and four ROUNDx instructions in Type 5 above are copied directly from Intel's SSE4.1; together with other Type 4 and Type 5 instructions these are the SSE5 instructions that do not contain the DREX byte. All the other Type 1-3 SSE5 instructions utilize the DREX byte to specify a 3rd distinctive (destination) operand and to offer access to XMM8-XMM16 registers (without & excluding the REX prefix).

In particular, the Type 1 (MAC) and Type 2 (CMOV/PERM) instructions are 3-way 4-operand operations, with destination is set to either source 1 or source 3. The fact that 3-way operation is allowed - even with destination equal to one of the sources - is instrumental in enabling flexible MAC and CMOV/PERM instructions. In the case of MAC, two multipliers and an accumulator must be specified; in the case of CMOV/PERM, two sources and a conditional predicate must be given. Without the ability to address 3 distinctive operands, these two types of accelerations are either impossible or done awkwardly (more on Intel's SSE4.1-way of doing it in a later part of this series).

What makes these two types of instructions, MAC and CMOV/PERM, which happily require 3 distinctive operands, so special? As previously said, the four conditional move & permutation instructions allow predicated transfer of data from any part of the source registers/memory to any part of the destination register, followed by one of seven optional operations. Just how many instructions are there in SSE/SSE2/SSE3 to perform similar and simpler tasks partially? Here is a quick list:
  • MOVAPD
  • MOVAPS
  • MOVDDUP
  • MOVSHDUP
  • MOVSLDUP
  • MOVDQA
  • MOVDQU
  • MOVHLPS
  • MOVLHPS
  • MOVQ
  • MOVSD
Of course this does not mean the four instructions in SSE5 will replace all the MOVs in SSE/SSE2 above, which are still useful for their simplicity (only 2 operands required) and possibly lower latency (no post-operation needed). However, it does illustrate how powerful and useful the PERM instructions in SSE5 can be - just imagine how hard it is to implement these operations in an SSE2-like style.

The MAC instructions turns out to be one of the "most-wanted" instruction accelerations. As shown in "Design issue in division and other floating point operations" by Oberman et al. in IEEE ToC, 1997, nearly 50% of floating-point multiplication results are consumed by a depending addition or subtraction. See the picture below, directly grabbed from the paper:
In other words, by combining multiplication with a depending addition/subtraction, we can eliminate 50% instructions following all multiplications. Until SSE5, it was impossible to truly fuse a multiplication with a depending add or subtract and take advantage of such acceleration.

Concluding Part 1.

As shown above, the SSE5 from AMD is indeed something very different from the previous x86 SIMD extensions from Intel. Some people even went so far to call it "AMD64-2", and the "top development" of the year; such enthusiasm, of course, is unduly.

Until now, AMD is still gathering community feedback and asking for community support on the SSE5 initiative. Apparently, SSE5 is still in development; it's a great proposal, but clearly not developed (yet). Also, the SSE5 instructions by themselves do not match the breadth and depth of AMD64, which not only expands x86 addressing space but also semantically changes the working of the ISA. SSE5, on the other hand, doesn't touch nor alter any bit of the x86-64 outside of its extending scope. However, as we will discuss in a later part, the direction pointed to by SSE5 can be used to further extend x86-64 in a more general and generic way rivaling the original AMD64.

Monday, September 10, 2007

Scalability counts!

As I have said in this article, Intel's new Core 2 line of processors have good cores but poor system architecture. The poor scalability of FSB means that Core 2, without extensive, expensive, and power-hungry chipset support, is only suitable for low-end personal enjoyment.

Take a look at this AnandTech benchmark. I'd note foremost that AnandTech is hardly an AMD-favoring on-line "journal"; thus we can expect its report to be at worst Intel-biased and at best neutral (which I'm hoping for here). In any rate, the benchmark picture is reproduced below:

The comparison between Barcelona (Opteron 2350 2.0GHz) and Clovertown (Xeon E5345 2.33GHz) couldn't be clearer: FSB is an outdated system architecture for today's high-end computing, and scalability does matter for server & workstation grade performance. While AMD's quad-core Opteron at 2.0GHz is slower than Intel's quad-core Xeon at 2.3GHz on single-socket test, the situation is reversed when going to a dual-socket setup, one that used by most workstations and entry-level servers.

The same phenomenon is also observed in this page where AMD's quad-core Opteron, at 17% slower clock rate, performs increasingly better than Intel's quad-core Xeon with more number of cores (picture reproduced below). Again, when it comes to server & workstation performance, scalability counts.

Friday, August 03, 2007

Not Everything about Memory is Bandwidth

The False Common Belief

There is this common belief among PC enthusiasts that bandwidth, or million transfers per second or megabytes per second, is the most important thing that a good memory system should aim for. Such a belief is so deep-rooted that even the professionals (i.e., AMD & Intel) began to calibrate & market their products based on the memory bandwidth values.

For example, take a look at this Barcelona architecture July update article. The first graph in that page, which seems to be an AMD presentation and is conveniently duplicated below, seems to suggest that all the memory enhancements in AMD's Barcelona (K10) over its predecessor (K8) are about "Increasing Memory Bandwidth".


The question is, do they really increase memory bandwidth? Lets take a look at the bullet points in the graph, from bottom to top.
  • The prefetchers. Prefetching does not increase memory bandwidth. On the contrary, it reduces available memory bandwidth by increasing memory bus utilization (search "increase in bus utilization" on the page).
  • Optimized Paging and Write Bursting. They both increase memory bus efficiency, which does not increase the bandwidth per se, although it helps improving the bandwidth effectiveness.
  • Larger Memory Buffer. A larger buffer can improve store-to-load forwarding and increase the size of write bursting. The buffer itself, however, does not increase memory bandwidth at all.
  • Independent Memory Channels. This certainly has no effect on memory bandwidth. Each of the two independent channels is half the width, resulting in the same overall bandwidth.
Thus, out of six bullet points, only two are marginally related to memory bandwidth. The bottom line: Barcelona still uses the same memory technology (DDR2) and the same memory bus width (128-bit), beyond which there is no more bandwidth to increase to!

However, one would be more wrong to think Barcelona's memory subsystem is not improved over its predecessor, because all the points above are nevertheless improvements, though not on increasing memory bandwidth, but on reducing memory latency. Intelligent memory prefetching can hide memory latency, as shown in the Intel article page linked above. Reduced read/write transitions due to write bursting and the larger memory buffer both can reduce memory latency considerably. The independent memory channels also reduces latency when multiple memory transactions are on-flight simultaneously - especially important for multi-core processing. In short, the memory subsystem of Barcelona is improved for lower latency, not higher bandwidth.

Why Does Barcelona Improve More Latency Than Bandwidth?

There are a few reasons that a general computing platform based on multiple levels of cache benefits more from lower memory latency. This is contrary to specialized signal processing or graphics processors where instruction branches (changes in instruction flow) and data dependencies (store-to-load forwarding) are few and rare. This fact is aptly described in the following "Pitfall" on page 501 of Computer Architecture A Quantitative Approach 3rd ed., Section 5.16, by Hennessy and Patterson:
  • Pitfall Emphasizing memory bandwidth in DRAMs versus memory latency. PCs do most memory access through a two-level cache hierarchy, so it is unclear how much benefit is gained from high bandwidth without also improving memory latency.
In other words, for general-purpose processors such as Athlon, Core 2 Duo, Opteron, and Xeon, what helps performance is not just the bandwidth, but more importantly the effective latency of their memory subsystem. This pitfall is promptly followed by its dual on the next page of the book, which on the other hand explains why most signal and graphics processors which require high memory bandwidth do not need multiple levels of cache like the general-purpose CPUs:
  • Pitfall Delivering high memory bandwidth in a cache-based system. Caches help with average cache memory latency but may not deliver high memory bandwidth to an application that needs it.
Memory Bandwidth Estimate for High-End Quad-Core CPUs

Still one may ask, is the memory bandwidth offered by say a DDR2-800 channel really enough for modern processors? It turns out that, at least for Intel's Penryn and AMD's Barcelona to come, it should be. To estimate the maximally required memory bandwidth, we assume a 3.33GHz quad-core processor with 4MB cache sustaining 3 IPC (instruction per cycle). Such a processor should be close to the top performing models from both AMD and Intel by the middle of next year. (See also the micro/macro-fusion article for Core 2's actual/sustainable IPC.)

First lets look at the data bandwidth. A 3.33GHz, 3 IPC processor would execute up to 10G I/s (giga-instructions per second). Suppose 1 out of 3 instructions has a load or store, which is supported by the fact both Core 2 and Barcelona have 6-issue (micro-op) engines and perform up to 2 loads or stores per cycle. Thus,

10G I/s * 0.333 LS/I = 3.33G LS/s (giga-load/store per second, per core)

Multiply this number by 4 cores, the total is 13.33G LS/s. According to Figure 5.10 of Computer Architecture AQA on page 416, a 4MB cache has miss rate about 1%. Lets make it 2% to be conservative. Thus the number of memory accesses going to the memory bus is

13.33G LS/s * 2% MA/LS = 0.267G MA/s (giga-memory accesses per second)

Each memory access is at most 16-byte, but mostly likely 8-byte or less in average. This makes the worst-case memory bandwidth requirement 0.267G*16 = 4.27GB/s, and the average-case 2.14GB/s. Note that a single channel of DDR2 memory can support up to 6.4GB/s, much more than the numbers above.

Now lets calculate the instruction bandwidth. Again, for 4 cores at 3.33GHz 3 IPC, there are 40G I/s. However, instructions usually have exceptional cache advantage. According to Figure 5.8 on page 406 of the same above textbook, a 64KB instruction cache has less than 1 miss per 1000 instructions. Assume (again quite conservatively) each instruction takes 5 bytes, this means the total memory bandwidth for fetching instructions is

40G I/s * 0.1% * 5 B/I = 0.2GB/s

Thus even under conservative estimation, the instruction-fetch bandwidth is negligible compared to the data load/store bandwidth. The conclusion is clear: the memory bandwidth of just one single DDR2-800 channel (6.4GB/s) is more than enough for the highest-end quad-core processor during the next 10 months to come. The problem, however, is not bandwidth, but latency.

Update 8/7/07 - Please take a look at the AMD presentation on Barcelona, page 8, where quad-core Barcelona is shown to utilize just 25% the total bandwidth of 10.6GB/s. Suppose this is obtained from a 2.0GHz K10 (one that was demo'd and apparently benchmarked by AMD), then, scaling up linearly, a fictional 3.3GHz K10 would reach about 41% utilization, or about 4.3GB/s. Notice how close this number is to my estimate above.

What About Core 2's Insatiable Appetite for FSB Speed?

A naturally raised question is that, if 6.4GB/s is more than enough for the highest-performing quad-core x86-64 processors in the next year or so, why is Intel raising the FSB (front-side-bus) speed above 1066MT/s (million-transfers per second) so aggressively to 1333MT/s and even 1600MT/s? Isn't 1066MT/s already offering more than 6.4GB/s bandwidth?

The reasons are two-fold:
  1. For Core 2 Quad, the FSB is not just used for memory accesses, but also I/O and inter-core communications. Since data transfer on FSB is most efficient with long trains of back-to-back bytes, such transfer-type transitions can greatly reduce effective bandwidth.
  2. Raising the FSB speed not only increases peak bandwidth, but also (more importantly) reduces transfer delay. A 400MHz (1600MT/s) bus will cut 1/3rd the data transfer time of a 266MHz (1066MT/s) bus.
In other words, due to the obsolete design of Intel's front-side-bus, the sheer value of peak memory bandwidth becomes insufficient to predict the memory subsystem's performance, where a potentially 10.6GB/s bus (1333MT/s * 8B/T) isn't even able to satisfy the need of a quad-core processor (3.33GHz, 3 IPC) requiring no more than 5 GB/s of continuous data access to/from the memory.

The Importance of Latency Reduction

To show how latency reduction is the more important reason to raise FSB speed, we will compare a dual-core system with two quad-core systems, one with a 2x wider memory bus, the other with a 1.5x faster memory bus. We will show that the faster FSB is more effective in bringing down the average memory access time, which is the major factor affecting a computer's IPC. More specifically, using the dual-core system as reference, suppose the following:
  1. The quad-core #1 system has the same bus speed but 2x the bus width (e.g., 128-bit vs. 64-bit). In other words, it has the same data transfer delay and 100% more peak memory bandwidth than the dual-core system.
  2. The quad-core #2 system has the same bus width but 1.5x the bus speed (e.g., 400MHz 1600MT/s vs. 266MHz 1066MT/s). In other words, it has 33% less data transfer delay and, consequently, 50% more peak memory bandwidth than the dual-core.
  3. The memory bus is time-slotted and serves the cores in round-robin. For the dual-core and quad-core #1 systems, each memory access slot is 60ns. For the quad-core #2 system, each slot is 40ns.
  4. Memory bandwidth utilization is 50% on the dual-core (2 out of 4 slots are occupied) and the quad-core #1 (4 out of 8 slots are occupied). It is 66.7% on quad-core #2 (4 out of 6 slots are occupied), calculated by 50% * (2x cores) / (1.5x bandwidth).
Note that the assumptions above are simplistic and optimistic. It does not take into account the reduced effective bandwidth & efficiency due to I/O and inter-core communications. When taken these two effects into account, the quad-core systems will perform much worse than they are estimated below.

Lets first calculate the average memory access latency of the dual-core system. When either core makes a memory request, it finds 3/4=75% of chance the memory bus is free, and 25% of chance it has to wait an additional 60ns for access. The effective latency is

60ns * 75% + (60ns+60ns) * 25% = 75ns

Thus in average, each memory access takes just 75ns to complete.

Now lets calculate the effective latency for the quad-core #1 system. When an arbitrary core makes a memory request, it finds only 5/8=62.5% of chance the memory bus is free, and 37.5% of chance it has to wait. The waiting time, however, is more complicated in this case, because there are C(8|3) = 56 cases how the slots are occupied. Skipping some mathematical derivations, the result is

60ns * (4+3+2+1*5)/8 = 105ns, 6 out of 56 cases
60ns * (3+2+2+1*5)/8 = 90ns, 30 out of 56 cases
60ns * (2+2+2+1*5)/8 = 82.5ns, 20 out of 56 cases

=> (105ns * 6/56) + (90ns * 30/56) + (82.5ns * 20/56) = 88.9ns

Thus even when we double the memory bandwidth, keep the same bus utilization, a quad-core system still induces 18.5% higher access latency than a dual-core system. Note that this is even in the case where memory utilization is as low as 50%. For higher utilization, the latency increase will only be worse. The conclusion is clear: increasing memory bandwidth is not enough to scale up memory performance for multi-core general-purpose processing.

Now lets look at the quad-core #2 system, where data transfer delay is reduced 33%, but memory width is the same and bus utilization is increased to 66.7%. When an arbitrary core makes a memory request, it finds just 3/6=50% of chance the memory bus is free, and 50% of chance it has to wait. The waiting time again is complicated as there are C(6|3) = 20 cases how the slots are occupied. Skipping again some mathematical derivations, we get

40ns * (4+3+2+1*3)/6 = 80ns, 4 out of 20 cases
40ns * (3+2+2+1*3)/6 = 66.7ns, 12 out of 20 cases
40ns * (2+2+2+1*3)/6 = 60ns, 4 out of 20 cases

=> (80ns * 4/20) + (66.7ns * 12/20) + (60ns * 4/20) = 68ns

The average memory access latency here is almost 10% lower than the dual-core case and 24% lower than the quad-core #1. The effect of higher memory bus utilization is completely offset by a lower data transfer delay. Again, for general-purpose multi-core processing, reducing memory access delay is much more important than increasing memory peak bandwidth.

Conclusion and Remark

Lets go back to the original (supposedly) AMD's presentation. Why does it say "increase memory bandwidth" all over the page? Probably because most people simply don't understand better, and to make them so an article like this one is probably necessary and not even sufficient. We seem to see AMD engineers trying so hard to twist the delicate bandwidth-latency relationship, push it and force it down to a form easily understood (yet probably not believed) by ordinary minds.

However, bandwidth is definitely not useless. It really depends on the workload. For streaming processing such as graphics and signal processing, bandwidth and throughput are everything, and latency becomes mostly irrelevant. You won't care whether a DVD frame is played to you 100 milliseconds after it was read out of a blu-ray disc, as long as the next frame comes within 15 milliseconds (70fps) or so. Yet 100 milliseconds is 300 million cycles of a 3GHz processor! For streaming applications, we certainly want continuous flow of high-bandwidth data, yet have millions of cycles of latency to spare.
Please Note: Anonymous comments will be read and respected only when they are on-topic and polite. Thanks.