New AMD processors quadruple memory bandwidth

Updated EddieC 0 Tallied Votes 349 Views Share

While most reporters and bloggers today are trumpeting the fine showing by chip-maker AMD in the semi-annual list of the world's , for me the real news is what AMD has planned for next year. In the first quarter of 2010, AMD is scheduled to begin shipping a processor that will include not only 12 cores, but also the vastly expanded memory bandwidth to support them.

Code-named Magny-Cours, the processors will offer four channels for DDR3 memory, twice that of its 6-core Opteron top-performer, which by the way powers what is today considered . AMD further boasts this week that five of the top 10 systems on the November list have AMD inside. DDR3 doubles the transfer rate of DDR2, effectively quadrupling bandwidth when available over four channels instead of two.

"A processor with four memory channels and 12 cores will be a very appealing option for film rendering and virtualization of high performance computing," said John Fruehe, director of business development at AMD's Server and Workstation Division, in a confidential group phone briefing Friday. It will be a natural, he said, for the types of things being done in the film industry, where Intel has a stronghold.

But the news isn't all upside; all of AMD's new processors, starting with the Magny-Cours, will require all new system hardware. However, AMD says that the socket infrastructure will support its 16-core processors coming in 2011 and beyond.

Dani AI

Generated

As pointed out, the Magny‑Cours parts (AMD Opteron 6100 family) were a deliberate architecture shift: AMD used a multi‑chip module (two dies per package) and exposed four independent DDR3 channels per G34 socket so the larger core counts would have more sustained memory bandwidth. This required new G34/C32 motherboards and platform silicon rather than drop‑in upgrades to older Opteron systems. (arstechnica.com)

That “quadrupling” sound bite needs a practical caveat. Four channels plus faster DDR3 raises theoretical throughput a lot, but real gains depend on DIMM speed, ranks and how the board populations are arranged; AMD and reviewers reported roughly 42.7 GB/s per G34 socket with DDR3‑1333 in typical configs, not a literal 4x in every workload. Expect large, memory‑bound HPC and render jobs to see the biggest wins; lightly threaded code will benefit much less. (theregister.com)

Virtualization implications were material. Because each Magny‑Cours package contains two dies with their own memory controllers, each die appears as a separate NUMA node to the OS/hypervisor; VMware ESX/vSphere treats those as distinct NUMA homes, so a VM that spans nodes (a “wide” VM) may suffer increased remote memory latency. Best practices with these systems are to size big VMs to fit a single NUMA node when possible, enable vNUMA for very large guests, avoid hot‑add CPU options that disable vNUMA, and validate core/memory layout against the hypervisor’s NUMA presentation. (frankdenneman.nl)

For procurement and tuning: buy G34 platform boards from vendors that published DIMM population rules and BIOS updates (several vendors shipped ready designs at launch), follow those population rules, and test target workloads under the hypervisor you intend to run. AMD also positioned the G34 platform to accept future Bulldozer/Interlagos 16‑core parts, so the socket had roadmap continuity for higher core counts. (channelpronetwork.com)

Be a part of the DaniWeb community

We're a friendly, industry-focused community of developers, IT pros, digital marketers, and technology enthusiasts meeting, networking, learning, and sharing knowledge.