AMD's CPU-to-GPU Infinity Fabric Detailed

(Image credit: AMD)

AMD is currently the only vendor with both x86 processors and discrete graphics cards under one roof, at least until Intel's Xe graphics roll out, giving Team Red some flexibility with its interconnect technology. This tech has been particularly useful in the world of high-performance computing (HPC), as evidenced by an AMD presentation at the Rice Oil and Gas HPC conference yesterday. 

AMD initially announced at its Next Horizon event in 2018 that it would extend its Infinity Fabric between the data center MI60 Radeon Instinct GPUs to enable a 100 Gbps link between GPUs, much like Nvidia's NVLink. But with its Frontier supercomputer announcement in May, AMD divulged that it would expand the approach to enable memory coherency between CPUs and GPUs.

Paul Alcorn
Editor-in-Chief

Paul Alcorn is the Editor-in-Chief for Tom's Hardware US. He also writes news and reviews on CPUs, storage, and enterprise hardware.

  • TechLurker
    So would this basically be a gestalt "mega APU" of sorts?
    Reply
  • JamesSneed
    I'm not sure how many people see this coming but desktop APU's in 3-4 years will be like this as well. When AMD is on TSMC's 3nm you have enough density and power savings to have a 8+ core CPU and the power of what was once a dedicated GPU all in one APU. If you have that you might as well have unified HBM memory on the APU as well.
    Reply
  • JayNor
    Does the AMD's Infinity Link heterogeneous solution support the asymmetric cache coherency feature of CXL that appears to be responsible for its rapid adoption?

    That CXL feature conceptually should remove interaction with caches of connected processors/gpus/fpgas/nnps using biased coherency bypass.

    https://www.nextplatform.com/2019/09/18/eating-the-interconnect-alphabet-soup-with-intels-cxl/

    Codeplay is porting Intel's dpc++ to run on NVDA's processors.

    https://codeplay.com/portal/02-03-20-codeplay-contribution-to-dpcpp-brings-sycl-support-for-nvidia-gpus
    NVDA and AMD have both joined the CXL Consortium, and Papermaster recently made positive comments about it.

    https://www.anandtech.com/show/15268/an-interview-with-amds-cto-mark-papermaster-theres-more-room-at-the-top
    Did AMD present strong advantages for using Infinity Fabric vs PCIE5/CXL for their heterogeneous GPU interconnect?
    Reply
  • Gomez Addams
    From the article : "...but Nvidia hasn't made any announcements about such wins, despite its dominating position for GPU-accelerated compute in the HPC and data center space. "

    Oh? Doesn't the Perlmutter system qualify? It has 112 V100s in it.
    Reply
  • Gomez Addams
    JamesSneed said:
    I'm not sure how many people see this coming but desktop APU's in 3-4 years will be like this as well. When AMD is on TSMC's 3nm you have enough density and power savings to have a 8+ core CPU and the power of what was once a dedicated GPU all in one APU. If you have that you might as well have unified HBM memory on the APU as well.

    Imagine a package like the 3970 or 3990 with half of those CPU chiplets being GPU chiplets. Then throw a few chiplets with GPU-CPU-unified HBM4 memory in there. That could be an amazing machine.
    Reply
  • digitalgriffin
    I looked at HSA years ago, and looking at AMD's scalable architecture, new memory types and made a very good guess where it was going. AMD was a big proponent of heterogeneous architectures working together. I read the leaves said this about 3 years back.

    About two years ago I said, "You're going to see an APU with a chiplette for CPU, a chiplette forGPU, an IO Die, and HBM package that is part of unified memory dedicated to graphics calls."

    My only miss was I predicted Zen 2 (Ryzen 3000) would come out of the gate this way. I was correct with chiplettes, but missed the GPU chiplette on package. But I'm betting we'll see something like this soon.

    While HSA is being less emphasized, one of the problems I had resolving (when I drew up block diagrams) was indeed cache coherency between chiplettes. I knew infinity fabric was the answer. But I didn't have the exact answer in terms of algorithms to keep from overloading it. That one took me a while to figure out.

    It's a similar problem dealing with GPU to GPU chiplettes. Everyone thought I was nuts. They likended it to SLI/Crossfire. But it isn't SLI/Crossfire with a unified memory architecture. It's just a matter of resolving tiles and sharing the data differences between them. Even NVIDIA posted a paper about how it wasn't practical. But they are all looking a lot closer at it now.
    Reply
  • TheEldest
    Gomez Addams said:
    From the article : "...but Nvidia hasn't made any announcements about such wins, despite its dominating position for GPU-accelerated compute in the HPC and data center space. "

    Oh? Doesn't the Perlmutter system qualify? It has 112 V100s in it.

    112 V100s is about 14 Petaflops or 0.014 Exaflops. The wins AMD has are for super computers in the 1-10 Exaflop range (70x - 700x more powerful).

    So I'd say, No, the perlmutter system doesn't qualify.
    Reply
  • Gomez Addams
    I believe there is one exaflop-class supercomputer in progress now - the El Capitan. The current fastest one, Summit, is powered by GV100s and gets 200 PF for double precision calculations and 3EFs for AI/TensorFlow calculations.
    Reply
  • Olle P
    JamesSneed said:
    ... desktop APU's in 3-4 years will be like this as well. ...
    You think it will take that long? I was expecting it to happen this year, with the Ryzen 4000-series (but it doesn't).
    Reply
  • d0x360
    Gomez Addams said:
    From the article : "...but Nvidia hasn't made any announcements about such wins, despite its dominating position for GPU-accelerated compute in the HPC and data center space. "

    Oh? Doesn't the Perlmutter system qualify? It has 112 V100s in it.

    No it doesn't even compare. AMD has selling many thousands of pieces of hardware for the world's fastest super computer. Having 112 v100's is nothing by comparison.

    AMD is also moving into data centers with both CPU's and GPU's and since nVidia doesn't make a CPU they can't compete in that market. Epyc is a better choice than anything Intel is offering. The price difference is huge and epyc has better performance. So much so that VMware is changing how they charge based on core count...oddly one of the tiers ends with the exact core count of Xeon...weird
    Reply