Nvidia Unveils Its Next-Generation 7nm Ampere A100 GPU for Data Centers, and It's Absolutely Massive

Nvidia A100
(Image credit: Nvidia)

The day has finally arrived for Nvidia to take the wraps off of its Ampere architecture. Sort of. While Ampere will inevitably make its way into some of the best graphics cards and find a place on our GPU hierarchy, today's digital GTC announcement is only about the Nvidia A100, a GPU designed primarily for the upcoming wave of exascale supercomputers and AI research. It's the descendant of Nvidia's existing line of Tesla V100 GPUs, and like Volta V100 we don't expect to see A100 silicon in any consumer GPUs. Well, maybe a Titan card—Titan A100?—but I don't even want to think about what such a card would cost, because the A100 is a behemoth of a chip.

(Image credit: Nvidia)
Latest Videos FromTom's Hardware
Jarred Walton
Senior Editor

Jarred Walton is a senior editor at Tom's Hardware focusing on everything GPU. He has been working as a tech journalist since 2004, writing for AnandTech, Maximum PC, and PC Gamer. From the first S3 Virge '3D decelerators' to today's GPUs, Jarred keeps up with all the latest graphics trends and is the one to ask about game performance.

  • tummybunny
    But can it run Crysis remastered?
    Reply
  • Jimbojan
    You may need to bring the Sun to power the system
    Reply
  • InvalidError
    One step closer to building the Jupiter Brain!
    Reply
  • bit_user
    @JarredWaltonGPU , thanks for the coverage, as always.

    ...but,
    For workloads that use TF32, the A100 can provide 312 TFLOPS of compute power in a single chip. That's up to 20X faster than the V100's 15.7 TFLOPS of FP32 performance, but that's not an entirely fair comparison since TF32 isn't quite the same as FP32.
    Right, a better comparison would be with the V100's fp16 tensor TFLOPS. Comparing it with general-purpose FP32 TFLOPS is apples vs. oranges, whereas TF32 vs fp16 tensor TFLOPS is maybe apples vs. pears.

    Nvidia says elsewhere that the third generation Tensor cores support FP64, which is where the 2.5X increase in performance comes from, but details are a scarce. Assuming the FP64 figure allows similar scaling as previous Nvidia GPUs, the A100 will have twice the FP32 throughput and four times the FP16 performance
    Um... don't assume that! Tensor cores are very specialized and support only a few, specific operations. You want to look at their spec for general-purpose fp64 TFLOPS.

    It seems likely the A100 will follow a similar path and leave RT cores for other Ampere GPUs.
    Agreed. Ray tracing is a different market segment. Even talking about photorealistic rendering for movies and such, it would make sense for them to take a page out of the Turing playbook and use the gaming GPUs for that.

    the A100 can be has multi-GPU instancing that allows it to be partitioned into seven separate instances.
    rats. You almost said "can has". I'd have enjoyed that.
    ; )
    Reply
  • bit_user
    InvalidError said:
    One step closer to building the Jupiter Brain!
    I think the bigger problems to solve are going to be power and cooling.

    On a related note, I'm reading it burns 400 W!
    Reply
  • JarredWaltonGPU
    bit_user said:
    @JarredWaltonGPU , thanks for the coverage, as always.

    ...but,

    Right, a better comparison would be with the V100's fp16 tensor TFLOPS. Comparing it with general-purpose FP32 TFLOPS is apples vs. oranges, whereas TF32 vs fp16 tensor TFLOPS is maybe apples vs. pears.


    Um... don't assume that! Tensor cores are very specialized and support only a few, specific operations. You want to look at their spec for general-purpose fp64 TFLOPS.


    Agreed. Ray tracing is a different market segment. Even talking about photorealistic rendering for movies and such, it would make sense for them to take a page out of the Turing playbook and use the gaming GPUs for that.


    rats. You almost said "can has". I'd have enjoyed that.
    ; )
    Nvidia has posted additional details, which I'm currently parsing and updating the article. There's a LOT to digest this morning!
    Reply
  • JarredWaltonGPU
    bit_user said:
    I think the bigger problems to solve are going to be power and cooling.

    On a related note, I'm reading it burns 400 W!
    Yes. Official specs were posted after the initial writeup, so I've fixed some things and added more details. RT cores and RTX are definitely not present.
    Reply
  • jeremyj_83
    @JarredWaltonGPU
    Isn't ECC a native feature on HBM2? If so then why would you need those extra stacks if not for redundancy?
    Reply
  • JarredWaltonGPU
    jeremyj_83 said:
    @JarredWaltonGPUIsn't ECC a native feature on HBM2? If so then why would you need those extra stacks if not for redundancy?
    I'm not positive on this, but ECC is an optional feature for HBM2 I think. I've found reports saying Samsung's chips have it... but is it always on? That's unclear. I have to think the sixth chip is simply for redundancy, which is still pretty nuts. Like, yields are so bad that we're going to stuff on six HBM2 stacks and hope the interfaces and everything else for at least five of them work properly? I asked for clarification on the six stacks -- TWICE -- but my question somehow got missed.
    Reply
  • jeremyj_83
    JarredWaltonGPU said:
    I'm not positive on this, but ECC is an optional feature for HBM2 I think. I've found reports saying Samsung's chips have it... but is it always on? That's unclear. I have to think the sixth chip is simply for redundancy, which is still pretty nuts. Like, yields are so bad that we're going to stuff on six HBM2 stacks and hope the interfaces and everything else for at least five of them work properly? I asked for clarification on the six stacks -- TWICE -- but my question somehow got missed.
    "And, of course, as a native feature of HBM2, it offers ECC memory protection."
    https://www.anandtech.com/show/15777/amd-reveals-radeon-pro-vii (bottom of the 6th paragraph)
    I just remembered where I had read the ECC reference to HMB2.

    Crazy idea but instead of redundancy could they use a stack as a large L3/4 cache of some sorts?
    Reply