Leak Suggests 'RTX 4090' Could Have 75% More Cores Than RTX 3090

Nvidia GeForce RTX 3080
(Image credit: Nvidia)

A leaker by the name of @davideneco25320 on Twitter has shared some very specific details about Nvidia's next-generation Ada (aka Lovelace) GPUs including SM counts and names of each new die. If his data is accurate (and given the recent Nvidia hack, it very well could be), Ada will be a massive upgrade over Ampere, the RTX 30-series, especially for the flagship GPU. As this is leaked data and cannot be completely trusted, take these results with a grain of salt.

Latest Videos FromTom's Hardware
Aaron Klotz
Contributing Writer

Aaron Klotz is a contributing writer for Tom’s Hardware, covering news related to computer hardware such as CPUs, and graphics cards.

  • hotaru.hino
    The RTX 3090 had over twice as many shaders as the RTX 2080 Ti, but certainly didn't perform twice as fast.
    Reply
  • ezst036
    All of the extra cores will be good news for miners.
    Reply
  • LolaGT
    A steal at $3000 U.S.

    I can see it now.
    Reply
  • USAFRet
    hotaru.hino said:
    The RTX 3090 had over twice as many shaders as the RTX 2080 Ti, but certainly didn't perform twice as fast.
    The same as in drive performance.
    2x benchmarks numbers does not equal 2x actual performance.
    Reply
  • blppt
    Hopefully we get a card that can actually handle raytracing this time. Third time's the charm?
    Reply
  • spongiemaster
    hotaru.hino said:
    The RTX 3090 had over twice as many shaders as the RTX 2080 Ti, but certainly didn't perform twice as fast.
    Nvidia changed the definition of a cuda core going from Turing to Ampere, so while it could have twice the performance in certain compute work loads, in games that required a combination of int and fp calculations, there weren't really twice as many execution units. Nvidia themselves stated that Ampere was a maximum of 1.7 times faster than Turing in rasterized graphics. Also memory bandwidth didn't come close to doubling between Turing and Ampere.
    Reply
  • hotaru.hino
    spongiemaster said:
    Nvidia changed the definition of a cuda core going from Turing to Ampere, so while it could have twice the performance in certain compute work loads, in games that required a combination of int and fp calculations, there weren't really twice as many execution units. Nvidia themselves stated that Ampere was a maximum of 1.7 times faster than Turing in rasterized graphics. Also memory bandwidth didn't come close to doubling between Turing and Ampere.
    Then let's look at some values. Or just one because I'm feeling lazy.

    Looking at page 13 (or 19 in the PDF) of the whitepaper on Turing, there's a graph of what games had a mix of INT and FP instructions. I'm just going to pick out Far Cry 5 from this. So if we take this graph, let's just assume for every 1 FP instruction, there were 0.4 INT instructions. And if we laid this on one of Turing's SMs, this means that if there are 64 FP instructions, there's only 25 INT instructions, or a utilization rate of ~70%. For Ampere, since there are some CUDA cores that are split between FP and INT, we can balance which one does what for better utilization. And doing some math, for the best utilization in Far Cry 5's example, 90 FP instructions plus 36 INT instructions, using 126 out of 128 of the CUDA cores vs 89.

    So right off the bat, without adding any more CUDA cores, Ampere should have about a 1.4x lead over Turing in this example. So how much does Ampere get in practice? 1.28x based on TechPowerUp's 4K benchmark.

    And sure Ampere doesn't have 2x the memory bandwidth, but it has almost double L1 cache, which should soak up the deficiency.

    Either way, my commentary is pointing out the odd conclusion the article seems to hint that NVIDIA only needs to add 75% more shaders for double the performance.
    Reply
  • sizzling
    blppt said:
    Hopefully we get a card that can actually handle raytracing this time. Third time's the charm?
    Not had any problems with Ray Tracing on my 3080.
    Reply
  • spongiemaster
    hotaru.hino said:
    Then let's look at some values. Or just one because I'm feeling lazy.

    Looking at page 13 (or 19 in the PDF) of the whitepaper on Turing, there's a graph of what games had a mix of INT and FP instructions. I'm just going to pick out Far Cry 5 from this. So if we take this graph, let's just assume for every 1 FP instruction, there were 0.4 INT instructions. And if we laid this on one of Turing's SMs, this means that if there are 64 FP instructions, there's only 25 INT instructions, or a utilization rate of ~70%. For Ampere, since there are some CUDA cores that are split between FP and INT, we can balance which one does what for better utilization. And doing some math, for the best utilization in Far Cry 5's example, 90 FP instructions plus 36 INT instructions, using 126 out of 128 of the CUDA cores vs 89.

    So right off the bat, without adding any more CUDA cores, Ampere should have about a 1.4x lead over Turing in this example. So how much does Ampere get in practice? 1.28x based on TechPowerUp's 4K benchmark.

    And sure Ampere doesn't have 2x the memory bandwidth, but it has almost double L1 cache, which should soak up the deficiency.

    Either way, my commentary is pointing out the odd conclusion the article seems to hint that NVIDIA only needs to add 75% more shaders for double the performance.
    Below is a comparison between 1/4 of a Turing SM and 1/4 of an Ampere SM


    As you can see, the number of FP32 CUDA cores doubled in Ampere, this is how Nvidia claims twice as many CUDA cores. However, you can also see not everything doubled. Whereas every CUDA core in Turing could concurrently perform 16x INT32 and 16x FP32 calculations, half of Ampere cores can either do 16x INT32 or 16FP32 instructions while the other half can only do 16x FP32 instructions. For purely fp32 workloads, you could see up to a theoretical doubling of performance, but for purely int32 workloads you wouldn't see any performance improvement at all. Because work loads are usually more floating point oriented than integer based (Nvidia claims 36 integer operations for every 100 floating point), Nvidia chose this layout that favors floating point performance. In the real world of mixed loads in games, you're going to see somewhere in between those 0 and 100% figures with Nvidia claiming a max of 70% improvement over Turing.
    Reply
  • digitalgriffin
    hotaru.hino said:
    The RTX 3090 had over twice as many shaders as the RTX 2080 Ti, but certainly didn't perform twice as fast.

    I agree. Even if you could develop a driver that could dish out the draw calls fast enough, there's a limit to the scheduler on the gpu being able to feed the cores. While more cores support more complex objects better, most objects aren't that complex. When you have 100 trees in the background you need them to be simple. So scenes are composed of many low to moderate complexity objects. This is also part of variable rate shading improvements. As a result most draw calls under utilize the GPUs full potential.

    This is part of the genius of unreal new engine. They reduce complexity of a scene so much in software by eliminating unseen details and reducing them. The math behind it is genius. Mesh reduction has never been a simple cs problem. They reduced it to a simple O(n*n) problem. Where n resents the number of layers of reduction.
    Reply