DeepSeek's AI breakthrough bypasses industry-standard CUDA for some functions, uses Nvidia's assembly-like PTX programming instead

Nvidia Hopper H100 GPU and DGX systems
(Image credit: Nvidia)

DeepSeek made quite a splash in the AI industry by training its Mixture-of-Experts (MoE) language model with 671 billion parameters using a cluster featuring 2,048 Nvidia H800 GPUs in about two months, showing 10X higher efficiency than AI industry leaders like Meta. The breakthrough was achieved by implementing tons of fine-grained optimizations and usage of Nvidia's assembly-like PTX (Parallel Thread Execution) programming instead of Nvidia's CUDA for some functions, according to an analysis from Mirae Asset Securities Korea cited by @Jukanlosreve.

What You Need to Know

The DeepSeek logo against a hexagonal textured background

(Image credit: DeepSeek)

Today: OpenAI boss Sam Altman calls DeepSeek 'impressive.' In 2023 he called competing nearly impossible.

Jan. 28, 2025: Investors panic: Nvidia stock loses $589B in value.

Dec. 27, 2024: DeepSeek is unveiled to the world.

Latest Videos FromTom's Hardware
Anton Shilov
Contributing Writer

Anton Shilov is a contributing writer at Tom’s Hardware. Over the past couple of decades, he has covered everything from CPUs and GPUs to supercomputers and from modern process technologies and latest fab tools to high-tech industry trends.

  • Gururu
    Have to give this one to the brilliant, resourceful and hard-working engineers over there. Is it always going to be high maintenance, even sustainable?
    Reply
  • DS426
    Even if it's difficult to maintain and implement, it's clearly worth it when talking about a 10x efficiency gain; imagine a $10 Bn datacenter only costing let's say $2 Bn (still accounting for non-GPU related costs) at the same AI training performance level. I believe we do need to focus more on optimizations than outright XPU compute performance, whether it's going a similar route as DeepSeek or other alternatives. I'd say this might also drive some changes to CUDA as NVIDIA obviously isn't going to like these headlines and what, $500B of market cap erased in a matter of hours?

    Broadly speaking, China seems to be impeccable at reverse engineering and than iterating over others, all at savings to both cost and time-to-market.
    Reply
  • JamesJones44
    I don't think it will hurt sales, even at 10x faster it still took 2 months if I read that right. Companies are likely to invest in hardware until that time becomes significantly less than 2 months.
    Reply
  • Gaidax
    There is a ton of low-hanging optimization fruit in this whole area that is still left to pick for sure.

    But in the end the industrial AI requirements are not going anywhere.
    Reply
  • vanadiel007
    It's scary to see AI being added to everything you use. Especially since it's not intelligent.
    Reply
  • Itsahobby
    Excellent article! Driving down the cost of AI is good for everybody. There will be more demands for Nvidia chips.
    Reply
  • Gaidax
    vanadiel007 said:
    It's scary to see AI being added to everything you use.
    Why?
    Reply
  • anoldnewb
    Gaidax said:
    Why?
    People should be concerned about rampant AI proliferation with out adequate safeguards because it is very prone to hallucinations. How can you be certain the output is reliable?
    People should have reason to be concerned were AI failure can harm people; for example, driving a semitruck at 70 MPH, automating air traffic control, flying airplanes, writing code for applications were failure can hurt people.
    Reply
  • Gaidax
    anoldnewb said:
    People should be concerned about rampant AI proliferation with out adequate safeguards because it is very prone to hallucinations. How can you be certain the output is reliable?
    People should have reason to be concerned were AI failure can harm people; for example, driving a semitruck at 70 MPH, automating air traffic control, flying airplanes, writing code for applications were failure can hurt people.
    So is the Internet. And heck it's FAR wilder at that too.

    Will you have some dumb answers from AI? Yes, but so will happen with your average Joe getting advice to drink bleach from his social media circle to cure a certain viral infection.

    Compared to nonsense you can read on the Internet from the "experts", AI is already far more curated and correct, and it will only get better, even if once in a while it will still fudge it up.

    In the end - the person in front of a display needs at the very least minimal understanding of what this notification means, or heck how Internet works at all.

    https://i.gyazo.com/9e0270540be1f3059d1ebaa229904000.png
    Reply
  • wsanders11
    The maintainability of their PTX compiler is not much of an issue. Before too long China will have their own chips rivalling "last year's model" NVidia, and they will be widely available and supported, at least outside the US and EU.
    Reply