Nvidia and Mistral AI's super-accurate small language model works on laptops and PCs

Mistral-NeMo-Minitron 8B art from Nvidia
(Image credit: Nvidia)

Nvidia and Mistral AI have released a new small language model that purportedly features "state-of-the-art" accuracy in a tiny footprint. The new LM is known as the Mistral-NemMo-Minitron 8B, a miniaturized version of NeMo 12B that has been pruned from 12 billion to 8 billion parameters.

The new 8 billion-parameter small language model was shrunken down through two different AI optimization methods, said Bryan Catanzaro, VP of deep learning research at Nvidia, in a blog post. The team behind the new LM used a process that combines pruning and distillation. "Pruning downsizes a neural network by removing model weights that contribute the least to accuracy. During distillation, the team retrained this pruned model on a small dataset to significantly boost accuracy, which had decreased through the pruning process."

Latest Videos FromTom's Hardware
Aaron Klotz
Contributing Writer

Aaron Klotz is a contributing writer for Tom’s Hardware, covering news related to computer hardware such as CPUs, and graphics cards.

  • Tom_Neverwinter
    I mean llama 3, 3.1 and most 8b models will run on cpu with 0 issues. I am literally running these on a orange pi 5 plus for fun. if they get model switching working so it can load whisperai, then unload it or be more efficient. I can then load a 8b model. process the data. unload the model. then load a xtts model output voice to the user and repeat. all on 8gb. my orangepi5plus has 16GB of ram. so I dont need to offload whisper the model or xtts but the cpu bottleneck even at 6TOPS is painfully slow at this time.
    Reply