Nvidia’s TiDAR experiment could speed up AI token generation using hybrid diffusion decoder — new research boasts big throughput gains, but limitations remain

MEMBER EXCLUSIVE
Nvidia chip
(Image credit: Getty Images / Antonio Bordunovi)

As the AI race between companies, nations, and ideologies continues apace, Nvidia has released a paper describing TiDAR, a decoding method that merges two historically separate approaches to accelerating language model inference. Language models produce text one token at a time, where a token is a small chunk of text, such as a word fragment or punctuation mark.

Each token normally requires a full forward pass through the model, and that cost dominates the speed and expense of running today’s AI systems. If a model can safely produce several tokens per step without losing quality, it could lead to faster response times, lower GPU hours, and reduced operating costs per request, all of which could add up to substantial savings for operators running large AI deployments, running the latest Nvidia hardware.

Latest Videos From
TOPICS
Luke James
Contributor

Luke James is a freelance writer and journalist.  Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.