Why you can trust Tom's Hardware
I was hoping to expand both the set of inference engines and AI models tested for this review, but I quickly ran into issues. The Ryzen AI Halo ships with a version of vLLM pre-installed, but it wasn't compatible with Qwen 3.6-35B-A3B when I tried to launch it, so I fell back to the old reliable llama.cpp.
We tested three models using llama.cpp: Qwen 3.6-35B-A3B, a relatively lightweight mixture-of-experts model that's been extremely popular of late, Google's Gemma 4 12B, another recent and relatively lightweight but dense model that activates all of its parameters per token, and gpt-oss-120B, a larger mixture-of-experts model that's been available for quite some time now. We used Unsloth's GGUF versions of these models in their Q4_K_M quantizations.
This time around, we're using llama-benchy as our benchmark harness. llama-benchy lets us get llama-bench-like performance results out of any model runner that can present us with an OpenAI-compatible endpoint, not just llama.cpp. That’s quite handy for comparing performance across inference engines.
First up, we'll look at latency and throughput performance with Qwen 3.6-35B-A3B:
View Original
View Original
And with the dense Gemma 4 12B:
View Original
View Original
And finally, with gpt-oss-120B:
View Original
View Original
At least with llama.cpp, the AI Halo's relative performance in single-user LLM serving versus GB10 isn't any different than what we saw several months ago when we took the Corsair AI Workstation 300 through its paces. While its tokens-per-second throughput is fine, albeit still slower than the Dell Pro Max GB10, its time-to-first-token latency falls far behind GB10 as context length grows.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
In a real-world scenario where llama.cpp can use (and is configured to use) its prompt caching features, these differences might not be as pronounced, but it's still important to note this worst-case behavior, especially for long-running coding workflows where context lengths can quickly grow.
Overall, the Ryzen AI Halo doesn't work any magic for Strix Halo inference performance compared to other implementations of this platform we've tested. While its tokens-per-second throughput is acceptable, its time-to-first-token latency can quickly rise to non-interactive levels with long contexts.
At the extremes of our testing, waiting two to four minutes for a model to start responding might be disruptive to an interactive workload like a coding assistant, even if the rate at which tokens flow is tolerable once they do start rolling.
I also ran the same check-in on ComfyUI image generation performance using the same basic Flux.2 Klein test we ran back in February, as well, and the same gulf that we saw in time-to-completion for ComfyUI work remains now. The GB10 GPU's larger shader complement chews through image generation far quicker than the Radeon 8086S.
Finally, we checked in on CPU performance with Geekbench 6. The AI Halo’s 16-core, 32-thread Zen 5 CPU holds an edge on the GB10’s 20-core Arm CPU complex in both single-threaded and multi-threaded performance, so tasks like code compilation could potentially be faster on the AI Halo. But that victory shouldn’t cause us to lose sight of the GB10’s all-around better AI performance.
Thermal performance and noise levels
We didn't want to tear down our Ryzen AI Halo to reveal its cooling system, but as we discussed at the beginning of this review, the system has vents on its front, top, and sides to allow for plenty of airflow, and you can see a decently sized copper heatsink through its rear vents.
Logging system temperatures on Linux is more difficult than it is on Windows, but we didn't see CPU or GPU temperatures higher than the mid-50 °C range when running the AI Halo through our typical workloads. All that suggests that this box is more than up to the task of cooling the chip inside.
As for noise, the Ryzen AI Halo runs its twin blower fans audibly even at idle, so it's always adding some amount of noise to a room. And under a ComfyUI generative workload, it gets significantly louder than the Dell Pro Max GB10 box we're using to represent that platform.
The noise signature of those fans is also a bit less refined than those in the GB10 boxes I’ve used. They have a notable high-pitched whine that's difficult to acclimate to when the system is placed on a desktop, whereas the GB10 systems I've used all just sound like moving air and are easy to ignore.
- MORE: Best Graphics Cards
- MORE: GPU Benchmarks and Hierarchy
- MORE: All Graphics Content
Current page: AI inference and image generation performance
Prev Page AMD Ryzen AI Halo Next Page Included software and playbooks
As the Senior Analyst, Graphics at Tom's Hardware, Jeff Kampman covers everything that has to do with graphics cards, gaming performance, and more. From integrated graphics processors to discrete graphics cards to the hyperscale installations powering our AI future, if it's got a GPU in it, Jeff is on it.
-
Neilbob Me when trying to understand the meaning and purpose of such devices (and A.I. in general).Reply
LKCi0gDF_d8
AMD just blowing bubbles, right along with N****a. I expect Intel will be next up with a meaningless overpriced plastic box. -
Pierce2623 Software compatibility is a negative compared to an ARM based machine?? Sure guys…. Also it seems only benchmarking some AI models and nothing else is purposely catering to Nvidia. AI hadn’t even started booming when we first heard about Strix Halo. So it clearly wasn’t designed solely for AI regardless of how they’ve advertised during the AI boom. If the AI boom hadn’t happened Strix Halo would be getting sold with 32GB for $1000. Knowing all that context, it seems weird to only test a few AI models.Reply -
Bigshrimp What are these devices for the price? I know this is a rhetorical question, but it's still funny that this was released. Costly for what it is...Reply -
suryasans AI writing by Google Gemini is much smarter and fairer than you. https://www.google.com/search?q=amd+ryzen+ai+halo+ai+developer+pc+vs+nvidia+gb10+power+consumption&sca_esv=4b04a474141781e0&rlz=1C1CHBF_enID1038ID1038&sxsrf=APpeQnsquHCA-jw-qqP98xNQhxFXMchnpg%3A1783360176978&ei=sOpLaqiwO_SVseMP0ty0qAw&biw=1600&bih=731&oq=amd+ryzen+ai+halo+ai+developer+pc+vs+nvidia+gb10+power+co&gs_lp=Egxnd3Mtd2l6LXNlcnAiOWFtZCByeXplbiBhaSBoYWxvIGFpIGRldmVsb3BlciBwYyB2cyBudmlkaWEgZ2IxMCBwb3dlciBjbyoCCAAyBRAhGKABSOhOUKIJWIk8cAF4AJABAJgBhwGgAasFqgEDNy4yuAEDyAEA-AEBmAIKoALIBcICCBAAGO8FGLADwgILEAAYgAQYogQYsAPCAgUQABjvBcICCBAAGIAEGKIEmAMAiAYBkAYFkgcDOC4yoAedE7IHAzcuMrgHxQXCBwM1LjXIBwyACAE&sclient=gws-wiz-serpReply -
maviz This is a completely useless piece of trash on a piece of shit ecosystem.Reply
I had the Ryzen with 32 GB and i will tell you the drivers are *USELESS*.
Even above that, the compatibility matrix is laughable...i had to sell both my CPU and my 6700 XT...and now i have a 5060 that is leaps and bounds faster and "just works"....and another 5070 TI which was not comparable to begin with. This box is hopelessly slow for any real AI work, priced like shit, performing like shit, i dont even know how they dare present this as any sort of evolution. You spend a few hundred more and get something 3 yrs ahead in the ecossystem and also MANY times faster...its laughable at best. -
JamesJones44 Reply
AI training and inference is all done on Linux and largely with ARM based servers believe it or not (they cost less and the CPU isn't all that relevant for AI workloads). The fact that x86 lags in this department isn't a surprise.Pierce2623 said:Software compatibility is a negative compared to an ARM based machine?? Sure guys…. Also it seems only benchmarking some AI models and nothing else is purposely catering to Nvidia. AI hadn’t even started booming when we first heard about Strix Halo. So it clearly wasn’t designed solely for AI regardless of how they’ve advertised during the AI boom. If the AI boom hadn’t happened Strix Halo would be getting sold with 32GB for $1000. Knowing all that context, it seems weird to only test a few AI models. -
dva852 Reply
AI is a CUDA world, at least for now. ROCm support is still spotty. x86/ARM is irrelevant.Pierce2623 said:Software compatibility is a negative compared to an ARM based machine??
AMD is leaning on open-source angle as it tries to breach CUDA moat. It also places higher focus on Windows platform, gunning for Windows devs as they move to AI. Nvidia is doing same with RTX Spark. Windows itself is moving to ARM. Lots of moving pieces.
You don't buy a $4K box to play games on an iGPU. Strix Halo was already considered "overpriced" at $2K--which admittedly was before-DC era. But repurposed for AI, it has gained new relevance, and higher valuation. IMO, not $4K high, but "higher."Pierce2623 said:Also it seems only benchmarking some AI models and nothing else is purposely catering to Nvidia.
At $4K vs DGX Spark's $4.7K, Strix AI is underwhelming. DGX has more compute, faster interconnect, better ecosystem. But the price is a bit irrelevant. It's a reference device, with official software (Win & Lin) stack, and most importantly, official support. It's meant to establish AMD's AI ecosystem. Think of it as placeholder for future iterations.
Even though Spark has faster compute, both it and Strix Halo are saddled with same memory bandwidth bottleneck. It's why the models used in AMD's benchmark (as well as THW ones) are either mixture-of-expert (read: low active parameter count), or small dense models. Even a midsized dense model would've brought both to a crawl. That's the real takeaway, that LPDDR5X is a makeshift solution to the AI bandwidth problem, and will need better solutions going forward.
Gorgon Halo also on LPDDR5X. Medusa Halo gets LPDDR6/X, in either late '27 or '28.
I looked on YT and "Ryzen AI Halo reviews" are popping up, so apparently today is "Halo AI's" official coming out party. Underlining the point is AMD's YT blurb, with Jack Huynh's mug trying reeeeaally hard to crack a smile. More happy thoughts, Jack.
For a more informative Ryzen Halo AI review, as well as providing a wider perspective,
Gz62bniDkpgView: https://www.youtube.com/watch?v=Gz62bniDkpg
After a forgettable foray in some high-end laptops, Strix Halo popped up in Shenzhen mini-PCs, and was immediatedly co-opted for AI use. The review below was 9 months ago, and gives a good perpective on Ryzen AI's WIP state at that point in time. Note also the $1.8K price.Pierce2623 said:AI hadn’t even started booming when we first heard about Strix Halo. So it clearly wasn’t designed solely for AI regardless of how they’ve advertised during the AI boom.
prIUKAbHlj8View: https://www.youtube.com/watch?v=prIUKAbHlj8 -
salgado18 Reply
For coding, for example: you set up a coding agent platform, like Claude Code, Codex or OpenCode. Then, you ask it to implement something or fix a bug. The platform will orchestrate AI agents and local tools to understand, plan, execute and test the task. That's the use case.Neilbob said:Me when trying to understand the meaning and purpose of such devices (and A.I. in general).
LKCi0gDF_d8
AMD just blowing bubbles, right along with N****a. I expect Intel will be next up with a meaningless overpriced plastic box.
Why such a box? Cloud-hosted AI models cost a lot of money if charged per token, and monthly subs have limits and are subject to price increases or vanish entirely. To run locally, even most gaming PCs struggle with such a heavy task that is running AI, let alone many of them in parallel. And setting up a central server with good hardware is crazy expensive and a big maintenance task.
Enter these boxes: you pu one on your desk, work on it or connected to it, and the agent platform uses its hardware to work, instead of cloud models. So your work is entirely local and contained to one PC.
I believe this is the future of PCs, and we're just seeing the first prototypes and proof-of-concepts. It's terrible value for gaming, regular working, media editing etc. But for local AI usage, it's a very tempting format. -
salgado18 Reply
What? I tested Claude Code and OpenCode, and it's just install -> set api key -> start prompting. Even with Ollama it's quick and easy.dva852 said:"For coding, for example: you set up a coding agent platform... and then it's immediately you spin up a containerized runtime, pipe your LLM through a RAG vector store, scaffold the MCP tool schema, YAML the CI/CD hooks, kubectl apply to the Kubernetes cluster--
FTFY o_O -
JamesJones44 The crazy part is this is starting to look like somewhat of a bargain. A year ago you could have gotten an 128GB M4 Max Mac Studio with double the memory bandwidth of the max+ 395 for $3700, now to get that level of performance you are in the 5.5+k range with GB 10 or Apple Silicon. Performance for local model operations is probably good enough for low to medium intensity tasks.Reply