AMD unveils industry-first Stable Diffusion 3.0 Medium AI model generator tailored for XDNA 2 NPUs — designed to run locally on Ryzen AI laptops
An offline image generator for XDNA 2 NPUs.
AMD, in collaboration with Stability AI, has unveiled the industry's first Stable Diffusion 3.0 Medium AI model tailored for the company's XDNA 2 NPUs which process data in the BF16 format. The model is designed to run locally on laptops based on AMD's Ryzen AI laptops and is available now via Amuse 3.1.
The model is a text-to-image generator based on Stable Diffusion 3.0 Medium, which is optimized for BF16 precision and designed to run locally on machines with XDNA 2 NPUs. The model is suitable for generating customizable stock-quality visuals, which can be branded or tailored for design and marketing applications. The model interprets written prompts and produces 1024×1024 images, then uses a built-in NPU pipeline to upscale them to 2048×2048 resolution, resulting in 4MP outputs, which AMD claims are suitable for print and professional use.
The model requires a PC equipped with an AMD Ryzen AI 300-series or Ryzen AI MAX+ processor, an XDNA 2 NPU capable of at least 50 TOPS, and a minimum of 24GB of system RAM, as the model alone uses 9GB during generation.
The key advantage of the model is, of course, that it runs entirely on-device; the model enables fast, offline image generation without needing Internet access or cloud services. The model is aimed at content creators and designers who need customizable images and supports advanced prompting features for fine control over image composition. AMD even provides examples. A prompt to draw a toucan looks as follows:
"Close up, award-winning wildlife photography, vibrant and exotic face of a toucan against a black background, focusing on the colorful beak, vibrant color, best shot, 8k, photography, high res."
To use the model, users must install the latest AMD Adrenalin Edition drivers and the Amuse 3.1 Beta software from Tensorstack. Once installed, users should open Amuse, switch to EZ Mode, move the slider to HQ, and enable the 'XDNA 2 Stable Diffusion Offload' option.
Usage of the model is subject to the Stability AI Community License. The model is free for individuals and small businesses with less than $1 million in annual revenue, though licensing terms may change eventually. Keep in mind that Amuse is still in beta, so its stability or performance may vary.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Follow Tom's Hardware on Google News to get our up-to-date news, analysis, and reviews in your feeds. Make sure to click the Follow button.
Anton Shilov is a contributing writer at Tom’s Hardware. Over the past couple of decades, he has covered everything from CPUs and GPUs to supercomputers and from modern process technologies and latest fab tools to high-tech industry trends.
-
-Fran- Ah, this is probable the enterprise version of that garbage they're trying to install in the driver package.Reply
Regards. -
usertests Reply
This isn't about Strix Halo, this is about XDNA2. That's included in Strix Point, Krackan, and even the newly announced Ryzen AI 5 330. So basic image generation capabilities using a 50 TOPS NPU will come to even sub-$400 laptops.abufrejoval said:With a 9GB model, you'd be far better off using a 12GB GPU, way faster and cheaper even in a laptop.
Strix Halo is great technology... at a terrible price
If it really needs 24 GB instead of 16 GB to run fast (or at all), that could be a problem in that segment given there are many systems with soldered RAM. Hopefully all laptops with "AI" in the processor name will come with over 12 GB, matching the Windows AI/Copilot requirement of 16 GB, but 24-32 GB could be rare. So you need 1-2 SODIMM slots to add more memory yourself. -
dalek1234 Reply
Given the specs, it's look like it's targeting professionals, who are willing to pay premium for a premium product. Besides, you can charge a premium when there is no competition in this space. Intel is asleep, and Nvidia doesn't have an x86 license.abufrejoval said:...
Strix Halo is great technology... at a terrible price
...
For the rest of us, a Strix Halo with same GPU, half CPU cores, and 32-64 MB of ram would make more sense. Supposedly those SKU are in the pipeline. I read today that some Chinese outfit is bringing those variants to Desktop. -
usertests Reply
The story is not about Strix Halo, it's about the XDNA2 NPU, which is also available in millions of Strix Point and Krackan laptops and mini PCs. All with around the same 50-55 TOPS of performance.dalek1234 said:Given the specs, it's look like it's targeting professionals, who are willing to pay premium for a premium product. Besides, you can charge a premium when there is no competition in this space. Intel is asleep, and Nvidia doesn't have an x86 license.
Maybe it can run faster on Strix Halo though. I don't know if the NPU itself benefits from the additional memory channels and bandwidth, or if's only the iGPU that needs it. Running Stable Diffusion on Strix Halo's iGPU may be better. -
DS426 Reply
Folks, easy to beat up anything with "AI" slapped on it, but let's step back a bit here.abufrejoval said:With a 9GB model, you'd be far better off using a 12GB GPU, way faster and cheaper even in a laptop.
Strix Halo is great technology... at a terrible price
It's designed to be cheaper than a dGPU, built on using a wider bus on commodity DRAM, but they charge HBM prices for it.
Sure, Strix Halo is expensive, but I don't know that faster and cheaper is always the case in the comparison being used here. As others already pointed out, regular Strix comes in a lot lower, so the trade-off could be cheaper and slower. Anyone doing serious AI inferencing should absolutely rely on a PC with a dGPU and ideally 16 GB of VRAM or more as that only adds a lot more models than can be ran entirely in VRAM rather shared with system RAM, or otherwise just a greater portion in VRAM. That said, the whole point of NPU's on laptops was to strike a balance of having some level of decent AI inferencing performance while maintaining decent battery life. You fire up a dGPU to about 100% utilization on a laptop and battery life falls off a cliff. If it's more of a workstation laptop that typically stays on a charger, different scenario.
I'm seeing laptops for about $600 on Amazon (ASUS, etc.) and elsewhere with Ryzen AI 5 340, 16 GB of RAM. 32 GB RAM models are significantly higher, which I agree is silly, but I think this is an OEM problem and not so much AMD pricing. In any case, those prices should eventually come down as they saturate marketplaces.
Also realize there are smaller models; the purview of this article is this specific AI model which recommends 32 GB of RAM. $600 and even lower laptops can run AI models for inferencing locally and at much faster speeds than other traditional laptops with faster CPU's but no NPU's. No, not everyone cares about this today, but the presumption is that it's only growing as it comes of age, and then the regret later on down the road would be not having an NPU. -
snemarch AMD should ditch Amuse and focus on improving whatever open-source project(s) instead.Reply
It's a steaming pile of garbage; locked-down wrt. model selection, and censorship that's even more silly than what OpenAI performs. -
usertests Reply
I've heard the RTX 3060 has only 100 TOPS. Nvidia doesn't list that but they have RTX 4060 at 242 TOPS: https://www.nvidia.com/en-us/geforce/graphics-cards/compare/abufrejoval said:I just cannot imagine the main model really running on the NPU, they must be using the GPU for the image generation and the NPU for upscaling, a bespoke solution that may be little more than a one-off tech demo, and with little chance of working with any other hardware, even their own a generation older or younger, because you'd have to redesign the workload split between so heterogeneous bits of hardware: there is no software stack to support that generically.
I guess it still has to be verified to what extent the NPU is being used here, but someone could do that and test the speed of generation too.
It should be noted that the Ryzen AI 5 330 only has 2 compute units, but the same XDNA2 as other models. So it's unlikely that the iGPU would be much help there.
I don't think AMD is ditching the NPU soon, although some would like them to. I expect we'll see an XDNA3 at 100 TOPS within the next two years. XDNA1 early adopters have 10-16 TOPS and no BF16 support so they are hosed. XDNA2+ will catch on as long Microsoft is pushing Copilot+. -
Bikki This is sorely needed, because windows fails to get any tangible impact on daily AI use, which left the NPU useless for most. Now at least normal people can have some fun generate images with SD 3.Reply
Noted that SD 3 is a very old model (released on june 2024) and have been superseded by newer models like SD 3.5 and flux.1 -
Bikki @abufrejoval Apple chip can do both generation and upscaling solely on the NPU. Is it because of your 10 TOPS that Amuse put the generation on GPU? I can't think of a reason why AMD's NPU can't do it.Reply -
DS426 Dang, this got intense, lol.Reply
Thank you for that deep-dive and elaboration, @abufrejoval . I don't mean to hype up NPU's... indeed Windows uses Phi Silica, a Small Language Model, to run things locally on Copilot PC's, as benefits are limited today. The two letters "AI" have been the biggest technology marketing term in many years, with things like "Copilot+" just being specific branding of this.
Anyways, I haven't tried Amuse on my Ryzen AI 375 HX 32 GB RAM work laptop, only on my Radeon RX 7900 rig at home. The difference in SDXL 3 and 3.5 models is noticeable, along with the different optimizations, fine-tuns, etc. I've played around with LM Studio far more, including with this HX 375 laptop. In a lot of cases, I don't see any NPU utilization according to Task Manager, which doesn't really matter anyways as the 890M is surprisingly strong for being an iGPU. Maybe the unified memory benefit as mentioned earlier? I can't say for certain.