Exploring Apple Silicon’s local AI performance with the Mac Studio and M4 Max — M4 Max beats GB10 and Strix Halo in decode throughput, but memory bandwidth isn't everything

To test LLM inference performance, we once again turn to llama.cpp, which includes native Metal/MLX support on Apple Silicon when built from source. We used the llama-benchy project as our benchmarking harness.

Mac Studio M4 Max

(Image credit: Tom's Hardware)

We tested three models of various memory footprints and architectures: Qwen 3.6-35B-A3B, a recent mixture-of-experts model; Gemma 4 12B, a small dense model, and OpenAI’s gpt-oss-120b, a large mixture-of-experts model. We used Unsloth’s four-bit quantizations of these models for our testing.

Out of the gate with Qwen 3.6-35B-A3B, the Mac Studio delivers prompt processing speeds that beat Strix Halo but are still behind GB10, especially with longer contexts.

Latest Videos FromTom's Hardware

Flip over to the throughput results, and the M4 Max pulls ahead of both GB10 and Strix Halo, turning in higher tokens-per-second throughput than both those systems at all context depths.

But despite having double the memory bandwidth of GB10 and Strix Halo on paper, the real-world improvement in token throughput on the M4 Max is just 25% on Qwen 3.6-35B-A3B. That’s certainly welcome, but it’s not nearly the speedup you might expect with twice the memory bandwidth on tap.

The dense Gemma 4 12B, on the other hand, really seems to let the Mac Studio’s bandwidth advantage shine. Tokens-per-second throughput scales more linearly with memory bandwidth with this model, and so the Mac Studio delivers a 1.8X increase in throughput over GB10 and a 2.26X increase over Strix Halo.

But prompt processing speeds are only in line with (or even slightly worse than) Strix Halo here, which isn’t a fantastic result for latency if that’s important to your application.

Finally, we’ll look at gpt-oss-120b. Prompt processing again lands between Strix Halo and GB10, and tokens-per-second throughput is far higher than with either of those systems. But the Mac Studio’s tokens-per-second throughput is, on average, 1.5x higher than GB10 and 1.6x higher than Strix Halo, despite its much higher memory bandwidth. Those are still impressive leads, but they show that memory bandwidth is just one component in judging the performance of an inference pipeline.

On the whole, our LLM inference testing shows that the M4 Max GPU can generate many more tokens per second than competing systems with large unified memory pools, but a memory bandwidth specification alone isn’t a good proxy for delivered performance. The actual boost you get from the M4 Max in LLM inference is highly dependent on the architecture of the model you want to run, so real-world benchmarking is still important.

Image generation performance

LLMs aren’t the only AI workload that one might want to run locally. Image and video generation also matter, and the performance of these diffusion models is generally bound by the compute throughput of the underlying GPU.

Mac Studio M4 Max

(Image credit: Tom's Hardware)

Running our Flux.2 Klein image generation test on the Mac Studio initially failed due to an apparent lack of GPU support for the FP8 data type that the ComfyUI workflow uses by default.

We were able to modify the ComfyUI workflow to use the base FP16 version of Flux.2 Klein on the Mac Studio, but that initial compatibility issue shows that you might not enjoy as smooth an image or video generation experience on Apple Silicon as you might on GB10 or even Strix Halo.

Compatibility issues aside, the M4 Max GPU requires even more time than the already pokey Radeon 8060S in this image generation test. Nvidia has repeatedly demonstrated the DGX Spark running as an image and video generation companion for Macs in its demos this year, and this result demonstrates why that pairing might make sense.

Image generation performance with the latest M5 GPU and its dedicated Neural Accelerators for matrix math might be better here, but we don’t have an M5-powered system at hand to compare against the M4 Max. Apple does claim that the M5 Max is “up to 3.8x faster” at image generation than the M4 Max, so take heed if this workload is important to you.

General CPU performance

We already gauged the CPU performance of the M4 Max in our initial review, but we wanted to stack it up against Strix Halo and GB10 directly to see how it performs against these popular platforms in the AI space. To establish a CPU performance ballpark for the M4 Max, we fired up the just-released Geekbench 7 and got to it.

Mac Studio M4 Max

(Image credit: Tom's Hardware)

In this benchmark, the M4 Max turns in a whopping 27% advantage in single-core performance and a 21% boost in multi-threaded performance versus the Ryzen Max+ 395’s 16-core, 32-thread Zen 5 x86 CPU.

And comparing Arm to Arm, the M4 Max delivers a 39% single-threaded advantage over the Cortex-X925 “big” core on the Nvidia GB10, along with a 22% advantage in multi-threaded work.

Those results are seriously impressive on their own, but Apple has even more advanced CPU cores in its arsenal now on board the M5 series. We’d expect even larger advantages in general use from those chips.

Geekbench is fine for a quick synthetic test, but I wanted to make sure that its results held in a real-world task, so I compiled the CPU-only version of llama.cpp from source across all three of our test systems using the Clang 21 compiler running across all available physical cores.

Mac Studio M4 Max

(Image credit: Tom's Hardware)

With all 16 of its CPU cores in the game, the M4 Max finishes this compilation job in about 37 seconds, anywhere from 33% to 37% faster than the Nvidia and AMD competition here. That’s even faster than the M4 Max’s Geekbench multi-threaded performance delta would suggest.

At least going by these high-level tests, Apple’s CPU performance among SoCs with large unified memory pools is best-in-class, and by a wide margin. And again, that’s with the last-gen M4, not the latest M5.

TOPICS
Jeffrey Kampman
Senior Analyst, Graphics

As the Senior Analyst, Graphics at Tom's Hardware, Jeff Kampman covers everything that has to do with graphics cards, gaming performance, and more. From integrated graphics processors to discrete graphics cards to the hyperscale installations powering our AI future, if it's got a GPU in it, Jeff is on it. 

  • Kindaian
    There are some significant issues with using a mac though, specially in terms of service setup and containers (no gpu passthrough for containers under macos due to cpu/hardware limitations).

    If those are not important to you, then the comparison is fair i would guess. BUT also one consideration is the price in relation to the performance.
    Reply
  • splus
    I don't get it - why would you test the old M4 Max Studio when there's a new M5 Max MacBook, which is better and faster, especially for AI? And, as you said, the M4 Max Studio configuration you tested basically isn't even available any more?
    Reply
  • JamesJones44
    splus said:
    I don't get it - why would you test the old M4 Max Studio when there's a new M5 Max MacBook, which is better and faster, especially for AI? And, as you said, the M4 Max Studio configuration you tested basically isn't even available any more?
    Prior to Apple limiting the M4 Max Studios RAM I would have argued 256 GB of VRAM would be better for a pro user than the M5 Max 128 GB. However, with Apple cutting that option, I agree, the M5 Max 128 GB MBP would have been a better comparison
    Reply
  • splus
    JamesJones44 said:
    Prior to Apple limiting the M4 Max Studios RAM I would have argued 256 GB of VRAM would be better for a pro user than the M5 Max 128 GB. However, with Apple cutting that option, I agree, the M5 Max 128 GB MBP would have been a better comparison
    It's good to have more memory, but this hardware and its memory bandwidth is simply too low for larger models. It would be too slow to be useful. 256 or 512 GB would be useful only with much faster chip (like M5 Max or Mx Ultra) and with much higher memory bandwidth. M4 Max doesn't have either. Same applies to the Spark and Ryzen 395. 128 GB is their practical and usable max limit.
    Reply
  • Dmtrii
    I don't understand how you're testing Geekbench multi-core! Either Windows is broken (I don't know, I haven't used it for 10+ years) or you chose the wrong power profile.

    Here's my result of Geekbench 7 for the Strix Halo 128GB (Framework, Linux, Performance profile) - https://browser.geekbench.com/v7/cpu/1327
    More than 30 000 for multi-core!

    The same was for Geekbench 6 - every time I got 15% better results in multi-core 🤷‍♂️
    Reply
  • Bikki
    @Jefferey Lamma.cpp does not natively support mlx, it supports metal but not Mlx. Ollama or Llm studio do.

    On Ollama new mlx backend, Qwen prefill speed increases by 1.5x and decode 2-3x on m4 pro compared to old version that uses llama.cpp backend.

    Ps: I’m currently an AI engineer, gladly contributr to AI article’s process unpaid. contact me at bik.huynguyen gmail
    Ref: https://ollama.com/blog/mlx
    Reply