The new Mac Studio comes with your choice of an M5 Max or M5 Ultra chip. The M5 Max is the more traditional of the two. It’s an 18-core CPU with 6 “super” cores and 12 performance cores, paired with a 32- or 40-core GPU, a 16-core neural engine, and up to 128 GB of unified memory. Bandwidth is where the two configurations split: the 32-core version runs at 460 GB/s, while the higher 40-core configuration hits 614 GB/s. Both retain neural accelerators in the GPU and hardware-accelerated ray tracing.
The M5 Ultra, on the other hand, is more interesting. The M5 Ultra is the first quad-chip Apple Silicon model. This means that instead of merging two Max chips like previous Ultra chips, Apple is combining two dual-chip M5 Max chips using a new UltraFusion interconnect, which the company says provides more than 4.4 TB/s of inter-chip bandwidth and more than six times the connection density of the previous design. The point of all this plumbing is to make four dies behave like a single processor, so the software doesn’t have to think about the seams.
There are two variants of the M5 Ultra. The base model has a 30-core CPU (10 super cores and 20 performance cores), a 64-core GPU with neural accelerators, and a 32-core neural engine. The high-end variant has a 36-core CPU and an 80-core GPU, along with up to 512GB of unified memory (available from October) and up to 16TB of storage. Memory bandwidth is 1.2 TB/s across the board – Apple says that’s 50% higher than the M3 Ultra’s roughly 819 GB/s, and it’s the number that matters most for the AI workloads this machine is expected to handle.
AI performance
So how does the M5 Ultra Mac Studio actually perform in AI? Well, definitely better than the M3 Ultra Mac Studio. I used both models for a series of local AI tasks, largely through LM Studio. On average, I found the M5 Ultra Mac Studio to perform 30-40% better than the M3 Ultra Mac Studio. For example, while I could get up to around 40 tokens per second with a 4-bit MLX optimized version of Qwen 3.8 27B on the M3 Ultra Mac Studio, the newer model typically got closer to 55 tokens per second. Likewise, while it almost always took more than a second for the first token on the M3 Ultra model, it was common to get closer to half a second on the M5 Ultra model. With lighter models like the Qwen 3.5 122B A10B, I get closer to 80 tokens per second on the M5 Ultra compared to 60 tokens per second on the M3 Ultra.
On slightly more demanding tasks like generating images in ComfyUI using a template like Z-Image Turbo, the results were even more impressive. I found the M5 Ultra generated images at least twice as fast as the M3 Ultra – and often three times faster, under the same conditions.
This can play a role in daily life beyond running local LLMs instead of using cloud services like ChatGPT or Claude. For example, I sometimes use small local templates to help me clean up dictation when writing with voice. I found that on the M5 Ultra, these admittedly light tasks were even faster than they already were on the M3 Ultra.
The improvements, of course, are due to a variety of factors, including memory bandwidth and increased GPU cores. Regardless, it’s indeed much faster for local AI tasks. For those who need even more performance, support for Thunderbolt 5 plus RDMA (a networking technique that allows machines to directly share memory without bogging down the processor) makes it possible to group multiple Mac Studios together. Apple claims that a cluster of four systems delivers up to three times better AI inference performance than a single unit. Of course, this means spending at least double the amount of cash, so this probably appeals more to businesses than individuals.
And the M5 Max? Well, I didn’t have a Mac Studio M5 Max to test during this review, but I did have a MacBook Pro M5 Max with enough memory to run the aforementioned Qwen 3.8 27B. The M5 Max remains a serious chip for local AI, and perhaps much more realistic for anyone who doesn’t want to spend the thousands of dollars required for an M5 Ultra machine. It actually works the same AI way as the M3 Ultra, getting similar tokens per second and taking a similar amount of time to process a prompt. That’s solid, considering how widely praised the M3 Ultra was when it released last year.
Graphics and media
Not everyone has bought into the local AI craze, and the M5 Ultra Mac Studio is excellent for many things besides running local LLMs. The M5 Ultra’s GPU supports second-generation dynamic caching (which allocates GPU memory on the fly rather than reserving worst-case amounts), hardware-accelerated mesh shading, and third-generation hardware ray tracing. For 3D and VFX work, these architectural changes are arguably more significant over the M3 Ultra than the modest CPU increase – the M3 Ultra also had an 80-core GPU cap, but a much older design.
The Media Engine is where the Ultra departs most from the M5 Max. It offers hardware acceleration for H.264, HEVC, ProRes and ProRes RAW, as well as AV1 decoding – and the Ultra contains two video decode engines, four video encode engines and four ProRes encode/decode engines, which is double the capacity of the M5 Max. Apple claims the machine can simultaneously play back up to 33 streams of 8K ProRes 422 at 30fps, up from 24 on the previous generation.
In practice, these upgrades enable incredibly impressive performance. I’ll be the first to admit that my graphics and video needs are by no means demanding, but in rendering and conversion tests, the M5 Ultra Mac Studio was blazingly fast, even compared to the M3 Ultra model. The basic videos I made exported instantly from Final Cut, and the more demanding projects I loaded also looked surprisingly fast. Basically, the M5 Ultra Max Studio is an incredible machine for video editors, photographers, and anyone who needs the best graphics performance you can currently get from a desktop computer.
