Why unified memory is the whole story on a Mac
A Windows or Linux PC has separate VRAM (on the GPU) and system RAM (for everything else). A Mac doesn't: the CPU, GPU, and Neural Engine all draw from one shared pool of unified memory. That means the number printed on the spec sheet, "16GB," "36GB," "128GB," is the actual ceiling for how large a model you can load, not a rough guide. Check the exact figure against a model's size in the hardware calculator before assuming a Mac "should" handle something.
The current Apple Silicon lineup for local AI
| Mac | Chip | Unified memory | Max bandwidth | Good for |
|---|---|---|---|---|
| MacBook Air | M4 | 16-32GB | 120GB/s | 8B models comfortably at 16GB; 13-14B at 24-32GB |
| Mac Studio | M5 Max | 36-128GB | 614GB/s | 32B-70B depending on configuration |
| Mac Studio | M5 Ultra | 96-512GB | 1.2TB/s | The tier where very large open-weight models become usable locally at all |
The M5 Mac Studio, announced August 25, 2026, is the newest entry here; most "run AI on a Mac" guides you'll find elsewhere predate its launch and still reference the M4 or M3 generation as current. One timing detail worth knowing if you're buying at the top end: the general lineup ships September 22, 2026, but the 512GB M5 Ultra configuration specifically isn't available until "late October," about a month later.
Capacity vs. speed: the real tradeoff
Unified memory's advantage is capacity: a 128GB Mac Studio can load models no consumer discrete GPU can touch, because no GPU ships with anywhere near that much VRAM. The tradeoff is bandwidth. Even the M5 Ultra's 1.2TB/s is well below what a high-end discrete GPU delivers from its dedicated VRAM, so token generation on a Mac can lag a well-matched NVIDIA setup even when both technically fit the same model. If you need maximum tokens-per-second on a model that fits either way, a discrete GPU usually wins. If you need to fit a model that simply won't fit on any reasonably priced discrete GPU, unified memory is often the only realistic option.
Which Mac for which model size
- 8B models: the base 16GB MacBook Air handles these comfortably at Q4, with room for everyday context lengths.
- 13-14B models: step up to 24GB or 32GB on the MacBook Air, or use a Mac Studio.
- 27-32B models: the M5 Max Mac Studio at 64GB or higher gives comfortable headroom.
- 70B and larger: this is Mac Studio territory specifically, at 64GB (M5 Max) or higher; the M5 Ultra's 96GB+ configurations open up the largest local-friendly open-weight models.
Run the exact model size through the hardware calculator and select "Apple Silicon Mac (unified memory)" to see the real number against your configuration.
Buying memory: it's expensive, and it's not upgradable later
Apple prices unified memory upgrades at a real premium, commonly cited around $25 per additional gigabyte, and unlike a desktop PC, you cannot add memory after purchase. If you're buying specifically for local AI, it's worth deciding your target model size tier before ordering, since going back for more memory means buying an entirely new machine.