Local AI On Apple Silicon uses 7X Less RAM

Local AI On Apple Silicon uses 7X Less RAM

The Only Bamboo
The Only Bamboo
2 影片觀看·2026年8月2日

Turbo Fieldfare runs a 26 billion parameter Gemma 4 model on a Mac in about 2GB of RAM, streaming most of the weights off the SSD, and it still gets 23 tokens per second on an M3 Max. We pull the design apart and show why it only works on Apple Silicon.

🔗 Relevant Links
https://github.com/drumih/turbo-fieldfare

❤️ More about us
Radically better observability stack: https://betterstack.com/
Written tutorials: https://betterstack.com/community/
Example projects: https://github.com/BetterStackHQ

📱 Socials
Twitter: https://twitter.com/betterstackhq
Instagram: https://www.instagram.com/betterstackhq/
TikTok: https://www.tiktok.com/@betterstack
LinkedIn: https://www.linkedin.com/company/betterstack

📌 Chapters:
0:00 A 26B model in 2GB of RAM
0:54 Running it locally on an M3 Max
1:23 How Gemma's mixture of experts works
2:14 Two piles: 1.35GB resident, 12.9GB on SSD
2:57 What happens when you generate a token
3:30 The router problem: you can't prefetch
3:58 Why this only works on a Mac
4:40 A file format the GPU reads directly
5:06 Hiding the disk read behind the shared expert
5:27 The LFU expert cache
6:16 Wrap up