Streaming 104 GB of MoE from Disk

Slotstream is a Swift/MLX tool that runs Qwen3.8-Flash-Next — a 125B MoE weighing 104 GB at 4-bit — on Apple Silicon Macs with far less RAM than the model weighs, by streaming routed expert weights on demand from NVMe SSD into a fixed cache-slot pool.

Read more →

Nativ: A Native Mac App for Running Frontier Models Locally

Prince Canuma, the author of MLX-VLM, shipped v0.0.1 of Nativ: a native SwiftUI app that turns an Apple Silicon Mac into a private, no-subscription local AI server supporting text, vision, audio, and code models via MLX — with OpenAI- and Anthropic-compatible inference endpoints built in.

Read more →