Two machines, two runtimes, one model key: 1.85x throughput from a heterogeneous local fleet
An AMD Strix Halo running llama.cpp on Vulkan and an M5 Max running MLX on Metal, serving the same model behind a single model name. Eight concurrent requests finished in 30 seconds instead of 55, aggregate throughput went from 28 to 53 tokens per second, and median latency dropped 2.66x. Inside: the fairness control that proves it is real capacity and not a misconfigured first box, the four routing defects it took to get there, and an honest account of what this does not speed up.