vLLM team’s startup Inferact runs Kimi K3 on 16 TPU v7 chips 57% faster than GB200
Inferact, founded by the original vLLM team, wrote a megakernel inference kernel for Google TPUs. Paired with DeepSeek’s DSpark speculative decoding, 16 TPU v7 chips served Kimi K3 at 709 tokens per second versus 452 on GB200 under the same setup, QbitAI reports. The code is open source.