Inferact's first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode, against 450 tokens/s for our GB200 baseline, both with DSpark speculative decoding. Without speculative decoding, the megakernels for K3 and Qwen 3.8 27B deliver roughly 1.4 to 2× the decode throughput of the GB200 baseline at batch sizes 1 through 8.