Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditionsMachine translation
Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.
Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.
The complete source text is not yet available.
Read at the original source