aifollow.news Search
Back 量子位 AI 报道
量子位 AI 报道· · Original publication time

Inferact releases TPU inference kernel; Kimi K3 test outpaces GB200 under stated conditionsMachine translation

Automatically verified and published · Generated and evidence-checked automatically; not reviewed by a human.

AI-assisted summary

Inferact's megakernel combines Kimi K3's 92 layers into one Pallas program and runs with DSpark speculative decoding. In a test using 16 TPU v7 Ironwood chips versus 16 GB200 chips, both with vLLM, throughput was 709 versus 452 tokens per second. The kernel is currently tailored to Kimi K3 and needs adaptation for other model architectures.

The complete source text is not yet available.

Read at the original source
Found an error? Send a correction