signal量子位2026-09-27
Google TPU Runs Kimi 57% Faster Than Nvidia GPU, Using DeepSeek Inference Framework
Inferact, a startup founded by the original vLLM team, ran Kimi K3 on 16 Google TPU v7 chips achieving 709 tokens per second, 57% faster than Nvidia GB200. They used a custom megakernel in Pallas and DeepSeek's DSpark framework, with an acceptance length of 6 and 8.5ms per decode step. The code is open-sourced.
- for who
- AI inference engineers and developers interested in TPU optimization
- why now
- Open-sourced TPU kernels let startups beat Nvidia GPUs by 57% on Kimi, immediately.
- what changes
- Developers can now use the open-sourced tpu-megakernels repo to run models like Kimi K3 on TPU faster, with optimizations benefiting the vLLM ecosystem.
- to do
- Explore the tpu-megakernels repository and test it with Kimi K3 or other supported models.
key points
- 16 TPU v7 hit 709 tokens/s on Kimi K3, beating GB200 by 57%
- Megakernel merges hundreds of kernels; DSpark adds speculative decoding with acceptance length 6
- Inferact from vLLM team raised $150M seed, code is open-sourced
#tpu#kimi#inference optimization#vllm#megakernel
score
score 9 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source