signalAI热榜2026-10-03
Baseten engineers test: LLM-generated inference engine up to 90% faster than vLLM
Baseten engineers used Claude Code with Fable 5 and the MetaInfer framework to auto-build an inference engine named VibeQwen for Qwen-3.6-35B-A3B with NVFP4 on a single B200. It achieved up to 90% faster single-stream decoding than vLLM 0.25.1, cutting first-token time from 28ms to 12ms, and 71% higher throughput at concurrency 32. The experiment consumed about 1.7 billion tokens and 200 B200 hours, with minimal human intervention.
- for who
- Engineers and researchers working on LLM inference optimization or evaluating automated engine generation.
- why now
- LLM-built inference engines now beat vLLM by 90%, enabling faster, cheaper AI deployments.
- what changes
- LLM-generated, task-specific inference engines can outperform general-purpose open-source engines like vLLM, making specialized optimization practical without manual kernel engineering.
- to do
- Use the MetaInfer repository and a capable agent like Claude Code with Fable 5 to generate a custom inference engine for a specific model and hardware, then benchmark against vLLM.
key points
- VibeQwen: 90% faster single-stream decoding than vLLM 0.25.1
- First token dropped from 28ms to 12ms; 71% higher throughput at concurrency 32
- Claude Code autonomously built engine over ~1 week using 1.7B tokens and 200 B200 hours
#inference optimization#LLM-generated engine#vLLM comparison
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source