29 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalAI热榜2026-10-03

Baseten engineers test: LLM-generated inference engine up to 90% faster than vLLM

Baseten engineers used Claude Code with Fable 5 and the MetaInfer framework to auto-build an inference engine named VibeQwen for Qwen-3.6-35B-A3B with NVFP4 on a single B200. It achieved up to 90% faster single-stream decoding than vLLM 0.25.1, cutting first-token time from 28ms to 12ms, and 71% higher throughput at concurrency 32. The experiment consumed about 1.7 billion tokens and 200 B200 hours, with minimal human intervention.

for who
Engineers and researchers working on LLM inference optimization or evaluating automated engine generation.
why now
LLM-built inference engines now beat vLLM by 90%, enabling faster, cheaper AI deployments.
what changes
LLM-generated, task-specific inference engines can outperform general-purpose open-source engines like vLLM, making specialized optimization practical without manual kernel engineering.
to do
Use the MetaInfer repository and a capable agent like Claude Code with Fable 5 to generate a custom inference engine for a specific model and hardware, then benchmark against vLLM.
key points
  • VibeQwen: 90% faster single-stream decoding than vLLM 0.25.1
  • First token dropped from 28ms to 12ms; 71% higher throughput at concurrency 32
  • Claude Code autonomously built engine over ~1 week using 1.7B tokens and 200 B200 hours
#inference optimization#LLM-generated engine#vLLM comparison
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source