signalHacker News Show HN2026-09-30
Show HN: Reflex Engine Beats Both Llama.cpp and vLLM on Cold-Start to TTFT
Reflex is a GGUF-native Rust and CUDA inference engine that compiles all CUDA kernels at build time, eliminating runtime JIT compilation and enabling ~456ms p50 cold start on a Tesla T4 for Qwen3-0.6B-Q4_K_M. In a real Runpod deployment, a cold invocation took ~42-44 seconds wall clock, of which Reflex's own load was ~1.6 seconds, the rest being platform provisioning. Compared to Runpod's official worker-vllm, Reflex was ~3.4x faster (~42-44s vs ~150s to servable state).
- for who
- Serverless GPU developers running bursty single-shot inference workloads.
- why now
- Cold-start speed cuts serverless GPU bills, beating vLLM 3.4x on real Runpod tests.
- what changes
- Inference engine selection shifts from maximizing concurrent throughput to minimizing time-to-first-token for per-second billed serverless.
- to do
- Clone the Reflex repository, build with CUDA toolkit, and deploy on Runpod to measure its cold-start advantage.
key points
- Reflex compiles all CUDA kernels at build time, avoiding JIT overhead for ~456ms p50 cold start on T4.
- On Runpod, Reflex took ~42-44s total, ~1.6s own load, vs vLLM's ~150s, a 3.4x speedup.
- Engine designed for serverless per-second billing: no batching, one request per process, platform handles queue and autoscaling.
#inference engine#cold start#GGUF#Rust#CUDA#serverless
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source