41 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalHacker News Show HN2026-09-30

Show HN: Reflex Engine Beats Both Llama.cpp and vLLM on Cold-Start to TTFT

Reflex is a GGUF-native Rust and CUDA inference engine that compiles all CUDA kernels at build time, eliminating runtime JIT compilation and enabling ~456ms p50 cold start on a Tesla T4 for Qwen3-0.6B-Q4_K_M. In a real Runpod deployment, a cold invocation took ~42-44 seconds wall clock, of which Reflex's own load was ~1.6 seconds, the rest being platform provisioning. Compared to Runpod's official worker-vllm, Reflex was ~3.4x faster (~42-44s vs ~150s to servable state).

for who
Serverless GPU developers running bursty single-shot inference workloads.
why now
Cold-start speed cuts serverless GPU bills, beating vLLM 3.4x on real Runpod tests.
what changes
Inference engine selection shifts from maximizing concurrent throughput to minimizing time-to-first-token for per-second billed serverless.
to do
Clone the Reflex repository, build with CUDA toolkit, and deploy on Runpod to measure its cold-start advantage.
key points
  • Reflex compiles all CUDA kernels at build time, avoiding JIT overhead for ~456ms p50 cold start on T4.
  • On Runpod, Reflex took ~42-44s total, ~1.6s own load, vs vLLM's ~150s, a 3.4x speedup.
  • Engine designed for serverless per-second billing: no batching, one request per process, platform handles queue and autoscaling.
#inference engine#cold start#GGUF#Rust#CUDA#serverless
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source