signalHacker News Show HN2026-09-21
Show HN: jevals - replacing LLM judges with typed Jev decisions
jevals replaces LLM-based agent evaluation with Jev's typed decision models, running eight evals (tool choice, groundedness, relevancy, injection, PHI, etc.) in one HTTP request at $0.00006 and 0.33s per trace via pip install, using Jev through TypeSafe/Vercel APIs or local open-weight models Kev (Qwen3) and Laya (ModernBERT). Unlike LLM judges that cost 6-11 round trips per sample and showed 92x-913x score variance in LangChain's comparison, Jev returns calibrated probabilities in a single forward pass at $0.042/M input tokens with p50 244ms latency, enabling evaluation on every trace and inside agent loops.
- why now
- jevals: new, $0.00006/request
- topic
- AI Tools & Agent Workflows
- source
- Hacker News Show HN
#Jev
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source