signalDEV Community2026-09-22
Designing an eval harness for prompt-injection detection: what measuring my defenses actually taught me
A 35-prompt golden set (25 attacks across five classes, 10 benign controls) with a two-layer judge (hardcoded counters plus LLM fallback for tool-blind classes) evaluated four defenses on a LangGraph travel-concierge agent running gpt-oss-120b. Instruction hierarchy alone cut indirect injection from 80% to 0% attack success, while the harness reported false positives alongside attack success to expose over-hardening costs. The runner persists per-sample JSON verdicts, resumes after rate-limit crashes, and uses fixed seed/temperature=0 for comparable defense runs.
- evidence
- 6 sources carry this story · confidence high
- topic
- AI Tools & Agent Workflows
- source
- DEV Community
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source