signalDEV Community2026-10-02
Tripwire: a guardrails layer for LLM apps, built for a friend
Tripwire is a lightweight guardrails library for LLM apps that checks user input and model output for prompt injection, secrets, and personal data. Initial tests caught 36% of attacks, rising to 38% on fresh attacks and 42% after rule additions, with 20% on a public benchmark. Adding Gemma as a second opinion improved detection to 92%.
- for who
- Developers building LLM applications who want independent safety checks without relying on the model alone.
- why now
- Hacktoberfest submission offers a lightweight, regex-based guardrails layer for LLM apps.
- what changes
- A developer can now plug in two-call guardrails into any LLM app, with optional Gemma second opinion, without building safety infrastructure.
- to do
- Install from GitHub and integrate guard.checkInput and guard.checkOutput in your own LLM application.
key points
- Initial version caught 36% of attacks, 38% on fresh ones, 42% after rule additions
- Adding Gemma as second opinion raised detection to 92% on fresh set
- Rules run in ~0.07 ms per sentence, 0.6 ms for 1,100 characters
#guardrails#llm safety#open source
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source