34 signals
HOlO V1 IS LIVEone ranked AI digest a day, scored in publicREAD HOW IT WORKS →HOlO V2 STARTSyour X account, your signals, every day
signalDEV Community2026-10-02

Tripwire: a guardrails layer for LLM apps, built for a friend

Tripwire is a lightweight guardrails library for LLM apps that checks user input and model output for prompt injection, secrets, and personal data. Initial tests caught 36% of attacks, rising to 38% on fresh attacks and 42% after rule additions, with 20% on a public benchmark. Adding Gemma as a second opinion improved detection to 92%.

for who
Developers building LLM applications who want independent safety checks without relying on the model alone.
why now
Hacktoberfest submission offers a lightweight, regex-based guardrails layer for LLM apps.
what changes
A developer can now plug in two-call guardrails into any LLM app, with optional Gemma second opinion, without building safety infrastructure.
to do
Install from GitHub and integrate guard.checkInput and guard.checkOutput in your own LLM application.
key points
  • Initial version caught 36% of attacks, 38% on fresh ones, 42% after rule additions
  • Adding Gemma as second opinion raised detection to 92% on fresh set
  • Rules run in ~0.07 ms per sentence, 0.6 ms for 1,100 characters
#guardrails#llm safety#open source
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source