Ship AI that survives contact.
Engineering-focused coverage of defensive AI. Guardrail architecture, classifier ensembles, model hardening, output filtering, refusal training, and the response patterns that hold under adversarial pressure in production systems.
Indirect Prompt Injection Explained: Defenses That Hold
read →// start here
The four guides most people need first
A guardrail stack is four decisions: what you block on the way out, what you record, how you know it still works, and which controls your architecture is missing entirely.
- Output filtering architecture for production LLMs The four-layer pipeline: deterministic guards, classifiers, schema validation, and selective LLM-as-judge.
- LLM audit logging: what to log, redact, and retain Field-level tiers, write-path redaction, and the six-month retention floor the EU AI Act sets.
- Monitoring LLM outputs in production Anomaly detection, output drift, and alert routing that does not page on noise.
- LLM guardrail benchmarks: build your own eval set Why vendor F1 scores do not transfer, and the four metrics that decide a guardrail purchase.
Not sure which controls apply to your deployment? The Guardrail Gap Analyzer maps application shape, trust boundary, and data sensitivity to the control set that applies, and scores every gap it finds.
Independent, specialist, and free to read
AI Defense publishes focused, sourced guides on a single topic. No paywall, no account, no ad tracking.
AI Defense — in your inbox
Defensive AI engineering — guardrails, hardening, response — delivered when there's something worth your inbox.
No spam. Unsubscribe anytime.