Rebuff
Prompt injection detection API for LLM applications
Verdict
Purpose-built for prompt injection detection — a real attack vector in RAG systems. Uses a canary token technique alongside an LLM classifier. Early-stage; combine with input sanitization.
Other Guardrails & Safety
- Guardrails AIStable
Add input/output validation and safety rails to LLM calls
- NeMo GuardrailsExperimental
NVIDIA toolkit for programmable guardrails via Colang language
- Llama Guard 4Stable
Meta's fine-tuned safety classifier for prompt and response screening — now natively multimodal
- Microsoft PresidioProduction
Data protection and anonymization for PII in LLM pipelines
- Lakera GuardStable
Real-time prompt injection and jailbreak protection API
- Azure AI Content SafetyProduction
Microsoft's content moderation API for text and images

