Llama Guard 4
Meta's fine-tuned safety classifier for prompt and response screening — now natively multimodal
Verdict
Production-deployable content moderation model, shipped alongside the Llama 4 herd with native image+text screening in a single dense model (no separate vision variant needed anymore). Run it as a sidecar to screen every input/output. MLCommons hazard taxonomy built in. Free, open-weight, and fast on a single GPU.
Other Guardrails & Safety
- Guardrails AIStable
Add input/output validation and safety rails to LLM calls
- NeMo GuardrailsExperimental
NVIDIA toolkit for programmable guardrails via Colang language
- RebuffExperimental
Prompt injection detection API for LLM applications
- Microsoft PresidioProduction
Data protection and anonymization for PII in LLM pipelines
- Lakera GuardStable
Real-time prompt injection and jailbreak protection API
- Azure AI Content SafetyProduction
Microsoft's content moderation API for text and images

