Anthropic Constitutional AI
Principle-based self-improvement for harmlessness and helpfulness
Verdict
Pioneering approach where the model critiques and revises its own outputs based on a set of principles. Built into Claude models. Influence on the field is massive even if not a standalone product.
Other Guardrails & Safety
- Guardrails AIStable
Add input/output validation and safety rails to LLM calls
- NeMo GuardrailsExperimental
NVIDIA toolkit for programmable guardrails via Colang language
- Llama Guard 4Stable
Meta's fine-tuned safety classifier for prompt and response screening — now natively multimodal
- RebuffExperimental
Prompt injection detection API for LLM applications
- Microsoft PresidioProduction
Data protection and anonymization for PII in LLM pipelines
- Lakera GuardStable
Real-time prompt injection and jailbreak protection API

