Braintrust
End-to-end LLM evaluation and experimentation platform
Verdict
Fastest SDK for building evaluation pipelines. Dataset versioning, scorer plugins, and team review workflows are polished. Growing fast in enterprise evaluations space.
Other Observability & Evals
- LangfuseProduction
Open-source LLM engineering platform — traces, evals, prompts
- Phoenix (Arize)Stable
Open-source AI observability with embedding and RAG tracing
- Weave (W&B)Stable
Weights & Biases' LLM tracing and eval framework
- HeliconeStable
LLM gateway with logging, caching, and cost analytics
- OpenLLMetryEmerging
OpenTelemetry-based observability for LLMs
- PromptLayerStable
Prompt version control, A/B testing, and analytics platform

