Phoenix (Arize)
Open-source AI observability with embedding and RAG tracing
Verdict
Strongest tool for RAG evaluation — UMAP visualisation of embeddings, retrieval quality scoring, and hallucination detection. Run locally or in Arize Cloud.
Other Observability & Evals
- LangfuseProduction
Open-source LLM engineering platform — traces, evals, prompts
- Weave (W&B)Stable
Weights & Biases' LLM tracing and eval framework
- HeliconeStable
LLM gateway with logging, caching, and cost analytics
- OpenLLMetryEmerging
OpenTelemetry-based observability for LLMs
- PromptLayerStable
Prompt version control, A/B testing, and analytics platform
- BraintrustEmerging
End-to-end LLM evaluation and experimentation platform

