Weave (W&B)
Weights & Biases' LLM tracing and eval framework
Verdict
Native W&B integration makes it ideal for teams already using wandb for ML experiments. Excellent trace/eval correlation. Strong Python SDK with minimal boilerplate.
Other Observability & Evals
- LangfuseProduction
Open-source LLM engineering platform — traces, evals, prompts
- Phoenix (Arize)Stable
Open-source AI observability with embedding and RAG tracing
- HeliconeStable
LLM gateway with logging, caching, and cost analytics
- OpenLLMetryEmerging
OpenTelemetry-based observability for LLMs
- PromptLayerStable
Prompt version control, A/B testing, and analytics platform
- BraintrustEmerging
End-to-end LLM evaluation and experimentation platform

