Promptfoo, RAGAS, DeepEval — frameworks and tools for systematic LLM evaluation at scale.
Five passes over the same idea, each from a different angle. Do them in order, or jump to whichever you need.
Evaluation frameworks provide structured tools for assessing LLM performance. Promptfoo enables prompt comparison and testing. RAGAS evaluates RAG pipelines. DeepEval provides unit-test-like assertions for LLM outputs. OpenAI Evals, LangSmith evaluators, and custom frameworks complete the ecosystem.