MMLU, HumanEval, GPQA, and beyond — how to interpret LLM benchmarks and what they actually measure.
Five passes over the same idea, each from a different angle. Do them in order, or jump to whichever you need.
LLM benchmarks attempt to quantify model capabilities across reasoning (MMLU, ARC), coding (HumanEval, SWE-Bench), math (GSM8K, MATH), and safety. Understanding what each benchmark measures, its limitations, and how contamination/gaming affects results is essential for making informed model selection decisions.