Skip to content
StrataHub

Datasets & Benchmarks · 2020

MMLU

A Berkeley PhD student assembled a 57-subject exam, from law to medicine, that became the LLM industry's report card.

In September 2020, a few months after GPT-3 stunned researchers with knowledge nobody had explicitly trained into it, Berkeley PhD student Dan Hendrycks and collaborators released a benchmark designed to measure that knowledge properly. They called it Massive Multitask Language Understanding, or MMLU.

The test contains close to 16,000 multiple-choice questions across 57 subjects, harvested from practice materials for real human exams: medical licensing, law, accounting, US history, college physics, professional psychology and more. Random guessing scores 25 percent, and the authors estimated expert-level human performance at roughly 90 percent.

The first results were humbling. GPT-3, the most capable model of its day, managed about 44 percent, barely halfway between guessing and expertise, with near-random performance on subjects like law and morality. That gap made MMLU the perfect yardstick for the scaling era.

For the next four years, virtually every major model announcement led with its MMLU score. GPT-4's 86.4 percent in 2023 was headline news, and the climb of open models up the MMLU ladder tracked the closing gap with proprietary labs.

Success bred obsolescence. By 2024 frontier models were clustering near 90 percent, researchers documented flawed questions and training-data contamination, and harder successors like MMLU-Pro and GPQA appeared. MMLU became a case study in Goodhart's law: when a measure becomes the target, it stops measuring, and the industry must build a harder exam.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.