Skip to content
StrataHub

Datasets & Benchmarks · 2021

HumanEval

OpenAI built it to measure Codex: 164 hand-written Python problems that became the exam for every code AI since.

In July 2021, OpenAI unveiled Codex, the model behind GitHub Copilot, in a paper by Mark Chen and colleagues titled Evaluating Large Language Models Trained on Code. To evaluate it they faced a contamination problem: any programming exercise scraped from the internet might already sit in the training data. So engineers and researchers hand-wrote 164 brand new Python problems, each a function signature, a docstring and a set of hidden unit tests.

The benchmark also changed how code generation was scored. Earlier work compared generated code to reference text with similarity metrics borrowed from translation, which rewarded code that looked right. HumanEval instead ran the code: a solution counts only if it passes the tests, measured by a metric called pass@k, the chance that at least one of k sampled attempts works.

The first results framed the coming decade. GPT-3 solved essentially none of the problems, while Codex solved about 28.8 percent in a single attempt and far more when allowed many samples, an early sign that generating and filtering multiple candidates could stand in for reliability.

For the next several years, pass@1 on HumanEval was the number every code model advertised, the way image models had once advertised ImageNet accuracy. Open and closed models alike climbed the ladder until frontier systems passed 90 percent.

Saturation revealed the benchmark's limits: 164 short, self-contained puzzles say little about real software work in sprawling codebases. Successors like SWE-bench moved the goalposts toward genuine engineering, but they all inherited HumanEval's core idea, that the only fair judge of generated code is whether it runs and passes the tests.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.