Skip to content
StrataHub

Datasets & Benchmarks · 2021

GSM8K

OpenAI hired writers to compose 8,500 grade school word problems, and the biggest models of 2021 flunked them.

By 2021 large language models wrote fluent essays but stumbled on arithmetic a ten-year-old could do. To measure the problem precisely, OpenAI researchers led by Karl Cobbe commissioned human writers to create 8,500 original grade school math word problems, each solvable in two to eight steps of basic arithmetic, released as GSM8K alongside the paper Training Verifiers to Solve Math Word Problems.

The deliberate simplicity was the point. These were not olympiad puzzles; they were problems about apples, allowances and train schedules. Yet GPT-3 failed the large majority of them, exposing a clean gap between linguistic fluency and multi-step reasoning.

The paper's own remedy was prophetic: instead of trusting a single answer, generate many candidate solutions and train a separate verifier model to rank them. Trading more computation at answer time for better reasoning would become one of the defining ideas of the decade.

GSM8K then became the stage for chain-of-thought prompting. In 2022, researchers showed that simply asking models to reason step by step before answering multiplied their scores, one of the most influential findings of the LLM era, and GSM8K was the benchmark where the effect was most vivid.

By 2024 frontier models scored above 95 percent, and a replication study with freshly written problems showed some models had partly memorized the test, so harder successors took over at the frontier. But GSM8K keeps its historical place: it is where the industry learned that reasoning could be measured, prompted and trained, the thread that leads directly to today's reasoning models.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.