Models · 2018
GPT-1
Before ChatGPT had a name, there was GPT-1, a quiet 2018 experiment in reading before answering.
In June 2018, OpenAI released a paper with an unglamorous title: 'Improving Language Understanding by Generative Pre-Training.' The model it introduced, later called GPT-1, launched a family that would eventually reshape the entire field.
The core idea was pre-training. Rather than train a fresh model for each language task, GPT-1 first read a large corpus of books, learning to predict the next word over and over. This gave it a broad, general sense of how language works before it saw any specific task.
GPT-1 was built on the Transformer architecture published a year earlier, using only its decoder half. After pre-training, a small amount of fine-tuning let it tackle tasks like answering questions or judging whether two sentences agreed in meaning.
By today's standards it was tiny, with around 117 million parameters, and it drew little public attention. But it validated a recipe, generative pre-training followed by fine-tuning, that would scale astonishingly well.
Each successor, GPT-2, GPT-3 and beyond, kept the same basic approach and simply grew larger, trained on far more text. The path from this modest paper to ChatGPT ran almost straight.
Related stories
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.