Skip to content
StrataHub

Models · 2018

BERT

BERT learned to read by playing a fill-in-the-blank game with billions of sentences.

In October 2018, Google researchers released BERT, short for Bidirectional Encoder Representations from Transformers. Within months it had set new records across a wide range of language understanding tasks and would soon help power Google Search itself.

BERT's defining trick was reading in both directions at once. Earlier language models read left to right, but BERT looked at the words before and after a gap simultaneously, giving it a richer sense of context.

To train this way, the researchers used a clever game called masked language modeling. They hid about 15 percent of the words in each sentence and asked the model to guess them, forcing it to learn grammar, meaning and world knowledge from raw text.

Like GPT, BERT was built on the Transformer, but it used the encoder half and was designed for understanding rather than generation. Fine-tuned versions excelled at question answering, sentiment analysis and more.

BERT arrived just months after GPT-1, and together the two marked a turning point: the era of large pre-trained language models had begun, and nearly every system that followed would build on their ideas.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.