Datasets & Benchmarks · 2012
The Titanic Dataset
Kaggle turned a 1912 shipwreck into the first machine learning problem for millions of beginners.
When the Titanic sank on 15 April 1912, about 1,500 of the 2,224 people aboard died. Because the disaster was so intensely investigated, an unusually complete record of the passengers survived: names, ages, ticket classes, fares, cabins and fates, later compiled by volunteer historians into resources like Encyclopedia Titanica.
In September 2012, the competition platform Kaggle repackaged those records as a permanent tutorial contest called Titanic: Machine Learning from Disaster. Competitors get 891 labeled passengers to train on and must predict survival for 418 more, using features such as sex, age, passenger class, fare and family aboard.
The data encodes a grim social lesson. The 'women and children first' evacuation order and the geography of class-divided decks mean that sex and ticket class dominate every model. The trivial rule that all women survive and all men perish scores about 76 percent, and beating that baseline by a meaningful margin is surprisingly hard.
That difficulty is exactly what makes it a good classroom. Titanic teaches the unglamorous core of applied machine learning: strong baselines, missing values (many ages are unknown), feature engineering from text like names and titles, and the danger of overfitting a small test set through repeated leaderboard submissions.
It also teaches a lesson about leakage: since the real outcomes are public history, a perfect score proves only that someone looked up the answers. Millions of people have made the Titanic dataset their first ever model, which arguably makes a century-old shipwreck the most common entry point into modern data science.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.