☀️ AI Morning Minute: MMLU
For years, this was the one AI test score that actually meant something. Now everybody aces it.
If you followed AI news at any point in the last few years, you saw this number without knowing its name. Every time a company bragged that its model “knows more” than the last one, MMLU was usually the scoreboard they were pointing at. It’s the test that turned “how smart is this AI” into a single percentage. And it’s a good example of how quickly a good test can wear out.
What it means
MMLU (Massive Multitask Language Understanding) is a giant multiple-choice quiz for AI models. About 16,000 questions spread across 57 subjects, everything from grade-school math to law, medicine, history, and computer science. Four choices each, pick the right one. The model’s final score is just its batting average across all of it. Higher percentage, more it got right.
When it came out in 2020, it was brutally hard for the AI of the day. That was the point. It covered so many subjects that a model couldn’t fake broad knowledge.
Why it matters
It became the number companies sold you on. For a stretch, MMLU was the headline stat in basically every model launch, the AI version of a car’s miles-per-gallon sticker. When you heard one model was “smarter” than another, this was often the receipt.
It shows how fast this stuff moves. In 2020, the best models scored around 32%, barely better than guessing. By 2026, the top ones cluster in the low 90s, all bunched within a few points of each other. When everybody scores an A, the test stops telling you who’s actually best.
It got so easy it kind of broke. Researchers found the test is now “saturated,” and there are even known wrong answers baked into it that cap the real ceiling around 95%. So the field moved on to harder versions like MMLU-Pro. Honestly, if you see a company bragging about a plain MMLU score today, that’s a bit of a tell they’re cherry-picking.
Simple example
Think about a spelling test that was genuinely tough for a third-grader. Great tool for a while. You could actually see who studied and who didn’t. But hand that same test to a room of adults, and everybody gets 95%. The test didn’t get worse. The people taking it just outgrew it, so a perfect score stops telling you anything about who’s the better speller.
That’s MMLU now. Still a fine test. The models just outgrew it.

