Article
MMLU Benchmark
MMLU, short for Measuring Massive Multitask Language Understanding, is a benchmark introduced to evaluate language-model accuracy across 57 academic and professional subj
MMLU Benchmark
MMLU, short for Measuring Massive Multitask Language Understanding, is a benchmark introduced to evaluate language-model accuracy across 57 academic and professional subjects.
What it measures
The subjects include areas such as elementary mathematics, United States history, computer science, law, and morality. The questions are multiple choice. The original researchers intended the benchmark to test the breadth and depth of a model's knowledge and problem-solving ability.
The paper also emphasized shortcomings in the models it evaluated. Performance varied substantially by subject, models often did not know when they were wrong, and results remained weak in several socially important areas.
Why a score needs context
An MMLU result belongs to a particular model and evaluation setup. Prompt wording, the number of examples supplied in context, scoring rules, model checkpoints, and implementation details can affect comparisons.
A single average also compresses performance across very different subjects. Two models with similar headline scores can have different strengths, weaknesses, costs, response behavior, and suitability for real work.
How E013 used it
E013 used MMLU as a simple way to compare Llama 3 with earlier or competing models. The shorthand of an SAT for models helps introduce the idea of a broad test, but MMLU is not a universal measure of intelligence or product quality.
The canonical article treats Meta's reported results as vendor benchmark evidence under the linked evaluation setup. It does not convert them into a permanent ranking.
Later evaluation context
MMLU-Pro introduced a harder and more reasoning-focused variant with additional answer choices and different questions. Its scores are not directly interchangeable with original MMLU scores.
The HELM MMLU record shows the value of standardized prompts, subject breakdowns, and access to prompts and predictions. That improves auditability but still does not turn MMLU into a complete product evaluation.
Editorial and verification notes
When citing an MMLU score, identify the exact model variant and evaluation configuration and link to the reported method. Do not compare numbers from different tables unless their setups are compatible. MMLU does not replace evaluation on the intended task. [[How to Read an MMLU Score]] provides the practical comparison method.
Sources
Follow the evidence.
- Introducing Llama 3.1ai.meta.com
- Measuring Massive Multitask Language Understandingarxiv.org
- YouTube episodeyoutu.be
- Introducing Muse Sparkabout.fb.com
- HELM MMLU recordcrfm.stanford.edu
- Introducing Our Open Mixed Reality Ecosystemabout.fb.com
- Muse Spark 1.1 action featuresabout.fb.com
- Android Open Source Projectsource.android.com
- Meta Llama 3 Community Licensegithub.com
- Meta Quest 3S announcementabout.fb.com
- NIST AI Risk Management Frameworknist.gov
- Meta company informationabout.meta.com
- Meta's Llama license is still not Open Sourceopensource.org
- MMLU implementation repositorygithub.com
- Introducing the Meta AI appabout.fb.com
- Meta Llama models repositorygithub.com
- MMLU-Proarxiv.org
- NIST Generative AI Profilenvlpubs.nist.gov
- Meta Llama 3 model cardgithub.com
- Meta 2025 full-year resultsinvestor.atmeta.com
- Meet Your New Assistant: Meta AIabout.fb.com
- Meta Horizon OS developer documentationdevelopers.meta.com
- Spotify episodeopen.spotify.com
- Meta generative AI privacy guidefacebook.com
- Introducing Meta Llama 3ai.meta.com