Research Note
MMLU Benchmark Design and Comparison Record
MMLU stands for Measuring Massive Multitask Language Understanding. The [original paper](https://arxiv.org/abs/2009.03300) introduced a multiple-choice evaluation spannin
In this article
MMLU Benchmark Design and Comparison Record
MMLU stands for Measuring Massive Multitask Language Understanding. The original paper introduced a multiple-choice evaluation spanning 57 academic and professional subjects.
The authors presented it as a test of broad knowledge and problem solving. Their reported models had uneven subject performance, frequently did not know when they were wrong, and remained weak in several socially important areas. The benchmark was not designed as a complete product-quality or safety score.
The official MMLU repository contains evaluation code, category mappings, and a historical leaderboard. A reported result still depends on the exact model, checkpoint, number of in-context examples, prompt format, scoring method, implementation, and date.
MMLU averages different subjects into one headline number. That improves summary comparison but hides dispersion. A model can gain in one domain and lose in another while the average moves little. It also says nothing directly about latency, cost, tool use, retrieval, long-context behavior, privacy, security, accessibility, or task-specific error costs.
Later research reinforces the need for method discipline. MMLU-Pro increased the number of answer options, removed some noisy or trivial questions, emphasized reasoning, and tested prompt sensitivity. It is a related benchmark, not an interchangeable score.
The HELM MMLU record demonstrates a more transparent comparison route by standardizing prompts and exposing subject breakdowns, raw prompts, and predictions. It does not make MMLU a universal capability measure.
E013 used MMLU as an accessible comparison device. Public copy should keep the exact table and evaluation setup attached to every model claim and avoid a permanent leaderboard.
Sources
Follow the evidence.
- Introducing Our Open Mixed Reality Ecosystemabout.fb.com
- Meet Your New Assistant: Meta AIabout.fb.com
- Meta Quest 3S announcementabout.fb.com
- Introducing the Meta AI appabout.fb.com
- Introducing Muse Sparkabout.fb.com
- Muse Spark 1.1 action featuresabout.fb.com
- Meta company informationabout.meta.com
- Introducing Llama 3.1ai.meta.com
- Introducing Meta Llama 3ai.meta.com
- Measuring Massive Multitask Language Understandingarxiv.org
- MMLU-Proarxiv.org
- HELM MMLU recordcrfm.stanford.edu
- Meta Horizon OS developer documentationdevelopers.meta.com
- MMLU implementation repositorygithub.com
- Meta Llama models repositorygithub.com
- Meta Llama 3 Community Licensegithub.com
- Meta Llama 3 model cardgithub.com
- Meta 2025 full-year resultsinvestor.atmeta.com
- NIST Generative AI Profilenvlpubs.nist.gov
- Spotify episodeopen.spotify.com
- Meta's Llama license is still not Open Sourceopensource.org
- Android Open Source Projectsource.android.com
- Meta generative AI privacy guidefacebook.com
- NIST AI Risk Management Frameworknist.gov
- YouTube episodeyoutu.be