Article
MMLU Benchmark: What the Score Measures and Misses
MMLU is a 57-subject multiple-choice benchmark. A score only makes sense with the exact model, prompt, shot count, scoring method, implementation, and date.
How to Read an MMLU Score
MMLU is a multiple-choice benchmark covering 57 academic and professional subjects. Its score summarizes performance under a particular evaluation setup.
It is not a universal intelligence score, a permanent model ranking, or proof that a product will work well on your task.
flowchart LR
A["Reported MMLU score"] --> B["Exact model and checkpoint"]
B --> C["Prompt and shot count"]
C --> D["Scoring and implementation"]
D --> E["Subject breakdown"]
E --> F["Evaluation date"]
F --> G["Only then compare"]
What MMLU measures
The original MMLU paper introduced a test spanning subjects such as elementary mathematics, United States history, computer science, law, medicine, and morality.
Questions are multiple choice. The original work aimed to test the breadth and depth of model knowledge and problem-solving ability across many tasks.
The authors also documented limitations in the models they evaluated. Results varied sharply by subject. Models frequently did not know when they were wrong. Performance remained weak in several socially important areas.
That history matters because the benchmark began as a way to expose shortcomings as well as measure progress.
Why the average is useful
One average makes broad comparison easier. A reader can quickly see whether a model performed near chance or handled a substantial portion of the test.
The official MMLU repository preserves evaluation code, category groupings, and historical results. It gives reviewers a method and artifact trail rather than only a chart image.
The average can also hide important differences. Two models with similar scores may have different subject strengths. One model may improve in humanities while another improves in science. The headline number can remain close even when their useful behavior differs.
What belongs beside every score
Name the exact model and checkpoint. A base model and an instruction-tuned model are different artifacts. A later checkpoint with the same family name can have a different context length, training mix, post-training method, and behavior.
Record the number of examples supplied in context, often called the shot count. Record the prompt format, answer extraction, scoring method, implementation, and any changes to the dataset.
Record the date. Benchmarks, model endpoints, evaluation libraries, and published tables change.
If two scores come from different methods, the difference may not represent a real capability gain.
What MMLU does not measure
MMLU does not directly measure latency, cost, energy use, tool calling, retrieval quality, long-context behavior, image understanding, data privacy, security, accessibility, or the quality of a complete assistant product.
It also does not measure the consequence of an error in your workflow. A wrong answer in a brainstorming tool and a wrong answer in a regulated decision process do not carry the same risk.
A product can score well and still fail because it cannot cite the needed source, follow a procedure, preserve permission boundaries, or stop when evidence is missing.
Later benchmark work
MMLU-Pro was designed as a more challenging and reasoning-focused variant. It expanded the answer choices, removed some noisy or trivial questions, and tested prompt sensitivity.
Its scores are not directly interchangeable with original MMLU scores. The benchmark, question set, and method changed.
The HELM MMLU record demonstrates another useful practice: standardized prompts, subject breakdowns, and transparent access to raw prompts and predictions. That improves auditability without turning one benchmark into a complete evaluation.
How E013 used MMLU
The surviving E013 outline uses MMLU as an accessible way to compare Llama 3 with earlier and competing models. The shorthand of an SAT for models helps explain a broad test.
The analogy should stop there. MMLU spans many subjects, but it is not an admissions test for production use. Meta’s launch table remains a vendor evaluation under disclosed settings.
The responsible public claim identifies the model variant and links to the exact evaluation record. It avoids mixing tables and does not convert one result into a timeless ranking.
The task-level check
Use MMLU to understand one part of a model’s general evaluation history. Then build a test for the work that matters.
A task evaluation should include representative inputs, difficult cases, missing information, conflicting sources, restricted data, expected abstention, error costs, and the point where a person reviews the result. Record the full system, not only the model.
E027, E029, and E031 add later Llama evaluation and safety context. E038 shows why claimed benchmark performance still needs reproducible implementation evidence.
The maintained [[MMLU Benchmark Research Profile]] owns the canonical benchmark identity. This page owns the practical reading method. AI assisted with research, drafting, and validation. Publication remains unauthorized pending benchmark, source, technical, accessibility, and founder review.
Sources
Follow the evidence.
- Introducing Llama 3.1ai.meta.com
- Measuring Massive Multitask Language Understandingarxiv.org
- YouTube episodeyoutu.be
- Introducing Muse Sparkabout.fb.com
- HELM MMLU recordcrfm.stanford.edu
- Introducing Our Open Mixed Reality Ecosystemabout.fb.com
- Muse Spark 1.1 action featuresabout.fb.com
- Android Open Source Projectsource.android.com
- Meta Llama 3 Community Licensegithub.com
- Meta Quest 3S announcementabout.fb.com
- NIST AI Risk Management Frameworknist.gov
- Meta company informationabout.meta.com
- Meta's Llama license is still not Open Sourceopensource.org
- MMLU implementation repositorygithub.com
- Introducing the Meta AI appabout.fb.com
- Meta Llama models repositorygithub.com
- MMLU-Proarxiv.org
- NIST Generative AI Profilenvlpubs.nist.gov
- Meta Llama 3 model cardgithub.com
- Meta 2025 full-year resultsinvestor.atmeta.com
- Meet Your New Assistant: Meta AIabout.fb.com
- Meta Horizon OS developer documentationdevelopers.meta.com
- Spotify episodeopen.spotify.com
- Meta generative AI privacy guidefacebook.com
- Introducing Meta Llama 3ai.meta.com