MMLU Explained: What the 57-Subject LLM Benchmark Actually Tests
MMLU covers 57 academic subjects with 15,908 multiple-choice questions from elementary to professional level, and remains one of the most widely cited LLM benchmarks despite significant criticisms.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
The Massive Multitask Language Understanding benchmark (arXiv:2009.03300) by Hendrycks et al. tests LLMs on 57 subject areas spanning STEM, humanities, social sciences, and professional domains. With 15,908 multiple-choice questions at difficulty levels from elementary to expert, it was designed to measure the breadth of knowledge a model has absorbed during pretraining.
The 57 Subjects
The subjects span a wide range:
- STEM: Abstract algebra, anatomy, astronomy, chemistry, computer science, electrical engineering, physics, mathematics
- Professional: Medical genetics, clinical knowledge, jurisprudence, accountancy, management, marketing
- Humanities: History, philosophy, prehistory, moral scenarios
- Social Sciences: Economics, psychology, sociology, geography, politics
- Miscellaneous: Miscellaneous knowledge spanning professional medicine, law, and business
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Scoring Methodology
Each question has four choices (A, B, C, D). The random baseline is 25%. MMLU is evaluated in two modes:
Few-shot (5-shot): Five example questions with answers are prepended to each test question. This tests whether the model can use the context format correctly and apply in-context learning.
Zero-shot: No examples. Tests pure parametric knowledge.
Score is reported as accuracy (fraction of questions answered correctly). Most papers use 5-shot.
Score Distribution Across Models
| Model | MMLU 5-shot |
|---|---|
| Random baseline | 25.0% |
| GPT-3 (2020) | 43.9% |
| GPT-3.5 | 70.0% |
| GPT-4 | 86.4% |
| Claude 3 Opus | 86.8% |
| Llama 3 70B | 82.0% |
| Human expert estimate | ~89.8% |
Running MMLU Evaluation
# Using EleutherAI's lm-evaluation-harness
pip install lm-eval
lm_eval --model hf \
--model_args pretrained=meta-llama/Meta-Llama-3-8B \
--tasks mmlu \
--num_fewshot 5 \
--device cuda:0 \
--batch_size 8 \
--output_path ./results/llama3-8b-mmlu
The Criticisms
Data contamination: Many MMLU questions appear verbatim in Common Crawl, meaning models may have memorized answers rather than demonstrating reasoning. Studies estimate contamination rates of 10-30% for some subjects.
Multiple-choice is not real-world: Production tasks require open-ended generation, not selecting from four options. A model can score 80% on MMLU while failing completely on open-ended versions of the same questions.
Inconsistent across papers: Different implementations (zero-shot vs few-shot, log-prob scoring vs generation, chain-of-thought or not) produce incomparable numbers. Papers cherry-pick the evaluation protocol that makes their model look best.
MMLU-Pro: A Harder Variant
MMLU-Pro (arXiv:2406.01574) addresses some criticisms by:
- Expanding to 10 choices instead of 4 (reducing random baseline to 10%)
- Filtering for harder questions with less ambiguity
- Requiring reasoning rather than pattern-matching
GPT-4 scores approximately 72% on MMLU-Pro versus 87% on MMLU - a more discriminating benchmark.
Further Reading

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
GPT-6 Astra Benchmarks, Pricing, and Safety: What Developers Need to Know
OpenAI's GPT-6 Astra delivers state-of-the-art coding and agentic performance, but costs 50% more than GPT-4o. This guide covers benchmarks, pricing, safety, and practical advice for developers deciding whether to upgrade.
What Is GPT-5.6 Sol Ultra Will Be in Codex? A Practical Overview
GPT-5.6 Sol Ultra is a rumored model optimized for code generation, integrated into Codex. We analyze the claims, potential capabilities, and what developers should expect.
Building reliable agentic AI systems: A Practical Overview
A practical guide to building reliable agentic AI systems covering structured outputs, observability, fallbacks, and cost controls with real code examples.
// discussion
Comments