AI Benchmark Vocabulary

MMLU, HumanEval, GPQA, MT-Bench, BIG-bench, HellaSwag — how benchmarks work and why saturation matters.

Key vocabulary

  • MMLU (Massive Multitask Language Understanding) — tests knowledge across 57 academic subjects.
  • HumanEval — measures code generation ability via functional correctness of Python solutions.
  • GPQA (Graduate-Level Google-Proof Q&A) — expert-level science questions designed to resist web search.
  • Benchmark saturation — when top models cluster near ceiling performance, making differentiation difficult.
  • Teaching to the test — concern that models are trained on or fine-tuned specifically for benchmark data.
0 / 22 completed
1 / 22
A colleague says “MMLU is saturating.” What does this mean?

Frequently Asked Questions

What will I practice in "AI Benchmark Vocabulary | Coders Lingo"?

This is an AI Model Evaluation Language exercise set. It walks through 22 scenario-based multiple-choice questions built around real usage of AI Model Evaluation Language terminology that IT professionals encounter on the job.

Is this exercise free to use?

Yes. Every exercise on CoderSlingo, including this one, is free to complete with no account, sign-up, or paywall.

How many questions are in this exercise?

This set contains 22 questions. Each one shows immediate feedback and a detailed explanation after you answer, so you learn the correct usage right away rather than waiting for a final score.