A hallucination is output that is fluent and confident but ungrounded — fabricated facts, invented citations, nonexistent API methods. Because LLMs are trained to produce probable next tokens, not to verify truth, they can generate convincing falsehoods. Evaluating and reducing hallucination is central to production LLM work. Techniques include retrieval-augmented generation (grounding answers in retrieved documents), measuring faithfulness (does the answer stay true to the provided context?), and citation requirements. Hallucination is especially dangerous in high-stakes domains like medicine or law.
2 / 5
What is an "eval set" (evaluation dataset) and why is it essential?
An eval set is a curated dataset of inputs — paired with reference answers or grading criteria — that you run your model/prompt against to measure quality objectively. It plays the role unit tests play in software: without it, prompt or model changes are guesswork ("it seems better"). A good eval set covers representative cases, edge cases, and known failure modes. Running it on every change catches regressions (a prompt tweak that fixes one case but breaks five others). Eval-driven development treats prompts and model choices as hypotheses validated against the eval set.
3 / 5
What is the "LLM-as-judge" evaluation technique?
LLM-as-judge uses a strong model (the judge) to score outputs against criteria — relevance, correctness, tone, faithfulness — replacing or augmenting expensive human evaluation. You give the judge the input, the output, and a rubric, and it returns a score and rationale. This scales evaluation to thousands of cases cheaply. Caveats: judges have biases (favoring longer answers, their own style, position bias), so practitioners calibrate the judge against human labels, use pairwise comparison instead of absolute scoring, and validate that the judge correlates with human judgment before trusting it.
4 / 5
In the RAGAS framework for evaluating RAG systems, what does "faithfulness" measure?
Faithfulness measures whether the answer is grounded in the retrieved documents — every claim in the answer should be supported by the context, with no fabrication. It is distinct from answer relevancy (does the answer address the question?) and context relevancy/precision (did retrieval fetch the right documents?). RAGAS decomposes RAG quality into these complementary metrics because a RAG system can fail in different ways: retrieving wrong docs, retrieving right docs but ignoring them, or answering off-topic. Faithfulness specifically targets hallucination relative to the provided context.
5 / 5
What is a "benchmark" like MMLU or HumanEval used for?
A benchmark is a standardized, public dataset and scoring method for comparing models on a capability. MMLU (Massive Multitask Language Understanding) tests broad academic knowledge across 57 subjects via multiple-choice questions. HumanEval measures functional code generation — does the produced code pass hidden unit tests? Benchmarks enable leaderboards and standardized comparison, but they have limits: contamination (benchmark data leaking into training), saturation (top models all near 100%), and gaming. They measure narrow proxies, so production systems still need task-specific eval sets reflecting real use.
What does the "LLM Evaluation" vocabulary exercise cover?
This exercise tests real IT vocabulary related to llm evaluation through 5 multiple-choice questions, each built from realistic workplace sentences rather than abstract definitions.
Is this vocabulary exercise free to use?
Yes. Every exercise on CoderSlingo, including this one, is completely free — no account, sign-up, or payment required.
How many questions does this exercise have?
This exercise has 5 questions. Each one shows a real-world sentence or scenario with multiple-choice options and an explanation once you answer.
What happens after I answer a question?
You'll see immediate feedback showing whether your answer was correct, along with a short explanation of why — then a button to move to the next question, and a full results screen at the end.
Can I retry the exercise if I get questions wrong?
Yes. Once you reach the results screen, click "Try again" to reset your answers and go through the exercise from the start as many times as you like.
Do I need to create an account to take this exercise?
No account is needed. Your answers are scored in your browser during the session — nothing is saved to a server, so you can jump straight in.
Is my progress saved if I leave the page?
No — progress within an exercise resets if you navigate away or reload. Each exercise is short enough to complete in a few minutes in one sitting.
Are these vocabulary exercises connected to other topics?
Yes — this module shares real-world context with 11 other vocabulary modules. See "Related vocabulary" below to keep building a connected skill set.
How is this different from reading a glossary or blog article?
Exercises like this one are active recall drills — you have to choose the correct term or phrasing yourself, which builds retention faster than passively reading a definition.
Where can I find more vocabulary exercises?
Browse the full Vocabulary exercises hub for hundreds of modules covering Agile, DevOps, security, databases, architecture, and more — organised by IT role and skill.