IR Evaluation
This lecture introduces the principles and methods used to evaluate Information Retrieval systems. It covers the Cranfield evaluation paradigm, test collections, relevance judgements, effectiveness measures such as precision, recall, MAP, and nDCG, and statistical significance testing. The lecture also highlights online evaluation in contrast to traditional offline evaluation. Emphasis is placed on experimentation as the foundation of modern IR research, and on how retrieval systems are compared and improved through systematic evaluation. Examples using the PyTerrier platform illustrate how modern retrieval experiments are conducted in practice. The lecture concludes with a brief overview of ongoing topics in IR evaluation, such as LLM-as-a-judge.