Data Management for IR

The lecture synthesizes the principles and practices of research data management (RDM) as foundations for transparent, reproducible, and sustainable science. It argues that data is the indispensable substrate of scientific reasoning and shows how organizational structures and technical infrastructures jointly enable robust workflows. Core topics include the intertwined research and data life cycles; the role of persistent identifiers (e.g., DOIs via CrossRef/DataCite) in ensuring citability, traceability, and long-term access; and practical RDM methods such as consistent folder structures, naming conventions, version control, and thorough documentation (README files, lab notebooks). The lecture surveys licensing choices for data and software. Reproducibility challenges, especially in machine learning and information retrieval, are addressed through strategies ranging from containerized “packaged” environments to transparent white-box workflows, alongside community initiatives like ACM Artifact Review and Badging