Speaker
Description
Scientific databases increasingly need to support more than data access. They also need to help researchers interpret measurements, understand prior uses of the data, reproduce published analyses, and identify new ways to reuse archived observations. This talk describes ongoing work around the Madrigal database aimed at building this broader layer of support. I will first describe recent progress on AI assisted data discovery, including tools for retrieving documentation, file header information, experiment metadata, API information, and synthesized records from papers that used Madrigal data. I will then discuss ongoing work to make these responses more precise through structured metadata queries, standardized access paths, and closer links between data products, variables, instruments, and scientific context.
The second part of the talk will focus on extending Madrigal from data discovery toward scientific use discovery. I will describe early efforts to represent published papers by their datasets, phenomena, methods, workflows, figures, and scientific claims, and to use these representations to identify related studies, useful case studies, methodological clusters, and possible extensions to archived data. I will also discuss plans for agent based workflow reconstruction, where selected papers are used to generate inspectable records of data access steps, code, reproduced outputs, failure modes, and training material. Madrigal provides a useful testbed because of its long history, diverse instruments, rich metadata, APIs, documentation, and large literature base. The broader goal is to develop a reusable repository intelligence layer that helps scientific databases expose what they contain, what has been done with their holdings, and what might be done next.