Speaker
Description
Measuring the impact of a scientific mission depends on knowing when its data are used, by whom, and how often. For decades, DOIs have been used to track citations of authors' publications; a similar mechanism can make an observatory or instrument a citable resource that serves as a connection to the various outputs it produces. SPASE's observatory and instrument records could support this, but minting a DOI requires an author list, which is notably lacking in the records and difficult to determine manually. This bottleneck, among others, stands between the community and DOI-based tracking of mission impact.
We present an LLM-based agent that automates the discovery of candidate authors for existing SPASE instrument and observatory records. Its core is a curation procedure derived from manual practice and refined through community feedback. It searches sources in a prioritized order to extract information: the SPASE record's own contacts, Calibration and Measurement Algorithms Documents (CMADs), sibling instrument and observatory records, data-provider sites, and the published literature. Keywords, publication timing, author count, and citation heuristics identify the mission and instrument description papers, and author order is then used to extract potential authors. It then applies an evidence-strength hierarchy that distinguishes authors from incidental mentions, producing a candidate author list with per-source provenance links and confidence levels for human review.
We describe the agent's architecture, its performance against manually validated test cases, and the central challenge of trust in AI-generated metadata. We close by outlining future directions, including extending this approach to other metadata fields in SPASE observatory, instrument, and dataset records to support impact measurement, citation, and discovery.