Speakers
Description
In order to link papers of heliophysics to their dataset, we need to retrieve certain informations in the text, such as the name of the measuring instrument, the time range of observation that is studied in the paper, as well as some other observational parameters. Moreover, in the objective of reproducing figures in papers for detecting errors in datasets, we aim at retrieving informations regarding the data transformation process, such as formulas, variable types, software or models used, and results that may be displayed in figures, tables or presented in the text.
To do so, we introduce an annotation schema and an LLM-based strategy to pre-annotate heliophysics papers in a zero-shot environment. We are currently reviewing those pre-annotations with the help of experts to create a high-quality training dataset for the Named Entity Recognition task.