DASH/IHDEA 2026

Europe/Dublin
Royal Irish Academy of Music, 36-38 Westland Row, Dublin D02 WY89
Description

 

Registration has been extended to end of Friday 11 September - this is the final extension!

The 2026 DASH / IHDEA Meeting will take place at the Royal Irish Academy of Music in Dublin, Ireland, from 5–9 October 2026, hosted by the Dublin Institute for Advanced Studies (DIAS), with a welcome reception on Sunday, 4 October. The DASH meeting will take place 5-7 October, followed immediately by the IHDEA meeting on 8-9 October.

The Data Analysis and Software in Heliophysics (DASH) workshop brings together scientific software and data practitioners in community discussions around topics of interest as proposed by community members. The goal of the International Heliophysics Data Environment Alliance (IHDEA) is to encourage the use of common standards and services in order to enable sharing of data and to enhance science.

Thank you to everyone who submitted DASH and IHDEA session proposals and abstracts, we received a fantastic range of ideas. The schedule for both the DASH and IHDEA meetings are now available on the Timetable page!

As always, we are seeking an interactive meeting, with plenty of time for discussion and networking. Informal working group discussions may also take place on Sunday 4 October, with details to be confirmed closer to the meeting.

See more here too for past DASH's : https://dash.heliophysics.net/

    • 17:00 19:00
      Welcome Reception 2h

      https://indico.dias.ie/event/2/page/23-welcome-reception

    • 08:30 09:00
      Registration 30m
    • 09:00 09:10
      Welcome 10m
    • 09:10 09:40
      Keynotes: Keynote 1
      • 09:10
        Keynote 1 30m
        Speaker: Shane Maloney (DIAS)
    • 09:40 10:40
      Session 1: General Session 1
      • 09:40
        SSTrack: An Automatic Sunspot Identification and Tracking Algorithm to Support the Measurement of Sunspot Rotation 15m

        Motions of sunspots cause the coronal magnetic field to become deformed with the effect that energy can be stored in the field. This energy can be released through events such as solar flares. Sunspots are known to rotate about their umbral centres, and this rotation contributes to the build-up of energy within an active region that can be released during an eruptive solar event. To understand the relationship between rotational forms of sunspot dynamics and solar activity, a large, unbiased statistical survey of sunspots is desirable.
        To generate such a statistical sample, a fully automatic sunspot identification and tracking method, SSTrack, is developed. This method is tested on a previously analysed four-month sample of active regions generated using a semi-automatic sunspot rotation tool. The new method is designed to be applied to long periods of observations so that large samples can be efficiently obtained, while working at a high-cadence to effectively capture fine-scale-behaviour, such as sunspot splitting and mergers. SSTrack is able to identify fifty-four of the fifty-six sunspots in the four-month sample as well as an additional forty-three sunspots not found by the semi-automatic method. It is able to identify many of the fifty-four commonly found sunspots earlier and track them for longer than the semi-automatic method. The rotation about umbral centres is calculated for each sunspot using the tracking data from both methods, and when considering only overlapping observations per sunspot the methods show good agreement. The semi-automatic method is found to detect greater rotation when a sunspot undergoes fragmentation due to the changing structure overly influencing the determination of the centre of the sunspot.
        SSTrack has since been applied to a two-year period of SDO continuum data leading up to solar maximum which has produced a sample containing 1246 sunspots in 390 active regions.

        Speaker: Charlotte Proverbs (Jeremiah Horrocks Institute, University of Lancashire)
      • 09:55
        Principal Component Analysis for In-Production Coronagraphic Calibration 15m

        The PUNCH (Polarimeter to UNify the Corona and Heliosphere) mission has three heliographic Wide-Field Imagers (WFIs) and one coronagraphic Narrow-Field Imager (NFI). Soon after launch, it was discovered that NFI has a very strong and unexpected dynamic stray light component. Driven by Earth-shine that has found a route through the optical system, this stray light is by far the dominant source of measured light. Showing a strong dependence on both the land mass and cloud cover under the satellite, this stray light varies significantly from image to image. This defeats the usual methods of removing stray light (e.g., running minimum subtraction) or ignoring stray light (e.g., difference imaging). One year after launch, the PUNCH team is nearing completion of the calibration pipeline for WFI, but NFI remains largely un-calibrated and un-used due to the difficulty of addressing this stray light. Early efforts showed that principal component analysis (PCA) is able to fit and remove the dynamic stray light (as well as the F corona) remarkably well, revealing that clear coronal signals are present and recoverable. This technique was automated and deployed in the QuickPUNCH pipeline (a low-latency pipeline intended for space weather forecasting), but the quality of the subtraction fell far below that of the initial proof of concept. Recently, the PCA approach has received renewed effort, with the goal of producing a reliable method for automated pipeline deployment. While not deep learning, PCA is a type of machine learning, and machine learning has not been commonly used for this sort of fundamental calibration in a production pipeline by corona- or heliographic missions. This presentation will present the method, show the results, discuss the various complications encountered, and highlight steps we've included to ensure we do not fit and remove the coronal signal of interest.

        Speaker: Sam Van Kooten
      • 10:10
        Characterizing Aviation Radiation Exposure Using Machine Learning and In-Flight Cosmic-Ray Muon Measurements 15m

        Cumulative exposure to ionizing radiation at aviation altitudes poses significant health risks for aircrews and, at higher altitudes, astronauts. Physics-based models are commonly used to estimate radiation levels during flight; however, they often do not fully capture the rapidly varying and complex nature of atmospheric radiation, limiting real-time prediction accuracy.To address this limitation, we explore machine learning (ML) approaches to improve the analysis and nowcasting of aviation radiation.
        Using newly compiled, ML-ready aviation radiation datasets, we train supervised ML models to identify nonlinear relationships between geospace environmental parameters and measured radiation effective dose rates. Our results show that a gradient boosting (XGBoost) model trained on the concurrent properties of the geospace environment improves radiation prediction accuracy by ~9% compared to the considered physics-based NAIRAS-v3 model. Feature importance analysis and Shapley Additive Explanations (SHAP) indicate key geospace parameters, including solar wind and solar polar fields, play a dominant role in controlling radiation variability at flight altitudes.
        In a complementary observational study, we examine the role of secondary cosmic-ray muons in aviation radiation environments at altitudes below 15 km. Atmospheric muon flux measurements obtained from a CubeSat prototype developed by the Nuclear Physics Group at Georgia State University are analyzed alongside radiation doses modeled by NAIRAS-v3. Correlation analysis demonstrates a strong, statistically significant positive relationship between measured muon counts per minute and modeled radiation dose rates (µSv/h), with a Pearson correlation coefficient of r = 0.93.

        Speakers: Mr Sanjib K C (Georgia State University, USA), Dr Viacheslav Sadykov (Georgia State University, USA)
      • 10:25
        Scientists need help selecting data-model comparison metrics: please embed this guidance in the analysis software 15m

        Comparing two number sets is a foundational component of the scientific process. We should not assume, however, that scientists know the best approach to conducting an appropriate data-model comparison for their particular need. There are many ways to compare two number sets, though, and it can be a confusing task to choose the right process for a particular situation. For example, certain metrics work well only when the distribution of data-model differences is Gaussian, such as root mean square error, correlation coefficient, and mean error. Other metrics work better when the number set distributions, or the distribution of their differences, is non-Gaussian. In addition, each metric only assesses a specific aspect of the data-model relationship. A complementary set of metrics should be employed to conduct a robust comparison. Depending on the scientist's objective in doing the comparison, some metrics could be far more important than others. This knowledge of the strengths and limitations of each metric, and the right combination of metrics for certain types of studies, is, not widely known and applied in scientific research. How many scientists actually work is that they use the metrics they already know, and often do not consider using other, potentially more appropriate ones. Scientists need guidance about the many metrics available to them. It is advocated that this guidance should be embedded into the analysis software.

        Speaker: Michael Liemohn (University of Michigan)
    • 10:40 11:00
      Coffee 20m
    • 11:00 12:30
      Session 2: AI Session 1 (Discovery & Metadata)
      • 11:00
        Linking Heliophysics Publications to Mission Data 15m

        Heliophysics papers frequently reference mission data informally rather than through formal citations, leaving data usage often untracked. We have created a pipeline that uses large language models to automatically extract data references from heliophysics papers. These data references consist of missions, instruments, and observed time ranges. For instance, a paper may describe using SOHO/LASCO observations to research CMEs, but not provide any clear citations. In this case, our system would identify that the paper uses the SOHO mission, LASCO instrument, and a time range used in the paper. The system has been manually validated against 5,000+ claims, achieving an aggregate precision of 99.88%. While initial extraction relied on proprietary models (GPT-5.X), our benchmarking found that open-source alternatives (gpt-oss-120B) perform comparably on this task. This has allowed us to scale to over 40,000 papers at one-tenth the cost. The extracted metadata has also been integrated into mission landing pages on the HelioData site, allowing for a high-level overview of the publications that use a mission’s data.

        Speakers: Aidan Scharnikow, Mr Anthony Buonomo
      • 11:15
        Toward DOI-Based Mission Impact Tracking: An LLM Agent for SPASE Authorship Metadata 15m

        Measuring the impact of a scientific mission depends on knowing when its data are used, by whom, and how often. For decades, DOIs have been used to track citations of authors' publications; a similar mechanism can make an observatory or instrument a citable resource that serves as a connection to the various outputs it produces. SPASE's observatory and instrument records could support this, but minting a DOI requires an author list, which is notably lacking in the records and difficult to determine manually. This bottleneck, among others, stands between the community and DOI-based tracking of mission impact.

        We present an LLM-based agent that automates the discovery of candidate authors for existing SPASE instrument and observatory records. Its core is a curation procedure derived from manual practice and refined through community feedback. It searches sources in a prioritized order to extract information: the SPASE record's own contacts, Calibration and Measurement Algorithms Documents (CMADs), sibling instrument and observatory records, data-provider sites, and the published literature. Keywords, publication timing, author count, and citation heuristics identify the mission and instrument description papers, and author order is then used to extract potential authors. It then applies an evidence-strength hierarchy that distinguishes authors from incidental mentions, producing a candidate author list with per-source provenance links and confidence levels for human review.

        We describe the agent's architecture, its performance against manually validated test cases, and the central challenge of trust in AI-generated metadata. We close by outlining future directions, including extending this approach to other metadata fields in SPASE observatory, instrument, and dataset records to support impact measurement, citation, and discovery.

        Speaker: Dr Disha Sardana (Heliophysics Data and Modeling Consortium)
      • 11:30
        An LLM-Agent Framework for Intelligent Scientific Data Discovery and Analysis at the National Space Science Data Center 15m

        The rapid growth of space science missions has led to an unprecedented increase in scientific data, engineering documents, metadata, and software resources. Efficiently organizing these heterogeneous resources and enabling intelligent access have become critical challenges for modern scientific research. At the National Space Science Data Center (NSSDC), Chinese Academy of Science, we are developing an AI-agent framework to support intelligent scientific data management, retrieval, and analysis for space science applications.

        Our framework adopts a multi-agent architecture in which specialized agents collaborate to accomplish complex scientific tasks through planning, reasoning, and tool invocation. Existing scientific data processing and analysis software have been encapsulated as Model Context Protocol (MCP) services, enabling LLM agents to directly discover and invoke domain-specific tools without requiring customized integrations. This service-oriented architecture significantly improves the flexibility and scalability of AI-assisted scientific workflows while allowing legacy scientific software to be seamlessly incorporated into agent-based systems.

        To enhance knowledge acquisition, we have also developed a Retrieval-Augmented Generation (RAG) framework that integrates heterogeneous scientific resources, including mission documents, metadata, technical reports, scientific publications, and archived datasets. Beyond conventional text retrieval, the system supports multimodal knowledge retrieval from both textual descriptions and scientific images, providing more accurate and context-aware responses to researchers. The combination of semantic retrieval and domain knowledge substantially improves the efficiency of scientific information discovery and reduces the effort required to locate relevant datasets and documentation.

        The integration of multi-agent collaboration, MCP-enabled scientific software, and multimodal RAG establishes a unified intelligent research environment for space science.

        Speaker: Fuli Ma (National Space science center, Chinese Academy of Science)
      • 11:45
        From Data Discovery to Idea Discovery: Papers, Workflows, and Agents as a New Layer for Madrigal 15m

        Scientific databases increasingly need to support more than data access. They also need to help researchers interpret measurements, understand prior uses of the data, reproduce published analyses, and identify new ways to reuse archived observations. This talk describes ongoing work around the Madrigal database aimed at building this broader layer of support. I will first describe recent progress on AI assisted data discovery, including tools for retrieving documentation, file header information, experiment metadata, API information, and synthesized records from papers that used Madrigal data. I will then discuss ongoing work to make these responses more precise through structured metadata queries, standardized access paths, and closer links between data products, variables, instruments, and scientific context.

        The second part of the talk will focus on extending Madrigal from data discovery toward scientific use discovery. I will describe early efforts to represent published papers by their datasets, phenomena, methods, workflows, figures, and scientific claims, and to use these representations to identify related studies, useful case studies, methodological clusters, and possible extensions to archived data. I will also discuss plans for agent based workflow reconstruction, where selected papers are used to generate inspectable records of data access steps, code, reproduced outputs, failure modes, and training material. Madrigal provides a useful testbed because of its long history, diverse instruments, rich metadata, APIs, documentation, and large literature base. The broader goal is to develop a reusable repository intelligence layer that helps scientific databases expose what they contain, what has been done with their holdings, and what might be done next.

        Speaker: Enrique Rojas Villalba (Massachusetts Institute of Technology)
      • 12:00
        Semantic Applications of Madrigal Data 15m

        Since 1960, MIT Haystack Observatory has supported atmospheric and geospace research through long running observational programs and the Madrigal distributed database, which provides access to incoherent scatter radar, total electron content, Fabry Perot interferometer, and other datasets. These resources have contributed to a large scientific literature, but identifying relationships among datasets, methods, phenomena, and conclusions across decades of publications remains difficult.

        We analyzed a corpus of 951 papers published between 1960 and 2026 that used data available through Madrigal. We developed an interactive application that combines large language models, text embeddings, knowledge graphs, and graph based retrieval augmented generation to catalog and visualize this literature. Each paper is represented by a paper level embedding, while structured representations capture the Madrigal datasets used, scientific methods, phenomena studied, and principal conclusions. Full text was extracted with PyMuPDF and divided into smaller passages to support more detailed concept extraction. Recursive clustering organizes the resulting representations at multiple levels, allowing users to move between broad research areas and narrower groups of related studies.

        Researchers can use the application to locate studies that use similar datasets or methods, trace how a topic has developed over time, identify representative case studies, and find methods applied to related problems in other research areas. The interface turns a static publication catalog into an explorable semantic map in which papers can be reorganized according to different scientific criteria.

        The talk will describe the application and present results from task based evaluations with students, researchers, and domain experts, including how their feedback informed revisions to the clustering criteria and interface. Although Madrigal provides the initial testbed, the open source software and workflow can be adapted to other scientific archives and document collections.

        Speaker: Ryan Cavanaugh (MIT Haystack Observatory)
      • 12:15
        LLM-Assisted Instrument Discovery in the Virtual Solar Observatory 15m

        The Virtual Solar Observatory (VSO) indexes instruments spanning X-ray through radio wavelengths, yet discovery remains largely keyword-dependent — requiring users to know provider names and instrument codes in advance. We present a prototype semantic search layer built on top of VSO that replaces exact-match lookup with embedding-based similarity search, enabling researchers to find relevant instruments through natural language queries and conceptual proximity rather than precise terminology.

        Each instrument record was encoded into a high-dimensional vector using a pre-trained large language model, with cosine similarity used to rank the most semantically related instruments for any given query. A Principal Component Analysis projection reveals natural clustering by wavelength band — with EUV, UV, X-ray, and radio instruments forming visually distinct groupings — confirming that the embedding space captures physically meaningful structure without VSO-specific supervision.

        Early results demonstrate that the model surfaces cross-mission instrument equivalents that keyword search would miss entirely, such as associating SDO/AIA EUV channels with complementary observations from STEREO/EUVI and Proba-2/SWAP. This work lays the groundwork for a conversational data discovery interface where heliophysicists can describe the data they need in scientific terms and receive ranked, cross-archive instrument recommendations — complementing the structured API and TAP access capabilities being developed in VSO 2.0.

        Speaker: Kaushal Patel (NASA/Columbus Technologies)
    • 12:30 14:00
      Lunch 1h 30m
    • 14:00 15:30
      Session 3: Ground-Based Heliophysics Data: Distribution, Discovery, and Science Enablement
      • 14:00
        The Whole Is Greater Than the Sum of Its Instruments: Synergizing Space- and Ground-Based Observations in Heliophysics 15m

        Addressing many of the most important questions in heliophysics requires observations that span the coupled Sun–Earth system. No single instrument or platform can fully capture the complexity of this interconnected environment; instead, scientific progress depends on integrating complementary observations from both space- and ground-based assets. While the heliophysics community has pursued this integration for decades, often through ad hoc collaborations, continued progress relies on sustained cooperation among data providers, mission teams, and instrument operators.
        Realizing the full scientific potential of these diverse observations requires accessible data products and analysis tools. Rapid access to quick-look visualizations enables researchers to efficiently identify intervals and events of interest, while standardized tools for exploration and analysis lower barriers to discovery. These capabilities are particularly important for engaging students and early-career researchers, helping to broaden participation in the field.
        In this talk, we present a compelling science case in which the community has combined observations from nearly every available instrument—and continues to incorporate new datasets—to advance our understanding of the coupled geospace environment. We also highlight the growing emphasis among satellite missions on coordinating observations with ground-based instrument networks to maximize scientific return. Although significant challenges remain, including fragmentation driven by funding agency priorities, the scientific value of integrated space- and ground-based observations is unequivocal. With expanding observational capabilities and increasing community coordination, the time is right to strengthen these partnerships and further bridge the gap between space- and ground-based heliophysics.

        Speaker: Bea Gallardo-Lacourt (NASA/CUA)
      • 14:15
        UCalgary Open Science Platform: Building for the Next Decade of Data Growth 15m

        In the era of large networks of heterogeneous sensors, the ability to efficiently use and understand data is paramount. The University of Calgary operates one of the largest ground-based heliophysics sensor networks, with 118 instruments across the globe producing, on average, 16 million data points and 1 million images per day. The group's instrumentation heritage dates to the 1980s, with the largest step change in data volume beginning in 2005 with ground-based all-sky imagers and magnetometers for the THEMIS mission. The archive currently holds >1.2 PB, with expected growth of 100–200 TB per year over the next decade.

        In this presentation, we provide an overview of capabilities developed over the past year for UCalgary's Open Science Platform (data.phys.ucalgary.ca). These include a substantial revamp of our web applications and visualization tools, new instrument deployments bringing additional data streams online, progress on HAPI and SPEDAS integrations to improve cross-network interoperability, and the initial design of a next-generation data system architected for the coming decade of growth. As sensor systems grow more complex (advanced operating modes, streaming data) and data volumes increase, the burden on users to keep pace with evolving distribution and utilization methods grows quickly. What underpins our data systems, and our data itself, is a devotion to simplicity for the user: providing the tools needed to effectively browse, discover, learn, and utilize our data. These updates continue that trajectory and position the platform for the next decade of network growth.

        Speaker: Darren Chaddock (University of Calgary)
      • 14:30
        From 37 Remote Sites to the Research Community: Lessons from Two Decades of Ground-Based Data Distribution 15m

        The University of Calgary now operates one of the largest ground-based heliophysics networks in the world: more than 110 instruments including all-sky imagers, riometers, magnetometers, and GNSS scintillation receivers — distributed across 37 sites in northern Canada, Alaska, Greenland, and Antarctica. Sustaining this network has forced us to solve, and repeatedly re-solve, the full chain of problems that sit between a photon at a remote field site and a figure in a paper: constrained and intermittent connectivity, heterogeneous instrument generations with incompatible data models, calibration provenance that must survive decades, and the tension between rapid quick-look availability and the stability expected of science-grade products.

        This presentation offers an operator's perspective on that chain, using our open data platform and the AuroraX conjunction-search and metadata services as concrete examples of what has worked and what has not. We discuss the decisions that proved durable — separating quick-look from science-grade product tiers, treating metadata as a first-class deliverable, and exposing data through programmatic interfaces alongside browsable archives — and the decisions we would make differently, including the real cost of retrofitting standards compliance onto established archives.

        Speaker: Emma Spanswick
      • 14:45
        Ground-based and space-based radio spectra in SciQLop: interoperability obstacles from the client side 15m

        Ground-based networks and space-based archives are served by largely disjoint client stacks: sunpy, radiospectra and Fido with per-station file conventions on one side; CDAWeb, AMDA and HAPI clients with ISTP metadata on the other. Comparing a type-III burst observed simultaneously by a ground spectrograph and a spacecraft receiver therefore requires manual work for each event. We report on serving both from a single interactive analysis environment, and on the interoperability obstacles this exposed. The work follows from the Lorentz Center workshop Bridging Gaps in Heliospheric Radio Data Analyses (Leiden, May 2026).

        SciQLop is an open-source, cross-platform environment for multi-mission in situ data analysis, with archive access through the Speasy library and an extensible plugin system. Its radio plugin serves 38 curated dynamic-spectrum products from space-based receivers spanning five decades: Wind/WAVES, STEREO-A and -B SWAVES, Solar Orbiter/RPW, PSP/FIELDS RFS, and the planetary receivers of Juno, Cassini, Galileo and Voyager 1 and 2. It serves ground-based e-CALLISTO, I-LOFAR (mode-357 beam-formed) and RSTN through radiospectra's Fido clients. All are exposed as time-windowed products that fetch and cache on demand as the user pans and zooms across a shared time axis. The approach is not specific to radio: an equivalent plugin serves FDSN seismic networks through ObsPy, and a SuperMAG provider being added to Speasy itself exposes that network's roughly 600 ground magnetometer stations as individual products, discoverable by IAGA code and returned in the local geomagnetic or the geographic frame.

        The integration exposed several obstacles of general interest. Ground-based products carry no ISTP-equivalent presentation metadata, requiring a separate plot-hints layer for them to render consistently with archive-served data. EOVSA spectrograms moved behind a registration wall during development, motivating an explicit "listed but not retrievable" product state rather than a silent absence. SuperMAG publishes no per-station temporal coverage, reporting instead which stations were operating over a requested interval, so a client cannot advertise a station's real extent without an extra query. Search attributes never fully identify a channel. A station attribute exists only in the affiliated package, sunpy's generic one having been removed, and the token separating concurrent signal chains at one site, such as an e-CALLISTO focus code, has no attribute at all, so channel identity is reconstructed for every instrument by re-filtering the returned rows on their columns. Getting this wrong is silent: at I-LOFAR, where every timestamp ships one file per linear polarisation, the two channels were merged into a single spectrogram, and no structural check could object because they share an identical frequency grid. I-LOFAR mode-357 files decode into three co-temporal frequency bands requiring reassembly, and carry their start time only in the filename, so a cache that renames files loses it.
        We close by asking what minimal agreed metadata ground-based networks would need to publish for a generic client to consume them as it already consumes CDAWeb and AMDA.

        Speaker: Alexis Jeandet (LPP-CNRS)
      • 15:00
        Future Directions for the HAPI Specification 15m

        HAPI is a widely adopted standard for accessing heliophysics and space weather time-series data. Its success comes from its simplicity: it is straightforward for data providers to implement and easy for client software to use. We present optional extensions to the specification that preserve this simplicity while enabling greater interoperability and automated, semantic interpretation of data. These include (1) a mechanism for expressing relationships among datasets through a new “linkages/” endpoint and (2) schemas that connect HAPI parameters to scientifically meaningful measurement types. The new extensions are designed to remain backward compatible with existing HAPI servers and clients. We will present early examples of these capabilities and discuss how they can support data discovery, comparison, and integration across distributed ground- and space-based archives.

        Speaker: Jon Vandegriff
      • 15:15
        Open Discussion 15m
    • 15:30 16:00
      Coffee 30m
    • 16:00 17:30
      Session 4: General Session 2 (Metadata, SPASE + FAIR)
      • 16:00
        Improving Data and Metadata Analysis in the CCMC Runs-on-Request (ROR) 15m

        Runs-on-Request (ROR) at the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) is a web-based system allowing users to perform customized simulations using over 60 Space Weather and Heliophysical models. The results of each simulation are stored in an interactive web archive, where users can analyze any of more than 43,000 complete simulations. Processing approximately 4000 user requests and generating 1 petabytes of simulation data a year, ROR provides unique opportunities for improving our understanding of space weather and its impacts on space exploration.

        In this presentation, we report on the recent improvements in ROR infrastructure, data pipelines, and user services. We also describe our efforts and lessons learned on interconnecting ROR with other partner systems and data streams.

        Speaker: Jack Topper (NASA Goddard Space Flight Center)
      • 16:15
        From Code Audit to Mission Software: Agentic AI Workflows for Inherited Scientific Repositories 15m

        Research Software Engineers frequently inherit scientific codebases they did not write, in unfamiliar domains, and on schedules that limit deep familiarization. Large Language Model–based coding agents, such as Codex, can accelerate this transition, but their value depends on how they are integrated with scientific verification and software-development practices. This presentation examines the use of agentic AI to assess, extend, and rewrite two scientific Python repositories supporting Europa Clipper mission planning.

        The first repository was an approximately 60,000-line implementation of the Hapke bidirectional reflectance model developed by a domain scientist. An AI-assisted audit traced equations through the implementation, compared the code with published model descriptions, and generated targeted questions for a subject-matter expert. During the initial audit, the agent identified an algebraic error that was independently verified against the scientific literature. The repository otherwise remained largely intact. Here, the agent’s principal value was accelerated code comprehension and preparation for expert review rather than autonomous code generation.

        The second repository, originally a data-ingestion wrapper, required a near-complete rewrite to support synthetic spectral generation, photometric and geometric calculations, and observation-planning workflows. Coding agents accelerated bug identification, implementation, testing, and refactoring. However, rapid AI-assisted development also introduced redundant utilities, inconsistent abstractions, and growing architectural complexity. Structured practices drawn from Git.Ship.Done. were subsequently used to organize work into documented requirements, plans, and executable tasks, improving continuity and making agent behavior easier to review.

        These case studies show that agentic AI can lower the barrier to working with unfamiliar scientific software, improve interactions with domain experts, and accelerate delivery of complex capabilities. They also demonstrate that agents can amplify weak architectural decisions and cannot replace scientific validation, code review, or deliberate software design. The presentation will discuss effective workflows, verification practices, and the continuing challenge of building maintainable and reproducible agent-assisted scientific software.

        Speaker: David Stephens
      • 16:30
        Status and prospective of the MASER ecosystem. 15m

        The study of radio emissions in the low-frequency range (kHz to tens of MHz) provides crucial insights into plasma processes, magnetic field interactions, and particle acceleration mechanisms. However, accessing and analyzing these complex datasets presents significant challenges for the Heliophysics scientific community. The MASER (Measuring, Analyzing & Simulating Emissions in Radio frequencies) portal addresses these challenges by providing a comprehensive suite of tools and data access solutions specifically designed for low-frequency radio astronomy applications.
        Our objective is to facilitate efficient data analysis workflows for researchers by implementing FAIR (Findable, Accessible, Interoperable, Reusable) principles throughout the MASER ecosystem. The portal combines multiple components to support the entire research pipeline: a repository system for data publication and online access with visualization capabilities, a set of python libraries (maser4py): maser-data library (within maser4py) for programmatic access and specialized analysis tools through maser-tools (currently in development). This integrated approach enables researchers to seamlessly transition from data discovery to advanced analysis.

        The maser4py library offers programmatic access to many low-frequency radio datasets while supporting open standards such as VESPA metadata for enhanced discoverability, TAP servers and TFCat catalogs. In an effort of interoperability, maser4py's integration with Astropy, SpacePy and XArray frameworks provides familiar tools for data manipulation and analysis. Significant collaborative efforts are underway to bridge gaps between existing tools and services. Notably, we are working to integrate maser4py within the PyHC environment for enhanced capabilities, and integrate Das2 server technology for efficient access to large-volume data from instruments like NenuFAR and SKA. We are also working on interfacing maser-data's capability to open complex low-frequency radio datasets with Sunpy tools like fido to retrieve those datasets. The interface with the NASA WebGeoCalc API enables automatic ephemeris calculations. While maintaining the codes, current development focuses on advanced radio analysis capabilities including goniopolarimetry and direction finding tools for use in official ground segments, and specialized tools for data providers including CDF/ISTP/PDS4 validation and label generation. The modular design of MASER ensures data storage and tool multiplicity do not create barriers, while enhancing the system robustness and flexibility for diverse research needs, notably in Heliophysics.

        Speaker: Lucas Grosset (LIRA - CNRS - Observatoire de Paris)
      • 16:45
        One API for Solar Events: The Helioviewer Project's events-api and event-tree 15m

        Solar event data lives across many services: for example, HEK features,
        CCMC DONKI flares and CMEs, FlareScoreboard prediction models, RHESSI
        flare lists, and WSA coronal-hole and magnetic-connectivity forecasts —
        each built for its own purpose, with its own format, positional
        descriptors, cadence, and update behavior. To simplify the management and
        display of these datasets on the same Sun in helioviewer.org, we needed
        all of them accessible in the same form. Normalizing data from each
        source comes with its own technical challenges. This includes varying
        descriptions of identity, shape, and coordinates, not just format:
        tracking predictions across reissues when records carry no stable
        identifier; representing event boundaries that are a single polygon in
        one source and many contours in another; and reconciling different
        coordinate frames (Stonyhurst, Carrington, helioprojective) into
        positions on the observed Sun. We present events-api, the Helioviewer
        Project's open-source answer: a service that collects each source on a
        schedule, normalizes everything into one event model with multi-polygon
        boundaries, rotates positions to the requested observation time, and
        serves it all through one REST API. helioviewer.org's event capabilities
        now uses events-api, and the corresponding @helioviewer/event-tree React
        component lets any React application embed the same events with a few
        lines of code. The result is a uniform, queryable event history across
        feature/event sources — one place to ask what happened on the Sun, and
        what was predicted.

        Speaker: Kasim Necdet Percinel (Columbus / NASA Goddard Space Flight Center)
      • 17:00
        Heliophysics Events Knowledgebase: Broadening Community Contributions of Events and Datasets 15m

        The Heliophysics Events Knowledgebase (HEK) provides a system for collecting and presenting heliophysical data based on the International Virtual Observatory Alliance (VOEvent) format and other standards. The HEK incorporates entries from distributed ground and space based solar observatories and event descriptions generated by automated algorithms, detailed data analysis and citizen scientists (while storing contact information and links to websites and papers for attribution purposes). This makes it possible to carry out integrated studies that span the full range of heliophysics and these resources are available to any interested researcher.

        We present improvements to the HEK focused on community contributions to the collection of metadata. New tools are enabling broader research community submissions of both events to the Heliophysics Events Registry (HER) and metadata for data sets to the Heliophysics Coverage Registry (HCR), highlighting the knowledgebase’s transition to an infrastructure model for heliophysics research support beyond the missions (Hinode, SDO, IRIS) that drove its initial development. We will discuss recent additions to our knowledgebase from a combination of the HEK team ingesting event lists from other sources and external use of new Python-based tools for contributing to the HEK. We will also present plans to add additional support for links, references, and inferences both between HEK entries and from HEK entries to other event lists and data repositories. This showcases the growth of the HEK beyond its historical focus on missions associated with the Lockheed Martin Solar & Astrophysics Laboratory to a more open and FAIR platform supporting a broad range of heliophysics data resources.

        Speaker: Dr Neal Hurlburt
      • 17:15
        Modernizing Data Browsing at the Space Physics Data Facility (SPDF): The Next Generation of Coordinated Data Analysis Web (CDAWeb) 15m

        CDAWeb has served as SPDF's flagship data access tool for decades, providing the heliophysics community with powerful visualization and access to over 3000 datasets. However, modern research demands—including large mission constellations, higher resolution data, and interactive exploration of complex multi-dimensional datasets—require a reimagined approach to data discovery and visualization.

        We are developing the next generation of CDAWeb: a modern, scalable data exploration platform that maintains CDAWeb's proven functionality and accessibility while expanding capabilities to meet current and future needs. Built on a modern technology stack and leveraging the Python in Heliophysics ecosystem, this platform will provide enhanced data discovery, advanced plotting capabilities, and improved support for visualizing multi-dimensional data such as 2D slices of plasma distribution functions.

        This presentation will discuss the technical challenges driving this modernization effort, visualization trade studies informing our design decisions, the current architecture and tooling infrastructure, and demonstrate early prototype functionality of the initial frontend and backend implementations. As an early-stage prototype, we welcome community feedback to help shape the future of this critical resource for heliophysics research.

        Speaker: Eric Grimes (Space Physics Data Facility)
    • 09:00 09:10
      Welcome 10m
    • 09:10 09:40
      Keynotes: Keynote 2
    • 09:40 10:40
      Session 5: Shadow Workflows and Instrumentation
      • 09:40
        ndcube: an endeavour towards interoperability and cross-field collaboration in astronomy 15m

        ndcube is a package for facilitating generalised coordinate-aware n-dimensional astronomical data analysis.  It does this by combining data, metadata, and coordinate information into unified data objects which can be used as coordinate-aware arrays.  Its biggest point of difference with its most similar package, xarray, is its support of the World Coordinate System (WCS) framework.  WCS is used throughout astronomy and can describe the coordinate frames of images, timeseries, spectra, polarisation measurements, simulation grids, custom coordinates and more, as well as any arbitrary combination of these.

        Despite WCS's ubiquity, there was no mature Python framework for performing generalised WCS-aware data analysis before ndcube.  Astropy NDData was an essential base upon which ndcube is built, but it provided limited generalised analysis tools. In this vacuum, observation- and coordinate-frame-specific data classes emerged (e.g. sunpy Map for 2-D solar images), each with their own API unable to be applied to other types of observations. This had the effect of silo-ing different subfields of astronomy into their own cliques, each with their own tools for performing the same analysis tasks. This increased friction for those who wanted to engage in multi-instrument studies or cross-field collaborations, as entirely distinct analysis workflows would have to be learned and merged. Moreover, it increased the development and maintenance burden for other packages as a zoo of similar but slightly different data containers had to be developed for each task.

        In this talk, we will discuss how ndcube overcame these challenges and now provides a generalised framework for analysing any WCS-based data.  This has led it to be adopted as a dependency by numerous Python packages, including IRISpy (NASA SMEX), XRTpy and EISpac (Hinode), DKIST user tools, specutils (JWST and beyond), the PUNCH data pipeline (NASA SMEX), and more.  This has substantially reduced their development and maintenance overheads.  We will further discuss the functionalities provided by ndcube and why it might be helpful to a still broader community.  Finally, we will discuss lessons learned from the development of ndcube, and some of the principles that promote interoperability and broader usability of packages across astronomy and heliophysics.

        Speaker: Daniel Ryan
      • 09:55
        The Missing Middle 15m

        Scientific software often grows out of necessity. Researchers build scripts, notebooks, and custom tools to solve immediate problems, but these workflows can become difficult to maintain, reproduce, and support. In this talk, I will discuss how our small software team works with scientists to introduce software engineering practices without disrupting how they work.

        Rather than asking scientists to completely change their workflows, we build around them. We introduced Git and self hosted Gitea through lightweight versioning approaches, including Git submodules, allowing researchers to continue developing their code while we handle integration into larger applications. We have also used AI tools to help translate between scientific code and production software.

        Once integrated, we treat scientific software as an external component that we can support without fully owning. We use containers for challenging environments such as IDL and Fortran on Macs, CI/CD pipelines for deployment, and services such as Celery to connect scientific tools with web applications. This allows us to add error tracking, reporting, and monitoring while helping scientists improve the reliability and accessibility of their software.

        I will share lessons learned from turning informal scientific workflows into sustainable systems while preserving the flexibility that made them useful in the first place.

        Speaker: Umer Salman (Southwest Research Institute)
      • 10:10
        Getting Scientists On Board with Cloud Portals 15m

        Cloud resources already exist that enable doing science on massive datasets and computationally large problems. There is also the pressure for collaboration and replicability, which clouds are already strong with. It requires some adjustment by scientists to learn cloud tools. Early adopters are willing to self-start, but most scientists need a push to spend precious science time in training up. We find two main approaches to tackle this. The (easy) technical solution is to hide the cloud coding while touting its capabilities, and the (hard) people solution is to devote time to workshops and training. JupyterHub cloud portals (like HelioCloud and Helio-Lite) make workflows nearly like the conventional 'on my laptop' experience, PyHC package devs and portals like Heliodata embed cloud APIs, and data portals like Heliodata include code stubs for accessing cloud data. Meanwhile, HelioCloud/Lite provides copious tutorials and videos (oft ignored) with hands-on demos and workshops that have proven effective. Yet currently we are still reliant on external pressures to motivate scientists: that there are science problems only solvable with cloud resources, that the push for stronger replicability will motivate users to adopt better practices. We share our experiences to spark discussion.

        Speaker: Dr Alex Antunes (JHUAPL)
      • 10:25
        Building an Observability Platform at the ESAC Science Data Center (ESDC) 15m

        The ESAC Science Data Centre (ESDC) develops, hosts and operates the science archives for the majority of ESA’s Space Science Missions, covering Astronomy, Planetary Science, Heliophysics and Human Robotic Exploration (HRE). Each archive aims to maximise the scientific exploitation of its datasets while ensuring their efficient long-term preservation.

        A key performance indicator used by decision makers for every archive has always been to report a set of basic usage metrics. Historically it has been difficult to identify and define common usage metrics across all missions due to varying factors. This has led over time to fragmented and inconsistent approaches for capturing statistics, with some legacy archives still relying on manual processes to produce the required reports.

        Likewise, understanding new trends in archive usage or obtaining deeper insights into our operational systems is often time-consuming and requires significant developer effort.

        As the ESDC transitions towards a common multi-mission platform providing services and user interfaces to access data from ESA Space Science missions, it becomes essential to have a homogeneous approach to observability across the ESDC. This will enable us to accurately capture user and system metrics, assess the availability and performance of applications at all levels, facilitate incident investigation, analyse user behaviour, identify operational and capacity risks, and drive the continuous improvement of our services.

        This presentation outlines the vision, architecture and implementation strategy for combining logs, metrics, health checks, synthetic monitoring and usage analytics into a shared, maintainable observability platform.

        Speaker: Mr Jonathan Cook
    • 10:40 11:00
      Coffee 20m
    • 11:00 12:30
      Session 6: Under the Hood: Lessons Learned in Mission Data Systems
      • 11:00
        PUNCH Mission Data System Reusability 15m

        The Polarimeter to UNify the Corona and Heliosphere (PUNCH) is a NASA small explorer mission studying the inner heliosphere via four satellites creating a synthetic observatory. In this talk, we use the standard "Mission Science Data Systems" template to discuss the status of the PUNCH mission. We highlight the science processing code design and functionality of the pipeline automation with the aim of identifying areas where code and design reuse are possible.

        Speaker: J. Marcus Hughes (Southwest Research Institute)
      • 11:15
        A Modular, Multi-Mission Science Data System for the Space Weather Science Operations Center 15m

        The Space Weather Science Operations Center (SWxSOC) is a multi-mission Science Operations Center currently supporting the HERMES, PADRE, REACH, and IMPAX missions. SWxSOC's Science Data System (SDS) is built around reusable, mission-configurable software and infrastructure, allowing new missions to inherit a working, low-cost, open-source pipeline instead of designing, building, and operating one from scratch. The entire hybrid cloud and on-premises environment is defined in Infrastructure as Code (IaC), so identical development and production deployments can be stood up on demand and evolved in version control alongside the mission software.

        The processing pipeline is organized as a chain of small, single-purpose serverless functions that communicate through a shared event queue. Each function is packaged as a container image, and a shared base image combined with a per-mission requirements layer lets the same infrastructure serve very different instruments with only configuration changes.

        Around this core, SWxSOC exposes a set of mission-agnostic building blocks that any mission on the platform can opt into. Per-mission time-series databases provide persistent housekeeping and science data storage, giving each mission a queryable historical record of instrument state and derived science metrics without operating its own database. Shared visualization dashboards sit on top of those stores so operators, instrument scientists, and science teams monitor the pipeline and analyze data products from the same live view. A notification service delivers alerts and processing summaries into mission operations channels. A real-time data streaming and alerting capability produces derived alerts that downstream mission tooling can react to. Standardized archive delivery connects each mission's pipeline to both the Solar Data Analysis Center (SDAC) and the Space Physics Data Facility (SPDF), closing the loop from raw binary data to publicly available science products.Alongside the platform, SWxSOC helps maintain a set of open-source Python packages, including standalone packages for ISTP CDF and SOLARNET FITS metadata templating and validation.

        Taken together, these software and infrastructure building blocks reduce duplicated effort across missions, make it practical to onboard new instruments quickly, and offer a template that other missions can adopt to stand up a modern Science Data System at low cost.

        Speaker: Andrew Robbertz (NASA Goddard Space Flight Center)
      • 11:30
        Lessons Learned from Building Multi-Mission Science Data Systems at the National Space Science Data Center of China 15m

        Mission Science Data Systems (SDS) are essential infrastructures for transforming spacecraft observations into accessible and usable scientific resources. However, developing and operating SDS for different space missions often requires addressing diverse mission objectives, payload configurations, data formats, processing pipelines, and operational constraints. This frequently results in mission-specific solutions, duplicated software development, and challenges in long-term maintenance and evolution.

        The National Space Science Data Center (NSSDC) of China has been developing and operating science data systems for multiple space science missions, supporting a broad range of scientific satellite programs in China. These systems cover the complete data lifecycle, including telemetry data ingestion, scientific data processing, product generation, quality control, metadata management, archival storage, and user-oriented data services. In addition to domestic space science missions, NSSDC has also supported international collaborative missions, including the SVOM mission jointly developed with the French National Centre for Space Studies (CNES) and the Solar wind Magnetosphere Ionosphere Link Explorer (SMILE) mission jointly developed between China and Europe.

        Through the development and operation of these multi-mission SDS, we have accumulated practical experience in designing reusable architectures, integrating heterogeneous scientific instruments, managing diverse data products, and establishing interoperable data services. We will present lessons learned from different mission scenarios, including challenges associated with mission-specific customization, software reuse, system scalability, international collaboration, and long-term operational sustainability.

        Speaker: Fuli Ma (National Space science center, Chinese Academy of Science)
      • 11:45
        PRIZM in the EZIE Science Data System: Testable L0A-L3 Product Generation and Validation 15m

        The Electrojet Zeeman Imaging Explorer (EZIE) uses polarized microwave measurements of molecular oxygen emission to infer magnetic perturbations associated with ionospheric currents. This contribution describes PRIZM—Polarized Radiative Transfer and Inversion for Zeeman Magnetic Imaging—the scientific-processing software developed to support and validate the connected EZIE product chain from L0A through L3. PRIZM carries observation geometry, ancillary geophysical models, processing configuration, and provenance across product levels while producing structured NetCDF outputs at each stage.

        Within this chain, L0A establishes the observation, spacecraft, and geometric context used by downstream processing. L0B represents the radiance and spectral-product layer, used for calibrating downlinked instrument data. L1 provides the calibrated science measurements used for inversion. L2 contains retrieved magnetic-field perturbations, associated uncertainties, and fit diagnostics, while L3 derives electrojet and current-related products from the retrieved magnetic information.

        The software integrates polarized radiative-transfer calculations, magnetic-field inversion, spacecraft geometry, ancillary models, product generation, and metadata handling within an installable Python package. A central design objective is to make both individual components and interfaces between product levels testable. Validation includes unit tests, synthetic cases with known injected perturbations, canonical comparisons with an independent radiative-transfer implementation, and field-level comparisons between generated and reference NetCDF products. These product-level comparisons reveal discrepancies in retrieved quantities, uncertainty estimates, metadata, coordinate conventions, and data organization that may not be identified through isolated algorithm tests.

        PRIZM also represents a transition from heritage observing-system simulation and research workflows toward reusable mission software. Key lessons include the need to validate interfaces as well as algorithms, preserve configuration and provenance across the full product chain, compare complete products rather than only numerical kernels, and retain executable reference cases during modernization. These practices provide a foundation for algorithm refinement, reprocessing, and scientific review, and may benefit other mission teams developing maintainable and traceable science-data systems.

        Speaker: David Stephens
      • 12:00
        Developing the Solar Orbiter Archive within the ESA Heliophysics Archive 15m

        Solar Orbiter is currently the most comprehensive ESA-led heliophysics space mission to observe the Sun and the heliosphere. The Solar Orbiter ARchive (SOAR) contains more than 5 million unique science files measured by ten instruments aboard the spacecraft. These data include high-resolution remote-sensing observations - such as the closest EUV images of the Sun ever taken and the world-first observations of the solar poles. In addition, the data contains state-of-the-art in-situ measurements of solar particles and the electromagnetic field at the spacecraft site - which allows to link phenomena on the Sun with their impact on the inner solar system.
        Over the last two years, the heliophysics archive team (composed of software developers and scientists) at ESA has developed a new Graphical User Interface (GUI) for the Solar Orbiter Archive within the multi-mission HelioPhysics Archive (HPA). This new HPA/SOAR GUI contains a comprehensive set of functionalities for the solar science community to search and explore the mission data. With the recent public release of HPA version 1.0 in June 2026, these functionalities have become now fully operational and include 1) advanced data search capabilities for searching by observation campaign, solar distance, or near real-time (“low-latency”) measurements of Solar Orbiter, 2) a new intuitive way of displaying search results in a graphical tree-view - organised by instrument, sensor, and data products, 3) data quick-look functionality and direct transfer of data products to external visualisation tools.
        These features have been developed with an agile methodology - including frequent testing and feedback from the targeted user community and mission stakeholders. We present the key features of the operational HPA/SOAR - and reflect on main challenges and best practices during the development process.

        Speaker: Nils Janitzek (ESA)
      • 12:15
        Data Pipelines using Prefect for SOLO SIS, LRO LAMP, IMAP CoDICE, and TRACERS ACI. 15m

        We have developed data pipelines using Prefect for Solar Orbiter SIS, LRO LAMP, IMAP CoDICE, and TRACERS ACI instruments.

        Prefect is an orchestrator for doing data pipelines, typically in the AI/ML realm, and it turns out to a huge benefit to doing our data pipelines for space science giving us observability and metrics. Although it is not perfect, there are some reasons why it can be advantageous to teams developing a pipeline when starting a project. Even more interesting is the case when a pipeline is fully developed and one wants to add operations to the pipeline. This talk or poster will talk about both types: for IMAP and TRACERS, the pipelines were developed from scratch and for LRO and SOLO, they were migrated to this method. We will discuss the pros and cons of this approach compared to other software including regular script files and other orchestrators such as Dagster and Airflow.

        Speaker: Joey Mukherjee
    • 12:30 14:00
      Lunch 1h 30m
    • 14:00 15:30
      Lightning Talks and Poster Lightning Talks
      • 14:00
        Making an Annotated and Curated Corpus of Heliophysics Literature for Named Entity Recognition 5m

        In order to link papers of heliophysics to their dataset, we need to retrieve certain informations in the text, such as the name of the measuring instrument, the time range of observation that is studied in the paper, as well as some other observational parameters. Moreover, in the objective of reproducing figures in papers for detecting errors in datasets, we aim at retrieving informations regarding the data transformation process, such as formulas, variable types, software or models used, and results that may be displayed in figures, tables or presented in the text.

        To do so, we introduce an annotation schema and an LLM-based strategy to pre-annotate heliophysics papers in a zero-shot environment. We are currently reviewing those pre-annotations with the help of experts to create a high-quality training dataset for the Named Entity Recognition task.

        Speakers: Mr Baptiste Cecconi (Paris Observatory - LIRA), Liza Fretel (Paris Observatory - LIRA)
      • 14:05
        A Real-Time, Multi-Source Alert and Decision-Support System for the FOXSI-5 Solar Flare Sounding Rocket Campaign 5m

        Triggered sounding rocket launches impose an unusual set of software requirements on heliophysics data systems: a decision must be made within minutes, from heterogeneous low-latency streams, with no opportunity to re-run the analysis. The first NASA sounding rocket solar flare campaign (FOXSI-4 and Hi-C Flare, 2024 April 17) showed that a purpose-built real-time alert system can deliver flare-optimized observations inside a 5 to 7 minute window (Vievering et al. 2026). We present the enhanced framework deployed for the FOXSI-5 campaign at the Poker Flat Research Range in spring 2026, and we discuss it as a software and data-environment case study.

        The system federates several independently developed tools behind a single operator display. ELSA (Early Large Solar flare Alert) applies launch trigger criteria derived from a parameter search over more than 10,000 historical GOES XRS flares, together with the Flare Anticipation Index. X-TOFF (X-ray Time of Flare Forecast) serves random forest predictions of remaining flare duration from streaming GOES data. WAFFLE (WKU Advanced Flare Forecasting aLgorithm tEam) produces near-real-time SDO/AIA high-temperature emission measure maps for spatial localization. OLAF (Online LASP Application for Flares) supplies a low-latency GOES proxy and short-term trend prediction from SDO/EVE/ESP. Advance target selection draws on the Kusano κ-scheme applied to SDO/HMI SHARP data.

        We describe the resulting architecture: how services distributed across five institutions were integrated, how data latency and gaps were handled, how model output was rendered for a human decision-maker under time pressure, and how the system was validated and operated on site at a remote range. We report campaign performance, discuss failure modes and lessons learned, and outline what would be needed to generalize this framework, through common interfaces and shared data standards, into a reusable resource for coordinated multi-instrument observations and operational space weather monitoring.

        Speaker: Milo Buitrago-Casas (University of California Berkeleu)
      • 14:10
        Open-Source Software for Plasma Emission Simulations and Diffusive Equilibrium Field line Distributions in the Io Plasma Torus 5m

        The torus is composed predominantly of sulfur and oxygen ions supplied by an extended atomic neutral cloud ultimately derived from $SO{_2}$ in Io's atmosphere. The major species of the torus are $S^{+}$, $S^{++}$, $S^{+++}$, $O^{+}$, and $O^{++}$. The emission is diagnostic of plasma conditions, and the majority of the emission is in the UV where a terawatt of energy is emitted. The emission is highly variable with different plasma distributions, 3D models, and viewing geometries. We provide an open-source software to produce thermal and non thermal 3D distributions of plasma and separate open-source software to then use the produced 3D plasma distributions to simulate emission at various wavelengths the resulting plasma would produce. The volume emission rates are integrated over the line of sight to produce emission brightnesses for comparison with observations. Given the viewing geometry and pointing of an instrument any line of sight can be simulated of the torus emission. Software is provided in Python, Fortran, C++, IDL, MATLAB, and Mathematica. CPU parallelized versions in Fortran and C++ using MPI are also provided. Further in single node applications a CPU parallelized Python version using multiprocessing is also provided. The codes can be found at https://doi.org/10.5281/zenodo.17809172 and https://doi.org/10.5281/zenodo.15623974.

        Speaker: Edward Nerney (Dublin Institute for Advanced Studies)
      • 14:15
        The Hassles of Managing Big Data 5m

        The performance and adoption of cloud-provided services has made petabyte-scale datasets from diverse missions increasingly available, yet accessibility without usability diminishes their scientific value. The FAIR data principles—Findable, Accessible, Interoperable, and Reusable—provide a strong framework for evaluating dataset utility. However, an equally critical and often underemphasized dimension is timeliness: the speed at which newly acquired data becomes available to users. A common question from users highlights this challenge: "Why does the archive only extend to a certain date, and when will newer data be available?"

        Addressing this gap requires overcoming significant logistical and technical barriers, particularly when replicating and maintaining over 1.5 petabytes of NASA mission data across heterogeneous archives. We present our experiences with HelioCloud in enhancing data availability timelines from months to days. We discuss the operational strategies, new utilities, challenges and results that highlight this shift while maintaining data integrity and accuracy. Improving data ingestion helps provide research opportunities for the curious minds of our HelioCloud users.

        Speaker: Omar Shalaby (NASA GSFC)
      • 14:20
        The Watchful Eye: Cloud Security, Monitoring, and Emerging Threats in Modern Scientific Infrastructure 5m

        As scientific workloads migrate to the cloud, securing research infrastructure against automated threats, opportunistic cryptominers, agentic bad actors, and even non-agentic LLMs is a growing challenge. This security must be balanced with the accessibility required for reproducible science.

        To address this, we share insights from operating cloud-based scientific platforms, detailing an automated monitoring pipeline that aggregates logs, detects anomalies, and generates daily security reports. These reports enable rapid patching against emerging threats and identify orphaned resources to reduce cloud costs.

        Furthermore, we explore integrating Large Language Models (LLMs) to transform unstructured logs into actionable insights, streamlining threat response. Ultimately, we provide practical, time-saving strategies to help researchers and administrators maintain secure, proactive environments without distracting from their core scientific work.

        Speakers: Omar Shalaby, Mr Jeffery Bradford (NASA GSFC)
      • 14:25
        Poster Lightning Talks 1h 5m
    • 15:30 16:00
      Coffee 30m
    • 16:00 17:30
      Poster Session
      • 16:00
        A New System to Categorize Mission Data at HDRL 1h 30m

        The current system of data processing levels at NASA Heliophysics Division has led to mission-specific definitions across the board, which is indicative of an inflexible system. Researchers therefore must understand each mission’s data processing level definitions to determine the best starting point for their science. Those efforts are further complicated in cross-science projects by different conventions in space and solar physics. A new system has been developed based on the feedback obtained at DASH last year, showcasing a community-managed vocabulary for processing steps and a simple scientist-friendly tagging system (e.g., “Blue Ribbon”) to support data discovery. We are actively seeking community feedback to refine this new categorization approach.

        Speaker: Rebecca Ringuette
      • 16:00
        AI-Assisted FinOps: Continuous Cost Monitoring and Optimization for Multi-Account NASA Heliophysics Science Division Cloud Infrastructure 1h 30m

        As NASA science workloads expand across cloud platforms, maintaining financial visibility and controlling operational costs has become increasingly important. Without continuous monitoring, budget variances and inefficient resource utilization are often identified only after billing cycles close, limiting opportunities for corrective action and making cost containment of cloud service use significantly harder.We present an AI-assisted Financial Operations (FinOps) framework that provides continuous cost visibility, automated reporting, anomaly detection, and infrastructure optimization across the HSDcloud environment, without requiring a dedicated FinOps person.
        The framework also measures the operational costs of AI services alongside traditional cloud infrastructure, enabling direct comparison between AI investment and realized infrastructure savings. The FinOps observability layer itself adds less than $1.50 per month in AWS overhead and is built entirely on existing infrastructure with no new servers or services required. In practice the use of this software is typically recovered quickly as even a single optimization recommendation is likely of greater value than the cost to run this software. This work demonstrates how AI-assisted FinOps improves fiscal stewardship, operational transparency, and long-term sustainability for NASA scientific cloud infrastructure while providing a practical model for cloud financial governance across research organizations.

        Speakers: Jeffery Bradford (NASA Goddard Space Flight Center), Omar Shalaby (ADNET Systems, Inc)
      • 16:00
        Automated Classification of Heliophysics Instrument Data Usage Using Large Language Models 1h 30m

        We developed and evaluated an automated approach for classifying the use of heliophysics instrument data across the scientific literature using a large language model (LLM). The system classifies data usage instances identified by an upstream pipeline, where each instance represents the use of observations from a specific mission, instrument, and observation period within a scientific paper. Each usage is classified according to two dimensions: provenance, determining whether a paper presents its own analysis of the data or reports results primarily derived from another work, and depth, determining whether the data are central to the scientific analysis or included only for illustration or context.

        The classifier was evaluated on 480 data usage instances from 72 heliophysics papers selected to represent complex and ambiguous cases of data usage. The LLM was provided only with extracted quotations and contextual information rather than full manuscripts, and results were compared with classifications obtained from full-paper analysis.

        Using extracted context alone, the classifier correctly identified both usage dimensions for 84% of cases. Providing full manuscripts increased accuracy to 90% but required substantially greater computational cost and introduced additional errors. Remaining misclassifications were primarily associated with limitations in extracted information rather than the LLM's classification ability.

        These results demonstrate the potential for LLM-based methods to enable scalable characterization of instrument data usage across the heliophysics literature.

        Speaker: Alexander Warder
      • 16:00
        Automated Multi-Wavelength Detection and Characterization of Solar Activity Features Using Kodaikanal Observatory Data 1h 30m

        The advancement of heliophysics research and data-driven space weather studies depends critically on the availability of large-scale, standardized, and physically validated solar observational datasets, yet such resources remain scarce for the historically significant Kodaikanal Solar Observatory (KSO) archives, which span over a century of continuous solar monitoring. This work addresses that gap by presenting an automated multi-wavelength analysis pipeline that transforms KSO's historical observations into standardized, science-ready datasets, enabling systematic investigations of solar activity and long-term solar variability.
        The pipeline processes high-resolution (4096 × 4096) Ca II K spectroheliograms to detect and characterize chromospheric plage regions using adaptive thresholding, edge detection, and contour-based segmentation. Accurate solar disk localization via circle-fitting algorithms and orientation correction ensures consistent spatial referencing across the full observational record. The methodology is extended to white-light and H-alpha data for automated detection of photospheric sunspots and chromospheric filaments, providing a unified multi-layer framework covering more than 100 years of daily solar observations. The framework enables consistent characterization of solar activity features across one of the world's longest continuous ground-based solar archives. Heliographic coordinates and physical areas are computed for all detected features after applying limb-darkening and foreshortening corrections, yielding reproducible measurements across wavelengths.
        The pipeline outputs standardized FITS-format science products containing feature identifiers, heliographic locations, timestamps, and area measurements, validated through comparison of computed feature areas against reference measurements and manual inspection of detected plage boundaries across representative observations. These datasets facilitate reproducible analysis, solar feature studies, activity cycle characterization, and integration into broader heliophysics research workflows, while remaining suitable for future machine learning applications in solar physics and space weather research. By transforming one of the world's longest solar observational records into standardized and reproducible science-ready datasets through open data standards and automated processing pipelines, this work directly supports international capacity building in space weather science (SDG 4, SDG 9) and contributes to the data-sharing infrastructure essential for global research-to-operations transitions (SDG 17). The framework demonstrates a scalable model for integrating archival solar observations from observatories in developing nations into modern heliophysics data systems and space weather research infrastructures.

        KEYWORDS: Kodaikanal Solar Observatory, Solar Feature Detection, Ground-Based Solar Data, Space Weather

        Speaker: Ms Hanshika Jain (Department of Physics and Electronics, Jain (Deemed-to-be University), Bengaluru, India)
      • 16:00
        Data Availability Statements: The Inclusion of Available Code and Software in JGR: Space Physics Articles 1h 30m

        One of the goals of the American Geophysical Union (AGU) Journals is to advance Earth and Space sciences. For this reason, having a data availability statement in AGU journals has evolved since 1993 and became a mandatory stand-alone section of any AGU manuscript by 2019. The data availability statement requires all authors to cite and make publicly available all data and software utilized in the research process, which includes code (e.g., Python, Jupyter Notebooks, R, MATLAB) used to perform data analysis and produce the manuscript’s figures. As stated in the AGU requirements for publication, these codes should be made available on a free and open platform and preserved in a repository, hence providing a citation with an appropriate DOI. In this study, we look at the Data Availability Statement sections of 55 papers published in the Journal of Geophysical Research (JGR): Space Physics during the month of July 2026. We find that only 27.3% of papers provided publicly available software and/or code, and this includes articles that provided just the source code developed by an external entity. Another recent study (Zhai et al., 2026) also found that only 38.1% of articles published in JGR: Space Physics between 2021 and early 2026 made software/code available in either a discipline-specific or general repository. These preliminary results help us understand that, although such software and code availability is mandatory, many papers have been published without it. This highlights how both authors and reviewers have a role in making the Data Availability Statement accessible and accurate according to AGU journal regulations. As we move forward and evaluate the Data Availability Statement sections of more papers to come leading up to October 2026, we expect the same low number of publicly available code and software utilized for the execution of space physics research.

        Speaker: Stephanie Colón Rodríguez (University of Michigan)
      • 16:00
        Gas-giant plasma interchange and the need for robust, multimodal identification of variable plasma signatures 1h 30m

        Jupiter and Saturn exemplify a unique magnetospheric paradigm defined by strong rotational driving and internal heavy-ion mass loading. In these systems, plasma accumulates into a dense inner torus before being propelled outwards via corotation-associated centrifugal forces. Magnetic flux is lost during the outward transport of the heavy plasma and is replaced via discrete interchange events (IEs), during which the relatively hot, tenuous, and “magnetically-buoyant” plasma of the outer inner-magnetosphere is transported inward. IEs manifest themselves through multiple possible in-situ signatures: events can encompass a depletion in low-energy (<~100 eV) plasma fluxes; an enhancement of higher-energy plasma fluxes; a sharp, few nT change in the magnetic field; and an enhancement of plasma wave activity across various wave types. IEs are foremost detectable by the first of these attributes, but any combination of these signatures may occur simultaneously over the 30s–few-minute IE period.

        While interchange is integral to global mass circulation at the gas giants, plausible instability-onset mechanisms and the subsequent inflow morphologies remain poorly constrained due to the transient, single-point nature of existing spacecraft observations. Statistical evaluations of IE intervals are thus critical to characterizing event-time properties and inferring the spatial extent of event occurrence. However, the variability of IE-associated signatures has resulted in individual studies developing and employing different detection routines which consider and prioritize different instrument measurements, and, consequently, all observational interchange analyses have considered largely distinct sets of events. In this study, we apply a standard multi-instrument appraisal to a broad compilation of Jovian IE intervals that were identified across multiple existing Juno-era surveys to 1) assess if surveys are in fact identifying physically-alike events and 2) develop recommendations for more robust and standardized event detection. Our appraisal includes heavy-ion plasma properties which have not previously been considered in statistical analysis. We find that events can be separated into distinct classes, each of which is better organized by a unique multi-signature subset and dominates in a specific spatial region. Directional flow and ion composition properties are shown to display distinct behaviors that make them useful for detecting and categorizing event intervals, including some that had not previously been identified with other methods. We suggest possible applications of machine learning (ML) to bridge gaps in selection methodology, but we also discuss how certain ML implementations can potentially exacerbate detection biases and discrepancies. Our work provides insights that may be used to improve how we define and identify diverse signatures of plasma dynamics across heliophysical domains.

        Speaker: Alexandra Roosnovo (University of Michigan)
      • 16:00
        Heliodata: A Unified Platform for Heliophysics Research Data 1h 30m

        Heliodata is a modern web application that provides seamless access to solar research datasets through a unified, browser-based platform. By integrating diverse data sources, including the SPASE metadata catalog, it serves as a centralized hub for discovering, managing, visualizing, and analyzing heliophysics data. The platform enables users to efficiently explore a broad range of datasets, utilize interactive solar visualization tools, and generate data plots without requiring specialized software installations. By simplifying access to heterogeneous datasets and providing an intuitive user interface, Heliodata enhances the accessibility and usability of heliophysics data, enabling researchers to conduct scientific investigations more efficiently and deepen their understanding of the Sun and its influence on the solar system.

        Speakers: Olawale Jaiyeola (DISH at HDRL), Bryan Stephenson, Zach Boquet
      • 16:00
        I-ALiRT Cloud Architecture and International Ground Station Integration 1h 30m

        The Interstellar Mapping and Acceleration Probe (IMAP) mission includes the Active Link for Real-Time (I-ALiRT) system to measure Space Weather phenomena. IMAP I-ALiRT continually broadcasts data 24/7 from the IMAP observatory in orbit about the L1 Sun-Earth Lagrange point, facilitated by NASA’s Deep Space Network (DSN) of ground stations as well as antenna partners across the globe. The IMAP Science Operations Center (SOC) at the Laboratory for Atmospheric and Space Physics (LASP) receives I-ALiRT raw data from ground stations and implements a real-time, low-latency processing pipeline. I-ALiRT utilizes AWS cloud resources to facilitate efficient data ingest and processing. Launch for the IMAP mission was September 24, 2025. Data became public on Feb 1, 2026.

        Speaker: Laura Sandoval (Laboratory for Atmospheric and Space Physics, CU Boulder)
      • 16:00
        Improved access of MLSO data via API 1h 30m

        The Mauna Loa Solar Observatory (MLSO) is expanding its web service API to provide solar activity observable in MLSO data, as well as filtering other MLSO data on this activity. Currently, data from our operational instruments, UCoMP and KCor, is available via the API. The API allows users to discover the available instruments and their products, to find files matching search queries, and to download the matching files.

        We provide Python, IDL, and command-line clients, along with the JSON responses from the GET requests to the web service. The API improves the ease of downloading files for multiple dates or for non-calendar date matching patterns, e.g., need one file of a particular type each day for a Carrington Rotation. The API can also be used in scripts to download the required data as needed.

        The API is now being used in MLSO tools and documentation, for example, tutorials for using the data from a given instrument use the API to retrieve example data. We are also updating our interactive tools for analyzing MLSO and other data to use the API.

        Near future plans include a Model Context Protocol (MCP) server to connect the data served by the API with AI applications, serving more datasets, and additional formats for products of current datasets, e.g. quicklook images.

        Speaker: Michael Galloy
      • 16:00
        Lessons learned from an operational heliophysics modelling pipeline 1h 30m

        Scientific models developed for research are commonly evaluated using selected events, with their input parameters fine-tuned for the specific task. Operational deployment, however, requires continuous execution with pre-determined parameters, while tolerating incomplete or inconsistent inputs. They furthermore need the ability to recover from infrastructure failures. The requirements for a research-focused model run and an operational system, even if based on the same model, are different. We report lessons learned from operating an automated heliophysics modelling pipeline connecting real-time data retrieval with MHD simulations, postprocessing, and web publication of the results.

        Operational experience showed that many significant failures arose not within the scientific models themselves, but in the surrounding operational infrastructure. Input data might be delayed, incomplete, or incompatible with assumptions. Additional interruptions were introduced by HPC frontend and network errors. Long-duration execution further exposed instabilities not observed or avoidable in selected research cases. At the same time, forecast usefulness required balancing model resolution and assumptions against runtime, queuing time, and publication latency.

        These experiences demonstrate the need for both consistent and timely input data, and constant monitoring and checkpointing of the system. Operational performance should be evaluated not only on model output but also using metrics such as run completion rate, input data age, manual-intervention rate, recovery time, and product latency. We conclude that operationalizing a scientific model is not equivalent to automating its execution. It requires both computational resilience, product latency and scientific validity as coupled design requirements. Continuous operation also provides a systematic stress that exposes hidden model assumptions and can guide subsequent research and model development.

        Speaker: Nikolett Biro (University of Michigan)
      • 16:00
        LISIRD: Discover, Visualize, and Download Solar Data on the Web 1h 30m

        The LASP Interactive Solar IRradiance Datacenter (LISIRD), https://lasp.colorado.edu/lisird/, is a website where researchers can discover, visualize, and download over 140 solar datasets from various missions, instruments, models, and laboratories. By offering openly available data via a simple web interface, LISIRD removes common barriers researchers encounter when accessing and analyzing solar data.

        LISIRD enables researchers to plot data interactively in the browser, save and share plot configurations via URL, and customize downloads by variable, time range, and format to produce analysis-ready data. These datasets are also available programmatically via LISIRD’s LaTiS and HAPI APIs, letting users pull data directly into their own notebooks and analysis pipelines. Together these capabilities lower the barrier to entry, inviting participation from scientists across disciplines and experience levels.

        This poster will demonstrate the key features of LISIRD, describe its current technology infrastructure, and outline future improvements. It will also encourage discussion and community collaboration to advance standards and interoperability for accessible data-analysis tools in heliophysics.

        Speaker: Hunter Leise (LASP)
      • 16:00
        Making Sense of Heliophysics Metadata for Legacy Datasets 1h 30m

        The original ISTP metadata recommendations for heliophsics data sets represented as CDF files have been widely adopted by many heliophysics data providers. Without such guidelines, it would be nearly impossible to build software products like SPEDAS or PySPEDAS, which can ingest data from a wide variety of providers and instrumentation and use the same analysis and visualization tools to work with it.

        The ISTP guidelines left significant room for interpretation by data providers. Many data sets are not fully compliant in various ways. Some data sets that are similar in theory can be represented differently by different providers. Our experience with SPEDAS and PySPEDAS is that a truly generic ISTP CDF reader is not yet (and may never be) possible: there will always need to be special cases and exceptions to handle the full range of metadata constructs that are already out "in the wild".

        We will present a selection of situations we've encountered during the development of SPEDAS and PySPEDAS, and some lessons learned that we can apply to the next generation of heliophysics metadata standards currently being developed.

        Speaker: Jim Lewis
      • 16:00
        Modernising ESA Science Archives: Towards a unified ecosystem of reusable tools and interfaces. 1h 30m

        The ESAC Science Data Centre (ESDC) develops, hosts and operates the science archives for the majority of ESA’s Space Science Missions, covering Astronomy, Planetary Science, Heliophysics and Human Robotic Exploration (HRE). Each archive aims to maximise the scientific exploitation of its datasets while ensuring their efficient long-term preservation.

        Over the decades, the software architecture of these archives has evolved through several generations of technology. Historically, each mission archive was developed independently, with bespoke backend services, proprietary data access mechanisms, and custom frontend applications. While this approach addressed mission-specific needs, it has led to significant challenges, including software obsolescence, increased maintenance effort, knowledge preservation challenges, and divergent archive evolution, particularly for archives approaching the legacy phase with reduced or no dedicated archiving resources.

        Aside from technical considerations, scientists moving between archives often need to adapt to different interfaces, workflows, and access mechanisms, leading to a fragmented user experience.

        To address these challenges, ESDC is progressively evolving towards the strategic integration of stand-alone mission archives into a unified, scalable platform: the ESDC Multi-Missions Data Services (EMDS). EMDS enables cross-disciplinary, unified data access through the homogenization of interoperability mechanisms such as the IVOA TAP protocol and HAPI, standardizing interfaces across scientific domains while enhancing scalability, maintainability, and usability. This approach will provide the foundation for new missions, such as SMILE and Proba-3, as well as missions from other scientific domains, while allowing existing archives to progressively migrate to EMDS, ensuring their data remains accessible and preserved for long-term scientific use.

        For graphical user interfaces, ESA science archives are moving towards a common Angular-based framework built around reusable components and widgets. Widgets are at the heart of our user interfaces and form the building blocks of the modern ESA science archives, such as the new HelioPhysics Archive (HPA) 1.0.0 released in June 2026.

        ESDC’s approach is to establish a shared widget library across services: a collection of mission-agnostic, reusable components that can be integrated into multiple archives while preserving mission-specific capabilities. This provides a consistent user experience and interaction model across scientific domains, reduces duplicated development and maintenance effort, accelerates delivery of new functionality, and enables shared innovation where improvements developed for one archive can benefit others.

        By combining EMDS, open standards, and a common widget ecosystem, ESA archives are evolving from isolated mission-specific systems into a sustainable, interoperable platform of shared capabilities. New missions can benefit from mature services and technologies from the outset, while existing archives can progressively adopt modern approaches, ensuring continued accessibility, preservation, and scientific exploitation of ESA’s data.

        This presentation outlines the vision, architecture, and implementation strategy behind the evolution of ESA Science Archives towards a sustainable, interoperable, and user-focused ecosystem.

        Speaker: Mr Jonathan Cook (Starion for ESA)
      • 16:00
        Plot Walk: Browsing 50 Years of Pre-Generated Heliophysics Plots 1h 30m

        The Space Physics Data Facility (SPDF) maintains an archive of over 17 million pre-generated plots containing ephemeris and scientific data from approximately 30 missions spanning five decades of heliophysics research. These plots, contributed by both SPDF staff and mission teams, represent a valuable visual record of heliospheric and magnetospheric observations from historical and ongoing missions. However, the heterogeneous nature of this archive—with inconsistent filename conventions, variable time range specifications, and diverse image dimensions—presented significant challenges for creating a unified browsing interface.

        We present Plot Walk, a web-based visualization system designed to provide seamless access to SPDF's diverse plot archive. This presentation will discuss the technical solutions we developed to address key challenges, including: parsing inconsistent temporal metadata from filenames (where some missions provide only start times, others only end times); handling plots with both fixed and variable time ranges; and displaying figures with widely varying dimensions and resolutions in a consistent interface.

        Plot Walk features an extensible architecture that allows new plot types to be easily integrated, caching for smooth sequential browsing, navigation controls for stepping through plot series, and inventory visualizations showing data availability. This tool enhances accessibility to SPDF's extensive data archive and demonstrates practical approaches for managing heterogeneous scientific visualization collections.

        Speaker: Eric Grimes (Space Physics Data Facility)
      • 16:00
        Running a Science Cloud without being a Cloud Expert 1h 30m

        Cloud computing is an increasingly important platform for scientific collaboration, interactive analysis, and reproducible research. However, deploying and maintaining cloud infrastructure often requires familiarity with a rapidly growing ecosystem of tools, creating a significant barrier for many scientists and research software engineers. A reproducible deployment makes it practical to rapidly create new science environments for collaborations, workshops, or project-specific efforts, while leveraging cloud bursting to temporarily scale compute resources for computationally intensive analyses without maintaining permanently provisioned infrastructure. HelioCloud was designed to reduce this complexity by providing a reproducible, open-source science cloud built on modern cloud-native technologies (OpenTofu, Helm, Kubernetes) while exposing only the concepts necessary for day-to-day operation. We discuss how researchers and developers can confidently contribute to and operate a science cloud without becoming cloud experts, while taking advantage of scalable computing, simplified deployments, and direct access to large scientific datasets.

        Speaker: Peter Shumate (Johns Hopkins University Applied Physics Lab)
      • 16:00
        SOLER Open-Source Python Tools for the Analysis of Energetic Solar Eruptions 1h 30m

        The recently expanded fleet of heliospheric spacecraft presents unique opportunities for exploring solar eruptive phenomena such as coronal mass ejections (CMEs) and solar energetic particles (SEPs) from multiple vantage points. However, the task of integrating diverse observations collected by different instruments across various spacecraft poses a notable challenge. To maximize the utilization of this data within the broader scientific community, the EU Horizon Europe project SOLER (Energetic Solar Eruptions: Data and Analysis Tools, https://soler-horizon.eu) offers a versatile array of tools. These tools, provided as open-source Python Jupyter Notebooks, cater to scientists with limited programming expertise. The offered functionalities start from automatic downloading & visualizing SEP intensity-time profiles and various other in-situ measurements by the whole heliospheric spacecraft fleet. Further analysis tools allow to determine and fit SEP event energy spectra as well as SEP Pitch-Angle Distributions (PADs) and first-order anisotropies, including methods for background subtraction. A final set of tools is dedicated to the analysis of SEP onset times, offering a regression method and a hybrid Poisson-CUSUM-bootstrapping approach. Here we provide an overview of the available toolkit and instructions on its utilization, which can be seamlessly accessed on the project’s dedicated JupyterHub server in the cloud.

        Speaker: Jan Gieseler (University of Turku)
      • 16:00
        SPEARHEAD Tools for High-Energy Particle Data Analysis 1h 30m

        The EU Horizon Europe project SPEARHEAD (SPEcification, Analysis & Re-calibration of High Energy pArticle Data, https://spearhead-he.eu) develops and provides open-access tools to improve the analysis of high-energy particle observations for heliophysics and space weather research. The toolkit includes Bowtie, which derives effective energies and geometric factors from instrument response functions of charged particle telescopes; FDAT, a graphical environment for identifying and modeling Forbush decrease events; G4VM, a pre-configured GEANT4 virtual machine with detailed spacecraft and instrument models; VDA, a notebook-based workflow for velocity dispersion analysis of solar energetic particle (SEP) events; and SEP-PACT, a Solar Energetic Particle Analysis and Calculation Tool that computes SEP event key parameters such as onset time, peak flux, peak time, and fluence for both electron and proton data of Solar Orbiter/HET, enabling consistent and efficient event characterization. Together, these tools enhance instrument calibration, event characterization, and simulation capabilities, supporting both scientific studies and operational applications. All tools are openly available through GitHub repositories and the project's own JupyterHub server, ensuring long-term accessibility for the heliophysics community.

        Speaker: Jan Gieseler (University of Turku)
      • 16:00
        Station-Resolved Analysis of Solar-Wind Coupling in SWMF Ground Magnetometer Simulations 1h 30m

        Global geomagnetic indices such as AE, AL, and SYM-H are widely used to study solar-wind–magnetosphere coupling, but they compress a spatially structured system into a small number of time series. This compression can obscure where, when, and how solar-wind information appears in local ground magnetic perturbations. In this project, we develop a station-resolved framework for analyzing local information transfer in the coupled solar-wind–magnetosphere–ionosphere system using the Al Shidi and Pulkkinen SWMF ground magnetometer simulation dataset. The
        dataset includes event-level SWMF log files with solar-wind/input quantities and simulated global quantities, as well as station files containing observed and simulated magnetic perturbation components at individual ground magnetometer stations.

        For each storm event, we align solar-wind driver variables with both observed and simulated station-level magnetic perturbations. We then compute traditional lagged correlation maps and compare them with information-theoretic measures such as mutual information and transfer entropy. This allows us to ask not only where the solar wind is correlated with local geomagnetic response, but where it provides predictive information beyond the local response’s own recent history. By applying the same workflow to observed and simulated station data, we can evaluate whether SWMF reproduces the spatial distribution, timing, and driver dependence of local geomagnetic information flow.

        In parallel with this classical analysis, we plan to explore whether similar information-flow and observability questions can be formulated using quantum-computing approaches, including small-scale implementations with Qiskit and IBM Quantum resources. This component will initially be exploratory: rather than claiming quantum advantage, we aim to test whether quantum or quantum-inspired representations can provide useful alternative ways to encode, compare, or classify solar-wind–magnetosphere coupling states. In this way, the project serves both as a station-level extension of traditional geospace model validation and a first step towards future quantum-assisted analysis of heliophysics data.

        Speaker: Gergely Koban (University of Michigan)
      • 16:00
        Staying Up to Date with Cloud Infrastructure 1h 30m

        Science workloads increasingly rely on cloud platforms to provide the elastic compute, storage, and networking needed for data- and compute-intensive research. By leveraging cloud-native technologies such as containers, Kubernetes, and managed services, scientific applications can scale from interactive prototyping to large parallel experiments while improving reproducibility and portability across institutions. This shift enables researchers to co-locate data and compute, automate complex workflows, and rapidly adopt accelerators and specialized hardware without owning physical infrastructure. At the same time, it introduces new challenges in performance variability, cost optimization, security, and compliance with data governance policies.

        One key and often overlooked aspect of working in the cloud is staying up to date with all the various tools and technologies. Kubernetes, which is the de facto orchestration platform for running containerized applications at scale, follows a four-month release cycle, where each release reaches end of life after 14 months. Organizations benefit from clear version policies that define supported releases, planned upgrade windows, and systematic testing of breaking changes in staging environments. Continuous learning—through hands-on labs, workshops, certifications, and community engagement—helps teams safely adopt new capabilities reliably.

        Speaker: Nicholas Lenzi (Johns Hopkins University Applied Physics Laboratory)
      • 16:00
        The SciQLop ecosystem: a layered open-source stack for in-situ heliophysics data 1h 30m

        Most of the effort in working with in situ space plasma data is not spent on the data itself. Finding one product among the tens of thousands an archive offers is genuinely hard, and once it is found, making sense of it means reading metadata by hand, working out which attribute holds the fill value, what the units are, which axis carries energy. Neither should be the scientist's job, and what comes next, exploring the data and recording what was found in it, should not require leaving the tool. The SciQLop project addresses this as a set of separate, independently useful components rather than as a single monolithic application. This poster presents the resulting stack and why it is split the way it is.

        AstraLint works the producer's side of that problem. It validates files against ISTP and PDS4 conformance suites, from the command line, in CI, or entirely in the browser, so that the metadata everything downstream depends on is correct before a file is published. CDFpp is a from-scratch, MIT-licensed, thread-safe C++20 CDF implementation with full read and write support and SIMD-accelerated epoch conversion, usable from C++, from Python as pycdfpp with zero-copy NumPy arrays and GIL-free I/O, and from JavaScript through a WebAssembly build with zero-copy typed arrays. Speasy presents AMDA, CDAWeb, SSCWeb, the Cluster Science Archive, CDPP's 3DView geometry service, HAPI servers and local archives behind one API, backed by a persistent on-disk cache and an optional server-side caching proxy; a SuperMAG ground-magnetometer provider is in progress. Everything it returns is a SpeasyVariable: values, a time axis, named axes, and the ISTP attributes that came with the data. That type is the stack's interface. Because FILLVAL, UNITS and the valid range travel with the array, a variable can clean itself, convert to a pandas dataframe or an astropy table, plot itself, and carry astropy units whenever the unit string is one astropy can parse. Every layer above speaks it, so nothing has to be converted between them. SciQLopPlots and its rendering engine draw multi-million-point time series and spectrograms on the GPU through Qt's RHI. jupyqt embeds JupyterLab inside a Qt application. tscat stores event catalogs locally, while cocat is a separate library that carries the same catalog model over Yjs CRDTs for real-time co-editing across institutions. SciQLop is the desktop application that composes them, adding a plugin system, an app store and a workspace model.

        Every layer is used on its own, in notebooks, in mission pipelines and in other groups' tools, and that constraint shapes the design: no layer may depend on the application above it, and each ships independently to PyPI with its own test suite and release cadence. Two of them also run with no installation at all, compiled to WebAssembly.

        We will show the architecture, where the standards actually bind, what the split costs, and the published studies produced with the stack. The poster is equally an invitation: tell us which layer would be useful to you without the rest.

        Speaker: Alexis Jeandet (LPP-CNRS)
      • 16:00
        The Virtual Solar Observatory 2.0: Connecting Heliophysics Data, Metadata, Services, and Users 1h 30m

        For more than two decades, the Virtual Solar Observatory has provided a common interface for discovering and accessing solar and heliophysics data held by distributed archives. VSO 2.0 is a major modernization of this service, replacing the original Perl- and SOAP-based system with a predominantly Python-based, REST-oriented architecture.

        This poster presents an overview of VSO 2.0 and the capabilities being developed for researchers, data providers, and software clients. These include a revised data model, an expanded Registry describing available data holdings and provider capabilities, richer metadata services, reusable search presets, and updated programmatic and web-based interfaces.

        We illustrate how the principal elements of VSO 2.0 connect users and software clients with distributed data providers, while maintaining a consistent approach to search and data access across heterogeneous archives. The poster will also describe the current status of the project and how the new architecture can support future instruments, datasets, metadata types, and community services.

        Together, these developments are intended to make VSO easier to extend, more transparent to users and providers, and better able to support metadata-driven discovery across the solar and heliophysics data ecosystem

        Speaker: Alisdair Davey (National Solar Observatory)
      • 16:00
        The Virtual Solar Observatory 2.0: New APIs, SunPy Integration, and TAP Access 1h 30m

        The modernization of the Virtual Solar Observatory as the Model-View-Controller-based VSO 2.0 application continues. We present new APIs for reporting operational health and usage metrics, together with services for extracting and visualizing selected physical observables and metadata values in High Performance Computing workflows. Examples include Stokes parameters from Solar Orbiter/PHI files and exposure-time metadata from Solar Orbiter and SDO/AIA files.
        We also describe provider-focused endpoints that expose available search options, sources, and instruments. Programmatic access is demonstrated through a new VSO 2.0 client for SunPy that supports Fido.search and Fido.fetch. The client uses Pydantic models to validate user requests before submitting them to VSO 2.0.
        Recent additions to the VSO 2.0 Registry describe the searchable characteristics of provider holdings in greater detail, including data levels, physical observables, datasets, and file types associated with providers, sources, and instruments. Finally, we outline a new Table Access Protocol server that will provide TAP-based access to holdings discoverable through VSO 2.0.
        Together, these capabilities provide finer-grained data discovery, improved operational visibility, and more flexible programmatic access, while establishing an extensible foundation for future VSO services.

        Speaker: Dr Ed Mansky
      • 16:00
        Towards Open Science: Open Data Practices of the Chinese National Space Science Data Center (NSSDC) 1h 30m

        Open science has entered a new stage of global consensus, critical for accelerating scientific progress, enhancing research impact, broadening the application of scientific outcomes, and fostering a healthy data ecosystem. This poster highlights the open data practices implemented by the National Space Science Data Center (NSSDC), as a national‑level data center.
        NSSDC has archived and integrated a multidisciplinary data resource system covering space physics, space environment, space astronomy, planetary science, space geoscience, and other related fields. All data held by NSSDC are assigned dual persistent identifiers (DOI and CSTR) during curation. To promote open data, NSSDC customizes data service systems for a series of satellite and ground‑based observation projects. To improve data discoverability, NSSDC has built a data retrieval service that offers cross-system, cross‑disciplinary, and distributed discovery of data resources. Meanwhile, data catalogues are synchronized to third‑party services through harvesting or registration via this retrieval service. By addressing the needs of the space science community, NSSDC aim to promote space science application.

        Speakers: Qi XU (National Space Science Center, Chinese Academy of Sciences), Mr xin xu
    • 09:00 09:10
      Welcome 10m
    • 09:10 09:40
      Keynotes: Keynote 3
    • 09:40 10:40
      Session 7: Science and Mission Planning Tools for Space Weather and Human Exploration
      • 09:40
        CCMC Tools to Enhance Scientific Analysis 15m

        The Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov) has developed many tools to aid in the scientific analysis of space weather. These tools leverage the many numerical simulations hosted at the CCMC while pulling in relevant observational data. These tools provide both context and scientific investigation to better understand space weather. CCMC tools are used operationally by several groups for space weather predictions, by many scientists for analysis of past space weather activity, and also for mission planning for future missions (for example GDC).

        Kamodo is an open source tool developed by the CCMC that leverages other Python in Heliophysics Community (PyHC) packages such as HAPI and SpacePy to bring in model output and observational data to create new visualizations and analysis. Kamodo can extract surfaces such as the bow shock and magnetopause locations from model runs to share with the SPDF Orbit Viewer or use in other analysis tools. Model output at satellite positions can be automatically extracted and compared with the observed values, automatically pulling satellite positions from SSCWeb and data from CDAWeb (HDRL/SPDF). Output from multiple ITM models can be extracted and compared to each other and to data.

        This presentation will highlight Kamodo, but also mention other CCMC tools and how to use them. Most of these tools utilize a web interface for easy access, but some also offer offline use for a more powerful and customized experience.

        Speaker: Darren De Zeeuw (The Catholic University of America / NASA Goddard Space Flight Center, CCMC)
      • 09:55
        End-to-End Software Suite for Space Weather Sub-L1 Multi-Spacecraft Mission Planning and CME Early Warning 15m

        Accurate forecasting of Earth-directed Coronal Mass Ejections (CMEs) and their geo-effectiveness requires advanced early warning capabilities that exceed the limitations of current single-point L1 monitors. Sub-L1 multi-spacecraft architectures offer a highly promising solution; however, architecting these missions requires robust, end-to-end simulation environments to evaluate payload configurations and ground-segment algorithms prior to deployment.

        In this contribution, we present a novel software pipeline designed to comprehensively simulate and analyze data from sub-L1 multi-spacecraft space weather missions. The software generates a realistic 2D spatial domain featuring an intermittent and turbulent background solar wind. This background is dynamically populated with key large-scale structures, including Corotating Interaction Regions (CIRs), Heliospheric Current Sheet (HCS) crossings, and transient CMEs. By modeling these interplanetary structures in 2D, the tool accurately reproduces the distinct, location-dependent solar wind conditions encountered by a distributed constellation of spacecraft.

        To bridge the gap between physical models and mission engineering, the pipeline incorporates instrument transfer functions, effectively translating the simulated space environment into realistic, telemetry-like in-situ sensor responses. Furthermore, the software features an integrated ground-segment analysis module tailored for multi-point data. This module demonstrates advanced operational techniques, including the use of magnetic helicity for robust automated CME detection. Crucially, the software implements geometric interferometry across the simulated spacecraft network to accurately estimate the propagation direction of CMEs—a vital measurement for predicting Earth-impact that is fundamentally impossible with traditional single-spacecraft observations.

        Ultimately, this software suite serves as a powerful mission planning tool. We will demonstrate how it can be utilized to define strict scientific requirements for CME early warning, optimize scientific payload configurations, and validate the most effective data analysis techniques for estimating CME propagation and geo-effectiveness at Earth.

        Speaker: Daniele Telloni (National Institute for Astrophysics - Astrophysical Observatory of Torino)
      • 10:10
        Designing and Implementing a Reusable Multi-Mission Timeline Framework 15m

        The Interface Region Imaging Spectrograph (IRIS) science operations team is nearing completion of a major redevelopment of the mission timeline creation software. This “Timeline Tool” is used for planning an instrument with 1-5 day timelines in LEO, often containing 10-15 planner-chosen observations per 24 hour period. We have strived to not just improve and refresh the aging timeline software for the 13-year-old mission, but to carefully formalize rules and data structures likely applicable to similar future missions and to modularize the planning software into general and mission specific components. One key goal is to enable relatively rapid and inexpensive adaptations to future space-based missions. We plan to convert and extend this too for the upcoming Multi-Slit Solar Explorer (MUSE), Extreme Ultraviolet High-throughput Spectroscopic Telescope (EUVST), and Chromospheric Magnetism Explorer (CMEx) programs; our presentation will cover the early stages of the adaptation for MUSE. We are also beginning to explore of how the redeveloped timeline web application could support automated “first drafts” of timelines to reduce the day to day manual labor of planning. These approaches to planning would use scripting coordinated observations from another instrument’s plan and/or modern LLM agents to interpret the various sources of instructions and construct a full science timeline via API calls to the timeline backend.

        Speaker: Ryan Timmons (Lockheed Martin)
      • 10:25
        DOFCAT: An Optical Flow Tool for Multi-Coronagraph CME Velocity Mapping and Physics-Driven Automated Detection 15m

        Coronal mass ejections (CMEs) are among the most energetic solar eruptions and primary drivers of space weather. Despite decades of study, their internal velocity distributions remain poorly characterised. Coronagraph-based studies of CMEs have traditionally relied on leading-edge tracking or geometric fitting methods, which provide limited information about the internal velocity structure of eruptions. We present DOFCAT (Dense Optical Flow CME Analysis Tool), an open-source Python-based software package that applies dense optical flow algorithms to coronagraph image sequences to generate spatially resolved, pixel-level velocity maps of CMEs and their substructures. Unlike conventional methods, DOFCAT requires no prior assumptions about CME geometry or morphology, making it broadly applicable across instruments and events.

        DOFCAT has been validated using high-resolution, high-cadence data from two next-generation space-based coronagraphs, ASPIICS onboard ESA's PROBA-3 and METIS onboard Solar Orbiter, as well as LASCO C2 onboard SOHO, providing complementary coverage of the middle corona (1.5–6.0 R☉), the critical region where CME impulsive acceleration and internal restructuring predominantly occur. The tool incorporates a dedicated pre-processing pipeline that includes background subtraction and edge-preserving noise smoothing while preserving feature edges to prepare coronagraph data for optical flow computation. We also developed a Gaussian-tapered Fourier (GTF) filter to suppress brightness flickering artefacts in ASPIICS running difference images, significantly improving velocity estimation stability. Validation across multiple structured CME events demonstrates that DOFCAT reliably captures internal velocity dispersion, front-core separation dynamics, and position-angle-dependent velocity gradients, which otherwise are inaccessible to traditional tracking approaches.

        A key prospective capability of DOFCAT is physics-driven automated CME detection. Analysis of optical flow maps reveals that CME passage through the coronagraphic field-of-view produces a characteristic and reproducible statistical signature in the frame-by-frame velocity distribution: a sharp rise in high-velocity pixel fraction relative to background levels, broadening of the velocity histogram, and subsequent return to background levels. This signature is image-independent and physically motivated, providing a robust basis for automated detection without relying on brightness thresholds or morphological assumptions. Building on this, we are currently developing a machine learning and AI-based CME detection model trained on velocity distribution signatures extracted from a large sample of CME events across multiple coronagraphs. This approach moves beyond traditional intensity-based cataloguing, enabling scalable, systematic, and physically interpretable CME identification across large coronagraph archives with direct applicability to real-time space weather monitoring pipelines.

        DOFCAT is publicly available as an open-source repository, designed for community use with data from existing and upcoming coronagraphs, including PUNCH and Vigil (at L5 vantage point). Standardised velocity map products from DOFCAT can serve as inputs for space weather forecasting pipelines, CME databases, and heliospheric propagation models. By providing a geometry-free, automation-ready framework for CME velocity analysis, DOFCAT addresses a significant gap in the current space weather instrumentation toolkit and contributes to the global effort to enable real-time, AI-enabled CME monitoring and forecasting.

        Speaker: Pritam Das (Aryabhatta Research Institute of Observational Sciences)
    • 10:40 11:00
      Coffee 20m
    • 11:00 12:30
      Session 8: Agentic AI and LLM Applications in Heliophysics
      • 11:00
        Democratizing Heliophysics Research with AI Agents 15m

        Modern heliophysics research is increasingly limited not only by model capability, but by fragmented workflows across literature, mission data archives, community scientific software, code, collaboration, and writing. Real research must connect physical interpretation with tools such as SPEDAS/PySPEDAS, mission data products, plots, manuscripts, and review processes. I argue that useful AI systems for scientific work should be designed as persistent, tool-using, auditable agents rather than one-off chat sessions.

        Using workflows around Parker Solar Probe, SPEDAS/PySPEDAS, and AI-assisted scientific communication as motivating examples, I present lessons from building LingTai, a local-first prototype runtime for persistent AI research workers. The focus is not a single SPEDAS bot, but the broader agent-harness layer required for scientific work: project memory, tool execution, provenance, delegation, human approval, and reproducible artifacts. I discuss how such harnesses can support code navigation, data-to-plot reproduction, figure and artifact management, manuscript/review workflows, and community software maintenance while leaving scientific judgment with domain experts. I close with candidate evaluation tasks for measuring whether AI agents can reliably support real heliophysics research.

        Speaker: zesen huang (university of california, los angeles)
      • 11:15
        Large Language Models as Heliophysics Software Development Tools 15m

        The SPEDAS/PyPSEDAS software development team has been ramping up our usage of chat-based and agentic LLM tools over the past year. The underlying models continue to quickly evolve and gain capability, but still require some careful prompting and context management to achieve the best results. We will present some of our recent experiences, lessons learned, and future plans for incorporating AI tools into our group's software development, architecture, and QA workflows.

        Speaker: Jim Lewis
      • 11:30
        Bringing AI-assisted Coding to HelioCloud with Jupyter AI 15m

        HelioCloud now includes the Jupyter AI extension, which brings frontier coding agents such as Codex, Claude, Copilot, Gemini, Goose, Mistral Vibe, and OpenCode into JupyterLab. Users can connect their agent of choice to Jupyter AI in HelioCloud and begin developing notebooks with AI assistance. In this presentation, we provide a real-time demonstration of Jupyter AI in HelioCloud using a practical heliophysics problem. We also introduce resources designed to help users become familiar with Jupyter AI and AI-assisted coding more broadly. These include a “getting started” tutorial notebook; a notebook reviewing Jupyter AI’s capabilities such as MCP server integration and tool use; and a notebook that evaluates a model’s effectiveness on common coding tasks, including test design and debugging. HelioCloud users without a paid or enterprise AI subscription will also find guidance for accessing free, yet still highly capable, agents. Together, the demonstration and supporting resources are intended to help heliophysics researchers begin using AI-assisted coding effectively within their existing computational workflows in HelioCloud.

        Speaker: Lisa Knowles (JHU APL)
      • 11:45
        Creating AI Agents for PyHC with Claude Code, Codex, and Beyond 15m

        This talk explores practical applications of agentic AI within the Python in Heliophysics Community (PyHC). We show how customizing off-the-shelf agentic coding harnesses can deliver 90% of the benefits of a purpose-built AI agent with 10% of the effort of creating one from scratch. We do this using instruction files, reusable skills, subagents, and MCP servers to adapt tools such as Claude Code, Codex, Copilot CLI, and Gemini CLI to PyHC-specific workflows. These customizations are embedded directly in project repositories, so users can immediately work with the agents in whatever compatible harness they already use or prefer.

        The talk demonstrates three agent-powered tools developed for PyHC: (1) the “PyHC Standards Evaluator,” an agent that automates the grading of packages against PyHC's development standards; (2) the “HSSI Metadata Extractor,” which parses software repositories to extract the metadata needed for submission to the Heliophysics Software Search Interface (HSSI); and (3) “PyHC-Chat,” an agent designed to answer questions about the PyHC ecosystem and its roughly 100 packages, help users identify and install the right packages for specific scientific tasks, write code with them, draft (executable) papers, and much more.

        As a broader takeaway, the approach behind PyHC-Chat—pairing a curated collection of resources with “lay of the land” instructions—is a reusable pattern attendees can apply to create customized agents for whatever organization, community, business, or specialized field matters to them.

        Speaker: Shawn Polson (LASP | PyHC)
      • 12:00
        Agentic LLM workflows inside SciQLop: data, literature, and the verification practices they require 15m

        Large-language-model applications in heliophysics generally operate outside the software in which analysis is performed: they answer questions about archives, generate scripts, or curate metadata. We report on an agent embedded in a running desktop analysis application, and on the verification practices required to use its output.

        SciQLop is an open-source PySide6 environment for in situ space plasma data, embedding a Jupyter kernel and reaching tens of thousands of products from CDAWeb, AMDA, the Cluster Science Archive and SSCWeb through the Speasy library. Its agent layer exposes a single MCP-shaped tool list, shared by several interchangeable backends, acting on the live session rather than on a sandbox: the kernel namespace shared with JupyterLab, the product tree and its path-to-data resolver, plot creation and figure read-back, and detached background jobs for long archive fetches. A second group of tools reaches the published literature, searching arXiv and NASA ADS and retrieving open-access full text from an arXiv identifier, DOI or ADS bibcode. Data and literature are consequently available within the same reasoning step, so a hypothesis can be framed from prior work, tested against archive data, and compared with published results without leaving the environment. Tools that modify state are gated and confirmed per call.

        Three studies built with this workflow are presented: a multi-mission survey of solar-wind radial evolution across the heliosphere, in which scaling exponents are measured from spacecraft conjunctions rather than from pooled statistics; an assessment of how well a single L1 monitor represents the solar wind reaching Earth, whose measured decorrelation length independently reproduces published flux-tube widths; and a temperature-anisotropy analysis of Saturn's inner magnetosphere from reconstructed Cassini/CAPS distributions, which returns a negative result at catalogued mirror-mode events.

        The transferable outcome is the verification discipline. Results that initially appeared striking were traced to incomplete instrument sampling, to a correlation that survived only until a confounding variable was controlled for, and to a calibration check that was true by construction. We conclude that agreement between an agent's result and expectation carries no information unless the check is independent of the procedure that produced the result, and we describe the practices we now apply by default.

        Speaker: Alexis Jeandet (LPP-CNRS, Paris, France)
      • 12:15
        Agentic Fault Analysis for CCMC ROR Simulations 15m

        The Space Weather and Heliophysics research community, supported by the Community Coordinated Modeling Center (CCMC, https://ccmc.gsfc.nasa.gov), provides a collaborative platform for space weather models and comparisons with data. The growing demand and increasing complexity of the modern models puts a practical limit on a manual diagnosis of model faults in otherwise automated pipelines, including CCMC’s Runs-on-Request (ROR) and Instant Runs (IR) services. To reduce this bottleneck, we have developed a prototype agentic AI system designed to streamline the error analysis and triage of simulation runs. Our preliminary tests show that the system can correctly recognize and suggest remediation for a number of complex simulation errors including numerical instability, grid mismatch, input configuration issues, and others.

        The system utilizes a Local Large Language Model (LLM) equipped with a suite of curated, diagnostic tools to investigate simulation failures. The system employs a multi-stage reasoning pipeline consisting of an Investigator (to build an evidence packet with exact file and line citations), a Critic (to attack unsupported causal claims and preserve contradictory evidence), and a Planner (to determine the root cause and suggest a discriminating solution). Furthermore, we implemented a custom context-budget manager that summarizes aged tool results, allowing the agent to parse massive log files without exceeding token limits.

        This presentation will detail the current triage architecture, demonstrate its operational effectiveness, and discuss the future transition from static metadata curation toward active, multi-agent scientific workflows.

        Speaker: Jack Topper (NASA Goddard Space Flight Center)
    • 12:30 14:00
      Lunch 1h 30m
    • 14:00 15:30
      Session 9: General Session 3 (Archives, Services & Platforms)
      • 14:00
        The Virtual Solar Observatory 2.0: Metadata-Driven Discovery and Access for Heliophysics Data 15m

        For more than two decades, the Virtual Solar Observatory has provided a common interface for discovering and accessing solar and heliophysics data held by distributed archives. VSO 2.0 modernizes this infrastructure by replacing the original Perl- and SOAP-based system with a predominantly Python-based, REST-oriented architecture.

        This presentation provides a high-level overview of VSO 2.0, with particular emphasis on its revised data model and Registry. The new data model provides a more consistent description of observations, datasets, instruments, providers, and services, while distinguishing between information used to construct searches and metadata returned to users. The Registry records both the available data holdings and the capabilities supported by individual providers.

        Together, the data model, Registry, APIs, programmatic clients, and redesigned web interface provide a more extensible platform for heliophysics data discovery and access. They also establish a foundation for richer metadata services, reusable search workflows, and the integration of new instruments, archives, and data products.

        We will summarize the current implementation, describe how the major components work together, and discuss how VSO 2.0 can contribute to the wider interoperable heliophysics data environment.

        Speaker: Alisdair Davey (National Solar Observatory)
      • 14:15
        The Heliophysics Software Search Interface: Current Status and Overview 15m

        Modern research workflows increasingly require side-by-side discovery of publications, data, and software. Existing efforts increasingly provide this interlinked discovery for publications (Science Explorer (Sci-X)) and data (HelioData). Current Heliophysics scientific software interfaces, however, are lacking in relevance for Heliophysics, connection to data and publications, and scope.

        This situation is further worsened by troubling gaps in publication infrastructure, resulting in citations to software in peer-reviewed publications being lost in tracking and attribution infrastructure. As a result, the science software funded by U.S. agencies, including NASA, is difficult to discover despite increasingly valiant attempts by researchers and publishers to adopt open science practices.

        The Heliophysics Software Search Interface (HSSI; https://hssi.hsdcloud.org/) platform is an open-source NASA ROSES23 HPOSS-funded effort to address these issues and improve Heliophysics software’s findability, discoverability, accessibility, citability, and searchability. Over a year into the effort, we have successfully completed our initial goal of creating the HSSI platform, including a landing page to search for science software in Heliophysics, a public resource registration form, as well as a new metadata structure that is tailored to Heliophysics specific needs and aligned with international standards. The landing page for the software search design is predominantly built upon NASA’s highly adopted Exoplanet Modeling and Analysis Center (EMAC) with a RestAPI added on top. We have successfully collaborated with EMAC to release the updated version of the code through LASP with the open-source Apache 2.0 license.

        This work will take viewers through an overview of the current HSSI website, including details on how to use the website, the state of our backend, the HSSI AI-powered Metadata Orchestrator and finally, efforts towards coordination with cross-disciplinary communities and tools.

        Speaker: Julie Barnum (Laboratory for Atmospheric and Space Physics and University of Colorado Boulder)
      • 14:30
        From Launch to Archive in Hours: The SMILE Mission 15m

        The Solar wind Magnetosphere Ionosphere Link Explorer (SMILE) is a joint mission between the European Space Agency (ESA) and the Chinese Academy of Sciences (CAS) dedicated to studying the interaction between the solar wind and Earth's magnetosphere. Beyond its scientific objectives, SMILE represents an important historic milestone for the ESAC Science Data Centre (ESDC): it is the first ESA science mission whose data were successfully received, processed and archived within hours of launch, demonstrating archive readiness from the earliest phases of mission operations.

        This achievement was enabled by the close collaboration between the ESA Science Operations Centre (SOC), archive teams, and mission stakeholders throughout development, testing, and operational preparation. Interfaces, data flows, and archive services were validated before launch, allowing the archive to support the mission from its very first day.

        SMILE data are managed through the ESDC Multi-Mission Data Services (EMDS) framework and made available through the HelioPhysics Archive (HPA), ESA's common access point for heliophysics missions. These archive services operate within the broader SMILE ground segment, supporting the flow of scientific data between the Chinese ground segment, the ESA Science Operations Centre (SOC), instrument teams across Europe, and the scientific community, while providing interoperable data access, visualization, and long-term preservation capabilities.

        This presentation describes the SMILE mission and archive architecture within a distributed ground segment involving ESA, CAS, and instrument teams across Europe. It presents the challenges and lessons learned and highlights the role of SMILE in the evolution of ESA's heliophysics data ecosystem.

        Speaker: Angela Carasa Ruiz
      • 14:45
        Using Free and Open-Source Tools to Manage Your Infrastructure 15m

        Whether you’re working with on-prem or cloud servers, you will need to give the right people access, install software dependencies, manage firewalls, set up domain names, perform system updates, and whatever other application-specific tasks are required to keep your services functioning. In this talk, I present the tools that The Helioviewer Project’s development team has adopted for managing our service infrastructure. These tools include Apache Superset for visualizing statistics, Sentry for application health monitoring, OpenTofu for managing cloud infrastructure, Docker for creating self-contained environments, and Ansible for deploying software. Using these tools reduces software development effort and takes off the mental load of remembering how each application is configured and deployed, allows applications to run autonomously, and moves any debugging efforts onto the application itself and away from its environment.

        Speaker: Daniel Garcia Briseno (Adnet Systems Inc. / NASA Goddard)
      • 15:00
        Escaping the Two Week Sprint 15m

        Research software engineers often work in environments where priorities shift quickly, projects overlap, and responsibilities are not always clearly defined. There is always another scientist to support, another bug to fix, or another deadline to meet. Early in my RSE career, I have found that surviving the day-to-day work is only part of the challenge. While this work is necessary, so is finding ways to intentionally grow, both individually and as a team, while still meeting the immediate needs of scientific projects.

        This talk explores tools and practices I have encountered outside of research software and adapted to navigate that balance. These include setting longer-term and time-bounded goals, improving ownership and knowledge sharing within teams, and advocating for better processes without just adding overhead.

        Rather than presenting a universal solution, I will share what I have tried, what has helped, and what questions remain as I continue building my own RSE toolbox. The goal is to start a conversation about how we can move beyond the next sprint while staying responsive to the unpredictable nature of research software.

        Speaker: Umer Salman (Southwest Research Institute)
      • 15:15
        SEPsando: a web-based sandbox for multipoint solar energetic particle and space weather events 15m

        Interpreting multipoint solar energetic particle (SEP) events requires reconciling spacecraft positions, magnetic connectivity, coronal mass ejection and flare properties, interplanetary shocks, and in situ plasma signatures across many missions, together with the interpretation of different models. That information already exists (CDAWeb, SSCWeb, OMNIWeb, DONKI, ISWA, Helioviewer, HEK, etc.) but lives behind a dozen independent interfaces, and researchers spend more time reconciling them in a 3D context and writing glue code than analyzing the event itself.

        We present SEPsando, a web-based visualization sandbox developed at NASA GSFC's Space Physics Data Facility. A 2D/3D scene shows the current heliophysics fleet with a scrubbable timeline; particles overlay as probability regions from their origin; magnetic connectivity to the Sun is computed on the fly with several methods. The backend is built on Flask with SunPy, Astropy, and pfsspy; the frontend uses Three.js for the interactive 3D scene and Chart.js for time series and maps. SEPsando currently integrates directly with Helioviewer, CCMC's DONKI and ISWA, the 4D Orbit Viewer, and IRAP's Connectivity Tool. The explicit design goal is to link the existing ecosystem's different APIs in a user-friendly environment.

        SEPsando is offered as a NASA-hosted service and as open-source software runnable locally. Ongoing work adds multipoint in-situ overlays, ENLIL and EUHFORIA representations, and deeper CDAWeb, OMNIWeb, and COHOWeb integrations.

        Speaker: Fernando Carcaboso
    • 15:30 16:00
      Coffee 30m
    • 16:00 17:00
      Session 10: General Session 4 (Datasets & Science Enablement)
      • 16:00
        Science-Ready Dataset of GAMERA-GL CME Simulations for Forward and Inverse Modeling of the Interplanetary Magnetic Field at L1 15m

        Understanding Coronal Mass Ejections (CMEs) and their impact on the geomagnetic environment is among the most critical questions of space weather. The recent advances in physics-based CME modeling resulted in the development of extensive simulation datasets covering a broad range of scenarios and allowing data-intensive techniques, such as machine learning, to assist with the exploration and understanding. In this work, we provide an update on the science-ready dataset constructed based on the existing grid of the GAMERA-GL simulations of the CMEs propagating in the inner Heliosphere. The dataset has three background solar wind options (corresponding to the solar activity minimum, and its rising and falling phases) and has the Gibson-Low flux rope of varying properties initiated at different locations, resulting in ∼23,000 complete simulation runs and ~7.4M unique timeseries of solar wind properties at hypothetical L1 locations. We discuss the applications of this dataset to two problems: (1) development of the surrogate model for the CME time series at L1 point, and a related problem of the forecasting of geoeffective CME properties, (2) development of the inverse model constraining magnetic field properties of the CME close to the Sun based on the L1 time series dynamics and CME geometry, (3) the problem of the CME arrival times. The results highlight how combining the extensive simulation grids and machine learning approach can help us understand the CME dynamics and enhance space weather forecasting.

        Speaker: Viacheslav Sadykov (Georgia State University)
      • 16:15
        Multi-spacecraft study of ICME-ICME interactions, ICME-in-sheath and associated GLE #77 during 11-12th Nov 2025 15m

        We examine the propagation, interaction, and turbulence properties of a multi-CME sequence erupted between 9–11 November 2025, using multi-spacecraft in situ observations from Aditya-L1 (MAG, ASPEX), WIND (SWE), MMS (FGM), and Solar Orbiter (MAG, EPD; positioned at 0.83 AU, 13.5° inclination), supplemented by ground-based neutron monitor data. X1.7 and X1.2 flares on 9–10 November (07:36 and 09:48 UT) produced CMEs with speeds of 625 and 1644 km/s, respectively; the faster ejecta overtook the preceding one, forming a complex merged structure whose shock arrived at L1 at ~23:48 UT on 11 November, driving an intense geomagnetic storm (Dst_min = −217 nT). Quasi-perpendicular shock parameters were derived from peak ion flux enhancements (Aditya-L1/ASPEX, Solar Orbiter/EPD) and triaxial magnetic field discontinuities (Aditya-L1/MAG, Solar Orbiter/MAG), followed by a turbulent sheath. Velocity, density, temperature, energetic-ion, and IMF signatures indicate clear ICME-ICME interaction. A subsequent X5.1 flare (10:04 UT, AR 14274) generated a 1845 km/s CME whose leading edge overtook the trailing magnetic cloud of the 10 November ejecta, producing an ICME-in-sheath configuration with compressed, turbulent plasma and anomalous fluctuations observed between 10:00–12:00 UT on 12 November. MMS-derived PSD analysis yields a kinetic-scale spectral index of ≈−2.602, intermediate between KAW- and whistler-type turbulence, indicating enhanced dissipation. Neutron monitor data confirm GLE #77, exhibiting a rare double-peak, anisotropic profile. ENLIL simulations corroborate the merged-ejecta formation and subsequent interaction with the preconditioned heliosphere, elucidating shock-driven particle acceleration mechanisms relevant to SEP generation.

        Speaker: Mr Harsh Bhati
      • 16:30
        CROCS and VALERIE: Web-Based Radio Diagnostics and Source Localization in the Heliosphere 15m

        Solar radio bursts provide remote signatures of energetic electrons and CME-driven shocks, but their interpretation often depends on coordinated multi-spacecraft observations and reliable estimates of radio-source locations. We present two NASA Goddard Space Flight Center web tools designed to support these analyses. Coordinated Radiodiagnostics of CMEs and Solar Flares (CROCS; https://parker.gsfc.nasa.gov/crocs.html) provides time-selectable displays of radio flux and polarization measurements from Parker Solar Probe, Solar Orbiter, STEREO, and Wind, enabling comparative analysis of solar eruptive activity across multiple observing platforms. Visual Analytics for Localizing Emission from Radio-Source Interplanetary Events (VALERIE; https://science.gsfc.nasa.gov/swaves/valerie.html) uses STEREO/WAVES direction-finding measurements to estimate the locations of interplanetary radio sources. Its stereoscopic mode triangulates sources observed simultaneously by STEREO-A and STEREO-B (Krupar et al., Solar Physics, 2014). A newly developed single-spacecraft mode uses STEREO-A measurements together with the wavevector-corrected ray-sphere method (WCRS; Krupar et al., ApJL, 2026). We describe the underlying methods, present representative event analyses, and demonstrate how CROCS and VALERIE support practical investigations of solar and interplanetary radio emission.

        Speaker: Dr Vratislav Krupar (UMBC/GPHI & NASA/GSFC)
      • 16:45
        Building Heliophysics Datasets with Citizen Science 15m

        Citizen science can turn complex heliophysics mission data into high quality, reusable datasets at scale. We present three citizen science projects built on heliophysics observations, spanning different stages of maturity: (1) a Magnetospheric Multiscale (MMS) project on differentiating magnetosheath types with a data paper ready for submission; (2) a second MMS project on boundary layers identification launched and collecting data; and (3) a THEMIS and TREx all sky imager project for auroral identification nearing launch.

        Across these projects, we use visual classification tasks to have volunteers identify boundaries and morphologies that are challenging to capture with automated methods alone. Methodological commonalities include: task decomposition tailored to non experts, intuitive data displays, detailed tutorials, and quality controls such as redundant classifications, expert benchmarks, and consistency checks. These design choices enable the construction of well documented labeled datasets suitable for traditional analysis and machine learning applications.

        We will compare lessons learned across the three development stages, showing how careful citizen science methodology can produce robust, community ready MMS and THEMIS data products and inform future heliophysics citizen science efforts.

        Speaker: Vicki Toy-Edens
    • 17:00 17:30
      Wrap-up 30m
    • 19:00 22:00
      Conference Dinner 3h

      https://indico.dias.ie/event/2/page/24-conference-dinner

    • 09:00 09:10
      Welcome 10m
    • 09:10 10:30
      Space Agencies and Institutions: Reports (new members)
      • 09:10
        Chinese Academy of Sciences 25m
        Speaker: Zimming Zou
      • 09:35
        Canadian Space Agency 25m
        Speaker: Bill Archer
      • 10:00
        JAXA report 15m
        Speaker: Yoshi Miyoshi
      • 10:15
        Report from IUGONET project 15m

        IUGONET (Inter-university Upper atmosphere Global Observation NETwork) is a collaborative project among Japanese universities that aimed at promoting the publication, sharing, and utilization of upper atmospheric data. These activities are carried out in accordance with the international standards recommended by the International Heliophysics Data Environment Alliance (IHDEA), including metadata models, protocols, and tools. In this presentation, we report our recent activities within the IUGONET project, including: (1) the publication of new datasets from IUGONET institutions and the creation of metadata based on the Space Physics Archive Search and Extract (SPASE) metadata model; (2) the development and public release of plugins for the Python-based data analysis tool, PySPEDAS; (3) the organization of data analysis workshops; and (4) the registration of Digital Object Identifiers (DOIs) using SPASE metadata. In addition, we will present the project's future plans.

        Speaker: Yoshimasa Tanaka
    • 10:30 11:00
      Coffee 30m
    • 11:00 12:40
      Space Agencies and Institutions: Reports
      • 11:00
        CNES/Observatoire de Paris report 15m
        Speaker: D Boucon (TBC) / B Cecconi
      • 11:15
        ESA report to IHDEA 2026 15m

        A lot has happened at the European Space Agency (ESA) regarding Heliophysics archives since the last IHDEA meeting held in Texas late 2025. A new version of the overarching HelioPhysics Archive (HPA) interface was setup. This version not contains a brand new Graphical User Interface for Solar Orbiter but also the SMILE mission archives and a first version of 3 legacy arcvhies. The HPA GitHub for the SPASE description of the heliophysics mission datasets now contains the description of all Cluster datasets and the description of datasets from three Solar Orbiter instruments. Highlights of these new features will be presented.

        Speaker: Arnaud Masson
      • 11:30
        ISRO report 15m
        Speaker: Sankar K
      • 11:45
        KASI report 15m
        Speaker: Ji-Hye Baek
      • 12:00
        NASA DISH report 10m
        Speaker: Brian Thomas
      • 12:10
        Updates from NASA's Solar Data Analysis Center 10m

        NASA's Solar Data Analysis Center (SDAC) supports the scientific analysis of NASA's free and open solar physics mission science data, enabling scientists to address the goals of NASA's Heliophysics Division (HPD). The SDAC accomplishes this by supporting science-enabling projects and curating the storage of NASA's free and open solar physics mission science data. The SDAC is part of NASA's Heliophysics Digital Resource Library (HDRL), a body that develops and implements coordinated data and scientific infrastructure strategies supporting NASA's HPD. This presentation will outline recent advances in SDAC-supported projects. This includes the Helioviewer Project, a data visualization and discovery effort, the Virtual Solar Observatory (VSO), a federated data search and download project, and SDAC Watch, a new effort to monitor and track data held at the SDAC. Finally, I will also describe a new effort to transfer and store petabytes of Solar Dynamics Observatory data to the SDAC.

        Speaker: Dr Jack Ireland (NASA Goddard Space Flight Center)
      • 12:20
        Annual Report of NASA’s Space Physics Data Facility 10m

        NASA’s Space Physics Data Facility (SPDF https://spdf.gsfc.nasa.gov) is the active and final archive for space physics data from NASA missions and joint missions with other national or international agencies. The data covers the space from 10 solar radii to the local interstellar medium (about 160 astronomical units), and the magnetosphere - ionosphere - thermosphere - mesosphere (M-ITM) of Earth and other applicable planets. SPDF serves the global science community with open access to about 800 TB of space physics data to enable correlative and collaborative research across discipline and mission boundaries with more than 100 operating and past missions/projects. Besides updating the self-describing Common Data Format (CDF) and the International Solar-Terrestrial Physics (ISTP) metadata guidelines, SPDF has been improving three main science-enabling services: Coordinated Data Analysis Web (CDAWeb), Satellite Situation Center Web (SSCWeb), and OMNI Web. In this annual report, we report the latest status of data archiving at SPDF and copying the data into HelioCloud for cloud-based collaborative data analysis. We highlight some new features and tools of SPDF, including the improved browser-based 4-D orbit viewer with simulated time-dependent bow shock and magnetopause, the improved browser-based ISTP metadata editor, a new solar energetic particle portal and propagation tool, a new capability of responding to queries using Large Language Models (LLMs) and Model Context Protocol (MCP). These updates facilitate the production, access, analysis, and citation of space physics data and other digit resources, therefore ultimately support Open Science and enhance the productivity for the science community.

        Speaker: Lan Jian (NASA Goddard Space Flight Center)
      • 12:30
        NASA CCMC 10m
        Speaker: Leila Mays
    • 12:40 14:00
      Lunch 1h 20m
    • 14:00 15:30
      IHDEA Working Group: Reports 1
      • 14:00
        IHDEA SPASE Working Group report 15m
        Speaker: Arnaud Masson/ Brian Thomas
      • 14:15
        IHDEA Standards Process Working Group report 30m
        Speaker: Brian Thomas
      • 14:45
        Updates from the IHDEA Data Access Working Group 15m

        Improved data access remains a key area in which standards can positively affect heliophysics research. As AI and machine learning become increasingly important for science, the ability to quickly ingest clean, well-structured data is essential, since data preparation often accounts for a substantial portion of analysis and machine-learning workflows. The Data Access Working Group is seeking to expand participation among data providers interested in standards-based approaches to data access. We are planning a community survey and will begin holding periodic meetings on topics related to identifying and promoting data access standards. We will also collaborate with the IHDEA Standards Process working group. The HAPI specification, in particular, is ready to move through a formal process toward adoption as an IHDEA standard. Finally, we will provide a brief overview of the status of data-access mechanisms at several relevant providers, including the VSO and ESAC.

        Speaker: Jonathan Cook
      • 15:00
        Report on the SPASE metadata registry for the ESA’s Heliophysics missions' activities 15m

        Improving data discoverability and interoperability is one of the ESAC Science Data Centre’s (ESDC) priorities for promoting the open exchange of data from ESA’s Science Archives, including the Heliophysics Archive (HPA). To accomplish this goal, ESA is targeting to mint a DOI per dataset available in the ESDC archives. Typically, each major domain of space science makes use of their metadata standard to seek for the necessary information to mint a DOI.

        The HPA makes use of the Space Physics Archive Search and Extract (SPASE, https://spase-group.org/) metadata model, created and maintained to meet the heliophysics community needs, to describe the calibrated and value-added datasets of instruments from all the ESA Heliophysics missions. The HPA makes all datasets’ SPASE resources publicly available via the HPA GitHub (https://github.com/HPA-ESDC-ESA-INT/SPASE/) after a review by the respective Principal Investigator (PI) team. Currently available in the HPA GitHub are nearly all instruments calibrated and derived datasets from Cluster, and three out of four in-situ instruments datasets for Solar Orbiter. In the near future, SPASE dataset descriptions will be made available for Solar Orbiter remote sensing instruments and progressively for SOHO instruments. Their generation started with a manual review of the automatic Common Data Format (CDF) to SPASE conversion by the ADAPT tool, developed at UCLA. It is now moving to a semi-automatic generation based on Python script accessing Table Access Protocol (TAP) server from the ESDC archives. SPASE to DataCite and Schema.org converters will then be used to mint the DOIs and to generate the SPASE landing pages' embedded metadata responsible for enhanced discovery via search engines, respectively.

        Speaker: Joana S. Oliveira (Telespazio UK for ESA)
      • 15:15
        Persistent IDentifiers (PIDs) for Science 15m
        Speaker: Rebecca Ringuette
    • 15:30 16:00
      Coffee 30m
    • 16:00 17:20
      IHDEA Working Group: Reports 2
      • 16:00
        Identifying spacecraft relative location in heliospheric environments 15m

        The IHDEA working group “Identifying spacecraft relative location in heliospheric environments” started its activity this year (2026).

        The motivation of this working group can be summarized in the following way. The plasma region-based location of a spacecraft is a key information for scientists working with in-situ measurements, both for supportive information and for data analysis. Coordinate systems are a standard way to locate spacecraft in space. However, due to the variability in time and space of plasma regions (interplanetary medium, plasmasphere, magnetospheric tail, …), coordinate systems do not suffice to locate the plasma region or the boundary layer (magnetopause, plasmapause, …) crossed by a spacecraft at any given time.
        Region classifications are performed using three different approaches: data visualisation, model dependent, or machine-learning based. Classifications are performed for each kind of missions, in various environments (e.g., Cluster, Arase, Juno, ...).
        The time needed to make the calibrated data available to the scientific community has improved greatly over the last decades. These data come with classical orbital information and at times with model predictions. However, the effective location (by regions) is not included in the data delivery plan of a scientific mission. The recent increasing number of region classification studies raises the question of a standard for these regions.
        A unified, classified dataset will enable us to use the richness of data from missions all over the solar system to understand fundamental plasma physics: the localized nature of in-situ data in dynamic systems creates large uncertainties that can obscure processes signatures. Statistical studies on large datasets are essential to improve the signal-to-noise ratio and detect physical signatures indiscernible in single-event analyses.

        The first results of the working group concerns the regions shall be included, how to identify, and to define these regions. The first level of classification is inside or outside the heliosphere. The second level of classification will include the interplanetary medium, magnetospheres, or cometary environments. The third level will include regions inside each of level 2 regions. At this stage we mainly focus on the definition of the transition region between inside and outside the heliosphere with the help of the NASA/Voyager spacecraft.

        Speaker: Benjamin Grison (Institute of Atmospheric Physics CAS)
      • 16:15
        Standardizing Reference Frame Terms in Heliophysics 15m

        We provide an update on the Reference Frame Standardization working group.

        Speaker: Dr Bob Weigel (GMU)
      • 16:30
        Heliophysics Semantics and Capabilities 15m
        Speaker: Baptiste Cecconi
      • 16:45
        Python in Heliophysics Community (PyHC) 15m
        Speaker: Julie Barnum
      • 17:00
        Science platforms coordination 20m
        Speaker: Shawn Polson
    • 17:20 17:30
      Announcements 10m
    • 09:00 09:10
      Welcome 10m
    • 09:10 10:30
      SPASE 3.0: Report and Discussion
      Convener: A. Koval/ R. Ringuette
      • 09:10
        SPASE 3.0 Data Model Development Status 1h 20m

        SPASE (Space Physics Archive Search and Extract) is the de facto metadata standard internationally used in Heliophysics primarily for observatories, instruments, and the data they produce. The development of SPASE 3.0, the next generation of the SPASE data model, is driven by the need to increase support for FAIR principles and to solve difficulties with creating and managing SPASE metadata. The SPASE 3.0 model is designed as a Data Catalog Vocabulary (DCAT) - Version 3 application profile, significantly extended with other RDF (Resource Description Framework) vocabularies and ontologies, primarily SOSA (Sensor, Observation, Sample, and Actuator), I-ADOPT (InteroperAble Descriptions of Observable Property Terminology), and QUDT (Quantities, Units, Dimensions and DataTypes), and with addition of a small number of SPASE-specific properties. We discuss the original functional requirements based on the community-accepted Space and Solar Physics use cases, the SPASE 3.0 model design and schema, resource description examples, and the improvements over the SPASE model current version, We will also discuss the timeline for progress and release of the SPASE 3.0 schema.

        Speaker: Andriy Koval (UMBC, NASA GSFC)
    • 10:30 11:00
      Coffee 30m
    • 11:00 12:30
      IHDEA Open Session: Oral Presentations 1
      • 11:00
        Pathways for enhancing FAIRness in Heliophysics 15m

        The OSTrails project (https://ostrails.eu/) is proposing a framework supporting open science through three pillars: Plan, Track and Access. The team at Observatoire de Paris is leading the Astronomy thematic pilot with the project, focussing here on the elements concerning the MASER service (Measuring, Analysing and Simulating Emissions in the Radio range, https://maser-lira.obspm.fr/), which deals with heliophysics and low frequency radioastronomy.

        The implementation of the three pillars is as follows:

        1. Plan: adopting and implementing a Data Management Plan tool, to streamline the data life cycle management and prepare the publication of the MASER datasets;
        2. Track: linking datasets, services, instruments, researchers and institutions in a knowledge graph, to enhance data discoverability as well as better measuring the impact of the MASER service;
        3. Assess: design adapted FAIR evaluation metrics and tests for astronomy, to assess the FAIRness of the research products published in MASER.

        We present the status of the OSTrails developments, how this applies to MASER and the astronomy community.

        This project has received funding from the European Union’s Horizon
        Europe framework programme under grant agreement No. 101130187.

        Speaker: Baptiste Cecconi (Observatoire de Paris)
      • 11:15
        Open solar data, data products, and tools at MEDOC 15m

        MEDOC (Multi-Experiment Data and Operation Centre), initially created as a European data and operation centre for the SoHO mission, has grown with data from other solar physics space missions, from STEREO to SDO, and now Solar Orbiter. In addition to observational data, MEDOC also provides datasets derived from observations (maps, catalogues...), tools for data analysis and interpretation, and numerical simulation results. We will present the current and future MEDOC interfaces, including APIs and VO services, data formats, implementation of DOIs, and how they contribute making MEDOC data Findable, Accessible, Interoperable, and Reusable (FAIR).

        Speaker: Barbara Perri (CEA - AIM)
      • 11:30
        Community Coordinated Modeling Center Collaborative FAIR Efforts 15m

        The Community Coordinated Modeling Center’s (CCMC) at NASA, was established in 2000 to facilitate space weather and space science research, and facilitate the transition of research to operational products. The CCMC hosts a collection of space weather models and coupled modeling systems for Run-on-Request (RoR) simulation services, and continuous real-time runs, and creates space weather dashboards, scoreboards, and validation platforms displaying these real-time runs and together with input and context data. This presentation provides an overview of CCMC’s current and future planned collaborative activities in implementing FAIR (Findable, Accessible, Interoperable, and Reusable) principles in order to facilitate access to simulation outputs across international boundaries. This is includes FAIR advancements in the RoR portal, Instant Runs, Continuous Runs available via the Integrated Space Weather Analysis (ISWA) System, Scoreboards, and in our current and planned collaborations with standards and services across many groups including the VSWMC, ROB, UKMO, HDRL (DISH, SDAC, SPDF) in interconnecting our services for the benefit of the heliophysics community.

        Speaker: Leila Mays
      • 11:45
        Proper intermodel comparison using event detection skill scores 15m

        Skill scores are useful secondary data-model comparison metrics that transform the value of another metric relative to that for a reference model. A well-known skill score for "continuous metrics analysis – those comparison techniques that employ the exact values of observations and model results – is prediction efficiency (PE), based on mean square error as the metric and the average of the observed values as the reference model. There is no equivalent skill score like PE for event detection analysis – those comparison techniques that convert the exact values into yes-no event status and create metrics from the contingency table counts. Here, we present two such options, one based on the proportion correct metric and another based on the critical success index metric. Like PE, these new skill scores use the observations as the reference model, which provides complete independence of the reference model from the accuracy of the new model. It is demonstrated that these skill scores provide context for model evaluation that is unique to other existing event detection metrics and valuable for the assessment, in particular with respect to comparing the new model's performance against an existing model. It is advocated that these new skill scores should be incorporated into space physics analysis software packages.

        Speaker: Michael Liemohn (University of Michigan)
      • 12:00
        The HelioData API: A Faceted, Vocabulary-Driven REST Interface for Unified Solar- and Space-Physics Dataset Discovery 15m

        Heliophysics dataset discovery has historically been partitioned by subdiscipline and by archive, forcing researchers and client developers to query solar and space-physics holdings through separate interfaces with incompatible query semantics. The HelioData API (https://api.heliophysics.net/api) addresses this with a single REST/JSON interface over the SPASE-described holdings unified by HDRL.

        The service exposes a browsable root of eight collections: datasets (7793 registered datasets) plus seven controlled vocabularies — observatories (3112), instruments (4494), quantities (1302), repositories (199), regions (24), measurementtypes (21), and spectralranges (12). Rather than relying on free-text search, each vocabulary endpoint returns stable identifiers (earth-magnetotail, MagneticField, numberflux-differential:alphaparticle) that compose directly into faceted dataset queries. The controlled vocabulary thus becomes the query grammar: a client enumerates the legal facet space before querying, eliminating the guesswork and brittle string matching typical of federated science search. The quantities vocabulary uses a compound quantity:species identifier scheme, and the region facets derive from HDRL's contribution of a ~200-concept heliophysics taxonomy to the Unified Astronomy Thesaurus, giving the facets an externally governed semantic basis. The 199 registered repositories span NASA, ESA, JAXA, NOAA, NSF, and university archives, making the index genuinely cross-agency.

        I will describe the service architecture, the SPASE-to-API mapping, and its use in an ongoing CCMC collaboration to resolve observational datasets relevant to a given model run. Plans to further map between the API concepts and other vocabularies will be discussed.

        Speaker: Dr Brian Thomas (NASA)
      • 12:15
        Open discussion on FAIR 15m
    • 12:30 14:00
      Lunch 1h 30m
    • 14:00 15:30
      IHDEA Open Session: Oral Presentations 2
      • 14:00
        Connecting Heliophysics SPASE Metadata: Building a Modernized System 15m

        Data and Integration Services for Heliophysics (DISH) at HDRL/NASA is modernizing its metadata pipelines by building more automated, interoperable systems. Rich SPASE records continue to play a central role in support of science discovery, with much of their metadata often sourced from other systems. We are modernizing these metadata acquisition processes by programmatically connecting SPASE to other metadata sources relevant to Heliophysics (e.g., FITS, HAPI, and ISTP). Even with modernized pipelines, significant additional work is needed to determine dataset-level fields and translate data-specific terminology to science-wide vocabularies. So, we will also present updates on AI tooling and ontological development that are beginning to address those needs with building momentum. Connecting the data, instruments, and observatories described in SPASE to other resources such as the software and research data that cite them is another important component needed to support science discovery. However, current external infrastructure supporting these citations are problematic, with many citations never having any effect. We will describe the status of these efforts and point towards improvements to come, including our new NASA Heliophysics Research Repository on Zenodo with semi-automated curation.

        Speaker: Rebecca Ringuette
      • 14:15
        Reducing Friction in HDRL’s SPASE Pipelines to Operationalize FAIR with mEditor 15m

        We present on the efforts at HDRL to improve our infrastructure to support Open Science. We focus on progress setting up our new, open-source content management system (CMS) for SPASE at HDRL using NASA-developed mEditor. Introducing mEditor into our pipeline will drastically improve the standardization of our SPASE metadata content and quality and ease of creating and editing SPASE records, which is notoriously difficult to do. mEditor provides a user-friendly interface and streamlined operation for editing SPASE records, made possible through an assembly of external scripts for a variety of automated workflows. This increased speed and efficiency in SPASE record modification/creation will lead to more DOIs, high quality HelioData pages, and discovery across the board for NASA Heliophysics data and resources.

        Speaker: Zach Boquet (DISH @ NASA GSFC's HDRL; Adnet Systems (United States))
      • 14:30
        Reducing Friction in HDRL’s SPASE Pipelines to Operationalize FAIR via Implemented SPASE Mappings 15m

        Mapping from SPASE metadata to Schema.org and DataCite improves Heliophysics infrastructure to support Open Science, primarily citation, discoverability and FAIR (Findability, Accessibility, Interoperability, and Reusability). This presentation is a follow-up on a previous poster given at the 2025 DASH meeting introducing these mapping scripts. We will summarize the work responsible for implementing these mappings into operations and the observed impact – including improving the ranking of our SPASE landing pages in internet searches, large increases in DOI creation and synchronization, and results from externally developed FAIR assessment tools. Following the SPASE-DataCite script’s open source release, both of these mapping tools will enable others to achieve similarly improved results.

        Speaker: Zach Boquet (DISH @ NASA GSFC's HDRL)
      • 14:45
        ISTP Metadata Guidelines: Current Status and Future Development 15m

        The International Solar-Terrestrial Physics (ISTP) Metadata Guidelines for definitive description of data in the self-describing scientific data formats, such as CDF and netCDF in Heliophysics, were originally developed at NASA Space Physics Data Facility (SPDF) as part of the ISTP science initiative for use by the Polar, Wind, Geotail and Cluster missions in the early 1990’s. The Guidelines were later adopted by the Interagency Consultative Group (IACG) for Space Science, and they are now used by most of the major heliophysics missions, including MMS, Van Allen Probes, Parker Solar Probe, Solar Orbiter and IMAP, and supported by numerous display and analysis tools, including CDAWeb, Autoplot, SpacePy and PySPEDAS/SPEDAS. The ISTP Metadata Guidelines are maintained openly in a publicly accessible, version-controlled GitHub repository (https://github.com/IHDE-Alliance/ISTP_metadata) that allows for contributions and engagement from the community. The baseline version 1.0.0 has also been published on Zenodo (https://doi.org/10.5281/zenodo.20790266). Though the ISTP Metadata Guidelines proved to be critical in enabling collaborative research in Heliophysics in the past 30 years, they are not always sufficient for definitive description of complex data from modern advanced instruments. They also do not provide adequate support for FAIR (Findability, Accessibility, Interoperability, and Reusability) principles. To enable the Guidelines future development, we have established the ISTP Metadata Guidelines Steering Committee made of experts in the field with extensive experience of using the ISTP metadata, including representatives/developers of systems/tools supporting the ISTP metadata, with all development proposals publicly discussed on GitHub. We are presenting the ISTP Metadata Guidelines current status and prioritized future development steps, and we welcome feedback from the wide community.

        Speaker: Andriy Koval (UMBC, NASA GSFC)
      • 15:00
        Open discussion on ISTP metadata 30m
    • 15:30 16:00
      Coffee 30m
    • 16:00 16:30
      Final IHDEA Session
      • 16:00
        Developing Heliophysics Standards and Best Practices for Modeling Software 30m
    • 16:30 17:00
      Final words / wrap-up 30m