Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “analysis workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

GeneLab: A Systems Biology Platform for Omics Analysis

NASA GeneLab is an open-access repository for omics datasets generated by biological experiments conducted in space or experiments relevant to spaceflight (e.g. simulated cosmic radiation, simulated microgravity, bed rest studies). The GeneLab Data Systems (GLDS) version 4.0 will be available on October 1st 2019, and will provide the latest in terms of professional state-of-the-art bioinformatics platform for the space biology and radiation community to upload their data into an omics data commons, to process their data with vetted standard workflows and to compare to existing analyses. Started in 2015 as a repository designed to archive omics data from space experiments, GeneLab has expanded its scope to all ionizing radiation omics experiments conducted on the ground and has put considerable effort in providing carefully characterized radiation metadata on all dataset. GeneLab is also providing processed data derived from the raw data covering a large spectrum of omics (genome, epigenome, transcriptome, epitranscriptome, proteome, metabolome) to help users explore important questions: 1) Which genes or proteins are expressed differently in space for various living organisms? 2) What specific DNA mutations or epigenetic changes happen in space or after exposure to ionizing radiation? and 3) How does genetics affect these responses? Processed data available on GeneLab are derived by standard data analysis workflows vetted by hundreds of scientists who volunteered to join one of the four GeneLab Analysis Working Groups (Animal AWG, Plant AWG, Microbe AWG, Multi-Omics AWG). In this presentation, we will discuss how to bridge the gap between irradiation studies performed on earth and biological experiments conducted in space since the early 1990's. We will discuss how radiation dosimetry was estimated for datasets derived from samples collected during the Space Shuttle era or on the International Space Station. Finally, we will address future strategies regarding dose monitoring in future missions into space, inter-agency efforts to unify data under one umbrella, and knowledge dissemination across the radiation research community and the space biology community.

open-science↗

GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics

Data package for Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the below citations for the data packages and associated manuscript. Please cite as: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. GLBRC Soil Yearlong Incubation 13C-SIP-Lipidomics. [Data Set] PNNL DataHub. doi: Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. MSV000097435: GLBRC soil yearlong incubation 13C-SIP-Lipidomics [Data Set] MassIVE. doi:10.25345/C57659T3K Rempfert KR, Bell SL, Kasanke CP, Kyle JE, Hofmockel KS. 2025. Lipids represent a dynamic, yet stable pool of microbially-derived soil carbon. In Prep This data package consists of compound-specific 13C SIP-lipidomics data from a yearlong tracer incubation experiment designed to investigate microbial lipid persistence in switchgrass bioenergy crop soils. In order to explore how lipid structure may modulate the persistence of C in soil lipids, we leveraged soils from two sites (Michigan - sandy texture, Wisconsin - silty texture) operated by the U.S. Department of Energy-funded Great Lakes Bioenergy Research Center (GLBRC). These sites had comparable climates, identical management practices, but contrasting soil textures, allowing us to assess the variability of lipid accrual or degradation in soils as well as provide insight regarding the degree to which edaphic properties may regulate the retention of soil lipids. Untargeted lipidomics analyses were performed to identify 13C-labeled lipids in the soil microbiome after long-term incubation. Soils were supplemented with 100 micrograms glucose per gram dry soil (99 atom % 13C or natural abundance for paired control) and incubated; samples were collected two months and one year after glucose addition. Lipid extracts (MPLEx) were analyzed by LC-MS/MS and identified using LIQUID. Calculation of isotopic enrichment of lipids was performed by targeted approach using TarMet to quantify lipid isotopologues and IsoCorrectoR to correct for natural abundance isotopes. Contents: Data package contents reported here are the first version and contain downstream analysis files for the raw LC-MS mass spectrometry files (.mzXML) deposited at the MassIVE database repository under accession MSV000097435 (80 experimental runs; 5.85 GB) | MassIVE DOI: 10.25345/C57659T3K. Support files include the additional data download 'Read Me' file containing data descriptor information. Reported data download contents are structured for compliance with project data sharing guidelines, community standards initiatives, and sponsor stakeholder policies supporting FAIR data principles. Data processing software, analysis tools, and data workflows are listed below corresponding to the host repository long-term location. Available Data Downloads (0.3 GB): "GLBRC soil yearlong incubation 13C-SIP-Lipidomics_readme.txt" - 'Read Me' data package content file (txt) "GLBRC_DataPackage_analysis files" - Data processing files (Rmd) and saved intermediate data processing outputs (rds, csv, xlsx) "GLBRC_13C_lipidomics_dataset.xlsx" - processed data in tabular format (xlsx) Linked Software: LIQUID LC-MS Analysis Software | 10.5281/zenodo.6459462 Lipid Mini-On Software Tools | 10.5281/zenodo.1492803 pmartR Omics Statistical Software | 10.5281/zenodo.6108667 xcms (v4.3.3) TarMet (v1.1.1) IsoCorrectoR (1.24.0) Funding Acknowledgments: This research was supported by an Early Career Research Program award funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research (OBER) Genomic Science program under FWP 68292, FWP 07880 and EMSL Exploratory Research Project 51095. A portion of this work was performed in the William R. Wiley Environmental Molecular Sciences Laboratory, a national scientific user facility sponsored by OBER and located at Pacific Northwest National Laboratory (PNNL). PNNL is a multi-program national laboratory operated by Battelle for the DOE under Contract DE-AC05-76RLO1830.

Rempfert, Kaitlin R [Pacific Northwest National La↗

pyNuMAD v.0.1

SAND2024-08606O The pyNuMAD software is used for managing wind turbine blade model data. pyNuMAD specializes in defining the geometry, materials, and boundary conditions for structural analysis of wind turbine blades. This includes loading in data files and providing an interface for users to make updates to the model. The software also features meshing functionality, which takes the blade model and creates a shell or brick mesh for use in finite element analysis. A typical user workflow might be: load in blade information from a yaml file, make adjustments to the blade properties, update the blade based on the adjustments, create a mesh of the blade, export this blade to another software for structural analysis. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Paquette, Joshua↗

MOOSE Reactor Module: An Open-Source Capability for Meshing Nuclear Reactor Geometries

The U.S. Department of Energy (DOE) Nuclear Energy Advanced Modeling and Simulation (NEAMS) program has developed numerous physics solvers utilizing the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) framework for multiphysics reactor analysis. These solvers require input finite element meshes representing the discretized spatial domain. Typically, reactor analysts turn to licensed tools for the creation of reactor geometry meshes. Recently, open-source functionality has been added to the MOOSE framework to mesh common reactor geometries and improve MOOSE-based nuclear reactor application user workflows. The new functionality is primarily contained in the new Reactor module of MOOSE and includes support for hexagonal pins, assemblies, and cores, extended Cartesian geometry support, options for modeling static and rotating control drums within a hexagonal assembly, core periphery triangulation, and automatic tagging of pin, assembly, plane, and depletion regions for easier post processing of physics results. A set of reactor geometry mesh builder objects further streamlines the construction of hexagonal and Cartesian cores and allows mapping of materials to regions during mesh generation. The meshes produced with the MOOSE Reactor module may be used directly within MOOSE-based applications or exported as Exodus II files for use in other finite element solvers. The tools have been demonstrated and verified using a variety of NEAMS physics solvers on a range of reactor applications, including a sodium-cooled fast reactor core analysis using Griffin, a fast reactor assembly thermal deformation analysis using MOOSE Tensor Mechanics, and a heat pipe–cooled microreactor coupled analysis using Griffin, Bison, and Sockeye. MOOSE’s Reactor module provides significant advantages compared to the use of external meshing tools when analyzing Cartesian and hexagonal reactor lattices using MOOSE-based applications: immediate accessibility (open-source) to the end user, low barrier to entry for new users, speed of mesh generation, volume preservation of meshed fuel pins, and simplification of analysis workflow when used in conjunction with MOOSE-based applications.

99 GENERAL AND MISCELLANEOUS↗

Depletion Benchmark Analysis on a Lead Fast Reactor Using PyARC/OpenMC

PyARC is a user-friendly fast reactor analysis tool that automates multiphysics workflows using the “extended suite” of Argonne Reactor Computation (ARC) codes by providing a single common input for model definition, code execution, and output post-processing. A lead fast reactor (LFR) benchmark model is used to perform depletion calculations using the newly integrated OpenMC depletion capability in PyARC, building on previous analysis using the ARC codes through PyARC and Serpent. Results for core lifetime k-effective, shutdown decay heat, and end-of-life heavy-metal inventory are compared to verify the PyARC/OpenMC integration against the PyARC/ARC workflow and Serpent for depletion analysis of LFR designs. The results show satisfactory agreement among all three methods, with remaining discrepancies largely attributable to differences in nuclear data libraries and decay-chain modeling detail rather than to fundamental modeling limitations.

Kiesling, Kalin R.↗

Sequence Design of Random Heteropolymers as Protein Mimics

Random heteropolymers (RHPs) have been computationally designed and experimentally shown to recapitulate protein-like phase behavior and function. However, unlike proteins, RHP sequences are only statistically defined and cannot be sequenced. Recent developments in reversible-deactivation radical polymerization allowed simulated polymer sequences based on the well-established Mayo–Lewis equation to more accurately reflect ground-truth sequences that are experimentally synthesized. This led to opportunities to perform bioinformatics-inspired analysis on simulated sequences to guide the design, synthesis, and interpretation of RHPs. We compared batches on the order of 10000 simulated RHP sequences that vary by synthetically controllable and measurable RHP characteristics such as chemical heterogeneity and average degree of polymerization. Our analysis spans across 3 levels: segments along a single chain, sequences within a batch, and batch-averaged statistics. We discuss simulator fidelity and highlight the importance of robust segment definition. Examples are presented that demonstrate the use of simulated sequence analysis for in-silico iterative design to mimic protein hydrophobic/hydrophilic segment distributions in RHPs and compare RHP and protein sequence segments to explain experimental results of RHPs that mimic protein function. To facilitate the community use of this workflow, the simulator and analysis modules have been made available through an open source toolkit, the RHPapp.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Assessing the Needs of NASA's Near Real-Time Earth Observation Products

"The 2017-2027 Decadal Survey for Earth Science and Applications from Space stated that NASA's Earth Science with planned implementation of applications provides sustained earth observations for societal benefits [1]. The Decadal Survey indicated that data latency is invaluable for time-sensitive applications including disaster risk reduction, wildland fire carbon emissions quantification, real-time measurements of the state of the hydrologic systems and many more. Data latency refers to the time between earth observation and data products available to users. During the past 13 years, NASA's Land, Atmosphere Near Real-Time Capability for Earth Observing Systems (LANCE) continues to provide free access to earth observation products that are made available much quicker than routine processing allows. The latency of most LANCE data products is Near Real-time (NRT) which is defined as less than three hours from satellite observations [2]. LANCE is managed by the Earth Science Data and Information System (ESDIS) Project at NASA Goddard Space Flight Center [3], and a User Working Group (UWG) is responsible for providing guidance to LANCE. LANCE data are used by direct users and brokers who add value to the data [4]. NASA Earth Applied Sciences Program (ASP) is one of the primary users of LANCE, which collaborates with partner organizations and provides support to scientists to solve problems in applications of earth observations. ASP promotes the use of LANCE NRT data products to demonstrate applications in decision making, facilitates end-user feedback to the science team to improve data products, and provides information on future demands for research. LANCE supports applications that need a rapid response including detecting wildland fires and volcanic eruptions, tracking smoke, ash and dust plumes, monitoring air quality and tracking extreme weather events such as hurricanes, landslides, and floods. To gather feedback regarding the availability, accessibility and actionability of NASA's NRT data products for societal benefit, three surveys and a few discussions with experts involved in the topic within ASP were conducted from the perspective of users. Feedback has been collected from users who are interested in using low latency NASA data within application communities of agriculture, disasters, water resources, health and air quality, ecological conservation, wildland fires and capacity building. Analysis-ready NRT data products in a variety of formats have been mentioned many times in the collected feedback, especially for applied users with little to no experience using research-grade earth observation products. Users prefer to have products that can be easily integrated into their existing workflows and take their analysis to the data. HDF5 is a commonly used data format for research, but typically requires some conversion to a more friendly format for applications and regular use in decision-making. Users prefer the GeoTIFF data format that can be directly ingested into a GIS mapping software and platform for data analysis and visualization. For example, LANCE’s fire, flood, SO2 and Black Marble Nighttime Blue/Yellow Composite data products have been integrated into NASA Disasters Mapping Portal, which is an GIS-based open data portal, for users in the disaster management community. There are 291 LANCE NRT layers available through GIBS and Worldview, where users can download a snapshot in GeoTIFF format. Operational users expect data to be processed as close to the user as possible. The collected feedback indicates that LANCE fire products within 3 hours latency would meet the needs of the wildland fire community. The ideal latency for volcanic application is 10-15 minutes. Users in Volcanic Ash Advisory Centers (VAAC) reported that the first forecast volcanic product should be issued within 75 minutes from the volcano eruption [5]. Overall, for disaster applications, data latency within 3 hours is useful while latency greater than 12 hours is not timely enough for operational use. Capacity building and training are critical for users to be able to access, interpret and use data products and tools for their decision making, especially for applied users with limited experience using earth observation products. LANCE data products have been used in a number of capacity building projects domestically and internationally [6]. As LANCE continues to bring new products into the system, users request training to utilize LANCE new and upcoming data products and capabilities in their applications. Due to the limitation of bandwidth and downstream flow paths, users in some developing countries need tools to select and download data for a specific area of interest instead of bulk downloads. The collected feedback also shows the lack of available SAR satellite low latency data products. The advantages of SAR to monitor conditions and changes on the ground through darkness, clouds, volcanic ash, and other atmospheric conditions, are appealing to low latency users. For example, terabytes of low latency but cloudy optical images are not helpful in rapidly identifying the extent of flood or fire impacts. LANCE could be complemented with low latency measurements via the upcoming NASA-ISRO Synthetic Aperture Radar (NISAR) mission [7]. Requests for higher spatial resolution products are expressed. A user from the wildland fire management community reported that products with 30-m spatial resolution could be used to detect small fires. The 30-m Landsat OLI fire data is now part of NASA’s Fire Information for Resource Management System (FIRMS) US/Canada [8]. Within the open and free NASA resources, LANCE disseminates NRT data products in a manner that allows them to be accessible and understandable to both scientific and applied users. In many application areas, latency plays an important or even decisive role where low latency earth observations help people to observe areas of interest, detect and track changes in the environment and make timely decisions. NASA’s Earth Applied Sciences Program promotes the use of LANCE NRT products and builds a bridge between application users and research teams. The collected feedback indicates data latency within 3 hours is useful for most of the applications, and shows the needs of user-friendly, analysis-ready products, and requests training on LANCE’s new and upcoming data products. User feedback has been provided to LANCE UWG for guidance and recommendations, and for translating findings into something actionable.

Tian Yao↗

Unsupervised Segmentation and Clustering Workflow for Efficient Processing of 4D-STEM and 5D-STEM Data

Four-dimensional scanning transmission electron microscopy (4D-STEM) enables mapping of diffraction information with nanometer-scale spatial resolution, offering detailed insight into local structure, orientation, and strain. However, as data dimensionality and sampling density increase, particularly for in situ scanning diffraction experiments (5D-STEM), robust segmentation of structurally consistent behavior across sequential measurements becomes essential for efficient and physically meaningful analysis. Here, we introduce a clustering framework that identifies crystallographically distinct domains from 4D-STEM datasets. By using local diffraction-pattern similarity as a metric, the method extracts closed contours delineating spatially contiguous regions. This approach produces cluster-averaged diffraction patterns that improve signal quality while reducing data volume by orders of magnitude, enabling rapid and accurate orientation, phase, and strain mapping. We demonstrate the applicability of this approach to in situ liquid-cell 4D-STEM data of gold nanoparticle growth. Our method provides a scalable and generalizable route for spatially coherent segmentation, data compression, and quantitative structure–strain mapping across diverse 4D-STEM modalities. The full analysis code and example workflows are publicly available to support reproducibility and reuse.

4D-STEM↗

A Dose of Reality: Radiation Analysis for Realistic Human Spacecraft

INTRODUCTION As with most computational analyses, a tradeoff exists between problem complexity, resource availability and response accuracy when modeling radiation transport from the source to a detector. The largest amount of analyst time for setting up an analysis is often spent ensuring that any simplifications made have minimal impact on the results. The vehicle shield geometry of interest is typically simplified from the original CAD design in order to reduce computation time, but this simplification requires the analyst to "re-draw" the geometry with a limited set of volumes in order to accommodate a specific radiation transport software package. The resulting low-fidelity geometry model cannot be shared with or compared to other radiation transport software packages, and the process can be error prone with increased model complexity. The work presented here demonstrates the use of the DAGMC (Direct Accelerated Geometry for Monte Carlo) Toolkit from the University of Wisconsin, to model the impacts of several space radiation sources on a CAD drawing of the US Lab module. METHODS The DAGMC toolkit workflow begins with the export of an existing CAD geometry from the native CAD to the ACIS format. The ACIS format file is then cleaned using SpaceClaim to remove small holes and component overlaps. Metadata is then assigned to the cleaned geometry file using CUBIT/Trelis from csimsoft (Registered Trademark). The DAGMC plugin script removes duplicate shared surfaces, facets the geometry to a specified tolerance, and ensures that the faceted geometry is water tight. This step also writes the material and scoring information to a standard input file format that the analyst can alter as desired prior to running the radiation transport program. The scoring results can be transformed, via python script, into a 3D format that is viewable in a standard graphics program. RESULTS The CAD model of the US Lab module of the International Space Station, inclusive of all the racks and components, was simplified to remove holes and volume overlaps. Problematic features within the drawing were also removed or repaired to prevent runtime issues. The cleaned drawing was then run through the DAGMC workflow to prepare for analysis. Pilot tests modeling transport of 1GeV proton and 800MeV/A oxygen sources show that reasonable results are converged upon in an acceptable amount of overall computation time from drawing preparation to data analysis. The FLUKA radiation transport code will next be used to model both a GCR and a trapped radiation source. These results will then be compared with measurements that have been made by the radiation instrumentation deployed inside the US Lab module. DISCUSSION Early analyses have indicated that the DAGMC workflow is a promising toolkit for running vehicle geometries of interest to NASA through multiple radiation transport codes. In addition, recent work has shown that a realistic human phantom, provided via a subcontract with the University of Florida, can be placed inside any vehicle geometry for a combinatorial analysis. This added functionality gives the user the ability to score various parameters at the organ level, and the results can then be used as input for cancer risk models.

Barzilla, J. E.↗

Advancing Concentrating Solar Thermal Modeling Using System Advisor Model (SAM)

Concentrating solar thermal (CST) technologies play a critical role in enabling dispatchable power and high-temperature industrial heat applications. Accurate and flexible modeling tools are essential for evaluating system performance, guiding technology research and development, and informing investment decisions. The National Laboratory of the Rockies's System Advisor Model (SAM) is a widely used techno-economic simulation platform for CST systems, providing detailed performance and financial modeling capabilities for multiple CST system configurations. SAM integrates physics-based performance models with financial analysis to simulate the behavior of complex energy systems under realistic operating conditions. For CST technologies (including tower, parabolic trough, and linear Fresnel), SAM enables hourly simulations using site-specific weather data that ensure feasible operating conditions and convergence of mass and energy between core system components (i.e., solar field, receiver, thermal energy storage, and power cycle). These capabilities allow researchers and developers to evaluate annual energy production, capacity factors, levelized cost of energy (LCOE), and system dispatch strategies. A key advantage of SAM lies in its flexibility for parametric analysis and large-scale computational studies. Users can vary system design parameters such as heliostat field layout, receiver dimensions, thermal energy storage capacity, power block sizing, and installation cost assumptions to investigate their impact on system performance and financial metrics. When combined with automated scripting through LK, SDKTool, or Python interfaces, SAM enables high-throughput simulation workflows that support sensitivity analysis, technology benchmarking, and optimization studies. These approaches are particularly valuable for next-generation CST concepts, where design spaces are large and system interactions are complex. Another important capability of SAM is its support for dispatch optimization and thermal energy storage modeling, which are central to the value proposition of CST technologies. The ability to simulate integrated storage and flexible power generation allows researchers to explore strategies that maximize grid value, improve capacity utilization, and enhance integration with variable resources such as photovoltaic and wind generation. This poster will present an overview of SAM's thermal system modeling capabilities including concentrating solar. Additionally, we will highlight new feature developments including: 1) implementing Google's OR-Tools optimization platform for faster and more robust dispatch optimization, 2) developing a new power load following controller for modeling behind-the-meter applications, 3) enabling direct modeling of CSP-PV hybrid systems with the inclusion of battery storage, and 4) developing a multi-receiver falling particle Gen3 system model.

14 SOLAR ENERGY↗

Complementary workflows for identifying one-hop network behavior and multi-hop network dependencies

A network analysis tool evaluates network flow information in complementary workflows to identify one-hop behavior of network assets and also identify multi-hop dependencies between network assets. In one workflow (e.g., using association rule learning), the network analysis tool can identify significant one-hop communication patterns to and/or from network assets, taken individually. Based on the identified one-hop behavior, the network analysis tool can discover patterns of similar communication among different network assets, which can inform decisions about deploying patch sets, mitigating damage, configuring a system, or detecting anomalous behavior. In a different workflow (e.g., using deep learning or cross-correlation analysis), the network analysis tool can identify significant multi-hop communication patterns that involve network assets in combination. Based on the identified multi-hop dependencies, the network analysis tool can discover functional relationships between network assets, which can inform decisions about configuring a system, managing critical network assets, or protecting critical network assets.

97 MATHEMATICS AND COMPUTING↗

GIScience in the era of Artificial Intelligence: a research agenda towards Autonomous GIS

The advent of generative AI exemplified by large language models (LLMs) opens new ways to represent and compute geographic information and transcends the process of geographic knowledge production, driving geographic information systems (GIS) towards autonomous GIS. Leveraging LLMs as the decision core, autonomous GIS can independently generate and execute geoprocessing workflows to perform spatial analysis. In this vision paper, we further elaborate on the concept of autonomous GIS and present a conceptual framework that defines its five autonomous goals, five levels of autonomy, five core functions, and three operational scales. We demonstrate how autonomous GIS could perform geospatial data retrieval, spatial analysis, and map making with four proof-of-concept GIS agents. We conclude by identifying critical challenges and future research directions, including fine-tuning and self-growing decision-cores, autonomous modelling, and examining the societal and practical implications of autonomous GIS. By establishing the groundwork for a paradigm shift in GIScience, this paper envisions a future where GIS moves beyond traditional workflows to autonomously reason, derive, innovate, and advance geospatial solutions to pressing global challenges. Meanwhile, we emphasize that as we design and deploy increasingly intelligent geospatial systems, we carry a responsibility to ensure they are developed in socially responsible ways, serve the public good, and support the continued value of human geographic insight in an AI-augmented future.

Autonomous GI↗

Streaming Large-Scale Microscopy Data to a Supercomputing Facility

Data management is a critical component of modern experimental workflows. As data generation rates increase, transferring data from acquisition servers to processing servers via conventional file-based methods is becoming increasingly impractical. The 4D Camera at the National Center for Electron Microscopy generates data at a nominal rate of 480 Gbit s -1 (87,000 frames s -1 ⁠), producing a 700 GB dataset in 15 s. To address the challenges associated with storing and processing such quantities of data, we developed a streaming workflow that utilizes a high-speed network to connect the 4D Camera’s data acquisition system to supercomputing nodes at the National Energy Research Scientific Computing Center, bypassing intermediate file storage entirely. In this work, we demonstrate the effectiveness of our streaming pipeline in a production setting through an hour-long experiment that generated over 10 TB of raw data, yielding high-quality datasets suitable for advanced analyses. Additionally, we compare the efficacy of this streaming workflow against the conventional file-transfer workflow by conducting a postmortem analysis on historical data from experiments performed by real users. Our findings show that the streaming workflow significantly improves data turnaround time, enables real-time decision-making, and minimizes the potential for human error by eliminating manual user interactions.

4D-STEM↗

Workflows Community Summit: Tightening the Integration between Computing Facilities and Scientific Workflows

Scientific workflows are used almost universally across science domains for solving complex and largescale computing and data analysis problems. The importance of workflows is highlighted by the fact that they have underpinned some of the most significant discoveries of the past decades. Many of these workflows have significant computational, storage, and communication demands, and thus must execute on a range of large-scale computer systems, from local clusters to public clouds and upcoming exascale HPC platforms. Managing these executions is often a significant undertaking, requiring a sophisticated and versatile software infrastructure. Historically, infrastructures for workflow execution consisted of complex, integrated systems, developed in-house by workflow practitioners with strong dependencies on a range of legacy technologies—even including sets of ad hoc scripts. Due to the increasing need to support workflows, dedicated workflow systems were developed to provide abstractions for creating, executing, and adapting workflows conveniently and efficiently while ensuring portability. While these efforts are all worthwhile individually, there are now hundreds of independent workflow systems. These workflow systems are created and used by thousands of researchers and developers, leading to a rapidly growing corpus of workflows research publications. The resulting workflow system technology landscape is fragmented, which may present significant barriers for future workflow users due to many seemingly comparable, yet usually mutually incompatible, systems that exist. In order to tackle some of the challenges described above, the DOE-funded ExaWorks and NSF-funded WorkflowsRI projects have organized in 2021 a series of events entitled the “Workflows Community Summit”. The third edition of the “Workflows Community Summit” explored workflows challenges and opportunities from the perspective of computing centers and facilities. This third summit builds on two prior summits (https://workflowsri.org/summits) that (i) established a high level vision for workflows research; and (ii) explored technical approaches for realizing that vision. The third summit brought together a small group of facilities representatives with the aim to understand how workflows are currently being used at each facility, how facilities would like to interact with workflow developers and users, how workflows fit with facility roadmaps, and what opportunities there are for tighter integration between facilities and workflows. This report documents and organizes the wealth of information provided by the participants before, during, and after the summit.

97 MATHEMATICS AND COMPUTING↗

Introduction to the Glenn Icing Computational Environment (GlennICE)

The NASA John H. Glenn Research Center at Lewis Field is developing the Glenn Icing Computational Environment (GlennICE) tool to aid those evaluating, designing and certifying aircraft, engines, and aircraft components for flight in icing conditions. This short course will walk through some of the underlying physics involved with GlennICE and how we achieve efficient 3D ice accretion predictions. After an introduction of GlennICE, an analysis of the Common Research Model High-Lift will be demonstrated to showcase the typical workflow for an aircraft icing analysis. Within this walkthrough, capabilities will be highlighted with future planned capabilities being discussed. Finally, we will showcase the impact GlennICE is having on NASA’s icing portfolio and how it is advancing aircraft icing research and safety.

Icing↗

Panorama 360 (Final Report)

This is the final technical report for the DOE-funded Panorama 360 project. Panorama 360 provided a resource for the collection, analysis, and sharing of performance data about end-to-end scientific workflows executing on DOE facilities. The work focused on workflows that include experimental data generation at DOE facilities. The main activities of Panorama 360 include the development of: 1. A distributed repository that stores different types of workflow execution data (e.g., point and time series performance traces at fine- and coarse-grained levels); 2. A set of open-source data capture, curation, and publishing tools fully integrated with a state-of-the-art workflow management system that automates data ingestion to the repository and enables users to discover, query, and process data from the repository; 3. A set of analysis algorithms and machine learning based tools to perform analysis and characterization of the gathered data, which can be used to detect anomalous performance or system faults; and 4. Best practices and recommendations for workflow evaluation, analysis, execution, and architectures.

97 MATHEMATICS AND COMPUTING↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗