Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

324 records · Page 18

Pseudonymization at Scale: OLCF’s Summit Usage Data Case Study

The analysis of vast amounts of data and the processing of complex computational jobs have traditionally relied upon high performance computing (HPC) systems, which offer reliable and efficient management of large-scale computational and data resources. Understanding these analyses’ needs is paramount for designing solutions that can lead to better science, and similarly, understanding the characteristics of the user behavior on those systems is important for improving user experiences on HPC systems. A common approach to gathering data about user behavior is to extract workload characteristics from system log data available only to system administrators. Recently at Oak Ridge Leadership Computing Facility (OLCF), however, we unveiled user behavior about the Summit supercomputer by collecting data from a user’s point of view with ordinary Unix commands.In this paper, we discuss the process, challenges, and lessons learned while preparing this dataset for publication and submission to an open data challenge. The original dataset contains personal identifiable information (PII) about the users of OLCF which needed be masked prior to publication, and we determined that anonymization, which scrubs PII completely, destroyed too much of the structure of the data to be interesting for the data challenge. We instead chose to pseudonymize the dataset, which reduced the linkability of the dataset to the users’ identities. Pseudonymization is significantly more computationally expensive than anonymization, and the size of our dataset, which is approximately 175 million lines of raw text, necessitated the development of a parallelized workflow that could be reused on different HPC machines. We demonstrate the scaling behavior of the workflow on two leadership class HPC systems at OLCF, and we show that we were able to bring the overall makespan time from an impractical 20+ hours on a single node down to around 2 hours. As a result of this work, we release the entire pseudonymized dataset and make the workflows and source code publicly available.

Maheshwari, Ketan↗

Design and Data Management System

The Design and Data Management System (DDMS) was developed to automate the NASA Engineering Order (EO) and Engineering Change Request (ECR) processes at the Propulsion Test Facilities at Stennis Space Center for efficient and effective Configuration Management (CM). Prior to the development of DDMS, the CM system was a manual, paper-based system that required an EO or ECR submitter to walk the changes through the acceptance process to obtain necessary approval signatures. This approval process could take up to two weeks, and was subject to a variety of human errors. The process also requires that the CM office make copies and distribute them to the Configuration Control Board members for review prior to meetings. At any point, there was a potential for an error or loss of the change records, meaning the configuration of record was not accurate. The new Web-based DDMS eliminates unnecessary copies, reduces the time needed to distribute the paperwork, reduces time to gain the necessary signatures, and prevents the variety of errors inherent in the previous manual system. After implementation of the DDMS, all EOs and ECRs can be automatically checked prior to submittal to ensure that the documentation is complete and accurate. Much of the configuration information can be documented in the DDMS through pull-down forms to ensure consistent entries by the engineers and technicians in the field. The software also can electronically route the documents through the signature process to obtain the necessary approvals needed for work authorization. The workflow of the system allows for backups and timestamps that determine the correct routing and completion of all required authorizations in a more timely manner, as well as assuring the quality and accuracy of the configuration documents.

Messer, Elizabeth↗

Where Is the Provenance? Ethical Replicability and Reproducibility in GIScience and Its Critical Applications

As replicability and reproducibility (R&R) crises develop within emerging convergent inquiry, ethical use of provenance information is central to the establishment and preservation of trust in critical applications of GIScience and geospatial technologies. Today large volumes of geospatial data are generated at high velocity from satellite sensors and unmanned aircraft systems, citizen sensors, geolocation-based data services, global navigation satellite systems, and so on. The extensive use of these data for applications such as disaster and humanitarian response raises the issue of R&R from competing perspectives of location privacy and geospatial data quality. Although geospatial data can be integrated and linked with contextual information to identify individuals’ movements, steps taken to ensure privacy can complicate the multiuser development of high-quality geospatial workflows. Provenance information as digital records of historical (retrospective) and potential future (prospective) geospatial processes is often overlooked, misunderstood, or inadequately addressed. We explore the relationship between provenance information, location privacy, and geospatial data quality in the context of R&R with a focus on disaster analytics. Here, we argue that in the era of big data and deep learning, GIScientists and associated institutions bear greater responsibility both for geospatial workflow quality and for location privacy. Given vastly heterogenous computational landscapes, we provide practical recommendations for ethically driven provenance and R&R research and development within the GIScience community and beyond.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

November 2021 Operational Highlights

Completed: 1. the initial tests of document workflows for Titan on the Red, which included standing up a software encryption capability to test data ingestion; 2. a software encryption capability in order to ingest the first data source by the end of the month; 3. a stand-alone computer for classified ontology data entry. Titan on the Red is an artificial intelligence/machine learning system to make digitizing, cataloging, and searching NSRC collections easier and more efficient. Created Online Vault backups, with one copy stored at LANL and one shipped to Lawrence Livermore National Laboratory. The Online Vault is a classified, searchable library of LANL’s nuclear weapons design and test history. Imported new Laboratory Directed Research and Development (LDRD) documents based on revised access categories. Bulk ingested nearly 10,000 documents via java-based PowerLoader. This application ingests metadata and content into the Online Vault to meet the requirements for the NSRC collections. Rehoused 300 linear feet of weapons physics documents in archival, acid-free storage folders and boxes.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Geologic hydrogen: From natural occurrences to anthropogenic generation – A review of fundamentals, potential, challenges and prospects

Growing demand for hydrogen is exposing the environmental and economic limits of reforming-based and carbon-managed supply chains, while the scale-up of electrolytic capacity remains capital-constrained. Geologic hydrogen, defined as molecular H₂ generated and stored within the Earth's crust offers a complementary, potentially lower-cost resource, yet exploration is still ad hoc. This review (1) revisits a global inventory of confirmed hydrogen seeps and subsurface occurrences; (2) analyzes the controlling reactions, migration pathways, and trapping conditions governing these occurrences; (3) proposes a process-based geologic hydrogen system concept analogous to, yet distinct from, the petroleum system; and (4) evaluates potential geologic hydrogen systems within the United States as a representative case study. Here, we contrast natural systems powered by serpentinization, mantle degassing or radiolysis with anthropogenic systems that stimulate the same reactions or convert in-situ hydrocarbons. Stable hydrogen accumulations require generation rates that outpace combined physical, chemical and microbial losses; the Bourakébougou field (Mali) exemplifies a self-recharging, free-gas reservoir sustained by meteoric-water serpentinization beneath an efficient caprock. Prospective geologic hydrogen resources are likely to occur in regions where iron-rich lithologies, deep-seated faults, and low-permeability sealing formations coexist. Applying this principle, we highlight three promising hydrogen play types in U.S. geological terrains: ophiolite belts (Appalachian and Californian regions), the Midcontinent Rift and the Lake Superior banded‑iron formations. Multiphysics numerical models and positive-unlabeled machine-learning workflows help to accelerate play screening and de-risk future production; yet, reaction kinetics, stimulation strategies, and full techno-economic and life-cycle assessments remain pivotal knowledge gaps.

Anthropogenic hydrogen generation↗

Automatic building energy model development and debugging using large language models agentic workflow

Building energy modeling (BEM) is a complex process that demands significant time and expertise, limiting its broader application in building design and operations. While Large Language Models (LLMs) agentic workflow have facilitated complex engineering processes, their application in BEM has not been specifically explored. This paper investigates the feasibility of automating BEM using LLM agentic workflow. Here, we developed a generic LLM-planning-based workflow that takes a building description as input and generates an error-free EnergyPlus building energy model. Our robust workflow includes four core agents: 1) Building Description Pre-Processing, 2) IDF Object Information Extraction, 3) Single IDF Object Generator Suite, and 4) IDF Debugging Agent. These agents divide the complex tasks into manageable sub-steps, enabling LLMs to generate accurate and reliable results at each stage. The case study demonstrates the successful translation of a building description into an error-free EnergyPlus model for the iUnit modular building at the National Renewable Energy Laboratory. The effectiveness of our workflow surpasses: 1) naive prompt engineering, 2) other LLM-based workflows, and 3) manual modeling, in terms of accuracy, reliability, and time efficiency. The paper concludes with a discussion on the interplay between foundational models and LLM agent planning design, advocating for the use of fine-tuned, specialized models to advance this field.

97 MATHEMATICS AND COMPUTING↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Conference presentation at Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗

Workflow for Process Automation of Soil Gas Results from an Automated Soil Gas-Sampling System for Application in Carbon Storage Projects

Extended abstract for Geoconvention, Calgary, Alberta, Canada, May 12–14, 2025. The Energy & Environmental Research Center (EERC) developed an automated workflow for processing soil gas measurements collected from the automated soil gas-sampling systems deployed across the project site. Raw soil gas measurements are collected from each station every 4 hours and automatically uploaded to a cloud database. The workflow begins by writing code to download the data to a workstation automatically, then the data are published to an online dashboard that visualizes the measurements in time-series plots and a process-based decision-making framework. This automated workflow accelerates the time from data acquisition to decision-making. It supports carbon storage project operators by preparing and delivering a live, standardized dataset for quick analysis and source attribution to provide assurance of containment and overall permit compliance.

02 PETROLEUM↗

Immersive Scientific Visualization of Molten-Salt Reactor Waste Characteristics Using Virtual Reality

Immersive visualization is changing how we explore, communicate, and understand complex scientific systems. In nuclear energy, an area in which data are often multidimensional, time-dependent, and difficult to interpret, virtual reality (VR) represents a powerful and intuitive informational medium. This work introduces a VR-based platform that visualizes the post-shutdown behavior and waste management lifecycle of molten-salt reactors (MSRs), a next-generation reactor type with unique operational and safety characteristics. The platform, built in Unity, is streamed on the Meta Quest 3 headset. It transforms high-fidelity simulation data into an interactive, immersive experience. Users can explore time-dependent reactor characteristics such as nuclide decay, which is a key factor for evaluating reactor waste strategies. The datasets were generated using the MOOSE (Multiphysics Object-Oriented Simulation Environment) framework and then processed through ParaView scripting for smooth integration into Unity. From a visualization standpoint, the platform emphasizes spatial storytelling, temporal exploration, and user-centered interaction. Users can navigate 3D reactor geometries, slice through volumetric data, and manipulate time to observe how physical phenomena evolve. Real-scale rendering and embodied interaction make the experience feel tangible. The interface is designed to be accessible, even to those without nuclear or simulation expertise. This lowers the barrier for stakeholders, policymakers, and the general public, while still supporting expert analysis and collaborative decision-making. This work shows how immersive visualization can function as both a scientific tool and a communication interface. By integrating simulation, processing, and visualization into a cohesive workflow, we offer a scalable framework for immersive scientific storytelling. The modular design supports future extensions to other reactor types and lifecycle stages, from shutdown to long-term storage, making the platform adaptable for both research and outreach.

99 - GENERAL AND MISCELLANEOUS↗

Understanding the Impact of Data Staging for Coupled Scientific Workflows

We report the rate of data generated by cutting-edge experimental science facilities and large-scale simulations enabled by current high-performance computing (HPC) systems has continued to grow at a far greater pace than the development of the network and storage capabilities on which these systems rely. To cope with this challenge, scientist are moving toward the creation of autonomous experiments and HPC simulations using machine learning. However, efficiently moving, storing, and processing large amounts of data away from the point of origin presents an incredible challenge. In-memory computing, in situ analysis, data staging, and data streaming are recognized viable alternatives to traditional file-based methods for transferring data between coupled workflows. However, the performance trade-offs and limitations for these methods are not fully understood when used in HPC applications. This article presents a comprehensive performance assessment of the current solutions for data staging when applied to applications that are not necessary I/O intensive which makes them not ideal candidates for these methods. Our study is based on experiments running at scale on Oak Ridge National Laboratory's Summit supercomputer using applications and simulations that cover typical computational motifs and patterns. We investigated the usability and cost/benefit trade-offs of staging algorithms for HPC applications under different scenarios and highlight opportunities for optimizing the dataflow between coupled simulation workflows.

97 MATHEMATICS AND COMPUTING↗

Urban building energy modeling (UBEM) tools: A state-of-the-art review of bottom-up physics-based approaches

Regulations corroborate the importance of retrofitting existing building stocks or constructing new energy-efficient districts. There is, thus, a need for modeling tools to evaluate energy scenarios to better manage and design cities, and numerous methodologies and tools have been developed. Among them, Urban Building Energy Modelling (UBEM) tools allow the energy simulation of buildings at large scales. Choosing an appropriate UBEM tool, balancing the level of complexity, accuracy, usability, and computing needs, remains a challenge for users. The review focuses on the main bottom-up physics-based UBEM tools, comparing them from a user-oriented perspective. Five categories are used: (i) the required inputs, (ii) the reported outputs, (iii) the exploited workflow, (iv) the applicability of each tool, and (v) the potential users. Moreover, a critical discussion is proposed, focusing on interests and trends in research and development. The results highlighted major differences between UBEM tools that must be considered to choose the proper one for an application. Finally, barriers of adoption of UBEM tools include the needs of a standardized ontology, a common three-dimensional city model, a standard procedure to collect data, and a standard set of test cases. This feeds into future development of UBEM tools to support cities’ sustainability goals.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

DOE Repository Metadata Profile (DRMP): A Metadata Framework for Advancing Interoperability and AI Readiness Across Scientific Repositories

The Department of Energy (DOE) funds a diverse and distributed ecosystem of repositories that steward scientific data, publications, and software across its research programs, user facilities, and national laboratories. While significant progress has been made in standardizing dataset-level metadata, the metadata describing repositories themselves (their identity, governance, access interfaces, policies, and technical capabilities) remains inconsistent and fragmented across DOE-funded systems. This variability limits discoverability, interoperability, automated validation, and AI-driven analysis, all of which are increasingly essential for modern scientific workflows. To address this gap, the DOE Data Curation Working Group (DCWG) developed the DOE Repository Metadata Profile (DRMP). The DRMP is a practical, community-driven framework that defines how repositories can describe themselves in a consistent, machine-actionable, and scalable manner. The DRMP is not a new metadata schema. Instead, it is a mapping profile and structured element set capturing the essential characteristics of DOE repositories. It harmonizes repository-level metadata across six widely adopted community schemas: RE3Data; DCAT-US v3; Schema.org; Dublin Core; DataCite 4.6; and PREMIS 3.0. This harmonization eliminates reinvention and enables interoperability within DOE and across the broader scientific ecosystem. A core objective of the DRMP is to reduce burden on repositories by allowing them to reuse their existing metadata through a Rosetta-style crosswalk rather than redesigning local implementations. The profile introduces a three-level conformance model that supports incremental adoption: • Level 1 – Minimum Viable Record (MVR): foundational identification elements required for workflows, project registration, and basic repository presence. • Level 2 – Interoperable: structured metadata enabling alignment with national and international discovery systems. • Level 3 – AI-Ready: enhanced provenance, policy transparency, fixity, semantic context, and capabilities that support automated reasoning, model training governance, and machine-assisted curation. To support implementation, the DRMP includes JSON Schema definitions, OpenAPI patterns, and MCP templates that allow repositories to publish machine-readable metadata directly within existing platforms. These resources are modular and lightweight, enabling adoption without major architectural change. Adopting the DRMP enables repositories to: • Enhance discoverability and interoperability by aligning identifiers, classifications, and descriptive elements across widely used schema standards. • Support federated discovery and cross-registration across DOE systems, Data.gov, and international catalogs. • Enable AI agents and workflow orchestration systems to interpret repository-level metadata within the American Science Cloud (AmSC) through Model Context Protocol (MCP)-based context publication. • Demonstrate alignment with DOE’s open science, stewardship, and FAIR data priorities. This guidance represents a community-driven step forward. Through voluntary adoption and continued feedback, the DRMP advances a cohesive, machine-actionable description of DOE repositories that supports FAIR data practices, preparing the infrastructure for AI-enabled research, and strengthening the discoverability and reuse of DOE’s scientific outputs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

2018 NISAR Applications Workshop: Forest and Disturbance; Workshop Report

Forest lands cover the globe and are important sources for providing ecosystem services including: carbon sequestration, biodiversity, timber, air and water quality. As such, counties around the world have dedicated programs for managing them. Accurate and timely information concerning the status of these forests (moisture, biomass, disturbance type, etc.) is essential to those Nations’ human and ecological health as well as economy. The joint NASA/US Forest Service workshop focused on arming forest land managers with observations and remote sensing information from the upcoming NASA-ISRO (Indian Space Research Organization) SAR (Synthetic Aperture Radar) (NISAR) satellite mission (expected to launch early 2022). Participants included representatives from different US Federal Agencies, private sector, and non-governmental organizations (NGO) that are key players in facilitating integration of Earth Observations (EO) into forest management and decision support workflows. They included scientists, technicians, and program managers with a responsibility for data acquisition and exploitation such as product development, delivery, and use, and capacity building. Discussions were held over two days to convey the broader forest and disturbance community information needs for various representative participants and programs and to facilitate the delivery of NISAR mission geospatial products and observational capabilities. Case studies were presented to demonstrate the current state of practice in the use of SAR remote sensing for applications of direct importance to forest and disturbance land management community. Eleven organizations presented their information requirements in response to a set of questions provided by the NASA team, then the NASA team responded by describing the degree to which NISAR could meet these requirements. Discussion ensued about needed data product specifications to increase utility (e.g., projection, latency, etc.), tools and capacity building. The general findings of this workshop were that (a) NISAR observations will be particularly useful to the global forest carbon and disturbance monitoring applications, but that certain data product design decisions (projections and radiometric and terrain corrections) need to be considered to increase utility; b) the biomass and disturbance detection algorithms meet many of the community needs, however there are other information products of value (e.g., soil moisture or disturbance classification, not just detection) and all products should be compliant with existing community standards for reporting uncertainty; c) providing SAR education to the community will be key specifically thinking about putting the information first and the SAR theory second, providing a simple guide of standard data processing steps (e.g., dB (decibel) to power conversion and speckle filtering); d) the community needs a user-friendly interface for finding free, archived data over their geographic regions of interest; e) user-friendly tools that connect to open-sources GIS (Global Information System) software (e.g., QGIS (Quantum GIS)) that include a graphical user interface (GUI) for SAR processing that enables both download and cloud processing. To integrate these findings and prepare the community before NISAR launches, it was suggested that there be a dedicated NISAR Forest and Disturbance Applications Working Group (as per the specifications in the NISAR Utilization Plan). After launch, it was decided that the community continue capacity building activities.

Stavros, Natasha↗

IDAES-PSE 2.6.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.6.0 Release Highlights Upcoming Changes IDAES will be switching to the new Pyomo solver interface in the next release. Whilst this will hopefully be a smooth transition for most users, there are a few important changes to be aware of. The new solver interface uses a different version of the IPOPT writer (“ipopt_v2”) and thus any custom configuration options you might have set for IPOPT will not carry over and will need to be reset. By default, the new Pyomo linear presolver will be activated with ipopt_v2. Whilst are working to identify any bugs in the presolver, it is possible that some edge cases will remain. IDAES will begin deploying a new set of scaling tools and APIs over the next few releases that make use of the new solver writers. The old scaling tools and APIs will remain for backward compatibility but will begin to be deprecated. New Models, Tools and Features New Intersphinx extension automatically linking Jupyter notebook examples to project documentation New end-to-end diagnostics example demonstrated on a real problem New complementarity formulation for VLE with cubic equations of state, backward compatibility for old formulation New solver interface with presolve (ipopt_v2) in support of upcoming changes to the initialization and APIs methods, with default set to ipopt to maintain backwards compatibility; this will deprecate once all examples have been updated New forecaster and parameterized bidder methods within grid integration library Updated surrogates API and examples to support Keras 3, with backwards compatibility for older formats such as TensorFlow SavedModel (TFSM) Updated costing base dictionary to include the 2023 cost year index value Updated ProcessBlock to include information on the constructing block class Updated Flowsheet Visualizer to allow visualize() method to return value and functions Bug Fixes Fixed bug in the Modular Property Framework that would cause errors when trying to use phase-based material balances with phase equilibria. Fixed bug in Modular Properties Framework that caused errors when initializing models with non-vapor-liquid phase equilibria. Fixed typos flagged by June update to crate-ci/typos and removed DMF-related exceptions Minor corrections of units of measurement handling in power plant waste/transport costing expressions, control volume material holdup expressions, and BTX property package parameters Fixed throwing >7500 numpy deprecation warnings by replacing scalar value assignment with element extraction and item iteration calls Testing and Robustness Migrated slow tests (>10s) to integration, impacting test coverage but also yielding a nearly 30% decrease in local test runtime Pinned pint to avoid issues with older supported Python versions Pinned codecov versions to avoid tokenless upload behavior with latest version Bumped extensions to version 3.4.2 to allow pointing to non-standard install location Deprecations and Removals Python 3.8 is no longer supported. The supported Python versions are 3.9 through 3.12 The Data Management Framework (DMF) is no longer supported. Importing idaes.core.dmf will cause a deprecation warning to be displayed until the next release The SOFC Keras surrogates have been removed. The current version of the SOFC surrogate model in the examples repository is a PySMO Kriging model.

AS↗

A Data Processing Workflow for Fixed-wing Drone Based Radiation Mapping in the Chornobyl Exclusion Zone (CEZ) - 20119

April 2020 marks the 34. anniversary of the high-profile radiological release from the Chornobyl Nuclear Power Plant (ChNPP). The release of radioactive material from reactor number four began on the 26 April 1986 and continued over a period of about 10 days, releasing approximately 1700 PBq of radioactive material (including 85 PBq of 137-Cs) into the environment. To this day, the accident remains the most significant release of radioactive material since civil nuclear power generation began. In the years since the accident, automated and remote radiation monitoring technologies have advanced significantly in their capabilities. One such example of this is the use of unmanned aerial vehicles (UAVs) in radiation mapping investigations. In April 2019, a team of scientists from the University of Bristol showcased a novel radiation mapping system within the CEZ, specifically aiming to map radiation over a large portion of the area immediately surrounding the ChNPP. Over six days of data collection, the system flew a total distance of 583.8 km, covering an area of 14.6 sq.km with an exceptional spatial resolution (sub 20 m/pixel). The work presented herein outlines and explains the data processing procedure to convert the raw data into 137-Cs activity (kBq/sq.m) and cesium-equivalent dose-rate (CED) at 1 m above ground level (μSv/hr). A demonstration of the validity of the method is demonstrated through the successful reduction of the raw data into a single linear relationship between the measured {sup 137}Cs net peak intensity and the {sup 137}Cs activity. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗

Verification and Validation of the Stakeholder Tool for Assessing Radioactive Transportation (START) – 22323

The U.S. Department of Energy (DOE) Office of Integrated Waste Management is planning for the eventual transportation, storage, and disposal of spent nuclear fuel (SNF) and high-level radioactive waste (HLW) from nuclear power plant and DOE sites. The Stakeholder Tool for Assessing Radioactive Transportation (START) is a web-based, geospatial decision-support tool developed for evaluating routing options and other aspects of transporting SNF and HLW, covering rail, truck, barge, and intermodal infrastructure and operations in the continental United States. The verification and validation (V&V) process is intended to independently assess START to provide confidence in the ability of START to accurately provide intended results. The V&V process checks the START tool using a variety of methods, ranging from independent hand calculations to comparison of START performance and results to those of other codes. The V&V activity was conducted independently from the START development team with opportunities to provide feedback and collaborate throughout the process. The V&V analyzed attributes of transportation routes produced by START, including route distance and both population and population density captured within buffer zones around routes. Population in the buffer zone, population density in the buffer zone, and route distance were all identified as crucial outputs of the START code and were subject to V&V tasks. Some of the improvements identified through the V&V process were standardizing the underlying population data in START, changing the projection of the population raster data, and changes to the methodology used for population density to improve its applicability for expected users. This collaboration also led to suggested improvements to some of the underlying shape file segments within START. A formal series of test routes went through the V&V process from START version 3.2.1 following the implementation of the recommended improvements to START. The V&V of the buffer zone population and distance reported in START were achieved by utilizing a custom workflow developed in QGIS. This platform was chosen because it is fully independent of START. The buffer zone population and route distances in START showed excellent agreement to independent V&V test results. Over 200 route test cases were run; in all cases, the percent difference in the population within the buffer zone was less than +/- 5%, and the majority of cases were below +/- 1%. These cases were also duplicated in ArcMap, which is Environmental Systems Research Institute’s (ESRI) GIS desktop application. The ArcMap results also show good agreement; however, they cannot be considered independent of START because START also uses ESRI software. The V&V tasks are currently addressing population density and estimated radiation dose in START as well as implementing further automation of the V&V process to allow larger suites of test routes to be run more easily and quickly by less experienced GIS users. These improvements to the V&V process will make it a valuable tool for the development team and the stakeholders so that each new iteration of START can easily be checked for quality and consistency. This paper and presentation will describe the V&V process, findings and implications, and ongoing V&V activities.

START, DOE, Transporation, dose, GIS↗