Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Stochastic Modeling Workflow to Generate Representative Geologic Variability in Training Dataset for SMART Initiative

The poster discusses the modeling workflow to generate ensemble of geologic realizations of the Illinois Basin Decatur Project (IBDP) site, based on available site characterization data and inherent uncertainty of those data, for use by project collaborators in DOE SMART Initiative (Phase 2) to build their forward modeling, history matching, and optimization workflows. This poster is summarized from the technical report for the SMART project submitted to U.S. DOE earlier this year.

Ganesh, Priya Ravi↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the FECM NETL Carbon Management Program Review Meeting 2024.

Creason, Christopher↗

Finite Element Analysis System Workflow Tools

A collection of MATLAB functions and class definitions called System Workflow Tools (SWFT) are available to semi-automate steps in the simulation process. Some of these steps are often simple and routine for smaller finite element models, but if done directly by an analyst can quickly become labor intensive, cumbersome, and error prone for larger, system level models. Some of SWFT’s capabilities demonstrated in this report includes writing Sierra input decks and processing Quantities of Interest (QOI) from results files. SWFT also writes scripts in order to utilize other software programs such as Cubit (separating system level CAD into subassemblies and components, creating nodesets and sidesets), DAKOTA (ensemble management), and ParaView (contour plots and animations). Detailed commands and workflows from mesh generation to report generation are provided as examples for analysts to utilize SWFT capabilities.

97 MATHEMATICS AND COMPUTING↗

Where are the Data? Automating a Workflow for Carbon Storage Data Gap Analyses

This presentation demonstrates a spatial analysis workflow to assess data availability for the many components of geologic carbon storage technical viability. The workflow relies upon a knowledge-data framework that links the different components of GCS technical viability to the data types needed for evaluation. Using this contextual information, a combination of data science methods (e.g., natural language processing) and spatial analyses are applied to identify areas where sufficient data exists for a given component. The results are aggregated into maps illustrating data density and spatial gaps across all technical viability factors and data categories, as well as the individual component and category level for a more nuanced understanding. Presented at the Geological Society of America Connects 2024 Annual Meeting in Anaheim, California, 22-25 September 2024.

Creason, Christopher↗

PyARC Status Report: New Integrations and Upgrades to the Fast Reactor Analysis Workflow Management Tool

PyARC was initially developed as an open source tool to support fast reactor analyses using the Argonne Reactor Computation (ARC) code suite as a part of the Nuclear Energy Advanced Modeling and Simulation (NEAMS) Workbench initiative in FY17. The goal of this initiative is to provide a common user interface for model generation, real-time validation, execution, output processing, and visualization for all integrated codes. This is accomplished through the reliance on tools available in the Workbench framework and runtime environment. While initially developed to support the ARC codes, PyARC was extended in FY22 to wrap other NEAMS and non-ARC codes, including Griffin and OpenMC, in the supported other neutronics workflows, and support users in the adoption of NEAMS-supported high fidelity analysis codes. Most recently, NUBOW-3D, a recently adopted ARC code, was integrated to support reactor bowing calculations as well. Integration of these codes into the NEAMS Workbench directly benefits the advanced reactor modeling community by: • Providing a set of controlled, maintained, documented and validated scripts to generate inputs, which promotes best practices, reduces the learning curve, and facilitates project collaboration. • Improving the user experience: the Workbench interface provides assistance for building an input through auto-completion, real-time validation, document navigation, and geometry and results visualization. • Automating complex calculations and workflows for reactor analysis. • Helping users transition to using high-fidelity NEAMS codes along-side the ARC codes. In FY22, a progress report was published that described the state of each of the tools integrated into PyARC. Since then, there have been many enhancements and upgrades to the existing integrations as well as entirely new code integrations as well. This report details all new integrations and major developments in PyARC since the version 2.0.0 release highlighted in the FY22 report.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

PV Degradation Modeling: Applying Geospatial Workflows with "PVDeg"

Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.

14 SOLAR ENERGY↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

A workflow for segmenting soil and plant X-ray computed tomography images with deep learning in Google’s Colaboratory

X-ray micro-computed tomography (X-ray μCT) has enabled the characterization of the properties and processes that take place in plants and soils at the micron scale. Despite the widespread use of this advanced technique, major limitations in both hardware and software limit the speed and accuracy of image processing and data analysis. Recent advances in machine learning, specifically the application of convolutional neural networks to image analysis, have enabled rapid and accurate segmentation of image data. Yet, challenges remain in applying convolutional neural networks to the analysis of environmentally and agriculturally relevant images. Specifically, there is a disconnect between the computer scientists and engineers, who build these AI/ML tools, and the potential end users in agricultural research, who may be unsure of how to apply these tools in their work. Additionally, the computing resources required for training and applying deep learning models are unique, more common to computer gaming systems or graphics design work, than to traditional computational systems. To navigate these challenges, we developed a modular workflow for applying convolutional neural networks to X-ray μCT images, using low-cost resources in Google’s Colaboratory web application. Here we present the results of the workflow, illustrating how parameters can be optimized to achieve best results using example scans from walnut leaves, almond flower buds, and a soil aggregate. We expect that this framework will accelerate the adoption and use of emerging deep learning techniques within the plant and soil sciences.

59 BASIC BIOLOGICAL SCIENCES↗

Integration of Open-Source URBANopt and Dragonfly Energy Modeling Capabilities into Practitioner Workflows for District-Scale Planning and Design

High-performance districts and communities offer opportunities for reducing energy use, emissions, and costs, and can be instrumental in helping cities achieve their climate goals. The design of such communities requires identification of opportunities early on and their re-evaluation throughout the planning process. There is a need for energy modeling tools that connect 3D Computer-Aided Design (CAD) platforms to simulation engines, enabling detailed energy analysis of districts within the workflows and tools used by practitioners. This paper introduces the Dragonfly and URBANoptTM combined toolset that supports the creation of urban models from a range of geometry formats typically used by designers and planners, and provides an integrated pathway to simulate district-scale energy systems. The toolset is piloted by a global architecture and master planning firm to evaluate several key urban-scale technical questions for the design of a district in Chicago. The findings indicate that, while energy savings can be achieved through traditional architectural studies and enhancements to individual building efficiency, the modeling toolset helps identify additional savings and insights that can be achieved when considering district-scale energy systems. Finally, this study demonstrates how the Dragonfly/URBANopt toolset can integrate with master planning workflows, thereby enabling an iterative performance-based design process.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Recommendations for Uniform Variant Calling of SARS-CoV-2 Genome Sequence across Bioinformatic Workflows

Genomic sequencing of clinical samples to identify emerging variants of SARS-CoV-2 has been a key public health tool for curbing the spread of the virus. As a result, an unprecedented number of SARS-CoV-2 genomes were sequenced during the COVID-19 pandemic, which allowed for rapid identification of genetic variants, enabling the timely design and testing of therapies and deployment of new vaccine formulations to combat the new variants. However, despite the technological advances of deep sequencing, the analysis of the raw sequence data generated globally is neither standardized nor consistent, leading to vastly disparate sequences that may impact identification of variants. Here, we show that for both Illumina and Oxford Nanopore sequencing platforms, downstream bioinformatic protocols used by industry, government, and academic groups resulted in different virus sequences from same sample. These bioinformatic workflows produced consensus genomes with differences in single nucleotide polymorphisms, inclusion and exclusion of insertions, and/or deletions, despite using the same raw sequence as input datasets. Here, we compared and characterized such discrepancies and propose a specific suite of parameters and protocols that should be adopted across the field. Consistent results from bioinformatic workflows are fundamental to SARS-CoV-2 and future pathogen surveillance efforts, including pandemic preparation, to allow for a data-driven and timely public health response.

60 APPLIED LIFE SCIENCES↗

Challenges for Implementing FAIR Digital Objects with High Performance Workflows

New types of workflows are being used in science that couple traditional distributed and high-performance computing (HPC) with data-intensive approaches, and orchestrate ensembles of numerical simulations and artificial intelligence (AI) models. Such workflows may use AI models to supplement computation where numerical simulations may be too computationally expensive, to automate trivial yet time consuming operations, to perform preliminary selections among intractable numbers of combinations in domains as diverse as protein binding, fine-grid climate simulations, and drug discovery.

97 MATHEMATICS AND COMPUTING↗

ADEPT: A Pedagogical Framework for Integrating Agentic AI with Deterministic Scientific Workflows

The integration of Large Language Models (LLMs) into scientific research promises to accelerate discovery, yet a significant gap remains between the dynamic reasoning of Artificial Intelligence (AI) agents and the static, deterministic nature of canonical scientific workflows. This paper introduces ADEPT (Agentic Discovery and Exploration Platform for Tools), a reference architecture and pedagogical framework explicitly designed to bridge this gap. ADEPT's primary mission is to provide a transparent, "glass-box" environment where researchers and engineers can learn to effectively wrap established scientific software (e.g., BLAST, Nextflow pipelines) and compose it into reliable, agent-driven workflows. We describe its modular, multi-server architecture, which leverages the Model Context Protocol (MCP) for tool serving, LangGraph for robust agentic orchestration, and a secure nsjail-based sandbox for safe code execution. By prioritizing architectural clarity, safety, and modularity, ADEPT serves as an extensible blueprint for building trustworthy AI-augmented systems and fosters the collaborative development necessary to responsibly employ agentic AI for science. We provide practical examples of how to adapt and extend this framework, highlighting its utility in workforce development and AI-readiness capabilities across research and development projects.

97 MATHEMATICS AND COMPUTING↗

Production processing and workflow management software evaluation in the DUNE collaboration

The Deep Underground Neutrino Experiment (DUNE) will be theworld’s foremost neutrino detector when it begins taking data in the mid-2020s.Two prototype detectors, collectively known as ProtoDUNE, have begun tak-ing data at CERN and have accumulated over 3 PB of raw and reconstructeddata since September 2018. Particle interaction within liquid argon time projec-tion chambers are challenging to reconstruct, and the collaboration has set upa dedicated Production Processing group to perform centralized reconstructionof the large ProtoDUNE datasets as well as to generate large-scale Monte Carlosimulation. Part of the production infrastructure includes workflow manage-ment software and monitoring tools that are necessary to eciently submit andmonitor the large and diverse set of jobs needed to meet the experiment’s goals.We will give a brief overview of DUNE and ProtoDUNE, describe the varioustypes of jobs within the Production Processing group’s purview, and discuss thesoftware and workflow management strategies are currently in place to meetexisting demand. We will conclude with a description of our requirements in aworkflow management software solution and our planned evaluation process.

Herner, Kenneth R.↗

Chimbuko: A Workflow-Level Scalable Performance Trace Analysis Tool

ABSTRACT Due to the sheer volume of data it is typically impractical to analyze the detailed performance of an HPC application running at-scale. While conventional small-scale benchmarking and scaling studies are often sufficient for simple applications, many modern workflow-based applications couple multiple elements with competing resource demands and complex inter-communication patterns for which performance cannot easily be studied in isolation and at small scale. This work discusses Chimbuko, a performance analysis framework that provides real-time, in situ anomaly detection. By focusing specifically on performance anomalies and their origin (aka provenance), data volumes are dramatically reduced without losing necessary details. To the best of our knowledge, Chimbuko is the first online, distributed, and scalable workflow-level performance trace analysis framework. We demonstrate the tool's usefulness on Oak Ridge National Laboratory's Summit system.

97 MATHEMATICS AND COMPUTING↗

2023 Real-time Optimization Workflow Status Update

Nuclear integrated energy systems are composed of a diverse set of energy generation sources and exist in dynamic and competitive electricity markets. With the inclusion of thermal energy storage, nuclear power systems can store heat for future use through various processes, such as water desalination or hydrogen generation. This heat storage can be managed in such a way as to economically optimize its usage. By combining real-time price data from the day-ahead and real-time markets with predictive and intelligent models, the charging and discharging of the thermal energy storage may be determined and optimized. This research details an approach through models and systems for the real-time optimization (RTO) of capacity allocation. Virtual models of the energy system and its physics phenomenon and component interactions provide intelligence to verification and prediction of operations. An optimization framework can use data generated from both a set of physical assets as well as the virtual models to predict future performance and create optimization and control workflows. Data warehouse technologies can be used to combine data across models, optimization workflows, market price data, and sensor data to intuitively store various types of data and provide integrations to physical control systems as well as user visualizations. Put together, these components can create a system for the RTO of nuclear integrated energy systems. Various virtual models and bench-scale physical demonstrations have been successfully performed and verified using this system. Larger scale testbeds that include thermal energy storage systems have been identified as future opportunities. A gap analysis that details the steps necessary to reach a physical demonstration at this scale is provided, along with conclusions on the current effort and future work.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Towards Lightweight Data Integration Using Multi-Workflow Provenance and Data Observability

Modern large-scale scientific discovery requires multidisciplinary collaboration across diverse computing facilities, including High Performance Computing (HPC) machines and the Edge-to-Cloud continuum. Integrated data analysis plays a crucial role in scientific discovery, especially in the current AI era, by enabling Responsible AI development, FAIR, Reproducibility, and User Steering. However, the heterogeneous nature of science poses challenges such as dealing with multiple supporting tools, cross-facility environments, and efficient HPC execution. Building on data observability, adapter system design, and provenance, we propose MIDA: an approach for lightweight runtime Multi-workflow Integrated Data Analysis. MIDA defines data observability strategies and adaptability methods for various parallel systems and machine learning tools. With observability, it intercepts the dataflows in the background without requiring instrumentation while integrating domain, provenance, and telemetry data at runtime into a unified database ready for user steering queries. We conduct experiments showing end-to-end multi-workflow analysis integrating data from Dask and MLFlow in a real distributed deep learning use case for materials science that runs on multiple environments with up to 276 GPUs in parallel. We show near-zero overhead running up to 100,000 tasks on 1,680 CPU cores on the Summit supercomputer.

Santos Souza, Renan↗

GeoThermalCloud for EGS – An Open-source, User-friendly, Scalable AI Workflow for Modeling Enhanced Geothermal Systems

Enhanced Geothermal Systems (EGS) offer a vast potential to expand the use of geothermal energy. Heat is extracted from this engineered system by injecting relatively cold water into subsurface fractures, which are in contact with hot dry rock, and brought back to surface through production wells. Creating EGS requires improving the natural permeability of hot crystalline rocks. In this short conference paper, we present a reproducible workflow for modeling EGS. Our workflow called the GeoThermalCloud (GTC) for EGS, leverages recent advances in machine learning, deep learning, and high-performance computing. This GTC framework is currently being made open-source, user-friendly, and reproducible through python scripts as well as Google Colab/Jupyter Notebooks. This GTC for EGS modeling scripts are made available at https://github.com/SmartTensors/GeoThermalCloud.jl/tree/master/EGS and will constantly be updated to cater for geothermal community. Current GTC framework provides scripts to train deep learning (DL) models for techno-economics and data worth analysis. The Geothermal Design Tool (https://github.com/GeoDesignTool/GeoDT.git), a fast and simplified multi-physics solver, is used to develop a database for training DL models. This short paper provides details on the scripts to curate, process, and train DL models. The scripts can easily be modified to train on databases generated by other popular open-source simulators such as PFLOTRAN, STOMP, TOUGH, and GEOSX or commercial software such as ResFrac and COMSOL.

15 GEOTHERMAL ENERGY↗