Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Workflow Management Systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Risk-Significant Adverse Condition Awareness Strengthens Assurance of Fault Management Systems

As spaceflight systems increase in complexity, Fault Management (FM) systems are ranked high in risk-based assessment of software criticality, emphasizing the importance of establishing highly competent domain expertise to provide assurance. Adverse conditions (ACs) and specific vulnerabilities encountered by safety- and mission-critical software systems have been identified through efforts to reduce the risk posture of software-intensive NASA missions. Acknowledgement of potential off-nominal conditions and analysis to determine software system resiliency are important aspects of hazard analysis and FM. A key component of assuring FM is an assessment of how well software addresses susceptibility to failure through consideration of ACs. Focus on significant risk predicted through experienced analysis conducted at the NASA Independent Verification & Validation (IV&V) Program enables the scoping of effective assurance strategies with regard to overall asset protection of complex spaceflight as well as ground systems. Research efforts sponsored by NASAs Office of Safety and Mission Assurance (OSMA) defined terminology, categorized data fields, and designed a baseline repository that centralizes and compiles a comprehensive listing of ACs and correlated data relevant across many NASA missions. This prototype tool helps projects improve analysis by tracking ACs and allowing queries based on project, mission type, domain/component, causal fault, and other key characteristics. Vulnerability in off-nominal situations, architectural design weaknesses, and unexpected or undesirable system behaviors in reaction to faults are curtailed with the awareness of ACs and risk-significant scenarios modeled for analysts through this database. Integration within the Enterprise Architecture at NASA IV&V enables interfacing with other tools and datasets, technical support, and accessibility across the Agency. This paper discusses the development of an improved workflow process utilizing this database for adaptive, risk-informed FM assurance that critical software systems will safely and securely protect against faults and respond to ACs in order to achieve successful missions.

IV&V↗

White Paper: Scalable Digital Twin Capabilities for Aging and Surveillance of Engineered Systems

This white paper presents a multi-year initiative to develop practical, secure, and scalable digital twin capabilities for engineered systems in aging and surveillance contexts—an approach pioneered at the National Nuclear Security Administration (NNSA) Lawrence Livermore National Laboratory (LLNL) that maps directly onto the needs and ambitions of the Navy for ship- and fleet-level digital twins. LLNL’s work in building part- and process-level digital twins for advanced manufacturing, with a vision to scale up to entire factory floors and, ultimately, enterprise-wide digital twins, offers an adaptable pathway for the Navy as it seeks to modernize lifecycle management, readiness, and predictive maintenance across ships and fleets. For our application, we integrate physics-based modeling with automated data ingestion, processing, and AI-driven calibration, creating hybrid models that are both interpretable and data responsive. We modernized legacy workflows, established centralized data infrastructure, automated experimental pipelines, and demonstrated end-to-end coupling of accelerated aging data with finite element simulations via optimization and surrogate modeling. The result is a generalizable framework that supports part-level digital twins today and lays the groundwork for future system-level twins suitable for Navy applications.

36 MATERIALS SCIENCE↗

Data Federation Challenges in Remote Near-Real-Time Fusion Experiment Data Processing

Fusion energy experiments and simulations provide critical information needed to plan future fusion reactors. As next-generation devices like ITER move toward long-pulse experiments, analyses, including AI and ML, should be performed in a wide range of time and computing constraints, from near-real-time constraints, between-shot analysis, and to campaign-wide long-term analysis. However, the data volume, velocity, and variety make it extremely challenging for analyses using only local computational resources. Researchers need the ability to compose and execute workflows spanning edge resources to large-scale high-performance computing facilities.We present Delta, a system to address data analysis challenges, including AI/ML, in fusion science, by leveraging the ADIOS I/O library and middleware, to support executing science workflows over the wide area network for near-real-time streaming. We discuss the data federation challenges in performing remote workflows, focusing on on-going research work in (1) managing, reducing, and streaming data to minimize I/O and data movement overheads, (2) decompressing and reorganizing data for analysis, and (3) executing workflows for automated data analysis. We introduce examples for deep-learning based data analysis for the fusion domain and demonstrate how we use Delta to construct end-to-end workflows for a fusion device in Korea, connecting a remote DOE facility in the USA. The capability demonstrated by this project is the basis for improving the state of the art for near-real-time data federation amongst remote facilities.

Choi, Jong Youl↗

Quantitative 14 N NMR with Monte Carlo Uncertainty Analysis of Nitrate/Nitrite in Alkaline Nuclear Waste

While monitoring of nitrate and nitrite concentrations is important for managing corrosion in nuclear waste systems, existing analytical methods are hindered by turbidity, spectral interference, and delays from sample handling. Here, we demonstrate quantitative 14 N nuclear magnetic resonance (qNMR) spectroscopy as a direct, matrix-tolerant approach for nitrate and nitrite detection at natural abundance. Monte Carlo resampling was integrated into the workflow to quantify random error, establish precision–time tradeoffs, and separate noise-limited uncertainty from systematic bias arising from shimming, transmitter offset, or excitation pulse conditions. Quantification of nitrate and nitrite were validated in controlled alkaline matrix challenges and in 18-component Hanford-type simulants. These results establish 14 N qNMR as a practical, uncertainty-bounded tool for monitoring redox-active nitrogen species in chemically complex environments and provide a generalizable framework for quantitative analysis of quadrupolar nuclei.

Graham, Trent R. [Pacific Northwest National Labor↗

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Baheri, Betis↗

Developing an Automated Microscopic Traffic Simulation Scenario Generation Tool

Traffic simulation is an effective tool for urban planners, traffic engineers, and researchers to study traffic. In particular, microscopic traffic simulation, which simulates individual vehicles’ movements within a transportation network, has demonstrated its importance in analyzing and managing transportation systems. However, integrating data from various sources, generating traffic scenarios, and importing information into traffic simulators to conduct microscopic simulations have always been a challenge. This paper presents a solution to overcome this challenge: RealTwin, a comprehensive tool for automated scenario generation for microscopic traffic simulation. Following a streamlined scenario generation and calibration workflow, RealTwin effectively bridges gaps between traffic data from various sources and traffic simulators, making microscopic traffic simulation more accessible for researchers and engineers across various levels of expertise. Using RealTwin to generate a real-world traffic scenario in Simulation of Urban Mobility (SUMO), VISSIM, and AIMSUN, RealTwin’s ability is demonstrated in the construction of realistic and consistent traffic scenarios in different simulators. Furthermore, this paper introduces and illustrates RealTwin’s capability for technology (e.g., autonomous vehicle) scenario generation. This feature can contribute to more comprehensive microscopic simulations, facilitating the analysis of potential effects of various technological innovations on mobility, energy efficiency, and safety. Finally, RealTwin is used to calibrate a simulation in SUMO. In conclusion, the calibration module enhances RealTwin’s ability to generate consistent simulations across different platforms and more realistic simulations that reflect real-world traffic operations.

autonomous vehicle↗

Calibration of urban building energy model using smart meter data for district peak load prediction

Urban building energy modeling (UBEM) is a powerful approach to assessing baseline building energy performance and retrofits with new technologies across building stocks in cities. However, the accuracy of UBEM is often constrained by the limited availability of reliable data about building characteristics and operations, such as envelope efficiency levels, HVAC system performance, and end-use load patterns. Existing research has performed UBEM calibration using annual or monthly energy consumption data, which falls short when higher-resolution time series applications are needed, such as peak load prediction for utility operation planning. This study presents a new framework for calibrating building energy models at urban scale using smart meter data, targeting the accurate prediction of summer peak electricity loads to support robust grid planning. The framework first integrates various data sources to enhance baseline input assumptions for building models, and then calibrates the baseline models through a pattern-matching approach. A case study using CityBES and two years of AMI data from over 9000 residential customers in Portland, Oregon, demonstrated the workflow and its effectiveness. The calibrated models achieved a daily peak load mean absolute percentage error of 2.6 % during the heatwave in the calibration year, and 2.0 % in the validation year using another year of AMI data. Using the calibrated models, we analyzed the demand flexibility potential of the district building stock as an application of UBEM calibration. The findings affirm the appropriate use of UBEM for peak electric load forecasting and demand side management at the utility distribution system level.

AMI data↗

Bridging Equipment Reliability Data and Risk Informed Decisions in a Plant Operation Context

Industry equipment reliability and asset management programs are essential elements that help ensure the safe and economical operation of nuclear power plants. The effectiveness of these programs is addressed in several industry-developed and regulatory programs. The Risk-Informed Asset Management (RIAM) project is tasked to develop tools in support of the equipment reliability and asset management programs at nuclear power plants. These tools are designed to create a direct bridge between component health/lifecycle data and decision making (e.g., maintenance scheduling and project prioritization). The goal of this article is to provide a guide for specific use cases that the RIAM project is targeting. We have grouped uses cases into three main areas. The first area focuses on the analysis of equipment reliability data with a particular emphasis on condition-based data, such as test/surveillance reports and component monitoring data. The second area focuses on the integration of equipment reliability into system/plant reliability models to determine system/plant health and identify the components that are critical to maintain an operational system. Lastly, the third area manages plant resources, such as maintenance activities and replacement scheduling using optimization methods. Here the primary focus is on supporting typical system engineer decisions regarding maintenance activity scheduling and component aging management. This is performed in a risk-informed context where the term “risk” is broadly constructed to include both plant reliability and economics. This framework combines data analytics tools to analyze equipment reliability data with risk-informed methods designed to support system engineer decisions (e.g., maintenance and replacement schedules, optimal maintenance posture) in a customizable workflow.

97 - MATHEMATICS AND COMPUTING↗

Portable, heterogeneous ensemble workflows at scale using libEnsemble

libEnsemble is a Python-based toolkit for running dynamic ensembles, developed as part of the DOE Exascale Computing Project. The toolkit utilizes a unique generator–simulator–allocator paradigm, where generators produce input for simulators, simulators evaluate those inputs, and allocators decide whether and when a simulator or generator should be called. The generator steers the ensemble based on simulation results. Generators may, for example, apply methods for numerical optimization, machine learning, or statistical calibration. libEnsemble communicates between a manager and workers. Flexibility is provided through multiple manager–worker communication substrates each of which has different benefits. These include Python’s multiprocessing, mpi4py, and TCP. Multisite ensembles are supported using Balsam or Globus Compute. We overview the unique characteristics of libEnsemble as well as current and potential interoperability with other packages in the workflow ecosystem. We highlight libEnsemble’s dynamic resource features: libEnsemble can detect system resources, such as available nodes, cores, and GPUs, and assign these in a portable way. These features allow users to specify the number of processors and GPUs required for each simulation; and resources will be automatically assigned on a wide range of systems, including Frontier, Aurora, and Perlmutter. Such ensembles can include multiple simulation types, some using GPUs and others using only CPUs, sharing nodes for maximum efficiency. We also describe the benefits of libEnsemble’s generator–simulator coupling, which easily exposes to the user the ability to cancel, and portably kill, running simulations based on models that are updated with intermediate simulation output. We demonstrate libEnsemble’s capabilities, scalability, and scientific impact via a Gaussian process surrogate training problem for the longitudinal density profile at the exit of a plasma accelerator stage. In conclusion, the study uses gpCAM for the surrogate model and employs either Wake-T or WarpX simulations, highlighting efficient use of resources that can easily extend to exascale.

Dynamic ensembles↗

DART-PFLOTRAN: An ensemble-based data assimilation system for estimating subsurface flow and transport model parameters

Ensemble-based Data Assimilation (EDA), based on the Monte Carlo approach, has been effectively applied to estimate model parameters through inverse modeling in subsurface flow and transport problems. However, implementation of EDA approach involves a complicated workflow that include setting up and executing ensemble forward model simulations, processing observations and model simulation results for parameter updates, and repeat for sequential or iterative EDA. To facilitate the management of such workflow and lower the barriers for adopting EDA-based parameter estimation in subsurface science, we develop a generic software frame-work linking the Data Assimilation Research Testbed (DART) with a massively parallel subsurface FLOw and TRANsport code PFLOTRAN. The new DART-PFLOTRAN leverages both the core data assimilation engines in DART and the computational power afforded by PFLOTRAN. In addition to the standard smoother and filtering options, DART-PFLOTRAN enables an iterative EDA workflow based on the Ensemble Smoother for Multiple Data Assimilation method (ES-MDA) to improve estimation accuracy for nonlinear forward problems. Here, we verify the implementation of ES-MDA in DART-PFLOTRAN using two synthetic cases designed to estimate static permeability and dynamic exchange fluxes across the riverbed, respectively, from continuous temperature measurements made across a depth profile. One-dimensional hydro-thermal simulations are performed in both cases to relate temperature responses with the parameters of interest. In the case of estimating dynamic parameters, we demonstrate the flexibility of DART-PFLOTRAN in automating sequential ES-MDA workflow, which will significantly reduce the time researchers spend on managing complex workflows in similar applications. Both studies yield accurate estimations of the parameters compared to their synthetic truth, while ES-MDA leads to more accurate estimation when a high level of nonlinearity exist between observed responses and unknown parameters. With a code base in Python and Fortran, DART-PFLOTRAN paves the way for applications in large-scale subsurface inverse modeling by automating the complex workflow of sequential ES-MDA that can be executed on various computing platforms.

97 MATHEMATICS AND COMPUTING↗

DIRAC current, upcoming and planned capabilities and technologies

DIRAC is the interware for building and operating large scale distributed computing systems. It is adopted by multiple collaborations from various scientific domains for implementing their computing models. DIRAC provides a framework and a rich set of ready-to-use services for Workload, Data and Production Management tasks of small, medium and large scientific communities having different computing requirements. The base functionality can be easily extended by custom components supporting community specific workflows. DIRAC is at the same time an aging project, and a new DiracX project is taking shape for replacing DIRAC in the long term. This contribution will highlight DIRAC’s current, upcoming and planned capabilities and technologies, and how the transition to DiracX will take place. Examples include, but are not limited to, adoption of security tokens and interactions with Identity Provider services, integration of Clouds and High Performance Computers, interface with Rucio, improved monitoring and deployment procedures.

97 MATHEMATICS AND COMPUTING↗

The Challenges of Releasing Human Data for Analysis

The NASA Johnson Space Center s (NASA JSC) Committee for the Protection of Human Subjects (CPHS) recently approved the formation of two human data repositories: the Lifetime Surveillance of Astronaut Health Repository (LSAH-R) for clinical data and the Life Sciences Data Archive Repository (LSDA-R) for research data. The establishment of these repositories forms the foundation for the release of data and information beyond the scope for which the data was originally collected. The release of clinical and research data and information is primarily managed by two NASA groups: the Evidence Base Working Group (EBWG), consisting of members of both repositories, and the LSAH Policy Board. The goal of unifying these repositories and their processes is to provide a mutually supportive approach to handling medical and research data, to enhance the use of medical and research data to reduce risk, and to promote the understanding of space physiology, countermeasures and other mitigation strategies. Over the past year, both repositories have received over 100 data and information requests from a wide variety of requesters. The disposition of these requests has highlighted the challenges faced when attempting to make data collected on a unique set of subjects available beyond the original intent for which the data were collected. As the EBWG works through each request, many considerations must be factored into account when deciding what data can be shared and how - from the Privacy Act of 1974 and the Health Insurance Portability and Accountability Act (HIPAA), to NASA s Health Information Management System (10HIMS) and Human Experimental and Research Data Records (10HERD) access requirements. Additional considerations include the presence of the data in the repositories and vetting requesters for legitimacy of their use of the data. Additionally, fair access must be ensured for intramural, as well as extramural investigators. All of this must be considered in the formulation of the charters, policies and workflows for the human data repositories at NASA.

Fitts, Mary↗

Collecting and Processing Earth Science Data Metrics at NASA ESDIS

Since the launch of Terra satellite in 1999, the number of Earth Science remote sensing data products created and distributed by NASA's Earth Observing System (EOS) Data and Information System (EOSDIS) has increased from a few hundred to nearly ten thousand. NASA's Earth Science Data and Information System (ESDIS) Metrics System (EMS) collects metrics on data ingest, archive, and distribution by its Distributed Active Archive Centers (DAACs) and the Science Investigator-led Systems (SIPS), known as Data Providers. These metrics are critical in helping NASA management as well as data producers in resource planning and gaining a wide range of knowledge of data users and data usage.EMS receives flat files, or log files of data archive, ingest, and distribution either in their raw format, such as Apache web logs, or text files of log records formatted by the Data Providers. Tens of millions of records are processed each day to extract metrics on data products, user information, distribution protocols and services, and so on. The metrics are then made available to designated parties.This presentation provides an overview of the EMS processing workflow and improvement efforts made in recent years to handle ever-increasing number of data records and new metrics requirements, discusses several key steps including mapping log records to data products and identifying user communities along with geo-distribution, and demonstrates typical metrics capabilities produced by the EMS system. Challenges and potential approaches to improve the system are also discussed.

Pan, Jianfu↗

Southwest Water Resources: Monitoring Surface Water Extents of Remote Stock Ponds in the Southwestern United States Using Earth Observing Systems for Enhanced Water Resources Management

Due to increasingly frequent and severe drought conditions in the southwestern US, land managers and livestock producers need to monitor stock ponds with increasing regularity. The ability to assess stock pond water levels with Earth observing satellite systems would enhance monitoring efforts of partners at the US Forest Service, Arizona Department of Game and Fish, and the Diablo Trust. This study employed Landsat 8 Operational Land Imager (OLI), Sentinel-1 C-band Synthetic Aperture Radar (C-SAR), and Sentinel-2 Multispectral Instrument (MSI) to monitor surface water extent for hundreds of critical stock ponds in Arizona. Using methods adapted from previously developed image processing workflows, this project conducted a time-series analysis to capture seasonal and interannual variations in surface water area between 2013 to 2021. In addition, end users can monitor the surface water extent of stock ponds through the developed Google Earth Engine software tool called Surface Water Identification and Forecasting Tool (SWIFT). SWIFT incorporates the Automated Water Extraction Index, Modified Normalized Difference Water Index, and Tasseled Cap-Wetness Index for optical imagery and the incidence angle, VV and VH polarization bands for Sentinel-1 imagery to detect small water bodies in the study area with an overall accuracy range of 88-93%. These tools will empower our partners to monitor the extents of water in their stock ponds remotely, enabling them to develop data-informed and sustainable management solutions for decades to come.

Rainey Aberle↗

Collaborative Resource Allocation

Collaborative Resource Allocation Networking Environment (CRANE) Version 0.5 is a prototype created to prove the newest concept of using a distributed environment to schedule Deep Space Network (DSN) antenna times in a collaborative fashion. This program is for all space-flight and terrestrial science project users and DSN schedulers to perform scheduling activities and conflict resolution, both synchronously and asynchronously. Project schedulers can, for the first time, participate directly in scheduling their tracking times into the official DSN schedule, and negotiate directly with other projects in an integrated scheduling system. A master schedule covers long-range, mid-range, near-real-time, and real-time scheduling time frames all in one, rather than the current method of separate functions that are supported by different processes and tools. CRANE also provides private workspaces (both dynamic and static), data sharing, scenario management, user control, rapid messaging (based on Java Message Service), data/time synchronization, workflow management, notification (including emails), conflict checking, and a linkage to a schedule generation engine. The data structure with corresponding database design combines object trees with multiple associated mortal instances and relational database to provide unprecedented traceability and simplify the existing DSN XML schedule representation. These technologies are used to provide traceability, schedule negotiation, conflict resolution, and load forecasting from real-time operations to long-range loading analysis up to 20 years in the future. CRANE includes a database, a stored procedure layer, an agent-based middle tier, a Web service wrapper, a Windows Integrated Analysis Environment (IAE), a Java application, and a Web page interface.

Wang, Yeou-Fang↗

Automated Network Services for Exascale Data Movement

The Large Hadron Collider (LHC) experiments distribute data by leveraging a diverse array of National Research and Education Networks (NRENs), where experiment data management systems treat networks as a “blackbox” resource. After the High Luminosity upgrade, the Compact Muon Solenoid (CMS) experiment alone will produce roughly 0.5 exabytes of data per year. NREN Networks are a critical part of the success of CMS and other LHC experiments. However, during data movement, NRENs are unaware of data priorities, importance, or need for quality of service, and this poses a challenge for operators to coordinate the movement of data and have predictable data flows across multi-domain networks. The overarching goal of SENSE (The Software-defined network for End-to-end Networked Science at Exascale) is to enable National Labs and universities to request and provision end-to-end intelligent network services for their application workflows leveraging SDN (Software-Defined Networking) capabilities. This work aims to allow LHC Experiments and Rucio, the data management software used by CMS Experiment, to allocate and prioritize certain data transfers over the wide area network. In this paper, we will present the current progress of the integration of SENSE, Multi-domain end-to-end SDN Orchestration with QoS (Quality of Service) capabilities, with Rucio, the data management software used by CMS Experiment.

Balcas, Justas↗

Optimizing Sample Collection and Accessibility through the Biospecimen and Tissue Sharing Collection (BTSC) Program

The Space Radiation Element (SRE) of the Human Research Program (HRP) is dedicated to establishing a robust biospecimen and tissue sharing collection (BTSC) program that enhances sample collection, tracking, access, distribution, and usability, with the goal of maximizing scientific return. By leveraging biospecimens and tissues from previous experiments, HRP effectively achieves its scientific objectives in characterizing and mitigating the human health impacts of spaceflight while optimizing resource utilization. To further improve the usability and accessibility of the current biospecimen archive, the project aims to expand upon NASA's existing resources and institutional knowledge, ensuring ongoing modernization. To facilitate seamless navigation of the program's workflow, an educational series on the BTSC program is provided to Principal Investigators (PIs). This comprehensive series equips PIs with crucial information on submitting their inventory via the BTSC Metadata Intake Form, ultimately leading to the public availability of their data on NASA's Life Science Portal (NLSP). Covering various aspects such as metadata submission instructions and backend processes for transferring metadata to the Laboratory Information Management System (LIMS), the series incorporates guidance from NASA's Biological Institutional Scientific Collection (NBISC) and Ames Life Sciences Data Archive (ALSDA). The BTSC program represents a significant stride towards enhancing the usability and accessibility of biospecimens for space research. By enabling NASA to deepen its understanding of the health implications of long-term spaceflight, this initiative plays a pivotal role in ensuring the safety and well-being of astronauts.

Shelita Renee Augustus↗

VerifyIO: Ensuring Correctness of Consistency Semantics in Parallel I/O

Abstract—High-performance computing (HPC) applications generate and consume substantial amounts of data, typically managed by parallel file systems. These applications access file systems either through the POSIX interface or by using highlevel I/O libraries. While the POSIX consistency model remains dominant in HPC, emerging file systems and popular I/O libraries increasingly adopt alternative consistency models that relax semantics in various ways, creating significant challenges for correctness and portability. This paper addresses these challenges by proposing a trace-driven I/O consistency verification workflow, implemented in our open-source tool, VerifyIO, which collects execution traces, detects data conflicts, and verifies proper synchronization against specified consistency models. Our extensive evaluation of 91 test case executions across three widely used I/O libraries with four I/O consistency models reveals critical consistency issues at both application and implementation levels.

Consistency Semantics↗