Engineering PapersSearch

SEARCH · Engineering Papers

Results for “trustworthy design”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

What Is the Agent Doing? Visualizing Agentic AI Querying Workflows

We explore how visualizations can help users understand what an AI agent is doing as it builds and runs queries over data. As part of the LinkQ system, a natural language interface for querying knowledge graphs with a large language model (LLM), we designed two complementary views: A State Diagram that shows where the agent is within a larger workflow, and a Live Action Display that gives real-time updates about the agent's current task. In a study with 14 practitioners, we found that these visuals helped participants build stronger mental models of the agent's behavior while also increasing their confidence in the system. However, we also observed that users sometimes trusted incorrect outputs simply because the agent appeared to be doing the "right" thing. Our findings point to both the value and risk of visualizing agent behavior in interactive AI systems.

97 MATHEMATICS AND COMPUTING

Robust Explanations using Diverse Adversarially Trained Ensembles, Multi-Modal Contrastive Learning, and Attribution-based Confidence Metrics

The primary objective of this project is to strengthen the trustworthiness of AI systems by designing algorithms that make their internal decision-making processes more understandable to human users. This involves creating clear, interpretable explanations for AI decisions and developing metrics to assess these explanations' validity and reliability. Significant progress has been achieved through (i) developing symbolic explanations, (ii) generating meaningful interpretive insights, (iii) establishing accuracy and confidence metrics, and (iv) devising methods to evaluate the knowledge boundaries of AI models. To date, the research findings have been shared in peer-reviewed publications, with accompanying scientific and technical information (STI) detailed below.

97 MATHEMATICS AND COMPUTING

AutoLabs: cognitive multi-agent systems with self-correction for autonomous chemical experimentation

The automation of chemical research through self-driving laboratories (SDLs) promises to accelerate scientific discovery, yet the reliability and granular performance of the underlying AI agents remain critical, under-examined challenges. In this work, we introduce AutoLabs, a self-correcting, multi-agent architecture designed to autonomously translate natural-language instructions into executable protocols for a high-throughput liquid handler. The system engages users in dialogue, decomposes experimental goals into discrete tasks for specialized agents, performs tool-assisted stoichiometric calculations, and iteratively self-corrects its output before generating a hardware-ready file. We present a comprehensive evaluation framework featuring five benchmark experiments of increasing complexity, from simple sample preparation to multi-plate timed syntheses. Through a systematic ablation study of 20 agent configurations, we assess the impact of reasoning capacity, architectural design (single- vs. multi-agent), tool use, and self-correction mechanisms. Our results demonstrate that agent reasoning capacity is the most critical factor for success, reducing quantitative errors in chemical amounts (nRMSE) by over 85% in complex tasks. When combined with a multi-agent architecture and iterative self-correction, AutoLabs approaches expert-authored reference procedures on the benchmark (F1-score > 0.89) on challenging multi-plate syntheses. These findings establish a clear blueprint for developing robust and trustworthy AI partners for autonomous laboratories, highlighting the synergistic effects of modular design, advanced reasoning, and self-correction to ensure both performance and reliability in high-stakes scientific applications. Code: https://github.com/pnnl/autolabs

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Hybrid learning techniques for scientific data reduction with performance guarantees

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Final report- UFL - RAPIDS2: A SciDAC Institute for Computer Science, Data, and Artificial Intelligence

The research initiatives supported by the U.S. Department of Energy (DOE) Grant DE-SC0022265 are fundamentally aimed at pioneering advanced machine learning (ML) techniques for scientific data compression within high-performance computing (HPC) environments. This comprehensive body of work addresses the critical challenge posed by the exponential growth of data generated by scientific simulations in domains such as fusion energy, climate modeling, and computational fluid dynamics (CFD). A core objective is to develop compression algorithms that achieve substantial data reduction—often by orders of magnitude—while rigorously ensuring the fidelity of both the primary data (PD) and scientifically crucial derived quantities of interest (QoI). The methodologies deployed under this grant integrate sophisticated deep learning architectures, prominently featuring autoencoders, advanced generative models like conditional diffusion, and hybrid learning techniques. Key innovations include the development of Guaranteed Autoencoders (GAE) and the Guaranteed Conditional Diffusion with Tensor Correction (GCDTC) framework, which provide explicit, instance-level error bounds on reconstructed data. Furthermore, specialized strategies such as nonlinear constraint satisfaction are employed to preserve the integrity of QoI, a vital requirement for the trustworthiness of downstream scientific analyses. This research also focuses on the design and implementation of scalable, GPU-accelerated software pipelines that seamlessly integrate into existing HPC workflows, ensuring both computational efficiency and practical applicability. The CAESAR framework, for example, unifies foundation and generative models to create an adaptive and efficient compression solution for spatio-temporal scientific data. Collectively, these efforts represent a significant advancement in mitigating the scientific data deluge, enabling more effective data management, accelerated scientific discovery, and optimized utilization of HPC resources.

97 MATHEMATICS AND COMPUTING

Virtual Reality for Shoot/No-Shoot Decision Training in Law Enforcement: A Literature Review and Research Agenda

Virtual reality (VR) can materially improve “shoot / no-shoot” (SNS) training by giving officers realistic, repeatable practice making high-stakes decisions under pressure. Traditional tools—live-fire ranges and video simulators—build basics, but they cannot adapt to each officer in real time or fully mirror the complexity of the field. VR closes that gap by creating immersive scenarios that are safer, more flexible, easier to scale across units, and able to capture objective performance data. SNS decisions are not just about marksmanship; they rely on perception, judgment, memory, and the ability to hold fire when a threat is uncertain. Effective training therefore needs realism, decision complexity, and branching outcomes that reflect the true consequences of choices. These elements strengthen recognition of hostile intent while reducing false positives and building the self-control required in ambiguous situations. VR brings specific advantages: dynamic environments, full-body interaction, and the ability to measure performance with precision—enabling targeted feedback and better transfer of learning to the street. At the same time, responsible deployment must address scenario quality (credible environments and behaviors), lawful decision models, and user wellbeing (appropriate stress levels, comfort, and safety). Sandia’s VIPER Lab is positioned to lead this work. The team combines human-performance science, AI/ML, and VR/AR development with a deep equipment bench (e.g., omnidirectional treadmill, eye-tracking, haptics, multiple HMDs). This ecosystem supports building and validating next-generation SNS training that is immersive, measurable, and trustworthy. Bottom line: Investment in VR-enabled SNS training that blends evidence-based design with careful validation and legal safeguards is expected to pay off in safer, more consistent decision-making and improved community trust, delivered through training that is practical to deploy at scale.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF

Assessing and Enabling Trustworthy Predictions for High-Consequence Decisions

Predictions from physics-based computational models provide critical information to inform high consequence decisions, e.g., engineering design decisions. The ability to assess the reliability of such predictions is therefore critical. However, to date, reliability assessment rely heavily on expert judgment and qualitative arguments. This report details the efforts of LDRD 233072 to develop quantitative methods to assess reliability of model predictions, especially in the context of simplifying assumptions that can impact their reliability.

42 ENGINEERING

Dependable classical-quantum computing systems engineering

Increasing evidence suggests quantum computing (QC) complements traditional High-Performance Computing (HPC) by leveraging its unique capabilities, leading to the emergence of a new, hybrid paradigm, QHPC. However, this integration introduces new challenges, with dependability–defined by reproducibility, resiliency, and security and privacy–emerging as a central concern for building trustworthy systems that provide an advantage to the users. This paper proposes a framework for dependable QHPC system design, organized around these three pillars. We identify integration challenges, anticipate roadblocks, and highlight productive synergies across QC, HPC, cloud platforms, and network security. Drawing from both classical computing principles and quantum-specific insights, we present a roadmap for co-design that supports robust hybrid architectures. Our approach offers concrete metrics for assessing dependability, provides design guidance for engineers working at the QC-HPC interface, and surfaces new engineering questions around complexity, scale, and fault tolerance. Ultimately, designing for dependability is key to realizing practical, scalable QHPC systems and accelerating the broader quantum ecosystem capable of translating quantum promises into actual application delivery.

HPC

Trusted Simulation: Considering Model Quality in the Context of User Trust

A high‐quality simulation model should help its users to easily and appropriately calibrate their trust in the model. Traditional evaluation metrics such as validation and robustness are necessary but insufficient for this task. Trust calibration depends on factors like the model's transparency, applicability to intended use, usability, reputation, and consideration of potential bias. This article proposes a framework for designing and evaluating system dynamics models by considering factors that contribute to the proper calibration of user trust. This framework takes inspiration from trusted artificial intelligence, broadening our traditional concept of model quality and explicitly focusing on what users need to consider a model trustworthy and to understand the model's relevance to its intended purpose. The trusted simulation framework can improve our integration of model quality activities throughout the modeling process, leading to more impactful and better‐targeted model design, development, and evaluation.

Naugle, Asmeret Bier [Sandia National Laboratories

Ecosystems for Scientific Computing in the Age of AI

Scientific computing is at an inflection point. Artificial intelligence (AI) is reshaping how scientific software is developed, how teams collaborate, how projects are governed, and how the next generation is trained. Drawing on insights from a 2025 workshop report, this article argues that the future of discovery will depend on agile, robust ecosystems built through socio-technical co-design—the intentional integration of technical and human systems. This perspective is essential for ensuring that future scientific computing remains trustworthy, sustainable, and scalable. It combines advances in AI, high-performance computing, and software with new models for cross-disciplinary collaboration, education, and workforce development. Key recommendations include building modular, trustworthy AI-enabled software ecosystems; enabling teams to integrate AI into scientific workflows while preserving human creativity, integrity, and rigor; and developing adaptive training pathways that keep pace with rapid technological change. By sharing these perspectives, we hope to stimulate broader community dialogue and encourage coordinated action.

AI

Dynamic Modeling and Characterization of Nuclear-grade Graphite

Idaho National Labs serves as the spearhead for many innovative energy solutions to the world's energy crisis. One such solution is the INL's Microreactor which is designed to deploy to extreme/remote environments where other sources of power are either unavailable or unreliable. In order to best design these energy solutions for their operational environments, it is crucial to understand how the design, components, and materials will respond to the environmental conditions. One key material in these innovative designs is a nuclear-grade graphite known as PCEA. This study examines the behavior of PCEA graphite under dynamic loading, similar to that which may occur in extreme environments. The objective is to characterize the dynamic behavior and produce an accurate, reliable constitutive model suitable for use in simulation tools such as INL's MOOSE. Graphite specimens were tested using a Split Hopkinson Pressure Bar (SHPB) to administer the dynamic compressive load. The SHPB was charged at various pressures to produce a range of strain rates on the material in compression. Data was acquired via strain gauges on the SHPB setup, from which stress, strain, and time data were collected. Analysis revealed the stress-strain behavior of the material as well as insights into the material behavior's relationship to strain rate. Further work must continue to characterize the various other dynamic behaviors of the material which will combine to create a substantially trustworthy constitutive model for this grade of nuclear-grade graphite. Ultimately, this will allow for realistic simulation of the material in reactor designs, allowing for prediction of design weaknesses and leading to improved designs for increased resilience, security, and reliability.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Deep Cyber-Physical Situational Awareness for Energy Systems: A Secure Foundation for Next-Generation Energy Management

This document provides the final report for the CYPRES project. The purpose is (1) to highlight and summarize its major accomplishments and (2) to provide guidance on how its outcomes have informed and can inform important additional research and technology transfer. The goal of CYPRES was the research, development, and demonstration of a security-oriented next generation cyber-physical EMS for electric power systems that detects malicious and abnormal events through the fusion of cyber and physical data. To achieve this, the CYPRES project team researched, developed, and built a prototype of the solution, referred to as the CYPRES EMS. The CYPRES EMS is a proof-of-concept cyber-physical platform that demonstrates the management of the energy system, communications, security, and cyber-physical grid modeling and analytics. As part of the capabilities of the CYPRES EMS, the team designed and developed a suite of power system applications for monitoring, risk analyses, detection, and control that are inherently cyberaware. At its core, the project aimed to research, develop, and demonstrate a security-oriented next-generation cyber-physical Energy Management System (EMS) capable of detecting malicious and abnormal events through the innovative fusion of cyber and physical data. This approach represents a fundamental shift from traditional EMS, reimagining how critical infrastructure can be protected through unified cyber-aware and physics-aware secure data flow pipelines. The project’s cornerstone deliverable, the CYPRES EMS, serves as a proof-of-concept cyber-physical platform that revolutionizes the management of energy systems, communications, security, and cyber-physical grid modeling and analytics. This prototype implements a comprehensive suite of power system applications for monitoring, risk analyses, detection, and control, all designed with inherent cyber awareness. The system’s architecture extends from end-devices in the field through to control center applications, establishing a secure and resilient control framework that addresses the challenges posed by diverse devices of unknown trustworthiness connecting to modern power systems. Through this innovative approach to deep cyber-physical situational awareness, the CYPRES project not only advances the state-of-the-art in energy infrastructure protection but also establishes a new paradigm for how EMS can be designed, deployed, and operated in an increasingly complex threat landscape. The findings and developments from this project provide crucial insights for stakeholders across the energy sector, offering a blueprint for enhancing the reliability and resilience of our nation’s critical energy infrastructure in the face of evolving cyber threats.

24 POWER TRANSMISSION AND DISTRIBUTION

Enabling end-to-end secure federated learning in biomedical research on heterogeneous computing environments with APPFLx

Facilitating large-scale, cross-institutional collaboration in biomedical machine learning (ML) projects requires a trustworthy and resilient federated learning (FL) environment to ensure that sensitive information such as protected health information is kept confidential. Specifically designed for this purpose, this work introduces APPFLx - a low-code, easy-to-use FL framework that enables easy setup, configuration, and running of FL experiments. APPFLx removes administrative boundaries of research organizations and healthcare systems while providing secure end-to-end communication, privacy-preserving functionality, and identity management. Furthermore, it is completely agnostic to the underlying computational infrastructure of participating clients, allowing an instantaneous deployment of this framework into existing computing infrastructures. Experimentally, the utility of APPFLx is demonstrated in two case studies: (1) predicting participant age from electrocardiogram (ECG) waveforms, and (2) detecting COVID-19 disease from chest radiographs. Here, ML models were securely trained across heterogeneous computing resources, including a combination of on-premise high-performance computing and cloud computing facilities. By securely unlocking data from multiple sources for training without directly sharing it, these FL models enhance generalizability and performance compared to centralized training models while ensuring data remains protected. In conclusion, APPFLx demonstrated itself as an easy-to-use framework for accelerating biomedical studies across organizations and healthcare systems on large datasets while maintaining the protection of private medical data.

Biomedical Research

On the Need to Align Intent and Implementation in Uncertainty Quantification for Machine Learning

Quantifying uncertainties for machine learning (ML) models is a foundational challenge in modern data analysis. This challenge is compounded by at least two key aspects of the field: (a) inconsistent terminology surrounding uncertainty and estimation across disciplines, and (b) the varying technical requirements for establishing trustworthy uncertainties in diverse problem contexts. In this position paper, we aim to clarify the depth of these challenges by identifying these inconsistencies and articulating how different contexts impose distinct epistemic demands. We examine the current landscape of estimation targets (e.g., prediction, inference, simulation-based inference), uncertainty constructs (e.g., frequentist, Bayesian, fiducial), and the approaches used to map between them. Drawing on the literature, we highlight and explain examples of problematic mappings. To help address these issues, we advocate for standards that promote alignment between the \textit{intent} and \textit{implementation} of uncertainty quantification (UQ) approaches. We discuss several axes of trustworthiness that are necessary (if not sufficient) for reliable UQ in ML models, and show how these axes can inform the design and evaluation of uncertainty-aware ML systems. Our practical recommendations focus on scientific ML, offering illustrative cases and use scenarios, particularly in the context of simulation-based inference (SBI).

Trivedi, Shubhendu [MIT] (ORCID:0000000312374301)

Trustworthiness and Trust: Identifying Factors that Drive Successful Human-AI Interaction in Nuclear Power Plant Applications

Emerging technologies such as artificial intelligence (AI) and machine learning (ML) are rapidly evolving and considered a promising tool for efficient and continued safe operations of the U.S. nuclear power plants (NPPs). Emerging AI techniques like large language models (LLMs) are one such technology that may support personnel at existing NPPs perform work more efficiently. For example, operators may query the current operational status of a power plant via a chat interface leveraging LLMs to access plant-related information in an interactive manner rather than manually collecting various sensor data for tasks such as surveillances or completing work orders. This is a fundamental shift in the way operators currently perform their tasks today. The literature of human-automation interaction indicates that trust is a crucial factor that drives successful interaction between a human operator and an automated system, like an AI-infused NPP application. This work presents the results of a literature review on key factors that relate to trust in AI/LLM technologies for NPP applications. The relevant literature of human factors and cognitive engineering has identified various factors related to trust including trustworthiness, performance characteristics, operator skill and perceived risk. This preliminary literature review will guide development and evaluation of models involving the identified factors influencing trust in AI and develop a framework for human-centered design for interface between humans and AI. By addressing trust, this work supports developing a technical basis for designing key characteristics of AI/LLM to support calibrated trust, which will ultimately support wide-scale adoption of AI/LLM technologies, as well as ensure safe, effective, and reliable use.

99 - GENERAL AND MISCELLANEOUS

Machine learning-accelerated discovery of heat-resistant polysulfates for electrostatic energy storage

The development of heat-resistant dielectric polymers that withstand intense electric fields at high temperatures is critical for electrification. Balancing thermal stability and electrical insulation, however, is exceptionally challenging as these properties are often inversely correlated. A traditional intuition-driven polymer design approach results in a slow discovery loop that limits breakthroughs. Here we present a machine learning-driven strategy to rapidly identify high-performance, heat-resistant polymers. A trustworthy feed-forward neural network is trained to predict key proxy parameters and down select polymer candidates from a library of nearly 50,000 polysulfates. The highly efficient and modular sulfur fluoride exchange click chemistry enables successful synthesis and validation of selected candidates. A polysulfate featuring a 9,9-di(naphthalene)-fluorene repeat unit exhibits excellent thermal resilience and achieves ultrahigh discharged energy density with over 90% efficiency at 200 °C. Its exceptional cycling stability underscores its promise for applications in demanding electrified environments.

Li, He

Designing a Comprehensive IDS Strategy for a Zero Trust Architecture Environment

Zero Trust Architecture or ZTA is a cybersecurity model for enterprises to structure their networked resources around to maintain total security externally and internally. In a Zero Trust environment, no part of the network is considered "trustworthy" and thus should be scrutinized and monitored extensively as is done in traditional "Trust But Verify" schemes at the network's perimeter. In this way, Zero Trust Architecture is a superior model for securing access to networked resources at the enterprise level. Fermilab, in pursuit of a better security posture, has decided to embrace this model of architecture for its network. Attaining this goal requires tremendous infrastructural, policy, and procedural adjustments that will affect all the lab's personnel and resources.

D'Antonio, Lucas

Equipment Self-Assessment Guide Checklist

This Equipment Self-Assessment Checklist is designed for asset owners and operators (AOOs) responsible for the deployment, operation, maintenance, or cybersecurity oversight of grid systems and digital energy technologies. It provides a structured inspection checklist for evaluating the security, integrity, and operational trustworthiness of equipment across substations, generation sites, distributed energy resources (DERs), and control environments.

32 - ENERGY CONSERVATION, CONSUMPTION, AND UTILIZA