Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “trustworthy”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Dynamic Model Agnostic Reliability Evaluation of Machine-Learning Models Integrated in Instrumentation & Control Systems

In recent years, the field of data-driven neural network-based machine learning (ML) algorithms has grown significantly and spurred research in its applicability to instrumentation and control systems. While they are promising in operational contexts, the trustworthiness of such algorithms is not adequately assessed. Failures of ML-integrated systems are poorly understood; the lack of comprehensive risk modeling can degrade the trustworthiness of these systems. In recent reports by the National Institute for Standards and Technology, trustworthiness in ML is a critical barrier to adoption and will play a vital role in intelligent systems' safe and accountable operation. Thus, in this work, we demonstrate a real-time model-agnostic method to evaluate the relative reliability of ML predictions by incorporating out-of-distribution detection on the training dataset. It is well documented that ML algorithms excel at interpolation (or near-interpolation) tasks but significantly degrade at extrapolation. This occurs when new samples are "far" from training samples. The method, referred to as the Laplacian distributed decay for reliability (LADDR), determines the difference between the operational and training datasets, which is used to calculate a prediction's relative reliability. LADDR is demonstrated on a feedforward neural network-based model used to predict safety significant factors during different loss-of-flow transients. LADDR is intended as a "data supervisor" and determines the appropriateness of well-trained ML models in the context of operational conditions. Ultimately, LADDR illustrates how training data can be used as evidence to support the trustworthiness of ML predictions when utilized for conventional interpolation tasks.

97 MATHEMATICS AND COMPUTING↗

ATTRACTOR: Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability

Autonomous systems (AS) are crucial to realizing the vision of new, complex transportation modes, such as advanced air mobility (AAM) and urban air mobility (UAM). A chief barrier to induction of AS into aviation is insufficient understanding of AS reliability in time-critical and safety-critical environments—an obstacle to certification. ATTRACTOR is aimed at building a basis for certification of classes of autonomous cyber-physical-human systems (CPHS) via establishing metrics and models of trustworthiness and trust in multi-agent team interactions, analyzable trajectories, explainability of computational algorithms (explainable artificial intelligence, or XAI), and persistent modeling and simulation, in the context of missions planning and operation. By “building a basis for certification,” we mean acquiring an understanding of when a system is trustworthy and developing computable means to estimate trustworthiness and trust in order to eventually inform functional requirements that contribute to certification. The outcomes are applicable not just to aviation but to all domains that rely on autonomous systems.

Autonomy↗

Clarifying trust of materials property predictions using neural networks with distribution-specific uncertainty quantification

It is critical that machine learning (ML) model predictions be trustworthy for high-throughput catalyst discovery approaches. Uncertainty quantification (UQ) methods allow estimation of the trustworthiness of an ML model, but these methods have not been well explored in the field of heterogeneous catalysis. Herein, we investigate different UQ methods applied to a crystal graph convolutional neural network to predict adsorption energies of molecules on alloys from the Open Catalyst 2020 dataset, the largest existing heterogeneous catalyst dataset. We apply three UQ methods to the adsorption energy predictions, namely k-fold ensembling, Monte Carlo dropout, and evidential regression. The effectiveness of each UQ method is assessed based on accuracy, sharpness, dispersion, calibration, and tightness. Evidential regression is demonstrated to be a powerful approach for rapidly obtaining tunable, competitively trustworthy UQ estimates for heterogeneous catalysis applications when using neural networks. Recalibration of model uncertainties is shown to be essential in practical screening applications of catalysts using uncertainties.

36 MATERIALS SCIENCE↗

Scalable workflow for evaluating and optimizing large language models

This work describes the improved workflow for evaluating open-source large language models (LLMs) for trustworthiness. The workflow facilitates the acquisition of LLMs, the generation of LLM responses, and the evaluation of the responses for their trustworthiness. As a use case, the workflow is employed to evaluate dense, quantized, and pruned Meta Llama3.1 LLMs for their truthfulness. The outcome of the project could set the stage for understanding and developing trustworthy models in the future projects.

97 MATHEMATICS AND COMPUTING↗

NASA's EOSDIS, Trust and Certification

NASA's Earth Observing System Data and Information System (EOSDIS) has been in operation since August 1994, managing most of NASA's Earth science data from satellites, airborne sensors, filed campaigns and other activities. Having been designated by the Federal Government as a project responsible for production, archiving and distribution of these data through its Distributed Active Archive Centers (DAACs), the Earth Science Data and Information System Project (ESDIS) is responsible for EOSDIS, and is legally bound by the Office of Management and Budgets circular A-130, the Federal Records Act. It must follow the regulations of the National Institute of Standards and Technologies (NIST) and National Archive and Records Administration (NARA). It must also follow the NASA Procedural Requirement 7120.5 (NASA Space Flight Program and Project Management). All these ensure that the data centers managed by ESDIS are trustworthy from the point of view of efficient and effective operations as well as preservation of valuable data from NASA's missions. Additional factors contributing to this trust are an extensive set of internal and external reviews throughout the history of EOSDIS starting in the early 1990s. Many of these reviews have involved external groups of scientific and technological experts. Also, independent annual surveys of user satisfaction that measure and publish the American Customer Satisfaction Index (ACSI), where EOSDIS has scored consistently high marks since 2004, provide an additional measure of trustworthiness. In addition, through an effort initiated in 2012 at the request of NASA HQ, the ESDIS Project and 10 of 12 DAACs have been certified by the International Council for Science (ICSU) World Data System (WDS) and are members of the ICSUWDS. This presentation addresses questions such as pros and cons of the certification process, key outcomes and next steps regarding certification. Recently, the ICSUWDS and Data Seal of Approval (DSA) organizations merged their Core Trustworthy Data Repositories Requirements and require that members be recertified every three years. Given the rigor with which NASA manages the ESDIS Project and the DAACs, the recertification through WDSDSA, while involving some additional work, is a relatively simple process.

Certification↗

A Persistent Simulation Environment for Autonomous Systems

The age of Autonomous Unmanned Aircraft Systems (AUAS) is creating new challenges for the accreditation and certification requiring new standards, policies and procedures that sanction whether a UAS is safe to fly. Establishing a basis for certification of autonomous systems via research into trust and trustworthiness is the focus of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR), a new NASA Convergent Aeronautics Solution (CAS) project. Simulation Environments to test and evaluate AUAS decision making may be a low-cost solution to help certify that various AUAS systems are trustworthy enough to be allowed to fly in current general and commercial aviation airspace. NASA is working to build a peer-to-peer persistent simulation (P3 Sim) environment. The P3 Sim will be a Massively Multiplayer Online (MMO) environment were AUAS avatars can interact with a complex dynamic environment and each other. The focus of the effort is to provide AUAS researchers a low-cost intuitive testing environment that will aid training for and assessment of decisions made by autonomous systems such as AUAS. This presentation focuses on the design approach and challenges faced in development of the P3 Sim Environment is support of investigating trustworthiness of autonomous systems.

Kelley, Benjamin N.↗

A Distributed Simulation-to-Flight Framework to Support Investigating Trust/Trustworthiness in Multi-Agent Systems

As autonomous systems continue to grow both in use and complexity, the necessity for robust and extensible simulation-to-flight frameworks is paramount for establishing an effective architecture for autonomous systems. Hardware test flights are time-consuming and cost prohibitive during early system design and development. Simulation environments can be useful tools to accelerate algorithm development and testing. However, transitions from simulation to flight (sim-to-flight) can be challenging, unless systems are designed with this transition in mind and with the necessary capabilities built into the architecture and framework. One of the objectives of Autonomy Teaming and TRAjectories for Complex Trusted Operational Reliability (ATTRACTOR) was to design and develop a distributed mixed-reality simulation environment to begin establishing a basis for certification of autonomous systems via research into trust and trustworthiness. ATTRACTOR’s objective was to construct computational concepts of trustworthiness and justifiable trust in multi-agent autonomous teams, to inform future certification of safety-critical and time-critical autonomous systems in aviation. In this paper, we present an autonomous systems architecture and development framework paired with a persistent distributed modeling and simulation (ModSim) environment for test and evaluation of autonomous systems. They were designed under ATTRACTOR in order to measure and establish trustworthiness and trust in single-and multi-agent human-machine systems whether these machines are fixed-wing general aviation, rotary-wing Unmanned Aerial Vehicles (UAVs), ground rovers, or even spacecraft. The Autonomous Entity Operational Network (AEON) framework enables autonomous system development with an easily extensible collection of libraries and plug-n-play nodes facilitated by the Data Distribution Service (DDS) communication protocol standard. The Baseline Environment for Autonomous Modeling (BEAM) simulation environment is a distributed mixed-reality Unity™-based environment built around the same DDS communication paradigm allowing for easy integration with AEON-based autonomous applications, enabling sim-to-flight with minimal configuration changes. Using AEON and BEAM, source code that runs in simulation ports directly to hardware and has successfully flown in the National Airspace System (NAS) at NASA LaRC many times over the lifetime of ATTRACTOR.

Benjamin N Kelley↗

Considerations when Implementing Shift Work in Nuclear Operations

Many domestic and international nuclear facilities have processes that require personnel to work on some type of shift. Shift work has many different schedules such as night, morning, swing, and rotating shift. Some common shifts are DuPont, Pitman, and Panama. There are many psychological and physiological impacts to a person working shifts. These impacts can affect work performance, safety, and security within an organization. A well-established nuclear organization relies on a strong security culture that implements a trustworthiness or reliability program. The IAEA defines this type of program as individuals meeting the highest standards of reliability, trustworthiness, and physical and mental suitability. Psychological and physiological impacts that apply to a Trustworthiness Program are reliability (an individual’s ability to adhere to security and safety rules and regulations) and physical and mental suitability. If shift work isn’t properly implemented in Facility Operations, then several issues or human factors can arise. An employee that suffers from a poorly designed shift, gets overworked, or doesn’t have the proper rest period implemented, may suffer from shift work disorder, acute fatigue, or cumulative fatigue. Introducing fatigue into Facility Operations can have severe consequences to security, safety, production, and cost. Another challenge of shiftwork within a facility is laziness and complacency. Creating an atmosphere where the employee is willing to admit fatigue to his or her supervisor is a challenge. It is up to the organization to do the needed research on shift work and to implement the best shift for their facility and personnel.

Stockwell, Brandon↗

Explainable Artificial Intelligence Technology for Predictive Maintenance

The domestic nuclear power plant fleet has relied on labor-intensive and time-consuming preventive maintenance programs, thus driving up operation and maintenance costs to achieve high-capacity factors. Artificial intelligence and machine learning can help simplify complex problems, such as diagnosing equipment degradation, to enable more effective decision-making. Benefits will be felt not only within existing analog and digital instrumentation and control, but also work processes, the integration of people with technology, and most importantly, the business case. Together, these hold promise to make nuclear power more efficient and reduce costs associated with operation and maintenance. While the artificial intelligence and machine learning technologies hold significant promise in the nuclear industry, there are challenges or barriers to their adoption. This report outlines the those different machine learning adoption barriers (categorized as historical, technical, economic, regulatory, and user) that the industry must overcome to realize the full benefits of artificial intelligence and machine learning capabilities for long-term economic sustainability. This report also provides solutions for some of these barriers by focusing on improving the explainability of machine learning to encourage trust from the end-user. Trust and explainability are essential to machine learning adoption. This report focuses on research-developed solutions to some of these barriers while analyzing a non-safety-related system, namely the circulating water system. This system frequently experiences waterbox fouling which our models preemptively diagnoses then explains to the operator how those conclusions were reached. This report presents and discusses the inherent trade-off between machine learning performance (in terms of accuracy) and explainability, where highly accurate machine learning methods (such as deep-learning) are the least explainable, and the most explainable methods (such as decision trees) are the least accurate. In addition, explainability of artificial intelligence techniques in terms of transparency and post-hoc metrics are discussed. This report outlines the importance of data novelty and value of new information in evaluating both the explainability and trustworthiness. Novelty detection helps to establish consistency or inconsistency of the new data with respect to the training data. On the other hand, value of information could be a part of the user-centric visualization recommendation system that request additional information to be collected, thereby strengthening the machine learning outcomes. During this project, a copyrighted user-centric visualization that aligns with a human-in-the-loop approach was developed. The user-centric visualization presents different levels of information and can be tailored as per user credentials to gain user confidence. One of the salient features of the user-centric visualization is it presents machine learning methods with explainability metrics. A simplified version of the user-centric visualization was presented to 32 users with varying levels of machine learning expertise. Feedback was solicited to test the hypothesis that the app contained sufficient explainability and that the users would trust the algorithm. Overall, the app was positively received, and the hypothesis was supported. This report discusses the trust-but-verify framework – a potential approach to build user trust artificial intelligence. The framework discusses trust from the human level to artificial intelligence level. The fundamental premise of the trust but verify framework is derived from an observation of nuclear safety culture (i.e., nuclear power plant personnel do not rely on a singular source of data to make a decision). This also ties back to the user-centric visualization that presents different levels of information to achieve both explainability and trustworthiness of artificial intelligence. Even so, the adoption of artificial intelligence and machine learning in the nuclear industry faces additional barriers, namely regulatory and stakeholder readiness. To overcome these challenges, new solutions must gain regulatory approval and cater to stakeholder needs. The Nuclear Regulatory Committee has a 5-year strategic plan which prepares them for reviewing artificial intelligence technologies in licensee submissions. Early and frequent engagement with the regulator is encouraged. Additionally, artificial intelligence solutions should incorporate human-in-the-loop considerations and offer explainability. Stakeholders must prepare by hiring or training staff to adapt to advancing technology in everyday plant tasks.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Uncertainty Quantification using Deep Ensembles for Decision Making in Cyber-Physical-Human Systems

In this paper and its companion, Differential Equation Approximation Using Gradient-Boosted Quantile Regression, Robison et al., we examine an approach to quantifying model uncertainty with the aim of increasing the trustworthiness of computational models in human-machine interactions. In Differential Equation Approximation Using Gradient-Boosted Quantile Regression, we focus on gradient-boosted decision trees, while in this one, we give more details about deep ensembles. Uncertainty quantification is crucial for building trustworthy autonomous decision-making agents in human-machine teams. There are two types of uncertainties: aleatoric and epistemic. The former is related to the inherent stochasticity (noise) of the process, whereas the latter is associated with the lack of knowledge or representation capability of models, such as neural networks. By lack of knowledge, we mean the model’s inability to accurately predict outputs for all possible inputs. The aleatory uncertainty can be estimated fairly easily with, for example, filters, whereas epistemic uncertainty is challenging to compute. This paper uses deep ensembles to quantify both aleatory and epistemic uncertainty. It can act as an uncertainty-aware surrogate transition model for decision-making frameworks. "Uncertainty-aware" means that the surrogate transition model should make predictions along with confidence in those predictions. In the context of decision-making, the transition models are ordinary differential equations (ODEs). Since ODEs can be simulated to make one-step or multi-step predictions, a good surrogate model for them should perform reasonably well in both modes. In a multi-step approach, the trajectory sampling method TS∞ was used to propagate uncertainty over multiple steps. The cartpole dynamical system was selected to demonstrate the ability of deep ensembles as good surrogate transition models for decision-making frameworks. The deep ensembles modeled the dynamics of cartpole ODEs and made uncertainty-aware predictions in single-step and multi-step transition modes.

CPH systems↗

Extreme sparsification of physics-augmented neural networks for interpretable model discovery in mechanics

Data-driven constitutive modeling with neural networks has received increased interest in recent years due to its ability to easily incorporate physical and mechanistic constraints and to overcome the challenging and time-consuming task of formulating phenomenological constitutive laws that can accurately capture the observed material response. However, even though neural network-based constitutive laws have been shown to generalize proficiently, the generated representations are not easily interpretable due to their high number of trainable parameters. Sparse regression approaches exist that allow for obtaining interpretable expressions, but the user is tasked with creating a library of model forms which by construction limits their expressiveness to the functional forms provided in the libraries. Here, in this work, we propose to train regularized physics-augmented neural network-based constitutive models utilizing a smoothed version of $L^0$-regularization. This aims to maintain the trustworthiness inherited by the physical constraints, but also enables interpretability which has not been possible thus far on any type of machine learning-based constitutive model where model forms were not assumed a priori but were actually discovered. During the training process, the network simultaneously fits the training data and penalizes the number of active parameters, while also ensuring constitutive constraints such as thermodynamic consistency. We show that the method can reliably obtain interpretable and trustworthy constitutive models for compressible and incompressible hyperelasticity, yield functions, and hardening models for elastoplasticity, using synthetic and experimental data. This work aims to set a new paradigm for interpretable machine learning models in the broad area of solid mechanics where low and limited data is available along with prior knowledge of physical constraints that the learned maps need to obey. This paradigm can potentially be extended to a broader spectrum of scientific exploration.

Data-driven constitutive models↗

Demystifying Cyberattacks: Potential for Securing Energy Systems With Explainable AI : Preprint

Modernization of energy systems has led to in- creased interactions among multiple critical infrastructures and diverse stakeholders making the challenge of operational decision making more complex and at times beyond cognitive capabilities of human operators. The state-of-the-art machine learning and deep learning approaches show promise of supporting users with complex decision-making challenges, such as those occurring in our rapidly transforming cyber-physical energy systems. However, successful adoption of data-driven decision support technology for critical infrastructure will be dependent on the ability of these technologies to be trustworthy and contextually interpretable. In this paper, we investigate the feasibility of implementing XAI for interpretable detection of cyberattacks in the energy system. Leveraging a proof-of-concept simulation use case of detection of a data falsification attack on a photovoltaic system using XGBoost algorithm, we demonstrate how Local Interpretable Model-Agnostic Explanations (LIME), a flavor XAI approach, can help provide contextual and actionable interpretation of cyberattack detection.

artificial intelligence↗

Ecosystems for Scientific Computing in the Age of AI

Scientific computing is at an inflection point. Artificial intelligence (AI) is reshaping how scientific software is developed, how teams collaborate, how projects are governed, and how the next generation is trained. Drawing on insights from a 2025 workshop report, this article argues that the future of discovery will depend on agile, robust ecosystems built through socio-technical co-design—the intentional integration of technical and human systems. This perspective is essential for ensuring that future scientific computing remains trustworthy, sustainable, and scalable. It combines advances in AI, high-performance computing, and software with new models for cross-disciplinary collaboration, education, and workforce development. Key recommendations include building modular, trustworthy AI-enabled software ecosystems; enabling teams to integrate AI into scientific workflows while preserving human creativity, integrity, and rigor; and developing adaptive training pathways that keep pace with rapid technological change. By sharing these perspectives, we hope to stimulate broader community dialogue and encourage coordinated action.

AI↗

ChatBLAS: The First AI-Generated and Portable BLAS Library

We present ChatBLAS, the first AI-generated and portable Basic Linear Algebra Subprograms (BLAS) library on different CPU/GPU configurations. The purpose of this study is (i) to evaluate the capabilities of current large language models (LLMs) to generate a portable and HPC library for BLAS operations and (ii) to define the fundamental practices and criteria to interact with LLMs for HPC targets to elevate the trustworthiness and performance levels of the AI-generated HPC codes. The generated C/C++ codes must be highly optimized using device-specific solutions to reach high levels of performance. Additionally, these codes are very algorithm-dependent, thereby adding an extra dimension of complexity to this study. We used OpenAI’s LLM ChatGPT and focused on vector-vector BLAS level-1 operations. ChatBLAS can generate functional and correct codes, achieving high-trustworthiness levels, and can compete or even provide better performance against vendor libraries.

Valero Lara, Pedro↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

High-Quality Revision of the Israeli Seismic Bulletin

Seismic bulletins, with trustworthy phase picks, origin times, and source locations are key for regional seismic studies, such as travel-time (TT) tomography, attenuation tomography, and anisotropy studies. To lay the groundwork for such studies in Israel, we revised the seismic bulletin of Israel and the surrounding area and obtained a trustworthy TT data set. From the earthquake and explosion bulletins of the Geophysical Institute of Israel, we compiled a starting data set of about 123,000 earthquakes and explosions that occurred during the past 40 yr. After screening out the poorly recorded events, we were left with a data set of ~38,000 well-recorded events. We then revised the remaining data set in two consecutive steps. In the first, we reviewed and updated station metadata, including changes in station metadata parameters over time. In the second step, we jointly relocated a list of selected seismic events, using the Bayesian hierarchical location software package (BayesLoc) of Myers et al. (2007) that performs joint relocation of multiple events. We observed striking dissimilarities between the spatial distributions of the newly relocated catalog and the initial locations. Although the depth distribution of the starting catalog is trimodal with peaks at 0, 5, and 10 km, the distribution in this study is unimodal, with a broad peak between 7.5 and 12.5 km. By differencing the observed arrival times and the origin times obtained through relocation with BayesLoc, we obtained a revised TT database that consists of 261,336 Pg, 132,876 Pn, 114,816 Sg, and 60,394 Sn arrivals, from a set of 30,458 jointly relocated seismic sources. In this work, we compared prerevision and postrevision TTs as a function of epicentral distance and concluded that the revised data set contains far fewer outliers and inconsistencies than the original data set. The revised TT data set may be used for seismic studies, such as TT tomography, attenuation tomography, and anisotropy studies.

58 GEOSCIENCES↗

SAGE Intrusion Detection System: Sensitivity Analysis Guided Explainability for Machine Learning.

This report details the results of a three-fold investigation of sensitivity analysis (SA) for machine learning (ML) explainability (MLE): (1) the mathematical assessment of the fidelity of an explanation with respect to a learned ML model, (2) quantifying the trustworthiness of a prediction, and (3) the impact of MLE on the efficiency of end-users through multiple users studies. We focused on the cybersecurity domain as the data is inherently non-intuitive. As ML is being using in an increasing number of domains, including domains where being wrong can elicit high consequences, MLE has been proposed as a means of generating trust in a learned ML models by end users. However, little analysis has been performed to determine if the explanations accurately represent the target model and they themselves should be trusted beyond subjective inspection. Current state-of-the-art MLE techniques only provide a list of important features based on heuristic measures and/or make certain assumptions about the data and the model which are not representative of the real-world data and models. Further, most are designed without considering the usefulness by an end-user in a broader context. To address these issues, we present a notion of explanation fidelity based on Shapley values from cooperative game theory. We find that all of the investigated MLE explainability methods produce explanations that are incongruent with the ML model that is being explained. This is because they make critical assumptions about feature independence and linear feature interactions for computational reasons. We also find that in deployed, explanations are rarely used due to a variety of reason including that there are several other tools which are trusted more than the explanations and there is little incentive to use the explanations. In the cases when the explanations are used, we found that there is the danger that explanations persuade the end users to wrongly accept false positives and false negatives. However, ML model developers and maintainers find the explanations more useful to help ensure that the ML model does not have obvious biases. In light of these findings, we suggest a number of future directions including developing MLE methods that directly model non-linear model interactions and including design principles that take into account the usefulness of explanations to the end user. We also augment explanations with a set of trustworthiness measures that measure geometric aspects of the data to determine if the model output should be trusted.

97 MATHEMATICS AND COMPUTING↗

Oak Ridge National Laboratory Pilot Demonstration of an Attestation and Anomaly Detection Framework using Distributed Ledger Technology for Power Grid Infrastructure

This report summarizes the design and pilot demonstration of a framework called Grid Guard that was created to provide increased data and device trustworthiness to electric grid devices by leveraging distributed ledger technology (DLT), specifically blockchain. Grid Guard contains a combination of core cryptographic methods such as the secure hash algorithm (SHA), and asymmetric cryptography, private permissioned blockchain, baselining configuration data, consensus algorithm (Raft) and the Hyperledger Fabric (HLF) framework. The system implements a low energy, fast, and robust enhancement to system trustworthiness within and across electric grid systems such as substations, control centers and metering infrastructures. Blockchain is a distributed database structured that provides a practically unalterable (immutable) timeline of stored transactions. By relying on hashing and the Raft consensus algorithm, if an entity tries to illegitimately alter a record at one instance of the database the other ledger nodes are not altered. They work to cross-reference each other and easily locate any incorrectly added data and remove it. The bulk raw data is stored in an off-chain storage (outside of the blockchain ledger) and a hash of this baseline data is stored in the Blockchain ledger via hashing windows of time-series and configuration data, after aggregation and filtering. The bulk off-chain data repository is then considered to be trust-anchored using the hashes stored in the blockchain. To secure the electric grid testbed devices and data, device configuration baselines were compared to those baselines that had been previously stored in the ledger. Statistical baselines for device configurations, network communication patterns, and high-speed sensor data are calculated and then stored off-chain and hashes stored in the ledger. Measurements such as three-phase voltage and current, frequency, breaker status, protection scheme settings, network configuration settings (and other device configuration artifacts) and network traffic features (packet interarrival times) are compared every minute or other selected time windows. During phase 1 of the Grid Guard DLT project different DLT technologies were studies, and an assessment was performed on DLT technology vulnerabilities, uses, and key characteristics. DLT consensus protocols were studies (e.g., RAFT, named after Reliable, Replicated, Redundant, And Fault-Tolerant). Also, cryptography, public, private and permissioned or permissionless systems were assessed. Grid Guard implements a permissioned private DLT. Consensus algorithm selection and choice of DLT implementation depended heavily on the use-case. For this use-case, parameters were selected to measure performance and existing tools for assessment. Benchmarking was performed theoretically and practically. During phase 2 hashed transactions/blocks were inserted into the ledger every second. During phase 2 of the Grid Guard DLT project, a prototype framework was developed and demonstrated for attestation of critical substation devices and data using precision timing systems that use PTP and IRIG-B protocols) on a testbed of operational devices that emulated a distribution substation, control center, and power metering infrastructure using real Operational Technology (OT). The testbed includes OT devices such as protective relays, human machine interfaces (HMI), and power meters. To determine when to collect and compare system and network baselines, an initial examination of an anomaly detection capability to identify malicious manipulation of data streams was conducted. The resulting anomaly detection was demonstrated in a set of experiments and leveraged to trigger device artifact attestation checks. Attestation checks occur against device configuration baselines when compared with the immutable blockchain-stored baselines, which provided a cryptographically supported means by which to store baselines. The electrical substation-grid testbed was created to test the Grid Guard framework. The testbed emulates the operations of a portion of a power grid and SCADA systems as closely as possible. The testbed integrates real protocols, mainly IEC 61850 standard protocols, such as the Sampled Value (SV) and the GOOSE protocols. The testbed also supports DNP3 and other layer 2 and layer 3 protocols such as Telnet, SSH, SFTP/FTP and other proprietary protocols needed to connect to industrial control system equipment. The testbed emulates real power conditions using the OpalRT hardware-in-the-loop (HIL) device which can create fault situations that cannot be easily tested on real systems. The electrical substation-grid testbed was created using real measurement, communication, and protection devices that electrical utilities commonly use.

24 POWER TRANSMISSION AND DISTRIBUTION↗