Trust Maturity Model for AI Systems
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
This presentation summarizes our work in the PrOMMiS project on benchmarking of data-driven optimization algorithms and their applications in self-driving laboratories. This work supports the broader project goal of accelerating the identification of promising separation methods and operating conditions for critical minerals separation processes. We present a systematic benchmarking study of 42 data-driven optimization algorithms on a broad collection of 502 test problems. The results identify BAM, GLCCLUSTER, and MULTIMIN as the most effective optimization solvers, with BAM showing the highest overall performance and solving more than 80% of the benchmark problems. The study also shows that no single solver consistently outperforms the others across all problem types, indicating that our future laboratory applications may benefit from using a small set of strong solvers rather than relying on a single method. The presentation also illustrates an in-silico chemical reactor case study showing that data-driven optimization methods can guide autonomous experimentation in a self-driving laboratory and identify optimal operating conditions within a small number of experiments. Overall, the results provide a basis for selecting efficient optimization methods and demonstrate the practical use of data-driven optimization in self-driving laboratory workflows.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Abstract AI-advised Decision Making is a form of human-autonomy teaming in which an AI recommender system suggests a solution to a human operator, who is responsible for the final decision. This work seeks to examine the importance of judgement and shared situation awareness between humans and automated agents when interacting together in the form of a recommender systems. We propose manipulating both human judgement and shared situation awareness by providing the human decision maker with relevant information that the automated agent (AI), in the form of a recommender system, uses to generate possible courses of action. This paper presents the results of a two-phase between-subjects study in which participants and a recommender system jointly make a high-stakes decision. We varied the amount of relevant information the participant had, the assessment technique of the proposed solution, and the reliability of the recommender system. Findings indicate that this technique of supporting the human’s judgement and establishing a shared situation awareness is effective in (1) boosting the human decision maker’s situation awareness and task performance, (2) calibrating their trust in AI teammates, and (3) reducing overreliance on an AI partner. Additionally, participants were able to pinpoint the limitations and boundaries of the AI partner’s capabilities. They were able to discern situations where the AI’s recommendations could be trusted versus instances when they should not rely on the AI’s advice. This work proposes and validates a way to provide model-agnostic transparency into recommender systems that can support the human decision maker and lead to improved team performance.
In this paper, an Artificial Intelligence-based (AI) system is proposed for an 11-level cascaded H-bridge multilevel inverter (MLI) with the aims of harmonic suppression and reliability enhancement. The system consists of three seamlessly integrated Neural Networks (NNs). First, a multilayer perceptron is used to generalize the optimal switching angles for selective harmonic elimination under non-equal DC voltages. Next, an autoencoder NN estimates the voltage sensor readings to address potential drifting. Finally, a perceptron NN detects inverter faults based solely on the output voltage of the MLI. Simulation scenarios were evaluated, and the results show that the proposed system provides a comprehensive solution for the robust operation of the MLI. The proposed solution is capable of minimizing the targeted harmonics orders with minimal impact on the fundamental voltage, even when the voltage sensor drifts. Furthermore, the inverter under fault conditions was successfully identified.
Energy systems are experiencing various changes that impact the distribution, use, and reliability of energy. Local utilities and municipalities must respond and adapt to these changes, moving towards a future energy system with modernized infrastructure and other targeted investments and policy decisions. However, planning for and enacting these advancements requires significant effort from experts and engineers to develop strategies that ensure a reliable and secure energy future. This includes characterizing the current energy infrastructure, identifying areas for development, and engaging with local community members. Emerging generative artificial intelligence techniques can alleviate pain points and help support the development of the next generation of energy systems. Here, in this article, we highlight on-going generative AI work in the areas of atmospheric modeling, building energy management, and distribution network design, and we propose a vision for the role of generative AI that considers opportunities and identifies challenges inherent to this technology.
High Performance Computing (HPC) applications rely heavily on code optimizations to achieve good performance on modern CPU and GPU architectures. Traditional Machine Learning auto-tuning approaches have demonstrated success in exploring high-dimensional spaces, but they often require expensive compile-run evaluations and lack adaptability for large HPC applications. The recent advances in Large Language Models (LLMs) and Agentic AI systems raise intriguing questions about the potential of these approaches to address specific optimization methodologies. This work aims to answer an essential question for the HPC community: “How Agentic AI Systems Compare to Traditional ML Autotuning Techniques?” To address this question, we present a comparative analysis between a traditional ML-based optimization approach and an Agentic AI system, evaluating their respective capabilities and limitations for loop-level optimization. In addition, we introduced a new Agentic AI system named LoopGen-AI using three different Large Language Models: GPT-4.1, Claude 4.0, and Gemini 2.5. A key finding is that LoopGen-AI achieves competitive per-formance with only a few program runs, the reasoning logs from the agents revealed that their decisions rely heavily on the combination of semantic understanding of the target kernel with dynamic feedback from the environment, highlighting a promising new dimension in performance tuning. In contrast, ML-based autotuners focus on statistical exploration, and require orders of magnitude more runs to reach peak performance. Additionally, our analysis shows that prompt engineering, particularly using Persona + Context Manager patterns, significantly impacts the effectiveness of Agentic AI. Our results indicate that while Agentic AI systems are not yet a complete replacement for ML-based autotuners, it can effectively complement traditional methods.
Despite its significant benefits in enhancing the transparency and trustworthiness of artificial intelligence (AI) systems, explainable AI (XAI) can unintentionally provide adversaries with insights into blackbox models, increasing their vulnerability to various attacks. In this paper, we develop a novel explanation-driven adversarial attack against blackbox classifiers based on feature substitution, called XSub. The key idea of XSub is to strategically replace important features (identified via XAI) in the original sample with corresponding important features of a different label, thereby increasing the likelihood of the model misclassifying the perturbed sample. XSub only requires a minimal number of queries and can be easily extended to launch backdoor attacks in case the attacker has access to the model's training data. Our evaluation shows that XSub is not only effective and stealthy but also low-cost, showcasing its feasibility across a wide range of AI applications.
Modernization of energy systems has led to in- creased interactions among multiple critical infrastructures and diverse stakeholders making the challenge of operational decision making more complex and at times beyond cognitive capabilities of human operators. The state-of-the-art machine learning and deep learning approaches show promise of supporting users with complex decision-making challenges, such as those occurring in our rapidly transforming cyber-physical energy systems. However, successful adoption of data-driven decision support technology for critical infrastructure will be dependent on the ability of these technologies to be trustworthy and contextually interpretable. In this paper, we investigate the feasibility of implementing XAI for interpretable detection of cyberattacks in the energy system. Leveraging a proof-of-concept simulation use case of detection of a data falsification attack on a photovoltaic system using XGBoost algorithm, we demonstrate how Local Interpretable Model-Agnostic Explanations (LIME), a flavor XAI approach, can help provide contextual and actionable interpretation of cyberattack detection.
Earth embeddings—vector representations of geographic locations indexed in space and time—are emerging as a unifying interface for geospatial AI. However, their quality depends not only on model design, but on how multimodal Earth observation (EO) data are spatially indexed, temporally aligned, and cross-modally associated during pretraining. We introduce MOSAIC-CONUS (Multimodal Observations with Spatially Aligned Imagery, Urban Points of Interest, In-Situ Measurements and Text Captions), a large-scale EO dataset over the contiguous United States, organized around 250,000 stratified point indices that serve as stable spatial keys across seven modalities: active radar, passive optical imagery, lidar-derived elevation, land cover, functional context, hydrometeorological measurements, and textual summaries. Unlike existing EO datasets, MOSAIC-CONUS introduces four contributions not jointly addressed in prior work: 1. an open-source, large-scale multimodal EO corpus structured around point-indexed data designed to support Earth embedding learning; 2. explicit radar-optical pairing tables spanning twelve temporal alignment regimes, formalizing cross-sensor alignment as a controllable variable for analyzing how temporal mismatch across modalities influences learned embeddings quality; 3. a benchmark suite spanning cross-modal retrieval, annual nightlights regression, and basin-held-out streamflow prediction, positioning MOSAIC-CONUS as a benchmark-ready resource for multimodal AI systems; and 4. a language-based embedding layer through co-registered textual summaries, enabling Earth embeddings to function as a queryable interface for agentic AI systems. The dataset and pairing protocols are publicly released.
This study aims to evaluate a thermal energy storage (TES) system integrated with an active insulation system (AIS) to form a TES + AIS integrated wall system as a partition and as a secondary cooling system to shift the peak load and reduce cooling energy consumption. To understand and demonstrate its cooling performance, the TES + AIS integrated wall system was installed in an office building in Oak Ridge, Tennessee. To investigate the effect of the TES + AIS integrated wall system in a typical office building and various climate zones, the US Department of Energy’s prototype office building model was modified to accommodate the proposed system. In this work, results showed that the minimum size of the proposed system to achieve energy savings varies depending on climate conditions. The minimum size of the proposed system for cooling energy saving and shifting peak cooling demand in climate zones 2, 3, and 5 is 29.7 m 2 , whereas it is 44.6 m 2 in zones 4 and 6. By installing the minimum size of the proposed system, 11.3 % to 16.4 % of cooling energy can be shifted during discharge hours in a representative summer day.
Resilience is largely defined as the ability to adapt or recover from adverse conditions, stresses, attacks, or compromises on systems that use or are enabled by digital resources. In Artificial Intelligence Management and Research for Advanced Networked Testbed Hub (AMARANTH), resilience is measured in the amount of time it took from the beginning of a testing period for the model to reach predictions outside of the original 95% confidence interval or using the Kullback-Leibler (KL) divergence theorem, the Population Stability Index (PSI), and traditional methods such as root mean squared error (RMSE) threshold. Artificial Intelligence (AI) model drift is of significant concern when deploying AI-integrated systems into critical and/or secure environments. Drift can impact resilience of the AI-integrated system post-deployment and requires consistent maintenance and upkeep to ensure the model is accurate and precise. To quantify model drift and predict the point when a model's drift becomes unacceptable, we describe using Kullback-Leibler (KL) divergence, Population Stability Index (PSI) and/or confidence interval width estimations to determine the point of failure and time to failure of a model post-deployment. Through simple code functions, the KL-divergence, PSI, confidence interval, and root mean squared (RMSE) point of failures can be used to derive when a model needs to be maintained as well as the impact of adversarial action through statistical means.
Manufacturing industries continue to face challenges in reducing waste, as upstream strategies such as source reduction and product redesign require a deeper understanding of processes compared to conventional recycling methods. Recent advancements in artificial intelligence (AI) and machine learning (ML) have opened new opportunities to integrate modern computational techniques with traditional waste minimization strategies. This paper explores AI-enabled approaches for product redesign, source reduction, and recycling that can significantly reduce waste generation while improving efficiency and sustainability. AI-driven material substitution and lightweighting in product design enable discovery of novel materials with optimized properties, reducing waste without compromising performance. Reinforcement learning models optimize process parameters, raw material specifications, and machine sequencing to minimize production losses, while Industrial Internet of Things (IIoT) systems paired with AI analytics enhance real-time waste tracking, predictive maintenance, and quality inspection. Furthermore, AI-based demand forecasting and production planning reduce overproduction and excess inventory, as demonstrated in industrial applications. In recycling, ML-powered pattern recognition and robotic sorting technologies achieve higher accuracy in waste segregation, directly improving recycling efficiency. Complementary solutions such as smart bins and AI-enabled waste pickup scheduling optimize collection logistics, reducing both costs and emissions. Although implementation requires upfront investment in infrastructure and training, the long-term benefits include higher material efficiency, reduced waste, improved product quality, and stronger sustainability outcomes across the supply chain. By leveraging AI-enabled systems, manufacturers can align waste minimization efforts with circular economy principles, creating scalable solutions for both industry and society.
The AI for Experimental Controls project team at Jefferson Lab has developed an AI system to control and calibrate a large drift chamber system in near-real time. The AI system will monitor environmental and experimental variables to recommend voltage settings that maintain consistent dE/dx gain and optimal resolution throughout the experiment. At present, calibrations are performed after data have been recorded and require a considerable amount of time and attention from experts. The calibrations currently require multiple iterations and depend on accurate tracking information. Our approach uses environmental data, such as atmospheric pressure and gas temperature, and beam conditions, such as the flux of incident particles, as inputs to a Gaussian Process Regression (GPR) model. For the data taken during the GlueX 2020 run period, the GPR is able to predict the existing gain correction factors to within 3.5%. This talk will briefly describe the development, testing, and future plans for this system at Jefferson Lab.
A compact laser source and a single sideband modulator used therein is disclosed. The compact laser source includes a seed laser and one or more channels, with each channel generating one or more output laser beams having corresponding different wavelengths. The compact laser source can be formed in whole or in part on a single optical motherboard to thereby minimize space and power requirements. By employing the disclosed single sideband modulator, harmonics in the generated output laser beams can be minimized. The compact laser source finds application in an atom interferometer (AI) system, which may be used to measure gravity, acceleration, or rotation of the AI system.
Abstract Objectives Because of the high-risk nature of emergencies and illegal activities at sea, it is critical that algorithms designed to detect anomalies from maritime traffic data be robust. However, there exist no publicly available maritime traffic data sets with real-world expert-labeled anomalies. As a result, most anomaly detection algorithms for maritime traffic are validated without ground truth. Data description We introduce the HawaiiCoast_GT data set, the first ever publicly available automatic identification system (AIS) data set with a large corresponding set of true anomalous incidents. This data set—cleaned and curated from raw Bureau of Ocean Energy Management (BOEM) and National Oceanic and Atmospheric Administration (NOAA) automatic identification system (AIS) data—covers Hawaii’s coastal waters for four years (2017–2020) and contains 88,749,176 AIS points for a total of 2622 unique vessels. This includes 208 labeled tracks corresponding to 154 rigorously documented real-world incidents.
Traditional building envelopes have passive insulation systems that cannot respond to dynamic changes in the environment. An Active Insulation System (AIS) consists of Active Insulation Materials (AIMs) that dynamically vary the thermal conductivity of the insulation system. Several researchers have evaluated the impact of AIS on building thermal and energy performance by using simulation tools. Up to 70% savings in annual heating and cooling energy and significant reductions in peak demand have been predicted for some climates with wall systems employing AIS. However, materials and assembly development still need a cost-effective product that achieves the required performance. Here, in this study, we present the process of developing an AIS that we will install in a test hut for its performance evaluation. Minimum performance criteria of the AIS system are developed based on R-low/R-high ratio, required time and efficiency to switch states, and cost estimates. The following steps during this study are creating the concept to meet the requirements, predicting the performance via simulations, developing the experimental setup for bench-scale testing, and finally, constructing a full-scale wall assembly and monitoring the performance when exposed to environmental chamber tests. The selected approach uses off-the-shelf products to create an AIS that can switch R-value between 0.98 ft 2 ·°F·h/BTU (0.173 m 2 ·K/W) and 5.81 ft 2 ·°F·h/BTU (1.02 m 2 ·K/W) and have a switching time of less than one minute between R-high and R-low.