Engineering PapersSearch

SEARCH · Engineering Papers

Results for “explainable machine learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Machine learning-powered data cleaning for LEGEND: a semi-supervised approach using affinity propagation and support vector machines

Neutrinoless double-beta decay ($0\nu\beta\beta$) is a rare nuclear process that, if observed, will provide insight into the nature of neutrinos and help explain the matter-antimatter asymmetry in the Universe. The large enriched germanium experiment for neutrinoless double-beta decay (LEGEND) will operate in two phases to search for $0\nu\beta\beta$. The first (second) stage will employ 200 (1000) kg of High-Purity Germanium (HPGe) enriched in 76 Ge to achieve a half-life sensitivity of 10 27 (10 28 ) years. In this study, we present a semi-supervised data-driven approach to remove non-physical events captured by HPGe detectors powered by a novel artificial intelligence model. We utilize affinity propagation to cluster waveform signals based on their shape and a support vector machine to classify them into different categories. We train, optimize, and test our model on data taken from a natural abundance HPGe detector installed in the Full Chain Test experimental stand at the University of North Carolina at Chapel Hill. We demonstrate that our model yields a maximum sacrifice of physics events of $0.024 ^{+0.004}_{-0.003} \%$ after data cleaning. Our model is being used to accelerate data cleaning development for LEGEND-200 and will serve to improve data cleaning procedures for LEGEND-1000.

artificial intelligence

Barriers to adopting artificial intelligence and machine learning technologies in nuclear power

Artificial intelligence and machine learning (AI/ML) technologies offer unique opportunities to transform nuclear plant operations and power generation. Benefits will be felt not only within existing analog and digital instrumentation and control, but also within work processes, the integration of people with technology and most importantly, the business case. The application of this new technology can help simplify complex problems and produce more effective decision-making, making nuclear power safer, more efficient, and more economically viable in the current energy market. Nonetheless, there are potential barriers to its adoption that must be overcome. The purpose of this paper is to categorize, review, and discuss barriers to AI/ML adoption within the nuclear power industry, with a focus on existing commercial reactors. Unique considerations for advanced reactors are also offered. Here we provide a comprehensive overview of the historical, technical, and business barriers that the industry faces, as well as stakeholder readiness, and end-user acceptance. We underscore the importance of user experience and offer potential solutions in overcoming each barrier. These include provisions for easier plant data access, a friendly regulatory environment, and investment in user trust and explainable AI.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill

Predictive models of the genetic bases underlying budding yeast fitness in multiple environments

Abstract The ability of organisms to adapt and survive depends on the effects of genes and the environment on fitness. However, the multigenic nature of fitness and genotype-by-environment interactions hinder our understanding of the genetic basis of fitness. Here, we established fitness prediction models for 35 environments using machine learning and existing fitness data and different genetic variant types for a Saccharomyces cerevisiae population. Models revealed that the predictive ability of genetic variants varied across environments, with copy number variants explaining the majority of fitness variation in most cases. Model interpretation showed that different variant types identified distinct gene sets associated with predictive variants. These gene sets were significantly enriched in experimentally validated genes affecting fitness in only a subset of environments, indicating that many genes influencing fitness remain unexplored. Notably, non-experimentally validated genes were more important than validated ones for fitness predictions. Gene contributions to predictions were both isolate- and environment-dependent, pointing to gene-by-gene and gene-by-environment interactions. Furthermore, models uncovered experimentally validated and novel candidate genetic interactions for a well-characterized stress, the fungicide benomyl. These findings highlight the feasibility of identifying the genetic basis of fitness by using different genetic variant types and offer novel targets for future functional analysis.

DNA copy number variations

Demonstration and Evaluation of Explainable and Trustworthy Predictive Technology for Condition-based Maintenance

The domestic nuclear power plant (NPP) fleet has historically relied on labor-intensive and time-consuming predictive maintenance (PdM) programs, thus driving up operation and maintenance (O&M) costs to achieve high-capacity factors. Artificial intelligence (AI) and machine-learning (ML) can help simplify complex problems such as diagnosing equipment degradation to enable more effective decision-making efforts. The benefits of AI will be felt through more efficient plant O&M, improved work processes, and better integration of people and technology. Together, these benefits hold the promise to make nuclear power more sustainable by reducing O&M costs while improving employee engagement. While AI and ML technologies hold significant promise for the nuclear industry, there are challenges or barriers to their adoption. Explainability and trustworthiness of AI are two salient challenges that need to be addressed for wider deployment of these technologies in NPPs. This research focuses specifically on addressing the explainability and trustworthiness of AI technologies to advance the human, technical, and organization (HTO) readiness levels in adopting a risk-informed PdM strategy at commercial NPPs. In addition, this approach can be adapted to enhance the acceptability of AI in other nuclear applications with a few application-specific modifications. The technical approach ensuring wider adoption of AI technologies was developed by Idaho National Laboratory (INL)—in collaboration with Public Service Enterprise Group (PSEG), Nuclear, LLC—by utilizing the circulating water system (CWS) at two PSEG-owned plant sites for demonstration. Focused user studies were performed in collaboration with subject matter experts (SMEs) from PSEG and other nuclear domains to enhance human and organization readiness by building trust in AI-informed technologies. VIsualization for PrEdictive maintenance Recommendation (VIPER)—a Battelle Energy Alliance, LLC, copyrighted software—was developed and expanded to provide a user-centric visualization by incorporating inputs from the collaborating utility, human factors engineering guidelines, and data analysts. The VIPER software enables users, who may be unfamiliar with ML in general, to be interactively engaged by asking technical questions about PdM, work orders, diagnosis results and their confidence levels, the kind of data being used, and the types of ML algorithms employed. This interactive engagement enhances explainability and builds trust. One of the enabling accomplishments was the integration of large language models (LLMs), both text-based and vision-based, in the VIPER software.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Understanding Twinning and Deformation in High Entropy Alloys

A combination of high strength and high ductility has been observed in multi-principal element alloys due to twin formation attributed to low stacking fault energy (SFE). In the pursuit of low SFE alloys, a key bottleneck is the lack of understanding of the composition–SFE cor- relations that would guide tailoring SFE via alloy composition. Using density functional theory (DFT), we show that dopant radius, which have been postulated as a key descriptor for SFE in dilute alloys, does not fully explain SFE trends across different host metals. Instead, charge density is a much more central descriptor. It allows us to (1) explain contrasting SFE trends in Ni and Cu host metals due to various dopants in dilute concentrations, (2) explain the large SFE variations observed in the literature even within a given alloy composition due to the nearest neighbor environments in “model” concentrated alloys, and (3) develop a machine learning model that can be used to predict SFEs in multi-elemental alloys. This model opens a possibility to use charge density as a descriptor for predicting SFE in alloys. Furthermore, a descriptor-less machine learning (ML) model based only on charge density images extracted from density functional theory (DFT) is developed to predict stacking fault energies (SFE) in concentrated alloys. The model is based on convolutional neural networks (CNNs) as one of the promising ML techniques for dealing with complex images and data. Identification of correct descriptors is a key bottleneck to develop ML models for predicting materials properties. Often, in most ML models, textbook physical descriptors such as atomic radius, valence charge and electronegativity are used as descriptors which have limitations because these properties change in concentrated alloys when multiple elements are mixed to form a solid solution. We illustrate that, within the scope of DFT, the search for descriptors can be circumvented by electronic charge density, which is the backbone of the Kohn-Sham DFT and describes the system completely. The performance of our model is demonstrated by predicting SFE of concentrated alloys with an RMSE and R2 of 6.18 mJ/m2 and 0.87, respectively, validating the accuracy of the proposed approach.

36 MATERIALS SCIENCE

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING

Explaining drivers of housing prices with nonlinear hedonic regressions

Housing markets play a critical role in shaping the spatial and demographic evolution of urban areas. Simulating housing price dynamics can enhance projections of future urban development outcomes. However, traditional hedonic regressions for housing prices, which neglect nonlinear interactions among explanatory variables, often exhibit limited predictive performance. While machine learning (ML) methods can provide a more flexible representation of the relationships between predictors, they are often regarded as “black boxes” due to their complexity and lack of transparency. Interpretable ML techniques provide a promising route by combining the flexibility of ML methods with approaches to analyze the relationships between inputs and outputs. In this study, we employ interpretable ML to analyze the patterns driving the housing market in Baltimore, Maryland, USA. We train an Artificial Neural Network (ANN) to predict Baltimore housing prices based on structural characteristics (e.g., home size, number of stories) and locational attributes (e.g., distance to the city center). We then conduct sensitivity and Partial Dependence Plot (PDP) analyses to interpret the fitted ANN model. We find that the ML model achieves higher predictive accuracy and explains 16 % more of housing price variance than a traditional linear regression model. The interpretable ML model also reveals more nuanced and realistic nonlinear relationships between housing sales price and predictors as well as interactive effects underlying Baltimore home price dynamics. For instance, while the linear model indicates a steady housing price increase over time, our interpretable ML model detects a post-2008 decline, with smaller properties experiencing the sharpest drop.

97 MATHEMATICS AND COMPUTING

Origin of enhanced performance when Mn-rich rocksalt cathodes transform to δ -DRX

Most Mn-rich cathodes are known to undergo phase transformation into structures resembling spinel-like ordering upon electrochemical cycling. Recently, the irreversible transformation of Ti-containing Mn-rich disordered rock-salt cathodes into a phase — named δ — with nanoscale spinel-like domains has been shown to increase energy density, capacity retention, and rate capability. However, the nature of the boundaries between domains and their relationship with composition and electrochemistry are not well understood. In this work, we discuss how the transformation into the multi-domain structure results in eight variants of Spinel domains, which is crucial for explaining the nanoscale domain formation in the δ -phase. We study the energetics of crystallographically unique boundaries and the possibility of Li-percolation across them with a fine-tuned CHGNet machine learning interatomic potential. Energetics of 16 d vacancies reveal a strong affinity to segregate to the boundaries, thereby opening Li-pathways at the boundary to enhance long-range Li-percolation in the δ structure. Defect calculations of the relatively low-mobility Ti show how it can influence the extent of Spinel ordering, domain morphology and size significantly; leading to guidelines for engineering electrochemical performance through changes in composition.

Anand, Shashwat

Enabling integrated AI control on DIII-D: a control system design with state-of-the-art experiments

We present the design and application of a general algorithm for Prediction And Control using MAchiNe learning (PACMAN) in DIII-D. Machine learning (ML)-based predictors and controllers have shown great promise in achieving regimes in which traditional controllers fail, such as tearing mode (TM) free scenarios, ELM-free scenarios and stable advanced tokamak conditions. The architecture presented here was deployed on DIII-D to facilitate the end-to-end implementation of advanced control experiments, from diagnostic processing to final actuation commands. This paper describes the detailed design of the algorithm and explains the motivation behind each design point. We also describe several successful ML control experiments in DIII-D using this algorithm, including a reinforcement learning controller targeting advanced non-inductive plasmas, a wide-pedestal quiescent H-mode ELM predictor, an Alfvén Eigenmode controller, a Model Predictive Control plasma profile controller and a state-machine TM predictor-controller. There is also discussion on guiding principles for real-time ML controller design and implementation.

machine learning

Evapotranspiration Partitioning Using Flux Tower Data in a Semi-Arid Ecosystem

Information about evapotranspiration (ET) and its components, that is, evaporation and transpiration, is crucial for a wide range of water and ecosystem management applications. However, partitioning ET into its two components is often challenging because of their spatiotemporal variabilities and lack of process understanding. This study developed a machine learning (ML) framework to shed light on ET processes and assess the relative importance of different drivers by incorporating hydrometeorology and biomass productivity variables. The Shapley Additive Explanations (SHAP) approach was applied to enhance explainability and rank the importance of ET drivers and their components. A total of 62 variables covering hydrometeorological and biomass productivity dimensions were considered from the Reynolds Creek Critical Zone Observatory (CZO) station in Idaho. The variable importance assessment identified the leading drivers individually for evaporation, transpiration and ET (soil water content for evaporation, vapour pressure deficit for transpiration and soil water content for ET). The results further highlighted the value of combining hydrometeorological and biomass productivity variables to achieve better predictability of ET processes.

54 ENVIRONMENTAL SCIENCES

Comparative modeling reveals the molecular determinants of aneuploidy fitness cost in a wild yeast model

Although implicated as deleterious in many organisms, aneuploidy can underlie rapid phenotypic evolution. However, aneuploidy will be maintained only if the benefit outweighs the cost, which remains incompletely understood. To quantify this cost and the molecular determinants behind it, we generated a panel of chromosome duplications in Saccharomyces cerevisiae and applied comparative modeling and molecular validation to understand aneuploidy toxicity. We show that 74%–94% of the variance in aneuploid strains’ growth rates is explained by the cumulative cost of genes on each chromosome, measured for single-gene duplications using a genomic library, along with the deleterious contribution of small nucleolar RNAs (snoRNAs) and beneficial effects of tRNAs. Machine learning to identify properties of detrimental gene duplicates provided no support for the balance hypothesis of aneuploidy toxicity and instead identified gene length as the best predictor of toxicity. Our results present a generalized framework for the cost of aneuploidy with implications for disease biology and evolution.

genic load

Machine learning-led semi-automated medium optimization reveals salt as key for flaviolin production in Pseudomonas putida

Although synthetic biology can produce valuable chemicals in a renewable manner, its progress is still hindered by a lack of predictive capabilities. Media optimization is a critical, and often overlooked, process which is essential to obtain the titers, rates and yields needed for commercial viability. Here, we present a molecule- and host-agnostic active learning process for media optimization that is enabled by a fast and highly repeatable semi-automated pipeline. Its application yielded 60% and 70% increases in titer, and 350% increase in process yield in three different campaigns for flaviolin production in Pseudomonas putida KT2440. Explainable Artificial Intelligence techniques pinpointed that, surprisingly, common salt (NaCl) is the most important component influencing production. The optimal salt concentration is very high, comparable to seawater and close to the limits that P. putida can tolerate. The availability of fast Design-Build-Test-Learn (DBTL) cycles allowed us to show that performance improvements for active learning are rarely monotonous. This work illustrates how machine learning and automation can change the paradigm of current synthetic biology research to make it more effective and informative, and suggests a cost-effective and underexploited strategy to facilitate the high titers, rates and yields essential for commercial viability.

59 BASIC BIOLOGICAL SCIENCES

Data and scripts associated with a manuscript analyzing ELM-FATES parameter sensitivity under pre-fire and postfire scenarios using machine learning

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Fire Severity-Dependent Shifts in Vegetation Parameter Sensitivity: A Pre- and Post-Fire Analysis Using ELM-FATES and Explainable AI” submitted to Journal of Advances in Modeling Earth Systems (Zahura et al. 2026). The study examines vegetation physiological parameters controlling pre-fire and post-fire vegetation dynamics. To support this analysis, 73 vegetation parameters in Functionally Assembled Terrestrial Ecosystem Simulator (FATES) (Fisher et al., 2018) , which is coupled with E3SM (Energy Exascale Earth System Model) land model (ELM, ELM-FATES), were perturbed using a Sobol sequence to generate 1,024 ensemble members for two plant functional types: needleleaf evergreen extratropical trees (NEET) and C3 grass. Simulations were conducted for the pre-fire period (2016) and post-fire period (2018–2023). Burn severity was represented by modifying the Nesterov index in FATES to 75,000, 150,000, and 300,000 for low, moderate, and high severity, respectively. A no-fire scenario was also included. Simulations were performed for 16 grid cells in the American River Watershed across different burn severities and plant functional types. XGBoost (eXtreme Gradient Boosting) models were trained using the parameter ensembles and ELM-FATES-simulated outputs, including leaf area index (LAI), gross primary productivity (GPP), aboveground biomass, vegetation evaporation, transpiration, and soil evaporation. Models were trained separately for each year and burn severity, followed by SHAP (SHapley Additive exPlanations) analysis to identify changes in dominant parameters after fire disturbance. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. The data package contains the ELM-FATES simulation data. The scripts and data related to the analysis will be added later. The inputs and outputs from ELM-FATES are inside the “FATES” folder. “FATES_domain_surface” contains the domain and surface netcdfs that were used to run ELM-FATES in the study area. “FATES_parameters” contains the 1024 ensembles that were generated using Sobol sequence. “FATES_outputs” folder contains ELM-FATES simulated variables. All files are .csv and .nc (NetCDF).

Aboveground biomass

Projected increases in tropical cyclone-induced U.S. electric power outage risk

Abstract While power outages caused by tropical cyclones (TCs) already pose a great threat to coastal communities, how—and why—these risks will change in a warming climate is poorly understood. To address this need, we develop a robust machine learning model to capture TC-induced power outage risk. When applied to 900 000 synthetic TCs downscaled from simulated historical and future climate conditions under a strong warming scenario, we find outage risk in the United States and Puerto Rico is expected to increase broadly by the end of the century, with some states seeing increases of 60% and higher. Further, we discover that rising rainfall rates will play an increasingly important role in TC-induced power outage risk as the climate changes, explaining more than 50% of the projected change in risk in some regions. These insights are important for guiding decision-makers in their future outage risk investment and mitigation plans.

Grid Resilience

FEW questions, many answers: using machine learning to assess how students connect food–energy–water (FEW) concepts

There is growing support and interest in postsecondary interdisciplinary environmental education which integrate concepts and disciplines in addition to providing varied perspectives. There is a need to assess student learning in these programs as well as rigorous evaluation of educational practices, especially of complex synthesis concepts. This work tests a text classification machine learning model as a tool to assess student systems thinking capabilities using two questions anchored by the Food-Energy-Water (FEW) Nexus phenomena by answering two questions (1) Can machine learning models be used to identify instructor-determined important concepts in student responses? (2) What do college students know about the interconnections between food, energy and water, and how have students assimilated systems thinking into their constructed responses about FEW? Reported here are a broad range of model performances across 26 text classification models associated with two different assessment items, with model accuracy ranging from 0.755 to 0.992. Expert-like responses were infrequent in our dataset compared to responses providing simpler, incomplete explanations of the systems presented in the question. For those students moving from describing individual effects to multiple effects, their reasoning about the mechanism behind the system indicates advanced systems thinking ability. Specifically, students exhibit higher expertise for explaining changing water usage than discussing tradeoffs for such changing usage. This research represents one of the first attempts to assess the links between foundational, discipline-specific concepts and systems thinking ability. These text classification approaches to scoring student FEW Nexus Constructed Responses (CR) indicate how these approaches can be used, in addition to several future research priorities for interdisciplinary, practice-based education research. Development of further complex question items using machine learning would allow evaluation of the relationship between foundational concept understanding and integration of those concepts as well as more nuanced understanding of student comprehension of complex interdisciplinary concepts.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION

Machine learning approaches for integrating multi-omics data to expand microbiome annotation (Final Technical Report)

We fulfilled all original three aims of the proposal. Following the earlier release (during the first phase of the project at Montana) of software that identifies and fills gaps in the annotation of metabolic proteins within bacterial genomes, we have nearly completed a second gap-filling tool that improves accuracy and explainability. We completed software for alignment-based annotation of protein coding DNA, allowing for coding frameshifts caused by sequencing error. Finally, we completed a neural embedding model for identifying similarities between protein sequences based on amino-wise latent vectors.

59 BASIC BIOLOGICAL SCIENCES

Machine Learning-Accelerated First-Principles Molecular Dynamics Reveals C–C Coupling Mechanisms toward Ethylene on Cu(100)

Here, the Cu(100) termination has been identified as the most effective facet for converting CO and CO 2 into ethylene. To enhance both the activity and selectivity of ethylene production, we perform machine-learning-accelerated, first-principles molecular dynamics simulations at 298 K in an explicit solvent at pH 7 to elucidate the C–C coupling mechanism─the critical reaction step in forming C 2+ products. Among the six potential C–C coupling pathways, the most feasible are CO* dimerization and CO – CHO* and CHO* – CHO* couplings. Using the computational hydrogen electrode method, we demonstrate that all three pathways are equally accessible at −0.6 V vs RHE. At a potential below −1.0 V vs RHE, the thermodynamic barriers for the CO – CHO* and CHO* – CHO* pathways become negligible. Our computational findings explain the experimental observations, particularly the absence of C 2+ products above −0.4 V vs RHE and the peaks in ethylene production near −0.6 and −1.0 V vs RHE. Since CHO* acts as a key intermediate common to both C–C coupling and CH 4 formation, we propose that suppressing CHO* hydrogenation would inhibit CH 4 pathways, thereby maximizing ethylene selectivity.

CO2 reduction